API design
A small asynchronous API can expose:Prevent server-side request forgery
A user-supplied URL can target private infrastructure. Validate the initial hostname and every redirect. Resolve DNS, reject loopback, link-local, private, multicast, and metadata ranges, then protect against DNS rebinding by controlling outbound networking rather than trusting one lookup. Block cloud metadata endpoints, private VPC ranges, internal service domains,file: URLs, and unsupported protocols. An application-level URL check alone is not a complete egress boundary.
Bound every resource
Set limits for:- page and total job duration,
- redirects and subresources,
- response body and download size,
- viewport and PDF dimensions,
- concurrent pages per Worker,
- jobs per tenant and per minute,
- output retention,
- browser restarts and retries.
Deploy the components
Create a Web server for the API and a Worker for the browser queue. Store durable job state in a database and outputs in private S3. Use signed download URLs with short expirations. Monitor queue age, browser crashes, navigation timeouts, memory, output size, and completion rate. A healthy HTTP API does not prove that browser Workers are making progress.Security limitation
The official Puppeteer Docker image expectsSYS_ADMIN for its documented sandbox mode, but Fargate restricts that capability. Disabling Chrome’s sandbox is not sufficient isolation for arbitrary hostile pages. Treat a public browsing API as an agent-sandbox problem and choose a runtime that can enforce the required browser sandbox and network policy.
Frequently asked questions
Can the API generate screenshots synchronously?
Can the API generate screenshots synchronously?
It can for tightly controlled internal pages, but asynchronous jobs handle slow sites, retries, quotas, browser crashes, and client disconnects more safely.
Where should generated PDFs and screenshots live?
Where should generated PDFs and screenshots live?
Store them in private S3 and return short-lived signed URLs. The container filesystem is ephemeral and disappears when a task is replaced.
Is blocking localhost enough to prevent SSRF?
Is blocking localhost enough to prevent SSRF?
No. Protect every redirect and resolved address, block metadata and private ranges, account for DNS rebinding, and enforce network-level egress restrictions.