> ## Documentation Index
> Fetch the complete documentation index at: https://docs.springwinter.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How to Build a Puppeteer Screenshot and PDF API on Springwinter

> Deploy a Puppeteer screenshot and PDF service on Springwinter with an HTTP control API, queue-backed browser workers, S3 output, quotas, and SSRF defenses.

A browser automation API should separate request handling from browser execution. Deploy the control API as a Springwinter Web server and deploy a bounded browser consumer as a Worker.

*Updated October 9, 2026.*

**Return a job ID immediately; never keep an internet-facing request open for an unpredictable browser session.**

## API design

A small asynchronous API can expose:

```text theme={null}
POST /captures
GET  /captures/:id
POST /captures/:id/cancel
GET  /captures/:id/download
```

The create endpoint validates the request, reserves quota, writes a job record, and publishes a queue message. The Worker creates the screenshot or PDF, uploads it to S3, and marks the job complete.

## Prevent server-side request forgery

A user-supplied URL can target private infrastructure. Validate the initial hostname and every redirect. Resolve DNS, reject loopback, link-local, private, multicast, and metadata ranges, then protect against DNS rebinding by controlling outbound networking rather than trusting one lookup.

Block cloud metadata endpoints, private VPC ranges, internal service domains, `file:` URLs, and unsupported protocols. An application-level URL check alone is not a complete egress boundary.

## Bound every resource

Set limits for:

* page and total job duration,
* redirects and subresources,
* response body and download size,
* viewport and PDF dimensions,
* concurrent pages per Worker,
* jobs per tenant and per minute,
* output retention,
* browser restarts and retries.

Use an operation allowlist. Do not expose arbitrary Chrome DevTools Protocol commands to untrusted customers.

## Deploy the components

Create a Web server for the API and a Worker for the browser queue. Store durable job state in a database and outputs in private S3. Use signed download URLs with short expirations.

Monitor queue age, browser crashes, navigation timeouts, memory, output size, and completion rate. A healthy HTTP API does not prove that browser Workers are making progress.

## Security limitation

The official Puppeteer Docker image expects `SYS_ADMIN` for its documented sandbox mode, but Fargate restricts that capability. Disabling Chrome's sandbox is not sufficient isolation for arbitrary hostile pages. Treat a public browsing API as an agent-sandbox problem and choose a runtime that can enforce the required browser sandbox and network policy.

## Frequently asked questions

<AccordionGroup>
  <Accordion title="Can the API generate screenshots synchronously?">
    It can for tightly controlled internal pages, but asynchronous jobs handle slow sites, retries, quotas, browser crashes, and client disconnects more safely.
  </Accordion>

  <Accordion title="Where should generated PDFs and screenshots live?">
    Store them in private S3 and return short-lived signed URLs. The container filesystem is ephemeral and disappears when a task is replaced.
  </Accordion>

  <Accordion title="Is blocking localhost enough to prevent SSRF?">
    No. Protect every redirect and resolved address, block metadata and private ranges, account for DNS rebinding, and enforce network-level egress restrictions.
  </Accordion>
</AccordionGroup>

## Sources and further reading

* [Puppeteer Docker guide](https://pptr.dev/guides/docker)
* [OWASP SSRF Prevention Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/Server_Side_Request_Forgery_Prevention_Cheat_Sheet.html)
* [Deploy a Puppeteer Worker](/blog/deploy-puppeteer-worker-springwinter)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.