For teams doing this to a whole site
Work through a site without writing a job runner
Point a durable crawl at a seed URL, or hand a batch a list, and collect every page's content and a picture of it.
The problem
One page is a request. Nine hundred pages is infrastructure: a queue, retries that do not double-charge, somewhere to put partial results, a way to answer "is it done yet", and a plan for the run that dies at page 600 because a worker timed out. Most of that is not the problem you set out to solve.
How domscout handles it
Point a crawl at a seed URL with the origins and path patterns it may follow, or hand a batch an explicit list. The job is durable: it survives a worker dying, each child reports its own result, quota is charged when a child starts rather than when the parent is created, and an idempotency key means a retried submission is the same run and not a second one.
Where it shows up
- First-time ingestion of a documentation site or knowledge base.
- Periodic refresh of a catalogue you already index.
- Bulk migration work where you need every page's content and a picture of it.
The request
POST /crawl. Crawl and batch start at Business. A crawl stops at 500 pages and depth 5, needs an explicit HTTPS origin allowlist, and always obeys robots.txt.
Go further
Other use cases: Web access for AI agents · Website change monitoring · Web archiving and evidence · SEO regression checks · Thumbnails and link previews · PDFs from your own HTML · Structured data from web pages · Link preview checks · Responsive layout checks
Give your product a browser.
Get clean web content and visual proof into your workflow in minutes.