Skip to main content
Crawl follows links from a start URL and returns Markdown for up to 500 pages. Bound the work with page, depth, path, and time limits.

Crawl a site

Read each item in results for its URL, Markdown, and metadata. The reference lists the full response and counters. Check partial when a deadline may have ended the crawl early.

Choose the workflow

Scope and failures

See scope and limits to restrict discovery and page content to control extraction. A missing start page returns 404 NOT_FOUND; unsupported start-page content returns 415 UNSUPPORTED_CONTENT. Crawl uses a weight of 10 in the API key’s rate-limit bucket. For a job that should continue beyond a single request, use a batch.