Extract Structured Data
Crawl a website and return schema-shaped data you can drop directly into products, agents, and workflows.
Authorizations
Bearer authentication header of the form Bearer <API_KEY>, where <API_KEY> is your api key.
Body
The starting website URL to crawl and extract from. Must include http:// or https://.
JSON Schema for the returned data object. Image fields such as image_urls or product_photos automatically make page image references available to extraction, so product data and photos can be returned in one call. TypeScript Zod users can pass a JSON Schema generated from a Zod object; Python users can pass the equivalent JSON Schema object.
Optional extraction guidance, such as which facts to prioritize or how to interpret fields in the schema.
2000When true, every returned value must be grounded in facts stated on the page; fields that cannot be supported by the page are returned as null/empty. When false (default), the model may make reasonable inferences and derivations from the page content (e.g. ideal customer, competitor analysis, recommendations) while keeping verifiable specifics (names, quotes, URLs, dates, metrics) faithful to the source.
When true, follow links on subdomains of the starting URL's domain.
Maximum number of pages to analyze for extraction. Hard cap: 50. Defaults to 5.
1 <= x <= 50Optional maximum link depth from the starting URL (0 = only the starting page). If omitted, there is no crawl depth limit.
x >= 0When true, iframe contents are included in Markdown before extraction.
Return cached scrape results if a prior scrape for the same parameters is younger than this many milliseconds. Defaults to 7 days (604800000 ms).
0 <= x <= 2592000000Optional browser wait time in milliseconds after initial page load for each crawled page.
0 <= x <= 30000When true, waits briefly for CSS and transition animations to settle before extracting each crawled page. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages.
Optional browser actions executed in order on the requested page after it loads, before links are discovered or additional pages are crawled. Requires a paid plan. When actions are provided and stopAfterMs is omitted, the crawl budget defaults to 110000 ms.
5Browser action discriminated by do. Each variant exposes only its applicable fields.
- Wait
- Perform
- Scroll
Soft time budget for the crawl in milliseconds. Min: 10000 (10s). Max: 110000 (110s). Defaults to 80000 (80s), or 110000 (110s) when browser actions are provided.
10000 <= x <= 110000Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).
1000 <= x <= 300000Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.
201 - 50Response
Successful response
Status of the response, e.g., 'ok'
The starting URL that was analyzed
List of URLs whose Markdown was used for extraction
Extracted data matching the request schema
Cache outcome for this response. Composite responses are hits only when every cache-controlled fetch contributing to the output was a hit; age_ms is the oldest contributing hit.
Metadata about the API key used for the request. Included in every response whenever a valid API key is provided, even when the response status is not 200.