The Golden Path
ID Handoffs
Extract -> Schema
Extract -> Tables
Extract -> Split -> Schema
Extract -> Charts
Async Chaining
When you setasync: true, wait for completion before passing the result to the next step.
Common Mistakes
Passing job_id where extraction_id is required
Passing job_id where extraction_id is required
A
job_id is for polling. After the job completes, read the completed result and pass its extraction_id into /schema, /split, /tables, or /charts.Expecting an extraction_id from /classify
Expecting an extraction_id from /classify
Classify is a routing step, not an extraction: it returns a
classification and pipeline_id, and persists nothing. Run the routed pipeline with the same file to produce the extraction_id that downstream steps need.Disabling storage before downstream steps
Disabling storage before downstream steps
Downstream steps need saved extraction artifacts. Keep storage enabled when you plan to chain.
Using Split when a page range is enough
Using Split when a page range is enough
If the target pages are always known, pass
pages, a table page_range, or a chart page_range. Use Split when topic location changes by document.Using Schema for a table-first workflow
Using Schema for a table-first workflow
Schema is great for named fields. Use Tables when preserving row and column relationships is the product.
Related
Pipeline Overview
See supported API pipeline shapes.
Moving from Platform to Production
Generate chained SDK calls from the Playground.