Skip to main content
POST
Extract rich content from a document

Overview

Pipeline Step 1 — Extract is where document processing begins: every downstream step consumes its extraction_id. After extraction, you can optionally split the document into topics, apply schema extraction to get structured data, or use tables for span-aware table extraction.Handling mixed document types? /classify can run before Extract to route each raw document to the right pipeline — it only chooses which pipeline runs, so extraction still happens here.
Extract document content as markdown, layout elements, table cells, words, figures, and optional extension output. When a response reaches the configured inline-size threshold (5 MB by default), the API returns a download URL instead of embedding the payload. Fetching that URL returns the same complete extraction result JSON you would receive inline. See Large Result Response below.
For large documents or batch processing workflows, set async: true to process asynchronously and poll for results via GET /job/jobId.
To process many files at once, use Batch Extract. It accepts an S3 prefix, local directory, or list of URLs and runs /extract on each file in parallel.

Async Mode

Set async: true to return immediately with a job ID for polling:
Async Response (200):
Use GET /job/{job_id} to poll for completion.

Request

Document Source

Provide the document using one of these methods:

Extraction Options

Refinement

refine is off unless you enable it. It accepts three forms:
Valid modes: refine.prompt accepts a string that applies to the whole refinement pass, or an object keyed by mode (e.g. {"tables": "...", "text": "..."}) to scope guidance per mode. Prompt keys must be a subset of the requested modes, and layout never accepts a custom prompt. When you use the boolean or array form, pass the guidance through the sibling refine_prompt field instead; refine.prompt takes precedence over refine_prompt when both are present. Refinement changes the normal markdown and bounding_boxes fields in place; it does not create a parallel response object.

Figure Processing

For spreadsheets, show_images: true returns embedded charts and images with workbook-specific metadata such as sheet_name, excel_range, chart_type, and source_ranges.

Extensions

Extensions add derived outputs while preserving the core response fields. The standard bounding_boxes.Words collection is always part of the extraction response. extensions.alt_outputs.wlbb (equivalently, the top-level word_level_bounding_boxes parameter) requests the additional extensions.altOutputs.wlbb word-box output used by existing WLBB workflows. Organizations can also have word-level output enabled by default for every extraction — contact support to configure an org-wide default.

Spreadsheet Options

Complete Processing Example

Storage Options

Control whether extractions are saved to your extraction library:

Response

The response structure varies based on document size to optimize for different use cases.

Standard Inline Response

When the response payload stays below the inline threshold, results are returned directly in the response body:

Response Fields

Partial Results (failed_pages)

Extraction completes with a partial result whenever at least one page succeeds. A job fails outright only when every page is unextractable. On a partial result:
  • failed_pages at the top level lists each unrecoverable page with a short reason.
  • A matching entry is appended to warnings.
  • The markdown carries a visible <!-- PAGE N FAILED EXTRACTION: ... --> placeholder at each hole, so page-oriented consumers keep their alignment.
  • Only successfully extracted pages are billed.
Treat failed_pages as the authoritative list of holes: if it is absent, every requested page was extracted.
If your organization was migrated from the default engine, responses keep the familiar shape documented on this page — including bounding_boxes with grouped Words, Tables[].cell_data with ids, spans, and confidences, and chunking output — and extraction bills at a flat 1 credit per page. Callers that explicitly request model="pulse-ultra-2" continue to receive today’s Ultra response format unchanged. To keep a request on the previous-generation engine, send model="pulse-ultra-1"; those responses carry average_word_confidence instead of confidence and no reading_order.

Large Result Response

When a response reaches the configured inline-size threshold (5 MB by default), or when force_url: true is supplied, the API returns a download URL instead of inlining the payload. Some rollback deployments may also offload documents over 70 pages. The downloaded JSON is the complete extraction result with the normal response shape.

Large Result Response Fields

Download and persist URL-backed results promptly. Anonymous Pulse result links are single-use and expire 1 hour after completion. Same-org authenticated requests can replay the Pulse URL while retention keeps the artifact available; presigned storage URLs follow their own expiry.

Handling Large Document Responses

Persist URL-backed results in your own storage after download. You can also keep storage.enabled and retrieve saved extractions from the Pulse Platform.

Example Usage

Core Processing

Use refine, refine_prompt, additional_prompt, and detect_selections to steer or refine the primary extraction result. These are separate from extensions, which add derived outputs.
refine also accepts a boolean ("refine": true runs the general pass) or a bare array of modes ("refine": ["tables", "formatting", "text", "layout"]); with those forms, pass refinement guidance through the sibling refine_prompt field. See Core Processing for precedence and defaults.

Basic Extraction

File Upload

Direct multipart uploads (file=) are limited to 100 MB. Larger uploads fail fast with HTTP 413 and error code FILE_TOO_LARGE_USE_URL. Files over 100 MB must be submitted via file_url pointing at your own hosted or presigned URL — file_url submissions are accepted at any size. Very large documents are sharded and processed transparently under a single job ID; see Working with Large Documents.

Structured Data (Extract → Schema)

Apply structured output with /schema after extraction. The resulting extraction_id lets you rerun or change schemas without processing the source document again.
Recommended two-step approach:

Document Metadata

Enable extensions.document_metadata to read native properties from the original file before conversion, rendering, or OCR. The option is a single boolean; Pulse returns every safely recoverable field for the detected format.
Absent metadata fields are omitted rather than returned as null. Metadata is evidence declared by the source file and is not independently verified. Original camera files may contain sensitive capture timestamps or GPS coordinates. See Document Metadata for format-specific behavior and implementation guidance.

Page Range and Chunking

With a pages= subset, every page number in the response — page-break markers, bounding boxes, tables, Words, and extension output — refers to the original document’s page numbers. Requesting pages="10-20" returns items labeled pages 10 through 20.

Footnote References

Enable extensions.footnote_references to detect footnote markers (e.g. *, †, 1) in body text and link them to their footnote text. Each result item carries the marker, the footnote’s text and location, the bounding-box IDs of the body blocks that cite it, and one entry per located marker occurrence with the marker’s own bounding box where available — enough to highlight both the footnote and every citation on the page.

Example Response

Footnote Reference Fields

Each references[] entry:
Footnote reference detection combines layout analysis with native PDF text extraction for accurate symbol identification. Supported markers include numbered (1, 2, 3), superscript and parenthesized forms (⁴, (1)), symbolic (*, †, ‡, §, ), and lettered (a, b, c) footnotes. Available for PDFs and images; text-based PDFs give the most precise marker positions.

Excel Spreadsheet Options

Spreadsheet table cells are returned under bounding_boxes.Tables[].cell_data. The default spreadsheet.cell_data_mode: "inline" returns each full cell array. Set it to "external" to return per-table JSONL artifact references when very large workbooks make inline cell arrays impractical.
Workbooks exported from claims systems, ERPs, and other automated pipelines often declare a “used range” that extends hundreds of thousands of rows past where the data actually ends. Set spreadsheet.only_data_rows: true and spreadsheet.only_data_cols: true to have Pulse trim those trailing empty “phantom” rows and columns before parsing. Surviving cells keep their original A1 coordinates, so any citation or bounding box that references a specific cell remains stable. Both flags default to false. See the extraction options above for the full reference.

Excel Charts and Embedded Images

When you set figure_processing.show_images: true on an Excel workbook, every embedded chart and image is collected from the workbook directly and returned under bounding_boxes.Images[]. Each entry carries a Pulse-hosted image_url you can fetch via results.getImage (or any HTTP client with your API key) to get the raw PNG/JPEG bytes.

Example bounding_boxes.Images Entry

See Bounding Boxes — Images Array for the full field reference and Get Result Image for the auth requirement on image_url.

Disable Storage

Authorizations

x-api-key
string
header
required

Body

multipart/form-data

Input schema for extraction requests. Provide either file (direct upload) or fileUrl (remote URL).

file
file

Document to upload directly. Required unless fileUrl is provided.

fileUrl
string<uri>

Public or pre-signed URL that Pulse will download and extract. Required unless file is provided.

model
enum<string>

Extraction model to use. pulse-ultra-2 runs Pulse Ultra 2; pulse-ultra-1 runs the previous-generation engine. If omitted or set to default, your organization's default backend is used (Pulse Ultra 2 for organizations upgraded to it), so send pulse-ultra-1 explicitly to keep a request on the previous engine.

Available options:
default,
pulse-ultra-1,
pulse-ultra-2
detectSelections
boolean

Pulse Ultra 2 only. Enables a specialized selection-mark detection pass that improves selected/unselected state accuracy for forms, checkboxes, radio buttons, handwritten checkmarks, X marks, and similar controls. Enabled by default when model is pulse-ultra-2; set to false to skip this pass. Passing true without model: pulse-ultra-2 returns a validation error.

extractionConfigId
string<uuid>

UUID of a saved extraction configuration (a "preset"). When provided, the server loads the saved configuration and applies its options on top of any inline parameters supplied in this request. Inline parameters always take precedence over preset values for the same field. Saved configs are managed via the platform UI or the input_extractions admin endpoints.

pages
string

Page range filter supporting segments such as 1-2 or mixed ranges like 1-2,5.

Pattern: ^[0-9]+(-[0-9]+)?(,[0-9]+(-[0-9]+)?)*$
forceUrl
boolean
default:false

When true, return the complete extraction result as a URL even if it is small. Spreadsheet extractions use URL delivery by default; set force_url: false to request inline spreadsheet output. URL delivery changes only the transport, not the result shape.

figureProcessing
object

Settings that control how figures and embedded visuals are processed. Applies to both PDFs/images (where figures are detected from layout) and spreadsheets (where charts and embedded images are read directly from the workbook). These options affect the markdown output and the bounding_boxes.Images[] array; they do not produce additional output fields elsewhere in the response.

extensions
object

Settings that enable additional processing passes or alternate output formats. Each enabled extension produces a corresponding output field under response.extensions.*.

spreadsheet
object

Settings for Excel/spreadsheet extraction. Controls handling of hidden rows, columns, and sheets, whether numeric cells are rendered using their display format or underlying raw value, where table cell metadata is returned (inline or as external artifacts), whether cell formatting is included, and optional trimming of empty phantom rows/columns past the last data-bearing cell. Applies to .xlsx, .xlsm, and .xls files. Accepts both camelCase and snake_case field names; spreadsheet_options is accepted as a legacy alias for this object.

storage
object

Options for persisting extraction artifacts. When enabled (default), artifacts are saved to storage and a database record is created.

async
boolean
default:false

If true, returns immediately with a job_id for polling via GET /job/{jobId}. Otherwise processes synchronously.

structuredOutput
object
deprecated

⚠️ DEPRECATED — Use the /schema endpoint after extraction instead. Pass the extraction_id from the extract response to /schema with your schema_config. This parameter still works for backward compatibility but will be removed in a future version.

schema
deprecated

(Deprecated) JSON schema describing structured data to extract. Use structuredOutput instead. Accepts either a JSON object or a stringified JSON representation.

schemaPrompt
string
deprecated

(Deprecated) Natural language prompt for schema-guided extraction. Use structuredOutput.schemaPrompt instead.

customPrompt
string
deprecated

(Deprecated) Custom instructions that augment the default extraction behaviour. Use figureProcessing or extensions instead.

chunking
string
deprecated

⚠️ DEPRECATED — Use extensions.chunking.chunkTypes instead. Comma-separated list of chunking strategies to apply (for example semantic,header,page,recursive). Still accepted for backward compatibility.

chunkSize
integer
deprecated

⚠️ DEPRECATED — Use extensions.chunking.chunkSize instead. Override for maximum characters per chunk when chunking is enabled.

Required range: x >= 1
extractFigure
boolean
default:false
deprecated

⚠️ DEPRECATED — Toggle to enable figure extraction in results.

figureDescription
boolean
default:false
deprecated

⚠️ DEPRECATED — Use figureProcessing.description instead. Toggle to generate descriptive captions for extracted figures.

showImages
boolean
default:false
deprecated

⚠️ DEPRECATED — Use figureProcessing.showImages instead. Embed base64-encoded images inline in figure tags in the output. Increases response size.

returnHtml
boolean
default:false
deprecated

⚠️ DEPRECATED — Use extensions.altOutputs.returnHtml instead. Whether to include HTML representation alongside markdown in the response.

thinking
boolean
default:false
deprecated

(Deprecated) Enables expanded rationale output for debugging.

Response

Extraction result. For documents under 70 pages the full result is returned inline. For larger documents and spreadsheet extractions the response can contain is_url: true and a single-use url to download the full result via GET /results/{jobId}.

Full extraction result returned by the synchronous /extract endpoint. Inherits all core fields and adds deprecated backward-compatibility fields.

markdown
string

Primary markdown content extracted from the document. Always present in the new format.

extensions
object

Output from enabled extensions. Each key corresponds to an extension that was enabled in the request under extensions.*. Only keys for enabled extensions are present.

bounding_boxes
object

Positional bounding-box data for text, titles, headers, footers, images, and tables. Images carries chart/image visuals (with image_url when figure_processing.show_images is enabled), Tables the detected tables, and Text/Title/Footer the paragraph/title/footer regions. Additional keys (e.g. markdown_with_ids, defined_names) round-trip without being typed.

extraction_id
string<uuid>

Persisted extraction ID. Present when storage is enabled (default). Use this ID with /split and /schema endpoints.

extraction_url
string

URL to view the extraction on the Pulse platform. Present when storage is enabled.

page_count
integer

Number of pages processed.

Required range: x >= 1
plan_info
object

Billing tier and cumulative usage information. Includes total_credits_used (primary billing metric) and pages_used (legacy compatibility).

warnings
string[]

Non-fatal warnings generated during extraction. Includes deprecation notices when legacy input parameters are used, as well as processing warnings (e.g. word-level bounding box limitations).

credits_used
number<float> | null

Number of credits consumed by this request. Only present when the organization has the credit billing system enabled.

html
string
deprecated

Deprecated — Use extensions.altOutputs.html instead. HTML representation of the extracted content. Present when the legacy returnHtml input was used.

chunks
object
deprecated

Deprecated — Use extensions.chunking instead. Document content split into chunks. Present when the legacy chunking input was used.

structured_output
object
deprecated

Deprecated — Only present when the deprecated structuredOutput input parameter was used. Use the /schema endpoint after extraction instead.

input_schema
object
deprecated

Deprecated — Echo of the schema that was applied. Only present when the deprecated structuredOutput input parameter was used.

schema_error
string
deprecated

Deprecated — Error message if schema processing failed via the deprecated structuredOutput input parameter.

content
string
deprecated

Deprecated — Alias for markdown. Included for backward compatibility with older SDK versions. Prefer markdown.

job_id
string
deprecated

Deprecated — Identifier assigned to the extraction job. Retained for backward compatibility.

metadata
object
deprecated

Deprecated — Additional metadata supplied by the backend. Retained for backward compatibility.