Overview
When extracting content with layout information, Pulse API returns bounding box coordinates for text, tables, and images. This spatial data enables precise document understanding and region-based extraction.Bounding Box Format
Bounding boxes are returned as normalized coordinates (0-1 range) in an 8-point format:- (x1, y1) = Top-left corner
- (x2, y2) = Top-right corner
- (x3, y3) = Bottom-right corner
- (x4, y4) = Bottom-left corner
Coordinates are normalized to 0-1 range, making them resolution-independent. To convert to pixels, multiply by the page width/height.
Response Structure
Thebounding_boxes object groups every detected layout element by its public category. Categories with no detected elements may be omitted.
Every detected layout element is routed into a public grouped category; list items, captions, footers, page numbers, formulas, and selection marks are not discarded. Clients do not need to parse provider-specific category names.
Markdown Fields
Use
bounding_boxes.markdown_with_ids when you need to correlate text positions with bounding boxes. Use the top-level markdown for clean content display or export.
Example Response
Here’s a real example of thebounding_boxes object from a workbook with an embedded chart, with figure_processing.show_images: true:
Field Descriptions
All
page_number and location.page values are 1-indexed original document page numbers. This holds even when the request used a pages= subset: extracting pages="10-20" produces items labeled pages 10 through 20 across text items, tables, Words, and extension output.Text Array
Each text element contains:id: Unique identifier (e.g.,txt-1) that links tomarkdown_with_idsviadata-bb-text-idcontent: The extracted text with prefix (e.g.,0a-NCRI)original_content: The clean extracted text without prefixbounding_box: 8-point coordinate array (may benullfor some document types)page_number: Page where the text appearsaverage_word_confidence: OCR confidence score (0-1)selected: Selection state whendetect_selectionswas enabled and the item represents a detected form control or marked-choice region
Title Array
Each title element contains:id: Unique identifier linking to markdowncontent: The title text with prefixoriginal_content: The clean title textbounding_box: 8-point coordinate arraypage_number: Page where the title appearsaverage_word_confidence: OCR confidence score (0-1)
Header Array
Each header element contains:id: Unique identifier linking to markdowncontent: The header text with prefixoriginal_content: The clean header textbounding_box: 8-point coordinate arraypage_number: Page where the header appearsaverage_word_confidence: OCR confidence score (0-1)
Footer Array
Each footer element contains:id: Unique identifier linking to markdowncontent: The footer text with prefixoriginal_content: The clean footer textbounding_box: 8-point coordinate arraypage_number: Page where the footer appearsaverage_word_confidence: OCR confidence score (0-1)
Images Array
Each image element represents a detected chart or embedded image. For PDFs and image inputs, entries are populated when figure detection runs. For spreadsheets, entries are populated for embedded charts and images directly read from the workbook.Fetching Visual Image Bytes
When you setfigure_processing.show_images: true on /extract, every chart/image entry comes back with an image_url pointing at GET /results/{jobId}/images/{filename}. Fetch it with your API key to get the raw PNG/JPEG bytes:
x-api-key authentication; there is no anonymous access.
Tables Array
The extraction engine decides which table regions exist. Pulse then runs CPU-only cell geometry detection inside those exact table crops. It never creates extra tables from regions the extraction engine did not identify. EachTables[] entry contains:
Each
cell_data[] item contains:
Words Array
Words[] is the standard word geometry collection used to preserve the extraction response contract. Each item contains content, page_number, confidence, and a polygon represented as [{"x": number, "y": number}, ...].
Selection Marks Array
When selection detection finds marked controls,SelectionMarks[] contains page_number, state (selected or unselected), confidence, and a normalized polygon.
Page Number Array
Each page number element contains:id: Unique identifiercontent: The page number textoriginal_content: The clean page number textbounding_box: 8-point coordinate arraypage_number: Page where it appearsaverage_word_confidence: OCR confidence score (0-1)
The
id field allows you to link bounding box elements to specific locations in the markdown_with_ids field via data-bb-text-id attributes.Footnote References
When you enableextensions.footnote_references in your extract request, the response includes an extensions.footnoteReferences array that uses bounding box IDs to link footnote markers to their in-text references.
Each entry contains:
symbol— the footnote marker (e.g.*,†,‡,1)footnoteTextId— theidof the footnote explanation, typically found in theFooterarrayreferenceTextIds— an array ofidvalues from theText,Title, orHeaderarrays identifying body paragraphs that contain the marker
footnoteTextId to look up the footnote’s position and content in bounding_boxes.Footer (or bounding_boxes.Text), and each entry in referenceTextIds to locate the citing paragraphs in bounding_boxes.Text, bounding_boxes.Title, or bounding_boxes.Header. This allows you to spatially highlight both the footnote and every place in the document that references it.
Footnote references are only available for PDF documents. See the Extract endpoint for usage examples.
Converting Coordinates
To convert normalized coordinates to pixel coordinates:Next Steps
Extract Endpoint
Enable bounding box extraction
Structured Output
Combine with structured data