> This page is for version v2025-04-21.
> For other versions, use one of these documentation indexes:
> - v2026-02-09 (default): https://docs.extend.ai/2026-02-09/llms.txt
> - v2025-04-21: https://docs.extend.ai/2025-04-21/llms.txt
> - v2024-12-23: https://docs.extend.ai/2024-12-23/llms.txt

> ## Documentation Index
> Fetch the complete documentation index at: https://docs.extend.ai/llms.txt
> Use this file to discover all available pages before exploring further.
>
> ## API version
> The current API version is `2026-02-09`, served at the site root (no version prefix in URLs).
> If this page URL contains `/2025-04-21/` or `/2024-12-23/`, you are reading an older API version.
> Prefer the current docs at https://docs.extend.ai/llms.txt unless the user explicitly needs that older version.
> Do not treat older-version pages as the source of truth for new integrations.

# Run Processor

POST https://api.extend.ai/processor_runs
Content-Type: application/json

Run processors (extraction, classification, splitting, etc.) on a given document.

**Synchronous vs Asynchronous Processing:**
- **Asynchronous (default)**: Returns immediately with `PROCESSING` status. Use webhooks or polling to get results.
- **Synchronous**: Set `sync: true` to wait for completion and get final results in the response (5-minute timeout).

**For asynchronous processing:**
- You can [configure webhooks](https://docs.extend.ai/2025-04-21/product/webhooks/configuration) to receive notifications when a processor run is complete or failed.
- Or you can [poll the get endpoint](https://docs.extend.ai/2025-04-21/developers/api-reference/processor-endpoints/get-processor-run) for updates on the status of the processor run.


Reference: https://docs.extend.ai/api-reference/endpoints/legacy/create-processor-run

## Authentication

- `Authorization` header (bearer token, required) — Bearer authentication of the form `Bearer <token>`, where token is your auth token.

## Request

### Headers

- `x-extend-api-version` ("2025-04-21", optional, default: 2025-04-21) — API version to use for the request. If you do not specify a version, you will either receive a `400 Bad Request` or be set to a previous legacy version. See [API Versioning](https://docs.extend.ai/2025-04-21/developers/api-versioning) for more details.

### Body (application/json)

This endpoint expects an object.

- `processorId` (string, required) — The ID of the processor to be run. Example: `"dp_Xj8mK2pL9nR4vT7qY5wZ"`
- `version` (string, optional, default: latest) — An optional version of the processor to use. When not supplied, the most recent published version of the processor will be used. Special values include: - `"latest"` for the most recent published version. If there are no published versions, the draft version will be used. - `"draft"` for the draft version. - Specific version numbers corresponding to versions your team has published, e.g. `"1.0"`, `"2.2"`, etc.
- `file` (ProcessorRunFileInput, optional) — The file to be processed. One of `file` or `rawText` must be provided. Supported file types can be found [here](https://docs.extend.ai/2025-04-21/product/general/supported-file-types).
- `rawText` (string, optional) — A raw string to be processed. Can be used in place of file when passing raw text data streams. One of `file` or `rawText` must be provided.
- `sync` (boolean, optional, default: false) — Whether to run the processor synchronously. When `true`, the request will wait for the processor run to complete and return the final results. When `false` (default), the request returns immediately with a `PROCESSING` status, and you can poll for completion or use webhooks. For production use cases, we recommending leaving sync off and building around an async integration for more resiliency, unless your use case is predictably fast (e.g. sub \< 30 seconds) run time or otherwise have integration constraints that require a synchronous API. **Timeout**: Synchronous requests have a 5-minute timeout. If the processor run takes longer, it will continue processing asynchronously and you can retrieve the results via the GET endpoint.
- `priority` (integer, optional, default: 50) — An optional value used to determine the relative order of ProcessorRuns when rate limiting is in effect. Lower values will be prioritized before higher values.
- `metadata` (map from string to any, optional) — An optional object that can be passed in to identify the run of the document processor. It will be returned back to you in the response and webhooks. To categorize processor runs for billing and usage tracking, include `extend:usage_tags` with an array of string values (e.g., `{"extend:usage_tags": ["production", "team-eng", "customer-123"]}`). Tags must contain only alphanumeric characters, hyphens, and underscores; any special characters will be automatically removed.
- `config` (ProcessorRunsPostRequestBodyContentApplicationJsonSchemaConfig, optional) — The configuration for the processor run. If this is provided, this config will be used. If not provided, the config for the specific version you provide will be used. The type of configuration must match the processor type.

## Response

### 200

Successfully created processor run

- `success` (boolean, required)
- `processorRun` (ProcessorRun, required)

## Errors

### 400 Bad Request Error

Bad Request

- `success` (boolean, optional)
- `error` (string, optional) — Error message

### 401 Unauthorized Error

Unauthorized

- `success` (boolean, optional)
- `error` (string, optional) — Error message

### 404 Not Found Error

Not Found

- `success` (boolean, optional)
- `error` (string, optional) — Error message

### 429 Too Many Requests Error

Too Many Requests - Processing runs limit exceeded

- `error` (string, required)

## Types

### ProcessorRunFileInput

Input file for running a single processor.

- `fileName` (string, optional) — The name of the file to be processed. If not provided, the file name will be inferred from the URL. It is highly recommended to include this parameter for legibility.
- `fileUrl` (string, optional) — A URL where the file can be downloaded from. If you use presigned URLs, we recommend an expiration time of 5-15 minutes. One of a `fileUrl` or `fileId` must be provided.
- `fileId` (string, optional) — Extend's internal ID for the file. It will always start with `file_`. One of a `fileUrl` or `fileId` must be provided. You can view a file ID from the Extend UI, for instance from running a parser or from a previous file creation. If you provide a `fileId`, any parsed data will be reused.

### ProcessorRunsPostRequestBodyContentApplicationJsonSchemaConfig

The configuration for the processor run. If this is provided, this config will be used. If not provided, the config for the specific version you provide will be used. The type of configuration must match the processor type.

### ProcessorRun

- `object` (string, required) — The type of response. In this case, it will always be `"document_processor_run"`.
- `id` (string, required) — The unique identifier for this processor run. Example: `"dpr_Xj8mK2pL9nR4vT7qY5wZ"`
- `processorId` (string, required) — The ID of the processor used for this run. Example: `"dp_Xj8mK2pL9nR4vT7qY5wZ"`
- `processorVersionId` (string, required) — The ID of the specific processor version used.
- `processorName` (string, required) — The name of the processor. Example: `"Invoice Processor"`
- `status` (enum, required) — The status of a processor run: * `"PENDING"` - The processor run has not started yet * `"PROCESSING"` - The processor run is in progress * `"PROCESSED"` - The processor run completed successfully * `"FAILED"` - The processor run failed * `"CANCELLED"` - The processor run was cancelled
  - Allowed values: `PENDING`, `PROCESSING`, `PROCESSED`, `FAILED`, `CANCELLED`
- `output` (ProcessorOutput, required) — The final output, either reviewed or initial. Conforms to the shape of output types and depends on the processor type and configuration shape.
- `reviewed` (boolean, required) — Indicates whether the run has been reviewed.
- `edited` (boolean, required) — Indicates whether the run results have been edited.
- `edits` (map from string to ExtractionOutputEdits, required)
- `type` (enum, required) — The type of processor: * `"CLASSIFY"` - Classifies documents into categories * `"EXTRACT"` - Extracts structured data from documents * `"SPLITTER"` - Splits documents into multiple parts
  - Allowed values: `CLASSIFY`, `EXTRACT`, `SPLITTER`
- `config` (ProcessorRunConfig, required) — The configuration used for this processor run. The type of configuration will match the processor type.
- `files` (list of File, required) — Details of the processed files. If this was a file generated from a splitter processor, this will be the sub file. See the File object for more details.
- `mergedProcessors` (list of ProcessorRunMergedProcessorsItems, required) — An array of processors that were merged to create this output. Will be an empty array unless this output was the result of a MergeExtraction step in a workflow.
- `url` (string, required) — The URL to view the processor run.
- `failureReason` (string, optional) — If the run failed, indicates the reason for failure.
- `failureMessage` (string, optional) — If the run failed, provides a detailed message about the failure.
- `metadata` (map from string to any, optional) — Any metadata that was provided when creating the processor run.
- `initialOutput` (ProcessorOutput, optional) — The initial output from the processor run. The type of output will match the processor type.
- `reviewedOutput` (ProcessorOutput, optional) — The output after review, if any.
- `usage` (DocumentProcessorRunCredits, optional) — These are usage credits for document processing (extraction, classification, or splitting). File parsing credits are tracked separately and can be retrieved from the [Get File](https://docs.extend.ai/2025-04-21/developers/api-reference/file-endpoints/get-file) endpoint. This field will not be returned for processor runs created before October 7, 2025, or for customers on legacy billing systems. For more details on how credits work, see our [Credits Guide](https://docs.extend.ai/general/how-credits-work).

### ClassificationConfig

- `type` (enum, required) — The type of configuration. Must be `"CLASSIFY"` for classification processors.
  - Allowed values: `CLASSIFY`
- `classifications` (list of Classification, required) — Array of possible classifications for the document.
- `baseProcessor` (enum, optional, default: classification_performance) — The base processor to use. For classifiers, this must be either `"classification_performance"` or `"classification_light"`. See [Classification Changelog](https://docs.extend.ai/model-versioning/classification/classification-performance) for more details.
  - Allowed values: `classification_performance`, `classification_light`
- `baseVersion` (string, optional) — The version of the `"classification_performance"` or `"classification_light"` processor to use. If this is provided, the `baseProcessor` must also be provided. See [Classification Changelog](https://docs.extend.ai/model-versioning/classification/classification-performance) for more details.
- `classificationRules` (string, optional) — Custom rules to guide the classification process in natural language.
- `advancedOptions` (ClassificationAdvancedOptions, optional) — Advanced configuration options.
- `parser` (ParseConfig, optional) — Configuration options for the parsing process.

### ExtractionConfig

- `type` (enum, required) — The type of configuration. Must be "EXTRACT" for extraction processors.
  - Allowed values: `EXTRACT`
- `baseProcessor` (enum, optional, default: extraction_performance) — The base processor to use. For extractors, this must be `"extraction_performance"`, `"extraction_light"`, or `"extraction_operator"`. See [Extraction Changelog](https://docs.extend.ai/model-versioning/extraction/extraction-performance) for more details. `"extraction_operator"` selects [Operator](https://docs.extend.ai/2025-04-21/product/extraction/operator).
  - Allowed values: `extraction_performance`, `extraction_light`, `extraction_operator`
- `baseVersion` (string, optional) — The version of the `"extraction_performance"`, `"extraction_light"`, or `"extraction_operator"` processor to use. If this is provided, the `baseProcessor` must also be provided. See [Extraction Changelog](https://docs.extend.ai/model-versioning/extraction/extraction-performance) for more details.
- `extractionRules` (string, optional) — Custom rules to guide the extraction process in natural language.
- `schema` (map from string to any, optional) — JSON Schema definition of the data to extract. Either `fields` or `schema` must be provided. See the [JSON Schema guide](https://docs.extend.ai/2025-04-21/product/extraction/schema) for details and examples of schema configuration.
- `advancedOptions` (ExtractionAdvancedOptions, optional) — Advanced configuration options.
- `parser` (ParseConfig, optional) — Configuration options for the parsing process.
- `fields` (list of ExtractionField, optional, deprecated) — Array of fields to extract from the document. Either `fields` or `schema` must be provided. We recommend using `schema` for new implementations.

### SplitterConfig

- `type` (enum, required) — The type of configuration. Must be "SPLITTER" for splitter processors.
  - Allowed values: `SPLITTER`
- `splitClassifications` (list of Classification, required) — Array of classifications that define the possible types of document sections.
- `baseProcessor` (enum, optional, default: splitting_performance) — The base processor to use. For splitters, this can currently only be `"splitting_performance"` or `"splitting_light"`. See [Splitting Changelog](https://docs.extend.ai/model-versioning/splitting/splitting-performance) for more details.
  - Allowed values: `splitting_performance`, `splitting_light`
- `baseVersion` (string, optional) — The version of the `"splitting_performance"` or `"splitting_light"` processor to use. If this is provided, the `baseProcessor` must also be provided. See [Splitting Changelog](https://docs.extend.ai/model-versioning/splitting/splitting-performance) for more details.
- `splitRules` (string, optional) — Custom rules to guide the document splitting process in natural language.
- `advancedOptions` (SplitterAdvancedOptions, optional) — Advanced configuration options.
- `parser` (ParseConfig, optional) — Configuration options for the parsing process.

### ProcessorOutput

The output from a processor run. The type of output will match the processor type.

### ExtractionOutputEdits

A record of edits made to the processor output.

- `originalValue` (any, optional) — The original value before editing.
- `editedValue` (any, optional) — The value after editing.
- `notes` (string, optional) — Any notes added during editing.
- `page` (double, optional) — The page number where the edit was made.
- `fieldType` (string, optional) — The type of the edited field.

### ProcessorRunConfig

The configuration used for this processor run. The type of configuration will match the processor type.

### File

- `object` (string, required) — The type of response. In this case, it will always be "file".
- `id` (string, required) — Extend's internal ID for the file. It will always start with `"file_"`. Example: `"file_xK9mLPqRtN3vS8wF5hB2cQ"`
- `name` (string, required) — The name of the file Example: `"Invoices.pdf"`
- `metadata` (FileMetadata, required)
- `createdAt` (string, required) — The time (in UTC) at which the file was created. Will follow the RFC 3339 format. Example: `"2024-03-21T15:30:00Z"`
- `updatedAt` (string, required) — The time (in UTC) at which the file was last updated. Will follow the RFC 3339 format. Example: `"2024-03-21T16:45:00Z"`
- `type` (enum, optional) — The type of the file
  - Allowed values: `PDF`, `CSV`, `IMG`, `TXT`, `DOCX`, `EXCEL`, `XML`, `HTML`
- `presignedUrl` (string, optional) — A presigned URL to download the file. Expires after 15 minutes.
- `parentFileId` (string, optional) — The ID of the parent file. Only included if this file is a derivative of another file, for instance if it was created via a Splitter in a workflow.
- `contents` (FileContents, optional)
- `usage` (FileCredits, optional) — These are usage credits for parsing the file. This field will not be returned for files processed before October 7, 2025, or for customers on legacy billing systems. For more details on how credits work, see our [Credits Guide](https://docs.extend.ai/general/how-credits-work).

### ProcessorRunMergedProcessorsItems

- `processorId` (string, required) — The ID of the merged processor. Example: `"dp_Xj8mK2pL9nR4vT7qY5wZ"`
- `processorVersionId` (string, required) — The ID of the specific processor version used.
- `processorName` (string, required) — The name of the merged processor. Example: `"Invoice Line Items Processor"`

### DocumentProcessorRunCredits

These are usage credits for document processing (extraction, classification, or splitting). File parsing credits are tracked separately and can be retrieved from the [Get File](https://docs.extend.ai/2025-04-21/developers/api-reference/file-endpoints/get-file) endpoint. This field will not be returned for processor runs created before October 7, 2025, or for customers on legacy billing systems. For more details on how credits work, see our [Credits Guide](https://docs.extend.ai/general/how-credits-work).

- `credits` (double, required) — The number of credits consumed specifically for document processing in this processor run. This does not include file parsing credits.

### Classification

- `id` (string, required) — Unique identifier for the classification. We recommend lowercase, underscore-separated format.
- `type` (string, required) — Type identifier for the classification.
- `description` (string, required) — A detailed description of the classification.

### ClassificationAdvancedOptions

- `context` (enum, optional, default: default) — The context to use for classification.
  - Allowed values: `default`, `max`
- `advancedMultimodalEnabled` (boolean, optional, default: false) — Enable advanced multimodal processing for better handling of visual elements during classification.
- `pageRanges` (list of PageRangesItems, optional) — Limit processing to the specified page ranges. See [Page Ranges](https://docs.extend.ai/2025-04-21/product/page-ranges).
- `fixedPageLimit` (integer, optional, deprecated) — Limit processing to a specific number of pages from the beginning of the document. See [Page Ranges](https://docs.extend.ai/2025-04-21/product/page-ranges).

### ParseConfig

Configuration options for the parsing process.

- `target` (enum, optional, default: markdown) — The target format for the parsed content. Supported values: * `markdown`: True markdown with logical reading order (headings, lists, tables, checkboxes). Best default for LLMs/RAG and enables section-based chunking. * `spatial`: Layout/position-preserving text that uses markdown elements for block types but is not strictly markdown due to whitespace/tabs used to maintain placement. Only page-based chunking is supported. Guidance: * Prefer `markdown` for most documents, multi-column reading order, and retrieval use cases * Prefer `spatial` for messy/scanned/handwritten or skewed documents, when you need near 1:1 layout fidelity, or for BOL-like logistics docs See “Markdown vs Spatial” in the Parse guide for details: /2025-04-21/developers/guides/parse#markdown-vs-spatial
  - Allowed values: `markdown`, `spatial`
- `chunkingStrategy` (ParseConfigChunkingStrategy, optional) — Strategy for dividing the document into chunks.
- `engine` (enum, optional, default: parse_performance) — The parsing engine to use. Supported values: * `parse_performance`: Full-featured parsing engine with highest accuracy (default) * `parse_light`: Lightweight parsing engine optimized for high-volume, cost-sensitive ingestion. Uses the new layout model with full layout support and the markdown target at lower cost and latency, but performs worse than `parse_performance` on lower-quality scans, harder handwriting, larger tables, non-Latin-based languages, and dense checkbox regions.
  - Allowed values: `parse_performance`, `parse_light`
- `blockOptions` (ParseConfigBlockOptions, optional) — Options for controlling how different block types are processed.
- `advancedOptions` (ParseConfigAdvancedOptions, optional)

### ExtractionAdvancedOptions

- `modelReasoningInsightsEnabled` (boolean, optional) — Whether to enable model reasoning insights.
- `advancedMultimodalEnabled` (boolean, optional) — Whether to enable advanced multimodal features.
- `citationsEnabled` (boolean, optional) — Whether to enable citations in the output.
- `arrayCitationStrategy` (enum, optional) — Granularity for array citations. This requires citationsEnabled=true and a base processor version that supports property-level array citations (extraction_performance ≥ 4.4.0).
  - Allowed values: `item`, `property`
- `reviewAgent` (ExtractionAdvancedOptionsReviewAgent, optional) — Configuration for the review agent that analyzes extraction results. When enabled, each field in the output metadata will include a `reviewAgentScore` (1-5) and may include additional `insights` of type `issue` or `review_summary` to help identify fields that may need manual review. Enabling the review agent incurs additional credits. To learn more, view the [Review Agent Documentation](https://docs.extend.ai/product/extraction/review-agent)
- `arrayStrategy` (ArrayStrategy, optional) — Strategy for handling large arrays in documents.
- `chunkingOptions` (ExtractChunkingOptions, optional)
- `excelSheetRanges` (list of ExcelSheetRange, optional) — Ranges of sheet indices to extract from Excel documents.
- `excelSheetSelectionStrategy` (enum, optional) — Strategy for selecting sheets from Excel documents.
  - Allowed values: `intelligent`, `all`, `first`, `last`
- `pageRanges` (list of PageRangesItems, optional) — Limit processing to the specified page ranges. See [Page Ranges](https://docs.extend.ai/2025-04-21/product/page-ranges).
- `documentKind` (string, optional, deprecated) — DEPRECATED - use extractionRules for all system prompts.
- `keyDefinitions` (string, optional, deprecated) — DEPRECATED - use extractionRules for all system prompts.
- `advancedFigureParsingEnabled` (boolean, optional, deprecated) — Whether to enable advanced figure parsing.
- `fixedPageLimit` (integer, optional, deprecated) — DEPRECATED - See [Page Ranges](https://docs.extend.ai/2025-04-21/product/page-ranges).

### ExtractionField

- `id` (string, required) — Unique identifier for the field.
- `name` (string, required) — Human-readable name for the field.
- `type` (enum, required) — The type of the field.
  - Allowed values: `string`, `number`, `currency`, `boolean`, `date`, `array`, `enum`, `object`, `signature`
- `description` (string, required) — Detailed description of the field, including expected content and format.
- `schema` (list of ExtractionField, optional) — Required when type is "array" or "object". Contains nested field definitions.
- `enum` (list of Enum, optional) — Required when type is "enum". List of allowed values.

### SplitterAdvancedOptions

- `splitIdentifierRules` (string, optional) — Custom rules for identifying split points.
- `splitMethod` (enum, optional, default: high_precision) — The method to use for splitting documents. `high_precision` is more accurate but slower, while `low_latency` is faster but less precise.
  - Allowed values: `high_precision`, `low_latency`
- `splitExcelDocumentsBySheetEnabled` (boolean, optional, default: false) — For Excel documents, split by worksheet.
- `pageRanges` (list of PageRangesItems, optional) — Limit processing to the specified page ranges. See [Page Ranges](https://docs.extend.ai/2025-04-21/product/page-ranges).
- `fixedPageLimit` (integer, optional, deprecated) — Limit processing to a specific number of pages from the beginning of the document. See [Page Ranges](https://docs.extend.ai/2025-04-21/product/page-ranges).

### ClassifierOutput

- `id` (string, required) — The unique identifier for this classification
- `type` (string, required) — The type of classification
- `confidence` (double, required) — A value between 0 and 1 indicating the model's confidence in the classification, where 1 represents maximum confidence
- `insights` (list of Insight, required) — Additional insights about the classification decision

### SplitterOutput

- `splits` (list of SplitterOutputSplitsItems, required)
- `isExternal` (boolean, optional)

### FileMetadata

- `pageCount` (double, optional) — The number of pages in the file. This is only set for PDF/DOCX files.
- `parentSplit` (FileMetadataParentSplit, optional) — The split metadata details. Only included if this file is a derivative of another file, for instance if it was created via a Splitter in a workflow.

### FileContents

- `rawText` (string, optional) — The raw text content of the file. This is included for all file types if the `rawText` query parameter is set to true in the endpoint request.
- `markdown` (string, optional) — Cleaned and structured markdown content of the entire file. Available for PDF and IMG file types. Only included if the `markdown` query parameter is set to true in the endpoint request.
- `pages` (list of FileContentsPagesItems, optional)
- `sheets` (list of FileContentsSheetsItems, optional)

### FileCredits

These are usage credits for parsing the file. This field will not be returned for files processed before October 7, 2025, or for customers on legacy billing systems. For more details on how credits work, see our [Credits Guide](https://docs.extend.ai/general/how-credits-work).

- `credits` (double, required) — The number of credits consumed specifically for parsing this file.

### PageRangesItems

- `start` (integer, optional) — The start page of the range.
- `end` (integer, optional, default: 750) — The end page of the range.

### ParseConfigChunkingStrategy

Strategy for dividing the document into chunks.

- `type` (enum, optional, default: page) — The type of chunking strategy. Supported values: * `page`: Chunk document by pages. * `document`: Entire document is a single chunk. Essentially no chunking. * `section`: Split by logical sections. Not supported for target=spatial.
  - Allowed values: `page`, `document`, `section`
- `options` (ParseConfigChunkingStrategyOptions, optional) — Additional options for the chunking strategy.

### ParseConfigBlockOptions

Options for controlling how different block types are processed.

- `figures` (ParseConfigBlockOptionsFigures, optional) — Options for figure blocks.
- `tables` (ParseConfigBlockOptionsTables, optional) — Options for table blocks.
- `text` (ParseConfigBlockOptionsText, optional) — Options for text blocks.

### ParseConfigAdvancedOptions

- `pageRotationEnabled` (boolean, optional, default: true) — Whether to automatically detect and correct page rotation.
- `pageRanges` (list of PageRangesItems, optional) — Limit processing to the specified page ranges. See [Page Ranges](https://docs.extend.ai/2025-04-21/product/page-ranges).
- `excelParsingMode` (enum, optional) — Controls how Excel files are parsed. * `basic`: Fast, deterministic parsing. * `advanced`: Enable layout block detection for complex spreadsheets. This mode incurs additional credits when enabled. For `.xls` files, `basic` mode is always used.
  - Allowed values: `basic`, `advanced`
- `excelSkipHiddenContent` (boolean, optional, default: false) — Whether to exclude hidden rows, columns, and sheets when parsing Excel files.
- `excelUseRawCellValues` (boolean, optional, default: false) — Whether to return raw calculated cell values instead of locale-formatted values when parsing Excel files. Useful when downstream processing needs the underlying numeric or unformatted data.
- `excelSkipCalculation` (boolean, optional, default: true) — Whether to skip formula recalculation when opening Excel workbooks. Significantly improves parsing speed for formula-heavy spreadsheets. Disable if cell values depend on volatile functions like NOW() or TODAY().
- `excelIncludeCellMetadata` (boolean, optional, default: false) — Whether to include spreadsheet cell provenance when parsing Excel files in advanced mode. When enabled, table cell block details include source cell references and formulas, text or heading block details can include source ranges, and HTML table output includes `data-cell` and `data-formula` attributes.
- `excelIncludeCellFormatting` (boolean, optional, default: false) — Whether to include spreadsheet cell formatting when parsing Excel files in advanced mode. When enabled, table cell block details include structured formatting such as bold, italic, font color, and background color, and HTML table output preserves inline cell styles.
- `verticalGroupingThreshold` (double, optional, default: 1) — Multiplier for the Y-axis threshold used to determine if text blocks should be placed on the same line or not (0.1-5.0, default 1.0). Higher values group elements that are further apart vertically. Only applies when the spatial target is set.
- `returnOcr` (ParseConfigAdvancedOptionsReturnOcr, optional) — Options for returning raw OCR data in the response.
- `alwaysConvertToPdf` (boolean, optional, default: false) — Whether to convert supported file types (images, Word documents, PowerPoint, Excel, HTML) to PDF before parsing. This can improve parsing quality for some file types and ensures spatial output with bounding boxes.
- `agenticOcrEnabled` (boolean, optional, default: false, deprecated) — Whether to enable agentic OCR corrections using VLM-based review and correction of OCR errors for messy handwriting and poorly scanned text. Deprecated - use `blockOptions.text.agentic` or `blockOptions.tables.agentic` instead for more granular control.

### ExtractionAdvancedOptionsReviewAgent

Configuration for the review agent that analyzes extraction results. When enabled, each field in the output metadata will include a `reviewAgentScore` (1-5) and may include additional `insights` of type `issue` or `review_summary` to help identify fields that may need manual review. Enabling the review agent incurs additional credits. To learn more, view the [Review Agent Documentation](https://docs.extend.ai/product/extraction/review-agent)

- `enabled` (boolean, optional) — Whether to enable the review agent for the extraction output.

### ArrayStrategy

- `type` (enum, required) — The strategy type for handling large array use cases. For most use cases this should be left null, or should be set to `large_array_heuristics`. If you're unsure of what your use case needs, reach out to the Extend team for help. - `large_array_heuristics`: Optimized for documents with very large arrays, but where latency matters. This strategy uses specialized heuristics around chunking, tables, and merging to handle large arrays efficiently over large documents. - `large_array_max_context`: Optimizes for accuracy over latency in documents with very large arrays. This strategy will do multiple passes through the entire document to ensure there is no context loss across any chunks/pages, maximizing accuracy for complex array extraction, but adding material latency. This strategy incurs additional extraction credits when enabled. - `large_array_overlap_context`: Balances accuracy and latency in documents with very large arrays. This strategy will always maintain surrounding/overlapping page context for every chunk/page that is extracted, eliminating array failure modes from context loss across page boundaries.
  - Allowed values: `large_array_heuristics`, `large_array_max_context`, `large_array_overlap_context`

### ExtractChunkingOptions

- `chunkingStrategy` (enum, optional) — The strategy to use for chunking the document.
  - Allowed values: `standard`, `semantic`
- `pageChunkSize` (integer, optional) — The size of page chunks.
- `chunkSelectionStrategy` (enum, optional) — The strategy to use for selecting chunks.
  - Allowed values: `intelligent`, `confidence`, `take_first`, `take_last`
- `customSemanticChunkingRules` (string, optional) — Custom rules for semantic chunking.

### ExcelSheetRange

- `start` (integer, required) — The starting sheet index (0-based).
- `end` (integer, required) — The ending sheet index (0-based). Must be greater than or equal to start.

### Enum

- `value` (string, required) — The value of the enum option.
- `description` (string, required) — Description of what this enum value represents.

### Insight

- `type` (enum, required) — The type of insight: - `reasoning`: Model reasoning about the extraction decision - `issue`: A potential problem or concern identified by the agentic confidence system - `review_summary`: A summary explanation from the agentic confidence system about why the field may need manual review
  - Allowed values: `reasoning`, `issue`, `review_summary`
- `content` (string, required) — The content of the insight.

### SplitterOutputSplitsItems

- `type` (string, required) — The type of the split document (set in the processor config), corresponds to the classificationId
- `observation` (string, required) — Explanation of the results
- `identifier` (string, required) — Identifier for the split document (e.g. invoice number)
- `startPage` (integer, required) — The start page of the split document
- `endPage` (integer, required) — The end page of the split document
- `classificationId` (string, required) — ID of the classification type (set in the processor config)
- `id` (string, required) — Unique ID for this split
- `fileId` (string, required) — File ID associated with this split
- `name` (string, optional) — Optional name for the split

### FileMetadataParentSplit

The split metadata details. Only included if this file is a derivative of another file, for instance if it was created via a Splitter in a workflow.

- `id` (string, required) — The ID of the split.
- `type` (string, required) — The type of the split.
- `identifier` (string, required) — The identifier of the split.
- `startPage` (integer, required) — The start page of the split.
- `endPage` (integer, required) — The end page of the split.

### FileContentsPagesItems

- `pageNumber` (integer, required) — The page number of this page in the document.
- `pageHeight` (double, optional)
- `pageWidth` (double, optional)
- `rawText` (string, optional) — The raw text content extracted from this page.
- `markdown` (string, optional) — Cleaned and structured markdown content of this page.
- `html` (string, optional) — Cleaned and structured html content of the page. Available for DOCX file types (that were not auto-converted to PDFs). Only included if the `html` query parameter is set to true in the endpoint request.

### FileContentsSheetsItems

- `sheetName` (string, required) — The name of the sheet.
- `rawText` (string, optional) — The raw text content of the sheet.

### ParseConfigChunkingStrategyOptions

Additional options for the chunking strategy.

- `minCharacters` (integer, optional) — Specify a minimum number of characters per chunk.
- `maxCharacters` (integer, optional) — Specify a maximum number of characters per chunk.

### ParseConfigBlockOptionsFigures

Options for figure blocks.

- `enabled` (boolean, optional, default: true) — Whether to include figures in the output.
- `figureImageClippingEnabled` (boolean, optional, default: true) — Whether to clip and extract images from figures.

### ParseConfigBlockOptionsTables

Options for table blocks.

- `targetFormat` (enum, optional, default: markdown) — The target format for the table blocks. Supported values: * `markdown`: Convert table to Markdown format * `html`: Convert table to HTML format
  - Allowed values: `markdown`, `html`
- `tableHeaderContinuationEnabled` (boolean, optional, default: false) — Whether to automatically copy table headers to headerless tables on subsequent pages when they have matching column counts. Useful for multi-page tables.
- `cellBlocksEnabled` (boolean, optional, default: false) — Whether to include individual table cell blocks in the output. When enabled, each cell in a table will be represented as a separate block with its own bounding box and content and will be `children` of the table block.
- `agentic` (ParseConfigBlockOptionsTablesAgentic, optional) — Options for agentic table processing using VLM-based review and correction. Enabling this incurs additional credits on pages where agentic table correction is triggered.
- `enabled` (boolean, optional, default: true, deprecated) — This option is deprecated and will have no effect. It will be removed in the next API version.

### ParseConfigBlockOptionsText

Options for text blocks.

- `signatureDetectionEnabled` (boolean, optional, default: true) — Whether an additional vision model will be utilized for advanced signature detection. Recommended for most use cases, but should be disabled if signature detection is not necessary and latency is a concern.
- `agentic` (ParseConfigBlockOptionsTextAgentic, optional) — Options for agentic text processing using VLM-based review and correction. Enabling this incurs additional credits on pages where agentic text correction is triggered.

### ParseConfigAdvancedOptionsReturnOcr

Options for returning raw OCR data in the response.

- `words` (boolean, optional, default: false) — Whether to include word-level OCR data in the response, including word bounding boxes and confidence scores. This meaningfully impacts the response size, consider using the signed url return format.

### ParseConfigBlockOptionsTablesAgentic

Options for agentic table processing using VLM-based review and correction. Enabling this incurs additional credits on pages where agentic table correction is triggered.

- `enabled` (boolean, optional, default: false) — Whether to enable agentic table processing.
- `customInstructions` (string, optional) — Custom instructions to guide the agentic table processing.

### ParseConfigBlockOptionsTextAgentic

Options for agentic text processing using VLM-based review and correction. Enabling this incurs additional credits on pages where agentic text correction is triggered.

- `enabled` (boolean, optional, default: false) — Whether to enable agentic text processing for OCR corrections.
- `customInstructions` (string, optional) — Custom instructions to guide the agentic text processing.

## Examples

**Request**

```json
{
  "processorId": "processor_id_here"
}
```

**Response**

```json
{
  "success": true,
  "processorRun": {
    "object": "document_processor_run",
    "id": "dpr_l39vTgFDiB13heVuMQnUa",
    "processorId": "dp_SmJyN3LMx9kW_YmFTxTha",
    "processorVersionId": "dpv_YrgxmNn83sAO0JChhmLpa",
    "processorName": "My Processor",
    "status": "PROCESSING",
    "output": {},
    "reviewed": false,
    "edited": false,
    "edits": {},
    "type": "EXTRACT",
    "config": {
      "type": "EXTRACT",
      "schema": {
        "type": "object",
        "required": [
          "name",
          "age"
        ],
        "properties": {
          "age": {
            "type": [
              "number",
              "null"
            ],
            "description": "The age of the person"
          },
          "name": {
            "type": [
              "string",
              "null"
            ],
            "description": "The name of the person"
          }
        },
        "additionalProperties": false
      },
      "advancedOptions": {
        "advancedMultimodalEnabled": true,
        "advancedFigureParsingEnabled": false,
        "chunkingOptions": {
          "chunkingStrategy": "standard",
          "chunkSelectionStrategy": "intelligent"
        }
      }
    },
    "files": [
      {
        "object": "file",
        "id": "file_0QyyVL9rrOd0_WllDDCNa",
        "name": "My File",
        "metadata": {},
        "createdAt": "2025-05-12T21:22:37.318000+00:00",
        "updatedAt": "2025-05-12T21:22:37.324000+00:00",
        "type": "PDF",
        "usage": {
          "credits": 10
        }
      }
    ],
    "mergedProcessors": [],
    "url": "https://dashboard.extend.ai/runs/dpr_l39vTgFDiB13heVuMQnUa",
    "initialOutput": {},
    "usage": {
      "credits": 3.5
    }
  }
}
```