> ## Documentation Index
> Fetch the complete documentation index at: https://docs.extend.ai/llms.txt
> Use this file to discover all available pages before exploring further.
>
> ## API version
> The current API version is `2026-02-09`, served at the site root (no version prefix in URLs).
> If this page URL contains `/2025-04-21/` or `/2024-12-23/`, you are reading an older API version.
> Prefer the current docs at https://docs.extend.ai/llms.txt unless the user explicitly needs that older version.
> Do not treat older-version pages as the source of truth for new integrations.

# Changelog

> Updates and improvements to the Extend platform.

Stay up to date on what's shipping in the Extend platform.

## Multifile evaluation set items

> Evaluation set items can now reference an ordered package of files via fileIds, and are evaluated as a single multifile extraction run on the 2026-02-09 API.

Evaluation sets can now hold **multifile items**, so you can measure the accuracy of an extractor that runs over a [package of files](/2026-02-09/extraction/multifile) rather than a single document.

* **`fileIds`** on `POST /evaluation_sets/{id}/items` — pass an ordered array of 2–50 unique file IDs instead of `fileId`. Exactly one of `fileId` or `fileIds` must be provided per item. Multifile items are only supported on evaluation sets attached to a JSON-schema extractor.
* **`files`** on `EvaluationSetItem` and `EvaluationSetItemSummary` — an ordered array of file summaries, populated for multifile items and `null` otherwise.
* **`file`** is now nullable — `null` for multifile items, matching the `file` / `files` convention on `ExtractRun`.

When you run the set, each multifile item is evaluated as one multifile extraction run with a single `expectedOutput` covering the whole package. See [Add a multifile item](/2026-02-09/evaluation/creating-evaluation-sets#add-a-multifile-item).

## Operator

**Operator** is now available as a third extraction base processor, `extraction_operator`. For long, table-heavy documents an extraction agent replaces the chunked pipeline: it plans, extracts repetitive structure programmatically, verifies totals and row counts against the source, and submits a schema-validated result. For other documents, the agent audits and corrects any fields the Review Agent flags with low confidence. Both paths require a single-file JSON Schema extraction.

**API**

* Select it with `"baseProcessor": "extraction_operator"` on `POST /extractors`, `POST /extractors/{id}`, `POST /extractors/{id}/versions`, inline `config` on `POST /extract` and `POST /extract_runs`, and inline workflow extractor configs.
* Settings Operator manages (Review Agent, multimodal, reasoning insights, array strategy, chunking) are applied automatically when a run is created.
* Run `usage.breakdown[].charges[]` can now include a base charge with `product: "extraction_agent"` and `unit: "usd"` charges with `product: "agent_usage"`.

**Pricing**

Operator runs bill a flat **2 credits/page** with no extraction surcharges, plus **Agent usage**: the agent's model provider cost converted to credits at your plan's rate.

Read the [Operator guide](/2026-02-09/extraction/operator) and [Operator Pricing](/2026-02-09/general/operator-pricing).

## Operator

**Operator** is now available as a third extraction base processor, `extraction_operator`. For long, table-heavy documents an extraction agent replaces the chunked pipeline: it plans, extracts repetitive structure programmatically, verifies totals and row counts against the source, and submits a schema-validated result. For other documents, the agent audits and corrects any fields the Review Agent flags with low confidence. Both paths require a single-file JSON Schema extraction.

**API**

* Select it with `"baseProcessor": "extraction_operator"` on `POST /extractors`, `POST /extractors/{id}`, `POST /extractors/{id}/versions`, inline `config` on `POST /extract` and `POST /extract_runs`, and inline workflow extractor configs.
* Settings Operator manages (Review Agent, multimodal, reasoning insights, array strategy, chunking) are applied automatically when a run is created.
* Run `usage.breakdown[].charges[]` can now include a base charge with `product: "extraction_agent"` and `unit: "usd"` charges with `product: "agent_usage"`.

**Pricing**

Operator runs bill a flat **2 credits/page** with no extraction surcharges, plus **Agent usage**: the agent's model provider cost converted to credits at your plan's rate.

Read the [Operator guide](/2026-02-09/extraction/operator) and [Operator Pricing](/2026-02-09/general/operator-pricing).

## Parse output metadata: page rotation, original dimensions, and file type

Parse runs now include a **`metadata`** object on `output`, alongside `chunks`:

* **`originalMimeType`** / **`finalMimeType`** — the file's media type before and after any format conversion.
* **`pages`** — an array of per-page entries, each with:
  * **`number`** — the page number, matching `metadata.page.number` on that page's blocks.
  * **`rotationApplied`** — degrees Extend rotated the page clockwise to make it upright. `0` if detection ran and the page was already upright, `null` if rotation detection was disabled for the run.
  * **`originalPageWidth`** / **`originalPageHeight`** — the file's true page size, independent of rendering resolution.
  * **`dpi`** — the DPI that `boundingBox`/`polygon` coordinates on the page are scaled to; multiply the original dimensions by `dpi / 72` to bring them into that same coordinate space.

`metadata` is `null` for parse runs that completed before this field was introduced, and `pages` is `null` when page-level metadata isn't available.

Bounding boxes and polygons are always returned in the corrected (upright) frame. To draw a citation or highlight on your own copy of the original file, use these fields to map coordinates back — see [Reconciling coordinates against your original file](/2026-02-09/parsing/response-format#reconciling-coordinates-against-your-original-file).

## Excel cell metadata and formatting in Parse

> Advanced Excel parsing can preserve source cell references, formulas, and structured cell formatting in parse output.

Advanced Excel parsing can now preserve source-cell provenance and formatting in parse output.

* Set **`advancedOptions.excelIncludeCellMetadata`** to return source cell references such as `B2` or merged ranges such as `A1:C1`; formula cells also include the source formula text.
* Set **`advancedOptions.excelIncludeCellFormatting`** to return structured cell formatting such as bold, italic, font color, and background color.
* For HTML table output, metadata is also emitted as `data-cell` and `data-formula` attributes, and formatting preserves inline cell styles.

Both options are off by default and require `advancedOptions.excelParsingMode: "advanced"`. See the [Parse configuration guide](/parsing/configuration) for setup details.

## Custom instructions for figure parsing

> Steer figure parsing with blockOptions.figures.customInstructions, injected into the vision model prompt that analyzes and summarizes each figure.

You can now provide custom instructions for figure parsing via `blockOptions.figures.customInstructions`. The instructions (up to 2,000 characters) are injected into the vision model prompt that analyzes and summarizes each figure, letting you steer descriptions toward your use case — for example, domain-specific terminology or details to always capture from diagrams.

Available on `parse_performance` >= `2.0.0` and `parse_light` >= `1.0.0`, with `blockOptions.figures.enabled: true`.

```json
{
  "config": {
    "blockOptions": {
      "figures": {
        "enabled": true,
        "customInstructions": "Describe all medical imaging figures using radiology terminology. Always note any visible measurement scales or annotations."
      }
    }
  }
}
```

In the dashboard, you can also configure figure custom instructions from the parser configuration panel in Studio.

## Processor cost preview and charge-level usage breakdown

> usage.breakdown entries can include a charges array itemizing cost drivers — billing product, unit, quantity, credits, and page-level detail.

Run responses now include a finer-grained credit breakdown. Each entry in `usage.breakdown` can now include a `charges` array that itemizes the cost drivers behind that run's credits — for example base processor, review agent, or page-specific add-ons like agentic text correction. Each charge lists the billing product, unit (`page` or `cell`), quantity, credits consumed, and applicable page numbers when billing is page-scoped.

For background on billing units, see [How credits work](/general/how-credits-work). For the run usage shape, see [`usage` on the extract run response](/api-reference/endpoints/extract/get-extract-run#response.body.usage).

* **API** — optional `charges` on each `usage.breakdown` entry with product, unit, quantity, credits, and page-level detail where applicable

## Cancel in-flight parse runs

> POST /parse_runs/{id}/cancel aborts an in-progress parse run on the 2026-02-09 API and sets its status to CANCELLED.

You can now cancel parse runs that are still processing. Call `POST /parse_runs/{id}/cancel` on the 2026-02-09 API to cancel an in-progress parse run and set its status to `CANCELLED`. Only runs with status `PROCESSING` can be cancelled.

* **`POST /parse_runs/{id}/cancel`** — abort an in-progress parse run (2026-02-09 API)

## Validation rules: INCLUDES and INCLUDESANY operators

> Workflow validation formulas add INCLUDES and INCLUDESANY operators for case-insensitive substring checks across array fields.

Workflow validation formulas now support **`INCLUDES`** and **`INCLUDESANY`** for checking whether text appears inside any element of an array field. Unlike **`CONTAINS`** and **`CONTAINSALL`**, which match whole array elements, these operators perform case-insensitive substring checks—useful when you need to verify that one or more keywords appear in fields such as line item descriptions.

See [Formulas](/workflows/formulas) for the full list of validation operators.

* **`INCLUDES(array, text)`** — returns true if `text` appears anywhere inside any element of `array`
* **`INCLUDESANY(array, value1, value2, ...)`** — returns true if any of the values appears anywhere inside any element of `array`

## OCR word confidence on parse chunk and block metadata

> Parse chunk and block metadata now include minOcrConfidence and avgOcrConfidence, the per-word OCR confidence across each region.

Parse runs now include **`minOcrConfidence`** and **`avgOcrConfidence`** on **chunk** and **block** metadata — the minimum and average per-word OCR confidence across the words in that region.

Both fields are returned on every chunk and block, and are **`null`** when word-level confidence isn't produced for that region. Values are in the range **0–1**.

For word-level confidence scores, set [`returnOcr.words`](/api-reference/endpoints/parse/create-parse-run) to `true`.

## Generate extractor JSON schemas via the REST API when you create an extractor

> Pass generate on POST /extractors with one to five sample files and Extend generates a JSON extraction schema from your examples.

For API version `2026-02-09`, you can optionally pass `generate` on [`POST /extractors`](/2026-02-09/api-reference/endpoints/extract/create-extractor) instead of supplying `config`. Provide one to five sample inputs as Extend file IDs or file URLs and Extend will generate a JSON extraction schema from those examples and return the extractor with the schema applied. Optionally add `generate.instructions` (up to 2,500 characters) to provide additional context about the document type or requirements for how values should be extracted. You cannot combine `generate` with `cloneExtractorId` or `config`.

* **`generate.files`**: 1–5 entries, each a file `{ id }` or `{ url }`.
* **`generate.instructions`**: optional free-text guidance to steer schema generation.

## Extraction Performance 4.8.0

> Extraction Performance 4.8.0 is the latest base processor version, upgrading the base models used for core extraction and large array strategies.

Extraction Performance **4.8.0** is now the latest base processor version. It upgrades the base models used for core extraction and large array strategies

## Extraction Performance 4.8.0

> Extraction Performance 4.8.0 is the latest base processor version, upgrading the base models used for core extraction and large array strategies.

Extraction Performance **4.8.0** is now the latest base processor version. It upgrades the base models used for core extraction and large array strategies

## Extraction Performance 4.8.0

> Extraction Performance 4.8.0 is the latest base processor version, upgrading the base models used for core extraction and large array strategies.

Extraction Performance **4.8.0** is now the latest base processor version. It upgrades the base models used for core extraction and large array strategies

## Extraction citation mode control (line, word, block)

> Set citationMode to line, word, or block so extraction citation polygons match the granularity you need; parse 2.0.0-beta now supports bounding box citations.

When bounding box citations are enabled on an extractor, you can set **`citationMode`** in **`advancedOptions`** to **`line`**, **`word`**, or **`block`** so citation polygons match the granularity you want. If you leave it unset, behavior matches what you have today (line-based citation processing plus block overlap handling across supported parse engines).

Extraction pipelines that use the **parse 2.0.0-beta** engine can now return bounding box citations for extracted values, so you are not limited to older parse versions when you need spatial references.

See [Citations](/extraction/response-format#citations) for how citations appear on extracted fields.

* **`citationMode`** — optional; configure in extractor advanced options in Extend Studio or on extractors (`line`, `word`, or `block`)

## Batch Extract and Batch Parse APIs

> New /extract/batch and /parse/batch endpoints queue thousands of files for background processing without running into rate limits.

Two new endpoints make bulk background processing easier:

* **`/extract/batch`** — queue thousands of files for extraction in the background without running into rate limits.
* **`/parse/batch`** — bulk background parse operations.

## Splitter Composer

> Composer can now optimize splitters automatically from eval sets, in addition to extractors and classifiers.

Composer is now available for splitters, so you can optimize them automatically from eval sets. Previously, Composer was only available for extractors and classifiers.

![](/_fern-img/e4180abcaf27e8aa658e273b1a83acb58601b36602f3ad57fb8910ea53f56599.webp)

If you're working on splitting, check out our [Splitter Benchmark](https://www.extend.ai/resources/document-splitting-benchmark) to see how we evaluate models.

## Formula handling in Parse

> The new parse engine can parse formula blocks as LaTeX, available as an advanced option for extracting math content from documents.

The new parse engine can now parse formula blocks as LaTeX. Enable it as an advanced option to extract math content directly from documents.

## Formula handling in Parse

> The new parse engine can parse formula blocks as LaTeX, available as an advanced option for extracting math content from documents.

The new parse engine can now parse formula blocks as LaTeX. Enable it as an advanced option to extract math content directly from documents.

## Workflows create/update API

> Create, update, and delete workflows and workflow versions entirely via API for programmatic, CI/CD-managed document pipelines.

You can now create, update, and delete workflows and workflow versions entirely via API. Previously, workflows could only be managed in the Extend dashboard. This makes it easier to generate workflows programmatically, check configs into code, manage them with CI/CD, and build them with agents like Claude Code.

```typescript
await client.workflows.update("workflow_abc123", {
    steps: [
      { "name": "trigger", "type": "TRIGGER", "next": [{ "step": "parse" }] },
      { "name": "parse", "type": "PARSE", "next": [{ "step": "split" }] },
      {
        "name": "split",
        "type": "SPLIT",
        "config": {
          "splitter": { "id": "spl_abc123", "version": "0.1" }
        },
        "next": [
          { "step": "extract_invoice", "classificationId": "cls_invoice" },
          { "step": "extract_receipt", "classificationId": "cls_receipt" },
          { "step": "review", "classificationId": "cls_other" }
        ]
      },
      { "name": "review", "type": "HUMAN_REVIEW", "next": [{ "step": "collect" }] },
      { "name": "collect", "type": "COLLECT", "next": [{ "step": "webhook" }] },
      { "name": "webhook", "type": "WEBHOOK_RESPONSE" }
    ]
});
```

See the [full documentation](/workflows/configuring-workflows).

_Showing the 20 most recent of 40 entries. Append `/llms.txt` to the changelog URL for the complete index._