Parashift/Dokumentli/Benchmark
Benchmark · as of 25 September 2026

Document AI put to the test: Dokumentli against five frontier models.

3365 real business documents, 5 datasets, 6 models, one task: extract structured fields with no prior knowledge of the layout. Dokumentli reaches 89.1 percent extraction accuracy, the best result in the field, and wins all five benchmarks against the strongest reference model.

3365 documents, 6 modelsMethod and exclusions disclosedZones Zurich, Frankfurt, Amsterdam
Benchmark 09/20266 models
Avg accuracy89.1 %
Benchmarks won5 of 5
Avg speed1.14 doc/s
Context budget8192 tokens
Reference modelsup to 64000
cANLS at threshold 0.8, macro average across 5 benchmarks
The result

Specialisation beats model size.

The field: Claude Opus 5.5 from Anthropic, GPT-6 Astra from OpenAI, Gemini 3.8 Flash from Google, Qwen3.8-27B from Alibaba and the open model Gemma 4 26B-A4B-IT. All five are general purpose models trained at vastly greater expense. Dokumentli is a compact vision language model specialised on European business documents, and it leads all five on accuracy.

89.1 %

average extraction accuracy, best of all six models

cANLS@0.8, 5 benchmarks
5 of 5

benchmarks won against Claude Opus 5.5, the strongest reference model

lead of 0.8 to 7.4 percentage points
1.9 ×

faster than the most accurate reference model

1.14 against 0.60 documents per second
1/8

of the context budget of the largest reference models

8192 against up to 64000 tokens
Every figure

Accuracy by benchmark and model.

BenchmarkDocumentsDokumentli 5.3Claude Opus 5.5GPT-6 AstraGemini 3.8Qwen3.8-27BGemma 4 26B
Invoice processingMixed accounts payable invoices41389.1 %86.1 %86.0 %85.5 %83.7 %71.1 %
E-commerceOrder and shipping documents147484.5 %77.1 %75.8 %76.4 %79.1 %63.1 %
LogisticsFreight and contract documents92782.2 %79.7 %79.8 %78.3 %75.4 %62.9 %
AutomotiveVehicle workshop invoices12696.8 %96.0 %95.8 %94.6 %93.2 %–
Real estateRental and utility statements42593.4 %87.5 %88.7 %85.8 %78.2 %70.3 %
Average336589.1 %85.3 %85.2 %84.1 %81.9 %66.9 %

cANLS at threshold 0.8, macro average per benchmark. Reference values come from human verified fields of real business documents. The largest lead is on order and shipping documents at 7.4 percentage points, the narrowest on vehicle workshop invoices at 0.8.

Speed

The fastest model in the field on average.

In production document processes throughput decides cycle time and infrastructure cost. Claude Opus 5.5 is the fastest reference model but stays clearly behind. GPT-6 Astra, Gemini 3.8 Flash and above all Qwen3.8-27B fall much further behind.

ModelDocuments per second
Dokumentli 5.31.14
Claude Opus 5.50.60
Gemma 4 26B0.34
GPT-6 Astra0.21
Gemini 3.8 Flash0.20
Qwen3.8-27B0.08

Averaged over all 5 benchmarks under identical conditions. Reference models via OpenRouter, Dokumentli through its own dedicated endpoint.

What follows

Accuracy, speed and operations are connected.

Runs on a single GPU

Thanks to its small context budget Dokumentli runs through the Dokumentli Node component on standard PC hardware with one GPU. No dedicated server infrastructure, no cloud connection, no document leaving the building.

Cost per document

Higher throughput per compute unit at leading accuracy means lower cost per processed document. That shows at volume, as in the e-commerce benchmark with 1474 and logistics with 927 documents.

European data sovereignty

Processing in the zones Zurich, Frankfurt and Amsterdam, on request entirely inside your own data centre. ISO 27001, SOC 2 Type 2, BSI C5. The model is built and operated in Europe, not merely hosted there.

Questions about the benchmark

Method, limits and pricing.

What does cANLS at a threshold of 0.8 measure?

cANLS stands for case-insensitive Average Normalized Levenshtein Similarity. Each extracted field is compared against a human verified reference value. A field counts as correct from a similarity of 0.8 upwards. This is a common variant of the ANLS industry standard for document extraction. Precision and recall on cell level and 95 percent confidence intervals were calculated in addition.

Was Dokumentli trained on this test data?

No. Dokumentli is specialised on business documents from the same industries and document types, but it was not trained on the benchmark datasets themselves. All 3365 test documents were unseen by the model.

Is this an independent test?

No, this is an internal benchmark run by Parashift. That is why the datasets, the metric, the model versions, the context budgets and every exclusion are disclosed. The five reference models were queried through standard APIs via OpenRouter with zero shot prompting, Dokumentli through its own dedicated endpoint.

Why is the Gemma 4 figure missing for the automotive benchmark?

Gemma 4 26B-A4B-IT was not re-tested against Dokumentli 5.3. The figures shown come from an earlier test run under the same conditions, and no measurement exists for the automotive benchmark.

How much compute does Dokumentli need?

Dokumentli works with a budget of 8192 tokens per request. The reference models were queried with up to 64000 tokens, roughly eight times the context per document, at lower or comparable accuracy. Through the Dokumentli Node component the model runs on standard PC hardware with a single GPU, without dedicated server infrastructure and without a cloud connection.

Where are documents processed?

In the zones Zurich, Frankfurt and Amsterdam, chosen by the customer. On request entirely inside your own data centre or on your own hardware. Parashift is certified to ISO 27001 and SOC 2 Type 2 and meets BSI C5.

How high is the automation rate in practice?

In productive Parashift processes the automation rate exceeds 90 percent. It measures the share of documents that pass through without human intervention and depends on confidence thresholds, validation rules and master data matching. It is measured per process and document type.

What does Dokumentli cost?

Dokumentli starts at 990 euros per month. The full Parashift platform with classification, separation, validation and evidence chain starts at 1590 euros per month.

The honest test is your own document stack.

Send us an anonymised set of your documents. You get back the extraction, the confidence values and the results field by field on your own paperwork, not on ours.

Dokumentli from 990 euros per month, platform from 1590 euros per month. Answer within one working day.