Parashift/Comparisons/Parashift vs. OpenAI
Comparison · OpenAI

Parashift vs. OpenAI

Both run on language models. With Dokumentli®, Parashift uses its own vision language model built for one thing: business documents. The difference lies in the model class, in how it is operated, and in the platform around it.

88.4% cANLS@0.8 in our own benchmarkRuns on a single 24 GB GPUProcessing in CH, DE or the EU
Benchmark 08/2026 · 3368 documentsMeasured
Dokumentli®88.4%
Gemini 3.6 Flash83.5%
GPT-5.6 Sol81.8%
Claude Opus 4.881.0%
cANLS@0.8 across five datasets, our own measurement
The real question

Not whether a language model. Which one, where it runs, and what sits around it.

The question «why not just use GPT for this» assumes an opposition that does not exist. Parashift works with a language model too. Dokumentli® is a small special-purpose vision language model, trained for exactly one job: classifying business documents and reading fields out of them.

A general-purpose frontier model can do that as well. It can also write code, reason and produce prose. That breadth has a price: accuracy on documents, speed, cost per page, and the freedom to decide where processing happens.

The difference sits on three levels. Which model does the work. Where it runs. And what else has to happen between document intake and the target system.

Three levels

Where the difference actually sits.

A specialist, not a generalist

Dokumentli® is trained on business documents and deliberately cannot do anything else. No open reasoning, no code, no conversation. That is why it is more accurate at its own job.

Measured, not claimed

In our own benchmark from August 2026 across 3368 real business documents, Dokumentli® reaches 88.4% cANLS@0.8, ahead of Gemini 3.6 Flash at 83.5% and GPT-5.6 Sol at 81.8%.

Runs on your own hardware

Dokumentli® runs on a single GPU with 24 GB of VRAM, on premise or in a private cloud. A hosted API service does not offer that choice.

Fix the processing location

Switzerland, Germany or the EU, contractually assured. For banks, insurers and hospitals that is the condition under which a project starts at all.

A platform, not an endpoint

Separation and classification come before extraction, then confidence checks, validation against master data and handover to the target system. A model call covers one of those steps.

Evidence chain per field

It is logged which document and which page a value came from, how certain it was and who corrected it. That is the basis for audit.

Where each fits

Which model is right when.

01 · REASONING

Frontier model

Reasoning, code, open conversation, creative writing. Dokumentli® is not built for any of that.

02 · ONE-OFF

Frontier model

One stack, one analysis, no audit duty. Fast and without setup.

03 · DOCUMENT STREAM

Dokumentli®

Daily throughput, identical results, confidence, handover to the core system.

04 · REGULATED

Dokumentli®

Processing location, evidence and logging are part of the requirement.

Side by side

General-purpose frontier model and Parashift with Dokumentli®.

TaskGeneral-purpose frontier modelParashift with Dokumentli®
Classify business documents and extract fieldsPossible, at lower accuracyCore task, trained for it
Accuracy in our benchmark 08/2026GPT-5.6 Sol: 81.8% cANLS@0.8Dokumentli®: 88.4% cANLS@0.8
Open reasoning, code, conversationCore taskNot provided for
On-premise operationNot provided forA single GPU with 24 GB VRAM
Fixing the processing locationDepends on the providerSwitzerland, Germany or the EU
Splitting stacks into documentsNot provided forPart of the processing
Confidence value per fieldNo calibrated valuePer field, with a threshold
Evidence for auditNo per-field logDocument, page, timestamp per value
Cost at high volumeGrows with every tokenCalculable per page or per node
The benchmark

Our own measurement, August 2026.

88.4%

cANLS@0.8, highest of seven models tested

Parashift benchmark 08/2026
3368

real business documents across five datasets

Parashift benchmark 08/2026
24 GB

of VRAM are enough for production use

Dokumentli® data sheet

Not an either-or.

Many customers use both: a frontier model for analysis, reasoning and text work, Dokumentli® for the document stream that has to reach the core system every day.

See all comparisons
Frequently asked

Questions about this comparison.

Is Parashift not just a language model as well?
It is, and that is the point. Dokumentli® is a vision language model, but a small and specialised one. It is trained on business documents and deliberately cannot code or reason. That restriction is what makes it more accurate and cheaper at its own job than a generalist.
Why would a smaller model be more accurate than GPT?
Because the task is narrow. In our own benchmark from August 2026 across 3368 real business documents from five datasets, Dokumentli® reaches 88.4% cANLS@0.8 and GPT-5.6 Sol 81.8%. No model had seen the test documents before.
What can Dokumentli® not do?
Open reasoning, programming, general knowledge questions outside the document, and creative writing. A frontier model is the right tool for all of that, and we say so.
Can we run the model ourselves?
Yes. Dokumentli® runs on a single GPU with 24 GB of VRAM, on premise or in a private cloud. No document content leaves your environment.
What is the difference between the model and the platform?
The model reads a document. The platform first splits the stack, classifies per page, checks confidence values, validates against master data and hands over structured data to the target system. A model call alone covers one of those steps.
How do costs behave at high volume?
Token-based billing grows with page count and answer length. Parashift bills per page, or per node when you run it yourself. At several thousand documents a day that matters more than the price per call.

See the difference on your own documents.

Book a no-obligation briefing: we run your documents through Dokumentli® and a frontier model and show the results side by side.

Reply the same working day · no sales pressure