Datatera Logo
DATATERA.ai
Evidence ledger / updated 14 Aug 2026

Document extraction model benchmark

Four extraction cases and exact provider routes. Current runs require runtime identity evidence; the original results remain visible with their historical run date and evidence status.

B1
CRM CSV
5 rows
B2
Financial CSV
332 rows
B3
Scanned invoice
OCR + tables
B4
Investor table
32 rows

Unified results

Historical and current models

The original 27-model run is the baseline. Ten newer models extend the same table; run date and evidence status distinguish historical results from the latest strict audit.

Model / routeB1B2B3B4ValidQualityValid timeEst. costRunEvidence
GPT-4.1 mini
gpt-4.1-mini
100%90%100%100%4 / 497.5%994s$0.0902026-03-14historical
Gemini 3.1 Flash
gemini-3.1-flash
100%90%100%100%4 / 497.5%579s$0.0342026-03-14historical
DeepSeek V3
openrouter/deepseek-v3
100%90%100%100%4 / 497.5%217s~$0.102026-03-14historical
Gemma 4 31B
google/gemma-4-31b-it
100%90%87.5%100%4 / 494.4%761.83s$0.0289982026-08-14audited
Step 3.5 Flash
openrouter/step-3.5-flash
100%90%87.5%100%4 / 494.4%171sFree2026-03-14historical
DeepSeek V3 Free
openrouter/deepseek-v3-free
100%90%87.5%100%4 / 494.4%182sFree2026-03-14historical
Kimi K2.5
openrouter/kimi-k2.5
100%90%87.5%100%4 / 494.4%212sNot recorded2026-03-14historical
MiMo V2 Flash
openrouter/mimo-v2-flash
100%90%87.5%100%4 / 494.4%221sNot recorded2026-03-14historical
GLM 5
openrouter/glm-5
100%90%87.5%100%4 / 494.4%246sNot recorded2026-03-14historical
Arcee Trinity
openrouter/arcee-trinity
100%90%87.5%100%4 / 494.4%259sFree2026-03-14historical
Kimi K2 (Groq)
groq/kimi-k2
100%90%87.5%100%4 / 494.4%270s$0.0392026-03-14historical
Qwen3 235B
openrouter/qwen3-235b
100%90%87.5%100%4 / 494.4%325s~$0.082026-03-14historical
Grok 4.1 Fast
openrouter/grok-4.1-fast
100%90%87.5%100%4 / 494.4%343sNot recorded2026-03-14historical
MiniMax M2.5
openrouter/minimax-m2.5
100%90%87.5%100%4 / 494.4%350sNot recorded2026-03-14historical
Claude Sonnet 4.6
openrouter/claude-sonnet-4-6
100%90%87.5%100%4 / 494.4%560s~$1.102026-03-14historical
GPT-5.4
gpt-5.4
100%90%100%80%4 / 492.5%495s$0.7152026-03-14historical
Qwen3.6 27B
qwen/qwen3.6-27b
100%90%87.5%Unscored3 / 492.5%*707.40s$0.0406072026-08-14partial
GPT-5.6 Terra
openai/gpt-5.6-terra
100%90%87.5%80%4 / 489.4%102.19s$0.2265982026-08-14audited
Muse Spark 1.2
meta/muse-spark-1.2
100%90%87.5%80%4 / 489.4%115.53s$0.2479492026-08-14audited
GLM 5.2
z-ai/glm-5.2
100%90%87.5%80%4 / 489.4%207.88s$0.1359882026-08-14audited
Gemini 3.7 Flash
google/gemini-3.7-flash
100%90%87.5%80%4 / 489.4%260.51s$0.1509342026-08-14audited
Muse Glimmer 30B
meta/muse-glimmer-30b
100%90%87.5%80%4 / 489.4%271.41s$0.0809122026-08-14audited
Mistral Large
openrouter/mistral-large
100%90%87.5%80%4 / 489.4%284s~$0.502026-03-14historical
GPT-4.1 Nano
gpt-4.1-nano
100%90%87.5%80%4 / 489.4%372s$0.0232026-03-14historical
Gemini 3.6 Flash
google/gemini-3.6-flash
100%90%87.5%80%4 / 489.4%433.51s$0.3188942026-08-14audited
Llama 4 Scout (Groq)
groq/llama-4-scout
100%90%87.5%73%4 / 487.6%210s$0.0092026-03-14historical
Llama 3.1 8B (Groq)
groq/llama-3.1-8b
100%90%87.5%73%4 / 487.6%226sFree tier2026-03-14historical
Llama 3.3 70B (Groq)
groq/llama-3.3-70b
100%90%87.5%73%4 / 487.6%289s$0.0222026-03-14historical
Qwen3 32B (Groq)
groq/qwen3-32b
100%90%87.5%73% flaky3 / 4 stable87.6%*315sFree tier2026-03-14partial
GPT-4.1
gpt-4.1
100%90%100%60%4 / 487.5%348s$0.4582026-03-14historical
GPT-5.3 Codex
gpt-5.3-codex
100%90%100%60%4 / 487.5%462s$0.7212026-03-14historical
GPT-5 Mini
gpt-5-mini
100%90%100%60%4 / 487.5%792s$0.1062026-03-14historical
Claude Sonnet 5
anthropic/claude-sonnet-5
100%90%87.5%60%4 / 484.4%398.53s$0.7842922026-08-14audited
Llama 4 Maverick (Groq)
groq/llama-4-maverick
100%90%87.5%60%4 / 484.4%225sFree tier2026-03-14historical
GPT OSS 120B (Groq)
groq/gpt-oss-120b
100%90%87.5%60%4 / 484.4%358s$0.0132026-03-14historical
GPT OSS 20B (Groq)
groq/gpt-oss-20b
100%90%87.5%60%4 / 484.4%407s$0.0142026-03-14historical
DeepSeek V4 Flash
deepseek/deepseek-v4-flash-0731
100%Unscored87.5%60%3 / 482.5%*237.53s$0.0036182026-08-14partial

* Quality is averaged across valid cases. Unscored or flaky cases remain labeled and are not converted to zero.

Scoring contract

Measured, not inferred

  1. 01 / identityThe requested OpenRouter route must appear in runtime audit evidence.
  2. 02 / executionAt least one audited LLM call must have handled the extraction.
  3. 03 / qualityGround-truth values, rows, and columns are checked per case.
  4. 04 / failuresFallbacks and failed cases remain visible but never receive a score.

Benchmark your own documents

Use the same evidence contract on a representative sample from your workflow.

Book a benchmark review