Evidence ledger / updated 14 Aug 2026
Document extraction model benchmark
Four extraction cases and exact provider routes. Current runs require runtime identity evidence; the original results remain visible with their historical run date and evidence status.
B1
CRM CSV
5 rows
B2
Financial CSV
332 rows
B3
Scanned invoice
OCR + tables
B4
Investor table
32 rows
Unified results
Historical and current models
The original 27-model run is the baseline. Ten newer models extend the same table; run date and evidence status distinguish historical results from the latest strict audit.
| Model / route | B1 | B2 | B3 | B4 | Valid | Quality | Valid time | Est. cost | Run | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|
GPT-4.1 mini gpt-4.1-mini | 100% | 90% | 100% | 100% | 4 / 4 | 97.5% | 994s | $0.090 | 2026-03-14 | historical |
Gemini 3.1 Flash gemini-3.1-flash | 100% | 90% | 100% | 100% | 4 / 4 | 97.5% | 579s | $0.034 | 2026-03-14 | historical |
DeepSeek V3 openrouter/deepseek-v3 | 100% | 90% | 100% | 100% | 4 / 4 | 97.5% | 217s | ~$0.10 | 2026-03-14 | historical |
Gemma 4 31B google/gemma-4-31b-it | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 761.83s | $0.028998 | 2026-08-14 | audited |
Step 3.5 Flash openrouter/step-3.5-flash | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 171s | Free | 2026-03-14 | historical |
DeepSeek V3 Free openrouter/deepseek-v3-free | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 182s | Free | 2026-03-14 | historical |
Kimi K2.5 openrouter/kimi-k2.5 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 212s | Not recorded | 2026-03-14 | historical |
MiMo V2 Flash openrouter/mimo-v2-flash | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 221s | Not recorded | 2026-03-14 | historical |
GLM 5 openrouter/glm-5 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 246s | Not recorded | 2026-03-14 | historical |
Arcee Trinity openrouter/arcee-trinity | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 259s | Free | 2026-03-14 | historical |
Kimi K2 (Groq) groq/kimi-k2 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 270s | $0.039 | 2026-03-14 | historical |
Qwen3 235B openrouter/qwen3-235b | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 325s | ~$0.08 | 2026-03-14 | historical |
Grok 4.1 Fast openrouter/grok-4.1-fast | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 343s | Not recorded | 2026-03-14 | historical |
MiniMax M2.5 openrouter/minimax-m2.5 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 350s | Not recorded | 2026-03-14 | historical |
Claude Sonnet 4.6 openrouter/claude-sonnet-4-6 | 100% | 90% | 87.5% | 100% | 4 / 4 | 94.4% | 560s | ~$1.10 | 2026-03-14 | historical |
GPT-5.4 gpt-5.4 | 100% | 90% | 100% | 80% | 4 / 4 | 92.5% | 495s | $0.715 | 2026-03-14 | historical |
Qwen3.6 27B qwen/qwen3.6-27b | 100% | 90% | 87.5% | Unscored | 3 / 4 | 92.5%* | 707.40s | $0.040607 | 2026-08-14 | partial |
GPT-5.6 Terra openai/gpt-5.6-terra | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 102.19s | $0.226598 | 2026-08-14 | audited |
Muse Spark 1.2 meta/muse-spark-1.2 | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 115.53s | $0.247949 | 2026-08-14 | audited |
GLM 5.2 z-ai/glm-5.2 | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 207.88s | $0.135988 | 2026-08-14 | audited |
Gemini 3.7 Flash google/gemini-3.7-flash | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 260.51s | $0.150934 | 2026-08-14 | audited |
Muse Glimmer 30B meta/muse-glimmer-30b | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 271.41s | $0.080912 | 2026-08-14 | audited |
Mistral Large openrouter/mistral-large | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 284s | ~$0.50 | 2026-03-14 | historical |
GPT-4.1 Nano gpt-4.1-nano | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 372s | $0.023 | 2026-03-14 | historical |
Gemini 3.6 Flash google/gemini-3.6-flash | 100% | 90% | 87.5% | 80% | 4 / 4 | 89.4% | 433.51s | $0.318894 | 2026-08-14 | audited |
Llama 4 Scout (Groq) groq/llama-4-scout | 100% | 90% | 87.5% | 73% | 4 / 4 | 87.6% | 210s | $0.009 | 2026-03-14 | historical |
Llama 3.1 8B (Groq) groq/llama-3.1-8b | 100% | 90% | 87.5% | 73% | 4 / 4 | 87.6% | 226s | Free tier | 2026-03-14 | historical |
Llama 3.3 70B (Groq) groq/llama-3.3-70b | 100% | 90% | 87.5% | 73% | 4 / 4 | 87.6% | 289s | $0.022 | 2026-03-14 | historical |
Qwen3 32B (Groq) groq/qwen3-32b | 100% | 90% | 87.5% | 73% flaky | 3 / 4 stable | 87.6%* | 315s | Free tier | 2026-03-14 | partial |
GPT-4.1 gpt-4.1 | 100% | 90% | 100% | 60% | 4 / 4 | 87.5% | 348s | $0.458 | 2026-03-14 | historical |
GPT-5.3 Codex gpt-5.3-codex | 100% | 90% | 100% | 60% | 4 / 4 | 87.5% | 462s | $0.721 | 2026-03-14 | historical |
GPT-5 Mini gpt-5-mini | 100% | 90% | 100% | 60% | 4 / 4 | 87.5% | 792s | $0.106 | 2026-03-14 | historical |
Claude Sonnet 5 anthropic/claude-sonnet-5 | 100% | 90% | 87.5% | 60% | 4 / 4 | 84.4% | 398.53s | $0.784292 | 2026-08-14 | audited |
Llama 4 Maverick (Groq) groq/llama-4-maverick | 100% | 90% | 87.5% | 60% | 4 / 4 | 84.4% | 225s | Free tier | 2026-03-14 | historical |
GPT OSS 120B (Groq) groq/gpt-oss-120b | 100% | 90% | 87.5% | 60% | 4 / 4 | 84.4% | 358s | $0.013 | 2026-03-14 | historical |
GPT OSS 20B (Groq) groq/gpt-oss-20b | 100% | 90% | 87.5% | 60% | 4 / 4 | 84.4% | 407s | $0.014 | 2026-03-14 | historical |
DeepSeek V4 Flash deepseek/deepseek-v4-flash-0731 | 100% | Unscored | 87.5% | 60% | 3 / 4 | 82.5%* | 237.53s | $0.003618 | 2026-08-14 | partial |
* Quality is averaged across valid cases. Unscored or flaky cases remain labeled and are not converted to zero.
Scoring contract
Measured, not inferred
- 01 / identityThe requested OpenRouter route must appear in runtime audit evidence.
- 02 / executionAt least one audited LLM call must have handled the extraction.
- 03 / qualityGround-truth values, rows, and columns are checked per case.
- 04 / failuresFallbacks and failed cases remain visible but never receive a score.
Benchmark your own documents
Use the same evidence contract on a representative sample from your workflow.