Benchmarks / Reconciliation
Line-haul Invoice Reconciliation
Read the write-up →Accuracy
99.2
best agent: 94.7
Score variance
0.00
std across reruns · agents 0.6–8.6
Verdict stability
96.7%
agents: 48.6–82.9%
Latency p90
80s
agents: 104–224s
Cadel vs a coding agent
| Contender | Accuracy | Score variance | Verdict stability | Robustness | Cost / invoice | Latency p90 |
|---|---|---|---|---|---|---|
| ★Cadel workflow | 99.2 | 0.00 | 96.7% | 95.8% | 1.00× | 80s |
| Codex · GPT-5.6 Sol ^ | 94.7 | 0.90 | 82.9% | — | 2.50× | 123s |
| Claude Code · Opus 5 | 94.1 | 0.56 | 77.1% | — | 4.86× † | 229s † |
| Codex · GPT-5.5 | 93.9 | 0.62 | 77.1% | — | 2.90× | 149s |
| Codex · GPT-5.6 Luna ^ | 93.2 | 2.03 | 68.6% | — | 0.57× | 104s |
| Claude Code · Opus | 93.1 | 0.97 | 71.4% | — | 2.94× | 155s |
| Claude Code · Sonnet 5 | 92.9 | 1.70 | 65.7% | — | 3.51× † | 243s † |
| Claude Code · Sonnet | 87.8 | 8.61 | 48.6% | — | 2.01× | 224s |
^ indicative: run through a bridge tool profile · † measured from the committed board .eval log
Does the extraction model matter here?
| Extraction model | Cases captured | Accuracy | Verdict stability (like-for-like) | Verdict stability (all groups) |
|---|---|---|---|---|
| ★gemini-2.5-flash | 60/60 | 99.4 | 96.7% | 96.7% |
| ★gemini-2.5-proSHIPPED | 59/60 | 98.4 | 89.7% | 86.7% |
| ★gpt-5.6-luna | 60/60 | 97.2 | 93.3% | 93.3% |
Where the score comes from
Extraction
Did it pull the right values off the page?
Shipped rank
7 / 64
Cadel headers
85.2%
Landscape recall
2 of 7 models
Cost gap
2.1×
| Front-end | Model | Balanced | Headers | Recall · landscape doc | Cost | Time |
|---|---|---|---|---|---|---|
| paddleocr | claude-haiku-4-5 | 0.948 | 0.907 | — | $0.190 | 970s |
| PRODUCT | gemini-2.5-flash | 0.935 | 0.870 | 100% | $0.188 | 313s |
| ★PRODUCTSHIPPED TODAY | gemini-2.5-pro | 0.926 | 0.852 | 100% | $0.386 | 246s |
| PRODUCT | claude-sonnet-4-6 | 0.891 | 0.815 | 93.0% | $0.518 | 385s |
| PRODUCT | gemini-2.5-flash-lite | 0.785 | 0.759 | 45.4% | $0.013 | 132s |
| PRODUCT | gpt-5.4-mini | 0.597 | 0.611 | 9.7% | $0.055 | 81s |
| PRODUCT | claude-haiku-4-5 | 0.521 | 0.537 | 2.2% | $0.266 | 214s |
| PRODUCT | gpt-5.4-nano | 0.414 | 0.333 | 1.1% | $0.028 | 140s |