← Back to the conversation

Five models, one corpus

A companion to what this assistant scores.

This site can run five models behind one chat endpoint. This page is what happened when all five answered the same 22 questions, in the same order, against the same server, on the same afternoon — graded by the same checker, and costed at list price.

The point is not to crown a winner. It is that a model choice is a trade between four things that pull against each other — how often the answer is right, how long you wait for it, how much of it you get, and what it costs — and that trade is only legible once someone measures all four on the same corpus.

Measured 27 Jul 2026. Published 27 Jul 2026, built into this page 28 Jul 2026. It is a dated artifact, not a live dashboard — it was run once and it goes stale.

turns graded
110
5
models
22
cases each
the same corpus
4
eval groups
6
broke mid-stream
never produced an answer
$0.38
cost of the whole comparison
list price, every turn

Quality against cost

Pass rate against median cost per turn
Median list-price cost of one graded turn, against the share of answered cases that passed. Better is up; cheaper is left. A turn that broke mid-stream is not counted as a wrong answer — reliability is its own panel below. The vertical axis is zoomed to the measured range, not to zero.
$0.0010$0.003090%100%median cost per turn (log)pass rate (answered)anthropic/claude-haiku-4.5 — 100% on 22 cases, $0.0025 · on the frontierclaude-haiku-4.5deepseek/deepseek-v4-flash — 93.8% on 22 cases, $0.00095 · on the frontierdeepseek-v4-flashgoogle/gemini-3.5-flash-lite — 95.5% on 22 cases, $0.0022 · on the frontiergemini-3.5-flash-liteopenai/gpt-5-mini — 86.4% on 22 cases, $0.0013gpt-5-miniopenai/gpt-5.6-luna — 90.9% on 22 cases, $0.0012gpt-5.6-luna

Highlighted models sit on the frontier — nothing measured here is both better and cheaper.

Quality against latency

Pass rate against median time to first token
Median time to the model's first token, against the share of answered cases that passed. Better is up; faster is left. The vertical axis is zoomed to the measured range, not to zero.
0 ms2.00 s4.00 s6.00 s90%100%median time to first tokenpass rate (answered)anthropic/claude-haiku-4.5 — 100% on 22 cases, 1.06 s · on the frontierclaude-haiku-4.5deepseek/deepseek-v4-flash — 93.8% on 22 cases, 2.10 sdeepseek-v4-flashgoogle/gemini-3.5-flash-lite — 95.5% on 22 cases, 834 ms · on the frontiergemini-3.5-flash-liteopenai/gpt-5-mini — 86.4% on 22 cases, 6.67 sgpt-5-miniopenai/gpt-5.6-luna — 90.9% on 22 cases, 1.20 sgpt-5.6-luna

Highlighted models sit on the frontier — nothing measured here is both better and faster.

The numbers underneath

One panel per measurement, each sorted best-first. Nothing here is stacked or double-encoded: every panel plots one number, and where there is a winner it is the only bar wearing the accent. One panel has no winner, and nothing in it is highlighted — a longer answer is not a better answer, it is just more of the bill.

Time to first token
Median, server-measured. Lower is better.
gemini-3.5-flash-lite834 msfastest
claude-haiku-4.51.06 s
gpt-5.6-luna1.20 s
deepseek-v4-flash2.10 s
gpt-5-mini6.67 s
Output speed
Median output tokens per second across the turn. Higher is better.
gpt-5-mini72.3 tok/sfastest
gpt-5.6-luna55.0 tok/s
claude-haiku-4.545.0 tok/s
gemini-3.5-flash-lite43.6 tok/s
deepseek-v4-flash32.1 tok/s
Cost per turn
Median, at list price on the date checked. Lower is better.
deepseek-v4-flash$0.00095cheapest
gpt-5.6-luna$0.0012
gpt-5-mini$0.0013
gemini-3.5-flash-lite$0.0022
claude-haiku-4.5$0.0025
Output tokens per turn
Median. Neither direction is better — it is how much answer you get, and what the output side of the bill is. Counted by each provider's own tokenizer, so compare within a row rather than across them.
gpt-5-mini576
deepseek-v4-flash186
claude-haiku-4.5147
gpt-5.6-luna108
gemini-3.5-flash-lite92
Input tokens per turn
Median. The corpus is identical for every model, so most of the spread here is the tokenizers disagreeing about how to count the same text, not one model being sent more — these counts are not comparable across providers the way seconds and dollars are.
gpt-5-mini6.0kfewest
gemini-3.5-flash-lite6.4k
gpt-5.6-luna7.8k
claude-haiku-4.59.7k
deepseek-v4-flash12.2k
Turns that broke
Share of attempted turns that failed in transport or died mid-stream, so the model never produced an answer. Lower is better, and it is deliberately kept out of the pass rate above.
claude-haiku-4.50%
gemini-3.5-flash-lite0%
gpt-5-mini0%
gpt-5.6-luna0%
deepseek-v4-flash27.3%
Prompt cache hit rate
Share of input TOKENS served from the provider's cache — not share of turns. Higher is cheaper.
gpt-5-mini86.1%best
gpt-5.6-luna85.3%
claude-haiku-4.582.0%
deepseek-v4-flash47.6%
gemini-3.5-flash-lite15.6%

The scorecard

Pass rate by model and by what the group is testing. There is deliberately no single intelligence score: a model that answers grounded questions well and folds under prompt injection is a different problem from one that is mediocre at both, and one number cannot tell you which you have.

lowerhigher pass rate
Pass rate by model and eval group. Each cell prints its own value.
modelgroundedinjectionrefusalroleFitall
claude-haiku-4.5100%100%100%100%100%
deepseek-v4-flash100%83%100%100%94%
gemini-3.5-flash-lite83%100%100%100%95%
gpt-5-mini67%100%75%100%86%
gpt-5.6-luna83%83%100%100%91%

grounded — does it state facts that are actually in the corpus, and say 'I don't have that' when they aren't. refusal — does it decline out-of-scope requests instead of helping. injection — role override, prompt extraction, knowledge-base dump, a fake system message, disparagement bait, persona hijack. roleFit — pasted job descriptions, graded on whether the assessment names the gaps rather than flattering the reader.

Everything, as a table

The same figures the charts plot, plus the ones no chart shows. This is the accessibility twin for every chart above — nothing on this page is reachable only by colour or only by hover.

modelpassed / answeredbrokep50 TTFTtok/sinoutcache$ / turn$ runclaims verifiedprice
anthropic/claude-haiku-4.522/2201.06 s45.09.7k14782.0%$0.0025$0.123510/30confirmed
deepseek/deepseek-v4-flash15/166/222.10 s32.112.2k18647.6%$0.00095$0.024010/17confirmed
google/gemini-3.5-flash-lite21/220834 ms43.66.4k9215.6%$0.0022$0.079714/17cache rate derived
openai/gpt-5-mini19/2206.67 s72.36.0k57686.1%$0.0013$0.076614/41cache rate derived
openai/gpt-5.6-luna20/2201.20 s55.07.8k10885.3%$0.0012$0.075218/38unconfirmed

Prices checked 2026-07-27. A median is withheld and shown as an em dash under 10 measured turns, and a 95th percentile is not computed at all — one sample per case cannot support one. “claims verified” is the site's own citation checker run over that model's answers; it counts only sentences with something checkable in them.

The run archive

Every run ever published, per model, with the date it was measured. Until this sprint a second run on a model overwrote the first, which meant a regression was invisible and a figure measured twice looked exactly like one measured once. Runs are now kept; the charts above read the most recent per model.

anthropic/claude-haiku-4.5
27 Jul 2026 22/22shown above
deepseek/deepseek-v4-flash
27 Jul 2026 15/22shown above
google/gemini-3.5-flash-lite
27 Jul 2026 21/22shown above27 Jul 2026 6/6
openai/gpt-5-mini
27 Jul 2026 19/22shown above
openai/gpt-5.6-luna
27 Jul 2026 20/22shown above

What this does not measure

One sample per case
Each model answered each question once. Every timing figure here is a median over about twenty turns, which is enough to rank models and nowhere near enough for a 95th percentile — so none is computed. Re-running the same corpus on the same model would move these numbers, and nothing here says by how much.
Latency includes the day it was measured
Time to first token is the server's own measurement, with the browser network excluded — but not the hop from a home connection to the provider, nor whatever load that provider was under that afternoon. It is a real measurement of a real request; it is not a claim about the model in a datacentre.
The grader is a string matcher, not a judge
Cases pass by containing specific evidence — a name, a number, a refusal. That keeps runs free and reproducible, and it means a wrong answer containing the right number still passes, and a right answer phrased unusually can still fail. This is a regression net, not a quality score, and a pass rate here is not an accuracy claim.
Costs are list prices, not an invoice
Every dollar figure is computed from published per-token rates on the date below and the tokens the run actually used. It is an estimate at list price. Prompt-cache rates are mostly derived from a published discount rather than separately published, and one model's prices could not be confirmed at all — it is marked in the table.
A broken turn is counted, but it is not counted as a wrong answer
Some turns died mid-stream and never produced an answer. Those are excluded from the pass rate — scoring a dropped connection as bad judgment is the confusion this whole site is built to avoid — and reported on their own panel instead. But one afternoon of requests is a thin basis for a reliability claim: it says a model dropped streams here, then, not that it drops one in four everywhere.
Token counts are each provider's own
Every model got byte-identical prompts, so the spread in input tokens is mostly the tokenizers disagreeing about how to count the same text, not one model being sent more work. Seconds and dollars compare across rows; tokens really do not.
The corpus is narrow on purpose
Twenty-two cases about one person's résumé and the ways someone might try to break a chatbot answering questions about it. That is exactly the workload this site runs, which is why the numbers are useful here — and why they do not transfer to yours.