Can Frontier Models Spot AI-Washing? We Tested 12 of Them
Can frontier models tell which AI companies actually ship AI? This is an example of a GTM eval, and why GTM teams need to regularly QA their AI. We tested 12 models - 9 frontier (Anthropic, OpenAI, Google) plus 3 open-weight (Kimi and GLM) - judging the same 243 funded, AI-branded companies on byte-identical frozen evidence.
- Fable 5 tops the board at 83.5% against human judgment.
- Sol is the deployment winner: the best-calibrated confidence on the board, enough to auto-process 85.7% of the list at 95%+ accuracy with no human review.
- A $22-per-thousand-accounts model statistically tied five models costing up to 3.4× more.
- Our most accurate judge disagrees with itself on 14% of companies run-to-run.
- We captured how significantly improved Kimi K3 is.
Why this eval exists
GTM stacks now route real decisions through model judgment - which of 10,000 accounts gets qualified, which companies enter a sequence, what an enrichment field claims - and almost none of those calls have an eval behind them. Models get picked on brand, price, and vibes, because public benchmarks measure math, code, and agentic capability, not judgment over marketing pages written to be believed. That gap is what this eval measures.
The setup
The task: given a frozen snapshot of a company’s website - homepage, product, careers, docs, engineering pages - classify how real its AI is on a 4-tier rubric: Shipped, Building, Wrapping, Washing.
Frozen means frozen. Every model sees identical bytes. No live browsing, no “one model got a better search result.” This is the part most public evals skip.
The dataset: 243 funded B2B software companies that claim AI on their homepage. It started larger and was trimmed in two disclosed QA passes - dead and mis-scraped evidence packs first, label hygiene second.
The panel: 9 frontier models - Fable 5, Opus 4.8, Sonnet 5, Haiku 4.5 (Anthropic); the GPT-5.6 line of Sol, Terra, and Luna (OpenAI); Gemini 3.1 Pro and 3.5 Flash (Google) - plus 2 open-weight models: GLM 5.2 (Zhipu) and Kimi K3 (Moonshot), which released two days before we swept it. The twelfth checkpoint is Kimi K2.6, K3’s predecessor: superseded, so it votes in no panel statistic and appears only in the generation-gap finding below. Twelve models swept, an eleven-model analysis panel.
The labels: four of our team labeled every company. Then the models exposed problems with those labels - the most useful mistake in the whole project, covered near the end. Final provenance: 133 rows judged by a human, 110 rows adjudicated by the model panel. Accuracy is only ever claimed on the 133 human-judged rows. The 110 panel-adjudicated rows are never scored for accuracy - that would be circular - but they still count for everything label-free.
No anchor model. No single model is the referee. Accuracy is scored against human labels; model-vs-model structure is scored against the whole panel. The reason for that rule is a finding of its own, below.
The price reflex failed a direct test
“Surely the expensive model is better.” We tested that sentence directly. It lost.
| Model | Accuracy vs human labels (n=133) | Cost per 1,000 accounts |
|---|---|---|
| Fable 5 | 83.5% | $128.90 |
| Gemini 3.5 Flash | 75.2% | $22.17 |
| Kimi K3 (open) | 73.5% | $45.14 |
| Sol | 72.9% | $35.15 |
| Sonnet 5 | 71.4% | $38.38 |
| Opus 4.8 | 69.9% | $76.34 |
| Gemini 3.1 Pro | 68.4% | $27.02 |
| GLM 5.2 (open) | 65.4% | $12.88 |
| Luna | 60.2% | $7.43 |
| Haiku 4.5 | 58.6% | $17.95 |
| Terra | 57.1% | $17.51 |
Fable 5 leads the board outright. Below it, six models are one statistical bunch - Flash, Kimi K3, Sol, Sonnet, Opus, Gemini Pro - with McNemar p = 0.14-1.00 on every pair. Flash is the cheapest of the six that tied, at $22.17 per 1,000 accounts. Paying up to 3.4× more bought nothing measurable.
Cheap is not the axis either. Flash decisively beats the models in its own price band (vs Luna p=0.007, vs Terra p=0.001), and the $13 GLM lands bottom-half. And even Fable’s lead over Flash and Sol is p = 0.05-0.07 - a lead, not a rout.
Then the falsification test. We gave the expensive models their maximum reasoning setting and re-ran, expecting the cheap model to fall back:
- Sol got worse. 72.9% to 70.5% (-2.4pp, p=0.55 - noise), while spending 3× the tokens and 1.7× the cost. At max effort its thinking spend roughly matches Fable’s, so even token-matched to the most accurate model it does not converge. The gap is a judgment prior, not a compute deficit.
- Terra genuinely improved: +9.6pp (p=0.019), the only significant config effect in the whole matrix. It still landed below Flash.
- Haiku moved +5.0pp (p=0.36, not significant) while shuffling 17% of its own verdicts. Flash itself moved +2.1pp (p=0.61) - already saturated at its default.
There is no universal “think harder” button. It is a per-model knob - one model up 10 points, two unmoved, one worse at 1.7× the price - and the only way to know which kind you have is to eval it.
Pick a failure direction, not a brand
The economics first: a human researcher does this job at roughly $833 per 1,000 accounts ($20/hr, 2.5 minutes each). The models did it for $7 to $129 - 6.5× to 112× cheaper, at seconds per account instead of minutes. The economics are a rout. They are also the least interesting column.
The interesting part is that models fail in different directions on the binary GTM call (qualified = Shipped or Building):
- GLM 5.2: 96.4% precision, 60.0% recall. It almost never cries wolf - and misses 40% of the real companies. At $0.063 per qualified find, it is the cheapest credible filter on the board. Haiku is harsher still: 95.7% precision, 48.9% recall. It misses half.
- Opus and Sonnet: 97.8% recall each. They find everything, and roughly 1 in 4 of their alarms is false.
Same evidence, opposite failure modes. Which one is “better” depends entirely on whether a false positive (poisoned outreach, wasted sequences) or a false negative (a real company you never contact) costs you more.
Then the question no pricing page answers: when the model says it is confident, is it right?
- Sol has the best calibration on the board (expected calibration error 0.049). Rank its calls by its own confidence and you could auto-process 85.7% of the list at 95%+ accuracy with no human review. Next best: Gemini Pro at 70.7%, Kimi K3 at 69.7%, Fable at 69.2%.
- Sonnet matches Opus-level accuracy with confidence you cannot route on: 3.0% auto. Effectively every row still needs a human.
- Flash is the best judge after Fable and a mediocre router: high flat confidence caps it at 42.1% auto.
The model with the best probabilities is not the most accurate one. And confidence is not a cross-model unit: Anthropic and open-weight models drop their confidence when they dissent from the panel majority, while OpenAI and Google models stay at 0.78+ even as the outlier. A 0.85 from Sol is routine; a 0.85 from Fable is near its ceiling. A threshold tuned on one lab breaks on another.
Why the $22 model holds up
Flash’s row invites the most suspicion, so it is the one we audited hardest. It holds, and the mechanism is measurable.
Exact-tier scoring rewards boundary calibration, not raw capability. Flash’s tier boundaries sit dead center of the 11-model panel (mean deviation -0.03). The pricier flagships miss by miscalibrating in opposite directions: Opus grades +0.40 tiers generous - the most generous judge tested, with 51 Shipped calls to Fable’s 38 - while Sol rounds down, calling Washing on 173 of 243 companies to Fable’s 134.
Flash’s price is budget-tier; its deliberation is not. 1,273 median reasoning tokens per call - 2.2× Fable, 7× Sol. “Cheap = shallow” is exactly the mapping this task falsifies. The only other model that combines dead-center boundaries with frontier-level deliberation, Kimi K3, is also the only other model in the top three. Neither ingredient works alone: Terra and Luna sit center but barely think; Haiku and GLM think hardest of anyone and grade from the harsh edge. Across all 11 models, the rank correlation between reasoning spend and accuracy is +0.06. Nothing.
And the same flip happened inside both flagship lines, where “Google just calibrates differently” can’t explain it:
- Sonnet outscored Opus, 71.4% to 69.9%. Not because it is smarter - the gap is exactly 2 companies at n=133, p=0.86 - but because Opus grades generous on a market whose final labels are ~59% Washing. 37 of Opus’s 40 misses on the human-judged slice landed above the human tier.
- Flash outscored Gemini Pro, 75.2% to 68.4% - and Pro failed in the opposite direction. 38 of its 42 misses landed below the human tier. Pro called Washing on 75 of the 133 human-judged rows; the humans said 45.
Both labs’ flagships lost to their own cheaper model, and they failed in opposite directions. That pair of flips is also the labels’ best defense: labels tilted harsh would have crowned Pro, labels tilted generous would have crowned Opus. Neither happened.
Price buys tokens. This task pays for calibration, and calibration does not follow the price list.
The judge disagrees with itself
This is the smartest result in the project.
Before trusting any AI-judge leaderboard, run the judge twice. We ran Fable 5 - the top scorer on the board - three times over the same 50 frozen packs. Same bytes, same rubric, same model. It flipped its own tier on 7 of 50 companies: 14%. Two runs of the same judge agree only 88-92%.
Why: thinking-enabled models sample a different reasoning path on every call. On genuinely borderline companies, different path, different answer.
The elegant part: it warns you. All 7 flips came from calls the model itself scored 0.50-0.55 confidence, the bottom of its range (stable calls spanned 0.45-0.95). Flipped companies also showed 4.6× the run-to-run spread in thinking length. The judge knows when it is torn.
The design consequence: this eval has no anchor model. Score ten rivals by “agreement with model X” and every number inherits X’s own 8-12% self-noise as a hard ceiling - and the fine ordering underneath is sampling variance wearing a ranking’s clothes.
There is one more way to see the same trap. Rank the models label-free - each against the plurality of the other ten - and the board flips: Gemini Pro jumps to 1st, Fable falls to 6th, Opus lands dead last. That board measures centrality, not accuracy. The confident conformists top it; the generous dissenters sink. Most published “LLM judge agreement” numbers are centrality wearing accuracy’s clothes. Every leaderboard is a choice of referee, and you should publish which one you chose.
Two clusters, split by temperament - not by lab
The honest output of this eval is two clusters, not a leaderboard. A top bunch - Flash, K3, Sol, Sonnet, Opus, Gemini Pro at 68.4-75.2%, with Fable above at 83.5% - and a bottom cluster genuinely apart: Luna, Haiku, Terra at 57.1-60.2%. GLM sits between.
The same split appears with zero labels, and the axis is not the one you would guess:
- Across all 55 model pairs, the five tightest agreements are all cross-lab: Gemini-GLM (κ 0.84), Gemini-Sol (0.81), Sol-GLM (0.79), Fable-Flash (0.77), Gemini-K3 (0.77).
- The loosest pair in the whole matrix is same-lab: Opus-Haiku, κ 0.39.
- Mean within-lab agreement is κ 0.64 against 0.60 cross-lab. Lab membership tells you almost nothing about how a model will judge. That matters here for a second reason: four of our eleven panelists are Anthropic models, so if lineage drove verdicts, every panel statistic in this post would be one lab voting as a bloc. Measured, it is not.
What actually predicts agreement is where a model draws the Washing boundary: Opus +0.40 tiers above the panel mean, Sonnet +0.25, Fable +0.14 on the generous side; Sol -0.11, Gemini Pro -0.17, GLM -0.22, Haiku -0.23 on the harsh side. And it is boundary definitions, not comprehension - every model lands within one tier of the panel verdict 87-98% of the time. Nobody misreads the pages. They disagree about where “real” starts.
You can watch the definitions collide in the rationales. On the rows where Sol and Fable disagree, Sol runs a nominalist checklist - “names no model” appears in 77% of its Washing rationales; a live AI feature is Washing until a provider is named. Fable judges by function: a mechanism identifiable in what the product does counts, named or not. Same cited facts, different definition of “identifiable.” The moment a company names its stack, they converge.
One version bump moved more than any lab gap
Moonshot released Kimi K3 on a Wednesday. We ran it on the frozen packs that Friday, next to its predecessor K2.6. Same company, same API, same config, one generation apart. K3 did not just improve - it un-learned its predecessor’s personality:
- Accuracy vs human labels: 56.8% (would have ranked last) to 73.5% (#3 on the board). Paired p=0.001 - the biggest single move anywhere in the data.
- Calibration: -0.33 vs the panel - the harshest model tested - to -0.02, dead center.
- K2.6 was a literal zero-false-positive gate: 100% precision at 44.4% recall. K3 traded it for balance: 85.0 / 75.6.
- Median output tokens: 3,046 to 1,421. Better judgment on half the thinking.
The stat that should worry anyone running open models in production: K3 agrees with its own predecessor at κ 0.56 - lower than its agreement with Gemini (0.77), Flash (0.76), or Fable (0.75). The “open-weight models share a skeptical temperament” cluster dissolved in one release. Temperament is a training-time property, and it can move a full cluster-width per generation.
Operationally: K3 qualifies 50 companies where K2.6 qualified 30. If this model line sat in your qualification pipeline, a version upgrade silently inverts its judgment on a fifth of your TAM. Pin your model versions. Re-eval on upgrade.
The category lesson: the open-vs-closed gap at the top of the board is 1.7 points (Flash 75.2 vs K3 73.5 - a tie). The spread across open-weight checkpoints is 16.7 points, roughly ten times wider. The license tells you nothing; the checkpoint tells you everything. And via API, open no longer means cheap: K3’s mid-frontier rates land it at $45.14 per 1,000 - twice Flash, which ties it (prices pending billing-page confirmation). The real open-weight economics case moves to self-hosting once the weights publish.
The token meter measures the account, not the model
Reasoning-token counts invite one wrong read and hide one useful one. We measured both directions.
Across models, spend is temperament, not quality. Default spend on identical packs runs 158 tokens (Terra) to 1,974 (Haiku) - a 12.5× spread - and its correlation with accuracy is +0.06. Every quadrant is occupied: frugal and right (Sol), frugal and wrong (Terra, Luna), deliberate and right (Flash, K3), deliberate and wrong (Haiku, GLM). The meter’s resting position is a vendor training choice - the budget tier substitutes test-time compute for parameter count, and flagships are tuned frugal because their tokens cost 10× and latency sells. Reading it as intelligence is the price reflex in a different costume.
Within one model, the meter is a difficulty detector - and it fired for 11 of 11. Every model thinks longer on exactly the companies the panel ends up fighting over: 1.4-3.2× its own unanimous-row median, with a pooled correlation between panel effort and verdict spread of ρ = 0.79 - the strongest correlation anywhere in this analysis. It is not pack length (ρ 0.33), and it beats self-reported confidence as a difficulty signal for 8 of 11 models. Confidence is a self-report with lab-specific manners; token spend is behavior.
The deployment move is free: token counts arrive on every API response. Log them per account and route the high-effort tail to a human. On this dataset, the five highest-effort companies all landed 2-3 tiers apart across the panel, and the effort floor is unanimous Washing. (For the cost side of reasoning tokens, we covered the billing math in GPT-5: reasoning tokens change the math.)
Latency per model
The sweep logged one more column per call: wall-clock latency. Ranked by median seconds per verdict, the board runs 2.7 seconds (Terra) to 43.9 (Kimi K3) - a 16× spread - and every model sits far inside the ~150 seconds a human researcher needs for the same account.
The ranking is mostly the token board read in wall-clock: how much a model reasons × how fast its provider serves it. The OpenAI trio is fastest because it thinks least. The two exceptions are the interesting rows:
- Flash breaks the “fast = shallow” mapping a second time: 1,273 reasoning tokens served in 6.8 seconds. Speed is not evidence of a small model, any more than price was evidence of a good one.
- The hosted open-weight tax shows up on the clock too. Kimi K3 thinks like Flash - 1,421 tokens to Flash’s 1,273 - and takes 6.5× the wall clock, with the longest tail on the board: p90 at 111 seconds, worst call 430. GLM is the same story, milder.
Operationally, the median is a throughput question, and throughput is a concurrency choice: at 8 concurrent calls the slowest judge cleared all 243 companies in ~35 minutes - a 10-hour day of human work. The column that actually disqualifies a model is the tail. If a verdict gates anything near-real-time - an inbound router, enrichment on form-fill - p90 is the number to eval, and it varies 22× across models that all “take seconds.” Honesty note: these are operational numbers from a single July sweep under live provider load, not a controlled latency benchmark, so treat the ranking as coarse.
The census: AI-washing is real, and even frontier models can’t consistently see through it
Everything above is about the judges. The other output of the eval is what they said about the market. Show 11 models the same 243 companies and ask which ones have real AI:
- The “this company’s AI is real” count ranged from 32 companies (Haiku) to 82 (Opus). Same pages, 2.6× spread.
- 113 companies - 46.5% - got a unanimous verdict from all 11 models. 96 of those unanimous verdicts are Washing. The panel was unanimous that a company actually shipped real AI exactly 4 times in 243. Near-consensus flattery is rare; near-consensus fraud-calling is not.
- On 48 companies - nearly 1 in 5 - at least one model said Shipped while another said Washing on the same evidence. The Wrapping-Washing boundary is the most contested line in the rubric.
- Score the market by strict panel plurality instead of unanimity and 164 of 243 - 67.5% - come out Washing.
If you build TAM lists from “AI company” filters, somewhere between a third and two-thirds of that list is marketing.
The biggest bug in the eval was our own labels
Our models agreed with each other and not with our humans. That is usually not the models.
Original model-vs-human agreement ran 25-41%, and the multi-model majority vote scored below every individual model - the statistical signature of noisy labels, not dumb models. On spot-check, the models were reading the rubric correctly and the humans had inflated washing companies upward. The humans used the “Building” tier 3 times in 243; the models used it 18-31 times.
The fix, and the honest cost of it:
- 81 contested rows we re-adjudicated by hand against the evidence - several against the panel; the last batch went 0 for 7 toward the top-scoring model.
- 110 rows where the QA-time panel overwhelmingly contradicted the original label - 85 unanimously, 25 at 8 of 9 - were relabeled to the panel majority. Machine-adjudicated rows are never scored for accuracy. They serve only the label-free analysis.
- 18 rows were deleted for label hygiene: their tiers were entered in a rushed format that flagged low-care labeling.
That leaves the split every number in this post respects: 133 human-judged rows where accuracy claims live, 110 panel-adjudicated rows that never touch an accuracy claim, 243 total for everything label-free.
The caveats are cheap and they buy trust: our adjudication was not blind (model votes were visible), each row has one human label so there is no inter-rater number yet, and the human slice skews more balanced than the full set because contested rows were likelier to get human attention. A blind double-label pass is the named next step.
The transferable lesson: if your models agree with each other but not with your ground truth, and majority-voting makes things worse - audit your labels before you audit your models.
Method notes, compressed
- 12 models were swept; the analysis panel is 11. K2.6 is superseded by K3 and appears only in the generation-gap comparison - keeping both would double-weight one vendor’s lineage in every consensus vote.
- Reasoning ran at each provider’s standard setting, then a max-setting pass bracketed the reachable range. You cannot force-match thinking budgets across labs.
- One real bug, disclosed: the first sweep ran the OpenAI models at ~zero reasoning and Haiku with thinking off. Re-running moved Haiku +14 points (artifact) and barely moved the OpenAI trio (genuine judgment).
- Evidence completeness dominates everything. One company’s engineering page was missing from its frozen pack: 9 of 11 models called it Washing. On the fixed pack, every model calls it qualified. And in a one-company probe on the most contested row, adding the founders’ bio page flipped the frontier models up to Shipped - citing ex-DeepMind pedigree the rubric explicitly says to ignore, while the cheap and open models held the line. One page moves verdicts harder than any model choice or thinking budget.
- The open-weight hosts ignore server-side JSON schemas, so structured output held by prompt instruction plus lenient parsing. Moonshot’s content filter refused the same 2 companies on both Kimi generations; GLM ran clean.
What’s next
v1 is deliberately closed-book: frozen evidence, no tools. v2 hands the same rubric to research agents - Exa, Parallel, Claygent-style - and lets them gather the evidence live, which tests the thing the frozen packs simulate: whether working longer actually reaches better evidence. Same harness, new axis.
And the standing play continues: every notable model release gets rerun on the frozen packs the same week. Kimi K3 got its row two days after launch. The next model gets one too.
If you route GTM decisions through a model
- Brand barely predicts judgment, and price predicted it inversely here. The only ranking that matters is on your task.
- “Turn up reasoning” is a per-model knob. +9.6 points for one model, nothing for two, worse at 1.7× the cost for one.
- Calibration decides automation, not accuracy. The model you can automate behind is the one whose confidence you can route on - and thresholds do not transfer across labs.
- Pin model versions. One generation flipped an open-weight model from harshest judge to dead center. The temperament you evaled is not stable across upgrades.
- Log reasoning tokens per account. A free difficulty detector that flags exactly the rows humans end up fighting over.
- Fix evidence before models. One missing page moved verdicts more than any model swap or thinking budget.
- Disclose label provenance. Claim accuracy only on rows a human actually judged, and say which rows those are.
If your qualification pipeline runs on a model nobody has ever scored, that is fixable in a week - the harness already exists.









