Benchmarks
We measure frontier models on GTM work.
Models qualify your accounts, research your prospects, and write your outbound. Almost nobody measures whether they're right. We do.
"The model feels good enough" is not a quality bar.
Most GTM teams running AI can't tell you whether it works. Someone read a few outputs at launch, they looked good, and it went live. But GTM failures are silent - a mis-qualified account or an invented research fact doesn't throw an error. It quietly costs you pipeline.
Our first shared eval put numbers on it. The best model agreed with human judgment 83.5% of the time, and the most accurate judge on the board disagreed with itself on 14% of companies from one run to the next. That is the machinery teams are pointing at their pipeline unmeasured.
More on why every GTM AI workflow needs a QA function: Evals for GTM Teams.
The evals.
ICP judgment
Can a model look at a company and judge whether it fits? We ran a full model panel against 243 funded AI companies on byte-identical frozen evidence, and claimed accuracy only on the 133 rows a human judged.
| Rank | Model | ||
|---|---|---|---|
| 1 | | 83.5% | $128.90 |
| 2 | | 75.2% | $22.17 |
| 3 | | 73.5% | $45.14 |
| 4 | | 72.9% | $35.15 |
| 5 | | 71.4% | $38.38 |
| 6 | | 69.9% | $76.34 |
| 7 | | 68.4% | $27.02 |
| 8 | | 65.4% | $12.88 |
| 9 | | 60.2% | $7.43 |
| 10 | | 58.6% | $17.95 |
| 11 | | 57.1% | $17.51 |
Click a column to sort. Accuracy is against the 133 human-judged rows, cost is per 1,000 accounts. One task at each provider's default reasoning setting - a snapshot, not a full benchmark.
Account research
The same rigor applied to the research step: how accurately do models profile an account - firmographics, stack, buying signals - and where do they confidently invent things?
Agentic search
Models with live web access through Exa and Parallel, measured against static-context baselines. Does search actually improve GTM judgment, and what does it cost?
Why us.
We build and run GTM systems on these models every day. The evals come out of delivery work, not hypotheticals.
Accuracy is scored against a human-judged set, never models grading each other. Methodology and test construction published in full.
Models change every few months. A benchmark from two generations ago tells you nothing about what to build on today.
Test before you build on it.
Choosing which model runs your enrichment, research, or qualification workflows - or judging a tool built on one? We run your actual use case through the candidates and hand you the data before you commit.
Book a Call