Claude Code Guide for GTM Teams (July 2026). Learn more

Benchmarks

We measure frontier models on GTM work.

Models qualify your accounts, research your prospects, and write your outbound. Almost nobody measures whether they're right. We do.

"The model feels good enough" is not a quality bar.

Most GTM teams running AI can't tell you whether it works. Someone read a few outputs at launch, they looked good, and it went live. But GTM failures are silent - a mis-qualified account or an invented research fact doesn't throw an error. It quietly costs you pipeline.

Our first shared eval put numbers on it. The best model agreed with human judgment 83.5% of the time, and the most accurate judge on the board disagreed with itself on 14% of companies from one run to the next. That is the machinery teams are pointing at their pipeline unmeasured.

More on why every GTM AI workflow needs a QA function: Evals for GTM Teams.

The evals.

Published

ICP judgment

Can a model look at a company and judge whether it fits? We ran a full model panel against 243 funded AI companies on byte-identical frozen evidence, and claimed accuracy only on the 133 rows a human judged.

Rank Model
1 Anthropic Claude Fable 5 83.5% $128.90
2 Google Gemini 3.5 Flash 75.2% $22.17
3 Moonshot Kimi K3 open 73.5% $45.14
4 OpenAI GPT-5.6 Sol 72.9% $35.15
5 Anthropic Claude Sonnet 5 71.4% $38.38
6 Anthropic Claude Opus 4.8 69.9% $76.34
7 Google Gemini 3.1 Pro 68.4% $27.02
8 Z.ai GLM-5.2 open 65.4% $12.88
9 OpenAI GPT-5.6 Luna 60.2% $7.43
10 Anthropic Claude Haiku 4.5 58.6% $17.95
11 OpenAI GPT-5.6 Terra 57.1% $17.51

Click a column to sort. Accuracy is against the 133 human-judged rows, cost is per 1,000 accounts. One task at each provider's default reasoning setting - a snapshot, not a full benchmark.

Read the write-up →
In development

Account research

The same rigor applied to the research step: how accurately do models profile an account - firmographics, stack, buying signals - and where do they confidently invent things?

Planned

Agentic search

Models with live web access through Exa and Parallel, measured against static-context baselines. Does search actually improve GTM judgment, and what does it cost?

Why us.

Practitioners

We build and run GTM systems on these models every day. The evals come out of delivery work, not hypotheticals.

Human-judged

Accuracy is scored against a human-judged set, never models grading each other. Methodology and test construction published in full.

Re-run

Models change every few months. A benchmark from two generations ago tells you nothing about what to build on today.

Test before you build on it.

Choosing which model runs your enrichment, research, or qualification workflows - or judging a tool built on one? We run your actual use case through the candidates and hand you the data before you commit.

Book a Call