Updated daily · Arena + Artificial Analysis
High-Volume Enrichment & Extraction
Clay-style enrichment at thousands of rows, lead scoring, intent classification, schema extraction. Cost and speed dominate; frontier IQ is wasted here.
One board from The GTM Model Leaderboard
High-Volume Enrichment & Extraction Leaderboard
Clay-style enrichment at thousands of rows, lead scoring, intent classification, schema extraction. Cost and speed dominate; frontier IQ is wasted here.
- 1Mercury 2InceptionFastest output measured795 tok/sWhy it wins
- 2Gemini 2.5 Flash-LiteGoogleLowest time to first token0.35s latencyWhy it wins
- 3Nova MicroAmazonCheapest blended price measured$0.03 / MWhy it wins
- 4Gemini 3.5 Flash-LiteGoogleSecond-fastest output, cheap tier417 tok/sWhy it wins
- 5Claude Haiku 4.5AnthropicBest cheap option in Claude stacksbalancedWhy it wins
Operator take · Ranked by fit, not a single benchmark. At 10,000 rows in Clay, a frontier model is a budget mistake: route volume work to a fast/cheap tier and reserve frontier models for the 5% of rows that matter.
Updated daily from the linked public leaderboards; last capture Jul 27, 2026. Elo and win-rate figures are preference-based measures, not task-completion guarantees. Always validate the top pick on your own representative work before routing production volume to it.
How to Read These Rankings
Four rules before you switch models
- Every ranking is task-specific. The #1 agent model is not the #1 writer, and neither is the right pick for 10,000 Clay rows.
- Scores come from public leaderboards (Arena, Artificial Analysis) with capture dates shown. Preference Elo measures which output people like, not whether work gets completed.
- Reasoning effort and harness matter: the same model at a different effort tier or in a different agent harness ranks differently.
- Cheap plus fast beats frontier for volume work. Route by job, not by brand loyalty.
We route these models inside production GTM systems every day. If you want help picking and wiring the right models into your own stack, that's what we do.