DemandLab

Updated daily · Arena + Artificial Analysis

The GTM Model Leaderboard

Which AI model wins each go-to-market job, right now. Eight leaderboards built from live public benchmark data: agents, outbound copy, research, GTM engineering, documents, strategy, enrichment, and long context.

Quick answer

The GTM Model Leaderboard, published by DemandLab, ranks AI models by the eight jobs go-to-market teams actually hire them for. There is no single best model: Claude Fable 5.1 (Max) leads agentic workflows while Claude Fable 5 leads outbound copy. Scores come from Arena and Artificial Analysis, recaptured daily, last updated September 23, 2026.

Compiled by Chris Arden, Fractional CMO, DemandLab · Updated on

Agentic GTM Workflows Leaderboard

Signal monitoring, enrichment pipelines, multi-step outbound automations, CRM agents. The model plans, calls tools, and completes work end to end.

  1. 1
    Claude Fable 5.1 (Max)
    Anthropic
    13.71%±1.72%
    Confirmed Success
  2. 2
    GPT-6 Astra (Max)
    OpenAI
    11.54%±2.10%
    Confirmed Success
  3. 3
    Claude Opus 5 (High)
    Anthropic
    10.25%±1.41%
    Confirmed Success
  4. 4
    Claude Opus 5 (Max)
    Anthropic
    10.16%±1.55%
    Confirmed Success
  5. 5
    Claude Fable 5 (High)
    Anthropic
    8.81%±1.25%
    Confirmed Success

Operator take · Same top five as last capture, but the scores were rescaled as Arena's vote volume grew. Claude Fable 5.1 Max still leads on Confirmed Success at 13.71%, with GPT-6 Astra Max second at 11.54%. Claude Opus 5 High and Opus 5 Max are now a near tie at 10.25% and 10.16%, so read #3 and #4 as one tier, and Fable 5 High holds fifth at 8.81%. Anthropic holds four of the five slots. The interval on the leader (±1.72%) now overlaps Astra's (±2.10%) less than before but the gap is only about two points, so this is a lead, not a lock. If your agentic GTM workflows run on Opus 5 today, Fable 5.1 Max is still the one worth a head-to-head test on cost and reliability, not just the top-line percentage.

Outbound Copy & Content Leaderboard

Cold email, LinkedIn sequences, landing page copy, newsletters, blog drafts. Voice match and edit distance matter more than raw IQ.

  1. 1
    Claude Fable 5
    AnthropicText arena Overall Elo 1507 ±5; Overall #1, Expert #1
    #1
    Creative Writing rank
  2. 2
    Claude Opus 4.6 (High)
    AnthropicElo 1505 ±4; Text arena Longer Query #1
    #2
    Creative Writing rank
  3. 3
    Gemini 3.8 Flash (High)
    GoogleElo 1496 ±19, swapped with 3.7 Flash for #3
    #3
    Creative Writing rank
  4. 4
    Gemini 3.7 Flash (High)
    GoogleElo 1495 ±18
    #4
    Creative Writing rank
  5. 5
    Claude Opus 4.7 (High)
    AnthropicElo 1488 ±7
    #5
    Creative Writing rank

Operator take · Google still holds two of the top five, but the two Gemini Flash entries traded places: 3.8 Flash High moved up to #3 (1496) and 3.7 Flash High slid to #4 (1495), a one-point gap that is well inside the margin of error, so treat it as a coin flip rather than a real change. The top two did not move: Fable 5 still leads Creative Writing, Text arena Overall, and Expert, which is the only model on this board clear enough to be worth a deliberate test. What is still worth noticing is which Gemini is winning here. Both Google entries in the top five are Flash models, the cheap fast tier, beating the Pro model at the writing job. If you have been paying Pro prices for draft copy on the assumption that bigger is better for prose, that assumption is not holding up. The workflow does not change: draft with a top-three model, then run your own voice pass. Judge the bake-off on edit distance from your own voice, because that is the only score that matters here.

Prospect & Market Research Leaderboard

Account research, funding-signal digests, competitive intel, pre-call briefs. Needs live web search plus defensible citations.

  1. 1
    GPT-5.6 Sol (xHigh)
    OpenAI
    1257±7
    Arena Elo
  2. 2
    Claude Opus 4.6 (Search)
    Anthropic
    1253±5
    Arena Elo
  3. 3
    GPT-5.5 (Search)
    OpenAI
    1242±5
    Arena Elo
  4. 4
    Claude Opus 4.7
    Anthropic
    1233±5
    Arena Elo
  5. 5
    Claude Fable 5
    Anthropic
    1230±8
    Arena Elo

Operator take · The most static board on this page, and it has now gone another three days without a single Elo point moving anywhere in the top five. GPT-5.6 Sol at xHigh still leads at 1257 with Claude Opus 4.6 Search four points back at 1253, which is a tie given intervals of 7 and 5, not a win. OpenAI holds first and third, Anthropic holds second, fourth, and fifth. The stillness is the signal: search-grounded quality is bounded by the index behind the model and by citation discipline, not by raw model IQ, so it does not lurch the way the reasoning boards do. First to fifth spans 27 points. That is not enough to justify moving your research workflow between stacks. Keep using whatever your tool already runs, Claude in Claude Code and Cowork, GPT in ChatGPT, and spend the effort on your source list instead.

GTM Engineering & Landing Pages Leaderboard

Building automations, Clay integrations, webhooks, prospect landing pages, internal tools. The builder work behind every AI GTM system.

  1. 1
    GPT-6 Astra (Max)
    OpenAIHeld #1
    1793±12
    Arena Elo
  2. 2
    Claude Fable 5.1 (Max)
    AnthropicSteady second, gap to #1 narrowed from 42 to 40 points
    1753±11
    Arena Elo
  3. 3
    Claude Opus 5 (Max)
    Anthropic
    1691±7
    Arena Elo
  4. 4
    GPT-6 Sol (Max)
    OpenAINew to the top five
    1689±19
    Arena Elo
  5. 5
    Qwen 3.8 Max
    Alibaba
    1671±12
    Arena Elo

Operator take · GPT-6 Astra still holds the top WebDev slot, 1793 against Claude Fable 5.1 Max's 1753, a 40-point gap that narrowed slightly from 42 last capture. Nothing is moving fast here. Opus 5 Max holds third at 1691, and GPT-6 Sol Max enters at fourth (1689), though its ±19 interval makes it a statistical tie with Opus 5 Max. Qwen 3.8 Max is fifth (1671) and Kimi K3 dropped out of the top five. If you build internal tools or Clay integrations on Claude today, there is no urgency to switch. Fable 5.1 Max is still excellent and stable. The Astra gap is worth a real evaluation only if raw WebDev score is what you are optimizing for.

Decks, Screenshots & Documents Leaderboard

Reading dashboards and campaign screenshots, parsing PDFs and contracts, pulling data out of decks and one-pagers.

  1. 1
    Claude Fable 5 (High)
    AnthropicOutside the Document top 5
    1310±8
    Vision Elo
  2. 2
    Qwen 3.8 (Max)
    AlibabaMoved up from #3 to #2; outside the Document top 5
    1302±8
    Vision Elo
  3. 3
    Claude Opus 4.7 (High)
    AnthropicOutside the Document top 5
    1301±7
    Vision Elo
  4. 4
    Claude Opus 4.7
    AnthropicOutside the Document top 5
    1300±7
    Vision Elo
  5. 5
    Claude Opus 4.6 (High)
    Anthropic#3 Document (1507)
    1299±7
    Vision Elo

Operator take · The Vision and Document boards held from last capture. On Vision, Fable 5 High still leads at 1310, with Qwen 3.8 Max (1302), Opus 4.7 High (1301), Opus 4.7 (1300) and Opus 4.6 High (1299) all within a few points, so the order below #1 is noise. On Document, Claude Opus 5 High leads at 1516, Claude Fable 5.1 Max is second at 1513, Opus 4.6 High and Opus 4.6 tie at 1507, and Fable 5 High is fifth at 1496. Practical read is unchanged: screenshots and charts go to Fable 5, and contracts and multi-page PDFs go to Opus 5 High.

Strategy & Hard Analysis Leaderboard

Pricing decisions, ICP definition, board-level analysis, expert-domain questions. The highest-stakes thinking you delegate.

  1. 1
    Claude Opus 5.5 (Max effort)
    AnthropicAA Intelligence Index 58
    #1
    Standing
  2. 2
    Claude Opus 5.5 (Xhigh effort)
    AnthropicAA Intelligence Index 56
    #2
    Standing
  3. 3
    Claude Opus 5.5 (High effort)
    AnthropicAA Intelligence Index 54
    #3
    Standing
  4. 4
    Claude Fable 5.1 (Max effort)
    AnthropicAA Intelligence Index 53
    #4
    Standing
  5. 5
    Claude Fable 5.1 (Xhigh effort)
    AnthropicAA Intelligence Index 53
    #5
    Standing

Operator take · New #1. Claude Opus 5.5 now sits at the top of the Artificial Analysis Intelligence Index, and it takes the first three slots across effort levels: Max at 58, Xhigh at 56, High at 54. Claude Fable 5.1 follows at 53 for both Max and Xhigh. GPT-6 Astra, which tied Fable 5.1 at 53 last capture, no longer shows in the top five on the card. The card did not surface cost per task today, so last capture's pricing figures are not repeated here. The practical read: the effort setting matters as much as the model, since Opus 5.5 High (54) beats Fable 5.1 Max (53). For board-level calls where a wrong answer costs more than API spend, test Opus 5.5 at max effort against whatever you run now, and check the per-task cost before you commit.

High-Volume Enrichment & Extraction Leaderboard

Clay-style enrichment at thousands of rows, lead scoring, intent classification, schema extraction. Cost and speed dominate; frontier IQ is wasted here.

  1. 1
    Celeris-1
    CelerisFastest output measured, down from 1,541 tok/s
    1,506 tok/s
    Why it wins
  2. 2
    Mercury 2.5
    InceptionSecond on throughput; Mercury 2 now third at 750 tok/s
    770 tok/s
    Why it wins
  3. 3
    Gemini 2.5 Flash-Lite
    GoogleLowest time to first token measured; North Mini Code second at 0.39s, Command A+ third at 0.42s
    0.30s latency
    Why it wins
  4. 4
    Llama 3.1 Instruct 8B
    MetaStill the price leader at $0.02 per million tokens, with Granite 4.2 3B tied and Nova Micro at $0.03
    $0.02 / M
    Why it wins
  5. 5
    Claude Haiku 4.5
    AnthropicBest cheap option in Claude stacks
    balanced
    Why it wins

Operator take · Ranked by fit, not a single benchmark. Celeris-1 reads 1,506 tokens per second today, inside the same noisy 1,200 to 1,900 band it has bounced around for weeks, and the runner-up slot now belongs to Mercury 2.5 at 770. Throughput readings on re-runs swing wider than the gap between most models here, so do not tune a pipeline to any single number. Latency and price leaders held: Gemini 2.5 Flash-Lite at 0.30s time to first token, Llama 3.1 8B at $0.02 per million tokens. Your Clay run is bounded by rate limits and row count long before tokens per second. The rule holds. At 10,000 rows a frontier model is a budget mistake: route volume work to a fast, cheap tier and reserve frontier models for the 5% of rows that matter.

Call Transcripts & Long Context Leaderboard

Synthesizing a quarter of Gong/Granola calls, full CRM histories, or an entire content library in one pass.

  1. 1
    Llama 4 Scout
    Meta
    10M tokens
    Context window
  2. 2
    Grok 4.20 (0309)
    xAI
    2M tokens
    Context window
  3. 3
    Gemini 1.5 Pro / 3 Pro
    Google
    1-2M tokens
    Context window
  4. 4
    Claude Fable 5
    AnthropicStrong recall quality, Arena Longer Query #2 behind Opus 4.6 High
    1M tokens
    Context window
  5. 5
    Grok 4.6 (High)
    xAI32.3s time to first token, so batch it rather than sitting on it
    500K tokens
    Context window

Operator take · Raw window size is not recall quality. A 10M-token window that loses the middle of your call archive is worse than a 1M window that doesn't. These rows are held again rather than rewritten, for the fourth capture running, because the context-window card keeps returning different orderings below the top two and none of them has held long enough to publish. Llama 4 Scout at 10M and Grok 4.20 at 2M are the two constants across every read. What is independently confirmed today is the recall side: Opus 4.6 at high effort still ranks #1 on Arena's Longer Query board with Fable 5 at #2, and that is the number that actually predicts whether a long transcript survives the trip. Grok 4.6 keeps its slot at 500K, but note its 32-second time to first token: fine for an overnight transcript synthesis job, painful if you are waiting on it live. Test with a needle from your own transcripts before committing.

Updated daily from the linked public leaderboards; last capture September 23, 2026. Elo and win-rate figures are preference-based measures, not task-completion guarantees. Always validate the top pick on your own representative work before routing production volume to it.

How to Read These Rankings

Four rules before you switch models

  • Every ranking is task-specific. The #1 agent model is not the #1 writer, and neither is the right pick for 10,000 Clay rows.
  • Scores come from public leaderboards (Arena, Artificial Analysis) with capture dates shown. Preference Elo measures which output people like, not whether work gets completed.
  • Reasoning effort and harness matter: the same model at a different effort tier or in a different agent harness ranks differently.
  • Cheap plus fast beats frontier for volume work. Route by job, not by brand loyalty.

Frequently Asked Questions

Common questions

What is the best AI model for go-to-market work?

There is no single best model, which is the point of ranking by job. Claude Fable 5.1 (Max) leads Agentic GTM Workflows; Claude Fable 5 leads Outbound Copy & Content; GPT-5.6 Sol (xHigh) leads Prospect & Market Research; GPT-6 Astra (Max) leads GTM Engineering & Landing Pages. Full boards for all 8 jobs are above, captured September 23, 2026.

Which AI model should I use for cold email and outbound copy?

Claude Fable 5 currently ranks first on the Arena Text arena, Creative Writing ranks for creative writing, as of September 23, 2026. DemandLab's working rule is to draft with a top-three model and then run your own voice pass: the model gets you most of the way and your edit does the rest.

Which AI model is cheapest for high-volume enrichment?

For volume work, cost and speed matter more than frontier intelligence. Ranked by fit, not a single benchmark. Celeris-1 reads 1,506 tokens per second today, inside the same noisy 1,200 to 1,900 band it has bounced around for weeks, and the runner-up slot now belongs to Mercury 2.5 at 770. Throughput readings on re-runs swing wider than the gap between most models here, so do not tune a pipeline to any single number. Latency and price leaders held: Gemini 2.5 Flash-Lite at 0.30s time to first token, Llama 3.1 8B at $0.02 per million tokens. Your Clay run is bounded by rate limits and row count long before tokens per second. The rule holds. At 10,000 rows a frontier model is a budget mistake: route volume work to a fast, cheap tier and reserve frontier models for the 5% of rows that matter.

Where does the GTM Model Leaderboard get its data?

From two public sources: Arena (arena.ai/leaderboard) for preference and agent rankings, and Artificial Analysis (artificialanalysis.ai/models) for intelligence, speed, latency, price, and context window. DemandLab captures them daily and maps each board to the go-to-market job it applies to. Scores are published unmodified, and nothing is estimated.

DemandLab routes these models inside production GTM systems every day. If you want help picking and wiring the right models into your own stack, that's what DemandLab does.