DemandLab

Updated daily · Arena + Artificial Analysis

High-Volume Enrichment & Extraction

Clay-style enrichment at thousands of rows, lead scoring, intent classification, schema extraction. Cost and speed dominate; frontier IQ is wasted here.

One board from The GTM Model Leaderboard

Quick answer

For High-Volume Enrichment & Extraction, Celeris's Celeris-1 leads at 1,506 tok/s Why it wins on the Artificial Analysis speed, latency, price leaders, ahead of Mercury 2.5. DemandLab recaptures this board daily; these numbers are from September 23, 2026. Rankings are job-specific, so the winner here is not the winner on the other seven boards.

Compiled by Chris Arden, Fractional CMO, DemandLab · Updated on

High-Volume Enrichment & Extraction Leaderboard

Clay-style enrichment at thousands of rows, lead scoring, intent classification, schema extraction. Cost and speed dominate; frontier IQ is wasted here.

  1. 1
    Celeris-1
    CelerisFastest output measured, down from 1,541 tok/s
    1,506 tok/s
    Why it wins
  2. 2
    Mercury 2.5
    InceptionSecond on throughput; Mercury 2 now third at 750 tok/s
    770 tok/s
    Why it wins
  3. 3
    Gemini 2.5 Flash-Lite
    GoogleLowest time to first token measured; North Mini Code second at 0.39s, Command A+ third at 0.42s
    0.30s latency
    Why it wins
  4. 4
    Llama 3.1 Instruct 8B
    MetaStill the price leader at $0.02 per million tokens, with Granite 4.2 3B tied and Nova Micro at $0.03
    $0.02 / M
    Why it wins
  5. 5
    Claude Haiku 4.5
    AnthropicBest cheap option in Claude stacks
    balanced
    Why it wins

Operator take · Ranked by fit, not a single benchmark. Celeris-1 reads 1,506 tokens per second today, inside the same noisy 1,200 to 1,900 band it has bounced around for weeks, and the runner-up slot now belongs to Mercury 2.5 at 770. Throughput readings on re-runs swing wider than the gap between most models here, so do not tune a pipeline to any single number. Latency and price leaders held: Gemini 2.5 Flash-Lite at 0.30s time to first token, Llama 3.1 8B at $0.02 per million tokens. Your Clay run is bounded by rate limits and row count long before tokens per second. The rule holds. At 10,000 rows a frontier model is a budget mistake: route volume work to a fast, cheap tier and reserve frontier models for the 5% of rows that matter.

Updated daily from the linked public leaderboards; last capture September 23, 2026. Elo and win-rate figures are preference-based measures, not task-completion guarantees. Always validate the top pick on your own representative work before routing production volume to it.

How to Read These Rankings

Four rules before you switch models

  • Every ranking is task-specific. The #1 agent model is not the #1 writer, and neither is the right pick for 10,000 Clay rows.
  • Scores come from public leaderboards (Arena, Artificial Analysis) with capture dates shown. Preference Elo measures which output people like, not whether work gets completed.
  • Reasoning effort and harness matter: the same model at a different effort tier or in a different agent harness ranks differently.
  • Cheap plus fast beats frontier for volume work. Route by job, not by brand loyalty.

Frequently Asked Questions

High-Volume Enrichment & Extraction: common questions

Which AI model is best for High-Volume Enrichment & Extraction?

As of September 23, 2026, Celeris-1 from Celeris ranks first for High-Volume Enrichment & Extraction at 1,506 tok/s Why it wins on the Artificial Analysis speed, latency, price leaders. Mercury 2.5 ranks second at 770 tok/s, and Gemini 2.5 Flash-Lite third at 0.30s latency. DemandLab recaptures the board daily, so check the date above before quoting a position.

Where do these AI model rankings come from?

High-Volume Enrichment & Extraction is scored from the Artificial Analysis speed, latency, price leaders (https://artificialanalysis.ai/models), measured by Why it wins. DemandLab captures the public results, maps them to the go-to-market job they apply to, and publishes them unmodified. No scores are estimated, blended, or adjusted, and any board that cannot be read on a given day is left unchanged rather than guessed.

How often is the GTM Model Leaderboard updated?

Daily. The last capture was September 23, 2026. Model leaderboards move faster than most buying cycles, and positions inside a confidence interval can flip overnight, so DemandLab treats a single day's order as one observation rather than a trend.

Should I switch to the top-ranked model?

Not automatically. Ranked by fit, not a single benchmark. Celeris-1 reads 1,506 tokens per second today, inside the same noisy 1,200 to 1,900 band it has bounced around for weeks, and the runner-up slot now belongs to Mercury 2.5 at 770. Throughput readings on re-runs swing wider than the gap between most models here, so do not tune a pipeline to any single number. Latency and price leaders held: Gemini 2.5 Flash-Lite at 0.30s time to first token, Llama 3.1 8B at $0.02 per million tokens. Your Clay run is bounded by rate limits and row count long before tokens per second. The rule holds. At 10,000 rows a frontier model is a budget mistake: route volume work to a fast, cheap tier and reserve frontier models for the 5% of rows that matter.

DemandLab routes these models inside production GTM systems every day. If you want help picking and wiring the right models into your own stack, that's what DemandLab does.