93%
of the top score for 30% of the cost
Same model, one effort setting lower. Claude Opus 5.5 at high effort scored 54 at $1.82 per task. At max effort it scored 58 at $5.98. The last four points cost more than three times as much.
ReportData as of October 3, 2026Seven-minute read
A plain-English look at what AI models cost, how good they are, and where they run. The numbers come from independent benchmarks. The point is simple: matching the model to the job beats sending every call to the top tier.
In plain English
Three findings
93%
of the top score for 30% of the cost
Same model, one effort setting lower. Claude Opus 5.5 at high effort scored 54 at $1.82 per task. At max effort it scored 58 at $5.98. The last four points cost more than three times as much.
86%
of the top score for 5% of the cost
GPT-6.1 Sol at high effort scored 50 at $0.32 per task. For most everyday agent steps, that gap in score never shows up in the result. The gap in the bill does.
79%
of the top score for 2% of the cost
MiMo-V2.6-Pro, the highest-scoring open-weight model on the board, scored 46 at $0.13 per task. Open weights also mean you choose where it runs.
The scoreboard
Selected models from the Artificial Analysis leaderboard. The bar shows each model's score as a share of the top result. The last column shows how many times cheaper each one is to run per task.
Claude Opus 5.5 (max effort)
Anthropic · Closed
58
Claude Opus 5.5 (high effort)
Anthropic · Closed
54
GPT-6 Astra (max)
OpenAI · Closed
53
Gemini 4 Argon (high)
Google · Closed
53
GPT-6.1 Sol (high)
OpenAI · Closed
50
MiMo-V2.6-Pro
Xiaomi · Open weights
46
GLM-5.3 (max)
Z AI · Open weights
45
GLM-5.3-Flash
Z AI · Open weights
42
Claude Sonnet 5.5 (medium)
Anthropic · Closed
41
DeepSeek V4.1 Flash (max)
DeepSeek · Open weights
39
GPT-6 Luna (high)
OpenAI · Closed
33
Source: Artificial Analysis LLM Leaderboard, Intelligence Index v4.3.2, read October 3, 2026. Cost per task is measured on their test tasks; your own tasks will cost more or less.
Match the model to the job
01
Contract review, multi-step planning, root-cause analysis, code changes across a large system.
What matters
Accuracy. A wrong answer here is expensive.
Usually a small share of an agent's calls, so the premium price stays contained.
Reach for
Top tier at a measured effort setting: Claude Opus 5.5 (high), GPT-6 Astra, Gemini 4 Argon.
02
Drafting replies, summarizing a record, filling a form, choosing the next tool to call.
What matters
Good-enough quality at a price you can run all day.
This is where most of the volume lives, and where most of the savings come from.
Reach for
Mid tier: GPT-6.1 Sol, Claude Sonnet 5.5, GLM-5.3.
03
Classifying tickets, tagging documents, extracting fields, routing requests.
What matters
Cost and speed per call, at thousands of calls a day.
GPT-6 Luna (high) ran at $0.03 per task, about 199 times cheaper than the top setting.
Reach for
Small and fast: GPT-6 Luna, DeepSeek V4.1 Flash, GLM-5.3-Flash.
04
Phone agents, website chat, anything a person waits on.
What matters
Time to the first word.
The top-scoring max-effort setting averaged over eleven minutes to first output in the same test. Right for a research task, wrong for a phone call.
Reach for
Low-latency settings: DeepSeek V4.1 Flash (0.95 seconds to first output), Claude Sonnet 5.5 at medium (1.23 seconds).
05
Patient records, financials, customer PII, anything under a compliance regime.
What matters
Where the data goes and who can see it.
Your prompts go to the provider you chose, under your contract, and never to the model's creator.
Reach for
Open-weight models served by a US provider you contract with, or deployed in your own cloud account.
A worked example
An illustration using the benchmark cost per task. Real workloads differ, which is why we measure yours before recommending anything. The shape of the result rarely changes: the bulk of the calls decides the bill.
Every call on Claude Opus 5.5 (max effort)
$5,980
10% hard reasoning on Claude Opus 5.5 (high)
$182
30% everyday steps on GPT-6.1 Sol (high)
$96
60% high-volume sorting on GPT-6 Luna (high)
$18
Matched to the job
$296
95% lower, with the hard work still on a top-tier model
Open-weight models, US-hosted
The model file itself is published. Any provider can run it, and so can your own cloud account. Your prompts go to whoever runs it under your contract, never to the company that trained it.
GLM-5.3 is served by 25 providers and DeepSeek V4.1 Flash by 22, including US companies such as Together AI, Fireworks, Baseten, Databricks, CoreWeave, DeepInfra, Modal. You can also deploy in your own AWS or Azure account.
Licenses vary, and some limit commercial use. Providers vary too: the same DeepSeek model ranged from $0.16 to $3.13 per task across providers, and GLM-5.3 endpoints kept between 94% and 100% of reference accuracy. Read each provider's data-retention terms, then test the endpoint you will actually use.
How we use this
A benchmark tells you which models are worth trying. It does not tell you which one handles your invoices, your tickets, or your customers. For every agent we build, we write an evaluation set from your real cases first, then run candidate models against it.
Each step gets the least expensive model that passes. The hard steps keep a top-tier model. The tests stay in place after launch, so when a new model ships or a price changes, switching is a measured decision instead of a guess.
Sources and method
Keep reading
Seven questions that show whether a delivery partner is set up to protect your launch. Ask every vendor, including us. Written for buyers who need to defend the choice to a CEO or a board.
Read →A mid-2026 industry report on where e-commerce actually stands: the real cost of replatforming, why look and feel still sells but no longer closes, and what the agentic shopping surge means for merchants. AI-referred retail traffic grew 393% year over year in Q1 and now converts 42% better than other channels. The highest-ROI investment in retail right now is not a new storefront. It is making the storefront you have legible to machines.
Read →Part 2 of the STRIVE Forum AI miniseries. A 90-minute hands-on workshop for small business owners choosing between the dozen AI tools landing in their inbox every week. Two short concept blocks, two embedded labs, and a four-question rubric to grade any AI output before sending it.
Read →Five levels, seven capability axes, and the four numbers your CFO will ask. The honest picture of what it takes to run AI agents in production, and how most teams overrate themselves by 1.5 levels.
Read →The order of operations we use with mid-market engineering teams that have been told to ship AI and do not know where to start. Six stages, named exit criteria, the anti-patterns that predict failure, and the first-90-days view that ties architecture, evaluation, and model economics into a coherent adoption sequence.
Read →The release-gate playbook for AI features. Covers the five evaluation dimensions, how to build a lean golden set, where LLM-as-judge is trustworthy and where it lies, rollout mechanics with named exit criteria, and the regression suite that keeps a shipped AI feature from quietly rotting in production.
Read →Practices
Voice and intake agents that run in production: they answer the phone, book meetings, and hand off to your team.
Book a call →