TL;DR — AI model pricing in 2026 ranges from $0.03 to $10.00 per million input tokens — a 375x spread. DeepSeek V4 Flash ($0.14/$0.28) is the cheapest frontier-class model. Gemini 3.1 Pro ($2/$12) is the cheapest US frontier model. Claude Fable 5 ($10/$50) is the most expensive. Output tokens cost 2-5x more than input at every provider. The best value is DeepSeek V4 Pro ($0.28/$0.87, 80.6% SWE-bench) at 18x lower cost than Claude Opus. Multi-model routing, prompt caching, batch processing, and discount gateways cut total spend by 60-80%.
AI Model Pricing Comparison 2026: Every Model Ranked by Cost and Quality
AI model pricing in 2026 is the most competitive it has ever been. The cost spread between the cheapest and most expensive models is 375x on input tokens. Capability differences between top and mid-tier models have narrowed to 2-5 percentage points. Price has become the primary differentiator for buyers (aikickstart 2026).
This guide provides the complete pricing landscape for every major AI model in 2026, with verified per-token costs, cost-per-task estimates, and strategies to reduce spending.
Complete AI Model Pricing Table — July 2026
All prices are USD per million tokens, verified against provider pricing pages and independent trackers.
Frontier Models (Premium Tier)
| Model | Provider | Input $/M | Output $/M | Cache $/M | Context | SWE-bench Pro | Best For |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $1.00 | 1M | 80.3% | Hardest multi-file engineering |
| GPT-5.5 Pro | OpenAI | $30.00 | $180.00 | N/A | 400K | 62.4% | Premium reasoning tier |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $0.50 | 1M | 69.2% | Coding, complex reasoning |
| GPT-5.5 | OpenAI | $5.00 | $30.00 | $0.50 | 1M | 58.6% | Agentic workflows |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $0.30 | 1M | 58.1% | Best value frontier (US) |
| GPT-5.4 | OpenAI | $2.50 | $15.00 | $0.25 | 1M | N/A | Production workhorse |
| Gemini 3.1 Pro | $2.00 | $12.00 | $0.20 | 2M | 54.2% | Reasoning, long context | |
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.15 | 1M | 48.2% | High-volume agentic |
Mid-Tier Models
| Model | Provider | Input $/M | Output $/M | Cache $/M | Context | Best For |
|---|---|---|---|---|---|---|
| GPT-5.3-Codex | OpenAI | $1.75 | $14.00 | $0.18 | 400K | Coding (OpenAI ecosystem) |
| GLM-5.2 | Z.AI | $1.40 | $4.40 | $0.26 | 256K | General-purpose (Chinese provider) |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $0.10 | 200K | Classification, routing |
| Kimi K2.6 | Moonshot | $0.95 | $4.00 | $0.16 | 256K | Math, coding (open-weight) |
| GPT-5.4 mini | OpenAI | $0.75 | $4.50 | $0.08 | 256K | Light tasks (OpenAI) |
| Mistral Large 3 | Mistral | $0.50 | $2.00 | N/A | 128K | EU data residency |
Budget Tier (Cost Leaders)
| Model | Provider | Input $/M | Output $/M | Cache $/M | Context | Best For |
|---|---|---|---|---|---|---|
| Gemini 3 Flash | $0.50 | $3.00 | N/A | 1M | Budget frontier | |
| DeepSeek V4 Pro | DeepSeek | $0.28 | $0.87 | $0.003 | 1M | Best value (80.6% SWE-bench) |
| Qwen 3.5 Plus | Alibaba | $0.40 | $2.40 | N/A | 256K | General-purpose (open-weight) |
| MiniMax M3 | MiniMax | $0.30 | $1.20 | $0.06 | 512K | General-purpose (open-weight) |
| Grok 4.3 | xAI | $0.20 | $0.60 | N/A | 2M | Large context budget option |
| GPT-5.4 Nano | OpenAI | $0.20 | $1.25 | N/A | 256K | Cheapest OpenAI current-gen |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | $0.003 | 1M | Cheapest frontier-class |
| Gemini 3.1 Flash-Lite | $0.10 | $0.40 | N/A | 1M | Cheapest US major provider | |
| GPT-4.1 Nano | OpenAI | $0.10 | $0.40 | N/A | 1M | Cheapest OpenAI (any gen) |
| LFM2 24B A2B | Together AI | $0.03 | $0.09 | N/A | 32K | Absolute cheapest (open-weight) |
Sources: morphllm.com pricing (June 2026), betterclaw.io LLM pricing guide (May 2026), aipricing.guru (July 2026, 128 tracked models), benchlm.ai pricing (July 2026), aikickstart.com comparison (June 2026).
Key pricing observations:
- Output tokens cost 2-5x more than input tokens at every provider. GPT-5.5 charges $5 input / $30 output — a 6x multiplier. Claude Fable 5 charges $10 input / $50 output — a 5x multiplier. This matters because output token counts are hard to predict in advance (brainroad 2026).
- DeepSeek V4 Pro's cache pricing is effectively free at $0.003/M — 99% off the already-low input price. For workloads with repeated context, DeepSeek's cache economics are unmatched.
- Gemini 3.1 Pro doubles its price above 200K tokens — from $2/$12 to $4/$18. Anthropic models have no long-context surcharge: a 900K-token request bills at the same rate as a 9K one (morphllm 2026).
- Claude Opus 4.7's new tokenizer generates up to 35% more tokens for the same text, inflating effective cost. Benchmark before migrating (betterclaw 2026).
Cost Per Task: What Does Real Work Cost?
Abstract pricing per million tokens is hard to reason about. Here is what common tasks actually cost:
Coding Tasks
| Model | Bug Fix | Feature | 20 Features/Month |
|---|---|---|---|
| DeepSeek V4 Pro | $0.05 | $0.21 | $4.20 |
| MiniMax M3 | $0.06 | $0.26 | $5.20 |
| Kimi K2.6 | $0.18 | $0.80 | $16.00 |
| GLM-5.1 | $0.26 | $1.15 | $23.00 |
| GPT-5.3-Codex | $0.37 | $1.54 | $30.80 |
| Gemini 3.1 Pro | $0.38 | $1.60 | $32.00 |
| Claude Sonnet 4.6 | $0.54 | $2.28 | $45.60 |
| Claude Opus 4.8 | $0.90 | $3.80 | $76.00 |
| GPT-5.5 | $0.95 | $4.00 | $80.00 |
| Claude Fable 5 | $1.80 | $7.60 | $152.00 |
Source: morphllm.com coding cost analysis (June 2026). Assumes ~5K input + 1.5K output per bug fix, ~20K input + 5K output per feature.
The spread is 36x between DeepSeek V4 Pro and Claude Fable 5 for the same coding work. For a team writing 20 features/month, that is $4.20 vs $152.00.
Chatbot Workload (1,000 conversations/day, 2K tokens each)
| Model | Monthly Cost | Annual Cost |
|---|---|---|
| Gemini 3.1 Flash-Lite | $12 | $144 |
| DeepSeek V4 Flash | $25 | $300 |
| Claude Haiku 4.5 | $120 | $1,440 |
| Gemini 3.1 Pro | $140 | $1,680 |
| Claude Sonnet 4.6 | $180 | $2,160 |
| Claude Opus 4.8 | $300 | $3,600 |
| GPT-5.5 | $350 | $4,200 |
| Claude Fable 5 | $600 | $7,200 |
Source: brainroad.com workload analysis (2026). The 50x cost difference between Gemini Flash-Lite and Claude Fable 5 for a simple chatbot is striking — and the quality difference for a simple chatbot is minimal.
Developer Coding Profile (Sonnet 4.6)
| Usage Profile | Tokens/Day | API Cost/Month | Subscription |
|---|---|---|---|
| Light (2-3 tasks/day) | ~1.2M in / 30K out | $36 | Claude Pro $20 |
| Daily-pro (full workday) | ~6M in / 150K out | $178 | Claude Max 5x $100 |
| Full-day agent (continuous) | ~20M in / 500K out | $594 | Claude Max 20x $200 |
Source: morphllm.com Claude cost analysis (June 2026). Assumes 75% cache hits. Subscription is cheaper than API for all profiles.
Pricing Quirks and Hidden Costs
Long-Context Surcharges
| Model | Below Threshold | Above Threshold | Threshold |
|---|---|---|---|
| Gemini 3.1 Pro | $2/$12 | $4/$18 | 200K tokens |
| Qwen 3.5 Plus | $0.40/$2.40 | $0.50/$3.00 | 256K tokens |
| MiniMax M3 | $0.30/$1.20 | $0.60/$2.40 | 512K tokens |
| Anthropic models | Same rate | Same rate | No surcharge |
For long-context workloads, Anthropic models have no surcharge — a 900K-token request bills at the same rate as a 9K one. Gemini doubles its price above 200K tokens. Always check long-context pricing before budgeting.
Tokenizer Inflation
Claude Opus 4.7's new tokenizer generates up to 35% more tokens for the same text compared to older models. This means your effective per-request cost may be 35% higher than the rate card suggests. Benchmark your actual token counts before migrating to Opus 4.7 (betterclaw 2026).
Output Token Multipliers
Output tokens cost 2-5x more than input tokens. This ratio matters more than most cost estimates account for:
| Provider | Input | Output | Multiplier |
|---|---|---|---|
| OpenAI (GPT-5.5) | $5 | $30 | 6x |
| Anthropic (Opus) | $5 | $25 | 5x |
| Google (Gemini Pro) | $2 | $12 | 6x |
| DeepSeek (V4 Pro) | $0.28 | $0.87 | 3x |
| DeepSeek (V4 Flash) | $0.14 | $0.28 | 2x |
DeepSeek has the lowest output multiplier (2-3x), making it especially cost-effective for workloads with high output token counts (code generation, long-form writing).
Cost Optimization Strategies
Strategy 1: Multi-Model Routing (60-80% savings)
Route 70% of traffic to budget models, 30% to frontier models:
from litellm import completion
def cost_optimized_route(prompt, task_type="simple"):
routes = {
"simple": "gemini-3.1-flash-lite", # $0.10/$0.40
"medium": "deepseek-v4-pro", # $0.28/$0.87
"complex": "claude-opus-4-8", # $5/$25
"coding": "claude-sonnet-4-6", # $3/$15
"bulk": "deepseek-v4-flash", # $0.14/$0.28
}
return completion(model=routes.get(task_type, "gemini-3.1-flash-lite"),
messages=[{"role": "user", "content": prompt}])
Strategy 2: Prompt Caching (50-90% input savings)
| Provider | Cache Discount | Cached Input $/M |
|---|---|---|
| DeepSeek | 99% off | $0.003 (V4 Pro) |
| 90% off | $0.15 (Gemini 3.5 Flash) | |
| Anthropic | 90% off | $0.50 (Opus), $0.30 (Sonnet) |
| OpenAI | 50-75% off | $0.50 (GPT-5.5) |
Strategy 3: Batch Processing (50% off)
OpenAI and Google offer 50% discounts for non-latency-sensitive workloads. If your job doesn't need to return in seconds — overnight processing, document analysis, bulk generation — batch mode is an easy 50% cut.
Strategy 4: Discount Gateways (40-80% off retail)
- Blackmagic AI: 48-74% off across 13+ providers. 20M GPT-5.5 tokens/month = $66 vs ~$250 retail.
- Hypereal AI: Claude Opus 4.7 at 32% below official, Claude Sonnet at 77% below.
- OpenRouter: Aggregates 100+ models with competitive pricing.
Strategy 5: Self-Hosting (Zero per-token cost above 30B tokens/month)
| Model | Hardware | Break-Even Volume |
|---|---|---|
| DeepSeek V4 Lite | 2x A100 | ~30B tokens/month |
| Llama 4 70B | 1x A100 | ~20B tokens/month |
| gpt-oss-120b | 1x H100 | ~25B tokens/month |
How to Choose the Right Price-Tier Model
Choose DeepSeek V4 Pro ($0.28/$0.87) as the default for cost-sensitive workloads that need frontier quality. At 80.6% SWE-bench Verified and 18x cheaper than Claude Opus, it is the best value model in 2026.
Choose Claude Sonnet 4.6 ($3/$15) as the default for production workloads needing US-provider frontier quality. Near-Opus quality at 40% lower cost. This is the model most production teams should default to.
Choose Gemini 3.1 Pro ($2/$12) for reasoning-heavy workloads and long-context tasks. 2.5x cheaper than Claude and GPT while leading on GPQA Diamond and offering a 2M context window.
Choose Gemini 3.1 Flash-Lite ($0.10/$0.40) for high-volume simple tasks. The cheapest capable model from a US major provider with a 1M context window and generous free tier.
Choose Claude Fable 5 ($10/$50) only for the hardest multi-file engineering tasks where 80.3% SWE-bench Pro justifies the premium. For everything else, Sonnet or Opus delivers 85% of the capability at a fraction of the cost.
For a deeper dive on the cheapest APIs, see our cheapest AI API comparison. For cost estimation tools, see our LLM cost estimation guide. For strategies to reduce spending, see our guide on reducing LLM spending.
FAQ
How accurate are AI model prices in 2026?
AI pricing changes weekly. Vendors run promotional rates, and "the same model" can mean different SKUs at different prices. The prices in this guide were verified against provider pricing pages and independent trackers (aipricing.guru, benchlm.ai, morphllm.com) in July 2026. Always confirm current rates before building budgets. Several figures from third-party comparisons do not check out against vendor pricing — we have used only verified figures in our tables (aikickstart 2026).
What is the cost difference between open-source and commercial AI?
At the API level, open-source models are 3-18x cheaper. DeepSeek V4 Pro ($0.28/$0.87) is 18x cheaper than Claude Opus 4.8 ($5/$25). Llama 4 70B on Together AI ($0.23/$0.40) is 22x cheaper. Self-hosted, open-source models have zero per-token cost but require GPU infrastructure ($4,500-8,000/month for a production deployment). The break-even is ~30B tokens/month. Below that, APIs are cheaper. See our open-source vs commercial AI guide for detailed TCO analysis.
How much does ChatGPT cost per month?
ChatGPT Free: $0 (GPT-5.5 with tight daily limits). ChatGPT Go: $8/month. ChatGPT Plus: $20/month (high GPT-5.5 limits, image tools, ~320-page context). ChatGPT Pro: $100 or $200/month (5x or 20x limits, full 1M context at $200). Claude Pro: $20/month (includes Claude Code and Cowork). Claude Max: $100 or $200/month (5x or 20x Pro usage). Google AI Pro: $19.99/month (Gemini 3.1 Pro, 1M context, Workspace). Google AI Ultra: $99.99/month (~20x limits, Spark agent, 20TB storage).
Is prompt caching worth it?
Yes, for any workload with repeated context. If you send the same system prompt, few-shot examples, or document context across multiple requests, caching reduces input costs by 50-90%. DeepSeek's cache pricing at $0.003/M is effectively free. Google's at $0.15/M saves 93%. Anthropic's at $0.50/M saves 90%. For a workload with 75% cache hits (typical for coding agents), caching reduces total input cost by 67-93%. The setup is minimal — just add a cache control header to your API calls.
What is the cheapest way to use Claude?
Claude Haiku 4.5 at $1/$5 per million tokens is the cheapest Claude API. For chat usage, Claude Pro at $20/month is cheaper than API for any usage above ~1M tokens/day. For coding, Claude Pro includes Claude Code (normally a separate product). For high-volume coding, Claude Max at $100/month covers ~6M tokens/day — cheaper than the ~$178/month API equivalent. For the absolute cheapest way to use Claude quality, route simple tasks to DeepSeek V4 Pro (which matches Claude on many benchmarks) and reserve Claude for complex tasks only.
Want a self-hosted AI company brain that does all of this out of the box?
Book a demo →