July 13, 2026

TL;DR — AI model pricing in 2026 ranges from $0.03 to $10.00 per million input tokens — a 375x spread. DeepSeek V4 Flash ($0.14/$0.28) is the cheapest frontier-class model. Gemini 3.1 Pro ($2/$12) is the cheapest US frontier model. Claude Fable 5 ($10/$50) is the most expensive. Output tokens cost 2-5x more than input at every provider. The best value is DeepSeek V4 Pro ($0.28/$0.87, 80.6% SWE-bench) at 18x lower cost than Claude Opus. Multi-model routing, prompt caching, batch processing, and discount gateways cut total spend by 60-80%.

AI Model Pricing Comparison 2026: Every Model Ranked by Cost and Quality

AI model pricing in 2026 is the most competitive it has ever been. The cost spread between the cheapest and most expensive models is 375x on input tokens. Capability differences between top and mid-tier models have narrowed to 2-5 percentage points. Price has become the primary differentiator for buyers (aikickstart 2026).

This guide provides the complete pricing landscape for every major AI model in 2026, with verified per-token costs, cost-per-task estimates, and strategies to reduce spending.

Complete AI Model Pricing Table — July 2026

All prices are USD per million tokens, verified against provider pricing pages and independent trackers.

Frontier Models (Premium Tier)

Model Provider Input $/M Output $/M Cache $/M Context SWE-bench Pro Best For
Claude Fable 5 Anthropic $10.00 $50.00 $1.00 1M 80.3% Hardest multi-file engineering
GPT-5.5 Pro OpenAI $30.00 $180.00 N/A 400K 62.4% Premium reasoning tier
Claude Opus 4.8 Anthropic $5.00 $25.00 $0.50 1M 69.2% Coding, complex reasoning
GPT-5.5 OpenAI $5.00 $30.00 $0.50 1M 58.6% Agentic workflows
Claude Sonnet 4.6 Anthropic $3.00 $15.00 $0.30 1M 58.1% Best value frontier (US)
GPT-5.4 OpenAI $2.50 $15.00 $0.25 1M N/A Production workhorse
Gemini 3.1 Pro Google $2.00 $12.00 $0.20 2M 54.2% Reasoning, long context
Gemini 3.5 Flash Google $1.50 $9.00 $0.15 1M 48.2% High-volume agentic

Mid-Tier Models

Model Provider Input $/M Output $/M Cache $/M Context Best For
GPT-5.3-Codex OpenAI $1.75 $14.00 $0.18 400K Coding (OpenAI ecosystem)
GLM-5.2 Z.AI $1.40 $4.40 $0.26 256K General-purpose (Chinese provider)
Claude Haiku 4.5 Anthropic $1.00 $5.00 $0.10 200K Classification, routing
Kimi K2.6 Moonshot $0.95 $4.00 $0.16 256K Math, coding (open-weight)
GPT-5.4 mini OpenAI $0.75 $4.50 $0.08 256K Light tasks (OpenAI)
Mistral Large 3 Mistral $0.50 $2.00 N/A 128K EU data residency

Budget Tier (Cost Leaders)

Model Provider Input $/M Output $/M Cache $/M Context Best For
Gemini 3 Flash Google $0.50 $3.00 N/A 1M Budget frontier
DeepSeek V4 Pro DeepSeek $0.28 $0.87 $0.003 1M Best value (80.6% SWE-bench)
Qwen 3.5 Plus Alibaba $0.40 $2.40 N/A 256K General-purpose (open-weight)
MiniMax M3 MiniMax $0.30 $1.20 $0.06 512K General-purpose (open-weight)
Grok 4.3 xAI $0.20 $0.60 N/A 2M Large context budget option
GPT-5.4 Nano OpenAI $0.20 $1.25 N/A 256K Cheapest OpenAI current-gen
DeepSeek V4 Flash DeepSeek $0.14 $0.28 $0.003 1M Cheapest frontier-class
Gemini 3.1 Flash-Lite Google $0.10 $0.40 N/A 1M Cheapest US major provider
GPT-4.1 Nano OpenAI $0.10 $0.40 N/A 1M Cheapest OpenAI (any gen)
LFM2 24B A2B Together AI $0.03 $0.09 N/A 32K Absolute cheapest (open-weight)

Sources: morphllm.com pricing (June 2026), betterclaw.io LLM pricing guide (May 2026), aipricing.guru (July 2026, 128 tracked models), benchlm.ai pricing (July 2026), aikickstart.com comparison (June 2026).

Key pricing observations:

  • Output tokens cost 2-5x more than input tokens at every provider. GPT-5.5 charges $5 input / $30 output — a 6x multiplier. Claude Fable 5 charges $10 input / $50 output — a 5x multiplier. This matters because output token counts are hard to predict in advance (brainroad 2026).
  • DeepSeek V4 Pro's cache pricing is effectively free at $0.003/M — 99% off the already-low input price. For workloads with repeated context, DeepSeek's cache economics are unmatched.
  • Gemini 3.1 Pro doubles its price above 200K tokens — from $2/$12 to $4/$18. Anthropic models have no long-context surcharge: a 900K-token request bills at the same rate as a 9K one (morphllm 2026).
  • Claude Opus 4.7's new tokenizer generates up to 35% more tokens for the same text, inflating effective cost. Benchmark before migrating (betterclaw 2026).

Cost Per Task: What Does Real Work Cost?

Abstract pricing per million tokens is hard to reason about. Here is what common tasks actually cost:

Coding Tasks

Model Bug Fix Feature 20 Features/Month
DeepSeek V4 Pro $0.05 $0.21 $4.20
MiniMax M3 $0.06 $0.26 $5.20
Kimi K2.6 $0.18 $0.80 $16.00
GLM-5.1 $0.26 $1.15 $23.00
GPT-5.3-Codex $0.37 $1.54 $30.80
Gemini 3.1 Pro $0.38 $1.60 $32.00
Claude Sonnet 4.6 $0.54 $2.28 $45.60
Claude Opus 4.8 $0.90 $3.80 $76.00
GPT-5.5 $0.95 $4.00 $80.00
Claude Fable 5 $1.80 $7.60 $152.00

Source: morphllm.com coding cost analysis (June 2026). Assumes ~5K input + 1.5K output per bug fix, ~20K input + 5K output per feature.

The spread is 36x between DeepSeek V4 Pro and Claude Fable 5 for the same coding work. For a team writing 20 features/month, that is $4.20 vs $152.00.

Chatbot Workload (1,000 conversations/day, 2K tokens each)

Model Monthly Cost Annual Cost
Gemini 3.1 Flash-Lite $12 $144
DeepSeek V4 Flash $25 $300
Claude Haiku 4.5 $120 $1,440
Gemini 3.1 Pro $140 $1,680
Claude Sonnet 4.6 $180 $2,160
Claude Opus 4.8 $300 $3,600
GPT-5.5 $350 $4,200
Claude Fable 5 $600 $7,200

Source: brainroad.com workload analysis (2026). The 50x cost difference between Gemini Flash-Lite and Claude Fable 5 for a simple chatbot is striking — and the quality difference for a simple chatbot is minimal.

Developer Coding Profile (Sonnet 4.6)

Usage Profile Tokens/Day API Cost/Month Subscription
Light (2-3 tasks/day) ~1.2M in / 30K out $36 Claude Pro $20
Daily-pro (full workday) ~6M in / 150K out $178 Claude Max 5x $100
Full-day agent (continuous) ~20M in / 500K out $594 Claude Max 20x $200

Source: morphllm.com Claude cost analysis (June 2026). Assumes 75% cache hits. Subscription is cheaper than API for all profiles.

Pricing Quirks and Hidden Costs

Long-Context Surcharges

Model Below Threshold Above Threshold Threshold
Gemini 3.1 Pro $2/$12 $4/$18 200K tokens
Qwen 3.5 Plus $0.40/$2.40 $0.50/$3.00 256K tokens
MiniMax M3 $0.30/$1.20 $0.60/$2.40 512K tokens
Anthropic models Same rate Same rate No surcharge

For long-context workloads, Anthropic models have no surcharge — a 900K-token request bills at the same rate as a 9K one. Gemini doubles its price above 200K tokens. Always check long-context pricing before budgeting.

Tokenizer Inflation

Claude Opus 4.7's new tokenizer generates up to 35% more tokens for the same text compared to older models. This means your effective per-request cost may be 35% higher than the rate card suggests. Benchmark your actual token counts before migrating to Opus 4.7 (betterclaw 2026).

Output Token Multipliers

Output tokens cost 2-5x more than input tokens. This ratio matters more than most cost estimates account for:

Provider Input Output Multiplier
OpenAI (GPT-5.5) $5 $30 6x
Anthropic (Opus) $5 $25 5x
Google (Gemini Pro) $2 $12 6x
DeepSeek (V4 Pro) $0.28 $0.87 3x
DeepSeek (V4 Flash) $0.14 $0.28 2x

DeepSeek has the lowest output multiplier (2-3x), making it especially cost-effective for workloads with high output token counts (code generation, long-form writing).

Cost Optimization Strategies

Strategy 1: Multi-Model Routing (60-80% savings)

Route 70% of traffic to budget models, 30% to frontier models:

from litellm import completion

def cost_optimized_route(prompt, task_type="simple"):
    routes = {
        "simple": "gemini-3.1-flash-lite",      # $0.10/$0.40
        "medium": "deepseek-v4-pro",             # $0.28/$0.87
        "complex": "claude-opus-4-8",            # $5/$25
        "coding": "claude-sonnet-4-6",           # $3/$15
        "bulk": "deepseek-v4-flash",             # $0.14/$0.28
    }
    return completion(model=routes.get(task_type, "gemini-3.1-flash-lite"),
                      messages=[{"role": "user", "content": prompt}])

Strategy 2: Prompt Caching (50-90% input savings)

Provider Cache Discount Cached Input $/M
DeepSeek 99% off $0.003 (V4 Pro)
Google 90% off $0.15 (Gemini 3.5 Flash)
Anthropic 90% off $0.50 (Opus), $0.30 (Sonnet)
OpenAI 50-75% off $0.50 (GPT-5.5)

Strategy 3: Batch Processing (50% off)

OpenAI and Google offer 50% discounts for non-latency-sensitive workloads. If your job doesn't need to return in seconds — overnight processing, document analysis, bulk generation — batch mode is an easy 50% cut.

Strategy 4: Discount Gateways (40-80% off retail)

  • Blackmagic AI: 48-74% off across 13+ providers. 20M GPT-5.5 tokens/month = $66 vs ~$250 retail.
  • Hypereal AI: Claude Opus 4.7 at 32% below official, Claude Sonnet at 77% below.
  • OpenRouter: Aggregates 100+ models with competitive pricing.

Strategy 5: Self-Hosting (Zero per-token cost above 30B tokens/month)

Model Hardware Break-Even Volume
DeepSeek V4 Lite 2x A100 ~30B tokens/month
Llama 4 70B 1x A100 ~20B tokens/month
gpt-oss-120b 1x H100 ~25B tokens/month

How to Choose the Right Price-Tier Model

flowchart TD Start["Which price tier?"] --> Q1{"Quality needed?"} Q1 -->|"Frontier quality"| Q2{"Which provider?"} Q2 -->|"Anthropic"| Claude["Claude Sonnet 4.6\n$3/$15\nBest value frontier"] Q2 -->|"OpenAI"| GPT["GPT-5.4\n$2.50/$15\nProduction workhorse"] Q2 -->|"Google"| Gemini["Gemini 3.1 Pro\n$2/$12\nCheapest frontier"] Q2 -->|"DeepSeek"| Deep["DeepSeek V4 Pro\n$0.28/$0.87\nBest value overall"] Q1 -->|"Good enough"| Q3{"Volume?"} Q3 -->|"High volume"| Flash["Gemini 3.1 Flash-Lite\n$0.10/$0.40\nCheapest US provider"] Q3 -->|"Very high volume"| DeepFlash["DeepSeek V4 Flash\n$0.14/$0.28\nCheapest frontier-class"] Q3 -->|"Absolute cheapest"| LFM["LFM2 24B on Together\n$0.03/$0.09\nOpen-weight"] Q1 -->|"Premium / hardest tasks"| Q4{"Budget?"} Q4 -->|"Yes"| Fable["Claude Fable 5\n$10/$50\n80.3% SWE-bench Pro"] Q4 -->|"No"| Opus["Claude Opus 4.8\n$5/$25\n69.2% SWE-bench Pro"]

Choose DeepSeek V4 Pro ($0.28/$0.87) as the default for cost-sensitive workloads that need frontier quality. At 80.6% SWE-bench Verified and 18x cheaper than Claude Opus, it is the best value model in 2026.

Choose Claude Sonnet 4.6 ($3/$15) as the default for production workloads needing US-provider frontier quality. Near-Opus quality at 40% lower cost. This is the model most production teams should default to.

Choose Gemini 3.1 Pro ($2/$12) for reasoning-heavy workloads and long-context tasks. 2.5x cheaper than Claude and GPT while leading on GPQA Diamond and offering a 2M context window.

Choose Gemini 3.1 Flash-Lite ($0.10/$0.40) for high-volume simple tasks. The cheapest capable model from a US major provider with a 1M context window and generous free tier.

Choose Claude Fable 5 ($10/$50) only for the hardest multi-file engineering tasks where 80.3% SWE-bench Pro justifies the premium. For everything else, Sonnet or Opus delivers 85% of the capability at a fraction of the cost.

For a deeper dive on the cheapest APIs, see our cheapest AI API comparison. For cost estimation tools, see our LLM cost estimation guide. For strategies to reduce spending, see our guide on reducing LLM spending.

FAQ

How accurate are AI model prices in 2026?

AI pricing changes weekly. Vendors run promotional rates, and "the same model" can mean different SKUs at different prices. The prices in this guide were verified against provider pricing pages and independent trackers (aipricing.guru, benchlm.ai, morphllm.com) in July 2026. Always confirm current rates before building budgets. Several figures from third-party comparisons do not check out against vendor pricing — we have used only verified figures in our tables (aikickstart 2026).

What is the cost difference between open-source and commercial AI?

At the API level, open-source models are 3-18x cheaper. DeepSeek V4 Pro ($0.28/$0.87) is 18x cheaper than Claude Opus 4.8 ($5/$25). Llama 4 70B on Together AI ($0.23/$0.40) is 22x cheaper. Self-hosted, open-source models have zero per-token cost but require GPU infrastructure ($4,500-8,000/month for a production deployment). The break-even is ~30B tokens/month. Below that, APIs are cheaper. See our open-source vs commercial AI guide for detailed TCO analysis.

How much does ChatGPT cost per month?

ChatGPT Free: $0 (GPT-5.5 with tight daily limits). ChatGPT Go: $8/month. ChatGPT Plus: $20/month (high GPT-5.5 limits, image tools, ~320-page context). ChatGPT Pro: $100 or $200/month (5x or 20x limits, full 1M context at $200). Claude Pro: $20/month (includes Claude Code and Cowork). Claude Max: $100 or $200/month (5x or 20x Pro usage). Google AI Pro: $19.99/month (Gemini 3.1 Pro, 1M context, Workspace). Google AI Ultra: $99.99/month (~20x limits, Spark agent, 20TB storage).

Is prompt caching worth it?

Yes, for any workload with repeated context. If you send the same system prompt, few-shot examples, or document context across multiple requests, caching reduces input costs by 50-90%. DeepSeek's cache pricing at $0.003/M is effectively free. Google's at $0.15/M saves 93%. Anthropic's at $0.50/M saves 90%. For a workload with 75% cache hits (typical for coding agents), caching reduces total input cost by 67-93%. The setup is minimal — just add a cache control header to your API calls.

What is the cheapest way to use Claude?

Claude Haiku 4.5 at $1/$5 per million tokens is the cheapest Claude API. For chat usage, Claude Pro at $20/month is cheaper than API for any usage above ~1M tokens/day. For coding, Claude Pro includes Claude Code (normally a separate product). For high-volume coding, Claude Max at $100/month covers ~6M tokens/day — cheaper than the ~$178/month API equivalent. For the absolute cheapest way to use Claude quality, route simple tasks to DeepSeek V4 Pro (which matches Claude on many benchmarks) and reserve Claude for complex tasks only.


Want a self-hosted AI company brain that does all of this out of the box?
Book a demo →