Post

AI API Pricing in July 2026: Claude vs GPT vs Gemini vs Grok vs DeepSeek

A snapshot of what the major AI APIs actually cost in July 2026 — per-token pricing for Anthropic, OpenAI, Google, xAI, and DeepSeek across frontier, mid, and budget tiers, plus the fine print that dominates real bills.

AI API Pricing in July 2026: Claude vs GPT vs Gemini vs Grok vs DeepSeek

If you’re building anything on top of an LLM (Large Language Model) API in mid-2026, the pricing landscape looks nothing like it did a year ago. New flagship models have shipped from every major provider in the last few months — Claude Fable 5, GPT-5.6, Gemini 3.5, Grok 4.5 — and the spread between the most and least expensive options has stretched to more than an order of magnitude. Picking a model on capability alone can leave you with a bill 8x higher than a competitor that would have handled your workload fine.

This post is a snapshot of what the major APIs actually cost right now, and — more usefully — the pricing mechanics that determine what you’ll really pay. All prices are USD per million tokens, input / output.

The Quick Take

  Cheapest usable Mid tier Frontier
Anthropic Haiku 4.5 ($1/$5) Sonnet 5 ($3/$15) Fable 5 ($10/$50)
OpenAI GPT-5.4 nano ($0.20/$1.25) GPT-5.6 Terra ($2.50/$15) GPT-5.6 Sol ($5/$30)
Google Gemini 3.1 Flash-Lite ($0.25/$1.50) Gemini 3.5 Flash ($1.50/$9) Gemini 3.1 Pro ($2/$12)
xAI Grok 4.1 Fast ($0.20/$0.50) Grok 4.3 ($1.25/$2.50) Grok 4.5 ($2/$6)
DeepSeek V4-Flash ($0.14/$0.28) V4-Pro ($0.435/$0.87)

If you just want the verdict: the frontier tier is where providers differentiate on price the most (Anthropic charges a premium, xAI undercuts aggressively), the budget tier is nearly free at DeepSeek’s rates, and for most workloads the number that actually matters is the output price — not the input price everyone quotes.

Frontier Tier: The Premium Spread

Provider Model Input Output
Anthropic Claude Fable 5 $10.00 $50.00
Anthropic Claude Opus 4.8 $5.00 $25.00
OpenAI GPT-5.6 Sol / GPT-5.5 $5.00 $30.00
xAI Grok 4.5 $2.00 $6.00
Google Gemini 3.1 Pro $2.00 $12.00

The story at the top is divergence. Anthropic’s Claude Fable 5 — the first of their new Mythos-class tier, positioned above Opus — is the most expensive widely available model at $10/$50, roughly 8x Grok 4.5’s output price. Anthropic is betting that frontier capability commands frontier pricing; xAI is betting the opposite, that aggressive pricing (plus up to $175/month in free API credits through their data-sharing program) buys market share.

Opus 4.8 and GPT-5.6 Sol sit in the middle at essentially matched pricing — $5 input, with OpenAI charging slightly more on output ($30 vs $25). If you’ve been treating “Opus vs GPT flagship” as a cost decision, it isn’t one anymore; pick on capability fit.

One wrinkle worth knowing: Gemini 3.1 Pro’s listed $2/$12 jumps to $4/$18 once your context exceeds 200K tokens. Long-context workloads on Gemini cost double what the headline number suggests.

Mid Tier: The Workhorse Zone

Provider Model Input Output
Anthropic Claude Sonnet 5 $3.00 $15.00
OpenAI GPT-5.6 Terra / GPT-5.4 $2.50 $15.00
Google Gemini 3.5 Flash $1.50 $9.00
xAI Grok 4.3 $1.25 $2.50
OpenAI GPT-5.6 Luna $1.00 $6.00

This is where most production traffic should live, and the pricing is tightly clustered — with one time-sensitive exception: Claude Sonnet 5 has introductory pricing of $2/$10 through August 31, 2026, after which it reverts to $3/$15. If you’re evaluating Sonnet 5 on cost right now, budget against the sticker price, not the promo.

Budget Tier: Racing to the Floor

Provider Model Input Output
Anthropic Claude Haiku 4.5 $1.00 $5.00
Google Gemini 3 Flash $0.50 $3.00
DeepSeek V4-Pro $0.435 $0.87
Google Gemini 3.1 Flash-Lite $0.25 $1.50
xAI Grok 4.1 Fast $0.20 $0.50
OpenAI GPT-5.4 nano $0.20 $1.25
DeepSeek V4-Flash $0.14 $0.28

DeepSeek remains the absolute floor. V4-Flash at $0.14/$0.28 costs 35x less on input than Opus 4.8 — and its cache-hit pricing is an almost comical $0.0028 per million tokens, about 2% of the base rate. For classification, extraction, routing, and other tasks where you don’t need frontier reasoning, running anything more expensive than this tier is money on fire.

Note that Anthropic’s Haiku 4.5, at $1/$5, is priced like a premium budget model — it’s the only entry here that costs more than some competitors’ mid tier. Whether that’s justified depends entirely on whether your task benefits from Haiku’s stronger reasoning relative to the Flash/nano class.

The Fine Print That Dominates Real Bills

The per-token tables above are the part everyone compares. These four mechanics are the part that determines what you actually pay.

Output price is the real price for agentic workloads. Anthropic and OpenAI both run 5–6x output-to-input multipliers; xAI runs 2–3x. Reasoning models bill their thinking tokens as output on every provider. An agent that reads a little and writes a lot — tool calls, chain-of-thought, code generation — has a bill dominated by the output column. Compare providers on that column, not the input price in the headline.

Prompt caching is table stakes, at roughly the same rate everywhere. Cached input bills at ~10% of the base input rate on Anthropic, OpenAI, and Google; xAI charges 25%; DeepSeek ~2%. If you’re running a chatbot or agent that resends a large system prompt or conversation history every turn, caching is the single biggest lever you have — often bigger than switching providers.

Batch discounts are a flat 50%. OpenAI, Google, and Anthropic all take half off for asynchronous batch jobs. Anything that doesn’t need a real-time response — nightly enrichment, bulk classification, report generation — should be running through a batch API.

Promos expire and context tiers bite. Sonnet 5’s intro pricing ends August 31. Gemini 3.1 Pro doubles past 200K context. xAI’s free credits require opting into data sharing. Every one of these changes the math, and none of them appear in a simple per-token comparison table.

Key Takeaways

  • The frontier spread is enormous: Claude Fable 5’s output tokens cost ~8x Grok 4.5’s. Premium pricing at the top is a deliberate strategy, not an accident.
  • For agentic and reasoning-heavy workloads, compare output prices — that’s where 5–6x multipliers and thinking tokens live.
  • Caching (~90% off repeated input) and batching (50% off async work) are near-universal and often matter more than model choice.
  • DeepSeek V4-Flash at $0.14/$0.28 is the floor; if your task doesn’t need frontier reasoning, the budget tier is effectively free.
  • Check the fine print: intro pricing (Sonnet 5 until Aug 31), long-context surcharges (Gemini past 200K), and free-credit programs (xAI) all move the real number.

Prices in this post were checked in mid-July 2026 against provider documentation and third-party trackers; this market reprices fast, so verify against the official pages before committing to a budget.

Resources

This post is licensed under CC BY 4.0 by the author.