Reference

AI coding model API prices (Claude Code and Codex)

What each model costs per million tokens at list price, as used by Claude Code and Codex CLI. This is the exact table Tallyhook uses to price every session, so it is kept current.

Updated

Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens. Claude Sonnet 5 costs $2 and $10. Claude Fable 5.1 costs $10 and $50. GPT-6 Astra, the Codex CLI default since September 2026, costs $10 and $50. Cache reads are far cheaper than fresh input, which is why long coding sessions cost less than their raw token counts suggest.

Claude models (Claude Code)

USD per million tokens. Claude bills prompt-cache writes at 1.25x input for the 5-minute cache and 2x input for the 1-hour cache; Claude Code uses both. Fast mode (Opus 5 and 4.8, Claude API only) costs $10 / $50 with the same cache multipliers on top.

ModelAPI idInputOutputCache readCache write 5mCache write 1h
Claude Fable 5.1claude-fable-5-1$10.00$50.00$0.25$12.50$20.00
Claude Mythos 5.1claude-mythos-5-1$10.00$50.00$0.25$12.50$20.00
Claude Fable 5claude-fable-5$10.00$50.00$1.00$12.50$20.00
Claude Mythos 5claude-mythos-5$10.00$50.00$1.00$12.50$20.00
Claude Opus 5.5 (fast mode)claude-opus-5-5:fast$8.00$40.00$0.40$10.00$16.00
Claude Opus 5.5claude-opus-5-5$4.00$20.00$0.20$5.00$8.00
Claude Opus 5 (fast mode)claude-opus-5:fast$10.00$50.00$1.00$12.50$20.00
Claude Opus 4.8 (fast mode)claude-opus-4-8:fast$10.00$50.00$1.00$12.50$20.00
Claude Opus 5claude-opus-5$5.00$25.00$0.50$6.25$10.00
Claude Opus 4.8claude-opus-4-8$5.00$25.00$0.50$6.25$10.00
Claude Opus 4.7claude-opus-4-7$5.00$25.00$0.50$6.25$10.00
Claude Opus 4.6claude-opus-4-6$5.00$25.00$0.50$6.25$10.00
Claude Opus 4.5claude-opus-4-5$5.00$25.00$0.50$6.25$10.00
Claude Opus 4.1claude-opus-4-1$15.00$75.00$1.50$18.75$30.00
Claude Opus 4claude-opus-4$15.00$75.00$1.50$18.75$30.00
Claude Sonnet 5claude-sonnet-5$2.00$10.00$0.20$2.50$4.00
Claude Sonnet 4.6claude-sonnet-4-6$3.00$15.00$0.30$3.75$6.00
Claude Sonnet 4.5claude-sonnet-4-5$3.00$15.00$0.30$3.75$6.00
Claude Sonnet 4claude-sonnet-4$3.00$15.00$0.30$3.75$6.00
Claude Haiku 4.5claude-haiku-4-5$1.00$5.00$0.10$1.25$2.00
Claude Haiku 3.5claude-3-5-haiku$0.80$4.00$0.08$1.00$1.60

OpenAI models (Codex CLI)

USD per million tokens, standard tier, short context. OpenAI bills cached input at the cache-read rate and has no separate cache-write charge. "Pro" models publish no cached rate, so cached input is priced as input. Some models bill long-context requests at a higher rate (for example 2x input), which this table does not model.

ModelAPI idInputOutputCached input
GPT-6 Astragpt-6-astra$10.00$50.00$1.00
GPT-5.6 Solgpt-5.6-sol$4.00$20.00$0.40
GPT-5.6 Terragpt-5.6-terra$2.00$12.00$0.20
GPT-5.6 Lunagpt-5.6-luna$0.20$1.20$0.02
GPT-5.5 Progpt-5.5-pro$30.00$180.00$30.00
GPT-5.5gpt-5.5$5.00$30.00$0.50
GPT-5.4 Progpt-5.4-pro$30.00$180.00$30.00
GPT-5.4 minigpt-5.4-mini$0.75$4.50$0.075
GPT-5.4 nanogpt-5.4-nano$0.20$1.25$0.02
GPT-5.4gpt-5.4$2.50$15.00$0.25
GPT-5.3 Codexgpt-5.3-codex$1.75$14.00$0.175
GPT-5.2 Codexgpt-5.2-codex$1.75$14.00$0.175
GPT-5.1 Codex minigpt-5.1-codex-mini$0.25$2.00$0.025
GPT-5.1 Codex Maxgpt-5.1-codex-max$1.25$10.00$0.125
GPT-5.1 Codexgpt-5.1-codex$1.25$10.00$0.125
GPT-5.2 Progpt-5.2-pro$21.00$168.00$21.00
GPT-5.2gpt-5.2$1.75$14.00$0.175
GPT-5.1gpt-5.1$1.25$10.00$0.125
GPT-5 Progpt-5-pro$15.00$120.00$15.00
GPT-5 Codex minigpt-5-codex-mini$0.25$2.00$0.025
GPT-5 Codexgpt-5-codex$1.25$10.00$0.125
GPT-5 minigpt-5-mini$0.25$2.00$0.025
GPT-5 nanogpt-5-nano$0.05$0.40$0.005
GPT-5gpt-5$1.25$10.00$0.125
codex-minicodex-mini$1.50$6.00$0.375
o3-proo3-pro$20.00$80.00$20.00
o3-minio3-mini$1.10$4.40$0.55
o3o3$2.00$8.00$0.50
o4-minio4-mini$1.10$4.40$0.275

How to turn tokens into a cost

For one request: input × input price + output × output price + cache reads × cache-read price + cache writes × cache-write price, each divided by 1,000,000. Claude Code logs all four counts for every response, and Codex CLI logs a running total per session. Two details change the answer materially:

  • Deduplicate streamed rows. Claude Code writes one log row per content block, each repeating the same usage, so summing rows overcounts. Count each message.id once, keeping the row with the most output.
  • Split cache writes by lifetime. The 1-hour cache costs 2x input, not 1.25x. Claude Code reports it separately as cache_creation.ephemeral_1h_input_tokens.

Tallyhook does this for every session on every developer's machine and attributes the result to a client. See how cost is calculated.

What about subscriptions?

On Claude Pro, Max, Team or Enterprise seats, and on ChatGPT plans, usage inside the plan has no per-token bill. A list-price figure is still the useful number: it is what the same work would cost on the API, and it is the defensible basis for billing a client for AI usage. See how to bill clients for AI coding agents.

Sources

Spot a price that changed? Email support@tallyhook.dev. Workspaces can also override any model's price for negotiated rates.