AI API Cost Calculator

Compare the cost of an LLM API call across every current model at once, sorted cheapest first.

LLM APIs bill per million tokens, priced separately for input and output — output typically costs 3-5x more than input. Published rates span $0.05 per million input tokens at the cheapest tier up to $30 at the most expensive, so which model you pick usually moves cost far more than how long your prompt is.

Prices verified against provider docs on September 2, 2026. Sorted by total cost — click any column header to re-sort.

GPT-5 nano

OpenAI · Context window inferred from GPT-5 mini

$0.05$0.40400K$0.105

GPT-4o mini

OpenAI

$0.15$0.60128K$0.195

GPT-5.6 Luna

OpenAI

$0.20$1.201,050K$0.34

GPT-5.4 nano

OpenAI · Context window inferred from GPT-5.4 mini

$0.20$1.25400K$0.35

GPT-4.1 mini

OpenAI

$0.40$1.601,000K$0.52

GPT-5 mini

OpenAI

$0.25$2.00400K$0.525

GPT-3.5 Turbo

OpenAI

$0.50$1.5016.385K$0.55

Gemini 3.5 Flash-Lite

Google · Context window inferred from Gemini 3.1 Pro / 3.8 Flash

$0.30$2.501,000K$0.65

Gemini 3.8 Flash

Google · Promotional rate through Dec 31, 2026; rises to $1.50/$7.50 after

$0.75$3.751,000K$1.125

GPT-5.4 mini

OpenAI

$0.75$4.50400K$1.275

Claude Haiku 4.5

Anthropic

$1.00$5.00200K$1.50

GPT-4.1

OpenAI

$2.00$8.001,000K$2.60

GPT-5.1

OpenAI

$1.25$10.00400K$2.625

GPT-5

OpenAI

$1.25$10.00400K$2.625

Claude Sonnet 5

Anthropic

$2.00$10.001,000K$3.00

GPT-4o

OpenAI

$2.50$10.00128K$3.25

GPT-5.6 Terra

OpenAI

$2.00$12.001,050K$3.40

Gemini 3.1 Pro

Google · Higher rate applies above 200K context

$2.00$12.001,000K$3.40

GPT-5.4

OpenAI · Higher rate applies above 272K context

$2.50$15.00400K$4.25

GPT-5.6 Sol

OpenAI

$4.00$20.001,050K$6.00

Claude Opus 5

Anthropic

$5.00$25.001,000K$7.50

GPT-5.5

OpenAI · Higher rate applies above 272K context

$5.00$30.001,050K$8.50

Claude Fable 5.1

Anthropic

$10.00$50.001,000K$15.00

GPT-5.5 Pro

OpenAI · Higher rate applies above 272K context

$30.00$180.001,050K$51.00

AI cost workflow

  1. Count tokens
  2. Check it fits
  3. Estimate cost

Found this useful? Share it

X LinkedIn Reddit WhatsApp

About AI API Cost Calculator

Enter your input/output token counts (use the Token Counter tool to get exact counts for your actual prompt) and how many requests you're planning — every model below updates its cost live, sorted cheapest-first by default, so you can see at a glance which model is the cheapest fit for your workload instead of checking one model at a time. Prices were verified directly against each provider's own pricing docs on September 2, 2026 — not third-party aggregators — and every model listed is a current, non-retired tier. Two simplifications worth knowing: GPT-5.5, GPT-5.5 Pro, GPT-5.4, and Gemini 3.1 Pro all charge a higher rate once a request crosses a context-length threshold (272K for the GPT-5.4/5.5 family, 200K for Gemini 3.1 Pro) — the table shows the lower/standard rate, flagged with a note on those rows. And this list intentionally shows one representative tier per price point rather than every point-release each provider ships — click "Verify current pricing" below the table for the full, current lineup straight from the source. Model pricing changes fast enough that this list needs revisiting regularly to stay accurate — treat the numbers as a strong starting point, not a locked-in quote for anything that matters.

Frequently Asked Questions

Verified directly against each provider's own docs (not a secondary aggregator) on September 2, 2026. AI API pricing changes fast — new model generations, promotional rates with expiration dates, and price cuts all happen within months, not years — so always confirm on the provider's live pricing page (linked below the table) before budgeting anything real.

A few current models (GPT-5.5, GPT-5.5 Pro, GPT-5.4, and Gemini 3.1 Pro) charge more per token once a single request crosses a length threshold — this calculator shows the lower, standard-tier rate for simplicity. If your typical request is longer than that threshold, your real cost on those specific models will be higher than shown here.

Every major LLM provider charges more for output tokens than input tokens, since generating text is more computationally expensive than reading it — output is typically 3-6x the input price across the models listed here.

Use the Token Counter tool to get an exact count for OpenAI models (via the real tiktoken library), or a labeled estimate for Claude and Gemini, then paste those numbers into the fields here.

No — this shows standard per-token pricing only. Every major provider offers a separate discounted rate for prompt caching (often 90% off cached input) or batch/asynchronous processing (typically 50% off both directions), which can cut real costs significantly for repetitive or non-time-sensitive workloads. Check the provider's pricing page for those separate rates.

Each provider ships far more point-releases and size variants than shown here. This list is intentionally trimmed to one representative entry per price/context tier per provider to stay legible and genuinely verifiable, rather than an exhaustive catalog that's harder to keep accurate.