Context Window Calculator

Check how much of a model's context window your text uses, and whether it fits.

A context window is the maximum number of tokens a model can handle in one request — the input prompt and the generated output combined. Current models range from GPT-3.5 Turbo's 16,385 tokens up to roughly 1,050,000 for the GPT-5.6 family. This tool shows how much of a given model's window your text actually fills.

Context windows verified against provider docs on September 2, 2026.

0 / 1,050,000 tokens0.0%

✓ Fits within the context window

AI cost workflow

  1. Count tokens
  2. Check it fits
  3. Estimate cost

Found this useful? Share it

X LinkedIn Reddit WhatsApp

About Context Window Calculator

Every LLM has a maximum context window — the total number of tokens (input plus expected output) it can process in one request — and while the gap has narrowed as most current flagship models settled around 1,000,000-1,050,000 tokens, it still varies a lot lower down the lineup. GPT-3.5 Turbo is the clear outlier at just 16,385 tokens; GPT-4o holds at 128,000; GPT-5, GPT-5.1, and the GPT-5.4 family sit at 400,000; GPT-4.1 and the current GPT-5.5/5.6 tiers reach roughly 1,000,000-1,050,000. Claude Haiku 4.5 caps at 200,000, while Claude Sonnet 5, Opus 5, and Fable 5.1 all share the full 1,000,000-token window. Gemini 3.1 Pro and Gemini 3.8 Flash both reach roughly 1,000,000 tokens too. Paste your document, codebase, or conversation history, pick a model, and this tool shows exactly how many tokens it uses — via the same real tiktoken tokenizer used in the Token Counter for GPT models, or a character-based estimate for Claude and Gemini, which don't publish an offline tokenizer — and what percentage of that model's window it fills. As a rough sense of scale: a 300-page novel (roughly 80,000-90,000 words) comes out to around 110,000-120,000 tokens using the standard ~0.75-words-per-token estimate — too big for GPT-3.5 Turbo without chunking, a tight fit inside GPT-4o's 128K window, and a small fraction of every other current model's window. This is useful for deciding whether to chunk a document, switch to a longer-context model, or trim your prompt before hitting a hard limit. Model context windows change with every new release faster than any static list can track, so treat these numbers as a starting point and verify against the provider's own docs before anything that matters.

Frequently Asked Questions

Yes — a model's context window is shared between the input prompt and the generated output. If your input uses 90% of the window, very little room is left for a long response.

They're reference values verified directly against provider docs, not third-party aggregators — but providers update context windows with new model releases faster than any static list can track, so check the provider's model documentation for the exact figure on their latest release.

Different model families use different tokenizers, so the same text can produce a different token count depending on which model you select.

GPT-3.5 Turbo's 16,385-token window reflects its older generation — it predates the long-context push that gave newer GPT, Claude, and Gemini models their much larger windows. It's the one model on this list where hitting the limit on a moderately long document is common.

At roughly 1,000,000 tokens, Gemini 3.1 Pro can hold around 750,000 words of text in one request — large enough for most small-to-medium codebases or a full novel-length manuscript, though very large monorepos can still exceed it. Note pricing on this model doubles above 200K tokens in a single request, separate from the context limit itself.