ChatGPT vs Gemini: price, context and limits compared
| GPT-5.5 | Gemini 3.1 Pro Preview | |
|---|---|---|
| Provider | OpenAI | |
| Context window | 1.05M | not stated |
| Max output | 128k | not stated |
| Input price | $5/1M tokens | $2/1M tokens |
| Output price | $30/1M tokens | $12/1M tokens |
| Figures as of | 2026-07-07 | 2026-07-07 |
| Cost of 1M in + 1M out | $35.00 | $14.00 |
ChatGPT, built by OpenAI, and Gemini, built by Google, are two of the most widely used assistants, and both companies publish the specifics this page sets side by side: what each flagship model costs to run, how much it can read in one go, and how much it can write back. This page does not rank the two on quality. It compares what’s measurable and leaves the rest to you.
What you can actually compare
Both vendors meter usage with a rate for incoming text and a separate, higher rate for generated text — that pattern, output priced above input, is standard across this category rather than a distinguishing feature of either company. One structural difference worth knowing: OpenAI applies a higher rate once a prompt crosses a certain length, and Google prices this Gemini model in tiers that shift with prompt size too — both vendors, in other words, treat very long prompts differently from short ones, just via different mechanisms. The table above carries the exact current figures.
Context window is the second axis: the amount of text a model can hold in view — instructions, documents, prior conversation — for a single exchange. Output limit is separate again, capping how much text comes back in one response rather than how much can be read in. Both figures are detailed in the table rather than repeated here, since they change on each vendor’s own schedule.
When the difference matters
Everyday, occasional use rarely surfaces a meaningful cost difference between the two — a normal day of questions costs very little regardless of which you use. The picture changes at volume: a service sending a large number of requests, especially ones with long prompts or long replies, will feel the input/output split and any length-based pricing tiers far more than a casual user ever would.
A large context window matters when a task needs everything in view simultaneously — a lengthy document, an extended back-and-forth, a big set of reference material — rather than split into smaller chunks fed one at a time. If your questions are short and self-contained, this ceiling is unlikely to matter either way.
Output limits matter specifically when you need one long response in a single pass — an extended piece of writing, a large block of generated code — rather than a short reply or a back-and-forth exchange.
What this page won’t tell you
This page compares published price and capacity only. It says nothing about which model reasons more reliably, writes more naturally, or handles ambiguous instructions better — those questions are genuinely contested among people who test these models closely, and the answers shift with every update from either company. For that kind of comparison, independent public evaluation leaderboards are a better source, and they’re built to track exactly that kind of change over time, which a static page cannot do responsibly.
For the current figures side by side, including other models beyond these two, use the interactive comparison tool.