AI model comparison: price and specs, side by side
Compare two AI models side by side — token cost at your own volumes, context window and output limits — from vendor-published figures. Not a performance ranking.
Full spec comparison
- Figures are what each vendor published on their official pricing/model page, captured on 2026-07-07. Prices change often — this dataset carries fast decay, so a stale copy fails Around’s build.
- Prices are per million tokens; the cost shown is simple arithmetic on the token volumes you enter. A token is a chunk of text, usually a word-piece — see the glossary.
- This tool compares published price and capacity, NOT performance, quality, or “which is better” — those are contested and out of scope.
- Around is built with Anthropic’s Claude and deliberately does not rank models. Values marked “not stated” were not published on the source page and are never guessed.
Cheapest models right now
Every model in the table, ranked by the cost of a one-million-token request in and one million out — the total of both published rates. Price only, as of 2026-07-07. Cheaper is not the same as better: this says nothing about quality, and Around (built with Anthropic’s Claude) deliberately does not rank models on capability.
| # | Model | Input /1M | Output /1M | 1M in + 1M out |
|---|---|---|---|---|
| 1 | Meta (Llama) Llama 3.1 8B Instant | $0.05 | $0.08 | $0.13 |
| 2 | Mistral Ministral 8B | $0.15 | $0.15 | $0.30 |
| 3 | DeepSeek DeepSeek-V4-Flash | $0.14 | $0.28 | $0.42 |
| 4 | Meta (Llama) Llama 4 Scout (17Bx16E) | $0.11 | $0.34 | $0.45 |
| 5 | Mistral Mistral Small 4 | $0.15 | $0.60 | $0.75 |
| 6 | Mistral Codestral | $0.30 | $0.90 | $1.20 |
| 7 | DeepSeek DeepSeek-V4-Pro | $0.435 | $0.87 | $1.305 |
| 8 | Meta (Llama) Llama 3.3 70B Versatile | $0.59 | $0.79 | $1.38 |
| 9 | Mistral Mistral Large 3 | $0.50 | $1.50 | $2.00 |
| 10 | Cohere Command R (03-2024) | $0.50 | $1.50 | $2.00 |
| 11 | Google Gemini 2.5 Flash | $0.30 | $2.50 | $2.80 |
| 12 | xAI grok-build-0.1 (grok-code-fast-1) | $1.00 | $2.00 | $3.00 |
| 13 | xAI Grok 4.3 | $1.25 | $2.50 | $3.75 |
| 14 | OpenAI GPT-5.4 mini | $0.75 | $4.50 | $5.25 |
| 15 | Anthropic Claude Haiku 4.5 | $1.00 | $5.00 | $6.00 |
| 16 | Mistral Magistral Medium | $2.00 | $5.00 | $7.00 |
| 17 | Mistral Mistral Medium 3.5 | $1.50 | $7.50 | $9.00 |
| 18 | Google Gemini 3.5 Flash | $1.50 | $9.00 | $10.50 |
| 19 | Anthropic Claude Sonnet 5 | $2.00 | $10.00 | $12.00 |
| 20 | Cohere Command R+ (08-2024) | $2.50 | $10.00 | $12.50 |
| 21 | Google Gemini 3.1 Pro Preview | $2.00 | $12.00 | $14.00 |
| 22 | Anthropic Claude Opus 4.8 | $5.00 | $25.00 | $30.00 |
| 23 | OpenAI GPT-5.5 | $5.00 | $30.00 | $35.00 |
| 24 | Anthropic Claude Fable 5 | $10.00 | $50.00 | $60.00 |
Not ranked (no published price for at least one direction): Google Gemini 2.5 Pro.
Open-weight models (e.g. Llama) are priced by a named host and vary by host; see the row’s source. Prices change often — the dataset behind this table is reviewed every 30 days, and a stale copy fails Around’s build.
Pick two models, enter how many tokens a typical request sends and receives, and this tool shows what each one would cost and how their published specifications line up. The calculator runs above; the sections below explain exactly what it is doing, and — just as importantly — what it deliberately does not do.
What it compares
Three kinds of number, and only these three: the price per million tokens each vendor charges for input and output, the size of the context window, and the maximum length of a single response. From the prices and the token volumes you enter, the tool multiplies out the cost of one request. That arithmetic is the whole calculation — there is no secret weighting, no adjustment, no house view baked in. If two models charge the same, they land in the same place.
Every figure carries the date it was read from the vendor’s own page. Where a vendor did not publish a number, the tool prints “not stated” rather than filling the gap with a guess. A blank is more honest than an invented figure, and a comparison built on guesses would be worth nothing.
What a token is, and why cost is measured in them
Models do not read words or characters; they read tokens — short chunks of text, usually a word or a fragment of one. A rule of thumb is that a thousand tokens is roughly seven or eight hundred English words, but it varies by language and content. Because a model’s work is counted in tokens, so is its price: you pay a rate per million tokens going in (your prompt) and a separate, usually higher, rate per million coming out (its reply). That is why the tool asks for input and output volumes separately — the split matters, and output is where the cost usually sits.
What it will not tell you
It will not tell you which model is “better”, and that is a deliberate line, not a missing feature. Benchmark scores and capability rankings are contested, they are cherry-picked by whoever is publishing them, and they date within weeks. There is a second reason we hold that line here in particular: Around is built using Anthropic’s Claude. A tool from us that ranked Claude against its competitors would be marking our own homework, and no amount of care would make that trustworthy. So this tool reports what the vendors themselves state — prices and capacities, dated and attributed — and stops there. The judgement about which model fits your task stays yours.
Trusting the numbers — and checking them
AI prices move faster than almost anything else Around tracks. New models arrive, rates are cut, and introductory prices lapse on a fixed date. Because of that, this dataset carries the shortest review cycle on the site: if the figures are not refreshed within a month, Around’s own build fails rather than serve you a stale price. Even so, treat the result as a well-sourced starting point, not a live quote — before you rely on a number for a real decision, check the current figure on the vendor’s pricing page. The tool links to the source behind every model.
Updated 7 July 2026