100k input tokens and 50k output tokens for a content batch.
Mistral calculator
Mistral Token Calculator
Compare Mistral's current general-purpose tiers with separate meters for prompts and generated text.
Prices verified: via OpenRouter
Convert tokens to cost
Use presets, share the exact inputs, and scan the live breakdown.
Examples
2,000 prompt tokens and a compact 300-token answer.
Mistral Small 4: $0.15 input / $0.60 output per 1M tokens.
How this Mistral Token Calculator works
Mistral's tiers are not a simple size ladder
Mistral Small 4 is the low-cost general tier at $0.15 input and $0.6 output per 1M tokens. Mistral Large 3, the open-weight flagship, is listed at $0.5 input and $1.5 output per 1M tokens. Mistral Medium 3.5 is the premium managed option at $1.5 input and $7.5 output per 1M tokens, aimed at reliable multi-tool work, coding, and complex agentic tasks. The names can be counterintuitive: Medium 3.5 currently costs more per token than Large 3. Select by endpoint capability and measured output quality, not the adjective alone. Small is suited to frequent production calls, Large offers a capable middle price, and Medium is the deliberate upgrade when its behavior earns the premium.
Output-heavy work changes the Mistral ranking
The output premium varies across this lineup. Large 3 output is three times its input rate, Small 4 output is four times input, and Medium 3.5 output is five times input. That asymmetry makes generated length unusually important on Medium: a compact prompt that launches a long coding or agent response can spend far more on completion than context. For extraction, classification, and routing, input dominates volume and Small remains especially economical. For drafting and code generation, cap unnecessary verbosity and measure actual completion tokens. The calculator keeps both sides separate so a model does not look cheap merely because its prompt rate is attractive.
Caching and batch discounts are intentionally excluded
Mistral advertises a substantial discount for cached input and a separate reduction for batch processing. Those modes depend on request structure and latency tolerance, so the central store records standard real-time, uncached text rates as the auditable baseline. A stable system prompt or shared document prefix can make caching valuable; offline evaluation, enrichment, and migration jobs may fit the Batch API. Neither discount should be applied blindly to interactive traffic. Multimodal media, OCR, audio, Agents API tools, regional inference endpoints, and enterprise uplifts have their own meters as well. Add them from production usage after estimating the core text model cost here.
When Mistral is cheap and when Medium gets expensive
Small 4 is the cheap Mistral choice for bulk multilingual text, lightweight agents, and applications that need a capable open model at high volume. Large 3 can be the value tier when Small misses quality requirements but Medium's managed premium is unnecessary. Medium 3.5 becomes expensive on long answers because its output rate is the highest in this group; use it where dependable multi-step execution or tool calling prevents failures elsewhere. For latency-tolerant jobs, batch processing may improve the comparison. For teams that can operate open weights efficiently, self-hosting Small or Large changes the cost model entirely and is not represented by a per-token API estimate.
Examples
Mistral Medium 3.5 · 1,000,000 input and 200,000 output tokens. From the central pricing store ($1.5 input and $7.5 output per 1M tokens), input costs $1.5, output costs $1.5, and the total is $3.
FAQ
Why does Mistral Medium 3.5 cost more than Large 3?
The names describe different product generations and deployment goals, not a permanent ascending price scale. Medium 3.5 is positioned as a premium agentic and multi-tool model in the current API lineup.
Does the Mistral calculator apply the Batch API discount?
No. It uses standard real-time rates. Batch processing is a separate, latency-tolerant mode and should be modeled only when the workload can actually use it.
Are Mistral cached input tokens included?
The estimate treats input as uncached. Mistral discounts eligible cached input, so a measured cache-heavy workload can cost less than the baseline shown here.