100k input tokens and 50k output tokens for a content batch.
Llama calculator
Llama Token Calculator
Estimate hosted Llama 4 token spend while keeping the distinction between open model weights and inference-provider pricing clear.
Prices verified: via OpenRouter
Convert tokens to cost
Use presets, share the exact inputs, and scan the live breakdown.
Examples
2,000 prompt tokens and a compact 300-token answer.
Llama 4 Scout: $0.10 input / $0.30 output per 1M tokens.
How this Llama Token Calculator works
Llama models do not have one universal API price
Meta releases Llama as open model weights, so there is no single Meta-direct per-token tariff that applies to every deployment. This page follows AP4's OpenRouter feed and currently offers Llama 4 Scout at $0.1 input and $0.3 output per 1M tokens and Llama 4 Maverick at $0.2 input and $0.8 output per 1M tokens. Scout is the efficiency tier with fewer experts and exceptionally long-context design; Maverick activates the same number of parameters per token but has a much larger expert pool for higher-capacity multimodal work. The displayed rates describe OpenRouter's managed route, not the cost of self-hosting or a promise that every upstream Llama host charges the same amount.
Scout and Maverick have different output multipliers
Hosted Scout output is priced at three times its input rate. Maverick's output rate is four times its input rate, so the gap between the two models widens as responses get longer. A search or classification task with a large prompt and a tiny label can keep Scout and Maverick relatively close in absolute dollars. A writing assistant, code generator, or visual analysis flow that emits extensive text gives Maverick's higher output meter much more weight. Count conversation history and retrieved context as input, then use actual completion usage for output; combining them into one token total hides the most important pricing difference in this family.
A hosted estimate is different from self-hosting economics
Running Llama weights on your own GPUs replaces a simple token tariff with hardware rental or purchase, idle capacity, batching efficiency, engineering time, and operations. A saturated deployment can beat hosted per-token pricing, while a lightly used endpoint can be far more expensive because GPUs accrue cost while idle. Quantization and provider-specific serving stacks also change throughput and quality. This calculator is most useful when consuming the synchronized OpenRouter route or comparing a hosted baseline against a self-hosted quote. It does not pretend to turn GPU-hours into tokens, and it excludes gateway features, image processing differences, caching, and provider-specific minimums.
When Llama is cheap—and when the comparison flips
Llama 4 Scout is attractive for high-volume assistants, multilingual processing, and long input workloads where an open model meets the quality bar. Maverick is the more expensive hosted choice, but it can justify its premium for more demanding multimodal or general-purpose tasks. The calculation can flip when a host changes, when strict data residency narrows the provider pool, or when self-hosted utilization is low. Compare the exact endpoint you will buy, not an abstract Llama average. Against premium proprietary models, hosted Scout is often inexpensive; against smaller specialized open models or a well-utilized internal cluster, even Scout may not be the cheapest operational answer.
Examples
Llama 4 Maverick · 1,000,000 input and 250,000 output tokens. From the central pricing store ($0.2 input and $0.8 output per 1M tokens), input costs $0.2, output costs $0.2, and the total is $0.4.
FAQ
Does Meta set the Llama token prices shown here?
No. Meta publishes the Llama model weights, while hosting companies set inference prices. These figures come from the central OpenRouter-synced route and are labeled as hosted estimates.
Does this calculate the cost of self-hosting Llama?
No. Self-hosting depends on GPU-hours, utilization, batching, quantization, power, and engineering overhead. Use the hosted result as a comparison point for a separate infrastructure model.
Which is cheaper, Llama 4 Scout or Maverick?
Scout has the lower synchronized input and output rates. Maverick costs more but may be worthwhile for workloads that benefit from its larger expert pool and higher-capacity multimodal performance.