Free Tool
LLM Model Cost Comparator
Compare costs across 10 models from 6 providers. Adjust token volume with the sliders — rankings update instantly. All calculations run in your browser with zero API calls.
Configure your usage
100K50M
25K12.5M
Choosing Groq Llama 3.1 8B Instant over Anthropic Claude Sonnet 4
saves 99%
$0.07/mo vs $6.75/mo at your volume
| Rank | Model | Provider | Input Cost | Output Cost | Total/mo |
|---|---|---|---|---|---|
| 1 | Llama 3.1 8B Instant | Groq | $0.05 | $0.02 | $0.07 |
| 2 | Gemini 2.0 Flash | $0.10 | $0.10 | $0.20 | |
| 3 | DeepSeek V4 Flash | DeepSeek | $0.14 | $0.14 | $0.28 |
| 4 | GPT-4o Mini | OpenAI | $0.15 | $0.15 | $0.30 |
| 5 | MiniMax M2.7 | MiniMax | $0.30 | $0.30 | $0.60 |
| 6 | Llama 3.3 70B Versatile | Groq | $0.59 | $0.20 | $0.79 |
| 7 | DeepSeek V4 Pro | DeepSeek | $0.55 | $0.55 | $1.10 |
| 8 | Claude Haiku 3.5 | Anthropic | $0.80 | $1.00 | $1.80 |
| 9 | GPT-4o | OpenAI | $2.50 | $2.50 | $5.00 |
| 10 | Claude Sonnet 4 | Anthropic | $3.00 | $3.75 | $6.75 |
How model routing reduces cost further
This comparison assumes you use one model for everything. In practice, our Inference Economics engagement applies intent routing — simple tasks go to the cheapest model, complex tasks go to the most capable. This typically reduces total cost by an additional 50-85% beyond these single-model estimates. Learn about Inference Economics →