| Your current API costs | ||||
|---|---|---|---|---|
| Provider & model | ||||
| Provider | ||||
| Model | ||||
| Input price | $/1M tok | |||
| Output price | $/1M tok | |||
| Performance | ||||
| Volume | ||||
| Requests/day | req/day | |||
| Avg input tok/req | tokens | |||
| Avg output tok/req | tokens | |||
| Calculation | ||||
| Daily input tok | = | |||
| Daily output tok | = | |||
| Mo. input tokens | = | |||
| Mo. output tokens | = | |||
| Mo. input cost | = | |||
| Mo. output cost | = | |||
| Mo. API total | = | |||
| Annual API | = | |||
| Requirements | |
|---|---|
| Data on your servers | |
| Latency | |
| Self-hosting costs | ||||
|---|---|---|---|---|
| Open-weight model | ||||
| Model | ||||
| Quantization | ||||
| Use case | ||||
| VRAM needed | ||||
| Performance | ||||
| GPU | ||||
| GPU model | ||||
| Number of GPUs | GPUs | |||
| Hosting | ||||
| GPU cost | ||||
| Hourly rate | $/hr | |||
| Hours/day | hrs | |||
| Maintenance | ||||
| Maintained by | ||||
| Contractor /hr | $/hr | |||
| Hours/month | hrs/mo | |||
| Calculation | ||||
| Mo. GPU cost | = | |||
| Mo. maintenance | = | |||
| Mo. self-host total | = | |||
| Annual self-host | = | |||
Benchmark comparison matrix
[V] = vendor self-reported.
[I] = independently measured (Vals AI, Artificial Analysis, Open LLM Leaderboard).
MMLU and GSM8K no longer distinguish between top models. GPQA Diamond and SWE-bench show the real differences.
| Model | MMLU | MMLU-Pro | GPQA (D) | HumanEval | MATH/GSM8K | Tier |
|---|---|---|---|---|---|---|
| Proprietary API models | ||||||
| Claude Opus 5 | N/A | 91.6 [I] | 93.4 [I] | N/A | N/A | frontier |
| GPT-5.6 Sol | N/A | N/A | 95.2 [I] | N/A | N/A | frontier |
| GPT-4o | 88.7 [V] | 74.7 [I] | 53.6 [V] | 90.2 [V] | 76.6/90.5 | legacy |
| GPT-4o Mini | 82.0 [V] | 64.0 [I] | 40.0 [I] | 87.2 [V] | 68.0/88.0 | budget |
| Gemini 2.5 Flash | ~88 [I] | 80 [I] | 81 [V] | N/A | AIME 73.3 | balanced |
| Haiku 4.5 | N/A | ~78 [I] | ~66 [I] | N/A | N/A | budget |
| Open-weight models | ||||||
| Kimi K3 (2.8T MoE) | N/A | N/A | 93.5 [V] | N/A | N/A | ~Opus 4.8 [V] |
| DeepSeek R1 | 90.8 [V] | 84.0 [V] | 71.5 [V] | N/A | 97.3/-- | ~Flash reasoning |
| DeepSeek V3 | 88.5 [V] | 75.9 [V] | 59.1 [V] | 82.6 [I] | 90.2/-- | ~GPT-4o |
| Llama 3.1 405B | 87.3 [V] | 73.3 [I] | 50.7 [V] | 89.0 [V] | 70.3/96.8 | ~GPT-4o |
| Qwen 2.5 72B | 86.1 [V] | 64 [V] | 49.1 [I] | 86 [V] | 83.1/95.8 | >GPT-4o Mini |
| Qwen 3 32B | ~85 [V] | 71.9 [I] | 54.6 [I] | -- [V] | 62/-- | ~4o Mini/Flash |
| Qwen 3 14B | 84.7 [V] | 70.5 [V] | 51.4 [V] | -- [V] | --/-- | >GPT-4o Mini |
| Llama 3.1 70B | 86.0 [V] | 66.4 [I] | 41 [I] | 80.5 [V] | 68/95 | ~GPT-4o knowl. |
| Phi-3 Medium | 78.0 [V] | ~40 est | low | 55 [V] | 53.1/91 | ~Mini (math) |
| Gemma 2 27B | 75.2 [V] | 49.4 [V] | 26.3 [V] | 51.2 [V] | 42.1/74.6 | ~Mini (pref.) |
| Gemma 2 9B | 71.3 [V] | 43.7 [V] | 24.8 [V] | 40.2 [V] | 36.4/68.6 | ~Haiku (chat) |
| Llama 3.1 8B | 69.4 [V] | 38 [I] | 29 [I] | 72.6 [V] | --/84.5 | ~Haiku/Mini |
| Phi-3 Mini | 68.8 [V] | ~30 est | low | 58 [V] | 41.3/82.5 | ~Haiku (math) |
| Mixtral 8x7B | 70.5 [V] | ~40 [I] | ~30 est | 40 [V] | 28.4/74 | GPT-3.5 |
| Mistral 7B | 62.5 [V] | ~30 est | ~28 est | 40 [V] | --/45 | GPT-3.5 |
Disclaimer: This is an estimation tool, not financial or technical advice.
All prices, benchmarks, and performance figures are approximate and change frequently.
Verify current pricing directly with each provider before making decisions.
If you spot an error or outdated number, please let us know at contact@cirqwit.com.
If you spot an error or outdated number, please let us know at contact@cirqwit.com.
