Thoughtful engineering.

Should you self-host your AI?

All prices pre-filled with verified August 2026 rates. Edit any cell.
Your current API costs
Provider & model
Provider
Model
Input price$/1M tok
Output price$/1M tok
Performance
Volume
Requests/dayreq/day
Avg input tok/reqtokens
Avg output tok/reqtokens
Calculation
Daily input tok=
Daily output tok=
Mo. input tokens=
Mo. output tokens=
Mo. input cost=
Mo. output cost=
Mo. API total=
Annual API=
Requirements
Data on your servers
Latency
Self-hosting costs
Open-weight model
Model
Quantization
Use case
VRAM needed
Performance
GPU
GPU model
Number of GPUsGPUs
Hosting
GPU cost
Hourly rate$/hr
Hours/dayhrs
Maintenance
Maintained by
Contractor /hr$/hr
Hours/monthhrs/mo
Calculation
Mo. GPU cost=
Mo. maintenance=
Mo. self-host total=
Annual self-host=
Benchmark comparison matrix

[V] = vendor self-reported.

[I] = independently measured (Vals AI, Artificial Analysis, Open LLM Leaderboard).

MMLU and GSM8K no longer distinguish between top models. GPQA Diamond and SWE-bench show the real differences.

ModelMMLUMMLU-ProGPQA (D)HumanEvalMATH/GSM8KTier
Proprietary API models
Claude Opus 5N/A91.6 [I]93.4 [I]N/AN/Afrontier
GPT-5.6 SolN/AN/A95.2 [I]N/AN/Afrontier
GPT-4o88.7 [V]74.7 [I]53.6 [V]90.2 [V]76.6/90.5legacy
GPT-4o Mini82.0 [V]64.0 [I]40.0 [I]87.2 [V]68.0/88.0budget
Gemini 2.5 Flash~88 [I]80 [I]81 [V]N/AAIME 73.3balanced
Haiku 4.5N/A~78 [I]~66 [I]N/AN/Abudget
Open-weight models
Kimi K3 (2.8T MoE)N/AN/A93.5 [V]N/AN/A~Opus 4.8 [V]
DeepSeek R190.8 [V]84.0 [V]71.5 [V]N/A97.3/--~Flash reasoning
DeepSeek V388.5 [V]75.9 [V]59.1 [V]82.6 [I]90.2/--~GPT-4o
Llama 3.1 405B87.3 [V]73.3 [I]50.7 [V]89.0 [V]70.3/96.8~GPT-4o
Qwen 2.5 72B86.1 [V]64 [V]49.1 [I]86 [V]83.1/95.8>GPT-4o Mini
Qwen 3 32B~85 [V]71.9 [I]54.6 [I]-- [V]62/--~4o Mini/Flash
Qwen 3 14B84.7 [V]70.5 [V]51.4 [V]-- [V]--/-->GPT-4o Mini
Llama 3.1 70B86.0 [V]66.4 [I]41 [I]80.5 [V]68/95~GPT-4o knowl.
Phi-3 Medium78.0 [V]~40 estlow55 [V]53.1/91~Mini (math)
Gemma 2 27B75.2 [V]49.4 [V]26.3 [V]51.2 [V]42.1/74.6~Mini (pref.)
Gemma 2 9B71.3 [V]43.7 [V]24.8 [V]40.2 [V]36.4/68.6~Haiku (chat)
Llama 3.1 8B69.4 [V]38 [I]29 [I]72.6 [V]--/84.5~Haiku/Mini
Phi-3 Mini68.8 [V]~30 estlow58 [V]41.3/82.5~Haiku (math)
Mixtral 8x7B70.5 [V]~40 [I]~30 est40 [V]28.4/74GPT-3.5
Mistral 7B62.5 [V]~30 est~28 est40 [V]--/45GPT-3.5
Anthropic pricing verified from official docs. Other pricing from secondary trackers (Finout, CloudZero, aipricing.guru, pricepertoken).
GPU rates: RunPod, Lambda, Vast.ai, Jarvislabs. Benchmarks: Trelis, Lyceum, CloudRift, DatabaseMart, Vals AI, Artificial Analysis.
All data as of August 2026.
Disclaimer: This is an estimation tool, not financial or technical advice. All prices, benchmarks, and performance figures are approximate and change frequently. Verify current pricing directly with each provider before making decisions.

If you spot an error or outdated number, please let us know at contact@cirqwit.com.
© 2026 Cirqwit  |  contact@cirqwit.com