QUICK ESTIMATE · ~30 SECONDS
Size your LLM deployment.
Pick a model, tell us your load. The estimate fills in on the right — tap any tile to flip and see the math.
01
✓ SETWhat model are you serving?
Sliding gallery of common models, or paste any HuggingFace model ID.
Google
Gemma 2 2B
2B
Google
Gemma 2 9B
9B
✓
Google
Gemma 2 27B
27B
Google
Gemma 3 12B
12BNEW
Google
Gemma 3 27B
27BNEW
Google
Gemma 4 2B
2BNEW
Google
Gemma 4 9B
9BNEW
Google
Gemma 4 27B
27BNEW
Meta
Llama 3 8B
8B
Meta
Llama 3 70B
70B
Meta
Llama 3.1 8B
8B
Meta
Llama 3.1 70B
70B
Mistral
Mistral 7B
7B
Mistral
Mixtral 8x7B
47B
Qwen
Qwen 2.5 7B
7B
RedHat
RedHat Model 7B
7B
NVIDIA
Nemotron 340B
340B
✓Selected: Gemma 2 27B · 27B params
OR PASTE A HUGGINGFACE MODEL ID
02
✓ SETLoad profile
Tap the pencil to adjust either value.
CONCURRENT USERS
10
custom
CONVERSATION LENGTH
8K
short
DEPLOYMENT TYPE