Skip to content
QUICK ESTIMATE · ~30 SECONDS

Size your LLM deployment.

Pick a model, tell us your load. The estimate fills in on the right — tap any tile to flip and see the math.

01

What model are you serving?

SET

Sliding gallery of common models, or paste any HuggingFace model ID.

Selected: Gemma 2 27B · 27B params
OR PASTE A HUGGINGFACE MODEL ID
02

Load profile

SET

Tap the pencil to adjust either value.

CONCURRENT USERS
10
custom
CONVERSATION LENGTH
8K
short
DEPLOYMENT TYPE
LIVE ESTIMATEtap a tile to see the math
ESTIMATED MONTHLYON-PREM
$3,672
4 × H100 · $1.26/hr · 730 hrs/mo
↺ SEE MATH
MONTHLY COST MATH
4 × $1.26/gpu-hr × 730 hrs
= $4k / month
On-prem pricing · amortized 36 months + electricity
↺ FLIP BACK
GPU COUNT
4× H100
42% utilized
↺ SEE MATH
GPU COUNT MATH
[134 ÷ 80GB]
= 4 GPUs
↺ FLIP BACK
TOTAL VRAM
134GB
model + KV cache
↺ SEE MATH
VRAM MATH
Weights: 54.0 GB
KV: 69.5 GB
= 134.3 GB
↺ FLIP BACK
MEMORY LAYOUT4 × H100 80GB
WeightsKV cacheFree
↺ SEE MATH
PER-GPU SPLIT
13.5 GB weights
17.4 GB KV cache
46.4 GB free headroom
Each user adds ~1.7 GB of KV cache.
↺ FLIP BACK
GPU