Skip to main content
PalCloud

Calculator · arithmetic shown

How much GPU memory does your model need?

The sum is simple — bytes per parameter × parameters, plus room for the KV cache when serving, or for gradients and optimiser state when training. This page shows every term so you can check it.
Precision
What are you running?

Estimated GPU memory needed

172.0 GB

  • Weights: 70B × 2 bytes = 140.0 GB
  • Runtime overhead: +20% of weights
  • KV cache allowance: 8 × ~0.50 GB = 4.0 GB

These are rules of thumb, not a guarantee: real usage depends on context length, batch size, attention implementation and whether you use LoRA, quantisation or offloading — any of which change the answer substantially.

How many GPUs that takes

GPUMemory eachGPUs neededPrices
GB300279 GB11 providers →
B300270 GB13 providers →
MI300X192 GB13 providers →
B200180 GB17 providers →
H200141 GB27 providers →
GH20096 GB21 providers →
RTX PRO 600096 GB25 providers →
H10080 GB315 providers →
L4048 GB41 providers →
L40S48 GB410 providers →
A10040 GB510 providers →
L424 GB83 providers →
Quadro RTX 600024 GB81 providers →
RTX 4000 Ada20 GB91 providers →

GPU memory figures come from each maker’s spec sheet, linked on the GPU page. Splitting a model across GPUs also costs some memory per GPU, so treat the count as a floor.

Compare GPU cloud prices → · What a run costs →

Know the GPU you need? Get quotes

Get Best Quotes