How much VRAM do I need? The arithmetic, then the card.
A 70B parameter model needs about 145 GB of VRAM to serve at BF16, 75 GB at FP8 and 40 GB at INT4, and GPU Rent Hub rents a dedicated card for each of those: $240 a month for the INT4 case on a 48 GB A40, $1,490 for FP8 on an 80 GB H100 PCIe, $1,590 for BF16 on a 192 GB MI300X, paid in Bitcoin, Ethereum, USDT, USDC, Monero, Solana or Litecoin with no identity verification. Every number below is arithmetic you can redo. None of it is a benchmark, because we will not quote one we cannot source.
How do you calculate how much VRAM a model needs?
Three terms. Weights, KV cache, runtime overhead. Two of them you can compute exactly.
The sizing block
VRAM required = (parameters × bytes per parameter) + KV cache + about 2 GB of runtime overhead. Bytes per parameter is 4 at FP32, 2 at BF16 or FP16, 1 at FP8, and 0.5 at INT4. The KV cache is 2 × layers × KV heads × head dimension × bytes per element × context length × batch size, where the leading 2 covers the key and the value. Runtime overhead is the CUDA context, the framework and allocator fragmentation, and 2 GB is a safe allowance for a single process.
Worked once: a 70B model in FP8 at 8,192 tokens and batch 1 is 70 GB of weights, 2.7 GB of KV cache and 2 GB of overhead, so 75 GB, which fits one 80 GB card with 5 GB spare.
Two conventions, fixed before you argue about the last gigabyte.
Parameters in billions times bytes per parameter gives gigabytes directly, on the decimal convention where a gigabyte is 109 bytes. Card capacities are quoted the same way. Tooling that reports gibibytes instead shrinks every figure here by about 7%, which is the reason to leave headroom rather than plan to the last gigabyte.
Size on resident parameters, not active ones. A mixture-of-experts model keeps every expert in VRAM and runs two per token: the memory of the whole, the speed of the part.
| Precision | Bytes/param | 7B | 70B |
|---|---|---|---|
| FP32rarely worth it for inference | 4 | 28 GB | 280 GB |
| BF16 / FP16the reference point | 2 | 14 GB | 140 GB |
| FP8Ada, Hopper, Blackwell, CDNA 3 | 1 | 7 GB | 70 GB |
| INT4 / NF4weight-only, any architecture | 0.5 | 3.5 GB | 35 GB |
How much VRAM do I need to run a 7B, 13B, 34B, 70B or 405B model?
Weights plus a KV cache at 8,192 tokens and batch 1, plus 2 GB. Change the context and the KV column moves, not the weights.
| Model | KV at 8k, batch 1 | BF16 total | FP8 total | INT4 total |
|---|---|---|---|---|
| 7B32 layers, GQA, 8 KV heads | 1.1 GB | 17 GB | 10 GB | 7 GB |
| 13B40 layers, MHA, no GQA | 6.7 GB | 35 GB | 22 GB | 15 GB |
| 34B60 layers, GQA, 8 KV heads | 2.0 GB | 72 GB | 38 GB | 21 GB |
| 70B80 layers, GQA, 8 KV heads | 2.7 GB | 145 GB | 75 GB | 40 GB |
| 405B126 layers, GQA, 8 KV heads | 4.2 GB | 816 GB | 411 GB | 209 GB |
The 13B row is the interesting one: it carries a larger KV cache than the 70B below it. Its shape uses multi-head attention with 40 key and value heads, where later models use grouped-query attention and 8. That change cut the cache five to eightfold and is why long-context serving became affordable. Sizing anything from before 2024, check the attention scheme before the parameter count.
Batch size multiplies the KV cache and nothing else. That 70B at batch 32 adds 85 GB of cache to the same 70 GB of weights, which is why an endpoint needs far more memory than one request at a time. Service sizing is on the GPU for LLM inference page.
How much VRAM does the KV cache use at long context?
Linearly with tokens and linearly with batch. At 128k the cache can cost more than the weights.
| Model shape | KV per token | 8k | 32k | 128k |
|---|---|---|---|---|
| 7B, 32 layers, GQAKV width 1,024 | 0.13 MB | 1.1 GB | 4.3 GB | 17.2 GB |
| 13B, 40 layers, MHAKV width 5,120 | 0.82 MB | 6.7 GB | 26.8 GB | 107.4 GB |
| 34B, 60 layers, GQAKV width 1,024 | 0.25 MB | 2.0 GB | 8.1 GB | 32.2 GB |
| 70B, 80 layers, GQAKV width 1,024 | 0.33 MB | 2.7 GB | 10.7 GB | 42.9 GB |
| 405B, 126 layers, GQAKV width 1,024 | 0.52 MB | 4.2 GB | 16.9 GB | 67.6 GB |
Where this bites.
A 70B at FP8 fits an 80 GB card at 8k with 5 GB spare. Push it to 32k and the cache alone is 10.7 GB, so it stops fitting. That is why the H200 SXM at 141 GB and the MI300X at 192 GB exist as products. They hold more context than an H100 SXM, which is a different thing from being faster.
How much VRAM do I need to fine-tune a model?
Full fine-tuning costs roughly 16 bytes per parameter. LoRA costs 2. QLoRA costs 0.5. The gap between those three numbers is the whole decision.
Full fine-tuning with Adam holds five things per parameter: BF16 weights (2 bytes), BF16 gradients (2), an FP32 master copy (4), and the Adam first and second moments (4 each). That is 16 bytes per parameter before activations. LoRA freezes the base, so you pay 2 bytes per parameter for the frozen model and optimiser state only on the adapter, typically half a percent of the parameters. QLoRA quantises that frozen base to 4 bits, so it costs 0.5 bytes per parameter and the adapter is unchanged.
| Model | Full FT, weights + Adam | Full FT, with activations | LoRA | QLoRA |
|---|---|---|---|---|
| 7B | 112 GB | 129 GB | 21 GB | 10 GB |
| 13B | 208 GB | 239 GB | 33 GB | 14 GB |
| 34B | 544 GB | 626 GB | 79 GB | 28 GB |
| 70B | 1,120 GB | 1,288 GB | 156 GB | 51 GB |
| 405B | 6,480 GB | 7,452 GB | 856 GB | 249 GB |
The full fine-tuning column is the naive figure, everything resident on the GPUs. ZeRO stage 3 with optimiser state offloaded to host memory cuts the GPU-side requirement substantially, and our 8-way nodes carry 2 TB of host RAM for exactly that. Offload trades PCIe bandwidth for VRAM, so it makes a run possible rather than fast. Method choice is on the GPU for fine-tuning page.
Read the last two columns against the first. A 70B full fine-tune wants an eight-card node at $11,450. The same 70B under QLoRA wants 51 GB, which is one A100 SXM 80GB at $990 a month. That is a twelvefold difference for a method that is worse but rarely twelve times worse. Most people asking this question have been quoted the full fine-tuning number and do not need it.
What GPU do I need for a 70B model, or any other size?
Model size down the side, method across the top. Each cell is the VRAM required and the cheapest thing in our fleet that holds it.
| Model | Inference BF16 | Inference FP8 | Inference INT4 | QLoRA | LoRA | Full fine-tune |
|---|---|---|---|---|---|---|
| 7B | 17 GBRTX A5000 $190 | 10 GBL4 $250 | 7 GBRTX A5000 $190 | 10 GBRTX A5000 $190 | 21 GBRTX A5000 $190, tight | 129 GBMI300X $1,590 |
| 13B | 35 GBA40 $240 | 22 GBL4 $250, tight | 15 GBRTX A5000 $190 | 14 GBRTX A5000 $190 | 33 GBA40 $240 | 239 GB4× A100 $3,643 |
| 34B | 72 GBA100 SXM $990 | 38 GBRTX 6000 Ada $460 | 21 GBRTX A5000 $190, tight | 28 GBA40 $240 | 79 GBA100 SXM $990, tight | 626 GB8× A100 $7,130, tight |
| 70B | 145 GBMI300X $1,590 | 75 GBH100 PCIe $1,490 | 40 GBA40 $240 | 51 GBA100 SXM $990 | 156 GBMI300X $1,590 | 1,288 GB8× MI300X $11,450 |
| 405B | 816 GB8× MI300X $11,450 | 411 GB8× MI300X $11,450 | 209 GB8× A40 $1,728 | 249 GB8× A40 $1,728 | 856 GB8× MI300X $11,450 | 7,452 GBfive nodes, quoted |
The cell to look at is 70B at INT4: 40 GB, held by an A40 48GB at $240 a month. Not a trick, just what four-bit quantisation and 48 GB of GDDR6 buy. Section 07 sets out what you gave up for it. Anything above one 8-way node is a cluster question, not a card question, and belongs on the multi-GPU cluster rental page.
Which GPU gives you the most VRAM per dollar?
Monthly price divided by gigabytes of VRAM, and memory bandwidth divided by monthly price. Nobody publishes either. Both decide more purchases than a TFLOPS figure does.
| GPU | VRAM | Bandwidth | Monthly | $ per GB/mo | GB/s per $/mo | Per GPU-hour |
|---|---|---|---|---|---|---|
| MI300XHBM3, CDNA 3 | 192 GB | 5.3 TB/s | $1,590/mo | $8.28 | 3.33 | $2.21 |
| B200HBM3e, Blackwell, 2-GPU minimum | 180 GB | 8 TB/s | $3,890/mo | $21.61 | 2.06 | $5.40 |
| H200 SXMHBM3e, Hopper | 141 GB | 4.8 TB/s | $2,290/mo | $16.24 | 2.10 | $3.18 |
| H100 SXMHBM3, NVLink 900 GB/s | 80 GB | 3.35 TB/s | $1,690/mo | $21.13 | 1.98 | $2.35 |
| H100 PCIeHBM2e | 80 GB | 2 TB/s | $1,490/mo | $18.63 | 1.34 | $2.07 |
| A100 SXMHBM2e, no FP8 | 80 GB | 2.04 TB/s | $990/mo | $12.38 | 2.06 | $1.38 |
| L40SGDDR6 ECC, FP8 | 48 GB | 864 GB/s | $520/mo | $10.83 | 1.66 | $0.72 |
| RTX 6000 AdaGDDR6 ECC, FP8 | 48 GB | 960 GB/s | $460/mo | $9.58 | 2.09 | $0.64 |
| RTX A6000GDDR6 ECC, no FP8 | 48 GB | 768 GB/s | $290/mo | $6.04 | 2.65 | $0.40 |
| A40GDDR6 ECC, no FP8 | 48 GB | 696 GB/s | $240/mo | $5.00 | 2.90 | $0.33 |
| RTX 5090GDDR7, FP8 | 32 GB | 1.79 TB/s | $540/mo | $16.88 | 3.31 | $0.75 |
| RTX 4090GDDR6X, FP8 | 24 GB | 1.01 TB/s | $390/mo | $16.25 | 2.59 | $0.54 |
| L4GDDR6, FP8, 72 W | 24 GB | 300 GB/s | $250/mo | $10.42 | 1.20 | $0.35 |
| RTX 3090GDDR6X, no FP8 | 24 GB | 936 GB/s | $230/mo | $9.58 | 4.07 | $0.32 |
| RTX A5000GDDR6 ECC, no FP8 | 24 GB | 768 GB/s | $190/mo | $7.92 | 4.04 | $0.26 |
Two different cards win the two columns, which is why both are published. The A40 holds a gigabyte for $5.00 a month, less than half what an L40S charges for the same 48 GB, because it is a 2020 card with 696 GB/s behind it. The RTX 3090 gives 4.07 GB/s per dollar, more than twice the H100 SXM, because it is an old consumer card with a fast bus and a low rent. Neither is the best card here. Both answer a specific question correctly.
The MI300X is the outlier. At $8.28 per gigabyte it is cheaper memory than an A100, an H100 or an H200, with 5.3 TB/s behind it, which is why it fills the large-model cells above. The cost is ROCm rather than CUDA: vLLM and PyTorch run, a hand-written CUDA kernel does not.
Is more VRAM always the right answer?
No. Three cases where buying capacity is the wrong purchase, with the arithmetic for each.
Bandwidth is the bottleneck, not capacity
Decoding one token reads every resident weight from VRAM once. Divide bandwidth by the model's resident size for a hard ceiling on single-stream tokens per second. A 7B at BF16 is 14 GB, so an A40 at 696 GB/s cannot pass 50 tokens per second and an RTX 5090 at 1.79 TB/s cannot pass 128. Both fit the model. One is 2.6 times faster at it.
Quantised on a small card beats full precision on a big one
A 34B at BF16 is 72 GB all in and needs an A100 at $990: 2.04 TB/s over its 68 GB of weights is a ceiling of 30 tokens per second. The same 34B at INT4 is 21 GB all in and fits an RTX 4090 at $390: 1.01 TB/s over its 17 GB of weights is a ceiling of 59. Faster, and 61% cheaper. What you gave up is quality on the quantised weights, which you evaluate rather than assume.
Two cheap cards, sometimes
Two A40 give 96 GB for $456 a month against one A100's 80 GB for $990. More memory, less than half the price. It works for batch inference with a replica per card, for QLoRA, and for layer-split serving at low concurrency. It fails under tensor parallelism and on any training step with an all-reduce, because the link between the cards is the part you did not buy.
The NVLink point, stated precisely.
An H100 SXM talks to its neighbours at 900 GB/s over NVLink. An A40 or RTX A6000 pair takes an NVLink bridge at 112 GB/s. An L40S has no NVLink and moves data over PCIe Gen4 peer-to-peer at a line rate of roughly 32 GB/s. An RTX 5090 has no NVLink either, but its link is Gen5, so roughly 63 GB/s. That gap is paid once per layer per token in a tensor-parallel run.
Splitting by layers instead of by tensors avoids most of that traffic and raises nothing. Per token you still read every byte of the model, on two cards in sequence rather than one alone. Tensor parallelism is what turns two cards into double the bandwidth, and it is exactly the mode that needs the fast link. That is the whole trade.
Worked comparison
- 2× RTX 5090
- 64 GB, 3.58 TB/s aggregate, $1,026/mo, PCIe Gen5 peer-to-peer
- 1× H100 SXM
- 80 GB, 3.35 TB/s, $1,690/mo, one memory space
- Verdict
- The pair wins on paper and loses on any model that does not fit in 32 GB per card
The cases where this page should send you elsewhere.
Our shortest term is seven days. If you need a card for an afternoon to check whether a 13B fits in 24 GB, use an hourly provider: a few hours there costs less than the $55 minimum week on an RTX A5000 here. We do not sell hours and have no plan to.
If you are choosing between two cards on speed rather than fit, stop reading tables and run your stack on both. Everything here is capacity arithmetic and a bandwidth ceiling. Neither predicts what your kernels, batch shape and scheduler will do, and we will not pretend otherwise with a throughput figure we cannot source.
If you need 6,480 GB for a 405B full fine-tune, you need a cluster, a fabric and a plan, not a rental page. Ask sales about multi-node capacity in DFW-1 Dallas, IAD-1 Ashburn or PDX-1 Hillsboro and expect a conversation about interconnect before a price.
- Learning, any model up to 13B
- RTX A5000 24 GB, $190/mo
- Serving 24B to 34B in production
- L40S 48 GB, $520/mo
- Fastest single card under $600
- RTX 5090 32 GB, $540/mo
- Cheapest 70B at INT4
- A40 48 GB, $240/mo
- 70B at BF16 on one card
- MI300X 192 GB, $1,590/mo
- Long context on 70B
- H200 SXM 141 GB, $2,290/mo
Deposit once, deploy in about ninety seconds, no email address and no identity document. See no KYC GPU hosting and how crypto payment works.
VRAM questions, asked the way people ask them.
How much VRAM do I need to run a 70B model?
How much VRAM do I need for a 7B model?
How do I calculate the KV cache size?
Is 24 GB of VRAM enough to fine-tune a model?
What GPU do I need for a 405B model?
Do two 24 GB cards work like one 48 GB card?
Which GPU has the cheapest VRAM per gigabyte?
Does quantising to INT4 hurt output quality?
Workload-specific sizing on GPU for LLM inference and GPU for fine-tuning. Full specs, stock and prices on the GPU catalogue.
You know the number. Now pick the card.
Fifteen models in DFW-1, IAD-1 and PDX-1, from $190 a month, weekly or monthly, paid in seven coins with no identity check.