GPU Rent HubVRAM guide

How much VRAM do I need? The arithmetic, then the card.

A 70B parameter model needs about 145 GB of VRAM to serve at BF16, 75 GB at FP8 and 40 GB at INT4, and GPU Rent Hub rents a dedicated card for each of those: $240 a month for the INT4 case on a 48 GB A40, $1,490 for FP8 on an 80 GB H100 PCIe, $1,590 for BF16 on a 192 GB MI300X, paid in Bitcoin, Ethereum, USDT, USDC, Monero, Solana or Litecoin with no identity verification. Every number below is arithmetic you can redo. None of it is a benchmark, because we will not quote one we cannot source.

Updated 2026-09-01
2 BBytes per parameter at BF16
$5.00Cheapest VRAM we rent, per GB per month
192 GBMost VRAM on one card, MI300X
15Cards to choose between
01 The formula

How do you calculate how much VRAM a model needs?

Three terms. Weights, KV cache, runtime overhead. Two of them you can compute exactly.

The sizing block

VRAM required = (parameters × bytes per parameter) + KV cache + about 2 GB of runtime overhead. Bytes per parameter is 4 at FP32, 2 at BF16 or FP16, 1 at FP8, and 0.5 at INT4. The KV cache is 2 × layers × KV heads × head dimension × bytes per element × context length × batch size, where the leading 2 covers the key and the value. Runtime overhead is the CUDA context, the framework and allocator fragmentation, and 2 GB is a safe allowance for a single process.

Worked once: a 70B model in FP8 at 8,192 tokens and batch 1 is 70 GB of weights, 2.7 GB of KV cache and 2 GB of overhead, so 75 GB, which fits one 80 GB card with 5 GB spare.

Two conventions, fixed before you argue about the last gigabyte.

Parameters in billions times bytes per parameter gives gigabytes directly, on the decimal convention where a gigabyte is 109 bytes. Card capacities are quoted the same way. Tooling that reports gibibytes instead shrinks every figure here by about 7%, which is the reason to leave headroom rather than plan to the last gigabyte.

Size on resident parameters, not active ones. A mixture-of-experts model keeps every expert in VRAM and runs two per token: the memory of the whole, the speed of the part.

PrecisionBytes/param7B70B
FP32rarely worth it for inference428 GB280 GB
BF16 / FP16the reference point214 GB140 GB
FP8Ada, Hopper, Blackwell, CDNA 317 GB70 GB
INT4 / NF4weight-only, any architecture0.53.5 GB35 GB
FP8 needs hardware support. The A100, A40, RTX A6000, RTX 3090 and RTX A5000 do not have it. INT4 weight-only quantisation runs on all of them.
02 Inference

How much VRAM do I need to run a 7B, 13B, 34B, 70B or 405B model?

Weights plus a KV cache at 8,192 tokens and batch 1, plus 2 GB. Change the context and the KV column moves, not the weights.

ModelKV at 8k, batch 1BF16 totalFP8 totalINT4 total
7B32 layers, GQA, 8 KV heads1.1 GB17 GB10 GB7 GB
13B40 layers, MHA, no GQA6.7 GB35 GB22 GB15 GB
34B60 layers, GQA, 8 KV heads2.0 GB72 GB38 GB21 GB
70B80 layers, GQA, 8 KV heads2.7 GB145 GB75 GB40 GB
405B126 layers, GQA, 8 KV heads4.2 GB816 GB411 GB209 GB
KV cache held at FP16 in every column. Head dimension 128 throughout. Check the layer count and KV head count in your model's config.json before you trust a row, because these are typical shapes, not universal ones.

The 13B row is the interesting one: it carries a larger KV cache than the 70B below it. Its shape uses multi-head attention with 40 key and value heads, where later models use grouped-query attention and 8. That change cut the cache five to eightfold and is why long-context serving became affordable. Sizing anything from before 2024, check the attention scheme before the parameter count.

Batch size multiplies the KV cache and nothing else. That 70B at batch 32 adds 85 GB of cache to the same 70 GB of weights, which is why an endpoint needs far more memory than one request at a time. Service sizing is on the GPU for LLM inference page.

03 Context

How much VRAM does the KV cache use at long context?

Linearly with tokens and linearly with batch. At 128k the cache can cost more than the weights.

Model shapeKV per token8k32k128k
7B, 32 layers, GQAKV width 1,0240.13 MB1.1 GB4.3 GB17.2 GB
13B, 40 layers, MHAKV width 5,1200.82 MB6.7 GB26.8 GB107.4 GB
34B, 60 layers, GQAKV width 1,0240.25 MB2.0 GB8.1 GB32.2 GB
70B, 80 layers, GQAKV width 1,0240.33 MB2.7 GB10.7 GB42.9 GB
405B, 126 layers, GQAKV width 1,0240.52 MB4.2 GB16.9 GB67.6 GB
FP16 cache, batch 1. Halve every figure for an FP8 cache. Multiply by batch size. KV width is KV heads multiplied by head dimension.

Where this bites.

A 70B at FP8 fits an 80 GB card at 8k with 5 GB spare. Push it to 32k and the cache alone is 10.7 GB, so it stops fitting. That is why the H200 SXM at 141 GB and the MI300X at 192 GB exist as products. They hold more context than an H100 SXM, which is a different thing from being faster.

Quantising the KV cache to FP8 is a separate lever from quantising the weights, and it is usually the cheaper one. It halves the cache with far less quality cost than dropping weights another step, and vLLM and TensorRT-LLM both support it.
04 Training

How much VRAM do I need to fine-tune a model?

Full fine-tuning costs roughly 16 bytes per parameter. LoRA costs 2. QLoRA costs 0.5. The gap between those three numbers is the whole decision.

Full fine-tuning with Adam holds five things per parameter: BF16 weights (2 bytes), BF16 gradients (2), an FP32 master copy (4), and the Adam first and second moments (4 each). That is 16 bytes per parameter before activations. LoRA freezes the base, so you pay 2 bytes per parameter for the frozen model and optimiser state only on the adapter, typically half a percent of the parameters. QLoRA quantises that frozen base to 4 bits, so it costs 0.5 bytes per parameter and the adapter is unchanged.

ModelFull FT, weights + AdamFull FT, with activationsLoRAQLoRA
7B112 GB129 GB21 GB10 GB
13B208 GB239 GB33 GB14 GB
34B544 GB626 GB79 GB28 GB
70B1,120 GB1,288 GB156 GB51 GB
405B6,480 GB7,452 GB856 GB249 GB
Assumptions stated so you can change them: gradient checkpointing on, sequence length 2,048, batch 1, rank-16 adapters on the attention projections at about 0.5% of parameters, 15% added to the full fine-tuning column for activations and fragmentation. Adapter optimiser state counted at 16 bytes per adapter parameter.

The full fine-tuning column is the naive figure, everything resident on the GPUs. ZeRO stage 3 with optimiser state offloaded to host memory cuts the GPU-side requirement substantially, and our 8-way nodes carry 2 TB of host RAM for exactly that. Offload trades PCIe bandwidth for VRAM, so it makes a run possible rather than fast. Method choice is on the GPU for fine-tuning page.

Read the last two columns against the first. A 70B full fine-tune wants an eight-card node at $11,450. The same 70B under QLoRA wants 51 GB, which is one A100 SXM 80GB at $990 a month. That is a twelvefold difference for a method that is worse but rarely twelve times worse. Most people asking this question have been quoted the full fine-tuning number and do not need it.

05 Lookup

What GPU do I need for a 70B model, or any other size?

Model size down the side, method across the top. Each cell is the VRAM required and the cheapest thing in our fleet that holds it.

ModelInference BF16Inference FP8Inference INT4QLoRALoRAFull fine-tune
7B 17 GBRTX A5000 $190 10 GBL4 $250 7 GBRTX A5000 $190 10 GBRTX A5000 $190 21 GBRTX A5000 $190, tight 129 GBMI300X $1,590
13B 35 GBA40 $240 22 GBL4 $250, tight 15 GBRTX A5000 $190 14 GBRTX A5000 $190 33 GBA40 $240 239 GB4× A100 $3,643
34B 72 GBA100 SXM $990 38 GBRTX 6000 Ada $460 21 GBRTX A5000 $190, tight 28 GBA40 $240 79 GBA100 SXM $990, tight 626 GB8× A100 $7,130, tight
70B 145 GBMI300X $1,590 75 GBH100 PCIe $1,490 40 GBA40 $240 51 GBA100 SXM $990 156 GBMI300X $1,590 1,288 GB8× MI300X $11,450
405B 816 GB8× MI300X $11,450 411 GB8× MI300X $11,450 209 GB8× A40 $1,728 249 GB8× A40 $1,728 856 GB8× MI300X $11,450 7,452 GBfive nodes, quoted
The cheapest single card that fits, at 8,192 tokens and batch 1; a multi-card set only where no single card holds it. This table answers one question, the cheapest card that holds the model, and the workload guides answer a different one, the cheapest card worth running it on, which is sometimes a different and more expensive card, with the reason given there. Tight means under 5 GB spare, or under 2 GB per card on multi-card rows. Multi-card prices include the volume discount of 5% at two cards, 8% at four and 10% at eight. Full fine-tuning rows assume NVLink or NVSwitch, because an all-reduce every step over PCIe peer-to-peer is not worth renting. Cheaper multi-card substitutions exist for most rows and section 07 explains when to take one. The 13B FP8 cell is the tightest: 22 GB leaves 2 GB spare on an L4, whose 300 GB/s is the slowest memory we rent, so the RTX 6000 Ada at $460 is the alternative when decode speed rather than fit is the constraint. The MI300X carries seven cells here and comes with three conditions: DFW-1 Dallas only, 8 cards free at the last catalogue update, and ROCm rather than CUDA, so vLLM and PyTorch run and a hand-written CUDA kernel does not. Specs and per-region stock on the GPU catalogue, discounts on the pricing page.

The cell to look at is 70B at INT4: 40 GB, held by an A40 48GB at $240 a month. Not a trick, just what four-bit quantisation and 48 GB of GDDR6 buy. Section 07 sets out what you gave up for it. Anything above one 8-way node is a cluster question, not a card question, and belongs on the multi-GPU cluster rental page.

06 Cost of memory

Which GPU gives you the most VRAM per dollar?

Monthly price divided by gigabytes of VRAM, and memory bandwidth divided by monthly price. Nobody publishes either. Both decide more purchases than a TFLOPS figure does.

GPUVRAMBandwidthMonthly$ per GB/moGB/s per $/moPer GPU-hour
MI300XHBM3, CDNA 3192 GB5.3 TB/s$1,590/mo$8.283.33$2.21
B200HBM3e, Blackwell, 2-GPU minimum180 GB8 TB/s$3,890/mo$21.612.06$5.40
H200 SXMHBM3e, Hopper141 GB4.8 TB/s$2,290/mo$16.242.10$3.18
H100 SXMHBM3, NVLink 900 GB/s80 GB3.35 TB/s$1,690/mo$21.131.98$2.35
H100 PCIeHBM2e80 GB2 TB/s$1,490/mo$18.631.34$2.07
A100 SXMHBM2e, no FP880 GB2.04 TB/s$990/mo$12.382.06$1.38
L40SGDDR6 ECC, FP848 GB864 GB/s$520/mo$10.831.66$0.72
RTX 6000 AdaGDDR6 ECC, FP848 GB960 GB/s$460/mo$9.582.09$0.64
RTX A6000GDDR6 ECC, no FP848 GB768 GB/s$290/mo$6.042.65$0.40
A40GDDR6 ECC, no FP848 GB696 GB/s$240/mo$5.002.90$0.33
RTX 5090GDDR7, FP832 GB1.79 TB/s$540/mo$16.883.31$0.75
RTX 4090GDDR6X, FP824 GB1.01 TB/s$390/mo$16.252.59$0.54
L4GDDR6, FP8, 72 W24 GB300 GB/s$250/mo$10.421.20$0.35
RTX 3090GDDR6X, no FP824 GB936 GB/s$230/mo$9.584.07$0.32
RTX A5000GDDR6 ECC, no FP824 GB768 GB/s$190/mo$7.924.04$0.26
Ranked by VRAM. Dollars per gigabyte is the monthly price divided by capacity. GB/s per dollar is memory bandwidth divided by the monthly price, so higher is better. Per GPU-hour is the monthly price divided by the 720 hours in a term, explained on monthly versus hourly GPU rental. The B200 price is per card, but it ships on an 8-way HGX baseboard and the smallest rentable unit is two cards, so the smallest B200 order is two cards: $7,780 a month at list, $7,391 with the 5% two-card discount.

Two different cards win the two columns, which is why both are published. The A40 holds a gigabyte for $5.00 a month, less than half what an L40S charges for the same 48 GB, because it is a 2020 card with 696 GB/s behind it. The RTX 3090 gives 4.07 GB/s per dollar, more than twice the H100 SXM, because it is an old consumer card with a fast bus and a low rent. Neither is the best card here. Both answer a specific question correctly.

The MI300X is the outlier. At $8.28 per gigabyte it is cheaper memory than an A100, an H100 or an H200, with 5.3 TB/s behind it, which is why it fills the large-model cells above. The cost is ROCm rather than CUDA: vLLM and PyTorch run, a hand-written CUDA kernel does not.

07 Counter-argument

Is more VRAM always the right answer?

No. Three cases where buying capacity is the wrong purchase, with the arithmetic for each.

01

Bandwidth is the bottleneck, not capacity

Decoding one token reads every resident weight from VRAM once. Divide bandwidth by the model's resident size for a hard ceiling on single-stream tokens per second. A 7B at BF16 is 14 GB, so an A40 at 696 GB/s cannot pass 50 tokens per second and an RTX 5090 at 1.79 TB/s cannot pass 128. Both fit the model. One is 2.6 times faster at it.

A ceiling, not a benchmark. You land below it.
02

Quantised on a small card beats full precision on a big one

A 34B at BF16 is 72 GB all in and needs an A100 at $990: 2.04 TB/s over its 68 GB of weights is a ceiling of 30 tokens per second. The same 34B at INT4 is 21 GB all in and fits an RTX 4090 at $390: 1.01 TB/s over its 17 GB of weights is a ceiling of 59. Faster, and 61% cheaper. What you gave up is quality on the quantised weights, which you evaluate rather than assume.

Cheaper and faster, worse output. Test it.
03

Two cheap cards, sometimes

Two A40 give 96 GB for $456 a month against one A100's 80 GB for $990. More memory, less than half the price. It works for batch inference with a replica per card, for QLoRA, and for layer-split serving at low concurrency. It fails under tensor parallelism and on any training step with an all-reduce, because the link between the cards is the part you did not buy.

96 GB for $456. The catch is the interconnect.

The NVLink point, stated precisely.

An H100 SXM talks to its neighbours at 900 GB/s over NVLink. An A40 or RTX A6000 pair takes an NVLink bridge at 112 GB/s. An L40S has no NVLink and moves data over PCIe Gen4 peer-to-peer at a line rate of roughly 32 GB/s. An RTX 5090 has no NVLink either, but its link is Gen5, so roughly 63 GB/s. That gap is paid once per layer per token in a tensor-parallel run.

Splitting by layers instead of by tensors avoids most of that traffic and raises nothing. Per token you still read every byte of the model, on two cards in sequence rather than one alone. Tensor parallelism is what turns two cards into double the bandwidth, and it is exactly the mode that needs the fast link. That is the whole trade.

Worked comparison

2× RTX 5090
64 GB, 3.58 TB/s aggregate, $1,026/mo, PCIe Gen5 peer-to-peer
1× H100 SXM
80 GB, 3.35 TB/s, $1,690/mo, one memory space
Verdict
The pair wins on paper and loses on any model that does not fit in 32 GB per card
If your model fits on one card, put it on one card. Every split you add is a synchronisation point, a failure mode and an evening of debugging device maps.
08 When not to rent from us

The cases where this page should send you elsewhere.

Our shortest term is seven days. If you need a card for an afternoon to check whether a 13B fits in 24 GB, use an hourly provider: a few hours there costs less than the $55 minimum week on an RTX A5000 here. We do not sell hours and have no plan to.

If you are choosing between two cards on speed rather than fit, stop reading tables and run your stack on both. Everything here is capacity arithmetic and a bandwidth ceiling. Neither predicts what your kernels, batch shape and scheduler will do, and we will not pretend otherwise with a throughput figure we cannot source.

If you need 6,480 GB for a 405B full fine-tune, you need a cluster, a fabric and a plan, not a rental page. Ask sales about multi-node capacity in DFW-1 Dallas, IAD-1 Ashburn or PDX-1 Hillsboro and expect a conversation about interconnect before a price.

TL;DR Start here
Learning, any model up to 13B
RTX A5000 24 GB, $190/mo
Serving 24B to 34B in production
L40S 48 GB, $520/mo
Fastest single card under $600
RTX 5090 32 GB, $540/mo
Cheapest 70B at INT4
A40 48 GB, $240/mo
70B at BF16 on one card
MI300X 192 GB, $1,590/mo
Long context on 70B
H200 SXM 141 GB, $2,290/mo
BTCETHUSDTUSDCXMRSOLLTC

Deposit once, deploy in about ninety seconds, no email address and no identity document. See no KYC GPU hosting and how crypto payment works.

09 Questions

VRAM questions, asked the way people ask them.

How much VRAM do I need to run a 70B model?
About 145 GB at BF16, 75 GB at FP8 and 40 GB at INT4, each at 8,192 tokens and batch 1. That is weights plus a 2.7 GB KV cache plus 2 GB of overhead. The FP8 figure fits one 80 GB card. The INT4 figure fits a 48 GB A40 at $240 a month.
How much VRAM do I need for a 7B model?
17 GB at BF16, 10 GB at FP8, 7 GB at INT4, at 8k context and batch 1. All three fit a 24 GB card, so an RTX A5000 at $190 a month covers any of them. The better question is bandwidth: that 7B has a decode ceiling of 55 tokens per second on an A5000 and 128 on an RTX 5090.
How do I calculate the KV cache size?
2 × layers × KV heads × head dimension × bytes per element × context length × batch size. For a 70B with 80 layers, 8 KV heads and head dimension 128, that is 0.33 MB per token at FP16, so 2.7 GB at 8k and 42.9 GB at 128k. An FP8 cache halves every figure.
Is 24 GB of VRAM enough to fine-tune a model?
For QLoRA yes, up to about 34B at short sequence lengths: a 13B needs roughly 14 GB and a 7B needs 10 GB. For LoRA in BF16, 24 GB covers a 7B at about 21 GB and nothing larger. For full fine-tuning with Adam, no. Even a 7B needs about 129 GB, because optimiser state costs 16 bytes per parameter.
What GPU do I need for a 405B model?
At INT4 it needs about 209 GB, which eight A40 hold for $1,728 a month, though at 696 GB/s each that is capacity rather than speed. At FP8 it needs 411 GB and at BF16 816 GB, both of which fit an 8× MI300X node at $11,450. Full fine-tuning needs roughly 7,452 GB and belongs on a cluster.
Do two 24 GB cards work like one 48 GB card?
For capacity, close enough. For speed, no. A model split across two cards still reads every byte per token, so layer splitting does not raise your ceiling. Tensor parallelism does, but it exchanges activations at every layer over the link between the cards: 112 GB/s on an NVLink bridge, roughly 63 GB/s over PCIe Gen5 and 32 GB/s over Gen4.
Which GPU has the cheapest VRAM per gigabyte?
The A40: 48 GB for $240 a month is $5.00 per gigabyte per month, the lowest ratio in our fleet. The RTX A6000 follows at $6.04, the RTX A5000 at $7.92, the MI300X at $8.28. An H100 SXM costs $21.13 per gigabyte because you are buying bandwidth and FP8 throughput, not capacity.
Does quantising to INT4 hurt output quality?
Yes, measurably, and by how much depends on the model and the method. Weight-only schemes such as AWQ and GPTQ lose noticeably less than naive rounding. The honest process is to quantise, run your own evaluation set, and decide whether the loss is worth the card you save. We rent both cards and hold no opinion.

Workload-specific sizing on GPU for LLM inference and GPU for fine-tuning. Full specs, stock and prices on the GPU catalogue.

You know the number. Now pick the card.

Fifteen models in DFW-1, IAD-1 and PDX-1, from $190 a month, weekly or monthly, paid in seven coins with no identity check.

Create an account