Choosing a GPU for LLM inference.
GPU Rent Hub rents dedicated GPU servers for LLM inference from $250 a month for an NVIDIA L4 to $2,290 a month for an H200 SXM, paid in Bitcoin, Ethereum, USDT, USDC, Monero, Solana or Litecoin, with no identity verification.
Two numbers pick the card: the VRAM the weights and KV cache need, and the memory bandwidth that sets tokens per second.
How much VRAM does an LLM need for inference?
Two figures decide the card here. The derivation belongs to the VRAM guide; the term that moves is the KV cache, and it moves with your traffic rather than with the model.
The two figures this page argues from
A 70B model in FP8 is 75 GB at 8,192 tokens and batch 1, and 2.7 GB of that is KV cache. Everything else, the weights-plus-cache-plus-overhead formula, the bytes per parameter at each precision and a row for every model size, is derived term by term in the GPU VRAM guide and is not repeated here. Those two figures do the work below: 75 GB is why an 80 GB card serves a 70B with 5 GB spare, and the 2.7 GB cache is the term carried forward into the 120B row of the table.
The cache term takes over at long context, where it scales with context length times batch size, not parameter count. At 128k and 32 concurrent requests it can exceed the weights. That is why the H200 SXM with 141 GB costs more than an H100 SXM running the same silicon.
| Model size | BF16, 2 B/param | FP8, 1 B/param | 4-bit, 0.5 B/param | Cheapest NVIDIA card that fits, FP8 |
|---|---|---|---|---|
| 7B–8BLlama 8B, Mistral 7B, Qwen 7B | 17 GB | 10 GB | 7 GB | L4, 24 GB |
| 13B–14B40 layers, MHA, large cache | 35 GB | 22 GB | 15 GB | L4, 24 GB, tight |
| 32B–34B | 72 GB | 38 GB | 21 GB | RTX 6000 Ada, 48 GB |
| 70B | 145 GB | 75 GB | 40 GB | H100 PCIe, 80 GB |
| 120Bsame formula, no guide row | 245 GB | 125 GB | 65 GB | H200 SXM, 141 GB |
| 405B | 816 GB | 411 GB | 209 GB | 8× MI300X, 1,536 GB |
Why does memory bandwidth matter more than TFLOPS for LLM inference?
Decoding one token reads every weight once. The card waits on memory, not on maths.
At a batch of one, decoding streams the whole weight set out of VRAM for every token and does about two floating-point operations per parameter, which at BF16 is one operation per byte loaded. Keeping an H100 SXM's tensor cores busy takes roughly 300, taking its 1,979 sparse FP16 TFLOPS as about 990 dense against 3.35 TB/s of bandwidth. The card is memory bound by a wide margin, so the ceiling on tokens per second is bandwidth divided by the bytes of weights read per token.
An 8B model in BF16 is 16 GB. On an RTX 4090 at 1.01 TB/s the ceiling is 63 tokens a second; on an H100 SXM at 3.35 TB/s it is 209. Those are spec-sheet ceilings and real decoding lands below them, by a fraction we will not quote because we cannot source one. The ranking holds. Prefill is the opposite case, compute bound, which is where FP8 tensor cores earn their keep.
| GPU | VRAM | Bandwidth | Ceiling, 8B in BF16 | Monthly | Per GPU-hour |
|---|---|---|---|---|---|
| B200Blackwell, HBM3e, 2-GPU minimum | 180 GB | 8 TB/s | 500 t/s | $3,890/mo | $5.40 |
| MI300XCDNA 3, HBM3 | 192 GB | 5.3 TB/s | 331 t/s | $1,590/mo | $2.21 |
| H200 SXMHopper, HBM3e | 141 GB | 4.8 TB/s | 300 t/s | $2,290/mo | $3.18 |
| H100 SXMHopper, HBM3 | 80 GB | 3.35 TB/s | 209 t/s | $1,690/mo | $2.35 |
| A100 SXMAmpere, no FP8 | 80 GB | 2.04 TB/s | 128 t/s | $990/mo | $1.38 |
| H100 PCIeHopper, HBM2e | 80 GB | 2 TB/s | 125 t/s | $1,490/mo | $2.07 |
| RTX 5090Blackwell, GDDR7 | 32 GB | 1.79 TB/s | 112 t/s | $540/mo | $0.75 |
| RTX 4090Ada, GDDR6X | 24 GB | 1.01 TB/s | 63 t/s | $390/mo | $0.54 |
| RTX 6000 AdaAda, GDDR6 ECC | 48 GB | 960 GB/s | 60 t/s | $460/mo | $0.64 |
| RTX 3090Ampere, no FP8 | 24 GB | 936 GB/s | 59 t/s | $230/mo | $0.32 |
| L40SAda, GDDR6 ECC | 48 GB | 864 GB/s | 54 t/s | $520/mo | $0.72 |
| A40Ampere, no FP8 | 48 GB | 696 GB/s | 44 t/s | $240/mo | $0.33 |
| L4Ada, 72 W | 24 GB | 300 GB/s | 19 t/s | $250/mo | $0.35 |
Does FP8 matter, and why is the A100 a worse inference card than its price suggests?
FP8 halves the weights. That halves the bytes read per token and doubles the model that fits.
Hopper, Ada Lovelace and Blackwell have FP8 tensor cores: the B200, H200, H100 SXM, H100 PCIe, L40S, RTX 6000 Ada, RTX 5090, RTX 4090 and L4. Ampere does not. The A100, A40, RTX A6000, RTX A5000 and RTX 3090 are FP8 not supported, and that is silicon, not a driver gap.
So a 70B model in FP8 is 75 GB and runs on one H100 SXM at $1,690 a month, $2.35 per GPU-hour. On Ampere it runs in BF16 at 145 GB, meaning two A100 SXM cards: $1,980 at list, $1,881 a month with the 5% two-card discount, which is $2.61 an hour for the endpoint against $2.35 for the single H100 SXM, with tensor-parallel traffic on every token. The cheaper card builds the more expensive endpoint.
Ampere is still worth renting
For fine-tuning, for 4-bit serving, and where 80 GB under $1,000 a month is the point. It is a poor default for a production token endpoint, and the $990 price tag hides that.
What is the best GPU for LLM inference at each model size?
Sorted by model size. One card per band, chosen on the arithmetic above.
| Serving this | Pick | VRAM | Weekly | Monthly | Per GPU-hour |
|---|---|---|---|---|---|
| 7B–8B, always on, low QPS72 W, single slot, FP8 | L4 | 24 GB | $70 | $250/mo | $0.35 |
| 7B–14B, latency matters1.01 TB/s, 265 in stock | RTX 4090 | 24 GB | $110 | $390/mo | $0.54 |
| 14B–24B in FP8fastest memory under $600 | RTX 5090 | 32 GB | $155 | $540/mo | $0.75 |
| 24B–34B in FP8ECC, 8-way nodes, RT cores | L40S | 48 GB | $148 | $520/mo | $0.72 |
| 32B in BF16, no quantisationHopper at 350 W | H100 PCIe | 80 GB | $420 | $1,490/mo | $2.07 |
| 70B in FP8, one cardthe production default | H100 SXM | 80 GB | $475 | $1,690/mo | $2.35 |
| 70B in BF16, one cardmost HBM per card | MI300X | 192 GB | $450 | $1,590/mo | $2.21 |
| 100B–120B, or 70B at 128kKV cache headroom | H200 SXM | 141 GB | $640 | $2,290/mo | $3.18 |
| 405B in FP8, one node8-way, NVSwitch 1.8 TB/s | 8× B200 | 1,440 GB | $7,850 | $28,010/mo | $4.86 |
Prices are identical in DFW-1 Dallas, IAD-1 Ashburn and PDX-1 Hillsboro, so choose the region on latency. The MI300X is Dallas only. Of the eight B200, six are in DFW-1 and two in PDX-1, with none in IAD-1.
What do I get on a GPU Rent Hub LLM inference server?
Bare metal, root, the serving stack preinstalled, nothing listening until you enable it.
vLLM 0.9 template
Ubuntu 24.04 with vLLM 0.9 and an OpenAI-compatible server unit installed. The unit ships disabled. You enable it and decide what may reach the port.
Ollama template
For a model answering in two minutes rather than a tuned server. Ollama is bound to localhost and you expose it yourself, usually over WireGuard.
20 TB outbound included
At roughly 100 bytes of JSON and SSE framing per streamed token, that is about 200 billion tokens a month. A text endpoint will not reach it.
100G port, 10G guaranteed
Shared 100G uplink with a 10 Gbps floor, one IPv4 and an IPv6 /64. A dedicated 100G port is $220 a month if you stream audio or images.
Always-on inference is the clearest case for a term: busy at unpredictable hours, intolerant of cold starts, holding weights a metered platform would pull again on every scale-up. The card is yours for 720 hours whether traffic arrives or not, and cannot be preempted. Crossover arithmetic on monthly versus hourly GPU rental, deployment detail in the documentation. Fund a balance in BTC, ETH, USDT, USDC, XMR, SOL or LTC: how crypto payment works, and what no-KYC GPU hosting means.
When should you not rent a dedicated GPU for LLM inference?
Three cases where a dedicated card is the wrong tool, and we would rather say so.
Second, bursty traffic. An endpoint answering a hundred requests on Tuesday and nothing until the next week keeps the card busy for minutes and rents it for 720 hours. Use an hourly provider. Our shortest term is seven days and there is no meter.
Third, a smaller card doing the job. Quantise a 34B model to 4-bit and it fits in 21 GB, which an RTX 3090 at $230 a month serves at 936 GB/s with 3 GB spare, faster memory than the cheaper RTX A5000 the guide's lookup names for the same cell. Do not rent an H100 SXM because the model has a large number in its name.
Rent from us if
The endpoint is always on, per-token billing has become your largest line, the weights and prompts must stay on hardware nobody else touches, or you want to pay in crypto without an identity check.
Do not rent from us if
You need a card for an afternoon, traffic is unpredictable and low, or you have not yet measured whether the model fits in 4-bit on a card a quarter of the price.
Training is a different sizing problem, dominated by optimiser states: see GPU for fine-tuning.
GPU for LLM inference, asked plainly.
What is the best GPU for LLM inference?
What does a 70B endpoint need beyond the weights?
Can I rent a GPU for vLLM?
Why is the A100 not a good LLM inference card?
How many tokens per second will a GPU produce?
Is a monthly rental right for an always-on LLM endpoint?
Do you include enough bandwidth for a token API?
Do I need to give ID to rent an LLM inference server?
Deeper sizing cases in the GPU VRAM guide, every card on the GPU catalogue, term discounts on the pricing page.
Put the endpoint on hardware nobody else touches.
Deploy vLLM on a dedicated card in DFW-1, IAD-1 or PDX-1 in about ninety seconds. Minimum deposit $20 equivalent, no email, no ID.