Your GPU. Exact local AI configurations. Real execution evidence.

Affiliate disclosure: We earn commissions when you shop through the links below at no additional cost to you.

Start with a source-backed VRAM estimate, then see tested Recipes for real workload evidence. Failures and unknowns stay visible.

See tested Recipes →

Estimate VRAM for your local LLM

GB

Select a GPU and model

What fits on a 24 GB GPU?

Sorted by minimum VRAM required. Green = fits comfortably. Yellow = tight. Red = won't fit.

How we calculate VRAM

1. Model weights

VRAM = params × bytes_per_param
FP16 = 2 bytes/param. Q8 = 1. Q4_KM ≈ 0.5. Q2 ≈ 0.3.

2. KV cache overhead

+2 × params × context_tokens × bytes_per_param
At 8k tokens: add ~0.5 GB per billion params (Q4).

3. CUDA overhead

~10-15% margin for kernel memory.
vLLM adds ~2-4 GB extra for paged attention.

Need more VRAM? Try cloud.

Rent an A100 or H100 by the hour. No commitment.

New models, new GPU requirements.

We update the calculator when new models drop. No spam.