How Much VRAM Do You Need to Run LLMs Locally?
The primary concern when deploying a local LLM is straightforward: does it fit within your GPU's memory? The answer hinges on the model's size, the quantization level applied, and the required context length. This guide provides a practical baseline for selecting the appropriate VRAM capacity.
How quantization affects VRAM
Quantization lowers the precision used to store model weights. Using lower-bit quantization reduces the model's footprint and VRAM consumption, albeit with a slight trade-off in output quality.
| Quant | Bits per weight | Typical use |
|---|---|---|
| Q8_0 | 8 | Very high quality |
| Q6_K | ~6.6 | Very good quality |
| Q5_K_M | ~5.5 | Good quality and size |
| Q4_K_M | ~4.5 | Good balance of size and quality |
| Q3_K_M | ~3.5 | Lower VRAM, more quality loss |
Q4_K_M is frequently selected when VRAM is constrained. If you have additional VRAM headroom, opting for Q5 or Q6 allows you to run the same model with less aggressive quantization.
Approximate VRAM by model size
The following figures provide rough estimates for model weights. Note that actual VRAM requirements are higher, as the runtime, KV cache, and context processing also consume memory.
| Model size | Q8_0 | Q6_K | Q4_K_M | Q3_K_M |
|---|---|---|---|---|
| 4B | ~5 GB | ~4 GB | ~3 GB | ~2.5 GB |
| 8B | ~9 GB | ~7 GB | ~5.5 GB | ~4.5 GB |
| 12B | ~13 GB | ~10 GB | ~8 GB | ~6.5 GB |
| 14B | ~16 GB | ~12 GB | ~9 GB | ~7.5 GB |
| 27B | ~30 GB | ~22 GB | ~17 GB | ~13 GB |
| 32B | ~36 GB | ~27 GB | ~20 GB | ~16 GB |
| 70B | ~80 GB | ~60 GB | ~42 GB | ~34 GB |
These are estimates rather than strict limits. Variations in model architecture and quantization formats can influence the actual memory footprint.
What different amounts of VRAM can run
| VRAM | Practical range | Current examples |
|---|---|---|
| 8 GB | Small models around 4B to 9B | Gemma 4 E4B, Qwen3.5 9B |
| 12 GB | Small to mid-sized models around 9B to 14B | Gemma 4 12B, Qwen3.5 9B |
| 16 GB | 12B to 27B with lower quantization | Gemma 4 26B-A4B, Qwen3.6 27B at Q4 |
| 24 GB | 27B to 35B at Q4 to Q6 | Qwen3.8 27B, Gemma 4 31B |
| 32 GB | 27B to 35B at higher quantization | Qwen3.8 27B, Gemma 4 31B |
| 48 GB | Large dense models at lower quantization | 70B-class models at Q3 to Q4 |
| 80 GB | Large dense models at higher quantization | 70B-class models at Q4 to Q6 |
These ranges apply to models whose weights can fully reside on the GPU. Large MoE models operate differently: while only a subset of parameters is active per token, the entire weight set must still be stored. Consequently, a model with 100B or more total parameters will not fit within a 100B VRAM budget simply because it activates fewer parameters.
MoE models
Mixture-of-Experts (MoE) models comprise multiple parameter groups known as experts. Since only specific experts are engaged for each token, inference can be more efficient than a dense model with an equivalent total parameter count.
However, inactive experts still occupy memory. Therefore, large MoE models may demand significantly more memory than their active parameter count implies. Very large models might require multiple GPUs or offloading to system RAM.
Context length also uses VRAM
Model weights represent only a portion of the total memory requirement. The KV cache expands as context length increases, meaning that running the same model with a 64K context can necessitate substantially more VRAM than a 4K context.
- Longer context windows consume more VRAM.
- The precision of the KV cache impacts memory usage.
- Batch size and the number of concurrent users also increase memory demands.
- Reserve sufficient VRAM for the runtime rather than allocating all capacity to model weights.
Practical tips
- Verify the actual size of the specific quantized model you intend to run.
- Avoid relying solely on the model file size as the exact VRAM requirement; ensure space remains for the KV cache and runtime.
- If a model exceeds available VRAM, parts of it can be offloaded to system RAM, though this typically slows down inference.
- For long-context or agentic workloads, budget more VRAM than the model weights alone require.
- Multiple GPUs can distribute the model load if a single GPU lacks sufficient VRAM.
Run it on DaDesktop
You do not need to purchase a GPU to run a local LLM. DaDesktop provides a cloud desktop with the required VRAM, enabling you to execute the model directly without owning the hardware.
Select the VRAM tier that suits your model, load it, and begin usage. There is no setup required, no hardware purchase, and no driver complications. View available GPUs for your options.