How accurate are the memory estimates?
They are planning-grade estimates, not runtime guarantees. We separate weights, KV cache, and overhead so you can inspect assumptions.
Use this AI model VRAM and RAM calculator to answer “can I run this model locally?” before you waste time on trial-and-error. Pick a model, set quantization, context length, and batch size, then compare fit results across NVIDIA, AMD, Apple Silicon unified memory, and CPU/RAM offload plans.
Pick a model, quantization, context length, batch size, and target hardware. Results apply to dedicated GPU VRAM, Apple unified memory, and CPU/RAM offload planning.
If repo id is unknown, calculator uses your custom params and labels results as estimated.
| Quant | Weights | KV | Overhead | Total |
|---|---|---|---|---|
| FP16 / BF16 | 50.29 GiB | 4 GiB | 3.74 GiB | 58.03 GiB |
| INT8 / Q8_0 | 25.77 GiB | 4 GiB | 2.78 GiB | 32.56 GiB |
| Q6_K | 19.96 GiB | 4 GiB | 2.52 GiB | 26.48 GiB |
| Q5_K_M | 17.13 GiB | 4 GiB | 2.43 GiB | 23.56 GiB |
| Q4_K_M | 14.46 GiB | 4 GiB | 2.46 GiB | 20.91 GiB |
| Q3_K_M / Q3 | 11 GiB | 4 GiB | 2.26 GiB | 17.26 GiB |
| Q2_K / Q2 | 7.86 GiB | 4 GiB | 1.98 GiB | 13.84 GiB |
| Hardware | Memory | Verdict | Est. tok/s |
|---|---|---|---|
| NVIDIA GeForce RTX 4090 24GB | 24 GiB | Fits | 37.26 |
| NVIDIA GeForce RTX 5090 32GB | 32 GiB | Fits | 66.25 |
| NVIDIA GeForce RTX 4080 SUPER 16GB | 16 GiB | Does not fit | 27.21 |
| NVIDIA GeForce RTX 4070 Ti SUPER 16GB | 16 GiB | Does not fit | 24.84 |
| AMD Radeon RX 7900 XTX 24GB | 24 GiB | Fits | 35.49 |
| NVIDIA A100 80GB PCIe | 80 GiB | Fits | 71.54 |
| Apple Silicon M3 Max (128GB unified memory) | 128 GiB | Fits | 12.61 |
| Apple Silicon M2 Ultra (192GB unified memory) | 192 GiB | Fits | 25.23 |
| 2× NVIDIA GeForce RTX 4090 (aggregate) | 48 GiB | Fits | 67.06 |
They are planning-grade estimates, not runtime guarantees. We separate weights, KV cache, and overhead so you can inspect assumptions.
Yes. We track total and active parameters separately: total parameters drive weight residency memory, while active parameters are used for decode-oriented throughput estimates.
Yes. The repository includes a script that pulls Hugging Face config.json and generates a typed model entry you can publish within hours.