ModelFitCheck

Will this LLM fit your GPU? Paste your specs or let us detect them β€” instant VRAM math, zero data sent.

57 models Β· 54 GPUs Β· updated Aug 2, 2026

πŸ”Look Up Any Model

Enter a HuggingFace model ID (e.g., meta-llama/Llama-3.1-8B) or an Ollama name (e.g., llama3.1:8b)

Your hardware

Detecting…

2 / 4 / 6 / 8 / 12 / 16 / 20 / 24 / 32 / 48 / 64

Models that fit

Context: 8k tokens β€Ή VRAM 0.0 GB

Best fits for your hardware

Nothing fits at this context/VRAM. Lower the context length or pick a smaller card.

ModelParamsFitⓐ
Kimi K3 8B2026-07

Kimi K3 8B dense model.

8B5.6 GB needed
Kimi K3 22B MoE2026-07
MoE architecture

Kimi K3 MoE architecture.

22B17.0 GB needed
Qwen 3.6 35B-A3B MoE2026-04
MoE architecture

Requested Qwen 3.6 35B MoE.

35B22.7 GB needed
Qwen 3.6 27B2026-04

Qwen 3.6 27B dense model.

27B17.7 GB needed
Google Gemma 3n E4B2025-06

Gemma 3n mobile-first, edge-optimized.

4B3.1 GB needed
Meta Llama 4 Scout 17B2025-04
MoE architecture

Llama 4 Scout MoE, 17B active params.

17B11.5 GB needed
Qwen 3 8B2025-04

Qwen 3 8B with thinking mode.

8B5.6 GB needed
Qwen 3 4B2025-04

Qwen 3 4B, compact reasoning model.

4B3.5 GB needed
Qwen 3 1.7B2025-04

Qwen 3 tiny, great for edge devices.

2B1.7 GB needed
Qwen 3 30B-A3B MoE2025-04
MoE architecture

Qwen 3 MoE, 3B active of 30B total.

30B19.2 GB needed
Google Gemma 3 12B Instruct2025-03

Gemma 3 mid, vision + text.

12B7.7 GB needed
Google Gemma 3 27B Instruct2025-03

Gemma 3 large, open 27B.

27B17.2 GB needed
Google Gemma 3 4B Instruct2025-02

Multimodal-capable 4B.

4B3.1 GB needed
Mistral Small 24B Instruct2025-01

Mistral Small 3 (24B).

24B15.7 GB needed
DeepSeek R1 Distill 7B2025-01

R1 reasoning distilled into 7B.

7B4.6 GB needed