Use-case guide

Best local vision LLMs in 2026

Best local multimodal and vision-language models for image understanding, OCR, document analysis and visual reasoning. Ranked from the LocalClaw model database with RAM requirements, quantization and links to static model pages.

Matching models
38
Best pick
MiniMax M3 (428B/23B active)
Primary signal
vision, multimodal
SEO query
best local vision LLM

Quick answer

For vision, start with MiniMax M3 (428B/23B active) if your hardware fits it. If not, choose the highest-ranked model that fits your RAM tier and preferred quantization.

Top local models for vision

#1

MiniMax M3 (428B/23B active)

428B (23B active) · 2048GB RAM · BF16 / custom runtime · Q:10 C:10 R:10 S:3

MiniMax native multimodal MoE with 1M context and MiniMax Sparse Attention. Around 428B parameters with 23B active. Built for long-context coding, cowork and agentic workflows, with local deployment via SGLang, vLLM or Transformers. Server-grade only.

chatcodereasoningagenticlong-contextmultimodal
#2

Bonsai 27B

27.3B (ternary / 1-bit) · 16GB RAM · Ternary Q2_0_g128 · Q:8 C:9 R:9 S:9

PrismML low-bit model derived from Qwen 3.6 27B. Official Apache 2.0 ternary (7.2GB deployed) and 1-bit (3.9GB) builds retain multimodal, reasoning and agentic capabilities through custom GGUF and MLX runtimes.

chatcodereasoningvisionagenticmultimodal
#3

Kimi K2.7 Code (1T MoE)

1T (32B active) · 1024GB RAM · BF16 / compressed-tensors · Q:10 C:10 R:10 S:2

Moonshot AI coding-focused agentic Kimi built on K2.6. 1T MoE with 32B active parameters, 256K context, MoonViT vision encoder and stronger long-horizon coding while reducing thinking-token usage by roughly 30% vs K2.6. Modified MIT. Server-grade only.

codereasoningagenticmultimodalquality
#4

Qwen 3.6 35B-A3B

35B (3B active, MoE) · 32GB RAM · Q4_K_M · Q:9 C:10 R:9 S:7

Qwen Team open-weight MoE for agentic coding and multimodal work. 35B total / 3B active, 262K native context, Apache 2.0, and strong GGUF availability through Unsloth and LM Studio-compatible artifacts.

chatcodereasoningvisionagenticpower
#5

Llama 4 Maverick (17B/400B MoE)

400B (17B active, 128 experts) · 384GB RAM · Q4_K_M · Q:10 C:10 R:10 S:2

Meta Llama 4 Maverick — 128-expert MoE flagship. Matches or beats GPT-4o and Gemini 2.0 Flash on reasoning, coding and multimodal benchmarks. 1M-token context. Server-grade hardware only. Llama 4 Community License.

chatvisionreasoningmultimodalquality
#6

Gemma 4 26B A4B

26B (A4B active) · 24GB RAM · Q4_K_M · Q:9 C:8 R:9 S:7

Gemma 4 MoE flagship-for-workstations: 26B total with ~4B active parameters. 256K context and excellent quality-per-watt for local inference. Apache 2.0.

chatcodereasoningpowermultimodalgeneral
#7

Gemma 4 31B

31B · 32GB RAM · Q4_K_M · Q:9 C:9 R:9 S:5

Largest Gemma 4 model for premium local quality. Strong coding and reasoning with 256K context and broad multilingual support. Apache 2.0.

chatcodereasoningqualitymultimodalgeneral
#8

Agents-A1

35B (3B active, MoE) · 32GB RAM · Q4_K_M · Q:9 C:8 R:9 S:5

InternScience Apache 2.0 agentic VLM. 35B-A3B MoE, 262K context, strong long-horizon search/tool-use benchmarks and official Q4_K_M GGUF artifacts for local workstations.

chatcodevisionagentreasoningpower
#9

Llama 4 Scout (17B/109B MoE)

109B (17B active, 16 experts) · 96GB RAM · Q4_K_M · Q:9 C:8 R:9 S:5

Meta Llama 4 Scout — natively multimodal MoE with 16 experts. 10M-token context window. Outperforms Gemma 3 and Mistral Small on most benchmarks at similar active cost. Llama 4 Community License.

chatvisionreasoningmultimodalpower
#10

Qwen 3 VL (32B)

32B · 32GB RAM · Q4_K_M · Q:9 C:7 R:9 S:5

Qwen 3 VL flagship open vision model. Competes with GPT-4o on MMMU, chart-QA and document reasoning. Native video understanding up to 1 hour. Apache 2.0.

visionchatmultimodalpowerquality
#11

Ministral 3 14B Instruct

14B · 16GB RAM · Q4_K_M · Q:8 C:8 R:8 S:7

Mistral AI larger Ministral 3 instruct model. Apache 2.0, official GGUF availability, better quality ceiling than the 3B/8B variants while staying practical on 16-32GB workstations.

chatvisionpowerreasoningmultilingual
#12

Llama 4 Maverick (17B/128E MoE)

17B active (400B total, 128 experts) · 320GB RAM · Q4_K_M · Q:10 C:10 R:10 S:1

Meta's largest open MoE. 17B active params across 128 experts (~400B total). Multimodal with exceptional image reasoning. Server-grade hardware required. Llama 4 License.

chatvisionquality
#13

Qwen 3 VL (8B)

8B · 12GB RAM · Q4_K_M · Q:8 C:6 R:8 S:7

Qwen 3 vision-language model. Strong OCR, document understanding, chart & UI reasoning. 128K context with native image+video inputs. Apache 2.0.

visionchatmultimodalstandard
#14

Agents-A1 4B

4B · 8GB RAM · Q4_K_M · Q:7 C:8 R:8 S:9

InternScience compact dense agent model with Apache 2.0 licensing, 262K context and official Q4_K_M GGUF artifacts for 8GB-class local assistants.

chatcodereasoningvisionagentictool-calling
#15

Gemma 4 12B

12B · 16GB RAM · Q4_K_M · Q:8 C:8 R:8 S:6

Google DeepMind 12B unified multimodal model. Text, image, audio and video inputs, 256K context, Apache 2.0, and a strong local sweet spot for 16-32 GB machines.

chatvisionaudiocodereasoningpower
#16

Qwen 2.5 VL (72B)

72B · 64GB RAM · Q4_K_M · Q:10 C:7 R:9 S:2

Qwen massive vision-language model. Exceptional image and video understanding at 72B scale. 72K context.

visionquality
#17

Gemma 3 (27B)

27B · 32GB RAM · Q4_K_M · Q:9 C:8 R:9 S:4

Google's flagship multimodal. Image + text understanding at an exceptional level.

chatvisionpowerqualitygeneral
#18

Llama 3.2 Vision (90B)

90B · 72GB RAM · Q4_K_M · Q:10 C:7 R:9 S:1

Meta's largest vision model. 128K context with powerful image reasoning and analysis. Requires significant hardware.

visionquality

How this ranking works

LocalClaw ranks models using their tags plus relative benchmark scores for speed, quality, coding and reasoning. The goal is a practical local setup recommendation, not a synthetic leaderboard.