Use-case guide

Best fast local LLMs for low-latency use in 2026

Best small and fast local LLMs for low-latency chat, laptops, edge machines and 8GB to 16GB RAM setups. Ranked from the LocalClaw model database with RAM requirements, quantization and links to static model pages.

Matching models
40
Best pick
Bonsai 27B
Primary signal
speed, light, edge
SEO query
fastest local LLM

Quick answer

For fast / small, start with Bonsai 27B if your hardware fits it. If not, choose the highest-ranked model that fits your RAM tier and preferred quantization.

Top local models for fast / small

#1

Bonsai 27B

27.3B (ternary / 1-bit) · 16GB RAM · Ternary Q2_0_g128 · Q:8 C:9 R:9 S:9

PrismML low-bit model derived from Qwen 3.6 27B. Official Apache 2.0 ternary (7.2GB deployed) and 1-bit (3.9GB) builds retain multimodal, reasoning and agentic capabilities through custom GGUF and MLX runtimes.

chatcodereasoningvisionagenticmultimodal
#2

Llama-3.1-Nemotron-Nano (4B)

4B · 6GB RAM · Q5_K_M · Q:7 C:6 R:8 S:10

⭐ Mac Mini M4 16GB top pick! NVIDIA fine-tune of Llama 3.1. Hybrid /think • /no_think mode — deep reasoning on demand, instant chat otherwise. ~80–120 tok/s on Apple Silicon Metal. 128K context. Apache 2.0.

chatlightspeedreasoning
#3

Nemotron 3 Nano (4B)

4B · 6GB RAM · Q5_K_M · Q:7 C:6 R:7 S:10

⭐ Mac Mini M4 16GB top pick! NVIDIA's hybrid model — distilled from 9B, keeps 95% of its quality. Hybrid attention + SSM layers = ~80–120 tok/s on Apple Silicon. Blazing fast, minimal RAM. NVIDIA Open Model License.

chatlightspeedreasoning
#4

LFM2.5-8B-A1B

8.3B (1.5B active) · 8GB RAM · Q4_K_M · Q:8 C:8 R:8 S:9

Liquid AI hybrid model built for on-device assistants. 8.3B total / 1.5B active, 128K context, tool use, GGUF, ONNX, MLX, llama.cpp and LM Studio support. Open-weight under LFM 1.0.

chatcodereasoningspeedstandardgeneral
#5

Agents-A1 4B

4B · 8GB RAM · Q4_K_M · Q:7 C:8 R:8 S:9

InternScience compact dense agent model with Apache 2.0 licensing, 262K context and official Q4_K_M GGUF artifacts for 8GB-class local assistants.

chatcodereasoningvisionagentictool-calling
#6

Qwen 3.5 MoE (35B/3B active)

35B (3B active) · 24GB RAM · Q4_K_M · Q:8 C:9 R:8 S:9

MoE gem — only 3B params active at inference. 19x faster than Qwen3-Max at 256K context. Best quality-per-watt of the series. Hybrid thinking mode. Runs on Mac Studio 32GB. Agentic coding standout.

chatcodereasoningpowerspeed
#7

MiniCPM5 1B

1B · 4GB RAM · Q4_K_M · Q:6 C:6 R:6 S:10

OpenBMB compact on-device LLM with Apache 2.0 licensing, 128K context, tool-calling focus and official GGUF plus MLX artifacts for laptops and edge devices.

chatcodereasoninglightspeedtool-calling
#8

Qwen 3.6 (6.7B)

6.7B · 8GB RAM · Q4_K_M · Q:7 C:7 R:8 S:9

Alibaba's hybrid-thinking micro-flagship. Toggles between instant answers and deep chain-of-thought reasoning on demand. 128K context, 29 languages, outperforms Qwen3-8B on reasoning benchmarks. Apache 2.0.

chatcodereasoningspeedgeneral
#9

Granite 3.3 (2B Instruct)

2B · 4GB RAM · Q5_K_M · Q:6 C:6 R:5 S:10

IBM ultra-efficient 2B. Best-in-class among small models for tool calling & structured output. Perfect for on-device RAG and agents. 128K context. Apache 2.0.

chatlightedgespeedcode
#10

Ling-2.6-flash (104B MoE)

104B (7.4B active) · 80GB RAM · Q4_K_M · Q:9 C:9 R:8 S:8

InclusionAI's MIT-licensed instruct MoE optimized for fast agent workloads. 104B total parameters, only 7.4B active, hybrid linear attention, 262K context and strong tool-use / multi-step execution with high token efficiency.

chatcodereasoningspeedquality
#11

Qwen 3.5 (2B)

2B · 4GB RAM · Q4_K_M · Q:5 C:5 R:4 S:10

Ultra-compact Qwen 3.5 with hybrid thinking mode and 256K context. Runs comfortably on 4 GB RAM — ideal for MacBook Air M1/M2, Windows laptops, and edge devices. Apache 2.0.

chatcodeedgespeed
#12

DeepScaleR (1.5B)

1.5B · 4GB RAM · Q5_K_M · Q:5 C:4 R:8 S:10

Tiny model beating o1-preview on math! Incredible reasoning-to-size ratio. 474K downloads.

reasoninglightspeed
#13

GLM 4.7 Flash

14B · 16GB RAM · Q5_K_M · Q:7 C:7 R:7 S:9

Zhipu AI's fast GLM model. 14B parameters optimized for quick responses with strong bilingual (CN/EN) capabilities. Efficient inference for everyday tasks. Apache 2.0.

chatcodepowerspeed
#14

Ministral 3 3B Instruct

3B · 4GB RAM · Q4_K_M · Q:6 C:6 R:6 S:9

Mistral AI compact multimodal instruct model. Apache 2.0, strong local app support through official GGUF, LM Studio, Ollama and llama.cpp artifacts. Practical on normal laptops.

chatvisionlightspeedgeneral
#15

Qwen 3.5 (4B)

4B · 6GB RAM · Q4_K_M · Q:6 C:6 R:6 S:9

Sweet-spot small model. Surprisingly capable for its size with hybrid thinking, 256K context and strong multilingual support. Runs on 8 GB RAM. The go-to for MacBook Air M4 16 GB. Apache 2.0.

chatcodereasoningspeedgeneral
#16

Ornith 1.0 9B GGUF

9B · 8GB RAM · Q4_K_M · Q:7 C:8 R:7 S:8

Compact Ornith 1.0 GGUF variant from DeepReinforce for agentic coding experiments on consumer hardware. MIT licensed and much more practical than the frontier 397B release.

chatcodereasoningspeedagentictool-calling
#17

Gemma 4 E2B

E2B · 6GB RAM · Q5_K_M · Q:6 C:5 R:6 S:9

Gemma 4 compact multimodal model for on-device usage. Supports text, image, audio, and video understanding with 256K context. Apache 2.0.

chatvisionspeededgemultimodalgeneral
#18

Qwen 3 (4B)

4B · 4GB RAM · Q5_K_M · Q:6 C:7 R:7 S:9

Alibaba's think-then-answer model. Built-in chain-of-thought reasoning at just 4B params.

chatcodelightspeedreasoning

How this ranking works

LocalClaw ranks models using their tags plus relative benchmark scores for speed, quality, coding and reasoning. The goal is a practical local setup recommendation, not a synthetic leaderboard.