Open-weight MoE

Laguna S 2.1

Poolside workstation-class coding MoE with 1M context, 118B total / 8B active parameters, OpenMDW-1.1 licensing and official Q4_K_M GGUF plus Ollama availability.

Large-memory workstation 96 GB RAM Q4_K_M Workstation local coding agent
Parameters
118B (8B active, MoE)
Minimum RAM
96 GB
Model size
75.2 GB
Quantization
Q4_K_M

Can Laguna S 2.1 run locally?

Laguna S 2.1 needs a serious workstation with large unified memory or high VRAM.

Search for laguna-s-2.1 in LM Studio or another GGUF-compatible runtime.

chatcodereasoningagentbeastlong-contexttool-calling

Install path

01
Check RAM fitMinimum 96 GB RAM. Start with the Q4_K_M quant.
02
Load the modelSearch laguna-s-2.1 in LM Studio.
03
Control locallyUse LocalClaw to manage models, agents, chat, channels and scheduled OpenClaw work.

Strengths

  • Official Poolside base weights and official GGUF conversion
  • 118B total / 8B active MoE design targets higher coding-agent quality than Laguna XS
  • 1,048,576-token context window with native interleaved reasoning support
  • Official Q4_K_M GGUF is about 75.2GB
  • Ollama availability plus Poolside llama.cpp fork guidance
  • Strong vendor-reported coding, SWE-bench and long-horizon benchmark results

Limitations

  • OpenMDW-1.1 is a custom license, not Apache 2.0 or MIT
  • Comfortable local use needs a 96GB+ workstation, and 1M context can need much more memory
  • Poolside currently points users to its Laguna llama.cpp fork for full GGUF support
  • The Q4_K_M GGUF is a very large download for ordinary laptops
  • Benchmark comparisons rely substantially on vendor-reported or third-party-reported numbers

Best use cases

  • Workstation local coding agent
  • Long-horizon software engineering tasks
  • Repository-scale reasoning
  • Terminal automation with tool calls
  • Private high-end code review
  • Large-context coding research

Capability profile

speed
3
quality
9
coding
10
reasoning
9

Technical notes

Developer
Poolside
License
OpenMDW-1.1
Context window
1,048,576 tokens
Architecture
118B total parameter Mixture-of-Experts language model with about 8B activated parameters per token, 256 routed experts plus one shared expert, grouped-query attention, and mixed sliding-window plus global attention layers.

This model fits these next steps

Hardware fit is based on LocalClaw's RAM tier, model size and quantization metadata. Always leave memory headroom for your OS and runtime.

Similar models to compare

Where to go next