Back to News
Models

OpenBMB Unveils MiniCPM5-2B, a Compact Dense Model Aimed at On-Device Deployment

The new 2.5-billion-parameter model posts competitive benchmark scores while retaining compatibility with standard inference frameworks.

September 7, 2026 · 1 min read · HowToPrompts Newsroom

OpenBMB has introduced MiniCPM5-2B, a dense causal language model packing roughly 2.52 billion parameters and supporting a native context window of 131,072 tokens. According to the project's model card, the release averages 53.9 across a suite of 34 benchmarks, outperforming Qwen3.5-4B's 51.1 average. The strongest gains appear in areas such as tool use, coding agents, and long-context retrieval tasks.

The training pipeline for MiniCPM5-2B combines 400 billion tokens of deep-thinking supervised fine-tuning with reinforcement learning from teacher models and on-policy distillation. A notable step merges outputs from 16 expert models into a single checkpoint, yielding a unified weight set. The team has also released the pre-training, SFT, and RL datasets, along with intermediate checkpoints labeled Base, Midtrain, and SFT-only, all under the permissive Apache 2.0 license.

Designed for local execution, the model's GGUF quantizations start at roughly 1.56 GB. Because it uses the standard LlamaForCausalLM architecture, it runs directly in popular engines such as vLLM, SGLang, llama.cpp, Ollama, and MLX without needing custom code forks. This compatibility, paired with the open release of training artifacts, positions MiniCPM5-2B as a flexible option for developers seeking a capable small model for edge environments.

Written in-house by the HowToPrompts newsroom, in our own words. The story was first reported by MarkTechPost.