> For the complete documentation index, see [llms.txt](https://brindha.gitbook.io/mylearning/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://brindha.gitbook.io/mylearning/tools/minicpm.md).

# MiniCPM

**What it is**

MiniCPM-V is a series of efficient multimodal large language models for edge devices that integrate advancements in architecture, training, and data. The 8B model outperforms GPT-4V, Gemini Pro, and Claude 3 across 11 public benchmarks, processes high-resolution images at any aspect ratio, achieves robust optical character recognition, and supports over 30 languages while running efficiently on mobile phones.&#x20;

**Developer**

Developed by OpenBMB, in collaboration with Tsinghua University's Natural Language Processing Lab (THUNLP) and ModelBest.

**OCR and vision efficiency**

MiniCPM-V 2.6 achieves state-of-the-art performance on OCRBench, surpassing proprietary models such as GPT-4o, GPT-4V, and Gemini 1.5 Pro. It produces only 640 tokens when processing a 1.8 million pixel image, which is 75% fewer than most models, directly improving inference speed, latency, memory usage, and power consumption.&#x20;

**Mobile deployment**

MiniCPM-V is the first end-deployable large language model supporting bilingual multimodal interaction in English and Chinese, and can be efficiently deployed on most GPU cards, personal computers, and even mobile phones.&#x20;

**Latest model: MiniCPM-V 4.5**

MiniCPM-V 4.5, built on Qwen3-8B and SigLIP2-400M with 8B total parameters, achieves an average score of 77.0 on OpenCompass, surpassing GPT-4o-latest, Gemini-2.0 Pro, and strong open-source models like Qwen2.5-VL 72B, making it the most performant multimodal model under 30B parameters. It also supports controllable hybrid fast and deep thinking modes.&#x20;

**Video understanding**

Powered by a new unified 3D-Resampler, MiniCPM-V 4.5 achieves a 96x compression rate for video tokens, where 6 video frames can be jointly compressed into 64 video tokens, enabling high-FPS and long video understanding without increasing computational cost.&#x20;
