MiniMax M3
1M context · native multimodal · flagship MoE for coding and agent workflows
MiniMax M3 is MiniMax's open-weight AI chat LLM released in June 2026, using a 428B-parameter MoE architecture (~23B activated), with up to 1M-token context and native multimodal text, image, and video input. With MiniMax Sparse Attention (MSA), MiniMax M3 balances speed and quality on long-context coding and Agent tasks, reaching frontier levels in coding, tool use, and long-horizon collaboration. Whether you are a developer, AI engineer, or enterprise team, you can learn about and integrate MiniMax M3 online for intelligent chat and automation workflows.
Context window
1M tokens
Total parameters
428B MoE
Active parameters
~23B
Input modalities
Text/image/video
💡 Why Choose
Why choose MiniMax M3?
MiniMax M3 is MiniMax's AI chat model officially released in June 2026, positioned around "long context + native multimodal + coding agents." It uses a 428B-parameter MoE architecture with ~23B active parameters per inference, offering up to 1M-token API context—fitting medium-to-large codebases, long documents, or full Agent conversation history into a single model view.
MiniMax M3's core architectural innovation is MiniMax Sparse Attention (MSA). Compared with traditional full attention, MSA reduces per-token compute at million-token scale to roughly 1/20 of predecessor M2, with ~9× Prefill and ~15× Decode acceleration—making "ultra-long context" a practical productivity feature rather than a theoretical spec. MiniMax M3 also mixes text, image, and video from the first training step, making it one of few open-weight chat models with native multimodal understanding.
On bikabika AI, you can explore MiniMax M3 core capabilities, thinking mode differences, typical use cases, and integration in one place to quickly assess whether it fits your coding, Agent, or multimodal chat needs—and try it online.
⚡ Features
MiniMax M3 core features
The six capabilities below form MiniMax M3's core competitive strengths in AI chat, coding, and agents.
1M-token ultra-long context
MiniMax M3 supports up to a 1M-token context window (officially guaranteed minimum 512K), loading large codebases, long documents, or multi-turn Agent history for repo-scale global reasoning.
Native multimodal input
MiniMax M3 mixes text, image, and video modalities from the first training step with deeper semantic fusion—for video understanding, visual Q&A, and cross-modal Agent tasks.
Frontier coding and Cowork capability
MiniMax M3 excels on long-horizon Agent benchmarks and coding—suited to code generation, review, refactoring, and multi-file collaborative development, and pairs with MiniMax Code and other Agent products.
MSA sparse attention acceleration
MiniMax M3 uses MiniMax Sparse Attention (MSA): at 1M context, Prefill is ~9× faster and Decode ~15× faster, with per-token compute about 1/20 of prior M2.
Switchable thinking mode
MiniMax M3 can turn thinking mode on or off: on for complex reasoning and Agent tasks; off for lower latency in chat completion and high-frequency calls.
Tool calling and Agent workflows
MiniMax M3 is optimized for tool use and Interleaved Thinking—re-evaluating state after tool execution—suited to multi-step automation and Computer Use apps.
Want to experience MiniMax M3 long-context coding yourself?
MiniMax M3 supports API and self-hosted deployment—start your first million-context AI chat experience now.
🔄 Compare
MiniMax M3 thinking mode comparison
MiniMax M3 supports per-request thinking mode switching—the comparison below helps you choose between complex agents and low-latency scenarios.
| Dimension | Thinking mode on | Thinking mode off |
|---|---|---|
| Best for | Complex reasoning, agents, long-horizon collaboration | Chat, code completion, high-frequency calls |
| Response latency | Higher (quality-first) | Low latency, high throughput |
| Reasoning depth | Deeper, suited to multi-step tasks | Direct output, speed-first |
| Agent performance | More stable tool calls and interleaved thinking | Better for lightweight interaction |
| Coding tasks | Complex refactors, cross-file analysis | Single-line/local completion |
| API pricing | Same as off mode | Same as on mode |
| Switch method | Enable thinking parameter | Disable thinking parameter |
| Recommended pairing | MiniMax Code Agent | Real-time chat bots |
💎 Highlights
MiniMax M3 technical highlights
- MiniMax M3 supports up to 1M-token context for repo-scale code and long-document analysis
- MiniMax M3's 428B MoE with ~23B activated balances capability density and inference efficiency
- MiniMax M3 MSA sparse attention sharply reduces compute and KV Cache pressure at million-token context
- MiniMax M3 is natively multimodal: unified modeling of text, image, and video input
- MiniMax M3 thinking mode switches per request for complex Agents vs. low-latency scenarios
- MiniMax M3 offers OpenAI/Anthropic-compatible APIs and open Hugging Face weights for self-hosting
🎯 Use Cases
MiniMax M3 use cases
From repo-scale coding to multimodal agents, MiniMax M3 delivers stable, efficient long-context AI chat and reasoning across the six scenarios below.
Developers · Architects
Large codebase understanding and refactoring
MiniMax M3's 1M-token context can load key modules of medium-to-large monorepos in one pass—cross-file dependency analysis, refactor design, and code review without chunked RAG.
- 1M tokens
- Codebase
- MiniMax M3
Content teams · Product managers
Long video and multimodal analysis
MiniMax M3 natively supports image and video input—combine with text instructions for video summarization, shot understanding, visual Q&A, and cross-modal extraction for media and creative workflows.
- Multimodal
- Video understanding
- MiniMax M3
AI engineers · Technical teams
Coding agents and automated development
MiniMax M3 reaches frontier levels on long-horizon agent benchmarks and coding tasks—pair with MiniMax Code and similar Agent products for multi-step development, testing, and deployment automation.
- Coding agents
- MiniMax Code
- Tool calling
Developers · Automation engineers
Long-horizon agent workflows
MiniMax M3 is optimized for interleaved thinking—re-evaluating state after tool execution, ideal for Agent pipelines needing multi-round tool calls, state tracking, and result aggregation.
- Agent workflow
- Tool calling
- MSA
Product teams · RPA engineers
Computer Use and desktop automation
MiniMax M3 supports Computer Use-style capabilities—combine visual input to understand UI state and plan actions, extending AI chat to desktop-level automation.
- Computer Use
- Desktop automation
- MiniMax M3
Enterprise IT · Ops teams
Enterprise API integration and self-hosting
MiniMax M3 offers OpenAI / Anthropic-compatible API and open Hugging Face weights—enterprises can choose cloud calls or private deployment per compliance needs.
- API integration
- Open weights
- Self-hosting
📖 Guide
How to use MiniMax M3
Follow the steps below to get started quickly and complete your first AI chat or Agent task with MiniMax M3.
-
Choose thinking mode
Choose thinking mode by task: enable for complex reasoning, Agents, and multi-step collaboration; disable for chat completion and low-latency scenarios.
-
Write your prompt
Write structured system prompts and user instructions with clear role, tool permissions, and output format—MiniMax M3 responds more stably to clear task descriptions.
-
Configure API parameters
For long-context tasks, pass complete materials in one go (codebase summaries, full documents, or chat history) to use the full 1M-token window.
-
Iterate and refine
Call via MiniMax API (OpenAI/Anthropic-compatible) or self-hosted weights; for Agents, fully retain assistant tool calls and reasoning chains in conversation history.
❓ FAQ
MiniMax M3 FAQ
Below are the most common questions about MiniMax M3 features, context, multimodal support, and integration.
What is MiniMax M3, and who is it for?
MiniMax M3 is MiniMax's open-weight AI chat LLM for coding, Agent workflows, and long-context multimodal tasks. It suits developers, AI engineers, technical teams handling large codebases or long documents, and enterprises wanting API access or self-hosted deployment.
How large is MiniMax M3's context window?
The MiniMax M3 API supports up to 1M tokens of context (input plus output combined), with an officially stated guaranteed minimum of 512K. Ideal for whole-repo analysis, long reports, and Agent sessions that need full history.
Which input modalities does MiniMax M3 support?
MiniMax M3 is a native multimodal model supporting text, image, and video input with text output—for visual Q&A, video understanding, and cross-modal Agent tasks with text instructions.
How do I switch MiniMax M3 thinking mode?
Control via the API thinking parameter: on for extra reasoning suited to complex Agents and coding; off for faster responses in chat and code completion. Modes can switch independently per request.
What upgrades does MiniMax M3 bring vs. MiniMax M2?
MiniMax M3 introduces MSA sparse attention: at 1M context, compute is about 1/20 of M2 with much faster Prefill/Decode; it also expands to native multimodal and stronger coding/Agent capability, leaping context to million-token scale.
How do I access MiniMax M3?
Call MiniMax M3 online via the official MiniMax API (OpenAI- and Anthropic-compatible formats), or download open weights from Hugging Face for self-hosting. Recommended parameters per official guidance: temperature=1.0, top_p=0.95.
💡 Tips
MiniMax M3 prompt and integration tips
Enable thinking mode for complex agents, cross-file coding, and long document analysis; disable for real-time chat and completion to reduce latency.
For Agent sessions, append full assistant replies (including tool calls and reasoning) to history so MiniMax M3 maintains coherent context.
For long-context tasks, pass complete materials in one request—fully use MiniMax M3's 1M-token window and avoid over-chunking that loses information.
Recommended API parameters per official guidance: temperature=1.0, top_p=0.95; keep output tokens within 131072 for more stable experience.
For multimodal tasks, specify visual elements and time ranges to focus on—MiniMax M3 responds more accurately to structured multimodal instructions.
Ready to try MiniMax M3?
Get started now and unlock your AI creative potential with MiniMax M3