bikabika AI
MiniMax June 1, 2026

MiniMax M3

1M context · native multimodal · flagship MoE for coding and agent workflows

MiniMax M3 is MiniMax's open-weight AI chat LLM released in June 2026, using a 428B-parameter MoE architecture (~23B activated), with up to 1M-token context and native multimodal text, image, and video input. With MiniMax Sparse Attention (MSA), MiniMax M3 balances speed and quality on long-context coding and Agent tasks, reaching frontier levels in coding, tool use, and long-horizon collaboration. Whether you are a developer, AI engineer, or enterprise team, you can learn about and integrate MiniMax M3 online for intelligent chat and automation workflows.

Context window

1M tokens

Total parameters

428B MoE

Active parameters

~23B

Input modalities

Text/image/video

💡 Why Choose

Why choose MiniMax M3?

MiniMax M3 is MiniMax's AI chat model officially released in June 2026, positioned around "long context + native multimodal + coding agents." It uses a 428B-parameter MoE architecture with ~23B active parameters per inference, offering up to 1M-token API context—fitting medium-to-large codebases, long documents, or full Agent conversation history into a single model view.

MiniMax M3's core architectural innovation is MiniMax Sparse Attention (MSA). Compared with traditional full attention, MSA reduces per-token compute at million-token scale to roughly 1/20 of predecessor M2, with ~9× Prefill and ~15× Decode acceleration—making "ultra-long context" a practical productivity feature rather than a theoretical spec. MiniMax M3 also mixes text, image, and video from the first training step, making it one of few open-weight chat models with native multimodal understanding.

On bikabika AI, you can explore MiniMax M3 core capabilities, thinking mode differences, typical use cases, and integration in one place to quickly assess whether it fits your coding, Agent, or multimodal chat needs—and try it online.

⚡ Features

MiniMax M3 core features

The six capabilities below form MiniMax M3's core competitive strengths in AI chat, coding, and agents.

1M-token ultra-long context

MiniMax M3 supports up to a 1M-token context window (officially guaranteed minimum 512K), loading large codebases, long documents, or multi-turn Agent history for repo-scale global reasoning.

Native multimodal input

MiniMax M3 mixes text, image, and video modalities from the first training step with deeper semantic fusion—for video understanding, visual Q&A, and cross-modal Agent tasks.

Frontier coding and Cowork capability

MiniMax M3 excels on long-horizon Agent benchmarks and coding—suited to code generation, review, refactoring, and multi-file collaborative development, and pairs with MiniMax Code and other Agent products.

MSA sparse attention acceleration

MiniMax M3 uses MiniMax Sparse Attention (MSA): at 1M context, Prefill is ~9× faster and Decode ~15× faster, with per-token compute about 1/20 of prior M2.

Switchable thinking mode

MiniMax M3 can turn thinking mode on or off: on for complex reasoning and Agent tasks; off for lower latency in chat completion and high-frequency calls.

Tool calling and Agent workflows

MiniMax M3 is optimized for tool use and Interleaved Thinking—re-evaluating state after tool execution—suited to multi-step automation and Computer Use apps.

Want to experience MiniMax M3 long-context coding yourself?

MiniMax M3 supports API and self-hosted deployment—start your first million-context AI chat experience now.

🔄 Compare

MiniMax M3 thinking mode comparison

MiniMax M3 supports per-request thinking mode switching—the comparison below helps you choose between complex agents and low-latency scenarios.

MiniMax M3 thinking mode comparison
DimensionThinking mode onThinking mode off
Best forComplex reasoning, agents, long-horizon collaborationChat, code completion, high-frequency calls
Response latencyHigher (quality-first)Low latency, high throughput
Reasoning depthDeeper, suited to multi-step tasksDirect output, speed-first
Agent performanceMore stable tool calls and interleaved thinkingBetter for lightweight interaction
Coding tasksComplex refactors, cross-file analysisSingle-line/local completion
API pricingSame as off modeSame as on mode
Switch methodEnable thinking parameterDisable thinking parameter
Recommended pairingMiniMax Code AgentReal-time chat bots

💎 Highlights

MiniMax M3 technical highlights

  • MiniMax M3 supports up to 1M-token context for repo-scale code and long-document analysis
  • MiniMax M3's 428B MoE with ~23B activated balances capability density and inference efficiency
  • MiniMax M3 MSA sparse attention sharply reduces compute and KV Cache pressure at million-token context
  • MiniMax M3 is natively multimodal: unified modeling of text, image, and video input
  • MiniMax M3 thinking mode switches per request for complex Agents vs. low-latency scenarios
  • MiniMax M3 offers OpenAI/Anthropic-compatible APIs and open Hugging Face weights for self-hosting

🎯 Use Cases

MiniMax M3 use cases

From repo-scale coding to multimodal agents, MiniMax M3 delivers stable, efficient long-context AI chat and reasoning across the six scenarios below.

Developers · Architects

Large codebase understanding and refactoring

MiniMax M3's 1M-token context can load key modules of medium-to-large monorepos in one pass—cross-file dependency analysis, refactor design, and code review without chunked RAG.

  • 1M tokens
  • Codebase
  • MiniMax M3

Content teams · Product managers

Long video and multimodal analysis

MiniMax M3 natively supports image and video input—combine with text instructions for video summarization, shot understanding, visual Q&A, and cross-modal extraction for media and creative workflows.

  • Multimodal
  • Video understanding
  • MiniMax M3

AI engineers · Technical teams

Coding agents and automated development

MiniMax M3 reaches frontier levels on long-horizon agent benchmarks and coding tasks—pair with MiniMax Code and similar Agent products for multi-step development, testing, and deployment automation.

  • Coding agents
  • MiniMax Code
  • Tool calling

Developers · Automation engineers

Long-horizon agent workflows

MiniMax M3 is optimized for interleaved thinking—re-evaluating state after tool execution, ideal for Agent pipelines needing multi-round tool calls, state tracking, and result aggregation.

  • Agent workflow
  • Tool calling
  • MSA

Product teams · RPA engineers

Computer Use and desktop automation

MiniMax M3 supports Computer Use-style capabilities—combine visual input to understand UI state and plan actions, extending AI chat to desktop-level automation.

  • Computer Use
  • Desktop automation
  • MiniMax M3

Enterprise IT · Ops teams

Enterprise API integration and self-hosting

MiniMax M3 offers OpenAI / Anthropic-compatible API and open Hugging Face weights—enterprises can choose cloud calls or private deployment per compliance needs.

  • API integration
  • Open weights
  • Self-hosting

📖 Guide

How to use MiniMax M3

Follow the steps below to get started quickly and complete your first AI chat or Agent task with MiniMax M3.

  1. Choose thinking mode

    Choose thinking mode by task: enable for complex reasoning, Agents, and multi-step collaboration; disable for chat completion and low-latency scenarios.

  2. Write your prompt

    Write structured system prompts and user instructions with clear role, tool permissions, and output format—MiniMax M3 responds more stably to clear task descriptions.

  3. Configure API parameters

    For long-context tasks, pass complete materials in one go (codebase summaries, full documents, or chat history) to use the full 1M-token window.

  4. Iterate and refine

    Call via MiniMax API (OpenAI/Anthropic-compatible) or self-hosted weights; for Agents, fully retain assistant tool calls and reasoning chains in conversation history.

❓ FAQ

MiniMax M3 FAQ

Below are the most common questions about MiniMax M3 features, context, multimodal support, and integration.

What is MiniMax M3, and who is it for?

MiniMax M3 is MiniMax's open-weight AI chat LLM for coding, Agent workflows, and long-context multimodal tasks. It suits developers, AI engineers, technical teams handling large codebases or long documents, and enterprises wanting API access or self-hosted deployment.

How large is MiniMax M3's context window?

The MiniMax M3 API supports up to 1M tokens of context (input plus output combined), with an officially stated guaranteed minimum of 512K. Ideal for whole-repo analysis, long reports, and Agent sessions that need full history.

Which input modalities does MiniMax M3 support?

MiniMax M3 is a native multimodal model supporting text, image, and video input with text output—for visual Q&A, video understanding, and cross-modal Agent tasks with text instructions.

How do I switch MiniMax M3 thinking mode?

Control via the API thinking parameter: on for extra reasoning suited to complex Agents and coding; off for faster responses in chat and code completion. Modes can switch independently per request.

What upgrades does MiniMax M3 bring vs. MiniMax M2?

MiniMax M3 introduces MSA sparse attention: at 1M context, compute is about 1/20 of M2 with much faster Prefill/Decode; it also expands to native multimodal and stronger coding/Agent capability, leaping context to million-token scale.

How do I access MiniMax M3?

Call MiniMax M3 online via the official MiniMax API (OpenAI- and Anthropic-compatible formats), or download open weights from Hugging Face for self-hosting. Recommended parameters per official guidance: temperature=1.0, top_p=0.95.

💡 Tips

MiniMax M3 prompt and integration tips

  • Enable thinking mode for complex agents, cross-file coding, and long document analysis; disable for real-time chat and completion to reduce latency.

  • For Agent sessions, append full assistant replies (including tool calls and reasoning) to history so MiniMax M3 maintains coherent context.

  • For long-context tasks, pass complete materials in one request—fully use MiniMax M3's 1M-token window and avoid over-chunking that loses information.

  • Recommended API parameters per official guidance: temperature=1.0, top_p=0.95; keep output tokens within 131072 for more stable experience.

  • For multimodal tasks, specify visual elements and time ranges to focus on—MiniMax M3 responds more accurately to structured multimodal instructions.

Ready to try MiniMax M3?

Get started now and unlock your AI creative potential with MiniMax M3