MiniMax M3.1 Flash Preview
Officialminimax/minimax-m3.1-flash-preview
MiniMax M3.1 Flash Preview — MiniMax's newest coding model, 1M context, thinking always on with five effort levels.
- Context window
- 1M
- Max output
- 524K
- Input modalities
- text · image
- Output modalities
- text
- Capabilities
- Vision · Reasoning · Tool calling · Coding · Prompt cache
- Upstream release
- 2026-09-27
- RouteMux name
- minimax/minimax-m3.1-flash-preview
- Official ID
- MiniMax-M3.1-Flash-Preview
Try MiniMax M3.1 Flash Preview here
Tune the request, see the current price ceiling, and inspect the real response without leaving this model page.
Sign in first to upload images to this vision model.
Your result will appear here
Enter an input on the left. The same RouteMux Playground execution and billing path powers this embedded test.
What is MiniMax M3.1 Flash Preview
MiniMax M3.1 Flash Preview is the first model in MiniMax's M3.1 line, released on September 27, 2026. MiniMax tunes it for speed and stability in everyday development work — bug fixes, feature builds, agentic coding — rather than for the hardest one-off problems.
It keeps the M3 footprint: a 1M-token context window, text and image input, tool calling and prompt caching, answering on OpenAI Chat Completions, OpenAI Responses and Anthropic Messages. What changes is thinking. It is always on and cannot be switched off, and the depth is set with a reasoning effort of low, medium, high, xhigh or max. When you omit it, MiniMax defaults to max.
MiniMax only offers this model through its subscription plans, not on its pay-as-you-go price list, so there is no official per-token price to compare against. It is also a preview: MiniMax may change or retire it.
MiniMax M3.1 Flash Preview API pricing
Source: minimax_m3_parity · Verified 2026-09-29
Cache read $0.006/M · Cache write —/M
| Tier | Input | Cache read | Cache write | Output |
|---|---|---|---|---|
| Standard ≤512K | $0.03 | $0.006 | — | $0.12 |
| Standard >512K | $0.06 | $0.012 | — | $0.24 |
Estimate cost
Limited time20% extra Credits on your first top-up — up to 200 bonus Credits
Applies to your first top-up within 72 hours of signing up. Minimum $20, once per account.
Key capabilities
Five effort levels
low / medium / high / xhigh / max. The default is max — set a lower level explicitly for latency-sensitive work.
1M-token context
With text and image input.
Tool calling and caching
Native tool calling and automatic prompt caching.
Three protocols
OpenAI Chat, OpenAI Responses and Anthropic Messages.
Interface and protocols
Use any of these IDs to call this model via the API.
minimax/minimax-m3.1-flash-previewMiniMax-M3.1-Flash-PreviewReplace the ROUTEMUX_KEY placeholder with your API key. Create one →
from openai import OpenAI
client = OpenAI(
base_url="https://api.routemux.com/v1",
api_key="<ROUTEMUX_KEY>",
)
completion = client.chat.completions.create(
model="minimax/minimax-m3.1-flash-preview",
messages=[{"role": "user", "content": "What is the meaning of life?"}],
)
print(completion.choices[0].message.content)Performance (last 30 days)
MiniMax M3.1 Flash Preview FAQ
›What is MiniMax M3.1 Flash Preview?
MiniMax M3.1 Flash Preview — MiniMax's newest coding model, 1M context, thinking always on with five effort levels.
›How large is the context window of MiniMax M3.1 Flash Preview?
MiniMax M3.1 Flash Preview supports a context window of up to 1M tokens, with up to 524K output tokens per request.
›How much does the MiniMax M3.1 Flash Preview API cost?
Through RouteMux, MiniMax M3.1 Flash Preview costs $0.03 for input and $0.12 for output (USD per 1M tokens), below the official list price.
›How do I call MiniMax M3.1 Flash Preview via API?
MiniMax M3.1 Flash Preview is OpenAI-compatible: point your base URL at RouteMux and set the model field to minimax/minimax-m3.1-flash-preview — no code changes needed.
›Which input and output modalities does MiniMax M3.1 Flash Preview support?
MiniMax M3.1 Flash Preview accepts text, image as input and produces text as output.
›When was MiniMax M3.1 Flash Preview released?
MiniMax M3.1 Flash Preview was released on September 27, 2026.
›Can I turn thinking off?
No. MiniMax rejects thinking disabled and an effort of "none" for this model with a 400 error. Use effort "low" for the lightest thinking.
›How do I set the effort level on each protocol?
OpenAI Chat Completions: reasoning_effort. OpenAI Responses: reasoning.effort. Anthropic Messages: output_config.effort. Thinking tokens are billed as output tokens.
›How is it different from MiniMax M3?
Same context window and protocols. M3.1 Flash Preview always thinks and offers two extra effort levels (xhigh and max); M3 lets you turn thinking off and stops at high. M3 is a stable release; this one is a preview.