MiMo-V2.6 notes

Architecture

Very simple main body. Standard transformer backbone with visual + audio encoders connected through lightweight projectors.

Repeated blocks that interleave local sliding-window attention (SWA, window size 128) with global attention (GA): N SWA blocks followed by a GA block.

But the first transformer block uses global attention and a dense FFN (instead of MoE).

No shared experts. Have MTP blocks too (dense).

MiMo-ViT

Some details here, but what is interesting is how they pretrain it. They pair it with a small pretrained LLM and optimize on normal cross-entropy for multimodal data understanding.

∴ no contrastive learning & stuff.

Over 4T vision tokens for pretraining it.

Audio encoder

20m hours. RVQ after some SWA + GA stuff & downsampling. See MiMo-Audio for the recipe.

RVQ again, nice.

Table 1 has all their numbers.
Table 1. MiMo-V2.6 model configurations
ConfigurationFlashPro
Main block
Layers (total / SWA / GA)48 / 39 / 970 / 60 / 10
Hidden size40966144
SWA heads (Q / KV)64 / 8128 / 8
Sliding window size128128
GA heads (Q / KV)64 / 4128 / 8
Head dimensions (QK / V)192 / 128192 / 128
Experts (total / activated)256 / 8384 / 8
Total parameters310B1.02T
Active parameters15B42B
MiMo-ViT
Layers (total / SWA / GA)28 / 24 / 4
Hidden size1280
Attention heads (Q / KV)32 / 8
Head dimension64
Patch size (T × H × W)2 × 16 × 16
Sliding window size (left / right)64 / 64
Spatial merge size2 × 2
Parameters681M
Audio tokenizer encoder
Layers (total / SWA / GA)24 / 12 / 12
Hidden size1024
Attention heads (Q / KV)16 / 16
Head dimension64
Mel bins128
Sliding window size128
Codebooks20
Parameters308M
Audio patch encoder
Layers6
Hidden size1024
Attention heads (Q / KV)16 / 16
Head dimension64
Attention group size4
Parameters127M
Speculative decoder
Layers (total / SWA / GA)5 / 5 / 05 / 5 / 0
Hidden size40966144
SWA heads (Q / KV)64 / 8128 / 8
Sliding window size10241024
GA heads (Q / KV)64 / 4128 / 8
Head dimensions (QK / V)128 / 128128 / 128

From Table 1 of the report, p. 6. MTP modules are excluded from main-block layer counts. Encoder counts include input embeddings, exclude projectors, and (for the audio tokenizer encoder) exclude EMA codebooks. Merged cells apply to both models.

MTP for speculative decoding gets 7 tokens ahead in a single forward pass (action chunk).

Pretraining

30T tokens for Pro (48T for Flash). First train text only, and later omni-modal all together with the pretrained encoders. Use AdamW.

Started at 32k context and then went to 256k.

Mid-training

For agentic tasks + long context, to make the RL better afterwards.

Start at 256k and then extend to 1m.

Switch from AdamW to Muown for the hidden weight matrices because AdamW is less efficient for the big RL batches they do later.

Use MXFP4 quantization-aware training (QAT) here.

RL

Big huge scale: 1,568 prompts × group size 16.

Objective:

ℒ(θ)=− 𝔼q∼⋃d𝒟d{oi}i=1G∼μθold(·|q) [1∑i=1G|oi| ∑i=1G∑t=1|oi| ri,tMi,tAilogπθ(oi,t|q,oi,<t)]

where r is the importance-sampling ratio, M is the token-level mask, and A is the advantage.

Per token,

ri,t=sg[πθ(oi,t)μθold(oi,t)]

They have separate clipping bounds for positive and negative advantages, adjusted based on policy entropy. Want to keep entropy high enough, but also not too high.

RL stages

Rollout. For every prompt q drawn from the union of task datasets ⋃d𝒟d, the rollout policy μθold generates a group of G candidate solutions {oi}i=1G.

Grading. They use agent eval to score the solutions and their problem-solving behaviour.

Training. They use the collected trajectories and resulting advantages to update the policy.

Compute breakdown for Pro:

43.8% rollout / 43.5% training / 12.7% grading.

Coding data

From GitHub, get PRs and their issues. If the issue describes the task well enough, use it as the task specification. Otherwise have an LLM reconstruct the task from the patch + repo context, without giving away the solution.

For vibe coding, they use requests from employees and the LLMs make the test cases. Also use existing code to derive tasks: describe its functionality and have the agent implement it.

Rerun a bunch, and they have audit agents.

General agent tasks

Use sandboxes for envs so they are easy to reset. Other agent(s) make the envs.

An agent takes human-made “seed” tasks and uses these to construct new tasks suited to the contents + tools in an env.

Rubrics evaluate completion: code checks for deterministic things, but also LLM checks for more open-ended content. A review agent revises the rubric based on rollouts.

Visual tasks

For open-ended visual tasks, once the rubrics stabilize, they use groups for grading: compare the rendered artifacts for the same query.

Cyber / OSS-Fuzz

Have to make sure that you trigger the correct thing + at the correct location. Not just any crash.

Multi-harness

Can’t use common ones like Codex: too many constraints & safeguards, and the model could learn to ignore these if it doesn’t help the reward. Also not modular.

∴ use lots of mini-harnesses:

system prompt + tools + context management.

Fixing (decreasing) reward hacking

Mid-training

Basically take the reward-hacked examples, have MiMo reflect on them & fix them, and then continue.

Makes the error and correction explicit, so it doesn’t just ignore it.

This seems like a simple way to help alignment: just have it in the data. But I guess the RL pressure might destroy it later.

Environment prep

Remove build logs & binaries & stuff, later git commits, etc.

Plus instructions against cheating, and container-level network isolation.

Hack agent

Give it examples of how to hack and have it probe etc.

Training-time auditing

The grader sets the effective reward for confirmed hacking trajectories to 0. Offline auditor checks also guide fixes to the environments.

Groupwise Agentic Grading

There are two main pieces:

Groupwise Reward Synthesis (GRS)

Take a subset of high-pass-rate tasks and make task-specific rubrics, which are then used in training with the test rewards to improve quality.

GRS produces solution rubrics and behaviour rubrics.

These are usually evaluated by a grader agent. It assigns:

  • a solution score Sisol for the quality of the implementation
  • a behaviour score Sibeh for how the agent approached & verified its work

So:

Ri=Ritest·Sisol·Sibeh

Ritest is binary.

Groupwise Advantage Redistribution (GAR)

For the rest of the code-agent tasks, i.e. mixed-outcome groups.

An online, SFT-trained grader jointly examines successful and failed trajectories in each group. Looks at all the guys. Can run tests etc. + look for cheats.

It ranks passing solutions and redistributes positive advantage towards higher-quality passing trajectories.

Confirmed hacks get effective reward 0 before recomputing group statistics. Let Ai=Ri−R¯, where R¯ is the group mean. Then:

𝒫={i:Ri=1},fi∈(0,1].

The uncapped redistribution is:

λ=∑i∈𝒫Ai∑i∈𝒫fiAi

and

Ai′={λfiAi,i∈𝒫Ai,i∉𝒫.

So it moves advantage mass amongst the good trajectories.

In practice they cap the common rescaling factor, then subtract the group mean again. Grading runs async.

They also penalize length relative to the group (once there is a certain amount of success).

They don’t reward bad / malformed tokens within a trajectory as highly. Change their advantage to make them more negative or less positive.

Actually, if Ai>0, they mask bad tokens and shift the weight elsewhere. If Ai<0, they upweight them. Total advantage mass for each sign stays the same when the scales aren’t clipped and the denominators are nonzero.

A bunch of basically regularization.

Also: freeze the MoE router, otherwise training messes up in RL.

MOPD2

Multi-Prefix Multi-Teacher On-Policy Distillation.

Used to combine teachers trained on different, often hard-to-verify, tasks.

Provide token-level supervision on the student rollout. Prefixes can come from SFT / teacher data; the student generates its own next turn from that history.

Good for hard-to-verify tasks: you can SFT a teacher on them, but not drag the student too OOD.

My shorthand for the token-level distillation signal:

logπT(yt|st)−logπθ(yt|st)

Finer-grained learning signals

Section 6.1: they organize the trajectories into a hierarchy:

Sample → Sequence → Context → Segment.

So: a prompt → a rollout / Agent Loop → a dialogue branch → a turn. Only model-generated turns contribute to the loss.

A penalty module can mask parts or shape their advantages. Rules can be handwritten or model-based. Also keeps infrastructure failures out of the training signal.

Not totally clear to me how all of this fits together, but pretty cool. Much more fine-grained rewards, which is nice and important—not just the scalar they construct.

Not one grader. Feels like they aren’t giving us that much detail on the exact choices here, which is interesting.

Harness / RL infrastructure

Harness pool has many harness instances and Agent Loops running in it, across processes. Each host process can carry many concurrent instances. Separate from that is the inference engine.

Need to be careful about the mismatch between training & inference engine: record the top-p candidate sets and MoE expert indices during rollout, then replay them during training.

The basic interaction is a loop that:

calls the inference engine on some history → scans the output for tool calls → appends results to the history → repeats.

The Agent Loop owns setup, interaction, reward evaluation, and cleanup. The harness drives it through a request endpoint.

It gets its initial prompt from the Sample Mixer, which basically handles scheduling + balancing how much work each source gets. Otherwise hard tasks / things that run longer won’t be represented correctly. Sample Mixer also ensures the correct task mixture, accounting for rollout duration and how often groups get filtered out.

Trajectories go to distributed storage. The scheduler mostly handles metadata for grouping; packers pull the needed data from storage for the trainer.


Nice! So: big RL + dense rewards. Much better than binary, and better than just a scalar!

Nice, they released a lot of the stuff. Pretty interesting.

Also funny: no fear of hacking the world.