Expert selectionby decoder depth

Early expert selection differs by modality. Cross-modal output transfer becomes clear between decoder layers 31 and 32.

TLDR

The router selects more of the same experts in deeper layers.

The router assigns each token to six experts. Here, modality means speech, image, or text. We compare the expert IDs across modalities.

In layers 2 to 9, the selected IDs differ by modality. Between layers 10 and 34, the lists become more similar. After layer 35, the router often selects similar lists for all three modalities.

Changing the early expert groups increases loss for every task. Output swaps locate another transition between layers 31 and 32.

At layer 32, matching-transcript routed outputs transfer better than geometry-matched different transcripts.

This shows cross-modal functional compatibility. It does not show full modality independence.

Question

At which layers does expert selection align across speech, image, and text?

Inkling-Small has one decoder for all three modalities. It has no separate vision or audio Transformer.

The router assigns each token to six of 256 experts. An expert is a small processing block inside the decoder.

We provided the same sentence as speech, a rendered image, and text.

We recorded expert selection and expert outputs at each decoder layer. We changed expert groups and swapped outputs at six selected layers.

These measurements separate expert addresses, output vectors, and effects on prediction loss.

Expert selection for one token

Select text, image, or audio. Then select one token. The chart lists the six selected experts at four layers.

LibriSpeech sample 1272-128104-0000
MISTER QUILTER IS THE APOSTLE OF THE MIDDLE CLASSES AND WE ARE GLAD TO WELCOME HIS GOSPEL

8 shown from 28 text tokens

A left mark indicates a leading space.

Selected tokenText token 15, “CL” with leading space

Read from top to bottom. Each block is one selected expert. Block width represents routing weight.

L2early layer
E323.7%
E14121.3%
E16319.6%
E12012.8%
E20712.3%
E6010.3%
6 experts selected
L21middle layer
E8619.2%
E19318.1%
E4516.9%
E13216.7%
E17115.5%
E5413.6%
L30middle layer
E10338.8%
E19127.0%
E10815.5%
E298.8%
E2095.3%
E2334.7%
L41late layer
E16325.7%
E5324.1%
E313.3%
E20113.1%
E1011.9%
E11711.9%
Interpretation

For this token, E3 has the highest weight at layer 2. E163 has the highest weight at layer 41. No expert ID repeats between consecutive shown layers.

The six saved weights add to 100% in each row. Color identifies the expert ID.

Convergence

Expert lists became more similar with depth.

We tested 18 transcripts. Expert selection differed by modality in layers 2 to 9.

In layers 10 to 34, matching transcripts produced more similar expert lists.

After layer 35, the same expert groups often received tokens from all three modalities.

This pattern is consistent with a transition from modality-specific processing to cross-modal content processing. Expert outputs test whether the late computation also aligns.

Expert-use profiles are more similar in later layers.

Each modality produces a 256-number profile at each layer. A score of 0 means different profiles. A score of 1 means identical profiles.

Expert-use similarity by decoder layerAt layer 2, all three scores range from 0.20 to 0.30. At layer 41, they range from 0.82 to 0.90. A red rule marks the start of the middle-depth increase at layer 10.modality splitlists convergesame experts across modalities0.000.250.500.751.00210203041decoder layerspeech + image0.90image + text0.84speech + text0.82routing similaritylayer 10
Similarity is 1 minus Jensen-Shannon divergence between normalized expert-use profiles. Bands show 95% bootstrap intervals across 18 transcripts. The red rule marks layer 10, where the middle-depth increase begins.
Test

We changed the experts at each stage.

Six transcripts identified the expert groups. Twelve different transcripts tested each change.

EarlyLayers 2-9

Modality separation. Expert selection differed by modality.

MiddleLayers 10-34

Matching transcripts. Expert lists were more similar for inputs with the same transcript.

LateLayers 35-41

Similar lists. Expert lists were often similar across modalities.

Route similarity alone does not establish causal importance. We blocked each group and measured the change.

We also swapped early groups between speech, image, and text.

The control condition changed other experts with similar overall use.

Early-group changes increased loss in all three tasks.

“Worse” marks a clear loss. “No clear change” includes small or uncertain differences.

TaskEarlyMiddleLate
Read speechWorseNo clear changeNo clear change
Read image textWorseNo clear changeNo clear change
Predict textWorseWorseNo clear change
Early-group changes increased loss in all three tasks. Middle-group changes increased only text-prediction loss. Late-group changes had no clear effect.
See exact effects and confidence intervals
Task
Earlymodality-specific experts
Middlesame-transcript experts
Latecommon experts
ASR
+0.218[+0.105, +0.332]control +0.027
+0.014[-0.027, +0.051]control +0.027
-0.007[-0.022, +0.003]control -0.011
OCR
+0.354[+0.205, +0.524]control +0.072
+0.000[-0.000, +0.001]control +0.007
-0.000[-0.000, +0.000]control +0.000
Text
+0.107[+0.012, +0.212]control +0.065
+0.212[+0.104, +0.335]control -0.097
+0.013[-0.009, +0.035]control -0.008
Values are changes in prediction loss. Positive values are worse. Brackets show 95% paired bootstrap intervals.

Same-transcript retrieval accuracy also decreased.

This score measures whether two inputs contain the same words. Higher values indicate better retrieval. The control value did not decrease.

Original model72.2%
Early experts removed52.8%
Early experts swapped51.4%
Control change87.5%

In four generated examples, speech word error rose from 4.8% to 10.0%. Image-text character error rose from 0% to 4.7%.

Computation

The same expert IDs can return different vectors.

At layer 41, modality pairs shared 68% to 70% of their routed expert IDs. The corresponding output vectors had 0.20 to 0.28 cosine similarity. #

An expert ID is an address. The output also depends on the vector that enters the expert. #

Routing similarity rises while output similarity falls.

speech + imagespeech + textimage + text
Routed expert addresses and output similarity in layers 35 to 41At layer 41, modality pairs share about 69 percent of routed expert IDs. Same-expert output cosine similarity ranges from 0.20 to 0.28.Same routed expert IDs0.68-0.70 at layer 410.000.250.500.751.0035363738394041Same-expert output similarity0.20-0.28 at layer 410.000.250.500.751.0035363738394041decoder layerSame addressesdo not meansame outputs.
Means across 18 transcripts at aligned answer positions. Each token uses six routed experts.

Shared experts account for most of the final update.

The shared-expert output has larger magnitude and higher cross-modal similarity. The routed-expert output is a smaller correction. #

Layer-41 outputSize relative to residualOutput similarity
Two shared experts0.46-0.490.85-0.87
Six routed experts0.025-0.0400.22-0.27

This scale difference explains the weak late ablation. Removing most routed weight changed only a small part of the layer update. #

Output transfer becomes clear between layers 31 and 32.

The swap replaced outputs at layers 21, 31, 32, 33, 36, and 41. The controls used another transcript or an equal-norm random vector. #

Input tokens, attention, router decisions, and every other layer output stayed unchanged. #

Layer 41 endpoint

Replaced outputDifferent transcriptRandom direction
Two shared experts+0.0318[+0.0003, +0.0763]+0.0362[+0.0017, +0.0783]
Six routed experts+0.0058[+0.0037, +0.0078]+0.0028[-0.0005, +0.0060]

Values show extra prediction loss relative to a matching cross-modal swap. Positive values mean that the matching swap caused less loss. Brackets show 95% intervals across eight transcripts. #

Final-state similarity after a shared-output swap

Matching transcript
0.938
Different transcript
0.864
Random direction
0.651

Matching geometry does not explain the layer-32 result.

We searched 73 transcripts for different-transcript donors with similar hidden states and routing profiles. The matched test used 166 token rows from 13 transcripts. #

The mean hidden-similarity difference was -0.0072. The mean routing-profile difference was +0.0040. We selected the rows before we measured prediction loss. #

Layer 32 after geometry matching

Replaced outputDifferent transcript minus matching transcript95% interval
Six routed experts+0.0307[+0.0077, +0.0580]
Two shared experts+0.0255[-0.0005, +0.0555]

The routed interval is above zero. A matching transcript supplies a more useful routed output than a geometry-matched different transcript. This shows cross-modal functional compatibility at layer 32. #

This result does not show full modality independence. The outputs can retain modality information that this test does not measure. #

Router similarity alone does not prove expert specialization. Wang, Hayou, and Nalisnick (2026) show that similar routing can follow hidden-state geometry rather than domain-specific expert functions. #

Xi Wang, Soufiane Hayou, and Eric Nalisnick. “The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise.” arXiv:2604.09780, 2026. #

Answer

Cross-modal output transfer becomes clear at layer 32.

Expert selection differed by modality in layers 2 to 9. Blocking these groups increased loss for every task.

Expert lists for matching transcripts became more similar in layers 10 to 34. Changes in this range increased only text-prediction loss.

At layer 31, only one control interval excluded zero for each output component. Both routed-output intervals excluded zero at layer 32.

Both shared-output intervals excluded zero at layer 33. The same criterion held for one component at layers 36 and 41.

Every tested depth from layer 32 onward met the criterion. The output component differed by depth.

At layer 32, the routed-output result remains clear after matching hidden-state direction and routing profile.

The late computation contains a shared transcript signal. It can still contain modality information.

Complementary pair

Drag the selected point. The linked point tracks its exact HSL complement.