2026-09-03
Latent MoE: Compressing the Experts Before You Route Them
NVIDIA's Nemotron 3 Super compresses tokens into a low-rank latent space before routing to experts. Same inference cost, four times more specialists. That changes what MoE architectures are for.