A sparse mixture-of-experts language model I wrote from scratch in pure PyTorch, in the spirit of Karpathy's makemore. The expert modules are the easy part. Most of the work went into the noisy top-k gating and the initialization, which are what decide whether an MoE trains at all. Meant as a readable reference for the architecture behind Mixtral, DBRX and Grok.
The NeurIPS 2024 tutorial on dynamic sparsity builds its mixture-of-experts notebook on this, says so in the notebook, and trains on data pulled straight from the repo. The M2L Summer School 2024 lists it as assigned reading.