
Diffusion from Zero: learning generative models from the ground up
Why I started this series as an electrical engineering student with no probability course behind me, how the parts fit together, and the resources that helped most.
Loading the interactive instruments.
Loading the interactive instruments.
THE SERIES
Building up to diffusion models and Diffusion Policy, one concept at a time.

Why I started this series as an electrical engineering student with no probability course behind me, how the parts fit together, and the resources that helped most.

Why a deep-learning series has to start with probability: what distributions, Bayes' theorem, expectation, and Jensen's inequality each contribute to generative models.

Maximum likelihood explains the losses we use every day: a Gaussian assumption gives MSE, a categorical one gives cross-entropy, and both connect to KL divergence.

p(x), p(z), p(x|z), p(z|x), q, and p-theta: what each expression in VAEs and diffusion models describes, explained with concrete examples instead of derivations.

How an autoencoder compresses and reconstructs its input, why the bottleneck matters, what reconstruction loss means in probability terms, and why a plain autoencoder is hard to sample from.

Why an autoencoder isn't enough for generation, how a VAE learns a latent distribution, what each term of the ELBO does, and how the reparameterization trick makes it trainable.
Following Calvin Luo's unified perspective in order: the ELBO, hierarchical VAEs, variational diffusion models, the three equivalent prediction targets, and the score-based view.

Sections 1 to 3 of Diffusion Policy: why behavior cloning needs multimodal action distributions, how DDPM becomes a visuomotor policy, and the CNN, transformer, and FiLM design choices.
Sections 4 to 9 of Diffusion Policy: modes and basins, why learning the score removes the normalizing constant, the benchmark and real-robot results, and the limits the authors state.