Abstract
This lecture introduces continuous generative modeling, diffusion and flow matching, as the dominant paradigm for visual synthesis, and traces its extension from static images to dynamic, embodied domains. We start with the core mathematical picture: forward noising processes, score/velocity-field learning, and the shared ODE/SDE view that unifies diffusion models and flow matching as a generative modeling framework that learns a velocity field to transform a source distribution into a target distribution through a continuous flow. From there, we move to video diffusion models and their emerging role as implicit world models, before turning to robotics: Diffusion Policy as a case study in using denoising processes to represent multimodal action distributions, and video generation as a substrate for planning and simulation in physical AI. The lecture closes with open questions around sampling efficiency, controllability, and what it means for a generative model to also be a world model.
Schedule
- Lecture: Mondays, 4pm–6pm (starts at 4 pm)
- Tutorial: Wednesdays, 12pm–2pm (starts at 12:15)
More information will follow soon.