Abstract
In my talk I will examine when learning to act in a dynamical system is no harder than supervised learning, despite feedback, distribution shift, and long-horizon consequences. We will see that reinforcement learning can achieve supervised-learning-type regret rates under three structural conditions: (i) imitation learning / offline RL when distribution shift and simultaneity bias are controlled (highlighting failure modes due to compounding error and practical design constraints), (ii) non-episodic online RL when models are identifiable via persistence of excitation - yielding regret bounds for nonlinear continuous-state-action systems through posterior sampling and certainty-equivalence principles, and (iii) decision-dependent optimization when the induced distribution dynamics are contractive, enabling stable online stochastic optimization with anticipation/steering in probability spaces. The ideas are illustrated across a range of embodied and cyber-physical systems (humanoid robot dancing, passive soaring flight, power-grid optimization, and magnetic manipulation), emphasizing a single message: with the right structural assumptions, learning in dynamical systems can be as statistically efficient as supervised learning.
Biographical Information
Michael Muehlebach leads the research group learning and dynamical systems at the Max Planck Institute for Intelligent Systems in Tuebingen. His group conducts fundamental research in online learning, physics-informed machine learning, control theory, and large-scale optimization. While the research is of fundamental nature, it is often evaluated in real-world experiments, which includes a ping-pong playing robot that is actuated by pneumatic artificial muscles, heavily underactuated balancing robots, magnetic manipulation systems, and flying robots. He won numerous awards including an Emmy Noether and Branco Weiss fellowship, as well as an ETH Medal and the HILTI prize for innovative research. He serves on the editorial board of Foundations and Trends in Machine Learning.