机器学习与数据科学博士生系列论坛(第一百零六期)—— Beyond Regret: Lyapunov Analysis of FTRL in Stochastic Bandits
报告人:詹景昕(色控
)
时间:2026-09-24 16:00-17:00
地点:腾讯会议:441-3293-3791
摘要:
Follow-the-Regularized-Leader (FTRL) is a central framework for regret minimization in multi-armed bandits. Its performance beyond cumulative regret, however, is less understood: how quickly does its sampling distribution concentrate on optimal arms, and how reliably can its loss estimates identify the best arm? These questions are challenging because importance-weighted feedback introduces large, state-dependent fluctuations.
This talk presents a Lyapunov approach to answering these questions in stochastic bandits, using the best-of-both-worlds (BOBW) algorithm 1/2-Tsallis-INF as the main example. Drawing on two recent works, we discuss the construction of suitable Lyapunov functions and the drift estimates that control the algorithm's behavior. The analysis yields polynomial bounds on best-arm identification error and convergence guarantees for the last iterate, without modifying the algorithm. We emphasize the intuition behind the method and its key technical ideas.
论坛简介:该线上论坛是由张志华教授机器学习实验室组织,每两周主办一次(除了公共假期)。论坛每次邀请一位博士生就某个前沿课题做较为系统深入的介绍,主题包括但不限于机器学习、高维统计学、运筹优化和理论计算机科学。