2026年第66期(总第1210期)
演讲主题:Tails of Two Sides——Conditions Inspired by Thompson Sampling for Multi-armed Bandit that Lead to Further Understandings and Improvements.
主讲人:杨健 罗格斯大学商学院管理科学与信息系统系教授
主持人:李建斌 教授、副院长
活动时间:2026年07月20日(周一)10:30-12:00
活动地址:管院大楼110教室
主讲人简介:
Jian Yang obtained his Ph.D. in Management Science from the University of Texas at Austin. After working for the Department of Mechanical and Industrial Engineering at New Jersey Institute of Technology, he is now a professor at the Department of Management Science and Information Systems, Rutgers Business School, Rutgers University. His research interests are in combinatorial optimization, logistics, production and inventory control, dynamic pricing, and game theory. At the present he is particularly interested in the role played by risk and ambiguity in dynamic inventory-price control and game-theoretical settings.
活动简介:
We may understand a policy for the multi-armed bandit (MAB) problem as using a contest between arm-specific indices to select the arm to be pulled in each period. Any policy is identifiable by its coefficients that help to form the indices, with Upper confidence bound (UCB), as well as Thompson sampling with Gaussian priors (TSg) and Beta priors (TSb) being no exceptions. Inspired by the latter two, we formulate conditions on the coefficients that lead to tight regret bounds. These conditions basically propound an llrr principle: left tails should be light and right tails should be of the right size. They also guide us to pay particular attention to asymmetric Gaussian-indexed policies AGI(sigma_l,sigma_r) whose coefficients take the form -sigma_l Z^-+sigma_r Z^+, with AGI(1,1) being just TSg. The latter's difference from the empirically more prominent TSb prompts us to generalize AGI further to AAGI policies that allow the magnitudes of index oscillations to depend on the various arms' potential rewards. Their good performances can be explained by another condition concerning a leading arm. Numerical experiments confirm that these policies, especially those with asymmetric tails, can perform well enough to beat TSb by large margins.