Recent Advances in Reward Modeling for Reinforcement Learning
Abstract:
Reinforcement learning (RL) has achieved remarkable success in robotics, games, and language model post-training, where reward signals are essential for effective agent training. In this talk, I will present our recent research on advanced reward modeling. This includes robustifying RL through transfer and weakly supervised learning, coping with interval-based rewards, and exploring diverse reward aggregation frameworks that move beyond the traditional discounted sum.
Short-bio:
Masashi Sugiyama received his Ph.D. in Computer Science from the Tokyo Institute of Technology, Japan, in 2001. After serving as an assistant and associate professor at the same institute, he became a professor at the University of Tokyo in 2014. Since 2016, he has also served as the director of the RIKEN Center for Advanced Intelligence Project.
His research interests include theories and algorithms of machine learning, such as weakly supervised learning, distribution shift adaptation, and reinforcement learning. He was a keynote speaker at ICLR 2023, and is currently an ICML board member and a NeurIPS advisory board member.