91亚色

Skip to main content Skip to local navigation

CAIS Seminar: Sharan Vaswani

Title: A Systematic Framework for Designing Policy Gradient Methods for Reinforcement Learning

Please register for the seminar. Registration is required.

My CAIS affiliation(Required)
If you wish to become a member, please complete the membership form at https://machform.osgoode.yorku.ca/machform/view.php?id=226829
I confirm my registration for the seminar.(Required)
Abstract: Reinforcement learning (RL) studies sequential decision-making problems in which an agent interacts with an environment, receives feedback in the form of rewards, and aims to learn a policy that maximizes long-term performance. RL has found applications in medicine, industrial control, robotics, and, more recently, for reasoning with large language models. Policy gradient (PG) methods are a widely used approach in RL, providing a direct way to optimize decision-making performance by using gradient ascent on the expected cumulative reward. Common PG algorithms improve the policy by iteratively optimizing surrogate objectives that approximate the true objective. While effective in practice, these surrogates are often designed in an ad hoc manner.

In this talk, we present a systematic framework for designing surrogate objectives for PG methods. The framework encompasses existing algorithms, explains common implementation heuristics, and provides a principled way to develop new methods. As聽an聽example, we introduce Softmax Policy Mirror Ascent (SPMA), a simple algorithm that is easy to implement, supports modern non-linear function approximation, and admits theoretical convergence guarantees in simplified settings. We demonstrate its empirical effectiveness on standard benchmarks, where it achieves performance comparable to or better than widely used state-of-the-art methods.

叠颈辞:听Sharan Vaswani is an Assistant Professor in the School of Computing Science at Simon Fraser University. Sharan obtained his PhD at the University of British Columbia and was a postdoctoral researcher at Mila - Quebec AI Institute and the University of Alberta. His research interests include designing better algorithms for sequential decision-making under uncertainty and optimization for machine learning.

Date

Jan 22 2026
Expired!

Time

11:00 am - 12:30 pm
QR Code