HomeNews & EventsEvents

AMSI-ANZIAM Lecture Tour, Newcastle

AMSI-ANZIAM Lecture Tour, Newcastle

Start Date

February 20, 2023

Registration

Closed

Image Credit: Australian Mathematical Sciences Institute (AMSI)

2023 ANZIAM Lecturer

The AMSI-ANZIAM Lecture Tour invites a distinguished international academic in an Applied Mathematical field to speak at universities across Australia after the conclusion of the ANZIAM Conference. It includes a series of talks including Specialist and Public lectures. The tour is organised biennially by AMSI and is supported by ANZIAM.

Date: 20 February 2023
Time: 2:00–3:00 pm AEDT
Venue: University of Newcastle
Title: Reinforcement Learning for Restless Bandits

Professor Konstantin Avrachenkov
National Institute for Research in Digital Science and Technology (INRIA)

Konstantin Avrachenkov received his Master degree in Control Theory from St. Petersburg State Polytechnic University (1996), Ph.D. degree in Mathematics from University of South Australia (2000) and Habilitation from University of Nice Sophia Antipolis (2010). Currently, he is a Director of Research at Inria Sophia Antipolis, France. He is an associate editor of the International Journal of Performance Evaluation, Probability in the Engineering and Informational Sciences, ACM TOMPECS, Stochastic Models and IEEE Network Magazine. Konstantin has co-authored two books “Analytic Perturbation Theory and its Applications”, SIAM, 2013 and “Statistical Analysis of Networks”, Now Publishers, 2022. He has won 5 best paper awards. His main theoretical research interests are Markov chains, Markov decision processes, random graphs and singular perturbations. He applies these methodological tools to the modeling and control of networks, and to design data mining and machine learning algorithms.

Talk Abstract: The Whittle index policy is a heuristic that has shown remarkably good performance and guaranteed asymptotic optimality when applied to the class of hard problems known as Restless Multi-Armed Bandit Problems (RMABPs). Some examples of applications of RMABPs are: machine maintenance, wireless channel scheduling, A/B testing and clinical trials, just to name a few. RMABP provides a classical example when a decision-maker needs to balance between exploration and exploitation. We present two approaches (tabular and neural network based) for learning the Whittle indices. The key feature of our approaches is the usage of two time-scales, a faster one to update the state-action Q-values, and a relatively slower one to update the Whittle indices. The neural network based approach computes the Q-values on the faster time-scale and is able to extrapolate information from one state to another, which makes the approach naturally scalable to environments with large state spaces. We present both the theoretical convergence analysis as well as illustrations by numerical examples.