Learning Reward Machines: A Study in Partially Observable Reinforcement Learning

Toro Icarte, Rodrigo Andrés; Klassen, Toryn Q.; Valenzano, Richard; Castro Anich, Margarita; Waldie, Ethan; McIlraith, Sheila A.

Learning Reward Machines: A Study in Partially Observable Reinforcement Learning

Files

Learning Reward Machines.pdf(79.34 KB)

Date

2023

Authors

Toro Icarte, Rodrigo Andrés

Klassen, Toryn Q.

Valenzano, Richard

Castro Anich, Margarita

Waldie, Ethan

McIlraith, Sheila A.

Abstract

Reinforcement Learning (RL) is a machine learning paradigm wherein an artificial agentinteracts with an environment with the purpose of learning behaviour that maximizesthe expected cumulative reward it receives from the environment. Reward machines(RMs) provide a structured, automata-based representation of a reward function thatenables an RL agent to decompose an RL problem into structured subproblems that canbe efficiently learned via off-policy learning. Here we show that RMs can be learnedfrom experience, instead of being specified by the user, and that the resulting problemdecomposition can be used to effectively solve partially observable RL problems. We posethe task of learning RMs as a discrete optimization problem where the objective is to findan RM that decomposes the problem into a set of subproblems such that the combinationof their optimal memoryless policies is an optimal policy for the original problem. Weshow the effectiveness of this approach on three partially observable domains, where itsignificantly outperforms A3C, PPO, and ACER, and discuss its advantages, limitations,and broader potential.

Keywords

Reinforcement learning, Reward machines, Partial observability, Automata learning

URI

https://doi.org/10.1016/j.artint.2023.103989
https://repositorio.uc.cl/handle/11534/74370

Collections

Artículos de revistas

Full item page