K-level Reasoning for Zero-Shot Coordination in Hanabi

Cui, Brandon; Hu, Hengyuan; Pineda, Luis; Foerster, Jakob N.

Computer Science > Artificial Intelligence

arXiv:2207.07166 (cs)

[Submitted on 14 Jul 2022]

Title:K-level Reasoning for Zero-Shot Coordination in Hanabi

Authors:Brandon Cui, Hengyuan Hu, Luis Pineda, Jakob N. Foerster

View PDF

Abstract:The standard problem setting in cooperative multi-agent settings is self-play (SP), where the goal is to train a team of agents that works well together. However, optimal SP policies commonly contain arbitrary conventions ("handshakes") and are not compatible with other, independently trained agents or humans. This latter desiderata was recently formalized by Hu et al. 2020 as the zero-shot coordination (ZSC) setting and partially addressed with their Other-Play (OP) algorithm, which showed improved ZSC and human-AI performance in the card game Hanabi. OP assumes access to the symmetries of the environment and prevents agents from breaking these in a mutually incompatible way during training. However, as the authors point out, discovering symmetries for a given environment is a computationally hard problem. Instead, we show that through a simple adaption of k-level reasoning (KLR) Costa Gomes et al. 2006, synchronously training all levels, we can obtain competitive ZSC and ad-hoc teamplay performance in Hanabi, including when paired with a human-like proxy bot. We also introduce a new method, synchronous-k-level reasoning with a best response (SyKLRBR), which further improves performance on our synchronous KLR by co-training a best response.

Comments:	Neurips 2021. 15 pages. 2 figures
Subjects:	Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
Cite as:	arXiv:2207.07166 [cs.AI]
	(or arXiv:2207.07166v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2207.07166
Journal reference:	Advances in Neural Information Processing Systems 2021. Vol 34. 8215--8228

Submission history

From: Brandon Cui Bicheng [view email]
[v1] Thu, 14 Jul 2022 18:53:34 UTC (12,303 KB)

Computer Science > Artificial Intelligence

Title:K-level Reasoning for Zero-Shot Coordination in Hanabi

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:K-level Reasoning for Zero-Shot Coordination in Hanabi

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators