arXiv:1612.08810 Abstract | arXiv Analytics

arXiv:1612.08810 [cs.LG]Abstract References Reviews Resources

The Predictron: End-To-End Learning and Planning

David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, Thomas Degris

Published 2016-12-28Version 1

One of the key challenges of artificial intelligence is to learn models that are effective in the context of planning. In this document we introduce the predictron architecture. The predictron consists of a fully abstract model, represented by a Markov reward process, that can be rolled forward multiple "imagined" planning steps. Each forward pass of the predictron accumulates internal rewards and values over multiple planning depths. The predictron is trained end-to-end so as to make these accumulated values accurately approximate the true value function. We applied the predictron to procedurally generated random mazes and a simulator for the game of pool. The predictron yielded significantly more accurate predictions than conventional deep neural network architectures.

Categories: cs.LG, cs.AI, cs.NE

Keywords: end-to-end learning, conventional deep neural network architectures, true value function, predictron accumulates internal rewards, multiple planning depths

Related articles: Most relevant | Search more

arXiv:2309.13077 [cs.LG] (Published 2023-09-21)

A Differentiable Framework for End-to-End Learning of Hybrid Structured Compression

Moonjung Eo, Suhyun Kang, Wonjong Rhee

arXiv:2002.05707 [cs.LG] (Published 2020-02-13)

A Framework for End-to-End Learning on Semantic Tree-Structured Data

William Woof, Ke Chen

arXiv:1805.06523 [cs.LG] (Published 2018-05-16)

End-to-end Learning of a Convolutional Neural Network via Deep Tensor Decomposition

Samet Oymak, Mahdi Soltanolkotabi