VALUE LANDSCAPE

A field instrument for the Bellman equation — paint a world, run synchronous value iteration, and watch the terrain and its policy settle.

V(s) = maxa Σs′ P(s′|s,a) · [ R(s,a,s′) + γ·V(s′) ]
Sweep0 Max Δ
−0.05 0 +0.05