Markov decision process (MDP)

The agent receives state s_t and reward r_t, then sends action a_t to the environment, which returns the next state s_(t+1) and reward r_(t+1)

Assumption: agent gets to observe the state

[Drawing from Sutton and Barto, Reinforcement Learning: An Introduction, 1998]

An MDP is defined by:

MDP(S,A,T,R,γ,H)

goal:

maxπ E[ ∑t=0H γt R(St,At,St+1)|π]

Find the policy that maximizes the expected sum of discounted rewards.

Example: Gridworld

A four-column, three-row Gridworld with Start at (1,1), a wall at (2,2), a robot at (3,1), a plus-one reward at (4,3), and a minus-one reward at (4,2)

Goal:

maxπ E[ ∑t=0H γt R(St,At,St+1)|π]

So, as written, this is not a value function: it does not explicitly condition on S0=s.

π:

A robot holds a Gridworld policy map with an action arrow in each nonterminal, unblocked cell