Markov decision process (MDP)
Assumption: agent gets to observe the state
An MDP is defined by:
- Set of states
- Set of actions
-
Transition function
Action noise is encoded in the transition probabilities: the chosen action can lead to different next states.
- Reward function
- Start state
- Discount factor
- Horizon
goal:
Find the policy that maximizes the expected sum of discounted rewards.
Example: Gridworld
Goal:
So, as written, this is not a value function: it does not explicitly condition on .