Deep RL Timeline

  1. Atari (DQN)

    [Deepmind]

  2. 2D locomotion (TRPO)

    [Berkeley]

    A red and purple two-legged walker on a checkerboard floor
  3. AlphaGo

    [Deepmind]

    AlphaGo versus Lee Sedol, with the Go board and player shown together
  4. 3D locomotion (TRPO+GAE)

    [Berkeley]

    A humanoid falls on a checkerboard floor at iteration 160 while learning 3D locomotion
  5. Real Robot Manipulation (GPS)

    [Berkeley, Google]

    A PR2 robot manipulates a wooden shape-sorting cube at iteration 15
  6. Dota2 (PPO)

    [OpenAI]

    OpenAI's Dota2 bot plays a one-on-one match in the middle lane
  7. DeepMimic

    [Berkeley]

    A montage of DeepMimic simulated character skills compared with reference motions
  8. AlphaStar

    [Deepmind]

    AlphaStar's StarCraft II MMR results compared with human player percentiles and Grandmaster league
  9. Rubik’s Cube (PPO+DR)

    [OpenAI]

    A black robotic hand with an orange and blue wrist holds a Rubik's Cube