Solved Exercises for Reinforcement Learning
Chapter-by-chapter exercises
- 1IntroductionTic-tac-toe, self-play, symmetries and the cost of pure greed.
- 2Multi-armed BanditsAction-value methods, non-stationary bandits, UCB, and gradient bandits.
- 3Finite Markov Decision ProcessesReturns, value functions, Bellman equations, and optimal policies.
- 4Dynamic ProgrammingPolicy evaluation, policy iteration, value iteration, and the gambler's problem.
- 5Monte Carlo MethodsPrediction, control, importance sampling, blackjack, and racetrack experiments.
- 6Temporal-Difference LearningTD prediction, SARSA, Q-learning, Expected SARSA, and windy gridworld.