Back to Resources

An advanced course examining reinforcement learning from Markov decision processes and multi-armed bandits to deep neural approaches. Students practice tabular methods such as Monte Carlo, temporal-difference learning, and Q-learning, then build deep Q-network, policy-gradient, and actor-critic agents, with attention to safety and ethical issues.
- Level
- Graduate
- Department
- MET
- Credits
- 4
- Prerequisites
- MET CS 767 or consent
- BU Bulletin
- View in the BU Bulletin
Instructors
Last verified: August 22, 2026
Suggest a correction