A NON-exhaustive list of RL algorithms¶
Source : OpenAI - A non-exhaustive, but useful taxonomy of algorithms in modern RL
RL Algorithms descriptions¶

Monte Carlo¶
- Every visit to Monte Carlo
Q-learning¶
- State–action–reward–state
SARSA¶
- State–action–reward–state–action
Q-learning - Lambda¶
- State–action–reward–state with eligibility traces
SARSA - Lambda¶
SARSA - State–action–reward–state–action with eligibility traces

DQN Deep¶
- Q Network
DDPG¶
- Deep Deterministic Policy Gradient
A3C Asynchronous¶
- Advantage Actor-Critic Algorithm
NAF¶
- Q-Learning with Normalized Advantage Functions
TRPO¶
- Trust Region Policy Optimization
PPO¶
- Proximal Policy Optimization
TD3¶
- Twin Delayed Deep Deterministic Policy Gradient
SAC¶
- Soft Actor-Critic
RL Algoritms comparison¶

Src : RG - James T. Graham
Resources¶
- OpenAI
-
Source : OpenAI - A non-exhaustive, but useful taxonomy of algorithms in modern RL
- Google AI
- Meta-AI
- Huggingface
-
Reinforcement Learning algorithms — an intuitive overview by Robert Moni - Great Link
- Introduction to Various Reinforcement Learning Algorithms. Part II (TRPO, PPO) by Kung-Hsiang, Huang (Steeve)
- papers RL: