Machine Learning & AI
Advanced
Q-Learning Update
Model-free reinforcement learning rule that learns action values from experience.
Formula
Variables
\alphaLearning rate
rReward
\gammaDiscount
s^{\prime}Next state
Example
The bracket is the temporal-difference error
Did You Know?
Deep Q-Networks used this rule to learn Atari games directly from raw pixels in 2013.
Share this formula
More in Machine Learning & AI
View allLinear Regression Model
BasicPredicts a continuous value as a weighted sum of input features plus a bias.
Gradient Descent Update
BasicIteratively moves parameters in the direction that most reduces the loss.
ReLU Activation
BasicRectified Linear Unit: outputs the input if positive, else zero.
Leaky ReLU
BasicA ReLU variant that lets a small gradient flow for negative inputs.
Tanh Activation
BasicSquashes input to the range (-1, 1); zero-centred activation.
Softmax
BasicTurns a vector of scores into a probability distribution that sums to 1.