Machine Learning & AI
Advanced
Scaled Dot-Product Attention
Lets each token weigh and aggregate information from all others — the heart of Transformers.
Formula
Variables
QQueries
KKeys
VValues
d_kKey dimension
Example
Scaling by sqrt(d_k) keeps gradients stable
Did You Know?
The 2017 paper "Attention Is All You Need" introduced this and launched the era of GPT and BERT.
Share this formula
More in Machine Learning & AI
View allLinear Regression Model
BasicPredicts a continuous value as a weighted sum of input features plus a bias.
Gradient Descent Update
BasicIteratively moves parameters in the direction that most reduces the loss.
ReLU Activation
BasicRectified Linear Unit: outputs the input if positive, else zero.
Leaky ReLU
BasicA ReLU variant that lets a small gradient flow for negative inputs.
Tanh Activation
BasicSquashes input to the range (-1, 1); zero-centred activation.
Softmax
BasicTurns a vector of scores into a probability distribution that sums to 1.