Loading...
My-Formulas
Complete Mathematics

Complete Mathematics

Master Transformer mathematics with complete formulas covering linear algebra, attention, embeddings, BERT, GPT, ViT, T5, LLMs, optimization, and modern Transformer architectures.

Level

advanced

Estimated Hours

15 hrs

Total Formulas

0

Total Lessons

0

Description

Complete Mathematics of Transformers

A comprehensive collection of mathematical concepts, formulas, and architectures used in modern Transformer models such as BERT, GPT, T5, LLaMA, Gemma, Qwen, DeepSeek, Mistral, Phi, and Vision Transformers.


Module 1 — Linear Algebra Foundations

Topics

  • Vectors
  • Matrices
  • Matrix Multiplication
  • Matrix Transpose
  • Identity Matrix
  • Inverse Matrix
  • Determinant
  • Rank
  • Trace
  • Dot Product
  • Matrix Norms
  • Vector Norms
  • Cosine Similarity
  • Eigenvalues
  • Eigenvectors
  • Singular Value Decomposition (SVD)

Example Formula

A×B=CA\times B=C

Module 2 — Probability & Statistics

Topics

  • Probability
  • Joint Probability
  • Conditional Probability
  • Bayes' Theorem
  • Random Variables
  • Probability Distribution
  • Expectation
  • Variance
  • Standard Deviation
  • Covariance
  • Correlation

Example Formula

P(AB)=P(BA)P(A)P(B)P(A|B)=\frac{P(B|A)P(A)}{P(B)}

Module 3 — Calculus for Deep Learning

Topics

  • Limits
  • Derivatives
  • Partial Derivatives
  • Chain Rule
  • Gradient
  • Jacobian
  • Hessian
  • Taylor Expansion

Example Formula

dydx\frac{dy}{dx}

Module 4 — Optimization Algorithms

Topics

  • Gradient Descent
  • Stochastic Gradient Descent
  • Mini-Batch Gradient Descent
  • Momentum
  • RMSProp
  • Adam
  • AdamW
  • Learning Rate Scheduling

Example Formula

θ=θηJ(θ)\theta=\theta-\eta\nabla J(\theta)

Module 5 — Neural Network Fundamentals

Topics

  • Linear Layer
  • Bias
  • Activation Functions
  • Sigmoid
  • Tanh
  • ReLU
  • GELU
  • Softmax
  • LayerNorm
  • Dropout

Example Formula

y=xW+by=xW+b

Module 6 — Embedding Mathematics

Topics

  • Vocabulary
  • Token Embedding
  • Position Embedding
  • Segment Embedding
  • Input Embedding
  • Learned Embedding

Example Formula

X=Etoken+EpositionX=E_{token}+E_{position}

Module 7 — Positional Encoding

Topics

  • Sinusoidal Encoding
  • Learned Positional Encoding
  • Relative Position Encoding
  • Rotary Position Embedding (RoPE)

Example Formula

PE(pos,2i)=sin(pos100002i/d)PE(pos,2i)= \sin\left(\frac{pos}{10000^{2i/d}}\right)

Module 8 — Self Attention

Topics

  • Query
  • Key
  • Value
  • Attention Score
  • Scaling
  • Attention Matrix
  • Attention Output

Example Formula

Attention(Q,K,V)=Softmax(QKTdk)VAttention(Q,K,V)= Softmax\left(\frac{QK^T}{\sqrt{d_k}}\right)V

Module 9 — Multi-Head Attention

Topics

  • Multiple Attention Heads
  • Head Projection
  • Concatenation
  • Output Projection

Example Formula

MultiHead=Concat(head1,,headh)WOMultiHead= Concat(head_1,\dots,head_h)W_O

Module 10 — Feed Forward Network

Topics

  • Dense Layer
  • Hidden Layer
  • GELU
  • Output Projection

Example Formula

FFN(x)=W2(GELU(W1x+b1))+b2FFN(x)=W_2(GELU(W_1x+b_1))+b_2

Module 11 — Residual Connections

Topics

  • Skip Connection
  • Add
  • LayerNorm

Example Formula

y=LayerNorm(x+F(x))y=LayerNorm(x+F(x))

Module 12 — Transformer Encoder

Topics

  • Encoder Block
  • Self Attention
  • Feed Forward
  • Residual Learning
  • Layer Normalization

Example Formula

Encoder(x)=LayerNorm(x+FFN(MHA(x)))Encoder(x)=LayerNorm(x+FFN(MHA(x)))

Module 13 — Transformer Decoder

Topics

  • Masked Attention
  • Cross Attention
  • Feed Forward
  • Output Layer

Example Formula

Decoder(Q,K,V)Decoder(Q,K,V)

Module 14 — Encoder–Decoder Architecture

Topics

  • Cross Attention
  • Sequence-to-Sequence
  • Translation

Example Formula

Attention(Qd,Ke,Ve)Attention(Q_d,K_e,V_e)

Module 15 — Output Projection

Topics

  • Vocabulary Projection
  • Logits
  • Softmax
  • Token Prediction

Example Formula

P=Softmax(HWT)P=Softmax(HW^T)

Module 16 — Loss Functions

Topics

  • Cross Entropy
  • Binary Cross Entropy
  • Label Smoothing
  • KL Divergence

Example Formula

L=yilog(pi)L=-\sum y_i\log(p_i)

Module 17 — Transformer Training

Topics

  • Backpropagation
  • Gradient Clipping
  • Warmup
  • Weight Decay
  • AdamW

Example Formula

W=WηWW=W-\eta\nabla W

Module 18 — Decoding Strategies

Topics

  • Greedy Search
  • Beam Search
  • Top-k Sampling
  • Top-p Sampling
  • Temperature Sampling

Example Formula

xt=argmaxP(xt)x_t=\arg\max P(x_t)

Module 19 — BERT Mathematics

Topics

  • Masked Language Modeling
  • Next Sentence Prediction
  • Encoder Stack

Example Formula

L=LMLM+LNSPL=L_{MLM}+L_{NSP}

Module 20 — GPT Mathematics

Topics

  • Causal Attention
  • Decoder Only
  • Next Token Prediction

Example Formula

P(xtx<t)P(x_t|x_{<t})

Module 21 — T5 Mathematics

Topics

  • Text-to-Text
  • Encoder-Decoder
  • Span Corruption

Example Formula

P(YX)P(Y|X)

Module 22 — Vision Transformer (ViT)

Topics

  • Patch Embedding
  • Image Tokens
  • CLS Token
  • Classification Head

Example Formula

X=Patch×WX=Patch\times W

Module 23 — Advanced Attention

Topics

  • RoPE
  • ALiBi
  • Flash Attention
  • Grouped Query Attention
  • Multi Query Attention

Example Formula

RoPE(Q,K)RoPE(Q,K)

Module 24 — Mixture of Experts (MoE)

Topics

  • Router
  • Experts
  • Sparse Activation
  • Gating

Example Formula

y=iGateiExperti(x)y=\sum_i Gate_i\cdot Expert_i(x)

Module 25 — Modern LLM Architecture

Topics

  • RMSNorm
  • SwiGLU
  • KV Cache
  • Grouped Query Attention
  • Rotary Embeddings

Example Formula

RMSNorm(x)=xmean(x2)+ϵγRMSNorm(x)= \frac{x}{\sqrt{mean(x^2)+\epsilon}}\gamma

Module 26 — Transformer Complexity

Topics

  • Time Complexity
  • Memory Complexity
  • FLOPs
  • Scaling Laws
  • Efficient Attention

Example Formula

O(n2d)O(n^2d)

Module 27 — Complete Transformer Pipeline

Topics

  • Tokenization
  • Embedding
  • Positional Encoding
  • Self Attention
  • Multi-Head Attention
  • Feed Forward
  • Encoder
  • Decoder
  • Linear Projection
  • Softmax
  • Token Generation

Example Pipeline

Input
   ↓
Embedding
   ↓
Position Encoding
   ↓
Self Attention
   ↓
Multi-Head Attention
   ↓
Feed Forward
   ↓
Encoder
   ↓
Decoder
   ↓
Linear
   ↓
Softmax
   ↓
Output Token

Instructor

Name: Ankit kushwaha

Email: ankitkushwaha909@gmail.com

SEO Information

SEO Title: Complete Mathematics of Transformers | Formulas, Attention & LLMs

SEO Description: Master Transformer mathematics with complete formulas covering linear algebra, attention, embeddings, BERT, GPT, ViT, T5, LLMs, optimization

Created: 7/22/2026, 6:57:52 AM

Updated: 7/22/2026, 10:54:42 AM

Formulas

Linear Algebra

Linear algebra is the mathematical foundation of Transformer models. Every embedding, attention score, projection layer, and feed-forward network is built using vectors and matrices.

10 minOrder #1Free

Probability and Statistics

Learn probability and statistics formulas used in Transformers, including Bayes' theorem, probability distributions, expectation, variance, covariance, and correlation.

30 minOrder #2Free

Calculus Deep Learning

Learn calculus formulas used in deep learning and Transformers, including limits, derivatives, partial derivatives, chain rule, gradients, Jacobian, Hessian, and Taylor expansion.

50 minOrder #3Free

Optimization Algorithms

Learn optimization algorithms used in deep learning and Transformers, including Gradient Descent, SGD, Momentum, RMSProp, Adam, AdamW, and learning rate scheduling formulas.

50 minOrder #4Free

Neural Network Fundamentals

Learn neural network fundamentals with formulas for linear layers, bias, activation functions, Sigmoid, Tanh, ReLU, GELU, Softmax, Layer Normalization, and Dropout used in Transformers.

59 minOrder #5Free

Embedding Mathematics

Learn Transformer embedding mathematics including vocabulary, token embeddings, positional embeddings, segment embeddings, input embeddings, and learned embeddings with complete formulas.

100 minOrder #6Free

Positional Encoding

Learn positional encoding in Transformer models with formulas for sinusoidal encoding, learned positional embeddings, relative position encoding, and Rotary Position Embedding (RoPE).

101 minOrder #7Free

Self-Attention

Learn the mathematics of Self-Attention in Transformers with complete formulas for Query, Key, Value, attention scores, scaled dot-product attention, attention weights, and output computation.

10 minOrder #8Free

Multi-Head Attention

Learn the complete mathematics of Multi-Head Attention in Transformers, including multiple attention heads, head projections, concatenation, output projection, and scaled dot-product attention formulas.

100 minOrder #9Free

Feed Forward Network

Learn the mathematics of the Feed Forward Network (FFN) in Transformer models, including dense layers, hidden layers, GELU activation, output projection, and complete FFN equations.

100 minOrder #10Free

Residual Connections

Learn the mathematics of Residual Connections in Transformer models, including skip connections, Add operation, Layer Normalization, residual learning, and complete formulas.

90 minOrder #11Free

Transformer Encoder

Learn the complete mathematics of the Transformer Encoder, including encoder blocks, self-attention, feed forward networks, residual learning, layer normalization, and encoder stack formulas.

90 minOrder #12Free

Transformer Decoder

Learn the complete mathematics of the Transformer Decoder, including masked self-attention, cross-attention, feed forward networks, residual connections, layer normalization, and output layer formulas.

91 minOrder #13Free

Encoder–Decoder Architect

Learn the complete mathematics of the Transformer Encoder–Decoder architecture, including cross attention, sequence-to-sequence learning, machine translation, encoder-decoder interaction, and output generation formulas.

10 minOrder #14Free

Output Projection

Learn the complete mathematics of Output Projection in Transformer models, including vocabulary projection, logits computation, Softmax probability distribution, and token prediction formulas.

90 minOrder #15Free

Loss Functions

Learn the complete mathematics of Transformer loss functions, including Cross Entropy, Binary Cross Entropy, Label Smoothing, KL Divergence, and training objective formulas.

10 minOrder #16Free

Transformer Training

Learn the complete mathematics of Transformer training, including backpropagation, gradient clipping, learning rate warmup, weight decay, AdamW optimization, and parameter update formulas.

90 minOrder #17Free

Transformer Decoding

Learn the complete mathematics of Transformer decoding strategies, including Greedy Search, Beam Search, Top-k Sampling, Top-p Sampling, Temperature Sampling

70 minOrder #18Free

BERT Mathematics

bert mathematics, bert formulas, masked language modeling, mlm, next sentence prediction, nsp, bert encoder, transformer encoder, bert architecture, transformer mathematics, deep learning mathematics, language models

90 minOrder #19Free

GPT Mathematics

Learn the complete mathematics of GPT, including causal attention, decoder-only Transformer architecture, next token prediction, autoregressive language modeling, and training objective formulas.

10 minOrder #20Free

T5 Mathematics

Learn the complete mathematics of T5 (Text-to-Text Transfer Transformer), including encoder-decoder architecture, span corruption, text-to-text learning, attention mechanisms, and sequence generation formulas.

50 minOrder #21Free

Vision Transformer (ViT)

Learn the complete mathematics of Vision Transformers (ViT), including patch embedding, image tokens, CLS token, positional embeddings, Transformer encoder, and image classification formulas.

50 minOrder #22Free

Advanced Attention

Learn the complete mathematics of advanced Transformer attention mechanisms, including Rotary Position Embedding (RoPE), ALiBi, FlashAttention, Grouped Query Attention (GQA), and Multi Query Attention (MQA).

90 minOrder #23Free

Mixture of Experts (MoE)

Learn the complete mathematics of Mixture of Experts (MoE), including routers, experts, sparse activation, gating networks, Top-k routing, load balancing, and expert output aggregation formulas.

59 minOrder #24Free

Modern LLM Architecture

Learn the complete mathematics of modern Large Language Model (LLM) architectures, including RMSNorm, SwiGLU, KV Cache, Grouped Query Attention (GQA), and Rotary Position Embedding (RoPE).

59 minOrder #25Free

Transformer Complexity

Learn the complete mathematics of Transformer computational complexity, including time complexity, memory complexity, FLOPs, scaling laws, efficient attention algorithms, and long-context optimization formulas.

90 minOrder #26Free

Complete Transformer Pipe

Learn the complete end-to-end Transformer pipeline, from tokenization and embeddings to attention, encoder, decoder, output projection, softmax, and token generation with mathematical formulas.

70 minOrder #27Free