Complete Mathematics of Transformers
A comprehensive collection of mathematical concepts, formulas, and architectures used in modern Transformer models such as BERT, GPT, T5, LLaMA, Gemma, Qwen, DeepSeek, Mistral, Phi, and Vision Transformers.
Module 1 — Linear Algebra Foundations
Topics
- Vectors
- Matrices
- Matrix Multiplication
- Matrix Transpose
- Identity Matrix
- Inverse Matrix
- Determinant
- Rank
- Trace
- Dot Product
- Matrix Norms
- Vector Norms
- Cosine Similarity
- Eigenvalues
- Eigenvectors
- Singular Value Decomposition (SVD)
Example Formula
Module 2 — Probability & Statistics
Topics
- Probability
- Joint Probability
- Conditional Probability
- Bayes' Theorem
- Random Variables
- Probability Distribution
- Expectation
- Variance
- Standard Deviation
- Covariance
- Correlation
Example Formula
Module 3 — Calculus for Deep Learning
Topics
- Limits
- Derivatives
- Partial Derivatives
- Chain Rule
- Gradient
- Jacobian
- Hessian
- Taylor Expansion
Example Formula
Module 4 — Optimization Algorithms
Topics
- Gradient Descent
- Stochastic Gradient Descent
- Mini-Batch Gradient Descent
- Momentum
- RMSProp
- Adam
- AdamW
- Learning Rate Scheduling
Example Formula
Module 5 — Neural Network Fundamentals
Topics
- Linear Layer
- Bias
- Activation Functions
- Sigmoid
- Tanh
- ReLU
- GELU
- Softmax
- LayerNorm
- Dropout
Example Formula
Module 6 — Embedding Mathematics
Topics
- Vocabulary
- Token Embedding
- Position Embedding
- Segment Embedding
- Input Embedding
- Learned Embedding
Example Formula
Module 7 — Positional Encoding
Topics
- Sinusoidal Encoding
- Learned Positional Encoding
- Relative Position Encoding
- Rotary Position Embedding (RoPE)
Example Formula
Module 8 — Self Attention
Topics
- Query
- Key
- Value
- Attention Score
- Scaling
- Attention Matrix
- Attention Output
Example Formula
Module 9 — Multi-Head Attention
Topics
- Multiple Attention Heads
- Head Projection
- Concatenation
- Output Projection
Example Formula
Module 10 — Feed Forward Network
Topics
- Dense Layer
- Hidden Layer
- GELU
- Output Projection
Example Formula
Module 11 — Residual Connections
Topics
- Skip Connection
- Add
- LayerNorm
Example Formula
Module 12 — Transformer Encoder
Topics
- Encoder Block
- Self Attention
- Feed Forward
- Residual Learning
- Layer Normalization
Example Formula
Module 13 — Transformer Decoder
Topics
- Masked Attention
- Cross Attention
- Feed Forward
- Output Layer
Example Formula
Module 14 — Encoder–Decoder Architecture
Topics
- Cross Attention
- Sequence-to-Sequence
- Translation
Example Formula
Module 15 — Output Projection
Topics
- Vocabulary Projection
- Logits
- Softmax
- Token Prediction
Example Formula
Module 16 — Loss Functions
Topics
- Cross Entropy
- Binary Cross Entropy
- Label Smoothing
- KL Divergence
Example Formula
Module 17 — Transformer Training
Topics
- Backpropagation
- Gradient Clipping
- Warmup
- Weight Decay
- AdamW
Example Formula
Module 18 — Decoding Strategies
Topics
- Greedy Search
- Beam Search
- Top-k Sampling
- Top-p Sampling
- Temperature Sampling
Example Formula
Module 19 — BERT Mathematics
Topics
- Masked Language Modeling
- Next Sentence Prediction
- Encoder Stack
Example Formula
Module 20 — GPT Mathematics
Topics
- Causal Attention
- Decoder Only
- Next Token Prediction
Example Formula
Module 21 — T5 Mathematics
Topics
- Text-to-Text
- Encoder-Decoder
- Span Corruption
Example Formula
Module 22 — Vision Transformer (ViT)
Topics
- Patch Embedding
- Image Tokens
- CLS Token
- Classification Head
Example Formula
Module 23 — Advanced Attention
Topics
- RoPE
- ALiBi
- Flash Attention
- Grouped Query Attention
- Multi Query Attention
Example Formula
Module 24 — Mixture of Experts (MoE)
Topics
- Router
- Experts
- Sparse Activation
- Gating
Example Formula
Module 25 — Modern LLM Architecture
Topics
- RMSNorm
- SwiGLU
- KV Cache
- Grouped Query Attention
- Rotary Embeddings
Example Formula
Module 26 — Transformer Complexity
Topics
- Time Complexity
- Memory Complexity
- FLOPs
- Scaling Laws
- Efficient Attention
Example Formula
Module 27 — Complete Transformer Pipeline
Topics
- Tokenization
- Embedding
- Positional Encoding
- Self Attention
- Multi-Head Attention
- Feed Forward
- Encoder
- Decoder
- Linear Projection
- Softmax
- Token Generation
Example Pipeline
Input
↓
Embedding
↓
Position Encoding
↓
Self Attention
↓
Multi-Head Attention
↓
Feed Forward
↓
Encoder
↓
Decoder
↓
Linear
↓
Softmax
↓
Output Token
