
Build & Train a GLM-5.3-Flash Model From Scratch with Python
Learn how to design, build, and train an advanced multimodal language model from scratch using modern techniques like mixture-of-experts, sparse attention, and reinforcement learning. This hands-on guide walks you through the entire lifecycle from tokenization and pre-training to post-training optimization and evaluation.
Video from @vukrosic
GitHub: https://github.com/vukrosic/glm-5.3-flash-from-scratch
Become AI researcher in 90 days: https://www.skool.com/become-ai-researcher-2669/about
Slides: https://github.com/vukrosic/glm-5.3-flash-from-scratch/blob/main/slides/slides.html
Blog: https://z.ai/blog/glm-5.3-flash
❤️ Support for this channel comes from our friends at Scrimba – the coding platform that's reinvented interactive learning: https://scrimba.com/freecodecamp
⭐️ Chapters ⭐️
- 00:00 Introduction & What We're Building
- 01:02 The Modern AI Researcher Role & Asking Research Questions
- 04:50 Tokenization & Byte-Level Vocabulary (Why Small Vocab Matters)
- 07:01 Embeddings & Transformer Forward Pass Overview
- 08:53 GLM-5.3 Architecture Overview & Model Specifications
- 10:25 Code Walkthrough: Embeddings & Token Representation
- 11:56 Manifold Constrained Hyperconnections (DeepSeek Residuals)
- 13:27 Output Projection & Weight Tying
- 14:27 RMSNorm & Normalization Layers
- 15:10 Positional Encodings (RoPE vs. NoPE) & Sparse Attention Indexer
- 18:30 Linear Attention (State-Space Memory) vs. Sparse Attention
- 20:46 Mixture of Experts (MoE) & Shared Experts
- 23:06 Adding Vision: Patch Embeddings & 2D RoPE
- 26:14 Pre-Training Pipeline, Loss & Optimization (AdamW)
- 28:05 Pre-Training Experiments: Data Diversity, Interleaving & Curriculums
- 30:27 Post-Training & Reinforcement Learning (RL) Setup
- 34:44 Designing Reward Functions & Group Relative Policy Optimization (GRPO)
- 37:37 Parameter-Efficient RL Updates & Freezing Layers
- 39:35 Evaluating RL Results: Task Gains & Regression Risks
- 40:47 RL Hyperparameter Experiments: Group Size, Temperature & Seeds
- 43:03 Summary & Advice for Aspiring AI Researchers
? Thanks to our Champion and Sponsor supporters:
? @omerhattapoglu1158
? @goddardtan
? @akihayashi6629
? @kikilogsin
? @anthonycampbell2148
? @tobymiller7790
? @rajibdassharma497
? @CloudVirtualizationEnthusiast
? @adilsoncarlosvianacarlos
? @martinmacchia1564
? @ulisesmoralez4160
? @_Oscar_
? @jedi-or-sith2728
? @justinhual1290
--
Learn to code for free and get a developer job: https://www.freecodecamp.org
Read hundreds of articles on programming: https://freecodecamp.org/news
Video from @vukrosic
GitHub: https://github.com/vukrosic/glm-5.3-flash-from-scratch
Become AI researcher in 90 days: https://www.skool.com/become-ai-researcher-2669/about
Slides: https://github.com/vukrosic/glm-5.3-flash-from-scratch/blob/main/slides/slides.html
Blog: https://z.ai/blog/glm-5.3-flash
❤️ Support for this channel comes from our friends at Scrimba – the coding platform that's reinvented interactive learning: https://scrimba.com/freecodecamp
⭐️ Chapters ⭐️
- 00:00 Introduction & What We're Building
- 01:02 The Modern AI Researcher Role & Asking Research Questions
- 04:50 Tokenization & Byte-Level Vocabulary (Why Small Vocab Matters)
- 07:01 Embeddings & Transformer Forward Pass Overview
- 08:53 GLM-5.3 Architecture Overview & Model Specifications
- 10:25 Code Walkthrough: Embeddings & Token Representation
- 11:56 Manifold Constrained Hyperconnections (DeepSeek Residuals)
- 13:27 Output Projection & Weight Tying
- 14:27 RMSNorm & Normalization Layers
- 15:10 Positional Encodings (RoPE vs. NoPE) & Sparse Attention Indexer
- 18:30 Linear Attention (State-Space Memory) vs. Sparse Attention
- 20:46 Mixture of Experts (MoE) & Shared Experts
- 23:06 Adding Vision: Patch Embeddings & 2D RoPE
- 26:14 Pre-Training Pipeline, Loss & Optimization (AdamW)
- 28:05 Pre-Training Experiments: Data Diversity, Interleaving & Curriculums
- 30:27 Post-Training & Reinforcement Learning (RL) Setup
- 34:44 Designing Reward Functions & Group Relative Policy Optimization (GRPO)
- 37:37 Parameter-Efficient RL Updates & Freezing Layers
- 39:35 Evaluating RL Results: Task Gains & Regression Risks
- 40:47 RL Hyperparameter Experiments: Group Size, Temperature & Seeds
- 43:03 Summary & Advice for Aspiring AI Researchers
? Thanks to our Champion and Sponsor supporters:
? @omerhattapoglu1158
? @goddardtan
? @akihayashi6629
? @kikilogsin
? @anthonycampbell2148
? @tobymiller7790
? @rajibdassharma497
? @CloudVirtualizationEnthusiast
? @adilsoncarlosvianacarlos
? @martinmacchia1564
? @ulisesmoralez4160
? @_Oscar_
? @jedi-or-sith2728
? @justinhual1290
--
Learn to code for free and get a developer job: https://www.freecodecamp.org
Read hundreds of articles on programming: https://freecodecamp.org/news
freeCodeCamp.org
Learn to code for free....
Stop choosing between learning coding fundamentals and learning to use AI. Do both.
freeCodeCamp.org
Are you familiar with the Zen of Python? Estefania breaks down the principles here.
freeCodeCamp.org