How DeepSeek Rewrote the Transformer [MLA]
994K views · Mar 5, 2025 · Science & Technology
Comments · 934
@WelchLabs · 1 year ago · pinned
Thanks to KiwiCo for sponsoring today’s video! Go to <a href="https://www.kiwico.com/welchlabs">https://www.kiwico.com/welchlabs</a> and use code WELCHLABS for 50% off your first monthly club crate or for 20% off your first Panda Crate!
92
@PaulScotti · 1 year ago
As an AI researcher who has read the DeepSeek papers already, this was a fantastic video explanation. Please make more videos along these lines!
985
@MirorR3fl3ction · 1 year ago
This is by far THE best AI deepdive Ive ever seen. I actually understand now not just how the attention arcitechure works but how and why DeepSeek's changes to the architechure result in such an incredible performance AND efficency improvement.
2K
@khoiduongminh5111 · 1 year ago
It's one of those ideas that look so so obvious when you see it, yet it is not at all! Very nice work from DeepSeek.
718
@lakshaymd · 1 year ago
The way you represent these large matrices as heat maps rather than grids of numbers helps a lot with reducing the overhead of processing the visuals. You're killing it with this series.
125
@largemoney522 · 1 year ago
In ML, it's easy to overuse tools like latent space encodings when simpler methods would be more appropriate, but what the deepseek team cooked up here is, as Welch says, a truly elegant approach to the problem at hand. Thanks for another great video!
378
@nick1f · 1 year ago
Sam Altman said that DeepSeek did not create anything new, but from your explanation, they actually improved the previous system a great deal and not only that, but they made everything public
150
@RevolutionofTime · 1 year ago
Honestly, the best visual explanations for KV Caching, and really good explanations for other points as well.<br>I can see a lot of effort. Good work 👏
441
@iadduk · 8 days ago
the best explanation of MLA on youtube
1
@Mranash-ir2yf · 1 year ago
That walkthrough of attention head patterns really demystified how transformers pull meaning out of context. I started adding these breakdowns to my notes in AI Carma for when I’m comparing model outcomes.
34
@Danji_Coppersmoke · 1 year ago
<a href="https://www.youtube.com/watch?v=0VLAoVGf_74&t=770">12:50</a> extra clarification. "Performance" in the table refers to "Quality" of output. Not traditional meaning of "speed" which is basically the same ignoring memory loading time.
33
@sh4ny1 · 1 year ago (edited)
<a href="https://www.youtube.com/watch?v=0VLAoVGf_74&t=412">6:52</a> a small additional insight:<br>i just would like to point out that in multiple head of attention mechanism we usually *don't have a separate Weight matrix * for each head to get K, Q and V but we divide the word embedding into chunks get a smaller separate K, Q,V matrix for each chunk for computing the dot product. that's why the word embedding lenght 'd' should be divisible by the total number of heads. <br>in chatgpt that is 768%12 == 0 and for deepseek R1 7168%128 == 0. this way we don't have a separate weigh matrix for each head. so that even if according to <a href="https://www.youtube.com/watch?v=0VLAoVGf_74&t=780">13:00</a> we have a same weight matrix for KQ but each head is processing only a subsection of KQ [num tokens d/num_heads ] sized matrix because we divided the original word embedding among multiple heads otherwise each head would be learning the somewhat similar things because the KQ is the same for all of them.<br><br>the idea of dividing word embedding into chunks and processing them in different heads stems from the fact that each word can have different concepts based on the context e.g the words surrounding it. for example the word 'Bank' can be either the financial instituion or the rive bank 'Bank' so if the word around the bank are related to fincanical instituion the a specific chunk of the word embedding for the word 'bank' would provide higher dot product so that values in final Output will have high value section after we concatenate the outputs from all the heads to create back a matrix of size [num tokens , d] where d is 768 in gpt2 and 7168 in DSR1<br><br>Thank you for such an amazing video btw !
109
Up next

How Transformers Work: Attention Explained Briefly & Simply
Artecha · 64 views

Update from Ukraine | Huge! Great news! Ballistic Missile FP-7 Used for the First Time
Denys Davydov · 212K views

AI can't cross this line and we don't know why.
Welch Labs · 2.5M views

The Misconception that Almost Stopped AI [How Models Learn Part 1]
Welch Labs · 674K views

How Transformers Became the LeBron James of AI
Papers in Motion · 9 views

How I released a game that has no assets
Zanzlanz · 1.1M views

DeepSeek V4's Secret: 98% Less Memory
Jia-Bin Huang · 110K views

Do AIs Learn the Same Thing?
Strad Slater · 1.7K views

The most cited paper of the century is a brilliant hack
Welch Labs · 1.6M views

The sorting algorithm that shouldn’t.
Stand-up Maths · 2.4M views

DeepSeek-V3 Explained by Google Engineer | Mixture of Experts | Multi-head Latent Attention | CUDA
Martin Is A Dad · 3K views

LLMs Don't Need More Parameters. They Need Loops.
NeuroDump · 313K views

It's Happening - Europe is Building an Impossible Fusion Reactor
Dr Ben Miles · 3.1M views

How do Graphics Cards Work? Exploring GPU Architecture
Branch Education · 7.6M views

Inside the World's Smartest Robot Brain [VLA]
Welch Labs · 201K views

DeepSeek's Insane Architecture Breakthrough [Engram Explained]
bycloud · 147K views

But how do AI images and videos actually work? | Guest video by Welch Labs
3Blue1Brown and Welch Labs · 2.1M views

Why AI Can Never Escape Turing's 1936 Proof
Universal Resilience with JT Yu · 1.7M views

Yann LeCun's $1B Bet Against LLMs [Part 1]
Welch Labs · 1.2M views

I Visualised Attention in Transformers
Gal Lahat · 377K views

LLMs Explained: Tokens, Embeddings, Transformers and More
Syntax · 875K views

But what is quantum computing? (Grover's Algorithm)
3Blue1Brown · 4M views

Why Does Diffusion Work Better than Auto-Regression?
Algorithmic Simplicity · 711K views

The Man Behind DeepSeek (Liang Wenfeng)
East Money · 705K views

How does AI actually work? Transformers explained
AI Search · 233K views

DeepSeek is a Game Changer for AI - Computerphile
Computerphile · 1.5M views