Ads skipped

How DeepSeek Rewrote the Transformer [MLA]

994K views · Mar 5, 2025 · Science & Technology

Comments · 934

  • @WelchLabs · 1 year ago · pinned

    Thanks to KiwiCo for sponsoring today’s video! Go to <a href="https://www.kiwico.com/welchlabs">https://www.kiwico.com/welchlabs</a> and use code WELCHLABS for 50% off your first monthly club crate or for 20% off your first Panda Crate!

    92

  • @PaulScotti · 1 year ago

    As an AI researcher who has read the DeepSeek papers already, this was a fantastic video explanation. Please make more videos along these lines!

    985

  • @MirorR3fl3ction · 1 year ago

    This is by far THE best AI deepdive Ive ever seen. I actually understand now not just how the attention arcitechure works but how and why DeepSeek&apos;s changes to the architechure result in such an incredible performance AND efficency improvement.

    2K

  • @khoiduongminh5111 · 1 year ago

    It&apos;s one of those ideas that look so so obvious when you see it, yet it is not at all! Very nice work from DeepSeek.

    718

  • @lakshaymd · 1 year ago

    The way you represent these large matrices as heat maps rather than grids of numbers helps a lot with reducing the overhead of processing the visuals. You&apos;re killing it with this series.

    125

  • @largemoney522 · 1 year ago

    In ML, it&apos;s easy to overuse tools like latent space encodings when simpler methods would be more appropriate, but what the deepseek team cooked up here is, as Welch says, a truly elegant approach to the problem at hand. Thanks for another great video!

    378

  • @nick1f · 1 year ago

    Sam Altman said that DeepSeek did not create anything new, but from your explanation, they actually improved the previous system a great deal and not only that, but they made everything public

    150

  • @RevolutionofTime · 1 year ago

    Honestly, the best visual explanations for KV Caching, and really good explanations for other points as well.<br>I can see a lot of effort. Good work 👏

    441

  • @iadduk · 8 days ago

    the best explanation of MLA on youtube

    1

  • @Mranash-ir2yf · 1 year ago

    That walkthrough of attention head patterns really demystified how transformers pull meaning out of context. I started adding these breakdowns to my notes in AI Carma for when I’m comparing model outcomes.

    34

  • @Danji_Coppersmoke · 1 year ago

    <a href="https://www.youtube.com/watch?v=0VLAoVGf_74&amp;t=770">12:50</a> extra clarification. &quot;Performance&quot; in the table refers to &quot;Quality&quot; of output. Not traditional meaning of &quot;speed&quot; which is basically the same ignoring memory loading time.

    33

  • @sh4ny1 · 1 year ago (edited)

    <a href="https://www.youtube.com/watch?v=0VLAoVGf_74&amp;t=412">6:52</a> a small additional insight:<br>i just would like to point out that in multiple head of attention mechanism we &nbsp;usually *don&apos;t have a separate Weight matrix * for each head to get K, Q and V but we divide the &nbsp;word embedding into chunks get a smaller separate K, Q,V matrix for each chunk for computing the dot product. &nbsp;that&apos;s why the word embedding lenght &apos;d&apos; should be divisible by the total number of heads. <br>in chatgpt that is 768%12 == 0 and for deepseek R1 7168%128 == 0. this way we don&apos;t have a separate weigh matrix for each head. so that even if &nbsp;according to <a href="https://www.youtube.com/watch?v=0VLAoVGf_74&amp;t=780">13:00</a> we have a same weight matrix for KQ but each head is processing only a subsection of &nbsp;KQ [num tokens &nbsp;d/num_heads ] sized matrix because we divided the original word embedding among multiple heads otherwise each head would be learning the somewhat similar things because the KQ is the same for all of them.<br><br>the idea of dividing word embedding into chunks and processing them in different heads stems from the fact that each word can have different concepts based on the context e.g the words surrounding it. for example the word &apos;Bank&apos; can be either the financial instituion or the rive bank &apos;Bank&apos; so if the word around the bank are related to fincanical instituion the a specific chunk of the word embedding for the word &apos;bank&apos; would provide higher dot product so that values in final Output will have high value section after we concatenate the outputs from all the heads to create back a matrix of size [num tokens , d] where d is 768 in gpt2 and 7168 in DSR1<br><br>Thank you for such an amazing video btw !

    109

Up next

LIVE

How Transformers Work: Attention Explained Briefly & Simply

Artecha · 64 views

LIVE

Update from Ukraine | Huge! Great news! Ballistic Missile FP-7 Used for the First Time

Denys Davydov · 212K views

LIVE

AI can't cross this line and we don't know why.

Welch Labs · 2.5M views

LIVE

The Misconception that Almost Stopped AI [How Models Learn Part 1]

Welch Labs · 674K views

LIVE

How Transformers Became the LeBron James of AI

Papers in Motion · 9 views

LIVE

How I released a game that has no assets

Zanzlanz · 1.1M views

LIVE

DeepSeek V4's Secret: 98% Less Memory

Jia-Bin Huang · 110K views

LIVE

Do AIs Learn the Same Thing?

Strad Slater · 1.7K views

LIVE

The most cited paper of the century is a brilliant hack

Welch Labs · 1.6M views

LIVE

The sorting algorithm that shouldn’t.

Stand-up Maths · 2.4M views

LIVE

DeepSeek-V3 Explained by Google Engineer | Mixture of Experts | Multi-head Latent Attention | CUDA

Martin Is A Dad · 3K views

LIVE

LLMs Don't Need More Parameters. They Need Loops.

NeuroDump · 313K views

LIVE

It's Happening - Europe is Building an Impossible Fusion Reactor

Dr Ben Miles · 3.1M views

LIVE

How do Graphics Cards Work? Exploring GPU Architecture

Branch Education · 7.6M views

LIVE

Inside the World's Smartest Robot Brain [VLA]

Welch Labs · 201K views

LIVE

DeepSeek's Insane Architecture Breakthrough [Engram Explained]

bycloud · 147K views

LIVE

But how do AI images and videos actually work? | Guest video by Welch Labs

3Blue1Brown and Welch Labs · 2.1M views

LIVE

Why AI Can Never Escape Turing's 1936 Proof

Universal Resilience with JT Yu · 1.7M views

LIVE

Yann LeCun's $1B Bet Against LLMs [Part 1]

Welch Labs · 1.2M views

LIVE

I Visualised Attention in Transformers

Gal Lahat · 377K views

LIVE

LLMs Explained: Tokens, Embeddings, Transformers and More

Syntax · 875K views

LIVE

But what is quantum computing? (Grover's Algorithm)

3Blue1Brown · 4M views

LIVE

Why Does Diffusion Work Better than Auto-Regression?

Algorithmic Simplicity · 711K views

LIVE

The Man Behind DeepSeek (Liang Wenfeng)

East Money · 705K views

LIVE

How does AI actually work? Transformers explained

AI Search · 233K views

LIVE

DeepSeek is a Game Changer for AI - Computerphile

Computerphile · 1.5M views

YouTube, with the door locked.

Aegis plays a clean stream instead of YouTube's player, so pre-roll ads, trackers, and fingerprinting never ride along. Drop Shields any time if you want the official player back.