Ads skipped

The most complex model we actually understand

887K views · Dec 20, 2025 · Science & Technology

Comments · 1.5K

  • @neelnanda2469 · 9 months ago · pinned

    Thanks for turning my work into such a gorgeous explanation! Much better than I've ever been able to communicate it

    2.7K

  • @Klapperklaus84 · 9 months ago

    95% of ML Researches give up on training one epoch before grokking occurs.

    4.1K

  • @mvsh · 9 months ago

    "It's all Fourier?!" "Always has been"

    1.3K

  • @jackkendall6420 · 9 months ago

    Seeing a sine wave emerge when you don't expect one always gives a sense of vertigo.

    631

  • @MuhsinunChowdhury · 9 months ago

    I gotta hand it to you, the animations in this video are masterfully done

    609

  • @DemiousStudios · 9 months ago

    <a href="https://www.youtube.com/watch?v=D8GOeCFFby4&amp;t=214">3:34</a> into the video before I realized I am not included in the &quot;we&quot; in the title.

    78

  • @mraxilus · 9 months ago

    I did something very similar in my 2018 research, &quot;Visualising the State Space Representations of LSTM Networks&quot;. The main difference being using dimensionality reduction instead of pairwise correlations, but I observed similar effects

    167

  • @harikishan5690 · 9 months ago

    I would also recommend taking a look at the question &quot; what happens when i continue training the model once it has grokked <br>this is where the paper &quot;Grokking and Generalization Collapse: Insights from HTSR theory&quot; finds that the model unlearns all that it has learnt during the grokking stage (Anti-grokking phase) , further strengthening the hypothesis that the model seems to random walk itself into this grokked region as hypothesized by Liu et al.

    13

  • @Patashu · 9 months ago

    <a href="https://www.youtube.com/watch?v=D8GOeCFFby4&amp;t=1213">20:13</a> I want to eat the Fourier transform waffle

    16

  • @angelenriquechavezponce1629 · 9 months ago

    What an outstanding piece of media, absolutely amazed by the clear cut visuals into the often blurry AI world. Kudos

    57

  • @taco_frog7 · 9 months ago

    Please release the book internationally

    281

  • @KoHaN7 · 9 months ago (edited)

    Hi, I just wanted to say I love your videos!<br>I also wanted to bring up some attention to a recent paper I just read about grokking, that might be of interest to you and the community. The paper in question is called &quot;Grokking at the edge of numerical stability&quot;. The peper tries to give an explanation to the phenomenon and also shows how you can controll the onset of the grokking learning phase. <br><br>The main idea, as I have understood it, is that grokking is a consequence of the model getting stuck in some sort of eigenstate of the gradiend direction, where the model has memorized the training data, but keeps minimizing the training loss function by simply saying the answer more confidently, without it being force to learn anything new. After all, cross entopy rewards not only the accucary but also the confidence of the model in the answer. In an ideal world, the model would keep doing this forever, never beeing forced to generalize. Lucky for us at some pioint the softmax is so extreme that it reaches numerical instability and the model is kicked out of this eigenstate, generating the onset of grokking. The authors find a clever way to visualize and also force the model to never go into this eigenstate to begin with, and therefore training and test accucary start raising at the same time. <br><br>I am looking forward to the time when the books will be available in the UK as I would like to buy several of them! 🚀<br>Keep up the good work, I am positive that your work is inspiring the next generation of ML Scientist and researchers in the field of AI!

    163

Up next

LIVE

Can humans make AI any better?

Welch Labs · 300K views

LIVE

How DeepSeek Rewrote the Transformer [MLA]

Welch Labs · 992K views

LIVE

Superexponential Strict-improvement Distances in Median Graphs

Evidence Press · 26 views

LIVE

What the Books Get Wrong about AI [Double Descent]

Welch Labs · 317K views

LIVE

Bill Gates: A.I. ‘Makes Nuclear Weapons Look Like Nothing’ | The Ezra Klein Show

The Ezra Klein Show and 2 more · 615K views

LIVE

Why a Magnet Doesn't Actually Run Out of Energy

Sleep On Physics · 123K views

LIVE

Reinventing Entropy | Compression is Intelligence Part 1

3Blue1Brown · 1.5M views

LIVE

Biggest Breakthroughs in Mathematics: 2025

Quanta Magazine · 806K views

LIVE

The moment we stopped understanding AI [AlexNet]

Welch Labs · 2.5M views

LIVE

The Dark Matter of AI [Mechanistic Interpretability]

Welch Labs · 304K views

LIVE

The AI Language We Can't Read: Neuralese ft. Rob Miles - Computerphile

Computerphile · 586K views

LIVE

What Did Ilya See?

Run The Numbers · 985K views

LIVE

Yann LeCun's $1B Bet Against LLMs [Part 1]

Welch Labs · 1.2M views

LIVE

The Quantum Mechanics Behind the Periodic Table

Perhaps Barksley and Bub Explains · 80K views

LIVE

Recursive Self-Improvement

Emergent Garden · 335K views

LIVE

The Closest Thing We Have to Alien Technology

Veritasium · 59M views

LIVE

The Misconception that Almost Stopped AI [How Models Learn Part 1]

Welch Labs · 674K views

LIVE

But how do AI images and videos actually work? | Guest video by Welch Labs

3Blue1Brown and Welch Labs · 2.1M views

LIVE

How AI Learned to Think

Art of the Problem · 212K views

LIVE

The most cited paper of the century is a brilliant hack

Welch Labs · 1.6M views

YouTube, with the door locked.

Aegis plays a clean stream instead of YouTube's player, so pre-roll ads, trackers, and fingerprinting never ride along. Drop Shields any time if you want the official player back.