[@DwarkeshPatel] What does the next training paradigm look like?
Link: https://youtu.be/20p5-kQXF_Q
Duration: 19 min
Transcript: Download plain text
Short Summary
This episode is a narration of a blog post from dwarkesh.com arguing that the entire AI industry is making one big research bet: training agents on millions of verifiable tasks across thousands of RL environments to achieve AGI. The author dissects the limits of current RLVR (real-world skills lack grindable outer-loop verification, AI is one-millionth as sample-efficient as humans, ~30–50% of lab compute goes to unused inference), then proposes alternatives like On-Policy Self-Distillation, "dreaming"/test-time training, and a 2027–2028 continual-learning scenario where week-long co-work sessions are distilled back into the base model.
![[@DwarkeshPatel] Summarizer](https://summaries.pages.dev/img/logo.webp)









