Episodes

  • When AI Research Starts Moving Faster Than Human Research - Zhengyao Jiang
    Sep 26 2026

    Weco let an AI coding agent rewrite the harness around another agent for eight days: its code, prompts and tools, while the underlying language model stayed fixed. Tim Scarfe asks Weco co-founder Zhengyao Jiang what the reported gains over two years of human engineering actually demonstrate.The discussion examines AIDE 85's generated code, held-out evaluation and the difficulty of separating useful discoveries from reward hacking. Jiang explains Weco's four levels of recursive self-improvement and compares the experiment with AlphaEvolve and the Darwin Gödel Machine.The limits matter as much as the gains. Jiang explains why the experiment did not establish that the system had become a better improver. The conversation closes with open-ended search, human-designed primitives and Parameter Golf: where does the next useful idea come from when the agent is searching inside a space that people designed?---TIMESTAMPS:00:00:00 Eight days of self-improvement: what counts?00:03:25 AIDE and the puzzle of useful spaghetti code00:08:38 Four levels of recursive self-improvement00:12:02 What AIDE 85 changed and how it was tested00:20:04 AlphaEvolve, Darwin Gödel Machine and the RSI claim00:26:21 Reward hacking and the limits of detection00:33:09 Open-ended search, harness tuning and creativity00:39:43 Parameter Golf and the limits of self-improvement---REFERENCES:organization:[00:00:30] Weco AIhttps://www.weco.ai/other:[00:00:33] AIDE²: The First Evidence of Recursive Self-Improvementhttps://www.weco.ai/blog/first-evidence-of-recursive-self-improvement[00:14:11] Faulty reward functions in the wildhttps://openai.com/index/faulty-reward-functions/[00:29:59] The Hugging Face incident and the road aheadhttps://openai.com/index/hugging-face-incident-and-the-road-ahead/tool:[00:03:29] AIDEhttps://github.com/WecoAI/aideml[00:04:29] MLE-benchhttps://github.com/openai/mle-bench[00:04:33] ALE-Benchhttps://github.com/SakanaAI/ALE-Bench[00:04:52] WeatherBench 2https://github.com/google-research/weatherbench2[00:08:18] ReActhttps://react-lm.github.io/[00:39:43] Parameter Golfhttps://github.com/openai/parameter-golfpaper:[00:20:08] AlphaEvolve: A coding agent for scientific and algorithmic discoveryhttps://arxiv.org/abs/2506.13131v1[00:21:35] Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agentshttps://arxiv.org/abs/2505.22954v3[00:23:45] Hyperagentshttps://arxiv.org/abs/2603.19461v1[00:27:01] SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agentshttps://arxiv.org/abs/2605.21384book:[00:33:14] Why Greatness Cannot Be Planned: The Myth of the Objectivehttps://link.springer.com/book/10.1007/978-3-319-15524-1---LINKS:https://app.rescript.info/share/3a9dc6189cb539c6a05fcc4f75c101b3PDF:https://app.rescript.info/api/public/sessions/9eda60ede2b31c92/pdf

    Show More Show Less
    44 mins
  • Why Scaling Prediction Cannot Create Intelligence - Alexander Mattick
    Sep 21 2026

    Alexander Mattick is a researcher at Fraunhofer IIS and a PhD researcher at the University of Technology Nuremberg (UTN), and a regular on Yannic Kilcher's Discord. He first came on MLST in 2022, after helping research the Yann LeCun and Randall Balestriero episode on interpolation.


    SPONSOR:

    ---

    Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open.

    Apply now: https://cyber.fund

    ---


    Alexander treats inference as the thread running through modern machine learning: once you have a model, what does it cost to get an answer out of it? He works through Monte Carlo, GFlowNets, energy-based models, diffusion, normalising flows and flow matching, with four short explainers he recorded himself. He is blunt about energy-based models: you can sample from them in principle, but it is rarely worth the compute. JEPA and "world model", he says, are closer to branding than to technical categories.


    Next: theories of deep learning, none of which he thinks predicts enough yet to guide practice, then reinforcement learning.


    ---

    0:00 Cold open: information is expensive

    0:51 Welcome back, Alexander Mattic

    2:08 Alexander's research background

    2:50 Inference: densities, sampling and Monte Carlo

    6:42 GFlowNets, energy functions and MCMC

    9:45 Explainer: energy-based models

    11:03 Why model a density at all?

    17:30 From learned energies to flow matching

    25:08 Explainers: diffusion and normalising flows

    28:33 Are energy-based models generative?

    33:22 JEPA, contrastive learning and collapse

    41:13 Why non-language modalities need flows

    44:51 Inference as search: branch and bound

    49:43 Q-learning and delayed consequences

    55:14 Flow matching, optimal transport, Fokker-Planck

    1:00:03 Explainer: flow matching

    1:01:49 AlphaFold, latents and scale versus architecture

    1:07:52 Two families of deep learning theory

    1:15:04 What a good theory would predict

    1:23:53 The manifold hypothesis and compression

    1:28:25 Is reward enough?

    1:32:01 Control theory versus reinforcement learning

    1:37:22 The Bitter Lesson and expensive information

    1:42:08 Constrained RL: the constrained MDP toolbox

    1:50:12 Creativity as constrained search

    1:55:44 Reality is protean: when abstractions hold

    2:00:32 What is a world model?

    2:04:38 Prediction is not control

    2:08:13 Robot demos, MPC and reliability


    ---

    REFERENCES:

    [6:55] GFlowNets (Bengio et al., 2021)

    https://arxiv.org/abs/2106.04399

    [38:46] Contrastive Self-Supervised Learning (Anand, 2020)

    https://ankeshanand.com/blog/2020/01/26/contrative-self-supervised-learning.html

    [38:56] LeJEPA (Balestriero and LeCun, 2025)

    https://arxiv.org/abs/2511.08544v3

    [47:10] RL for Node Selection in Branch-and-Bound (Mattick)

    https://openreview.net/forum?id=0ez68a5UqI

    [56:20] Flow Matching for Generative Modeling

    https://arxiv.org/abs/2210.02747v2

    [1:12:41] Disentangling feature and lazy training in deep neural networks

    https://arxiv.org/abs/1906.08034v4

    [1:31:05] Reward is enough (Silver)

    https://doi.org/10.1016/j.artint.2021.103535

    [1:35:12] Learning ReLU networks to high uniform accuracy is intractable (Berner et al.)

    https://arxiv.org/abs/2205.13531v2

    [1:40:20] Dota 2 with Large Scale Deep RL

    https://arxiv.org/abs/1912.06680v1

    [1:45:41] Constrained Update Projection for Safe Policy Optimization (Yang et al., 2022)

    https://arxiv.org/abs/2209.07089

    [1:46:11] SafeMPO (ICLR 2026)

    https://openreview.net/forum?id=1m0EU6QXj6

    [1:50:17] Why Creativity Cannot Be Interpolated

    https://archive.mlst.ai/paper/why-creativity-cannot-be-interpolated/

    [1:51:39] Invalid Action Masking (Huang and Ontañón)

    https://arxiv.org/abs/2006.14171

    [2:00:04] Probability Theory: The Logic of Science (Jaynes, 2003)

    https://www.cambridge.org/core/books/probability-theory/9CA08E224FF30123304E6D8935CF1A99

    [2:01:53] Training Agents Inside of Scalable World Models (Hafner et al., 2025)

    https://arxiv.org/abs/2509.24527v1

    [2:03:43] World Models (Ha and Schmidhuber, 2018)

    https://arxiv.org/abs/1803.10122v4

    Show More Show Less
    2 hrs and 14 mins
  • How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
    Sep 15 2026

    The car making a left turn at the start of this episode was never filmed. Cosmos 3 generated it. Ming-Yu Liu, who leads the Cosmos research at NVIDIA, explains how one model can describe a video, generate one, and produce robot actions.


    He walks Tim through the architecture. A vision language model reasons one token at a time; its weights then initialise a bidirectional diffusion generator for video, audio and action, and a shared temporal position scheme lines up signals that run at different rates. Ming-Yu treats "world model" as a set of tools, not one definition: forward dynamics, inverse dynamics and policy, trained together under a capacity limit so that each helps the others. He also explains why plentiful first-person human video carries over to robots, which have far less data of their own, and why a Cosmos model post-trained on the DROID dataset is a good starting point for pick-and-place policies.


    The most practical thread is testing. A neural simulator does not need accurate success rates. It only needs to rank policy A above policy B the way the real world would, so a team can narrow down which checkpoints deserve a real trial. Cosmos Dreams applies that closed-loop idea to driving and robotics, and Ming-Yu argues that humanoids around children and pets make safety matter even more than it does for cars. The conversation ends on the Super, Nano and Edge sizes (Edge targets Jetson Thor, Orin and DGX Spark) and where to find the open weights, code and data.


    This episode is a paid partnership with NVIDIA.


    Learn more about Cosmos: https://nvda.ws/4cJoY1S

    Explore Cosmos Lab: https://research.nvidia.com/labs/cosmos-lab/cosmos3/


    ---

    TIMESTAMPS:

    00:00:00 A road that was never filmed

    00:02:28 Inside Cosmos 3: reasoning and generator towers

    00:05:02 World models: dynamics, policy and one clock

    00:08:59 Learning robot skills from human video

    00:11:06 Ambiguous tasks and system 2 planning

    00:12:53 Neural simulators for policy verification

    00:16:41 Cosmos as a starting point for robot policies

    00:19:00 Cosmos Dreams and robot safety

    00:22:04 Super, Nano and Edge model sizes

    00:24:24 Open models, the Cosmos repo and feedback


    ---

    REFERENCES:

    tool:

    [00:00:13] Cosmos 3 (NVIDIA Cosmos Lab project page)

    https://research.nvidia.com/labs/cosmos-lab/cosmos3/

    [00:18:27] NVIDIA Cosmos GitHub repository

    https://github.com/NVIDIA/cosmos

    [00:22:05] Cosmos3-Edge model card

    https://huggingface.co/nvidia/Cosmos3-Edge

    [00:22:15] Cosmos3-Super model card

    https://huggingface.co/nvidia/Cosmos3-Super

    [00:22:16] Cosmos3-Nano model card

    https://huggingface.co/nvidia/Cosmos3-Nano

    [00:22:50] NVIDIA Jetson Thor

    https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/

    [00:22:52] NVIDIA Jetson Orin

    https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/

    [00:22:53] NVIDIA DGX Spark

    https://www.nvidia.com/en-us/products/workstations/dgx-spark/

    [00:24:42] Cosmos 3 collection on Hugging Face

    https://huggingface.co/collections/nvidia/cosmos3

    other:

    [00:01:07] Cosmos-Dreams closed-loop simulators (NVIDIA SIGGRAPH 2026 blog)

    https://blogs.nvidia.com/blog/siggraph-news-2026/

    paper:

    [00:08:54] Cosmos 3: Omnimodal World Models for Physical AI

    https://arxiv.org/abs/2606.02800

    [00:17:43] DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    https://arxiv.org/abs/2403.12945


    ---

    RESCRIPT: https://app.rescript.info/share/e2385948cf465f0d6a2c0930150fc3ab

    Show More Show Less
    26 mins
  • Speech Recognition Is Not a Solved Problem — Pavan Muddireddy
    Sep 14 2026

    Pavankumar Reddy Muddireddy leads audio research at Mistral AI. He joins Tim Scarfe for a deep technical tour of Voxtral — and explains why the frontier of deployed voice is still a cascade of specialised models rather than one end-to-end system.


    IN PARTNERSHIP WITH MISTRAL AI:

    ---

    This episode was produced in partnership with Mistral AI.

    Mistral AI: https://mistral.ai/

    ---


    The conversation opens on architecture. Voxtral Chat feeds a 3B Ministral text trunk with continuous embeddings from an audio encoder, passed to the decoder as direct token input rather than through cross-attention as in Whisper, so the model can answer questions about emotion, timing and who spoke when without an intermediate transcript to lose them. The real-time model becomes a dual-stream decoder that consumes audio and emits text at once, at a target delay down to 160ms, with slower streams in parallel for anything that can wait for more context.


    On generation, Pavan explains why Voxtral TTS predicts continuous latents rather than discrete codec tokens, traces the lineage from SoundStream through EnCodec to Mimi's split of semantic and acoustic codebooks, and places FSQ and flow matching in it. Tim presses on the priors underneath: why a mel spectrogram instead of raw waveform, what noise augmentation buys, and when acoustic overfitting becomes somebody's fine-tuning problem. Then the failure modes. Diarisation is emitted autoregressively inside the transcript rather than by a separate head, which makes streaming diarisation fragile — less context, late speaker changes, invented extra speakers. And because the architecture commits to what it has already predicted, one out-of-distribution mistake compounds into looping or skipped segments, which is what DPO corrects: the negative supervision pre-training and SFT cannot give.


    The last third is the argument Tim keeps returning to. Customers running voice agents over millions of sessions describe scaffolding, not a solved problem, with a sharp drop outside the top few languages. Cascades survive because each component stays separately adaptable, observable and constrainable. And voice alone is cognitive debt: absorbing information and deciding in one serial stream is harder than glancing at a menu. Voice becomes ubiquitous beside a screen, not instead of one.


    ---

    TIMESTAMPS:

    00:00:00 Cold open

    00:00:46 Why Mistral moved into audio

    00:09:27 Inside Voxtral: trunk, encoder, dual streams

    00:20:22 Speech that works in real time

    00:30:52 How a voice becomes tokens

    00:39:59 Flow matching, FSQ and the new codec

    00:52:51 When speech models lose the speaker

    01:03:23 Correcting hallucinations with preferences

    01:12:12 Controlling synthetic speech

    01:20:06 Why cascades still win

    01:29:25 Speech in the wild

    01:33:46 Audio models as interfaces

    01:37:54 Why voice still needs a screen


    ---

    REFERENCES:

    paper:

    [00:01:42] Mistral 7B

    https://arxiv.org/abs/2310.06825

    [00:09:38] Voxtral

    https://arxiv.org/abs/2507.13264

    [00:14:41] Whisper: Robust Speech Recognition

    https://arxiv.org/abs/2212.04356

    [00:19:11] Voxtral Realtime

    https://arxiv.org/abs/2602.11298

    [00:21:52] Delayed Streams Modeling (Kyutai)

    https://arxiv.org/abs/2509.08753

    [00:30:52] Voxtral TTS

    https://arxiv.org/abs/2603.25551

    [00:32:38] SoundStream neural audio codec

    https://arxiv.org/abs/2107.03312

    [00:34:59] Flow Matching for Generative Modeling

    https://arxiv.org/abs/2210.02747

    [00:37:03] EnCodec: High Fidelity Neural Audio Compression

    https://arxiv.org/abs/2210.13438

    [00:37:42] Moshi and the Mimi codec

    https://arxiv.org/abs/2410.00037

    [00:39:05] Finite Scalar Quantization (FSQ)

    https://arxiv.org/abs/2309.15505

    [01:03:33] Direct Preference Optimization (DPO)

    https://arxiv.org/abs/2305.18290

    dataset:

    [00:46:14] Mozilla Common Voice

    https://commonvoice.mozilla.org/en/datasets

    organization:

    [00:50:47] Hugging Face

    https://huggingface.co/

    Show More Show Less
    1 hr and 42 mins
  • How Replication Could Teach Machines What Good Science Looks Like — Edward Hughes
    Sep 11 2026

    Can a machine learn the judgement that separates a plausible-looking result from a faithful experiment? Edward Hughes, Chief Scientist and co-founder of Inherent, joins Tim Scarfe to argue that creativity is not optimisation, and that the missing capability in AI is choosing which questions are worth asking.


    SPONSOR:

    ---

    Cyber Fund built the Monastery to help founders ship products that were impossible a year ago.

    Apply now: https://cyber.fund

    ---


    Edward makes the case that Move 37 was innovative rather than creative, and that the field, not the individual, decides what counts as a discovery. That reframing runs through Csikszentmihalyi, Deutsch and exaptation into open-endedness, where deceptive goals and imperfect world models turn out to be the point rather than the problem. The second half turns to the paper: Replica, a task space built by redacting figures from real papers, and Faraday, a 27-billion-parameter model trained to steer a frontier coding agent that then beats the frontier on held-out replications.


    ---

    TIMESTAMPS:

    00:00:00 Cold open: Move 37, Faraday and collective intelligence

    00:01:08 Sponsor: CyberFund

    00:01:46 Inherent's $50M raise and the road from string theory

    00:09:14 Three timescales of learning: weights, context, culture

    00:13:47 Move 37 was innovative, not creative: the field decides

    00:20:39 Creativity as satisficing: the urinal and evolution

    00:25:06 Exaptation and the Tristan chord: creativity in context

    00:30:56 Coherence for whom? Deutsch's hard-to-vary explanations

    00:35:53 Why copying is creative: Deutsch and the constraint engineer

    00:42:27 Societies of agents and the strong Moravec paradox

    00:45:51 Evaluate in hindsight: from Lean proofs to climate change

    00:51:56 Picbreeder, local goals and why discovery needs deception

    00:57:21 Spaghetti proofs, translation layers and superhuman Go

    01:00:37 Does nature compress? Naturalness and real patterns

    01:07:36 Why replicate? Replica's redacted figures and Faraday

    01:12:31 Faraday beats Codex, Claude and GLM 5.2 on held-out tasks

    01:15:31 Replication to innovation: how the Transformer happened

    01:18:26 Deep replication: what Faraday learns from Voyager and GNoME

    01:23:37 Can the AI scientist cheat? Goodharting the judge

    01:29:09 Inside Replica: scale-down, 8xB300 runs, per-task rubrics

    01:34:11 The RL crisis: getting GRPO to work with per-turn credit

    01:39:43 Weights vs harnesses: AlphaEvolve, DGM and EvoTune

    01:45:45 The recursive company: agents cross a phase transition

    01:50:35 Collective intelligence and the electric dynamo

    01:55:46 What replaces OKRs? Incumbents and the burden of knowledge


    ---

    REFERENCES:

    MLST Creativity Article:

    https://archive.mlst.ai/read/why-creativity-cannot-be-interpolated


    organization:

    [00:01:47] Inherent

    https://inherentlabs.ai/

    other:

    [00:20:51] Marcel Duchamp, Fountain

    https://www.tate.org.uk/art/artworks/duchamp-fountain-t07573

    [00:05:19] Human-Timescale Adaptation in an Open-Ended Task Space (Adaptive Agent)

    https://arxiv.org/abs/2301.07608

    [00:06:05] The AI Scientist

    https://arxiv.org/abs/2408.06292

    [00:12:13] Training AI Scientists to Replicate Research (Replica and Faraday)

    https://arxiv.org/abs/2608.13331

    [01:44:46] Evolutionary Principles in Self-Referential Learning

    https://people.idsia.ch/~juergen/diploma.html

    [01:59:33] Are Ideas Getting Harder to Find?

    https://www.nber.org/papers/w23782

    book:

    [00:16:04] Creativity: Flow

    https://search.worldcat.org/title/254487436

    [00:26:22] Why Greatness Cannot Be Planned

    https://link.springer.com/book/10.1007/978-3-319-15524-1

    [00:33:03] The Beginning of Infinity

    https://www.penguinrandomhouse.com/books/293575/the-beginning-of-infinity-by-david-deutsch/

    [01:55:47] Laws of Knowledge

    https://www.penguin.co.nz/books/the-infinite-alphabet-9780241655672


    (Full list refs on YT/rescript)

    ---

    RESCRIPT:

    https://app.rescript.info/session/670296ba913761d0?share=6281911cac9bdbff637f10819d4d1e5c

    Show More Show Less
    2 hrs and 2 mins
  • AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen
    Sep 8 2026

    Could slowing AI development make superintelligence safer? Daniel Kokotajlo and Thomas Larsen of the AI Futures Project join Tim Scarfe to examine AI 2040: Plan A, a proposal to buy time before AI exceeds human control.


    SPONSOR:

    ---

    Cyber Fund built the Monastery to help founders ship products that were impossible a year ago.

    Apply now: https://cyber.fund

    ---


    After revisiting AI 2027 and the limits of forecasting, they ask what happens when AI can automate research and sustain an economy without human workers. Tim challenges the case for general models and asks whether intelligence alone explains power. Plan A proposes an initial pause to build safety infrastructure, then cautious development up to the strongest AI that can still be reliably controlled. The discussion tests the distinction between control and alignment, the case for public AI research, and whether the US and China could enforce a slowdown. It ends with the evidence that would change their forecasts.


    ---

    TIMESTAMPS:

    00:00:00 AI 2040: a slower route to superintelligence

    00:01:34 Sponsor: Cyber Fund

    00:02:12 From OpenAI to AI 2027

    00:06:58 Forecasts, war games and self-fulfilling prophecies

    00:17:44 Why AI sceptics are changing their minds

    00:23:04 When AI can replace its own researchers

    00:28:45 Could an AI economy grow without human workers?

    00:37:32 One general model or a society of specialists?

    00:47:43 Brains, machines and collective intelligence

    00:56:12 Plan A: buy time at the controllable frontier

    01:00:02 Why control buys time but cannot replace alignment

    01:06:36 Why AI research should be public

    01:10:32 Can the US and China enforce an AI slowdown?

    01:19:04 Why AI policy debates miss the technology

    01:21:56 Is AI normal technology? The remaining disagreement


    Many thanks to James Wilken-Smith for helping with show research.


    ---

    REFERENCES:

    other:

    [00:00:01] AI 2040: Plan A

    https://ai-2040.com/

    [00:03:27] AI 2027

    https://ai-2027.com/

    [00:13:47] Scenario Scrutiny for AI Policy

    https://blog.aifutures.org/p/scenario-scrutiny-for-ai-policy

    [00:33:11] The 2028 Global Intelligence Crisis

    https://www.citriniresearch.com/p/2028gic

    [01:00:40] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    https://www.redwoodresearch.org/research/hugging-face-incident

    [01:09:21] The Hugging Face incident and the road ahead

    https://openai.com/index/hugging-face-incident-and-the-road-ahead/

    [01:22:01] AI as Normal Technology

    https://www.normaltech.ai/p/ai-as-normal-technology

    [01:22:51] Common Ground between AI 2027 & AI as Normal Technology

    https://asteriskmag.substack.com/p/common-ground-between-ai-2027-and

    person:

    [00:19:43] Geoffrey Hinton

    https://www.cs.toronto.edu/~hinton/

    [00:20:07] Ryan Greenblatt

    https://www.lesswrong.com/users/ryan_greenblatt

    [00:26:06] Elon Musk

    https://www.tesla.com/elon-musk

    tool:

    [00:21:46] ARC-AGI-3

    https://arcprize.org/arc-agi/3

    [00:21:53] AlphaGo and Move 37

    https://deepmind.google/research/alphago/

    [00:39:41] Claude

    https://claude.com/product/overview

    [00:39:58] NVIDIA H100 GPU

    https://www.nvidia.com/en-us/data-center/h100/

    paper:

    [00:24:42] Training AI Scientists to Replicate Research

    https://arxiv.org/abs/2608.13331v1

    [01:27:19] Validity of the single processor approach to achieving large scale computing capabilities

    https://www.cs.cmu.edu/~18742/papers/Amdahl1967.pdf

    book:

    [00:28:52] Bullshit Jobs: A Theory

    https://www.simonandschuster.com/books/Bullshit-Jobs/David-Graeber/9781501143335

    organization:

    [01:05:09] Redwood Research

    https://www.redwoodresearch.org/


    ---

    RESCRIPT:

    https://app.rescript.info/public/share/33d1a58fa8f307ae7dfd504d4fdaa9d5

    Show More Show Less
    1 hr and 30 mins
  • Designing How AI Grows — Tom McGrath
    Sep 2 2026

    Tom McGrath is co-founder and Chief Scientist at Goodfire, and a former Google DeepMind researcher. He joins Tim Scarfe to ask what neural networks actually learn, whether their internal representations converge on structures in the world, and whether interpretability can extract new scientific knowledge rather than merely explain model outputs.


    Beginning with AlphaZero and learned modularity, the conversation moves into neural geometry: concept manifolds, reusable computation inside Llama, and why activation steering can fail when it pushes a model off-manifold. McGrath then makes the case for intentional design, using interpretability as part of the training loop. They examine controlled generalisation, features as rewards, predictive data debugging, and the uncomfortable fact that a model may recognise a hallucination or reward hack and still produce it.


    The discussion closes on grader awareness, oversight and collusion between adaptive agents, then returns to sparse autoencoders. SAEs are useful, McGrath argues, but they may fracture the higher-dimensional structures networks actually use. This episode was made with support from Goodfire.


    ---

    TIMESTAMPS:

    00:00:00 Introduction: Can interpretability speed-run science?

    00:02:03 The invisible grader

    00:06:51 What AlphaZero learned from the world

    00:12:24 Interpretability as a control loop

    00:21:54 The forbidden method and safer interventions

    00:37:36 Why models catch hallucinations too late

    00:46:19 Debug the dataset before training

    00:50:44 Why neural networks become modular

    00:55:57 Finding the geometry inside a network

    01:02:55 Why steering falls off the manifold

    01:12:10 A reusable calculator inside Llama

    01:17:19 From abstractions to goals

    01:25:28 Reward hacking, oversight and collusion

    01:37:23 Are sparse autoencoders dead?


    ---

    REFERENCES:

    paper:

    [00:05:45] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

    https://arxiv.org/abs/2502.17424v7

    [00:11:05] Acquisition of Chess Knowledge in AlphaZero

    https://arxiv.org/abs/2111.09259

    [00:25:30] Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning

    https://arxiv.org/abs/2507.16795

    [00:29:30] Persona Vectors: Monitoring and Controlling Character Traits in Language Models

    https://arxiv.org/abs/2507.21509

    [00:41:14] Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability

    https://arxiv.org/abs/2602.10067

    [00:47:03] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

    https://arxiv.org/abs/2606.12360

    [01:00:26] Do Sparse Autoencoders Capture Concept Manifolds?

    https://arxiv.org/abs/2604.28119

    [01:03:04] Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior

    https://arxiv.org/abs/2605.05115

    [01:14:20] Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts

    https://arxiv.org/abs/2605.01148

    [01:29:35] Measuring Reward-Seeking via Contrastive Belief Updates

    https://arxiv.org/abs/2607.18966v1

    other:

    [00:15:44] Intentional Design

    https://www.goodfire.com/blog/intentional-design

    [00:56:12] The World Inside Neural Networks

    https://www.goodfire.com/research/the-world-inside-neural-networks

    [01:37:28] A Pragmatic Vision for Interpretability

    https://www.alignmentforum.org/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-interpretability


    ---

    RESCRIPT:

    https://app.rescript.info/share/846cfee4131b664fd09209cc3b98018e

    Show More Show Less
    1 hr and 40 mins