Machine Learning Street Talk (MLST) cover art

Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

By: Machine Learning Street Talk (MLST)
Listen for free

Welcome! We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-depth analysis. Our approach is unrivalled in terms of scope and rigour – we believe in intellectual diversity in AI, and we touch on all of the main ideas in the field with the hype surgically removed. MLST is run by Tim Scarfe, Ph.D (https://www.linkedin.com/in/ecsquizor/) and features regular appearances from MIT Doctor of Philosophy Keith Duggar (https://www.linkedin.com/in/dr-keith-duggar/).Machine Learning Street Talk (MLST)
Episodes
  • When AI Research Starts Moving Faster Than Human Research - Zhengyao Jiang
    Sep 26 2026

    Weco let an AI coding agent rewrite the harness around another agent for eight days: its code, prompts and tools, while the underlying language model stayed fixed. Tim Scarfe asks Weco co-founder Zhengyao Jiang what the reported gains over two years of human engineering actually demonstrate.The discussion examines AIDE 85's generated code, held-out evaluation and the difficulty of separating useful discoveries from reward hacking. Jiang explains Weco's four levels of recursive self-improvement and compares the experiment with AlphaEvolve and the Darwin Gödel Machine.The limits matter as much as the gains. Jiang explains why the experiment did not establish that the system had become a better improver. The conversation closes with open-ended search, human-designed primitives and Parameter Golf: where does the next useful idea come from when the agent is searching inside a space that people designed?---TIMESTAMPS:00:00:00 Eight days of self-improvement: what counts?00:03:25 AIDE and the puzzle of useful spaghetti code00:08:38 Four levels of recursive self-improvement00:12:02 What AIDE 85 changed and how it was tested00:20:04 AlphaEvolve, Darwin Gödel Machine and the RSI claim00:26:21 Reward hacking and the limits of detection00:33:09 Open-ended search, harness tuning and creativity00:39:43 Parameter Golf and the limits of self-improvement---REFERENCES:organization:[00:00:30] Weco AIhttps://www.weco.ai/other:[00:00:33] AIDE²: The First Evidence of Recursive Self-Improvementhttps://www.weco.ai/blog/first-evidence-of-recursive-self-improvement[00:14:11] Faulty reward functions in the wildhttps://openai.com/index/faulty-reward-functions/[00:29:59] The Hugging Face incident and the road aheadhttps://openai.com/index/hugging-face-incident-and-the-road-ahead/tool:[00:03:29] AIDEhttps://github.com/WecoAI/aideml[00:04:29] MLE-benchhttps://github.com/openai/mle-bench[00:04:33] ALE-Benchhttps://github.com/SakanaAI/ALE-Bench[00:04:52] WeatherBench 2https://github.com/google-research/weatherbench2[00:08:18] ReActhttps://react-lm.github.io/[00:39:43] Parameter Golfhttps://github.com/openai/parameter-golfpaper:[00:20:08] AlphaEvolve: A coding agent for scientific and algorithmic discoveryhttps://arxiv.org/abs/2506.13131v1[00:21:35] Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agentshttps://arxiv.org/abs/2505.22954v3[00:23:45] Hyperagentshttps://arxiv.org/abs/2603.19461v1[00:27:01] SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agentshttps://arxiv.org/abs/2605.21384book:[00:33:14] Why Greatness Cannot Be Planned: The Myth of the Objectivehttps://link.springer.com/book/10.1007/978-3-319-15524-1---LINKS:https://app.rescript.info/share/3a9dc6189cb539c6a05fcc4f75c101b3PDF:https://app.rescript.info/api/public/sessions/9eda60ede2b31c92/pdf

    Show More Show Less
    44 mins
  • Why Scaling Prediction Cannot Create Intelligence - Alexander Mattick
    Sep 21 2026

    Alexander Mattick is a researcher at Fraunhofer IIS and a PhD researcher at the University of Technology Nuremberg (UTN), and a regular on Yannic Kilcher's Discord. He first came on MLST in 2022, after helping research the Yann LeCun and Randall Balestriero episode on interpolation.


    SPONSOR:

    ---

    Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open.

    Apply now: https://cyber.fund

    ---


    Alexander treats inference as the thread running through modern machine learning: once you have a model, what does it cost to get an answer out of it? He works through Monte Carlo, GFlowNets, energy-based models, diffusion, normalising flows and flow matching, with four short explainers he recorded himself. He is blunt about energy-based models: you can sample from them in principle, but it is rarely worth the compute. JEPA and "world model", he says, are closer to branding than to technical categories.


    Next: theories of deep learning, none of which he thinks predicts enough yet to guide practice, then reinforcement learning.


    ---

    0:00 Cold open: information is expensive

    0:51 Welcome back, Alexander Mattic

    2:08 Alexander's research background

    2:50 Inference: densities, sampling and Monte Carlo

    6:42 GFlowNets, energy functions and MCMC

    9:45 Explainer: energy-based models

    11:03 Why model a density at all?

    17:30 From learned energies to flow matching

    25:08 Explainers: diffusion and normalising flows

    28:33 Are energy-based models generative?

    33:22 JEPA, contrastive learning and collapse

    41:13 Why non-language modalities need flows

    44:51 Inference as search: branch and bound

    49:43 Q-learning and delayed consequences

    55:14 Flow matching, optimal transport, Fokker-Planck

    1:00:03 Explainer: flow matching

    1:01:49 AlphaFold, latents and scale versus architecture

    1:07:52 Two families of deep learning theory

    1:15:04 What a good theory would predict

    1:23:53 The manifold hypothesis and compression

    1:28:25 Is reward enough?

    1:32:01 Control theory versus reinforcement learning

    1:37:22 The Bitter Lesson and expensive information

    1:42:08 Constrained RL: the constrained MDP toolbox

    1:50:12 Creativity as constrained search

    1:55:44 Reality is protean: when abstractions hold

    2:00:32 What is a world model?

    2:04:38 Prediction is not control

    2:08:13 Robot demos, MPC and reliability


    ---

    REFERENCES:

    [6:55] GFlowNets (Bengio et al., 2021)

    https://arxiv.org/abs/2106.04399

    [38:46] Contrastive Self-Supervised Learning (Anand, 2020)

    https://ankeshanand.com/blog/2020/01/26/contrative-self-supervised-learning.html

    [38:56] LeJEPA (Balestriero and LeCun, 2025)

    https://arxiv.org/abs/2511.08544v3

    [47:10] RL for Node Selection in Branch-and-Bound (Mattick)

    https://openreview.net/forum?id=0ez68a5UqI

    [56:20] Flow Matching for Generative Modeling

    https://arxiv.org/abs/2210.02747v2

    [1:12:41] Disentangling feature and lazy training in deep neural networks

    https://arxiv.org/abs/1906.08034v4

    [1:31:05] Reward is enough (Silver)

    https://doi.org/10.1016/j.artint.2021.103535

    [1:35:12] Learning ReLU networks to high uniform accuracy is intractable (Berner et al.)

    https://arxiv.org/abs/2205.13531v2

    [1:40:20] Dota 2 with Large Scale Deep RL

    https://arxiv.org/abs/1912.06680v1

    [1:45:41] Constrained Update Projection for Safe Policy Optimization (Yang et al., 2022)

    https://arxiv.org/abs/2209.07089

    [1:46:11] SafeMPO (ICLR 2026)

    https://openreview.net/forum?id=1m0EU6QXj6

    [1:50:17] Why Creativity Cannot Be Interpolated

    https://archive.mlst.ai/paper/why-creativity-cannot-be-interpolated/

    [1:51:39] Invalid Action Masking (Huang and Ontañón)

    https://arxiv.org/abs/2006.14171

    [2:00:04] Probability Theory: The Logic of Science (Jaynes, 2003)

    https://www.cambridge.org/core/books/probability-theory/9CA08E224FF30123304E6D8935CF1A99

    [2:01:53] Training Agents Inside of Scalable World Models (Hafner et al., 2025)

    https://arxiv.org/abs/2509.24527v1

    [2:03:43] World Models (Ha and Schmidhuber, 2018)

    https://arxiv.org/abs/1803.10122v4

    Show More Show Less
    2 hrs and 14 mins
adbl_web_anon_alc_button_suppression_t1
No reviews yet