The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · Sam Charrington

The Evolution of Reasoning in Small Language Models with Yejin Choi

·1 hr 6 min·4 clips
Even with higher temperature, Yejin Choi says Llama, ChatGPT, and DeepSeek R1 can behave strikingly similarly.
1. The TWIML AI Podcast episode centers on Yejin Choi’s work on reasoning in small language models. 2. Sam Charrington interviews Yejin Choi, a Stanford University professor and senior fellow in the Computer Science Department and the Institute for Human-Centered AI. 3. The episode asks how small language models can reason better without relying only on the largest GPU-heavy systems. 4. Choi says her earlier work focused on common sense, knowledge, reasoning, and natural language generation. 5. She now emphasizes large language models, small language models, large reasoning models, small reasoning models, and pluralistic norms and values. 6. Choi links that shift to democratizing generative AI for academics who cannot buy as many GPUs. 7. She argues that the field has poured far more investment into scaling up than into exploring smaller counterparts. 8. Choi describes several routes to smaller models, including quantization, pruning neurons, and hybrid architectures like Mamba Hybrid from NVIDIA. 9. She also points to better post-training data as a major lever, especially data beyond what the internet already provides. 10. Choi says that supervised fine-tuning, reinforcement learning, and expert data from lawyers or former International Math Olympiad winners are already part of the pipeline. 11. She says synthetic data needs careful prompting, iteration, and revision because a simple request to ChatGPT can produce repetition. 12. Choi gives hard math solutions as a concrete case where synthetic data can create qualitatively new examples that do not exist on the internet. 13. She describes using reinforcement learning with verifiers to explore candidate solutions and keep the ones that check out. 14. Choi explains that the Artificial Hive Mind paper found intra-model homogeneity and inter-model homogeneity after post-training. 15. She names Llama, ChatGPT, and DeepSeek R1 as models that can produce strikingly similar behavior on open-ended questions. 16. Choi says that concern led to spectrum tuning, a post-training method from her former student Taylor Sorenson that aims to preserve a spectrum of outputs. 17. The conversation stays technical and interview-driven, with Choi answering long, method-focused questions from Charrington. 18. The tone stays analytical, with repeated examples from math reasoning, post-training, and data curation. 19. Listeners interested in small language models, reasoning, and AI training pipelines would likely get the most from it. 20. Listeners wanting a light overview without model-training detail may skip it.

As heard by us

A measured look at model homogeneity, spectrum tuning, and AI's mixed human effects.

The piece takes a clear technical concern and keeps it in focus: models can stay less diverse than expected even across repeated runs, and Yejin Choi uses that to explore both intra-model and inter-model homogeneity.

Read the full review in PlayNext →

Why you'd press play

You want a grounded take on why small models can still sound alike.

Read the full recommendation in PlayNext →
Listen to the show on