Super Data Science: ML & AI Podcast with Jon Krohn · Jon Krohn

963: Reinforcement Learning for Agents, with Amazon AGI Labs’ Antje Barth

·51 min·3 clips
An agent that's reliable 60% of the time is, in practice, 0% useful.
1. Super Data Science: ML & AI Podcast with Jon Krohn centers this episode on Amazon AGI Labs' NovaAct and reinforcement learning for agents. 2. Jon Krohn interviews Antje Barth, a technical staff member at Amazon AGI Labs, a developer-relations lead, and a multi-time O'Reilly author. 3. The episode asks what makes a browser agent reliable enough to move from a playground into production. 4. Barth says NovaAct had a research preview since March last year and launched as a GA service at AWS re:Invent. 5. She says developers can start in a free playground at nova.amazon.com/act without an AWS account. 6. The playground takes a website URL and natural-language instructions such as signing up for an event or filling out a form. 7. She says the browser environment is embedded in the playground so users can watch the agent act without extra setup. 8. The system shows reasoning traces while the agent runs, which helps developers debug instructions and troubleshoot failures. 9. NovaAct generates a Python script from the natural-language workflow and lets developers move it into an IDE or SDK. 10. The IDE extension includes live preview and unified tracing so the browser stays inside the coding environment. 11. Barth says a 60% reliable agent is effectively 0% useful for real-world production use cases. 12. She says Amazon's goal is to reach 90% and above on workflow execution, not just impressive demos. 13. NovaAct uses reinforcement-learning-based web gyms that replicate common UI tasks such as form filling and shopping workflows. 14. She says agents self-play through thousands of tasks, fail, retry, and learn better ways to complete the job. 15. Barth says the approach helps agents generalize when buttons move, icons change, or sites switch between light mode and dark mode. 16. She says public benchmarks like Work Arena and RealBench matter, but customer-specific evaluations on real company tasks matter more. 17. The tone is technical and practical, with Jon Krohn asking process questions and Barth answering with deployment details. 18. The format stays interview-driven, with product walkthroughs, benchmark discussion, and examples from Amazon customers and internal teams. 19. Listeners building AI agents, QA workflows, or enterprise automation will get the most from this episode. 20. Listeners looking for broad AI news without product-level detail may skip this one.

As heard by us

A grounded case for making AI agents reliable in everyday work before calling them useful.

The episode treats a useful AI agent as something people can trust in ordinary work, not just admire in a demo. Antje Barth, joining from Amazon AGI Labs in San Francisco, and host John Krohn keep the discussion anchored in expense reports, web research, meeting summaries, and…

Read the full review in PlayNext →

Why you'd press play

You want AI agents that help with real work and feel like a digital co-worker, not just a tool.

Read the full recommendation in PlayNext →
Listen to the show on