The Agentic Quality Podcast
PODCAST GUEST · CONTEXTQAWalked through how a team tests AI systems where the same input can produce a different output every time. Instead of validating fixed outputs, the approach shifted to testing behavior — whether the right services were called, the right rules were followed, and the system behaved correctly even when the final result changed. Deployment velocity grew from roughly 4 per month to 4 per week, eventually reaching 1,000+ deployments in a year with no manual intervention. The episode also covers agents in production, debugging with agents, human-in-the-loop decisions, trust boundaries and observability.