Almog Baku has shipped more than 30 AI agents into production and founded the 7,000-strong GenAI Israel community. Now he's building infrastructure to help teams do root-cause analysis on agents once they're live. His argument in this episode: most teams treat building the agent as the whole job. Realistically, it’s half of it.The other half is designing, in advance, how you'll learn why it fails.
Hear him outline:
Why he tells clients to warn their CEO in advance that the first production version will be "an epic failure," and builds the learning infrastructure around that assumption
The LLM Triangle Principles: an SOP modeled on a real domain expert, paired with the right model, prompting technique, and data- in that order, not model-first
Why the Valentine's Day game he built for his wife in four days turned out only "moderate at best," and what that taught him about skipping domain expertise
Why he treats eval scores as clues that point to a trace worth reading
The difference between agentic workflows he'll let run loose (personal tools like Claude Code) and production pipelines that need staged quality
The "credit assignment" problem: when a chain of sub-agents fails, figuring out which one is to blame, and why he thinks it's currently unsolved at the infrastructure level
:quality(80))