Andy Spezzatti
I work on how to know whether AI agents actually work.
At Zapier I build the evaluation infrastructure for agentic automation: simulated worlds that agents act in, end-to-end scenarios grounded in real outcomes, and the tooling that turns those measurements into release decisions.
I've been circling the same question for a while. At Genentech I built simulation engines so a portfolio of three hundred clinical trials could be planned before reality arrived. At Taltrics, which I co-founded, we shipped multi-agent systems to enterprises and post-trained them to stop hallucinating tool calls. The instrument keeps changing; the question — how do you trust a system before it meets the world? — hasn't. Along the way I published on trustworthy AI in healthcare.
Now — world simulation, eval-driven development, benchmarking agent performance. Sep 2026
Views here are my own.
- Aug 2026Which branch would you bet onA leaderboard reports scores. A decision needs a preference and an uncertainty, and frontier scenarios are exactly where you have neither by default.
- Jun 2026Eval-Driven DevelopmentAgents move decisions out of control flow and into the model. That part of the specification can only be recovered by watching the system act, which makes the eval the place where it gets written.
- May 2026Known-bad candidatesAn eval that has never been shown a wrong answer has not been shown to measure anything.
- Apr 2026Five of five is not a statisticWhat repeated eval runs can tell you, what they cannot, and why the two get confused.
- Dec 2022Priors for a portfolioA model's job is to be honest about what it doesn't know, in a form a decision can use. Three years of simulating clinical trials taught me what that means.
- 2025 –Eval infrastructure for agentic automationZapier — simulated worlds and release gates for an agent platform used by 1.5M+ people.
- 2022 – 25Scoping and pricing a deal on the callTaltrics — agents that take an RFP to a priced proposal in minutes, post-trained so the number is right.
- 2019 – 22Simulating a clinical-trial portfolioGenentech / Roche — a discrete-event engine behind planning for 300+ trials in 70 countries.
- 2023Sailing the Data Sea to Advance Research on the Sustainable Development GoalsThe Ethics of Artificial Intelligence for the Sustainable Development Goals, Springer
- 2022
- 2021Co-design of a trustworthy AI system in healthcare: deep learning based skin lesion classifierFrontiers in Human Dynamics
- Feb 2026Build-Along Workshop on Agentic AIZapier × Gamma
- Sep 2024AI and the Future of Global SustainabilitySocially Conscious AI Podcast
- Jun 2022
- Jul 2020