Conversational AI · Product system
Voice AI Evaluation Suite
A full-stack evaluation workspace for turning unpredictable voice-agent behavior into repeatable scenarios, traces, quality signals, and release decisions.

The challenge
Start with the product question.
Conversational agents can sound convincing in a demo and still fail unpredictably in production. Teams need a shared way to test customer scenarios, inspect behavior, compare versions, and decide whether a release is ready.
What I built
A working system to learn from.
I built a locally hosted workflow for authoring personas and scenarios, running multi-trial evaluations, using model-based judges, inspecting transcripts and tool activity, and comparing quality, performance, and cost signals. It supports chat as well as voice-agent testing through Twilio and LiveKit.
Product decisions
The thinking behind the build.
- 01
Model non-determinism as a product quality problem, not just an engineering detail.
- 02
Separate simulation, execution, scoring, and analysis so failures are easier to diagnose.
- 03
Pair automated judgment with human review and disagreement analysis.
- 04
Make traces useful for release conversations across product and engineering.
This is an independent project shared to demonstrate product thinking and hands-on exploration. It does not represent confidential employer work or commercial performance claims.

