Conversational AI · Product system

Voice AI Evaluation Suite

A full-stack evaluation workspace for turning unpredictable voice-agent behavior into repeatable scenarios, traces, quality signals, and release decisions.

Independent project2026
Voice AI Evaluation Suite dashboard showing quality and latency signals

The challenge

Start with the product question.

Conversational agents can sound convincing in a demo and still fail unpredictably in production. Teams need a shared way to test customer scenarios, inspect behavior, compare versions, and decide whether a release is ready.

What I built

A working system to learn from.

I built a locally hosted workflow for authoring personas and scenarios, running multi-trial evaluations, using model-based judges, inspecting transcripts and tool activity, and comparing quality, performance, and cost signals. It supports chat as well as voice-agent testing through Twilio and LiveKit.

Product decisions

The thinking behind the build.

  1. 01

    Model non-determinism as a product quality problem, not just an engineering detail.

  2. 02

    Separate simulation, execution, scoring, and analysis so failures are easier to diagnose.

  3. 03

    Pair automated judgment with human review and disagreement analysis.

  4. 04

    Make traces useful for release conversations across product and engineering.

About this project

This is an independent project shared to demonstrate product thinking and hands-on exploration. It does not represent confidential employer work or commercial performance claims.

Next projectInvestment Advisor