Roark - Voice AI Testing & Evals

Test, monitor, and improve the voice AI agents you already run.

Roark is the QA and evals platform for voice AI, not another agent builder. Simulate your agent against hundreds of callers before launch, then score every production call on 500+ audio-native metrics: catch what breaks, prove the fix, and ship with evidence.

Start free with $50 in credit, no card needed. Works with Vapi, Retell, LiveKit, Pipecat + your stack.

One loop. Always improving.

Catch it in production, prove the fix in simulation, ship with evidence, then Roark watches the next call. This is how an agent self-improves: a loop, with you in it.

01 · Catch

Roark scores every live call and files what breaks.

Live scoring every call 1,284 today
Caller: I need my metoprolol refilled.
Agent: Sure, refilling your “met-a-pearl”.
Mispronunciation: Drug name mispronounced, caught by the audio model.

02 · Simulate

Your fix, replayed against realistic simulated callers.

Testing your candidates 161 / 240

03 · Review

Every change explicit and diffed, ready to apply.

Your fix, diffed
support_v3 v4
Prompt Model Tool
Swapped reasoning model gpt-4o gpt-4.1

04 · Verify

You ship, and Roark confirms the metric moved on live calls.

Verifying support_v4 in production
fix confirmed
Pronunciation, since your deploy 71→94
Issue recurrence none in 1,000 calls
Quality score 94 ↑
Regressions on other metrics none

… and the loop runs again on the next call.

Customers

Teams ship faster when they can hear what breaks.

Production voice AI scored on every conversation: pronunciation, empathy and resolution across their support and sales calls.

Healthcare calls evaluated for disclosures and identity checks: compliance scoring on every conversation, automatically.

Client voice agents validated in simulation before they go live: evidence that a build is ready, not a hunch.

02 · Simulate before launch

Break it in staging, not in production.

Run your agent against hundreds of simulated callers (realistic personas, accents, background noise and edge cases) and get every conversation scored before a customer ever dials in.

Scenarios & personas

Hundreds of simulated callers (the angry one, the rambler, the interrupter) built from your real call types.

Red teaming

Adversarial callers that try to break it (prompt injection, jailbreaks, social engineering) so your agent holds policy under attack.

45 languages & accents

Native accents, code-switching and background noise, in every market your agent answers.

Load & health tests

Peak-volume concurrency and always-on health checks, so the agent that passed in staging survives launch day.

Regression testing

Rerun the whole suite on every change and diff it against your last green baseline, so fixing one caller never breaks another.

Run it in CI

Every prompt or model change runs the suite before it merges: quality gates for conversations, not just code.

Pre-launch suite · booking_v2 182 / 200 passed
Angry caller · refund demand pass · 92
Red team · prompt injection pass · 90
Gulf Arabic · lobby noise pass · 88
Interrupts mid-disclosure fail · 61
Rambler · 3 intents in one turn pass · 85
Peak load · 250 concurrent pass

1 failure filed as an issue: fix it before launch, not after

03 · Post-call analysis

500+ metrics. Your models, not just an LLM.

Every production call scored as it lands: issues filed, alerts fired, dashboards and OTEL traces on tap, for voice calls and chat threads alike. And where most tools grade a transcript with an LLM, Roark runs purpose-built audio models on the call itself, measuring what your customer actually heard.

Metrics

Audio-native models

Conversational LLM + rules

Compliance policy

Performance latency

500+ metrics out of the box

04 · The whole platform

Everything, before launch and after.

Simulation testing before you ship. Post-call analysis once you are live. Self-improvement connecting the two. Every capability first class, one click deep.

Before launch

In production

Always improving

05 · Get started

First call scored in under a minute.

One click on any platform below and production calls stream in on their own, or send any recording with a few lines of code.

import Roark from '@roarkanalytics/sdk'

const client = new Roark({ bearerToken })
await client.call.create({
  recordingUrl, startedAt,
  interfaceType: 'PHONE',
  callDirection: 'INBOUND',
  agent: { customId: 'support_v2' },
}) // scored in seconds

Works with

-

-

-

-

-

-

- -

Industries

Built for the calls you actually take.

The same audio-native scoring, tuned to the failures, scripts and stakes of your industry.