Roark - Voice AI Testing & Evals
Test, monitor, and improve the voice AI agents you already run.
Roark is the QA and evals platform for voice AI, not another agent builder. Simulate your agent against hundreds of callers before launch, then score every production call on 500+ audio-native metrics: catch what breaks, prove the fix, and ship with evidence.
Start free with $50 in credit, no card needed. Works with Vapi, Retell, LiveKit, Pipecat + your stack.
One loop. Always improving.
Catch it in production, prove the fix in simulation, ship with evidence, then Roark watches the next call. This is how an agent self-improves: a loop, with you in it.
01 · Catch
Roark scores every live call and files what breaks.
Live scoring every call 1,284 today
Caller: I need my metoprolol refilled.
Agent: Sure, refilling your “met-a-pearl”.
Mispronunciation: Drug name mispronounced, caught by the audio model.
02 · Simulate
Your fix, replayed against realistic simulated callers.
Testing your candidates 161 / 240
- verify_identity toolpass
prompt v4 · verify before refund winner · 96
model · gpt-4.1 pass
03 · Review
Every change explicit and diffed, ready to apply.
Your fix, diffed
support_v3 v4
Prompt Model Tool
Swapped reasoning model gpt-4o gpt-4.1
04 · Verify
You ship, and Roark confirms the metric moved on live calls.
Verifying support_v4 in production
fix confirmed
Pronunciation, since your deploy 71→94
Issue recurrence none in 1,000 calls
Quality score 94 ↑
Regressions on other metrics none
… and the loop runs again on the next call.
Customers
Teams ship faster when they can hear what breaks.
Production voice AI scored on every conversation: pronunciation, empathy and resolution across their support and sales calls.
Healthcare calls evaluated for disclosures and identity checks: compliance scoring on every conversation, automatically.
Client voice agents validated in simulation before they go live: evidence that a build is ready, not a hunch.
02 · Simulate before launch
Break it in staging, not in production.
Run your agent against hundreds of simulated callers (realistic personas, accents, background noise and edge cases) and get every conversation scored before a customer ever dials in.
Scenarios & personas
Hundreds of simulated callers (the angry one, the rambler, the interrupter) built from your real call types.
Red teaming
Adversarial callers that try to break it (prompt injection, jailbreaks, social engineering) so your agent holds policy under attack.
45 languages & accents
Native accents, code-switching and background noise, in every market your agent answers.
Load & health tests
Peak-volume concurrency and always-on health checks, so the agent that passed in staging survives launch day.
Regression testing
Rerun the whole suite on every change and diff it against your last green baseline, so fixing one caller never breaks another.
Run it in CI
Every prompt or model change runs the suite before it merges: quality gates for conversations, not just code.
Pre-launch suite · booking_v2 182 / 200 passed
Angry caller · refund demand pass · 92
Red team · prompt injection pass · 90
Gulf Arabic · lobby noise pass · 88
Interrupts mid-disclosure fail · 61
Rambler · 3 intents in one turn pass · 85
Peak load · 250 concurrent pass
1 failure filed as an issue: fix it before launch, not after
03 · Post-call analysis
500+ metrics. Your models, not just an LLM.
Every production call scored as it lands: issues filed, alerts fired, dashboards and OTEL traces on tap, for voice calls and chat threads alike. And where most tools grade a transcript with an LLM, Roark runs purpose-built audio models on the call itself, measuring what your customer actually heard.
Metrics
- Pronunciation: 94
- Empathy: 84
- Instruction following: 97
- Response time: 88
- Disclosures: ✓
- Identity check: ✓
- + 60 more scored
Audio-native models
- Pronunciation
- Accent clarity
- Emotion
- Vocal stress
- Pace & pauses
- Interruptions
Conversational LLM + rules
- Resolution
- Empathy
- Task success
- Hallucination
- Repetition
- Tone
Compliance policy
- Disclosures
- PII exposure
- Identity check
- Script adherence
Performance latency
- Time-to-first-word
- Turn latency
- ASR WER
- Barge-in handling
500+ metrics out of the box
04 · The whole platform
Everything, before launch and after.
Simulation testing before you ship. Post-call analysis once you are live. Self-improvement connecting the two. Every capability first class, one click deep.
Before launch
- Simulation testing
- Scenarios & personas: Hundreds of simulated callers built from your real call types
- Red teaming: Adversarial callers probing for jailbreaks, prompt injection and off-policy answers
- Multilingual testing: 45 languages with native accents, code-switching and noise.
- Load testing: Hundreds of concurrent callers before launch day does it for you.
- Health checks: Always-on probes that catch an outage before a customer does.
- Regression testing: Every change diffed against your last green run so nothing quietly breaks.
- CI/CD gates: Every prompt or model change runs the suite before it merges.
In production
- Post-call analysis
- 500+ metrics: Audio-native, conversational, compliance and latency, on every call.
- Pronunciation testing: Drug names, SKUs and people, verified from the waveform.
- Custom metrics: Your rubric, scored on every call automatically.
- Human review & ground truth: Label real calls, set the correct answer, and measure how much each metric agrees.
- Issue tracker: Failures filed and clustered, tracked across deploys.
- Alerts & monitors: Thresholds on any metric, pinged to Slack or webhook.
- Tracing & dashboards: OTEL spans per turn and score trends by agent and version.
Always improving
- Self-improvement
- Prompt optimizer: Evidence-grounded prompt edits, drafted from the failing calls.
- Ground-truth tuning: Tune a metric from your labels until the model agrees with your team.
- Model, voice & infra fixes: Suggested swaps and settings when wording was never the problem.
- Fix verification: Replay your fix against the exact callers that broke it.
- Live verification: The metric watched from the first call after your deploy.
05 · Get started
First call scored in under a minute.
One click on any platform below and production calls stream in on their own, or send any recording with a few lines of code.
import Roark from '@roarkanalytics/sdk'
const client = new Roark({ bearerToken })
await client.call.create({
recordingUrl, startedAt,
interfaceType: 'PHONE',
callDirection: 'INBOUND',
agent: { customId: 'support_v2' },
}) // scored in seconds
Works with
-
-
-
-
-
-
- -
Industries
Built for the calls you actually take.
The same audio-native scoring, tuned to the failures, scripts and stakes of your industry.