Human review & ground truth for voice AI metrics - Roark

Ground truth for every metric.

Your team reviews real calls, sets the correct answer, and Roark measures how closely each metric agrees. Then it tunes the metric to match your judgment. Evals you can defend, because the people who know the work calibrated them.

Start free with $50 in credit, no card needed. Works with Vapi, Retell, LiveKit, Pipecat + your stack.

Metrics your experts can vouch for.

An LLM judge is a guess until a human checks it. Roark turns your team’s review into ground truth, scores every metric against it, and tunes until the model and your experts agree.

Your team sets the ground truth

Reviewers open a call with its transcript and audio and mark the correct value for every metric. For segment and turn metrics they can anchor the label to the exact moment. Assign several reviewers and Roark rolls their answers into one settled truth, sending real disagreements to adjudication.

Review · refund_call_04283 of 3 labeled

Metric Model You
Identity verified Pass Fail
Empathy 3 / 5 4 / 5
Resolution Resolved Resolved

transcript + audio in view · two reviewers · disputes settled inline

See how much each metric agrees

For every metric, Roark compares the model score to your team’s ground truth: an agreement rate, Cohen’s kappa, and the exact calls where they diverge. You learn which evals to trust and which need work, in numbers you can put in front of a customer.

Alignment · support_v2128 reviewed

14 disagreements queued · Identity verified is the one to fix

Tune the metric to match you

Confirmed labels become examples the metric learns from, the disagreements first. Re-score the reviewed set and watch alignment climb. Nothing changes in production until you publish the version that agrees with your team.

Tune · Identity verified draft v2

…and every metric gets more trustworthy with each review.

Everything a labeling workflow needs.

Review is a team sport. Roark handles the mechanics so your experts spend their time judging calls, not wrangling a spreadsheet.

Multiple reviewers

Assign a call to several reviewers. The first answer settles it; a second, differing answer opens a dispute.

Inline adjudication

Disagreements surface as Disputed and are resolved in place, so one answer is always the authoritative truth.

Per-metric rubric

Give reviewers the exact rubric for each metric, so labels stay consistent across your whole team.

Moment-level labels

Anchor a label to a quote in the transcript for segment and turn metrics, not just the call as a whole.

Inter-annotator agreement

Measure how much your reviewers agree with each other, and catch a rubric that needs sharpening.

Any metric, any type

Works on everything Roark scores: pass/fail, numeric scores, categories, audio-native and custom alike.

First call scored in under a minute.

One click on any platform below and production calls stream in on their own, or send any recording with a few lines of code.

import Roark from '@roarkanalytics/sdk'

const client = new Roark({ bearerToken })
await client.call.create({
  recordingUrl, startedAt,
  interfaceType: 'PHONE',
  callDirection: 'INBOUND',
  agent: { customId: 'support_v2' },
}) // scored in seconds

Node · Python, plus a REST API for CI/CD and webhooks the instant a call is scored.

Enterprise-grade from day one: annual pen tests, SSO/SAML, role-based access, configurable retention.

Bring a recording. We’ll score it live.

See your own agent measured on the audio it actually produced, in the demo, in real time. Stop guessing whether your voice AI works.