Testing voicemail detection in outbound voice agents (and iOS 26 Call Screening) - Roark

Answering-machine detection is the piece of your outbound voice agent nobody thinks about until it burns a week of dial minutes leaving half-messages to humans. It sits before the first LLM token, decides which of two disjoint scripts to run, and rides on top of an audio classifier that has to be right in the first three seconds of the call. Get it wrong and every downstream metric, connect rate, cost per pickup, callback rate, is quietly poisoned.

The wrinkle is that "who picked up" stopped being a binary a while ago. A production outbound voice agent in 2026 has to distinguish humans, voicemail greetings, IVR menus, hold music, telecom messages ("your call is being connected"), and now a growing share of carrier-side and OS-level screeners. A LiveKit Agents user running thousands of answered calls per day , a share that has been climbing as more callers turn screening on. If your AMD test plan is still "does it detect voicemail: y/n", you are testing 2019.

Why AMD keeps regressing in production

AMD failures look benign on a dashboard and expensive on the bill. The failure modes:

Each of these produces its own downstream damage. And each is invisible unless you specifically test for it.

How an outbound voice agent classifies the call before speaking

What "correct" behavior actually looks like

Before you can test AMD, you have to write down the correct behavior for every pickup class, not just human vs. machine. This is the table I use as the acceptance spec on outbound programs:

Pickup class Correct agent behavior Common bug
Human, immediate hello Start the conversation within your latency budget Silence while sync AMD finishes
Human, slow "helloooo?" (elderly, noisy line) Classify as human, still start conversation Classified as voicemail, agent hangs up
Standard voicemail greeting Wait for the beep, leave pre-generated message with callback Talks over greeting, message truncated
Custom voicemail (long, music, kids talking) Detect voicemail confidently, still wait for beep Detects late, leaves partial message
IVR / auto-attendant ("press 1 for sales") Do not leave a message; either navigate or hang up cleanly Leaves voicemail into a menu
iOS 26 / Pixel Call Screening Briefly state purpose, honor screener timeout, hang up Full pitch delivered to a bot
Telecom hold ("your call is being connected") Wait, do not speak, retry classification Starts pitch to a hold announcement
No answer, ring-out Hang up, log as no-answer for retry cadence Agent stays on the line until max duration
Fax tone Immediate hang up Long silence until timeout

If your test plan does not cover every row, your production is running on hope. And note that "leave voicemail" is not a single behavior. Twilio's DetectMessageEnd mode is specifically for waiting until the end of the greeting to trigger the callback, because otherwise your message starts mid-greeting. That is a real config knob you have to test both settings of.

The AMD stack you are actually testing

Modern voice-AI platforms have moved past classic DSP-based AMD, but the moving parts multiplied rather than shrunk. On Pipecat, the VoicemailDetector runs a parallel pipeline with a classifier LLM whose only job is to output CONVERSATION or VOICEMAIL, gating TTS output until the decision lands. On Retell, voicemail and IVR detection is a first-class agent setting that runs continuously within a timeout window and the docs claim under 30 ms of added latency. On Vapi, the current recommended path is LLM-based detection via function calling, with the Twilio AMD path marked legacy. LiveKit does not ship a built-in detector, which is why the community request to add one explicitly cites the iOS 26 screening problem.

The consequence: your AMD is now typically a small LLM classification running in parallel with your main conversation LLM, holding TTS in a gate, with a timeout that decides how long you wait before defaulting one way or the other. That is five knobs, not one, and every knob deserves a regression test.

Building a real AMD test plan

Testing AMD by dialing your own cell phone and letting it go to voicemail is not a test plan. It is a smoke check. A real plan looks like this.

1. Enumerate personas, not scenarios

Voicemail detection is a caller-side property. Every test case is defined by "what the called party sounds like in the first ten seconds". You need personas for:

Each of these is a persona in the simulation harness, and each gets an expected agent behavior from the acceptance table above.

2. Test over real phone calls, not text transcripts

AMD is an audio problem. If you evaluate the classifier by feeding it strings, you will not catch the classes of failure that actually break production: greeting cadence, background noise, codec artifacts on the PSTN leg, TTS-generated voicemail greetings that sound "too clean". You want the test caller to dial your agent over real telephony (or WebRTC where your agent expects it), play the persona's audio, and let your production stack ingest exactly what it will ingest in production.

This is the reason we built Roark to dial voice agents over PSTN and WebRTC rather than mock the audio. Simulations built from personas, scenarios, and run plans let you cover the whole AMD matrix, in 45 languages and accents, without babysitting a phone. Every persona can carry its own voice, accent, pace, and background noise environment.

3. Score more than "did it detect"

The pass/fail bar for an AMD scenario is a compound assertion:

Every one of these is a metric, and each has its own threshold. This is where audio-native scoring matters: the difference between "leaves a voicemail" and "leaves a voicemail whose first three seconds are cut off" is invisible in the transcript and obvious on the waveform. Roark scores every call on audio-native metrics like pace, pauses, and pronunciation, not just transcript checks, which is what catches "the callback phone number came out as five-five-five, one-two-three, four-five-six-seven" versus mumbled.

Illustrative AMD regression on a live outbound call

4. Wire the tests into CI, not release day

AMD regresses on every change to the voicemail classifier prompt, every LLM version bump, every TTS provider swap, every silence-threshold tune. The rule I use: if it takes a code change or a config change to modify AMD behavior, that change ships behind a green suite.

Roark's simulations can be triggered over HTTP so a CI job kicks off the AMD suite on every PR, and inbound and outbound are both supported. The important part is that "AMD works" becomes a gate, not a vibe check.

5. Replay production failures as regression tests

The AMD failures worth writing tests for are the ones you have already seen. Every time a call in production ends up in the wrong branch, capture the recording and turn it into a fixture. Roark supports exactly this loop: production call replay lets you capture real calls and replay them against updated agent logic, turning yesterday's false positive into a permanent regression test. Do this for a month and your AMD suite starts to reflect your real caller distribution, not a synthetic one.

Illustrative pre-launch AMD simulation suite result

What to score on live traffic

Simulation catches the failures you thought to design for. Live scoring catches the ones you did not. On production outbound traffic you want continuous scores on:

Every one of these is something Roark scores automatically once calls are ingested, with issues filed the moment a call fails a threshold. That is the difference between finding AMD regressions on Monday and finding them in the QBR three weeks later.

The one paragraph version

Outbound voice agents live and die on the first three seconds of the called-party audio. If your test coverage for that window is "we tried a couple of voicemails once", the current wave of Call Screening rollouts, custom greetings, and multilingual voicemail systems is going to eat your connect rate quietly and expensively. Write the acceptance table. Turn each row into a persona. Dial them over real telephony. Score the audio, not the transcript. Replay every production miss into the suite. Gate the branch on green. Then, and only then, are you actually testing AMD.