The Zeroic Journal

AI voice agent architecture: the four parts that decide if it ships

Where each layer of a production voice agent breaks on live phone calls, and the order we'd build them in, from voice agents we run in production.

On this page 7 sections
  1. Cascaded pipeline vs speech-to-speech: start cascaded
  2. Speech-to-text fails on names, numbers, and languages
  3. Turn detection decides whether the agent feels human
  4. Text-to-speech is where latency hides
  5. Telephony is the layer demos skip
  6. Autonomy comes from integrations, not the model
  7. Build order: from demo to production

A production AI voice agent architecture isn’t one model. It’s speech-to-text, turn detection, text-to-speech and telephony, held together by an orchestration layer and wired into the systems it acts on. Whether it ships depends on how each of those fails on live calls, and on the integrations: they’re what turn a good conversation into finished work.

That matters right now because the demo has become trivial. Open-source frameworks like LiveKit Agents and Pipecat will get you a talking agent on a phone number in an afternoon. Getting from that afternoon to an agent that answers an estate agency’s phones at 9pm on a Sunday, in the caller’s first language, while writing the viewing request into the right CRM record, is the part nobody demos. That’s the part we spend our time on. This is how the stack breaks down.

Cascaded pipeline vs speech-to-speech: start cascaded

There are two ways to build a voice agent. A cascaded pipeline transcribes the caller (STT), sends text to an LLM, and synthesizes the reply (TTS). A speech-to-speech model takes audio in and puts audio out, with no text in the middle.

Speech-to-speech sounds more natural and skips two hops. It also hides everything. When a caller says the agent booked the wrong appointment, a cascaded pipeline gives you the exact transcript, the exact LLM input, the exact tool call. A speech-to-speech model gives you a recording and, at best, a side transcript of audio the model reasoned over directly.

For an agent that takes actions - books, cancels, writes to a CRM, quotes a policy - you need that text layer. Our rule:

  • Cascaded wins when the agent calls tools, when you need call transcripts for compliance or QA, and when you need to swap one component (a better STT for Arabic, a cheaper LLM for routing) without rebuilding the rest.
  • Speech-to-speech wins for open-ended conversation where tone matters more than actions - companionship, coaching, language practice.
  • The hybrid is where the field is heading. OpenAI’s own realtime agents reference repo uses a realtime model to handle greetings and information-gathering while a text model supervises “tool calls and more challenging responses”. The same repo notes that stitched pipelines often land at response latency of “1.5s or longer” - which is why the cascade needs engineering. Chaining three API calls won’t get you there.

Pick cascaded as the backbone. Move to the hybrid once your eval set shows where naturalness is costing you calls.

Speech-to-text fails on names, numbers, and languages

STT is the first place a real call diverges from the demo. Demo callers speak clearly into a laptop mic. Real callers are on a bus, on speakerphone, spelling a street name, reading a postcode, or switching between English and their first language mid-sentence.

Three failure surfaces matter:

  • Proper nouns and alphanumerics. Names, street names, postcodes, booking references, email addresses. Generic STT models are weakest exactly where your CRM lookup needs precision. Feed the recognizer a keyword list per agency (their branch names, their streets, their listing references) wherever the provider supports biasing, and confirm critical values back to the caller in a structured way - “that’s S-W-1-A, 1-A-A?”.
  • Language variance. OpenAI’s Whisper README is blunt about it: “Whisper’s performance varies widely depending on the language,” and it publishes per-language error rates to prove it. The same is true of every provider. A single English benchmark tells you nothing about Punjabi or Romanian.
  • Streaming vs batch. Whisper’s reference implementation processes audio in a 30-second sliding window. That’s fine for transcribing a finished call and wrong for a live one. Production voice needs a streaming recognizer that emits partial transcripts as the caller speaks, so the LLM can start working before the sentence ends.

There’s no single “best” STT vendor. What works is a per-language evaluation built from your own call recordings, and the freedom (cascaded again) to route different languages to different recognizers.

Turn detection decides whether the agent feels human

If you only get one layer right, get this one right. Turn detection answers a deceptively simple question: has the caller finished talking?

The naive answer is a silence threshold. Wait for N milliseconds of quiet, then respond. It fails in both directions. Set it short and the agent interrupts every caller who pauses to remember their budget. Set it long and every reply arrives after an awkward beat of dead air. Callers read both as “this is a robot” and start talking over it, which makes everything worse.

Production stacks now split the job in two:

  1. Voice activity detection (VAD) answers “is someone speaking right now?” It has to be fast and cheap because it runs on every audio frame. Silero VAD, which ships as a core dependency in Pipecat and is the VAD plugin LiveKit recommends, processes a 30ms+ chunk in under 1ms on a single CPU thread, from a model around two megabytes, trained on corpora covering over 6,000 languages.
  2. End-of-turn detection answers “are they done?” This needs meaning as well as energy. Smart Turn is an open, audio-native model built for exactly this: it reads raw audio rather than a transcript so it can use “grammar, tone and pace of speech”, supports 23 languages, ships as an 8MB quantized CPU model, and runs in under 100ms on most cloud instances. LiveKit Agents has moved the same way: its current turn detector is an audio model that combines what is said with acoustic cues like intonation, pitch and rhythm, and its older text-based detector is deprecated.

Then there’s barge-in - the caller interrupting the agent. When VAD fires while the agent is speaking, the orchestrator has to stop TTS playback immediately, flush any queued audio, cancel the in-flight LLM generation, and record what the caller actually heard, so the conversation history reflects reality rather than the full sentence the model wrote. Miss that last step and the agent will later reference information the caller never heard.

Two details that separate usable from polished:

  • Not every sound is an interruption. “Mm-hm” and “yeah” while the agent is talking are backchannels. The caller isn’t asking the agent to stop. Treat very short utterances during agent speech as acknowledgments unless they contain content.
  • Filler buys time honestly. When a tool call will take a second - a CRM lookup, a calendar check - say “let me check that slot” first. Silence during a lookup is the most common cause of callers asking “hello?”.

Text-to-speech is where latency hides

TTS rarely fails loudly. It fails quietly, by adding hundreds of milliseconds you don’t notice until you measure the whole round trip.

The rules we hold to:

  • Stream end to end. The LLM streams tokens; the orchestrator chunks them at sentence or clause boundaries; TTS starts synthesizing the first clause while the LLM is still writing the second. An agent that waits for the full LLM response before speaking will always feel slow, whatever models you pick.
  • Write for the ear. LLM output formatted for a screen - bullet points, markdown, “1)”, URLs, prices as “£1,250pcm” - sounds terrible read aloud. The system prompt has to ask for spoken-form output, and a normalization step should catch what slips through: currency, dates, postcodes, phone numbers.
  • One voice per language, tested per language. A voice that sounds natural in English can sound flat or mispronounce names in another language. Test it in the same per-language eval as STT.

Telephony is the layer demos skip

Browser demos run over WebRTC with clean, wideband audio. Phone calls arrive over the public telephone network via a SIP trunk or a provider’s media stream, usually as narrowband 8kHz audio (Twilio’s Media Streams, for example, always delivers 8kHz mu-law), with its own jitter and its own failure modes. Pipecat ships serializers for telephony providers such as Twilio, Telnyx, Plivo and Vonage, and LiveKit Agents plugs into LiveKit’s telephony stack for inbound and outbound calls. Both exist because a phone line is a different problem from a browser mic.

What the telephony layer has to own:

  • Codec and sample-rate handling. Evaluate your STT on phone-quality audio. Accuracy on 8kHz calls can be noticeably worse than the vendor’s headline number.
  • Warm transfer. When the agent escalates, it has to bridge the caller to a human with context - who they are, what they need, what they asked - so they don’t land in a queue and start over.
  • Out-of-hours behavior. The whole point of an AI receptionist is the 7pm and weekend calls. Outside business hours, “transfer to a human” becomes “log a callback with a structured summary in the CRM”, and that path needs the same testing as the happy path.
  • Recording and consent. Call recording, disclosure that the caller is speaking to an AI, and data retention are jurisdiction-specific. Design them in on day one; retrofitting them into a live phone system is miserable.

Autonomy comes from integrations, not the model

Here’s the part most posts on this miss. A voice agent that answers questions well but can’t act is a voicemail with better manners. Autonomy comes from the systems of record.

A caller asking “can I move my Thursday appointment to Saturday morning?” needs the agent to find the booking, check what’s free, check the right person’s calendar, rebook the slot, and write the lead back to the CRM. Every one of those is a tool call against a live system with its own rate limits, stale data, and the occasional timeout. Design for that:

  • Tools return structured errors, and the agent has a script for each. “The calendar didn’t respond” should become “I’ll have the team confirm Saturday by text within the hour” and never a hallucinated confirmation.
  • Every interaction writes back. If the call isn’t in the CRM, from the business’s point of view it didn’t happen.
  • Split agents by job. OpenAI’s reference repo recommends sequential handoffs between specialized agents because putting “all instructions and tools in a single agent” can degrade performance. A receptionist that routes, a negotiator that qualifies, and a property manager that logs maintenance each carry a smaller, sharper prompt and toolset.

Build order: from demo to production

If you’re taking a voice agent from demo to production, work in this order:

  1. Choose cascaded as the backbone. You need transcripts and tool calls you can inspect. Revisit speech-to-speech later for specific conversational flows.
  2. Build the eval set before you tune anything. Fifty recorded calls per language, each labeled with what should have happened. Every later decision gets checked against it.
  3. Fix turn detection next. Fast VAD plus a semantic or audio-native end-of-turn model, with barge-in that truncates history to what the caller heard.
  4. Stream everything and write for the ear. Clause-level TTS streaming, spoken-form output, normalized numbers and postcodes.
  5. Wire the systems of record. CRM, calendar, bookings. Structured tool errors with scripted recoveries.
  6. Design the escalation path. Warm transfer in hours, logged callback out of hours, and a clear rule for which topics always go to a human.
  7. Only then add languages. One at a time, each gated on its own eval results.

If you already have a voice agent live and the numbers aren’t moving, an AI Audit will rank which of these layers is costing you the most calls. If you’re starting from zero, the MVP Sprint is built for exactly this kind of product. Everything we offer is on the services page.

Frequently asked questions

Should a production voice agent use a speech-to-speech model or a cascaded STT-LLM-TTS pipeline?
For anything that books appointments, writes to a CRM, or needs an audit trail, start cascaded. Text at every handoff means you can log, replay, and evaluate each call. Speech-to-speech models win on naturalness and are worth adding for the conversational layer, but the pattern OpenAI's own reference repo uses is a realtime model for chit-chat with a text model supervising tool calls, rather than the realtime model alone.
What is the hardest part of a voice agent to get right?
Turn detection. Deciding when the caller has finished speaking is the single biggest driver of whether the agent feels rude, slow, or natural. Silence thresholds alone either cut people off mid-thought or add dead air after every sentence. Pair a fast VAD with a semantic or audio-native end-of-turn model.
How do you make a voice agent work in multiple languages?
Treat each language as its own eval target. Speech recognition accuracy varies widely by language - Whisper's own documentation says so and publishes per-language error rates. Test STT, the turn detector, and the TTS voice per language with real call recordings before you switch a language on, and keep a per-language fallback to a human or a callback.
What should happen when the AI voice agent doesn't know the answer?
It should say so, capture the details, and hand off - a warm transfer during business hours, a logged callback outside them. An agent that guesses on price, availability, or legal questions costs you more than one that escalates. Design the fallback path up front as a feature of the product.
Can an AI receptionist handle most inbound calls on its own?
Yes, when the domain is bounded and the agent is wired into the systems of record. How much it handles on its own comes from the integration work - CRM reads and writes, calendars, booking data - as much as from the model.

References

  1. LiveKit Agents - open-source framework for realtime voice agents (GitHub)
  2. Pipecat - open-source framework for voice and multimodal agents (GitHub)
  3. Smart Turn - open-source audio-native turn detection model (GitHub)
  4. Silero VAD - pre-trained voice activity detector (GitHub)
  5. OpenAI Whisper - README and per-language performance (GitHub)
  6. OpenAI Realtime Agents - chat-supervisor and handoff patterns (GitHub)
  7. LiveKit Docs - Turn detector
  8. Pipecat Docs - SileroVADAnalyzer
  9. Twilio Docs - Media Streams WebSocket messages

More from the journal

Hear from Prashant within 24 hours.

Prashant Abbi Partner @ Zeroic

Building a voice agent that has to survive real inbound calls? We build real-time voice pipelines and the integration layer that lets them finish the job, not just answer. An AI audit ranks which layer is costing you the most calls.

See AI audits, or book a 30-minute call