All posts

· 9 min read

Building Call AI: a real-time voice agent for banks with Gemini Live and Twilio

How I built a real-time voice agent for a loans and deposits desk with Gemini Live, FastAPI and Twilio Media Streams: the audio pipeline, barge-in, code-enforced authorization and a call lifecycle that never cuts off the goodbye.

case-studyvoice-aigemini-livefastapipythontwiliowebsockets
Building Call AI: a real-time voice agent for banks with Gemini Live and Twilio

Phone banking is a good fit for voice AI. The domain is narrow, there is structured data behind it, and most callers just want one number: the EMI, the balance, the due date. I built Call AI to see how far a single native-audio model could take a loans and deposits desk for a fictional lender, and to find out which parts of a voice agent are model problems and which are plain engineering problems.

Most of the interesting work turned out to be on the second list.

You can try the live demo or read the source on GitHub. The customer data is mock, and the lender is fictional.

What it does

Call AI answers a call, verifies the caller, and answers questions about their loans and deposits. It works in English and Hindi, and follows the caller if they switch mid-call. The same agent runs in two places: a browser demo and a real phone number through Twilio.

  1. Greeting. It introduces itself as an AI assistant for the lender and says the call may be recorded.
  2. Verification. It asks for the last 4 digits of the registered mobile number and the date of birth. Nothing else is discussed before that, not even whether an account exists.
  3. Answers. EMI amount, next due date, outstanding balance, overdue amount, savings balance, and fixed deposit rate and maturity. If the caller has several loans, it asks which one by type, never by ID.
  4. Escalation. A request for a human, a complaint, hardship, suspected fraud, anything out of scope, or three failed attempts to understand the caller ends in a handoff. More on what "handoff" means below.
  5. Goodbye. The agent ends the call itself once the caller is done.

Architecture at a glance

Browser mic  ──► WebSocket (browser) ──┐
                                        ├──► CallSession ◄──► Gemini Live
Phone ──► Twilio ──► WebSocket (media) ─┘     (one session per call)
                                                   │
                       tools: verify_customer, check_loan_emi,
                       check_account_balance, escalate_to_agent, end_call

The backend is FastAPI. Every WebSocket creates a CallSession that opens one Gemini Live connection and runs two asyncio tasks: one pumps audio from the caller into the model, the other reads everything the model sends back (audio, transcripts, tool calls, interruptions). When either task finishes, the other is cancelled. Calls share no mutable state, so each one is independent.

The frontend is Next.js with an AudioWorklet for microphone capture, plus an architecture page that explains the system to anyone who opens the demo.

One agent, two transports

The browser and Twilio disagree about almost everything: framing, codec, sample rate, how you stop playback. So the agent logic never touches either. It talks to a small Transport interface with five operations: receive audio, send audio, clear playback, send an event, and finish.

The browser transport sends binary PCM frames and JSON events. The Twilio transport speaks Twilio's JSON protocol and converts audio on the way in and out. Everything above that line, the prompt, the tools and the call lifecycle, is shared. It also made testing much cheaper, because tests can plug in a fake transport instead of a phone.

The audio pipeline

Voice projects tend to go wrong in the audio plumbing rather than in the model. The formats here:

  • Gemini Live takes 16 kHz PCM16 mono in and returns 24 kHz PCM16 out.
  • Browser, upstream: the page creates its AudioContext at 16 kHz so the browser does the resampling. An AudioWorklet converts Float32 samples to Int16 and posts 40 ms chunks (1,280 bytes) over the WebSocket. Echo cancellation, noise suppression and auto gain are requested from getUserMedia.
  • Browser, downstream: each 24 kHz chunk becomes an AudioBuffer scheduled to start exactly when the previous one ends, which keeps playback gapless without a hand-rolled jitter buffer.
  • Twilio, upstream: 8 kHz μ-law frames of 20 ms (160 bytes) are decoded to linear PCM and resampled from 8 kHz to 16 kHz.
  • Twilio, downstream: 24 kHz to 8 kHz, encoded to μ-law, base64-encoded and sent as a media event for the right stream.

The resampler keeps its state between frames. Resampling each 20 ms frame in isolation starts every chunk from zero, which leaves audible clicks at the frame boundaries. All of this runs on the standard library's audioop, so there is no NumPy or SciPy in the audio path. The catch is that audioop was removed in Python 3.13, which is why the backend is pinned to Python 3.12.

Barge-in

People interrupt voice agents constantly. Gemini detects the interruption on the server and flags it, so detection was never the hard part. The hard part is the audio that is already in flight.

In the browser, that means stopping every scheduled audio source and resetting the playback clock. On Twilio it means sending a clear event, because Twilio buffers audio on its side and will keep playing the old sentence after the caller has talked over it. A test covers that barge-in clears playback.

Turn-taking itself is left to the model's default voice activity detection. That is a trade-off: it works well, but I have no say over the thresholds.

Authorization lives in code, not in the prompt

The prompt tells the agent to verify the caller before anything else. If that were the only protection, a good enough social-engineering attempt could talk the model out of it. So the rule is enforced in code as well:

  • The data tools never accept a customer ID from the model. They read it from a per-call context that only verify_customer can set.
  • Every data tool returns not_verified when that context is empty.
  • Verification matches on the last 4 digits of the mobile number plus the date of birth. A failure returns a generic "the details did not match", never which field was wrong.
  • After 3 failed attempts the tool returns a locked state and tells the model to escalate instead of retrying.

Even if a caller convinces the model to "just look up Priya's loan", there is no way to ask a tool for a different customer. The model can say anything; it can only do what the tools allow.

Tools designed for a model, not a human

There are five tools: verify_customer, check_loan_emi, check_account_balance, escalate_to_agent and end_call. A few choices made the conversations much better:

  • Results are written as instructions. When a caller has two loans and asks about "my EMI", the tool returns needs_clarification with the available loan types. The model's next sentence is then "Which one, the home loan or the personal loan?" instead of a guess. When nothing matches, the tool returns the types that do exist.
  • Dates are computed, not stored. The next EMI date comes from today's date and the loan's EMI day, including year rollover and short months. The mock data never goes stale, and today's date is injected into the system prompt.
  • Grounding is mandatory. The prompt requires a tool call before the agent states any number, date, balance or status.
  • Tool failures are survivable. An exception inside a tool becomes an error result the model can talk about, not a crashed call.

Ending a call is harder than starting one

Hanging up sounds trivial until the goodbye gets cut off mid-word. The agent can request an ending through end_call or escalate_to_agent, but the server owns the lifecycle:

  1. After the tool response goes back to the model, the session enters an ending state and emits an escalated event once, if relevant.
  2. It waits for the model to actually speak the goodbye, meaning at least one audio part arrives after the ending started, and then for the turn to complete.
  3. A 20-second watchdog force-finishes the call if the model never produces a goodbye.
  4. In the browser, the page waits for the queued audio to finish playing, plus 800 ms, before tearing the connection down.
  5. On Twilio, the server sends a named mark and waits up to 10 seconds for Twilio to echo it back. That echo means the caller has heard everything.

Finishing is idempotent, so a race between the watchdog and the normal path cannot close the call twice.

The "I'm connecting you now" bug

In early testing the agent told callers "I'm transferring you to a human agent now". There was no transfer to make, because the system has no way to connect a call to a person.

I fixed it in two places. Escalation is now an honest callback: the agent says a human will call back, then ends the call. And the prompt has a hard rule never to say or imply that something was done unless the matching tool has just returned success.

Guardrails in the prompt

The prompt covers the things a bank would care about:

  • It says it is an AI if asked.
  • Short replies: one or two sentences, one question at a time, no lists or markdown, since all of it is spoken.
  • It will not accept or repeat OTPs, PINs, CVVs, card numbers or full account numbers.
  • It will not advise, recommend products, negotiate rates, promise waivers or take payments.
  • For an overdue EMI it stays factual and kind, and never mentions legal or credit consequences.
  • It refuses to reveal or paraphrase its own instructions, and ignores "ignore previous instructions" style attempts.
  • Amounts are spoken in the Indian style (lakh), and dates naturally.

The prompt also includes around ten short example exchanges, so the model hears what a good spoken reply sounds like and not only the rules.

Latency

For a phone call, the number that matters is the gap between the caller finishing a sentence and the first audio byte coming back. I measured it on a development machine, with two runs per model, so treat these as indicative and not as a benchmark:

  • An earlier native-audio model: roughly 3 to 4 seconds.
  • Two newer Live models: about 1.35 seconds and about 1.1 seconds.

The prompt asks for one short filler such as "One moment, let me check that" before a tool call, so a lookup does not sound like dead air. The code does not enforce it, so it is a prompt behaviour, not a guarantee.

Testing without a phone

The tests use fake transports and fake Gemini sessions, which means they run without a network connection or a phone. They check behaviours that matter on a real call:

  • The call ends only after the goodbye audio has played.
  • An escalation notifies exactly once, even if the model calls the tool twice.
  • Barge-in clears playback.
  • Twilio frames come out at the right length after resampling.
  • A customer can never read another customer's records.
  • Next-due dates roll over correctly across month and year ends.

What these tests cannot tell me is whether the model behaves. Automated conversation evals are the next thing I want to add.

What I would change next

This is a demo, and it has the gaps of one. The main ones:

  • Reconnects. If the Gemini session drops, the call ends. The server's go-away signal is only logged, and session resumption is not wired up.
  • Call limits. There is no maximum call duration or silence timeout.
  • PII in logs. Transcripts and tool results are printed to stdout, including balances. Real use needs structured logging with redaction.
  • Webhook security. Twilio signature validation is not implemented yet, and the browser WebSocket has no authentication or rate limiting.
  • Lockout. The three-attempt limit resets when the caller hangs up and calls back. It should be tracked per phone number.
  • Escalation. On the phone path it currently logs the reason. A real version would create a ticket or schedule the callback.
  • Deployment. A containerised setup behind a reverse proxy with TLS is on the roadmap.

The stack

  • Backend: Python 3.12, FastAPI, uvicorn, the Google GenAI SDK (Gemini Live), pydantic-settings, uv.
  • Telephony: Twilio Media Streams, with ngrok for local testing.
  • Frontend: Next.js, React, TypeScript, Web Audio API with an AudioWorklet.
  • Tests: pytest with fake transports and fake model sessions.

What I took from it

  • Put the rules that must hold in code. The prompt can ask for good behaviour, but only code can guarantee it.
  • Audio plumbing and the call lifecycle took more time than the AI part: sample rates, buffers, flushing, and ending a call without clipping the goodbye.
  • Tool results are prompts too. How a result is worded shapes the model's next sentence as much as the system prompt does.

If you want to see it work, the demo is live, and the code is on GitHub.

Found this useful? Share it.

© 2025 bhagat.dev