Of the five kinds of work we are taking to market, the Indic voice floor is the one where the case for a room is clearest. Three things line up there at once. The market already prices this work per interaction, so the work is measurable. The models that carry the work are small, so the whole stack fits in one pod. And the regulator has just made every recovery call a recorded, retained, inspectable object with a named custodian. When the bank's auditor asks where the recording is, the answer has to be an address, and a room in a named building is one.
The US pattern: priced per outcome, and the model is a sliver of the minute
The contact centre was the first place the American software market stopped selling seats. In December 2024 Sierra set out outcome-based pricing for its agents: the customer pays when a conversation is resolved and, in most cases, nothing when it is not. Genesys moved the same way through a token model: one token lists at a dollar in the United States, an automated call summary costs one fiftieth of a token, about two cents, and an agentic interaction consumes 1.2 tokens, about a dollar twenty. Nobody on either side counts language-model tokens; they count summaries, interactions and resolutions.
Underneath the price, the model is a small part of the cost of a voice minute. On 8 July 2026 Inworld published a worked cost model for a cascaded voice agent, speech to text, then a language model, then text to speech, at July 2026 list prices. Across the stacks it priced, a conversation minute runs from about $0.007 to $0.091. The cheapest, a realtime transcriber, a 26-billion-parameter open model and a realtime synthesiser, lands at $0.00686 a minute, and the split inside it is the part worth reading.
| Component | Cost a minute | Share |
|---|---|---|
| Speech to text | $0.001667 | 24% |
| Language model, 2,340 tokens in and 100 out | $0.000198 | 3% |
| Text to speech, 400 characters | $0.0050 | 73% |
| Whole minute | $0.00686 | 100% |
The language model is under 3% of the minute; transcription is about 24% and synthesis about 73%. Two consequences follow. A pod that hosts only the language model changes little of the minute's cost, because the model was a small part of it: the pod has to carry transcription, synthesis and the overnight work to change the cost structure of a minute at all. And because the token arithmetic is light, the constraint on the day shift is not how many tokens the machine can produce but how quickly it produces the first one.
What an Indic floor actually runs, by day and by night
A collections and service floor in India has two shifts, and the regulator drew the line between them. Since August 2022 a lender and its recovery agents may not call a borrower before 8 a.m. or after 7 p.m.; the directions of August 2026 restate the window as 08:00 to 19:00. The day is real-time collections and service calls in Hindi, Tamil, Telugu and the floor's other languages, each one a stream: audio in, a transcript forming as the caller speaks, a small model answering with the account's context in front of it, speech out. Every stream has a latency budget of a few hundred milliseconds, and the whole day is a fight to stay inside it.
At seven the floor goes quiet and the machine does not. The night's work is to transcribe every recording the day produced, automated and human, then summarise and score each against the rules the floor is held to: was the call inside the window, was the borrower told it was being recorded, was there anything a reviewer would call intimidation, what was the disposition. Floors used to sample this work; a machine can do it to every call.
| Day, 08:00 to 19:00 | Night, 19:00 to 08:00 | |
|---|---|---|
| Work | Live calls in Indic languages; assist on human calls | Transcribe, summarise and score every recording |
| Bound by | First-token latency; small batches | Throughput; large batches |
| What fails if it is wrong | The caller hears a pause | The morning's compliance report is late |
The two shifts are the same eight accelerators in two regimes: the day latency-bound, many short streams at small batch; the night throughput-bound, a queue at large batch. A room that runs only the day is a room half used, and our unit economics do not close on one: the modelled steady state is 60% of sellable capacity, and a mill that cannot run overnight cannot reach it.
Why the bank must be able to name the room
Three regimes touch the recordings, and only one is the data-protection statute.
The first is the outsourcing regime. The Reserve Bank's Directions on Outsourcing of Information Technology Services of April 2023 cover banks and NBFCs and name application development, data-centre services and cloud computing outright. An agreement under them must provide for storage of data only in India as applicable, give the regulated entity the right to audit the provider and its sub-contractors, and ensure the provider grants unrestricted access to the data and the relevant business premises. An auditor who may walk into the premises needs premises to walk into.
The second is older and narrower: since the April 2018 circular on storage of payment-system data, payment data has to sit in a system in India, and most compliance teams treat a collections call about a repayment as inside that line.
The third is the Digital Personal Data Protection Act. A voice recording is data about an identifiable individual, so it is personal data under section 2(t). But section 16 restricts transfer only to countries the Centre notifies, and as of mid-2026 none has been, so the Act alone does not keep a recording in the building. The sector rules do, and the August 2026 directions make them bite: from 1 January 2027 every recovery call must be recorded, the borrower told, and the recording kept for six months or until any litigation about it ends, whichever is later. A recording is now a regulated record, and a regulated record has a custodian and a location.
What ten kilowatts can and cannot do for real-time voice
What it can do. The models on a voice floor are small. An eight-billion-parameter model quantised weighs about 4.5 GB, and a mill carries 1,536 GB, so the whole stack, transcriber, language model, synthesiser and an index of the floor's own scripts, fits in one box with memory to spare. The night shift is what the machine was built for: a queue, large batches, eight accelerators flat out.
What it cannot do. Real-time voice is governed by the first token, not the last. A turn has to start speaking within a few hundred milliseconds, and staying inside that budget means small batches, which means far fewer tokens a second than a batch job would serve. The transcriber and the synthesiser take their share of the same silicon, and the synthesiser is the heavier. So the calls a pod holds at once are set by latency, a fraction of what batch throughput would imply. A frontier-class model per call in real time is not in the pod's range. And one pod is not redundant: a platform fronting a national floor must keep a second lane, the cloud it already uses, for the day the room is down.
We have not published a concurrent-call figure and will not until we have measured one on our own accelerators with transcription and synthesis resident. Every public voice benchmark we have read measures the model alone, and the model, as the Inworld arithmetic shows, is the easy part.
What we will measure before quoting
Before quoting any floor we will run its own recordings and scripts through a pod and measure five things: first-token and full-turn latency at rising concurrency with the whole stack resident; the concurrent calls the pod holds inside the latency budget, the day figure; calls transcribed and scored per hour at night, the night figure; word-error rate on the floor's own Indic audio; and the energy the day and the night work draw at the meter. Then the proposal is based on that workload and its deployment, with quality and total cost set beside what the floor runs today, before anyone decides.
That is why the voice floor needs a named room. The work is measurable, the stack fits, the night is real, and the regulator has just given every recording an owner who has to say where it is.