Skip to content
AI for businessFor building ownersHow it worksPricingResearchCompany
Let’s talk
AI for businessFor building ownersHow it worksPricingResearchCompanyLet’s talk
On this pageSources

Research9 September 2026

Why the voice floor needs a named room

Voice AI is priced per resolution, the model is a sliver of a minute, and from 2027 every recovery call is recorded. Why a bank must name the room.

6 min read
  • voice
  • BFSI
  • use cases

Of the five kinds of work we are taking to market, the Indic voice floor is the one where the case for a room is clearest. Three things line up there at once. The market already prices this work per interaction, so the work is measurable. The models that carry the work are small, so the whole stack fits in one pod. And the regulator has just made every recovery call a recorded, retained, inspectable object with a named custodian. When the bank's auditor asks where the recording is, the answer has to be an address, and a room in a named building is one.

The US pattern: priced per outcome, and the model is a sliver of the minute

The contact centre was the first place the American software market stopped selling seats. In December 2024 Sierra set out outcome-based pricing for its agents: the customer pays when a conversation is resolved and, in most cases, nothing when it is not. Genesys moved the same way through a token model: one token lists at a dollar in the United States, an automated call summary costs one fiftieth of a token, about two cents, and an agentic interaction consumes 1.2 tokens, about a dollar twenty. Nobody on either side counts language-model tokens; they count summaries, interactions and resolutions.

Underneath the price, the model is a small part of the cost of a voice minute. On 8 July 2026 Inworld published a worked cost model for a cascaded voice agent, speech to text, then a language model, then text to speech, at July 2026 list prices. Across the stacks it priced, a conversation minute runs from about $0.007 to $0.091. The cheapest, a realtime transcriber, a 26-billion-parameter open model and a realtime synthesiser, lands at $0.00686 a minute, and the split inside it is the part worth reading.

One voice minute on the cheapest cascaded stack, July 2026 list prices, from the Inworld worked model
ComponentCost a minuteShare
Speech to text$0.00166724%
Language model, 2,340 tokens in and 100 out$0.0001983%
Text to speech, 400 characters$0.005073%
Whole minute$0.00686100%
One voice minute on the cheapest cascaded stack, July 2026 list prices, from the Inworld worked model

The language model is under 3% of the minute; transcription is about 24% and synthesis about 73%. Two consequences follow. A pod that hosts only the language model changes little of the minute's cost, because the model was a small part of it: the pod has to carry transcription, synthesis and the overnight work to change the cost structure of a minute at all. And because the token arithmetic is light, the constraint on the day shift is not how many tokens the machine can produce but how quickly it produces the first one.

What an Indic floor actually runs, by day and by night

A collections and service floor in India has two shifts, and the regulator drew the line between them. Since August 2022 a lender and its recovery agents may not call a borrower before 8 a.m. or after 7 p.m.; the directions of August 2026 restate the window as 08:00 to 19:00. The day is real-time collections and service calls in Hindi, Tamil, Telugu and the floor's other languages, each one a stream: audio in, a transcript forming as the caller speaks, a small model answering with the account's context in front of it, speech out. Every stream has a latency budget of a few hundred milliseconds, and the whole day is a fight to stay inside it.

At seven the floor goes quiet and the machine does not. The night's work is to transcribe every recording the day produced, automated and human, then summarise and score each against the rules the floor is held to: was the call inside the window, was the borrower told it was being recorded, was there anything a reviewer would call intimidation, what was the disposition. Floors used to sample this work; a machine can do it to every call.

The two shifts of a collections and service floor
Day, 08:00 to 19:00Night, 19:00 to 08:00
WorkLive calls in Indic languages; assist on human callsTranscribe, summarise and score every recording
Bound byFirst-token latency; small batchesThroughput; large batches
What fails if it is wrongThe caller hears a pauseThe morning's compliance report is late
The two shifts of a collections and service floor

The two shifts are the same eight accelerators in two regimes: the day latency-bound, many short streams at small batch; the night throughput-bound, a queue at large batch. A room that runs only the day is a room half used, and our unit economics do not close on one: the modelled steady state is 60% of sellable capacity, and a mill that cannot run overnight cannot reach it.

Why the bank must be able to name the room

Three regimes touch the recordings, and only one is the data-protection statute.

The first is the outsourcing regime. The Reserve Bank's Directions on Outsourcing of Information Technology Services of April 2023 cover banks and NBFCs and name application development, data-centre services and cloud computing outright. An agreement under them must provide for storage of data only in India as applicable, give the regulated entity the right to audit the provider and its sub-contractors, and ensure the provider grants unrestricted access to the data and the relevant business premises. An auditor who may walk into the premises needs premises to walk into.

The second is older and narrower: since the April 2018 circular on storage of payment-system data, payment data has to sit in a system in India, and most compliance teams treat a collections call about a repayment as inside that line.

The third is the Digital Personal Data Protection Act. A voice recording is data about an identifiable individual, so it is personal data under section 2(t). But section 16 restricts transfer only to countries the Centre notifies, and as of mid-2026 none has been, so the Act alone does not keep a recording in the building. The sector rules do, and the August 2026 directions make them bite: from 1 January 2027 every recovery call must be recorded, the borrower told, and the recording kept for six months or until any litigation about it ends, whichever is later. A recording is now a regulated record, and a regulated record has a custodian and a location.

What the room does and does not settle

We are not the buyer's lawyers; its compliance team decides what the register says. A hyperscaler region in India satisfies residency. The outsourcing directions ask for more: a named provider, an audit right and access to premises. The room is how the register gets an address, not a substitute for the buyer's policy.

What ten kilowatts can and cannot do for real-time voice

What it can do. The models on a voice floor are small. An eight-billion-parameter model quantised weighs about 4.5 GB, and a mill carries 1,536 GB, so the whole stack, transcriber, language model, synthesiser and an index of the floor's own scripts, fits in one box with memory to spare. The night shift is what the machine was built for: a queue, large batches, eight accelerators flat out.

What it cannot do. Real-time voice is governed by the first token, not the last. A turn has to start speaking within a few hundred milliseconds, and staying inside that budget means small batches, which means far fewer tokens a second than a batch job would serve. The transcriber and the synthesiser take their share of the same silicon, and the synthesiser is the heavier. So the calls a pod holds at once are set by latency, a fraction of what batch throughput would imply. A frontier-class model per call in real time is not in the pod's range. And one pod is not redundant: a platform fronting a national floor must keep a second lane, the cloud it already uses, for the day the room is down.

We have not published a concurrent-call figure and will not until we have measured one on our own accelerators with transcription and synthesis resident. Every public voice benchmark we have read measures the model alone, and the model, as the Inworld arithmetic shows, is the easy part.

What we will measure before quoting

Before quoting any floor we will run its own recordings and scripts through a pod and measure five things: first-token and full-turn latency at rising concurrency with the whole stack resident; the concurrent calls the pod holds inside the latency budget, the day figure; calls transcribed and scored per hour at night, the night figure; word-error rate on the floor's own Indic audio; and the energy the day and the night work draw at the meter. Then the proposal is based on that workload and its deployment, with quality and total cost set beside what the floor runs today, before anyone decides.

That is why the voice floor needs a named room. The work is measurable, the stack fits, the night is real, and the regulator has just given every recording an owner who has to say where it is.

Sources

  1. 01Outcome-based pricing for AI agentsSierra, 10 December 2024You pay only when the software achieves specific, valuable outcomes; unresolved conversations in most cases carry no charge.
  2. 02Genesys Cloud tokens modelGenesys Cloud Resource Center, 12 July 20261.2 tokens per Agentic Virtual Agent interaction; 50 AI summaries per token.
  3. 03Voice Agent Cost Per Minute 2026: a worked cost modelInworld AI, 8 July 2026Cascaded voice stacks cost $0.007 to $0.091 a minute; in the cheapest stack the language model is under 3 per cent of the minute, speech recognition and speech synthesis the rest.
  4. 04Outsourcing of Financial Services: responsibilities of regulated entities employing recovery agentsReserve Bank of India, RBI/2022-23/108, 12 August 2022Regulated entities and their recovery agents shall not call a borrower or guarantor before 8:00 a.m. or after 7:00 p.m.
  5. 05Responsible Business Conduct amendment directions, 2026 (recovery calls)Reserve Bank of India, RBI/2026-2027/223, 6 August 2026, effective 1 January 2027Recovery calls must be recorded, the borrower informed, and recordings preserved for six months or until related litigation concludes; contact only between 08:00 and 19:00 unless the borrower authorises otherwise.
  6. 06Reserve Bank of India (Outsourcing of Information Technology Services) Directions, 2023Reserve Bank of India, RBI/2023-24/102, 10 April 2023, effective 1 October 2023Scope includes application development, data centre and cloud services; data storage only in India as per extant requirements; the regulated entity's right to audit the provider and its sub-contractors, with unrestricted access to data and relevant business premises.
  7. 07Storage of Payment System DataReserve Bank of India, DPSS.CO.OD.No 2785/06.08.005/2017-18, 6 April 2018The entire data relating to payment systems is to be stored in a system only in India.
  8. 08Digital Personal Data Protection Act, 2023 (No. 22 of 2023)Ministry of Law and Justice, Government of India, 11 August 2023Personal data is any data about an individual identifiable by or in relation to it (s.2(t)); transfer outside India may be restricted to countries the Central Government notifies (s.16).
  9. 09The mill: Pod-10/K, the volume SKUAzita unit model, September 2026A ten-kilowatt factory. 1,536 GB of memory in one mill, enough to hold a 671-billion-parameter model. Electrical intake 415 V three-phase, twenty amps a phase, off the board that already feeds the floor; no new connection. 14.0 to 14.4 kW at the meter, not ten: the room has to reject that heat. Water: zero, air-cooled, no structural work and no plant room.
  10. 10What an eight-billion-parameter model weighsAzita analysis, derived; retrieval economics, Azita analysis, September 20264.5 GB for an eight-billion-parameter model, quantised. 1,536 GB in one mill.
  11. 11Utilisation, modelledAzita proof-of-concept model, September 202660% of sellable capacity contracted is the modelled steady state; operating band 45 to 65%; ceiling 80%

Every number on this site is listed with its citation on the sources page. Figures marked modelled are outputs of our model, not measurements.

Earlier · 8 September 2026Why an engineering team buys a room
LaterThis is the latest post.

Let’s find the right fit.

AI for your business

Tell us what you want to run and where your data needs to stay.

Discuss your requirements

Put your building to work

Share a few details about your property. We’ll help you understand its potential and what comes next.

Assess your property

AI factories in buildings that already have power.

Platform

  • How it works
  • RackQuilt
  • Token Mills
  • AzExchange

For buyers

  • AI for business
  • Voice agents for BFSI floors
  • Private coding pods
  • Hospitals and diagnostics
  • VFX assist for post houses
  • Promo and product clips
  • Creators
  • Agencies
  • Engineering
  • Finance
  • Healthcare & BFSI

For buildings

  • Estates
  • For building owners

Tools

  • Pricing
  • Room illustrations

Proof

  • Trust & compliance
  • Sources
  • Research

Company

  • Company
  • Team
  • Careers
  • Press
  • Investors
  • Contact
AI enquirieshello@azitalabs.comProperty enquirieshello@azitalabs.com

© 2026 Azita Labs Private Limited

PrivacyTermsHosting terms

New Delhi, India