Skip to main content
VARTA

Cost calculator methodology

The cost calculator exists to be checked, not trusted. This page explains, in full, how it turns four sliders into two monthly totals: where every rate comes from, what the synthetic call it prices is built from, the two places the model gives VARTA the benefit of the doubt, and what it leaves out entirely.

Where the rates come from

Every provider rate below is mirrored from the same pricing tables the VARTA product uses internally for its own per-call cost ledger (cost_model.py). That mirror was last synced on 27 Aug 2026 -- see “What the mirror date does and doesn't mean” below before reading that date as a verification date. Each row underneath carries its own, independent asOf, which is when that specific figure was last checked against the linked source -- rows are not all checked on the same day, and several predate the mirror sync by months.

Text-to-speech -- per 1,000 characters

ProviderRateSourceChecked
elevenlabs_flash_v2_5$0.050elevenlabs.io14 May 2026
elevenlabs_turbo_v2_5$0.100elevenlabs.io14 May 2026
elevenlabs_multilingual_v2$0.180elevenlabs.io14 May 2026
openai_tts_1$0.015openai.com14 May 2026
openai_tts_1_hd$0.030openai.com14 May 2026
sarvam_bulbul_v2$0.040sarvam.ai14 May 2026
cartesia_sonic$0.065cartesia.ai14 May 2026
deepgram_aura$0.015deepgram.com14 May 2026
azure_neural$0.016azure.microsoft.com14 May 2026
google_wavenet$0.016cloud.google.com14 May 2026
google_chirp3_hd$0.030cloud.google.com14 May 2026

Language models -- per 1,000,000 tokens, in / out

ModelInputOutputSourceChecked
gpt-4o$2.50$10.00openai.com14 May 2026
gpt-4o-mini$0.15$0.60openai.com14 May 2026
gpt-5.4$2.50$15.00developers.openai.com20 Jul 2026
gpt-5.4-mini$0.75$4.50developers.openai.com20 Jul 2026
gpt-5.4-nano$0.20$1.25developers.openai.com20 Jul 2026
gpt-5.4-pro$30.00$180.00developers.openai.com20 Jul 2026
gpt-5.5$5.00$30.00developers.openai.com20 Jul 2026
gpt-5.5-pro$30.00$180.00developers.openai.com20 Jul 2026
gpt-5.6-sol$5.00$30.00developers.openai.com20 Jul 2026
gpt-5.6-terra$2.50$15.00developers.openai.com20 Jul 2026
gpt-5.6-luna$1.00$6.00developers.openai.com20 Jul 2026
claude-haiku-4-5$1.00$5.00www.anthropic.com14 May 2026
claude-sonnet-4-6$3.00$15.00www.anthropic.com14 May 2026
claude-opus-4-7$15.00$75.00www.anthropic.com14 May 2026
gemini-2.5-flash$0.30$2.50ai.google.dev14 May 2026
gemini-2.5-flash-lite$0.10$0.40ai.google.dev10 Jul 2026
gemini-2.5-pro$1.25$10.00ai.google.dev10 Jul 2026
gemini-2.0-flash$0.10$0.40ai.google.dev10 Jul 2026
gemini-3.5-flash$1.50$9.00ai.google.dev20 Jul 2026
gemini-3.1-flash-lite$0.25$1.50ai.google.dev20 Jul 2026
gemini-3.1-pro-preview$2.00$12.00ai.google.dev20 Jul 2026
gemini-3-flash-preview$0.50$3.00ai.google.dev20 Jul 2026

Speech-to-text -- per minute (streaming)

ProviderRateSourceChecked
deepgram_nova_3VARTA's fixed default$0.0077deepgram.com15 Jul 2026
elevenlabs_scribe_v2$0.0065elevenlabs.io15 Jul 2026
sarvam_saaras_v2$0.0060docs.sarvam.ai15 Jul 2026
openai_whisper$0.0060openai.com14 May 2026
google_chirp$0.0240cloud.google.com14 May 2026

Telephony -- per minute, outbound

CarrierRateSourceChecked
twilio_us_outboundVARTA's fixed default$0.0140www.twilio.com14 May 2026
exotel_in_outbound$0.0120exotel.com14 May 2026
plivo_in_outbound$0.0110www.plivo.com14 May 2026

VARTA's own speech-to-text and telephony vendor are fixed product defaults (deepgram_nova_3 and twilio_us_outbound, marked above) -- they are not part of the baseline stack you pick on the calculator, so they price the same regardless of which comparison you choose. VARTA's text-to-speech and language-model rates are not fixed: the calculator prices VARTA on the same TTS and LLM vendor rates as whichever baseline stack you selected, so the comparison isolates architecture -- caching and bounded context -- rather than mixing in a vendor-price difference at the same time. That is a deliberate choice: an earlier version of this model priced VARTA on a fixed premium stack regardless of the baseline, which let a cheaper baseline's vendor discount overwhelm VARTA's architectural saving and produce a negative number for the site's own default comparison.

The provisional-pricing caveat

These figures are recorded as vendor list prices in the source cost model and have not been independently re-verified here — treat each as a starting point, not a confirmed-current rate. Each provider row is dated (asOf) to when it was last checked against its sourceUrl — some rows are older than others. List prices are used on purpose: a vendor offering a negotiated discount would only widen the savings gap this calculator shows, so this comparison is conservative rather than inflated. Refresh each rate against its source before citing it in a published case study.

That caveat is live, not a formality left over from an earlier draft: as of this page's last build, these figures are still flagged as unverified vendor list prices rather than confirmed-current ones.

What the mirror date does and doesn't mean

27 Aug 2026 is the date this page's rate tables were last copied out of the product's own pricing module and matched to be identical to it -- not the date any individual rate was itself checked against the vendor named in its “Source” column. Those are two different dates, and this page keeps them separate on purpose: the mirror date tells you the copy is faithful to the source module; each row's own “Checked” date tells you how stale that particular figure might be. Several rows above are dated months before the mirror date -- that is expected, not an error, and is exactly what the provisional-pricing caveat is for.

The FX rate

Every USD line item is converted to INR once, at the total, using a pinned rate of 84.00 = $1.00, dated the same 27 Aug 2026 as the rest of this mirror. The public calculator deliberately does not call a live FX feed: a rate that moves between two visits to the same shared link would make the comparison harder to reproduce, not easier, for the one page whose job is letting a reader check the arithmetic. Converting once at the total (rather than converting each component and summing) avoids compounding rounding error across five line items.

What the synthetic call trace assumes, and why

The calculator only collects four aggregate inputs -- calls per month, average call length, cache-hit rate, language -- but the real cost estimators price a call turn by turn. To bridge that gap, a synthetic per-turn trace is invented whose aggregate shape matches those four inputs, so the same turn-by-turn arithmetic that prices a real recorded call can price a hypothetical one. Every constant below is a modelling assumption, not a measurement pulled from real call logs -- treat each as a documented, disagreeable judgement call, not a verified fact about real traffic.

Pace of a call
6 conversation turns per minute of call time -- about one agent-customer exchange every twenty seconds.
Agent turn length
120 characters of spoken reply -- roughly one short sentence, the length of a typical workflow prompt or confirmation.
Customer turn length
15 tokens of transcribed reply -- a short yes/no, a date, or a name.

Cache hits are spread across agent turns deterministically, by index, so that the same inputs always produce the exact same trace and the same total -- a public calculator that returned a different number on refresh would undermine the trust the whole page exists to build. Nothing in the trace synthesiser calls a random-number generator.

The two assumptions that favour VARTA

A page that claims to be auditable is not auditable if it only discloses the assumptions that run against the thing it is selling. Two modelling choices here make VARTA look cheaper than a real call may be -- both are disclosed plainly, in the same place as every other assumption, not buried:

  1. A bounded, 340-token system prompt. Every answered turn charges VARTA's LLM input as a fixed 340 tokens plus the preceding customer turn -- the measured size of VARTA's own routing prompt, used here as a lower bound on VARTA's real per-turn context, not a measurement of everything a live call's context can hold. The naive baseline, by contrast, is charged its full, growing conversation history on every turn -- which is the architectural difference this whole comparison exists to show, but it also means VARTA's side of the ledger is anchored to a best case.
  2. One LLM call per answered turn. The model charges exactly one LLM call for every VARTA turn that was not served from cache, though a real turn's recorded call log can include more than one LLM call -- for example, a classification step and a separate routing step. There is no measured figure to anchor a second invented constant for how many extra calls a typical turn makes, so rather than guess, this is disclosed as a known simplification instead of modelled: it is a genuine undercount of VARTA's cost, not a conservative one.

Neither assumption is applied to the baseline stack. Both push VARTA's modelled cost down and therefore push the reported saving percentage up -- so read the number on /cost as closer to a ceiling on VARTA's advantage than a guarantee of it.

What this model deliberately excludes

  • Any per-minute price for VARTA itself, at any point. VARTA is not billed by the minute here or anywhere else -- a time-based rate would hide the effect of caching, which is the entire point of this comparison.
  • Negotiated or volume-discounted vendor pricing. Every rate above is a public list price; a vendor discount on either side would only widen the gap this calculator shows, so using list prices throughout is conservative rather than inflated.
  • One-time or fixed costs -- engineering time, integration work, platform or support fees, compliance overhead. This model only prices the five per-call, per-minute line items in the breakdown table: text-to-speech, language model, speech-to-text, telephony, and the WebRTC media-transport layer.
  • Human-agent labour cost. This is a comparison between two automated-call architectures, not automation against a human contact centre.
  • Taxes, regulatory surcharges, or carrier-specific fees layered on top of a per-minute telephony rate.
  • Any effect of language on unit cost. The language selector changes which baseline stack is suggested by default (an Indic stack for Hindi and mixed calls, a US stack for English) -- a genuine, sourced change of which vendor rates apply -- but it does not scale characters or tokens, since inventing a per-language multiplier would be false precision on the one page whose job is auditability.

What this shares with production, and what it doesn't

What this calculator shares with a real call is the rate tables and the per-unit cost helpers: the TTS, LLM, STT and telephony rates above are mirrored from cost_model.py, and that module's per-unit functions -- tts_cost_inr, llm_cost_inr, and stt_cost_inr -- are exactly what a live call calls, inline, to price each TTS synthesis, each LLM call and each transcribed minute as it happens. Those are the functions actually wired into production, in app/varta/l4_reasoning.py, app/routers/sessions.py and app/core/state.py.

What this calculator does not share with production is the two functions that turn a whole trace into a total. estimateVarta and estimateBaseline on this page are a structural port of estimate_varta_cost and estimate_baseline_cost -- two functions that live in cost_model.py alongside the per-unit helpers, kept function-by-function parallel to the Python originals on purpose and checked against golden values generated by running that real Python module, not reimplemented from a description of it. But nothing in the live call path invokes them: in production they run from the benchmark harness and from that module's own test suite, not from a real call. A real call's total is instead the running sum the per-unit helpers accumulate onto the session as the call happens, turn by turn -- the same arithmetic shape these two aggregating functions reproduce over an invented trace, checked line-for-line against goldens rather than assumed to agree, but not the identical call path a real call takes.

One more gap worth naming here rather than leaving for a reader to find: this calculator's LLM rate table is the one llm_cost_inr actually bills a call against. A separate part of the product tracks a different kind of LLM spend -- the cost of generating or editing a workflow with VARTA's AI co-author -- against its own, independently maintained rate table. That authoring-time tracker has nothing to do with what a call costs and this calculator does not draw on it, but it is a second LLM price list in the same codebase, so if you go looking for “the” LLM rates and find two, this is why.