AI Voice Agent Cost Per Minute: Retell vs Vapi vs Bland

You get quoted a base rate of $0.05 per minute and build your budget around it. Then the first month's invoice arrives at three times that, because the LLM tokens, the text-to-speech, and the telephony pass-through all bill separately. This is the honest all-in math before you commit to a platform.
Quick answer: The real AI voice agent cost per minute lands between $0.10 and $0.33 all-in for platforms like Retell, Vapi, and Bland — not the $0.05-$0.07 base rate they advertise. The gap comes from stacking four separate charges: the platform's orchestration fee, LLM tokens, text-to-speech, and telephony (Twilio) pass-through. Multiple 2026 pricing analyses (ainora.lt, cekura.ai, cloudtalk.io) put Retell's blended cost at $0.13-$0.31/min in production.
What does an AI voice agent actually cost per minute?
An AI voice agent costs $0.10 to $0.33 per minute all-in once you stack every component, versus the $0.05-$0.07 base rate the platform quotes you. The advertised number is only the orchestration layer — the software that connects the pieces. Everything the agent actually needs to hold a conversation gets billed on top.
Every minute of a live call stacks these charges:
- Platform orchestration — the base rate Retell, Vapi, or Bland advertises. Roughly $0.05-$0.07/min.
- LLM tokens — the model reasoning and generating replies. Depends on model choice and how much context you pass per turn.
- Text-to-speech (TTS) — the voice itself. ElevenLabs premium voices cost more than default options; this is often the second-largest line item.
- Speech-to-text (STT) — transcribing the caller. Usually bundled, sometimes not.
- Telephony — the actual phone connection, almost always Twilio pass-through at roughly $0.014/min inbound.
Add those and a "$0.07/min" agent using a premium voice and a frontier LLM can bill $0.20-$0.30/min in real production. The three independent 2026 analyses agree: Retell's blended real-world cost sits around $0.13-$0.31/min.
Retell vs Vapi vs Bland: how the pricing models differ
All three charge a base orchestration rate near $0.05-$0.09/min and pass through LLM, voice, and telephony separately. The difference is how much they bundle and how transparent the bill is. None of them is the "cheap" option once you run real volume; they trade convenience against control.
| Platform | Advertised base | Blended real-world (2026 estimates) | Model |
|---|---|---|---|
| Retell | ~$0.07/min | $0.13-$0.31/min | Base + pass-through LLM/TTS/telephony |
| Vapi | ~$0.05/min | $0.10-$0.30/min | Base + pass-through, granular add-ons |
| Bland | ~$0.09/min all-in | $0.09-$0.20/min | More bundled, less componentized |
Bland markets a more bundled all-in number, which looks cheaper on the label but gives you less control over which LLM and voice you run. Vapi is the most granular — powerful if you want to tune every component, more surface area to get surprised by. Retell sits in the middle and is our platform of choice as a Gold Retell partner, because the orchestration is reliable enough to run a business phone line unattended.
The mistake buyers make is comparing base rates. Compare the blended number with your model and voice choices plugged in, or you're comparing marketing labels.
The per-component breakdown you can plug your own volume into
To estimate your true cost, price each component separately, multiply by your monthly call minutes, and add telephony. This is the math no vendor comparison hands you because the answer depends on your voice and model choices, not theirs.
Estimate a single minute
Work bottom-up for one minute of talk time:
- Base platform rate — start at $0.06/min (a reasonable Retell/Vapi midpoint).
- LLM — a fast frontier model runs roughly $0.02-$0.06/min depending on context size per turn. Text-heavy prompts and long call histories push this up. Anthropic's Sonnet 5 launched with introductory API pricing of $2/M input and $10/M output through August 31, 2026, which improves the economics of chatty, text-heavy agent turns.
- TTS voice — default voices run cheap; ElevenLabs-tier premium voices add $0.05-$0.10/min. This is where "just use the nicer voice" quietly doubles your bill.
- STT — often bundled; budget $0.01-$0.02/min if it isn't.
- Telephony — Twilio at ~$0.014/min inbound.
Sum a realistic mid-tier build: 0.06 + 0.04 + 0.07 + 0.015 + 0.014 ≈ $0.20/min. That's your honest number, not the $0.06 on the pricing page.
Multiply by real volume
Take the law-firm voice agent we built, "Emily", which handled 571 calls hello-to-booked over 8 months. At an average of 4 minutes per call, that's roughly 2,284 minutes. At $0.20/min all-in, the raw platform cost is about $457 over 8 months, while the agent returned $10K-$25K/mo to the firm and caught all 176 after-hours calls with a 96.5% self-serve rate. The per-minute cost is almost a rounding error against the revenue when the agent actually books clients. We break down those numbers in the Emily case study.
When a flat managed price beats per-minute
A flat managed price wins when you can't predict volume, don't want to babysit four vendors' billing, or need the agent to reliably book revenue rather than just be cheap. Per-minute pricing is only "cheaper" if you value your time at zero and never hit a spike month.
Per-minute makes sense when:
- Your call volume is low and steady.
- You have the technical time to tune LLM, TTS, and telephony yourself.
- You're a builder who wants full control (see how to build a voice agent you can sell).
A flat build-and-manage price makes sense when:
- You're a business owner who wants a working phone line, not a billing dashboard.
- Volume is spiky — after-hours surges, seasonal rushes — and a per-minute bill could balloon.
- The agent's job is revenue: booking clients, catching missed calls. Every missed call is a business's most ready-to-buy lead, and the missed-call gap costs more than any per-minute rate.
The honest tradeoff: DIY per-minute can hit $0.10-$0.20/min if you tune it well, but you own the setup, the retries, the failovers, and the 2am debugging. A managed build costs more per minute on paper and saves you from becoming the quality check. We break down that decision in DIY vs hiring an agency.
Common pitfalls when pricing a voice agent
The three mistakes that blow up a voice agent budget are all invisible on the pricing page. Watch for them before you commit.
- Comparing base rates instead of blended rates. The $0.05 label is orchestration only. Always price LLM, TTS, and telephony on top.
- Ignoring retries and failed calls. Dropped calls, retries, and tool-call loops still burn tokens and minutes. Real cost depends on prompt size, tool calls, and retries — budget a 10-20% overhead.
- Picking the premium voice by default. ElevenLabs-tier TTS can be the single biggest line item. Test whether a cheaper voice books just as many calls before you pay for the premium one.
- Optimizing cost over conversion. A $0.10/min agent that books nothing is more expensive than a $0.30/min agent that catches every after-hours lead. Judge on booked outcomes, not per-minute price — the KPI benchmarks tell you if it's working.
FAQ
How much does an AI voice agent cost per minute?
An AI voice agent costs $0.10-$0.33 per minute all-in once you stack platform orchestration, LLM tokens, text-to-speech, and telephony. The $0.05-$0.07 base rate platforms advertise covers only the orchestration layer, not the components the agent needs to hold a conversation.
Is Retell cheaper than Vapi or Bland?
Not meaningfully — all three land in the same $0.10-$0.31/min blended range once you add LLM, voice, and telephony. Bland bundles more into a higher base rate, Vapi is the most granular, and Retell sits in the middle with reliable orchestration, which is why we run it as a Gold Retell partner.
Why is my voice agent bill higher than the advertised rate?
Because the advertised rate is the platform's orchestration fee only, and it bills LLM tokens, text-to-speech, speech-to-text, and Twilio telephony separately on every minute. A premium ElevenLabs voice and a frontier LLM can easily push a "$0.07/min" agent to $0.20-$0.30/min.
Does a voice agent's per-minute cost matter more than what it books?
No — a voice agent that books clients justifies its per-minute cost many times over. Emily cost roughly $457 in platform time over 8 months while returning $10K-$25K/mo to a law firm, making the per-minute rate almost a rounding error against revenue.
Want the exact platform, prompts, and setup we use to run a business phone line unattended? Join the free Claude Community to see the builds, or if you'd rather have it done for you in under 30 days, get a custom voice agent built.
About Terrell Gentry
Founder at 6omb
Terrell is the founder of 6omb and runs Claude Community, the #1 Skool community for Voice AI agents. Over 16 months his team has built 100+ AI agent systems delivering $10M+ in business value, including voice agents like Emily, which booked 453 new clients for a law firm in 8 months. He is a Y Combinator Startup School alum (SUS20) and a Gold Retell partner.
You might also like

AI Answering Service Cost vs Hiring a Receptionist
AI answering service cost vs hiring a receptionist: honest math tied to call volume, per-minute pricing, and what a human still does better.

AI Voice Agent for Dental Offices: Missed Calls to Bookings
An AI voice agent for dental office phone lines catches missed calls and books appointments 24/7. Real hello-to-booked flow, PMS reality, and honest limits.

Custom AI Agent Cost for Small Business (2026 Real Ranges)
Custom AI agent cost for a small business in 2026: real price ranges by what the agent does, plus why the API bill is noise vs integration labor.
Join 10k+ founders going AI-first with Claude
The Claude Masterclass, 50+ copy-paste Claude Code skills, agent-building workshops, and a community actively building the same thing you are. Free for now.
Join the free community