top of page
davydov consulting logo

DAVYDOV CONSULTING BLOG

HOME  >  NEWS  >  POST

Google Launches Gemini 3.8 Live: What Real-Time Voice AI Means for Your Business

21 hours ago
6 min read

On 15 September Google released two new voice models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These are not chatbots that read their answers aloud. They are native speech-to-speech models, which means they listen and talk directly, without converting everything to text in the middle. Google is pitching them as the engine for production voice agents, the kind that answer your support line or walk a new employee through onboarding.

The launch matters because of the numbers behind it. The Extended Thinking variant now sits at the top of the main independent quality leaderboard for voice AI, just ahead of OpenAI and xAI, and Google priced the standard model at roughly one sixth of what its closest rival charges. When the quality leader is also the price leader, businesses that were waiting on voice AI have fewer reasons left to wait.

In this article we look at what the two models actually do, what the benchmarks and prices say, and what a business that builds or orders digital products should take from all of it. Let's sink in!

Business professional speaking to a Gemini voice AI assistant on her smartphone in a modern office

What Google Actually Released

Gemini 3.8 Live is the workhorse of the pair. Google built it for scale and cost efficiency, the situations where a company needs thousands of parallel conversations and cannot pay premium rates for each minute. It handles real-time audio in and out, processes visual input at near real time, and switches automatically between 97 languages, even in the middle of a conversation. A caller can start in English, drop into Polish, and the model follows without being asked.

Gemini 3.8 Live Extended Thinking is the premium variant. Its trick is that it reasons and speaks at the same time. Older voice systems either went silent while they worked through a complex request, which callers read as a dropped line, or they skipped the thinking and gave shallow answers. Extended Thinking fills the gap naturally, with phrases like "Let me check that", while the actual multi-step work runs in the background.

Both models support what Google calls asynchronous function calling. The model can fire off an API request, a database lookup or a booking action in the background and keep the conversation going while the result comes back. They also handle alphanumeric data with more precision than previous generations, so confirmation codes and order numbers survive the trip. Both are available through the Gemini Live API and Google AI Studio from day one, with enterprise previews in Gemini Enterprise and a rollout into Google Workspace for the Extended Thinking variant.


The Numbers Behind the Launch

Voice AI has had a credibility problem for years, and it usually showed up in benchmarks. This launch is the first time Google can point at an independent leaderboard and claim the top spot. On the Speech to Speech Quality Index from Artificial Analysis, Gemini 3.8 Live Extended Thinking scores 82.6, ahead of GPT-Live-1 Astra at 81.5 and Grok Voice Think Fast 2.0 at 81.3. The margins are small, but the direction matters.

The agentic benchmarks tell a similar story. On τ-Voice, which measures whether a voice agent actually completes tasks rather than just chatting about them, Extended Thinking reaches 68.6 percent against 67.9 for OpenAI and 56.5 for xAI. On Sierra's banking leaderboard, a deliberately hard test of regulated, multi-step customer tasks, it manages 35.1 percent, with OpenAI at 32 and xAI's realtime model at 16.5. On pure audio reasoning, Big Bench Audio, it hits 97.7 percent.

One honest note here. A 35 percent completion rate on realistic banking tasks means the best voice agent on the market still fails almost two thirds of the time on genuinely hard workflows. These models are ready for well-scoped jobs, not for replacing a whole contact centre on their own.

Customer support team working alongside AI voice agents with live audio waveform screens

Pricing That Changes the Conversation

The quality numbers grab headlines, but the pricing is what will move budgets. Gemini 3.8 Live costs $0.005 per minute of audio input and $0.018 per minute of output, which works out to roughly $0.84 per hour of input audio. The Extended Thinking variant lands around $3.50 per hour. Compare that with $4.80 for Grok Voice Think Fast 2.0 and $5.83 for GPT-Live-1 Astra, and the gap is hard to ignore.

For a business, this changes the shape of the calculation. At the old price levels, a voice agent handling long support calls at scale was a serious line item, and many projects died in the spreadsheet stage. At under a dollar per hour of listening, pilots that looked marginal a quarter ago suddenly look cheap against the cost of a human answering the same routine questions.

There is a catch worth knowing. The models are hosted API only. There is no self-hosted option, so your audio goes through Google's infrastructure and your costs follow whatever Google decides to charge next year. That is a dependency, and it belongs in any serious evaluation, especially for companies in regulated sectors.


Where It Fits Against OpenAI and xAI

The competitive picture is tighter than the marketing suggests. OpenAI's GPT-Live-1 Astra remains within about one point on quality, and its ecosystem around realtime agents is mature. xAI is behind on the hard agentic tests but competitive on raw conversational quality. What Google has done is undercut both on price while nudging ahead on quality, and that combination is aimed squarely at businesses making a platform choice this year.

Google also arrives with distribution the others cannot match. The models are shipping into Search Live, Gemini Live for subscribers and Google Workspace, which means millions of users will get used to talking to this exact voice stack in their daily tools. When your customers already talk to Gemini in their browser, a Gemini-powered agent on your phone line feels less strange to them.

The integration ecosystem is also broader than you might expect at launch. Agora, LiveKit, LangChain, Pipecat, Vercel and several other platforms supported the models on day one, so the usual pattern of waiting six months for tooling to catch up does not really apply here.

Small business team listening to a multilingual AI voice assistant during a team meeting

What This Means for Businesses Building Digital Products

The practical question is what to do with this. Our advice is to start with one narrow, high-volume conversation your business already has. Appointment booking, order status, basic triage before a human takes over, first-line answers about opening hours and pricing. These are exactly the well-scoped tasks where current completion rates are strong, and the per-minute economics now work even for a mid-sized company.

Second, take the multilingual support seriously. Automatic switching between 97 languages is not a gimmick if you serve customers across markets. A small e-commerce company can now offer phone support in a dozen languages without hiring for a single one of them, which was simply not a realistic option before.

Third, design for the handover from the start. The benchmark numbers say the same thing we tell clients about every agent project: the system will fail on hard cases, so the product decision is what happens when it does. A voice agent that hands a caller to a human with full context is an asset. One that loops and frustrates is a liability that costs you customers.

And finally, if you already run a voice agent on OpenAI or on one of the older text-to-speech pipelines, this launch is the trigger to re-run your cost model. Switching providers is real work, but a five-times price difference on input audio pays for a lot of migration effort, and the partner tooling makes the technical side smaller than it used to be.


Final Notes

Gemini 3.8 Live is the moment voice AI pricing stopped being the blocker. Google now leads the quality table and undercuts everyone on cost, and the models shipped with real tooling support rather than a waitlist. The technology still has clear limits on complex tasks, and the hosted-only model creates a dependency that deserves eyes-open evaluation.

Still, the direction is obvious. Voice is becoming a normal interface for business software, the way chat became one two years ago. The companies that benefit first will be the ones that pick one concrete conversation, ship a narrow agent, and learn on real callers while their competitors are still watching demos. If you are weighing up where voice AI could fit in your product or your support operation, that first narrow pilot is the right conversation to have, and we are happy to help you scope it.

 
 
 

Comments


​Thanks for reaching out. Some one will reach out to you shortly.

CONTACT US

bottom of page