xAI's Grok Voice just claimed the top position in agentic performance benchmarks. On July 29, Elon Musk announced the milestone alongside the release of Grok Voice Think Fast 2.0 — a model that improves on its predecessor across speed, transcription accuracy, and real-world voice task completion. Here are the five numbers that define what this model actually delivers.

1. 56.5% on τ-voice Bench — the agentic score that earned the #1 title
The headline claim rests on a single benchmark: τ-voice Bench for Agentic Performance. Grok Voice Think Fast 2.0 scored 56.5% on that test, placing it ahead of its own predecessor (52.1%), GPT-Realtime-2.1 High (45.7%), and Gemini 3.1 Flash High (37.7%), according to xAI's release data. Agentic performance measures how well a voice model can complete multi-step tasks autonomously — not just transcribe speech, but act on it. A nearly 9-point gap over the nearest non-xAI competitor is a meaningful lead, though benchmarks always reflect the specific tasks chosen by their designers.
2. 0.70 seconds — time to first audio
Latency is the silent killer of voice AI. A model can be accurate and capable, but if users hear a half-second pause before every response, the experience degrades quickly. Think Fast 2.0 posts a Time to First Audio of 0.70 seconds, which xAI is positioning as production-grade for real-time conversational use. For context, that's the kind of responsiveness that separates a voice assistant that feels natural from one that feels like a phone tree. Whether that figure holds up under real-world network conditions and concurrent load remains to be seen.
3. 82.9% — overall AA Speech-to-Speech Quality Index
Beyond agentic task completion, xAI also tracks a broader Speech-to-Speech Quality Index that covers naturalness, comprehension, and output fidelity across the full voice pipeline. Think Fast 2.0 scores 82.9% on that composite measure. This is the number that matters most for everyday conversational use — it captures whether the model sounds right and understands correctly across varied inputs, not just on structured benchmark tasks. xAI hasn't published the full methodology for this index, so it's worth treating it as a directional signal rather than a universal standard.
4. 1.5–2.0x better transcription accuracy than leading alternatives
Transcription accuracy — how correctly the model converts spoken audio to text — is the foundation everything else builds on. According to xAI, Think Fast 2.0 delivers a 1.5–2.0x improvement over Deepgram Nova 3 and ElevenLabs Scribe v2, and a 1.4x improvement over Think Fast 1.0. The more striking figure: in noisy environments, that gap reportedly widens to approximately 10x over competing transcription services. If that holds up under independent testing, it would represent a substantial practical advantage for real-world deployments where background noise is the norm rather than the exception.
5. $0.08 per minute — the price point for developers
Performance benchmarks matter less if the pricing makes deployment impractical. xAI has set Think Fast 2.0 at $0.08 per minute of audio — a figure that will determine how quickly developers integrate it into applications. The grok-voice-latest API endpoint is scheduled to transition from Think Fast 1.0 to Think Fast 2.0, meaning existing integrations will automatically gain the new model's capabilities without a code change. For Tesla owners who use Grok through the vehicle or the app, the practical question is how quickly these improvements surface in the in-car experience — that timeline hasn't been specified.
The benchmark lead is real, but benchmarks are a starting point. The more telling signal will come from how Think Fast 2.0 performs in the hands of developers building production voice applications — and whether the transcription accuracy advantage in noisy environments translates to noticeably better in-car performance for Tesla owners using Grok day to day.
Sources & reporting notes
The links below identify the material source records used for this report.
- @elonmusk on X (2026-07-29T20:12:25.000Z) — Direct source
Source links are preserved as published or accessed. See our editorial standards and corrections policy.
The BASENOR Editorial Desk covers Tesla, SpaceX, and related technology, curating reporting from primary sources — official accounts, regulatory filings, and software release data. Every article passes source-record and fact-checking review before publication. About the newsroom.
This report was curated by the BASENOR Editorial Desk from the sources listed above. Read our editorial standards or email editorial@basenor.com to report an error.









