Skip to main content

Voice that listens, adapts, and gets more done | July 2026

GPT-Realtime-2.1 is now available in Open Beta for Bolt Voice Agents using the OpenAI realtime voice path with all Agent voices that say "BETA"

Written by Kootayba Chihabi

Released: July 2026

Overview Header

This upgrade replaces GPT-Realtime-1.5 with a more capable speech-to-speech model built to understand spoken context, reason through more complex requests, respond naturally, and use tools during a live conversation.

The result is a voice experience that should feel more adaptive and consistent, from the first greeting through a complex request or specialist handoff.

Details Header

More natural conversations

GPT-Realtime-2.1 is better at understanding the full meaning of a spoken request, not just transcribing the words

Agents can respond with more natural pacing and expression, while adjusting their tone to better fit the moment. That might mean speaking calmly while resolving an issue, responding more empathetically when a caller is frustrated, or sounding more upbeat when confirming a successful action.

The model is also better equipped to:

  • Handle corrections and interruptions without losing the conversation

  • Recover when the caller changes their request

  • Follow more detailed voice instructions

  • Maintain context through longer conversations

  • Respond with a tone that better fits the caller and situation

A larger context window also gives Bolt Voice Agents more room to carry information through longer, more complex workflows without losing track of what has already happened.

Stronger reasoning for agents that do real work

GPT-Realtime-2.1 brings stronger reasoning, instruction following, and tool use to live voice conversations.

This gives Bolt Voice Agents a better foundation for conversations where they need to do more than answer a simple question. They can reason through a request, follow the agent’s configured instructions, retrieve information, use Custom Skills, and move the conversation forward while the caller remains on the line.

The model also has stronger recognition of specialized terms and proper names, which can help with school-specific language such as program names, campus terminology, staff names, and departments.

It is also better at recognizing and repeating information that mixes letters and numbers, including:

  • Student and application IDs

  • Confirmation codes

  • Phone numbers

  • Email addresses

  • Course and program codes

  • Information spelled aloud by the caller

Better interruptions and background-noise handling

GPT-Realtime-2.1 is designed to distinguish more accurately between meaningful speech, silence, and background noise.

That means the agent should be less likely to treat an incidental sound as a caller interruption, while still responding when the caller genuinely speaks over the agent, corrects something, or changes direction.

This matters because natural conversation is rarely perfectly turn-based. Callers pause, rethink a question, speak over an answer, or correct a detail. The upgraded model gives Bolt Voice Agents a better foundation for handling those moments without losing context or restarting the interaction.

More natural and reliable agent handoffs

Voice handoffs between Bolt Agents now include more natural transition pacing, so the conversation does not switch agents abruptly.

We also improved multi-specialist routing. When a caller changes topics after being transferred, the current agent can hand the call to another configured specialist rather than becoming stuck with the first receiving agent.

For example, a caller can move from an admissions question to housing, then ask about financial aid, with each request routed according to the institution’s configured Hand Off to Agent skills.

Language continuity has also been improved during agent handoffs and tool progress messages. Once a conversation is taking place in a particular language, the agent will maintain that language throughout its answers, acknowledgements, and tool-related updates.

Faster responses at the moments that matter

OpenAI reports at least a 25% reduction in latency across its Realtime voice models through improved caching.

In practical terms, this is intended to reduce the unusually slow responses that can make a live conversation feel delayed or unresponsive.

It does not eliminate every wait. Actions that rely on external systems, message delivery, or longer-running tools may still take additional time. We are continuing to improve the short acknowledgements and progress updates callers hear while those actions are running.

Benefit Header

Available now in Open Beta

GPT-Realtime-2.1 is now available for Bolt Voice Agents using the OpenAI realtime voice experience with all Agent voices that say "BETA".

The feature will remain in Open Beta while we continue validating it across more institutions, agents, voices, languages, accents, Custom Skills, and call patterns.

Areas we are continuing to improve include:

  • Tool calling across a broader range of customer workflows

  • Progress updates during longer-running actions

  • Evaluation coverage for future model upgrades

  • Consistency across different voices and calling conditions

From the first greeting through a complex request or specialist handoff, GPT-Realtime-2.1 gives Bolt Voice Agents a more natural, responsive, and capable foundation for live student conversations.

Did this answer your question?