Realtime TTS-2
A new frontier voice model that feels as human as it sounds.
Realtime TTS-2 from Inworld AI is a new generation of voice model built for realtime conversation. It hears the full audio of the exchange, picks up the user's tone, pacing and emotional state, then takes voice direction in plain English the way developers prompt an LLM. It holds one voice identity across over 100 languages. Available today via the Inworld API and the Inworld Realtime API as a research preview.
Voice AI that actually feels human.
"Inworld's TTS-2 marks a real step forward in emotionally expressive voice synthesis. When combined with the conversational intelligence of LiveKit agents, it enables interactions that feel genuinely human — responsive, nuanced, and alive in ways that feel natural."
David Zhao · Co-Founder & CTO, LiveKit
"I've never seen steering work like this before TTS-2. The output is extremely natural and faithful to the steering prompt, even when it's hyper-specific. The biggest battle you fight with TTS is feeling bland, stale, and robotic — this level of steering unlocks a whole new axis to keep the experience fresh."
Creston Brooks · Co-founder & CTO, Luvu
"We believe language learning should have no borders. TTS 2.0 just made that a lot more real."
Dimitri Dekanozishvili · Co-founder, Talkpal
"Inworld is finally closing the gap between 'impressive' and 'actually believable' with TTS 2.0. When your character speaks and you forget it's AI, that's when the story becomes real."
Louis Muk · CEO, Isekai Zero
"We've had early access to Inworld TTS-2 for a few days and we're all blown away. The expressiveness, language steering, and multi-lingual support are genuinely impressive."
Nash Ramdial · Developer Relations, Stream
"Inworld just made voice AI feel genuinely human across 100+ languages. Partnering with them means we can help bring that experience to kids around the world, safely and compliantly."
Kieran Donovan · CEO, k-ID
"AI Native games need characters you can deeply connect with. TTS 2 is a significant advance in helping make that future a reality."
Nick Walton · CEO, Latitude
"Realtime TTS-2 pushes further on a dimension VoiceRun customers care about: directability."
Nick Leonard · CEO & Co-Founder, VoiceRun
Voice AI was built for audiobooks. We rebuilt it for conversation.
Realtime TTS 1.5 already ranks #1 on the Artificial Analysis Speech Arena, ahead of Google and ElevenLabs. Quality is solved. So we asked the next question: what does voice AI sound like when it is built for the way humans actually talk to each other? In realtime, mutual, alive to the moment?
Realtime TTS-2 is built from the ground up for realtime conversation. It listens to the prior turns of the exchange, so your tone and pacing carry forward.
What makes it sound conversational.
Beyond the four capabilities of voice direction, conversational awareness, crosslingual, and advanced voice design, several smaller tools further enhance conversational quality:
01 · Non-verbal markers.
A laugh in the right place lands harder than a paragraph.
02 · Disfluencies.
Real uh and um, in the right places.
03 · Voice cloning.
Bring a real voice in. Use it everywhere.
04 · Stability modes.
Dial expressiveness up or down.
Built for natural realtime conversation
A real conversation isn't just words. It's the tone someone uses, the pause before they answer, the energy they carry into a sentence. Most voice agents stitch a pipeline together and lose all of that signal at every handoff. We built each layer ourselves and pass the full audio context, the user's state, and the conversation history through one persistent connection.
Frequently asked questions
What is Realtime TTS-2?
Realtime TTS-2 is a new generation of voice model from Inworld AI built for realtime conversation. It hears the full audio context of the exchange and the user's emotional state, tone, and pacing.
What languages does Realtime TTS-2 support?
Realtime TTS-2 is expanding to over 100 languages with on-the-fly switching inside a single generation, preserving the speaker's voice identity.
How do I steer the voice?
Voice direction is a natural-language string on the request, the same way you prompt an LLM.
Is voice cloning supported?
Yes. Voice cloning is a two-step API call that allows you to prototype a voice in seconds.
What are the three voice design modes (Expressive, Balanced, Stable)?
Advanced Voice Design ships with three stability modes. Expressive is the most creative, Balanced is the default, and Stable is the most consistent.