Inworld launches TTS-2 models for expressive, low-latency AI character voices
Simulated-people products in other fields — what works, what fails · Wednesday, 2 September 2026
Why it matters
For Inworld-style game characters and other simulated people, controllable prosody, pauses, and nonverbal cues provide a practical layer for making a persona feel responsive without relying only on dialogue text. The launch also shows that believable voice behavior is constrained by unit economics and latency; the reported results are useful signals, but the claims are vendor-reported and do not establish whether the voices remain bounded, non-creepy, or consistent with a character’s deeper persona.
What happened
Inworld launched Realtime TTS-2 and the lower-latency TTS-2 Flash through its API, adding natural-language controls for tone, pacing, pauses, and nonverbal sounds such as laughter, breathing, and sighs. The models support text-described voices, authorized voice cloning from five to 15 seconds of audio, more than 100 languages for delivery steering, and voice consistency across more than 200 languages and 500 dialects. Inworld reports sub-100ms latency for TTS-2 and 25ms time to first byte for Flash, while customer case studies claim lower speech costs and improved usage or retention in products including Talkpal and Status.
Players & places
- Inworld
- LiveKit
- Latitude
- Talkpal
- Wishroll
- Kylan Gibbs