Short answer
Voice AI company ElevenLabs has announced its next generation of speech models, offering wider language support and finer expression control. The new architecture allows voice cloning from a much shorter audio sample and keeps voice identity more consistent across long passages of text for voice assistants.
Highlights
- The new models now cover a wider set of languages than the previous version.
- Users can now clone a voice using just a 10-second audio sample.
- ElevenLabs' annualized revenue run rate has climbed markedly since the start of the year.

2 min readEditor-in-chief: Uğur Deniz İlhan
ElevenLabs has announced two new speech models called ElevenLabs v4 and v4 Turbo, according to TechCrunch. The new versions bring more expression control, lower latency for voice agents, and support for more than 90 languages.
What's new?
The company released its v3 model last year; for the v4 generation it adopted an entirely new architecture. As a result, users can now clone a voice using just a 10-second audio sample. The model keeps voice identity more consistent across longer chunks of text and adjusts expression based on the context of the text as it reads it aloud. The inline expression tags introduced with v3 have been expanded in v4, letting users stack multiple tags in sequence.
Which languages improved?
The previous version supported 70 languages; ElevenLabs says it has now raised that number to 90. The company said it observed the biggest quality jump in Japanese, Brazilian Portuguese, Mandarin and Cantonese. The new model can also start generating audio as soon as the underlying LLM starts generating its answer, shortening wait times in voice assistants.
What should product teams in Turkey do?
For Turkey-based teams building call-center automation and voice assistant products, lower latency and cloning from a shorter sample could cut costs, particularly in multilingual customer-service scenarios. The rapid growth of ElevenLabs' enterprise calling business suggests competing models will move in the same direction; teams choosing a voice provider are advised to weigh not just language coverage but conversational quality criteria such as latency and how well the model handles interruptions.
Frequently asked
- How much audio is needed to clone a voice with v4?
- According to ElevenLabs, the new architecture means users can now clone a voice using just a 10-second audio sample.
Sources
- TechCrunch ·
Follow UNIT Journal
What's new in search, AI and technology, in your feed every day.


