Blog · 2026-09-30
Tensorial TTS v2: 43 languages, 29 more on trial, and a voice that starts in 70 milliseconds
Our multilingual speech model in production: what it speaks, how it streams, and how we separate what it learned from what it borrows.
Tensorial TTS v2 is our multilingual speech synthesis model, now in production. It speaks 43 languages with two voices, Juliana and Marco Antonio, in every one of them, and publishes 29 more experimental languages. Above all, it was designed for one thing: to start talking now. Hear every language on the product page.
Very low latency
In a conversation with a voice agent, nobody notices how good the voice is if it arrives late. So the model synthesizes in short windows and streams each one the moment it is ready. On a 6.4-second English sentence we measured in our test setup (an Apple M4, CPU only, two threads, warm model) about 70 ms to the first audio chunk, against roughly 540 ms to receive the whole clip. The compute cost is around 0.1 of the speech duration, which means one CPU thread synthesizes about ten times faster than real time, with no GPU.
These are local measurements. In production, time to first audio adds your network round trip; the API streams in six formats (mp3, opus, aac, flac, wav and pcm) with incremental encoding.
43 stable languages
The stable languages were trained and evaluated. They range from English, Portuguese (Brazil and Portugal), Spanish (Spain and Mexico), French, German and Italian to Japanese, Korean, Mandarin, Cantonese, Hindi, Bengali, Tamil, Arabic (three regional variants), Hebrew, Persian, Turkish, Swahili and Icelandic. Text normalization is complete in each one: numbers, currency, units, dates and times are read the way a person would read them.
On 2026-09-29 an automatic judge scored two sentences per language. Across the 43, the transcription error stayed low, except for Mandarin and Cantonese, whose metric is inflated because the judge writes in a different variant of the script. This is triage, not a verdict: the weakest languages in the test were Thai, Lithuanian, Marathi, Hebrew, Persian, Finnish and Cantonese, and they still need review by native speakers.
29 experimental languages, honestly
The model can also speak languages it never saw in training. It receives the real language's phonemes and borrows the conditioning of a close trained language. The content comes out right, but the prosody is the borrowed language's: where it fits, it sounds native; where it does not, it carries an accent. That is why these languages are kept apart in the API, in their own field, and marked as experimental on the page. They passed an automatic gate, and none has been reviewed by a native speaker.
- Catalan, Basque, Croatian, Serbian, Bulgarian, Albanian and Macedonian
- Kazakh, Georgian, Armenian, Azerbaijani and Uzbek
- Malay, Nepali, Punjabi, Gujarati, Kannada, Malayalam and other Indian languages
- Māori, Hawaiian, Papiamento, Haitian Creole and Quechua
How to use it
The API is compatible with POST /v1/audio/speech. You pick the model, the voice
and the language, ask for stream: true and play the chunks as they arrive. The
language is never guessed: guessing was measured at 82.7% accuracy on short texts, and the
errors fall between similar languages, which would produce fluent audio in the wrong language.
curl https://<your-endpoint>/v1/audio/speech \
-H "Authorization: Bearer $TENSORIAL_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "tensorial-vits-plbert-istft-43-3.0.0",
"input": "Hello! This is Tensorial.",
"voice": "vox_a64467727c6f9ed4",
"language": "en",
"stream": true}' --output speech.mp3
Request access
We are opening access gradually. Hear the languages and, if you want to use Tensorial TTS in your product, get in touch.