Every language.
Speaking in milliseconds.
A neural voice that speaks 43 languages with two voices, plus 29 experimental ones, and starts talking before the sentence is even finished.
It starts talking while it is still thinking.
Text to speech is only useful in a conversation if the wait disappears. Tensorial TTS synthesizes in short windows and streams each one the moment it is ready, so the listener hears the start of a sentence while the end is still being computed.
- Streaming, not buffering. Audio leaves in small chunks, in all six formats, with incremental encoding.
- Light on hardware. Two seconds of speech cost about 150 ms of CPU. It runs on commodity ARM servers.
- Built for agents and live voice. Short utterances, the common case in conversation, are exactly where the first-chunk time matters most.
A 6-second sentence, time until you hear it
Measured locally on an Apple M4 (CPU, ONNX, 2 threads), warm, 6.4 s of English audio. Service latency adds your network round trip.
Hear every language.
The same sentence in every language, synthesized by the model. Press play, switch the voice, search by name.
Stable, and honestly experimental.
We separate what the model was trained on from what it only borrows, so you know what you are shipping.
43 stable languages
Trained and evaluated. From English, Portuguese, Spanish and Mandarin to Hindi, Arabic (three regional variants), Japanese, Swahili and Icelandic. Full text normalization: numbers, currency, units, dates and times.
29 experimental languages
From Catalan and Basque to Kazakh, Georgian, Māori and Quechua. Same full text frontend; only the model is experimental. Published separately in the API so you can opt in on purpose.
Two voices, everywhere
Juliana and Marco Antonio speak every language, stable or experimental. Japanese mixed with Latin-alphabet words is handled in one sentence.
A drop-in speech endpoint.
A plain HTTP API compatible with POST /v1/audio/speech. Point your existing client at it, choose a language and a voice, and stream the audio back.
- Six formats · streamed incrementally
- 22.05 kHz · or 48 kHz with bandwidth extension
- Keys, quotas and a web console
# stream Portuguese speech to a file curl https://<your-endpoint>/v1/audio/speech \ -H "Authorization: Bearer $TENSORIAL_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "tensorial-vits-plbert-istft-43-3.0.0", "input": "Olá! Esta é a Tensorial.", "voice": "vox_a64467727c6f9ed4", "language": "pt-BR", "response_format": "mp3", "stream": true }' --output fala.mp3
Give your product a voice in any language.
Tell us what you are building and we will get you access.
Figures: latency measured on the v2 model, warm, on local CPU hardware; first-audio time in production adds network. Stable languages were checked with an automatic judge and two sentences per language; experimental languages were not reviewed by native speakers.