Choose a text-to-speech model
Start with the result you need: English narration, quick CPU-based speech, wider language support, or a model that accepts a voice reference. Models marked ready can be used here now.
Not sure what your machine can handle? Run the compatibility check.
Start with what matters to your project
These are practical starting points, not universal rankings. Open the model page to check voices, languages, download size, and browser requirements before generating a full script.
Ready to use in your browser
Kokoro-82M
English speech with 28 voices and CPU or WebGPU paths
- Size
- 88 MB CPU / 310 MB WebGPU
- Languages
- 2
- License
- Apache-2.0
- Good for
- English voice selection
Piper
Fast speech generation across a wide range of languages
- Size
- 27-115 MB per voice
- Languages
- 24
- License
- MIT (core)
- Good for
- Fast and multilingual
KittenTTS Nano
A small English model for lightweight devices
- Size
- 57 MB
- Languages
- 1
- License
- Apache-2.0
- Good for
- Smallest download
Kyutai Pocket TTS
Streaming English speech with optional voice references
- Size
- 189 MB
- Languages
- 1
- License
- CC-BY-4.0 (weights), MIT (code)
- Good for
- Streaming and voice reference
Research and future models (7)
These records document additional browser experiments and server-class models. They are not part of the four reviewed model guides above unless they become usable here and pass the publication gate.
Experimental and planned models
Supertonic-3
One multilingual model with 31-language coverage
- Size
- 398 MB
- Languages
- 18
- License
- OpenRAIL-M (weights), MIT (code)
- Good for
- Multilingual WebGPU
Chatterbox
Expressive English speech with zero-shot voice references
- Size
- 1550 MB
- Languages
- 1
- License
- MIT
- Good for
- Large WebGPU model
MOSS-TTS-Nano
Browser-oriented voice references across 20 languages
- Size
- 684 MB
- Languages
- 20
- License
- Apache-2.0
- Good for
- Experimental multilingual
Chatterbox Multilingual
Large multilingual speech model with voice references
- Size
- 1600 MB
- Languages
- 23
- License
- MIT
- Good for
- Experimental WebGPU
OuteTTS-1.0 (0.6B)
Compact speech language model with one-shot voice references
- Size
- 582 MB
- Languages
- 12
- License
- Apache-2.0
- Good for
- Experimental pipeline
Orpheus-3B
Large expressive speech model tracked as a future experiment
- Size
- 2230 MB
- Languages
- 1
- License
- Apache-2.0
- Good for
- Research reference
Models that need a separate desktop or server setup
Models already downloaded on this device are managed under installed models.