Most open-source text-to-speech guides rank models by how they sound. That is the hardest thing to verify and the fastest thing to go stale. This one checks two things you can confirm yourself from the links below: what each licence actually permits, and whether anyone still ships code.
Both checks turned something up. Six of the seventeen are not open source in the sense the phrase implies, including the fastest-growing text-to-speech project we track. And the second-most-starred project here has been dead for two and a half years while a maintained fork carries on somewhere almost nobody links to.
How these seventeen were chosen: ten text-to-speech projects from Breakwave's board, five widely cited ones we do not track, and two successor projects this check turned up. It is not a ranking and it is not exhaustive. Our board also carries MOSS-TTS, supertonic, mlx-audio, sherpa-onnx and espeak-ng, which are not compared here.
Six of these are not open source, in two different ways
Two ship under licences that restrict who may use them and for what. Four more have freely licensed code and restricted weights, which is the trap that catches people, because the badge on a repository page describes only the code.
The licence that calls itself open source
index-tts has 22.7k stars and is
the fastest-growing text-to-speech project on our board. GitHub's licence
detector reports NOASSERTION, meaning it could not match the file to a
licence it recognises. That reads as a blank, so it is easy to skip.
The file is the bilibili Model Use License Agreement, and §1.4 defines the licensed "Model" as both the weights and the final code, so it governs the whole repository. Worth reading before you build:
- §2.2 requires a separate written licence if your products or services, or an affiliate's, had over 100 million monthly active users last month or over RMB 1 billion in revenue last year.
- §3.4(c) forbids using it to improve any other AI model, except its own derivatives or non-commercial models.
- §4.2 restricts deployment in high-risk scenarios "such as" medical diagnosis, autonomous driving, military applications, critical infrastructure and automated decision-making. The list is illustrative rather than closed. The clause both prohibits ("you must ensure ... are not deployed") and shifts liability to you if you deploy anyway.
Section 2.3 calls this "an open-source license". Under the Open Source Definition it is not one. Clause 6 forbids restricting a field of endeavour, which §4.2 does, and clause 5 forbids discriminating against persons or groups, which §2.2 does by licensee size.
For most readers this is still usable, and it is a licence you have to read
rather than assume. The repository also ships a separate DISCLAIMER file
limiting use to research, study and lawful creative applications, and
barring unauthorised commercial use of synthesised voices. Note the licence
names the model "bilibili indextts2" while the repository hosts more than
one release, so check which one you are pulling.
The one that blocks commercial use outright
fish-speech has 32.2k stars,
more than index-tts, and also reports NOASSERTION. Its
Fish Audio Research License
requires "a separate written license agreement from Fish Audio" for any
commercial purpose.
On the practical question of whether you can use something at work, both
restrict you, in different ways. fish-speech's licence blocks commercial use
at any size unless you negotiate one. index-tts's licence charges you nothing
below bilibili's thresholds, but its separate DISCLAIMER limits use to
research, study and lawful creative work, and bars unauthorised commercial
use of synthesised voices. Neither is a licence you can adopt unread.
The code is free, the weights are not
coqui-ai/TTS is MPL-2.0, which permits commercial use, though it is weak copyleft rather than permissive: modifications to covered files must be released under the MPL when you distribute them. The weights are a separate question. XTTS-v2 ships under the Coqui Public Model License, whose opening lines say it "allows only non-commercial use of a machine learning model and its outputs". Read the badge and you would ship it. Read the model card and you would not.
ChatTTS splits the same way: code under
AGPLv3+, model under CC BY-NC 4.0, which its README limits to educational
and research use. The maintained coqui fork, idiap/coqui-ai-TTS, inherits
the same split from its parent.
F5-TTS is the one most likely to catch
you out, because its repository badge says MIT and nothing on the page
contradicts it. Its README's own licence section does: "Our code is released
under MIT License. The pre-trained models are licensed under the CC-BY-NC
license due to the training data Emilia." The
model card confirms cc-by-nc-4.0.
The code is genuinely MIT; the thing you actually run is not.
CosyVoice is a counter-example.
Apache-2.0 on the repository, and its
model card states
apache-2.0 too, so both halves agree. It has moved org, from
FunAudioLLM to QwenAudio.
Some licences also restrict cloning particular voices independently of the
general terms. index-tts's DISCLAIMER is an example, barring synthesis of
public figures without authorisation.
Is anyone still shipping?
Last commit is the committer date on each project's default branch, read
from the GitHub API on 12 August 2026. The default branch is often not
main: it is dev for coqui-ai/TTS and idiap/coqui-ai-TTS, develop for
PaddleSpeech, and
master for espnet, Kokoro-FastAPI, chatterbox and piper.
| Project | Stars | Last commit | Licence (code / weights) | Status |
|---|---|---|---|---|
| index-tts | 22.7k | 2026-08-12 | bilibili MULA (both) | Active |
| Kokoro-FastAPI | 5.3k | 2026-08-12 | Apache-2.0 | Active (wrapper) |
| espnet | 9.9k | 2026-08-11 | Apache-2.0 | Active (toolkit) |
| piper1-gpl | 5.1k | 2026-08-09 | GPL-3.0 | Active (successor to piper) |
| F5-TTS | 15.1k | 2026-07-23 | MIT / CC-BY-NC | Active |
| GPT-SoVITS | 60.8k | 2026-07-22 | MIT (both) | Active |
| chatterbox | 26.0k | 2026-07-21 | MIT | Active |
| PaddleSpeech | 12.7k | 2026-06-12 | Apache-2.0 | Active (toolkit) |
| idiap/coqui-ai-TTS | 2.3k | 2026-06-10 | MPL-2.0 / CPML | Active (fork of coqui) |
| fish-speech | 32.2k | 2026-06-09 | Fish Audio Research | Active |
| CosyVoice | 22.7k | 2026-05-25 | Apache-2.0 (both) | Active |
| ChatTTS | 39.8k | 2026-04-10 | AGPLv3+ / CC BY-NC | Slowing |
| rhasspy/piper | 11.3k | 2025-08-26 | MIT | Archived, moved |
| hexgrad/kokoro | 8.4k | 2025-08-06 | Apache-2.0 | Stale ~12 months |
| WhisperSpeech | 4.6k | 2025-06-08 | MIT | Stale ~14 months |
| MeloTTS | 7.6k | 2024-12-24 | MIT | Stale ~20 months |
| coqui-ai/TTS | 45.9k | 2024-02-10 | MPL-2.0 / CPML | Dormant, forked |
Reading the licence column. A slash means the weights ship under a different licence from the code and we opened both. A cell marked (both) means one licence covers code and weights, and we checked that too. Everywhere else we checked the repository licence only, so if you are shipping commercially, open the model card yourself before relying on it.
Four have gone more than a year without a commit: coqui-ai/TTS, MeloTTS,
WhisperSpeech and hexgrad/kokoro. rhasspy/piper is close at 351 days and
is a special case, below.
Falsifiable: every date above is the default branch's HEAD, and every licence in that column is a file you can open from the link beside it. If coqui-ai/TTS ships a commit tomorrow, its row is wrong.
Piper moved, it did not die
The archived or dormant rows are the ones guides get wrong, and both of the big ones have somewhere else to go.
rhasspy/piper is archived, and a
number of guides list it as abandoned. Its README is one line:
"Development has moved:
https://github.com/OHF-Voice/piper1-gpl".
The successor installs as pip install piper-tts and describes itself as
"a fast and local neural text-to-speech engine that embeds espeak-ng for
phonemization". Nothing was relicensed: the archived MIT code stays MIT, and
the successor is GPL-3.0 because it embeds espeak-ng, which is GPL. If you
ship proprietary software, that change matters.
Coqui is dormant; its fork is not
coqui-ai/TTS last shipped in February
2024, and Coqui, the company behind it, shut down around then. Its README
says none of this and still advertises XTTS-v2 as current. The work
continues in idiap/coqui-ai-TTS, a
maintained fork from Idiap Research Institute whose last commit was June
2026 and which publishes the coqui-tts package on PyPI. Its most recent
PyPI release is older than that, so check the tag you are installing.
So the original is genuinely dormant and the project is not. If you want XTTS-v2, the fork is the live route, and its weights carry the same non-commercial licence as the parent's.
Kokoro: the package we track is a wrapper, not the model
Worth untangling, because the two are easy to confuse. hexgrad/kokoro is the model, Apache-2.0 and 8.4k stars, and it has not had a commit in about twelve months. Kokoro-FastAPI is a Docker wrapper around it, also Apache-2.0, 5.3k stars, and it shipped on the day we measured. A wrapper being actively maintained tells you nothing about the model inside it. Our board tracks the wrapper and not the model, which is a gap on our side.
GPT-SoVITS is the biggest here, and MIT on both halves
GPT-SoVITS has 60.8k stars, more than any other project in the table and a third more than the dormant coqui-ai/TTS. Six projects here carry an MIT repository licence, so that alone distinguishes nothing. What does is that its weights are MIT as well, which we checked. Along with CosyVoice it is one of only two projects on this page where we opened both halves and found no carve-out. It shipped in July 2026.
Chatterbox is the plain one
chatterbox is MIT from Resemble AI, a commercial voice company, and it shipped in July 2026. Its model card is MIT too, so there is no weights carve-out. A commercial vendor releasing more freely than several of the community projects on this page is worth noticing.
F5-TTS, which sits next to it in the table on the same MIT badge and the same July 2026 commit, is the counter-case described above: its weights are not MIT. It also came out of a paper, arXiv:2410.06885, if you want the method.
A trap in the data everyone uses
If you automate a staleness check, do not use the GitHub API's pushed_at
field. It records a push to any branch, including ones nobody merged, so it
overstates how maintained a project looks. Across these projects it
disagrees with the real default-branch commit date on four, and for
coqui-ai/TTS it reads 2024-08-16 against a true last commit of 2024-02-10,
six months adrift. Read the default branch's HEAD instead.
Star count is a poor guide on its own, and it fails hardest on the project it matters most for: coqui-ai/TTS is the second-highest-starred entry here and the least actively developed.
Then the ordinary questions
Streaming. The split that decides most real applications. Batch models generate a whole utterance before you hear anything, which suits narration and anything pre-rendered; streaming models emit audio as they generate, so time to first audio governs how a conversation feels.
Seven document it: CosyVoice claims bi-directional streaming at around 150ms latency, coqui-ai/TTS and its idiap fork claim XTTS streaming under 200ms, and fish-speech, ChatTTS, PaddleSpeech and Kokoro-FastAPI all document streaming output. espnet is the one to watch out for: its README's streaming claims are all for recognition and enhancement, never synthesis.
The others do not mention streaming in their README, which is not the same as not supporting it. These are documentation checks, not tests.
Hardware. Ask what a model needs for your workload rather than what the demo used, and work out cost per hour of audio before committing to self-hosting. A hosted API is sometimes cheaper than the GPU time.
Languages. Coverage varies enormously. XTTS-v2's model card lists seventeen languages, while Meta's MMS work reports 1,107. Outside the majors, verify quality yourself.
Cloning. Test with your own reference audio. Demos use clean studio recordings and your source material probably is not.
If you just want the shortlist
Apply the article's own two checks to the table and most of it falls away. Commercially usable, and committed within the last twelve months:
| Project | Licence | Streaming documented | Note |
|---|---|---|---|
| CosyVoice | Apache-2.0, both halves | Yes, ~150ms claimed | Code and weights both checked, both Apache-2.0 |
| chatterbox | MIT, both halves | Not documented | Code and weights both checked |
| GPT-SoVITS | MIT, both halves | Not documented | Most-starred here at 60.8k |
| Kokoro-FastAPI | Apache-2.0 | Yes | A wrapper; the model it wraps is 12 months stale |
| PaddleSpeech | Apache-2.0 | Yes, TTS and ASR | A toolkit rather than one model |
| espnet | Apache-2.0 | ASR only | A toolkit rather than one model |
| piper1-gpl | GPL-3.0 | Not documented | GPL blocks proprietary distribution |
Ruled out on licence: fish-speech and ChatTTS for commercial use, index-tts
because its DISCLAIMER limits use to research, study and creative work,
coqui-ai/TTS and its idiap fork because XTTS-v2's weights are
non-commercial, and F5-TTS because its weights are CC-BY-NC despite the MIT
badge on the repository.
Ruled out on maintenance: hexgrad/kokoro, WhisperSpeech and MeloTTS.
rhasspy/piper is inside twelve months at 351 days but archived, and its
successor piper1-gpl is in the list above.
Weights were opened for CosyVoice, chatterbox, GPT-SoVITS, F5-TTS, index-tts, ChatTTS and the two coqui repositories. For the rest of this shortlist we checked the repository licence and nothing else, so "commercially usable" means the code is, and the weights are your check to make.
This is a sort of the table above, not a quality judgement. We have not listened to any of them.
What we have not measured
We have not listened to any of these models, or installed any of them. No audio quality ranking, no latency benchmarks, no VRAM figures, and no claim about whether a stale project still runs on a current Python or CUDA. Everything above is licence text and commit history.
We also do not track all of these. Seven are absent from our board: chatterbox, piper, hexgrad/kokoro, MeloTTS, WhisperSpeech, and both successor projects, piper1-gpl and idiap/coqui-ai-TTS. For those seven the commit dates and licences above come straight from GitHub rather than from our own tracking. The Kokoro gap described earlier, where we track the wrapper and not the model, is the one worth fixing on our side.
For the projects we do track, the live voice AI ranking shows which are gaining attention fastest right now. It is a composite of growth signals and it will not tell you whether a project is maintained, which is why this page uses commit dates for that.
The part most guides leave out
Synthetic voice carries obligations other model categories do not.
Being open source does not exempt you. The EU AI Act's Article 2(12) carves free and open-source AI systems out of most of the Act, but not out of Article 50. If you ship synthetic speech in the EU, the transparency duties apply whatever your licence says.
Those duties, under Article 50, have applied since 2 August 2026, and two of them are worth separating. Article 50(2) requires providers of systems generating synthetic audio to mark outputs in a machine-readable format so they are detectable as artificially generated. Article 50(4) requires deployers to disclose content constituting a deep fake. The provider marking duty is the broader of the two. Work out which applies to you.
Beyond the law, two things are worth doing anyway. Cloning a real person's voice without permission is a problem regardless of what a licence allows, and that includes colleagues who agreed to a recording for something else. And logging or watermarking what you generate means that if your system is ever misused, you can show what it did and did not produce.
None of this is legal advice, and if you are deploying voice cloning at scale you should get some.
Live rankings: voice AI. How we research and correct these pages: editorial standards.