Quick Verdict
IndexTTS is a controllable zero-shot text-to-speech system released by Bilibili’s speech team. The only official code channel identified by the maintainers is index-tts/index-tts; the project explicitly disclaims unofficial sites and services. As of July 20, 2026, the repository was active and its main line focused on IndexTTS2. The maintainers also state that repository history was reset, so old local clones should be removed and cloned again rather than merged blindly.
IndexTTS2 conditions generation on a speaker reference while separating speaker identity from emotion. Emotion can come from another audio prompt, an eight-dimension vector, the synthesis text, or a separate natural-language description. Its mixed Chinese-character and pinyin representation can correct some polyphonic and uncommon pronunciations. Two caveats materially affect selection. First, the paper and README describe precise duration control, but the release notice still says that functionality is not enabled in the current release. Second, this repository is not Apache-2.0: its root uses the custom bilibili Model Use License Agreement for the published IndexTTS2 model, weights, and final code.
IndexTTS is best evaluated as a local, controllable Chinese-first synthesis system for teams with supported GPU infrastructure and mature rights review. It is not permission to reproduce an identifiable person’s voice. Every speaker prompt should be owned by the speaker or covered by explicit, documented authorization, and synthetic identity should be disclosed wherever listeners could otherwise be misled.
Best For
- Studios with performer consent that need Chinese or English narration drafts, character dialogue, or controlled emotional delivery.
- Production teams that want to keep speaker prompts and generated audio in a managed local environment.
- Researchers evaluating autoregressive zero-shot TTS, cross-lingual behavior, emotion separation, pinyin control, and duration modeling.
- Engineers able to maintain Git LFS,
uv, large checkpoints, CUDA 12.8 or newer, and a secured local WebUI or Python service.
For a fuller few-shot data-preparation and fine-tuning workbench, compare GPT-SoVITS. Teams prioritizing a managed commercial API over local operation can consider ElevenLabs, while still applying voice-consent controls.
Key Features
- Zero-shot speaker conditioning: synthesize new text from a speaker reference without training a separate model for every voice. Similarity remains dependent on recording quality, language, range, and script.
- Separated speaker and emotion prompts: preserve an authorized timbre while conditioning expression from a different emotional reference.
- Multiple emotion interfaces: use emotional audio, an eight-value emotion vector, automatic interpretation of the synthesis text, or a separate text description;
emo_alphacontrols influence. - Chinese pinyin control: valid pinyin annotations can correct selected polyphonic characters and uncommon readings. The official guide warns that arbitrary consonant-vowel combinations are not supported.
- Natural duration today, precise duration in the method: current generation supports natural autoregressive length. The published research proposes explicit token-length control, but the downloadable release does not yet enable it.
- Local WebUI and Python inference: official setup uses
uv, checkpoints from Hugging Face or ModelScope, a local WebUI, and theIndexTTS2Python interface. - Optional performance paths: FP16, flash-attention acceleration,
torch.compile, and DeepSpeed are available depending on installed extras. DeepSpeed can be slower on some systems, so benchmark rather than enabling every flag by default. - Legacy path: IndexTTS1 and 1.5 guidance remains available for teams maintaining an earlier pipeline, while IndexTTS2 is the current focus.
Use Cases
A responsible studio can record or license one clean speaker reference, keep that timbre fixed, and use separate authorized emotion prompts or descriptions to create variants for game dialogue, course narration, or advertising review. Pinyin annotations can improve names, places, and polyphonic characters, but a native-language reviewer should listen to every deliverable. Because exact duration control is not enabled in the current release, video workflows should budget for editing, time adjustment, alternate takes, or rerecording instead of promising frame-accurate dubbing.
Create a voice-asset register before production. It should capture the rights holder, authorization text, languages and channels allowed, validity period, revocation process, reference hash, checkpoint, operator, and output destination. Public audio should include clear AI-voice disclosure. Generated speech must not be used for unauthorized identity use, impersonation, fraud, or harassment, and it must not serve as an authentication factor.
For calling or interactive systems, determine recording-consent rules for the jurisdictions of all participants and disclose both the AI identity and recording before capture. Payment, identity, healthcare, legal commitments, disputes, and abuse signals need independent verification and a clear human takeover path. Do not provide workflows for extracting a voice from public videos or bypassing speaker verification; public audibility is not authorization.
Pricing
IndexTTS1, 1.5, and 2 code and checkpoints can be downloaded through official channels without a software subscription. Real cost includes NVIDIA GPU capacity, storage, recording rights, review labor, security, and maintenance. The current setup guide supports uv as the reliable dependency path and recommends CUDA Toolkit 12.8 or newer. DeepSpeed and flash-attention can require additional Windows work.
Licensing is a primary constraint. The current root license is Bilibili’s custom model agreement, not Apache-2.0. It grants a limited license for the published IndexTTS2 model, weights, and final code, but organizations whose products or affiliates exceeded 100 million monthly active users in the preceding month, or RMB 1 billion annual revenue in the preceding year, must request a separate written license. Distribution requires notices, the agreement, and downstream conditions; the license also contains model-improvement and high-risk-use restrictions. Third-party dependencies, speaker recordings, emotion prompts, and personality or performer rights remain separate. The official README directs commercial cooperation questions to [email protected].
Pros
- Speaker identity and emotion are controlled separately rather than being inseparable properties of one prompt.
- Audio, vector, inferred-text, and descriptive-text emotion modes offer practical creative control.
- Mixed Chinese-character and pinyin input can solve selected pronunciation problems.
- Official WebUI, Python API, ModelScope, and Hugging Face routes make local evaluation straightforward for a prepared engineering team.
- FP16 and optional acceleration paths allow teams to trade memory, speed, compatibility, and quality.
- The canonical repository is active and clearly identifies its current IndexTTS2 focus and legacy IndexTTS1 path.
Cons
- Precise duration control remains disabled in the current release despite being a headline research contribution.
- The officially supported setup centers on
uv, large downloads, and NVIDIA CUDA 12.8+, with extra Windows complexity for acceleration packages. - Most demonstrated strengths concentrate on Chinese and English; every target language and cross-lingual direction needs its own acceptance test.
- The custom license has scale thresholds, downstream obligations, and use restrictions that require legal review.
- Zero-shot speaker reproduction creates impersonation risk, while the project does not provide a complete consent registry, abuse service, watermarking program, or approval workflow.
- An open local WebUI is not production security; teams must add authentication, upload validation, encryption, isolation, retention controls, and auditability.
Alternatives
| Tool | Best fit | Main difference |
|---|---|---|
| GPT-SoVITS | Few-shot fine-tuning and multilingual TTS workflow | More integrated data tools; MIT code, though model and dependency terms still need review |
| CosyVoice | Multilingual and streaming speech research | Different model and deployment ecosystem with separate license evaluation |
| Fish Audio | Open models combined with a hosted platform | More platform-oriented, with different data and billing boundaries |
| ElevenLabs | Fast commercial voice API integration | Less operations work, recurring service cost, and third-party data processing |
| Applio | RVC singing and speech conversion | Transforms existing audio instead of following the same zero-shot text-to-speech route |
FAQ
Is IndexTTS2 precise duration control available now?
The official README still marks it as not enabled in the current release. Research and demos are not substitutes for acceptance testing of downloadable code, so plan production around natural-duration generation.
Is IndexTTS licensed under Apache-2.0?
No. The current official repository uses Bilibili’s custom Model Use License Agreement, and GitHub identifies it as a nonstandard license. Old descriptions claiming Apache-2.0 should not guide deployment.
Does the custom license permit commercial use?
It grants conditional rights but includes organization-scale thresholds, downstream-distribution duties, model-improvement restrictions, and high-risk provisions. Obtain legal review and contact the official team when the threshold or cooperation needs apply.
Can a public video be used as a speaker prompt?
Not merely because it is public. Obtain explicit permission that covers model conditioning, generation, language, channel, duration, and derivative artifacts, and provide revocation and complaint procedures.
Is an NVIDIA GPU required?
The current official setup and acceleration guidance center on CUDA 12.8 or newer, making NVIDIA the supported practical path. Memory and speed vary with precision and acceleration flags and must be measured on the intended device.
How can a team reduce abuse and leakage risk?
Isolate the WebUI, authenticate operators, validate files, record authorization and provenance, encrypt prompts and outputs, disclose synthetic identity, block high-risk impersonation, and require independent verification plus human review for sensitive actions.
Bottom Line
IndexTTS2 is compelling because it separates an authorized speaker identity from emotional direction and adds useful Chinese pronunciation controls. It should not be selected on the assumption that every paper feature, especially precise duration, is already released. Run current code on the target GPU, evaluate each language, pronunciation, latency, and memory requirement, and have legal counsel review the custom license and its thresholds. With explicit speaker consent, AI disclosure, secure local operation, output provenance, abuse controls, and human escalation, it can support controlled local narration without turning zero-shot synthesis into an impersonation workflow.