Qwen3-TTS (Alibaba)
Text-to-speech models: design a voice from a text description or clone one from a short clip, 10 languages.
Voiceover service for ads and videos.
Apache-2.0 model. Needs a GPU for speed.
Apache-2.0 for the code (LICENSE file read); model weights also declared apache-2.0 on the Hugging Face model card checked (Qwen/Qwen3-TTS-12Hz-1.7B-Base)
Yes: code and the checked weights are both Apache-2.0, so hosting it for a paying client is permitted; the practical restriction is voice-cloning consent law, which neither the README nor the model card addresses.
Not archived, but last push was 2026-03-17 (about 7 months before 2026-10-07) - looks like a release repo rather than an actively developed one. No official website on the GitHub repo. Needs an NVIDIA GPU in practice (examples use device_map='cuda:0'; FlashAttention 2 'recommended to reduce GPU memory usage'); exact VRAM requirement: not found in README or model card. Models are 0.6B and 1.7B parameters. Does 3-second voice cloning from a reference clip - get written consent from the voice owner; no consent/responsible-use guidance was found in the README or model card. Only one of the five model cards was opened for the weights licence. Alibaba API: Qwen3-TTS-VC listed for Singapore and China (Beijing) regions.
Advertisers and video producers
Voiceover service for ads and videos.
Alibaba Cloud Model Studio sells Qwen3-TTS as an API. The only price found: voice creation (voice cloning) 'Billed at $0.01/voice', with a Singapore-region free quota of 1,000 voices within 90 days of activation. Per-character speech-synthesis price: not found on the page opened.
ElevenLabs (inferred: the README only cites ElevenLabs once, in a benchmark table - it does not call itself an alternative). ElevenLabs entry paid plan: Starter, $6/month standard (page showed a $1 first-month promotion; $5/month on annual billing).