A Nepali fine-tune of Fun-CosyVoice3-0.5B · model card
The model has no voice of its own. Only the LLM was fine-tuned and the vocoder is frozen, so the voice you hear is whichever reference clip you give it. Record about 10 seconds of Nepali, or upload a WAV.
Example — loads the reference clip, its transcript and the text
Reference voice — record ~10 s of Nepali, or upload a WAV
Transcript of that clip (Nepali)
Tone
Trained on ~3.5 s clips; long text run together sounds rushed.
Notes
Pronunciation is the strength — 0.294 CER against a 0.307 human floor.
Pause structure is the weakness. The tone tag moves speaking rate and pitch range, but the model does not place breath pauses; the sentence split above is a workaround, not a fix.
Danda, zero-width joiners and long digit runs are normalised for you — without that the output is silently wrong.
Generation runs on your own ZeroGPU quota, not the owner's.
First run after idle pays a cold start while the model streams to GPU.