Nepali Call Centre Voice — Experimental

A Nepali fine-tune of Fun-CosyVoice3-0.5B · model card

The model has no voice of its own. Only the LLM was fine-tuned and the vocoder is frozen, so the voice you hear is whichever reference clip you give it. Record about 10 seconds of Nepali, or upload a WAV.

Example — loads the reference clip, its transcript and the text
Reference voice — record ~10 s of Nepali, or upload a WAV Transcript of that clip (Nepali)
Mode
Tone

Trained on ~3.5 s clips; long text run together sounds rushed.

Notes

  • Pronunciation is the strength — 0.294 CER against a 0.307 human floor.
  • Pause structure is the weakness. The tone tag moves speaking rate and pitch range, but the model does not place breath pauses; the sentence split above is a workaround, not a fix.
  • Danda, zero-width joiners and long digit runs are normalised for you — without that the output is silently wrong.
  • Generation runs on your own ZeroGPU quota, not the owner's.
  • First run after idle pays a cold start while the model streams to GPU.