Upload your audio, run Whisper to get timestamped lyrics, then download the JSON and refine it (fix wrong words, split/merge lines, correct timings) — paste it into Gemini or edit manually. Then go to Step 2 to generate the video.
Whisper model
Language
tiny ~15 s · small ~1 min · medium ~3 min (all on CPU)
Pinning the language (Hindi/Gujarati) improves accuracy
Fine-tuned models auto-selected when available (e.g. Gujarati + small)