Drop in an audio file and its lyrics as plain text. Three models will
time every line against the isolated vocal, score each one, and hand you
only the handful that need an ear.
Add a track. An audio file and a .txt — one line
per lyric line, a blank line between sections.
Wait about five minutes. The vocal is separated once and
cached; the progress log runs while it works.
Check the timings. Every line gets a score, and the words two
models disagree about are queued for you to settle by ear.
Everything runs on this machine. Nothing is uploaded, and the models
download once on the first run.