INDEPENDENT RESEARCH · 2026 · PAPER IN PREPARATION FOR TISMIR

Transposed
twins.

Melody models are scored on how well they predict held-out tunes. But a folk tune, a hymn or a chorale melody often exists in many versions, in different keys and with small changes. I built a transposition-invariant search for those versions, ran it over eight melody corpora and their standard splits, and measured what the leaks do to the numbers papers report.

Loading melodies…
01

How often a test melody already sits in the training set

A melody is reduced to the sequence of its pitch intervals and rhythm ratios, which does not change under transposition or tempo. Candidate pairs come from MinHash over short phrases; every pair is then checked exactly. A twin shares at least half its phrases, or twelve consecutive intervals, with another melody; a copy shares nine tenths of its phrases.

CorpusSplitTest with a twin in trainingTransposed copyNote
JSB Choralesofficial test split36%13%43% of test chorales share at least half their soprano with a training chorale
PDMX (public domain scores)random 80/10/1050%39%6.2% keep a transposed copy even inside PDMX's own deduplicated subset
Essen folksongsrandom 80/10/1013%8%variants of the same tune collected in different villages
Nottinghamrandom 80/10/104.4%1.5%the standard Boulanger split leaks 2.4%
Hooktheoryofficial test split0.7%0.3%song-level split; leaks are covers and misspelled artist names
MuseDataofficial split, top voice2.4%0%
POP909random 80/10/100.3%0.2%
Piano-midi.deofficial split, top voice0%0%

JSB Chorales, a benchmark in sequence-modeling papers for over a decade, is the starkest case. Bach harmonized many hymn tunes more than once, so 36% of its test chorales sing a soprano that is also sung in training, usually in the same key with a different harmonization. The polyphonic fingerprint does not see it (0% overlap); the melody does.

02

What the leak is worth: retrain without the twins

The clean way to price a leak is to remove it and retrain. Each model is trained twice: on the standard training split, and on the same split minus the 25 training chorales that twin a test chorale. The change on leaked test pieces, net of the change on clean ones, is what the leak was worth.

Model (soprano lines, 3 seeds)Leak benefit, nats per note
4-gram count model0.555
Transformer, 0.46M parameters0.011
Transformer, 1.8M parameters0.026
Transformer, 7.4M parameters0.060

Two things follow. On the standard test set a 4-gram count model, which copies phrases, scores 2.28 nats per note and beats every transformer (the best scores 2.37). On the clean test chorales the larger transformers beat it (2.51 against 2.83). The leak reverses the ranking. And among transformers the benefit grows with size, so the leak flatters exactly the larger models that papers report as improvements.

The same holds on the full four-voice benchmark, in the units published papers use: the leak is worth 0.055 nats per frame to a 0.11M-parameter model and 0.25 to a 1.26M one. Published models on this benchmark differ by similar amounts.

03

Do melody models copy their training data?

PDMX is public domain, so its melodies can be shown next to whatever a model produces. I split it by twin family, so no test melody has a relative in training, and trained the same transformers on two versions of the training pool: raw, with every duplicate kept, and deduplicated, one melody per family. Both versions also hold 360 synthetic “canary” melodies inserted between 1 and 32 times.

Up to 7.4M parameters, no model reproduces a canary from its opening, even one inserted 32 times, and no model reproduces a real training melody. The reason is the transposition augmentation almost every melody model uses: it shows each melody in a new key every pass. Switch it off and the same model clearly prefers the canaries it has seen, more strongly the more often it saw them.

Free samples do share long runs with training melodies, 22% of samples at 1.8M parameters, but those runs are repeated notes, trills and looped arpeggios that occur in thousands of scores. The longest real melodic passage any sample shares with a training tune is 15 to 20 notes. So the leak in section 02 inflates scores with no memorization you could detect by looking for copies. Pick a sample and hear its longest borrowed passage next to the melody it comes from.

Loading samples…
04

Limits

  • The Session and IrishMAN are not analysed: The Session's licence forbids processing its tunes with language-model tools.
  • Essen melodies are counted but never published here, because its licence forbids redistribution.
  • The audit measures melodic identity, not musical quality; a twin can be a deliberate variant rather than an error.