Transposed
twins.
Melody models are scored on how well they predict held-out tunes. But a folk tune, a hymn or a chorale melody often exists in many versions, in different keys and with small changes. I built a transposition-invariant search for those versions, ran it over eight melody corpora and their standard splits, and measured what the leaks do to the numbers papers report.
How often a test melody already sits in the training set
A melody is reduced to the sequence of its pitch intervals and rhythm ratios, which does not change under transposition or tempo. Candidate pairs come from MinHash over short phrases; every pair is then checked exactly. A twin shares at least half its phrases, or twelve consecutive intervals, with another melody; a copy shares nine tenths of its phrases.
| Corpus | Split | Test with a twin in training | Transposed copy | Note |
|---|---|---|---|---|
| JSB Chorales | official test split | 36% | 13% | 43% of test chorales share at least half their soprano with a training chorale |
| PDMX (public domain scores) | random 80/10/10 | 50% | 39% | 6.2% keep a transposed copy even inside PDMX's own deduplicated subset |
| Essen folksongs | random 80/10/10 | 13% | 8% | variants of the same tune collected in different villages |
| Nottingham | random 80/10/10 | 4.4% | 1.5% | the standard Boulanger split leaks 2.4% |
| Hooktheory | official test split | 0.7% | 0.3% | song-level split; leaks are covers and misspelled artist names |
| MuseData | official split, top voice | 2.4% | 0% | |
| POP909 | random 80/10/10 | 0.3% | 0.2% | |
| Piano-midi.de | official split, top voice | 0% | 0% |
JSB Chorales, a benchmark in sequence-modeling papers for over a decade, is the starkest case. Bach harmonized many hymn tunes more than once, so 36% of its test chorales sing a soprano that is also sung in training, usually in the same key with a different harmonization. The polyphonic fingerprint does not see it (0% overlap); the melody does.
What the leak is worth: retrain without the twins
The clean way to price a leak is to remove it and retrain. Each model is trained twice: on the standard training split, and on the same split minus the 25 training chorales that twin a test chorale. The change on leaked test pieces, net of the change on clean ones, is what the leak was worth.
| Model (soprano lines, 3 seeds) | Leak benefit, nats per note |
|---|---|
| 4-gram count model | 0.555 |
| Transformer, 0.46M parameters | 0.011 |
| Transformer, 1.8M parameters | 0.026 |
| Transformer, 7.4M parameters | 0.060 |
Two things follow. On the standard test set a 4-gram count model, which copies phrases, scores 2.28 nats per note and beats every transformer (the best scores 2.37). On the clean test chorales the larger transformers beat it (2.51 against 2.83). The leak reverses the ranking. And among transformers the benefit grows with size, so the leak flatters exactly the larger models that papers report as improvements.
The same holds on the full four-voice benchmark, in the units published papers use: the leak is worth 0.055 nats per frame to a 0.11M-parameter model and 0.25 to a 1.26M one. Published models on this benchmark differ by similar amounts.
Do melody models copy their training data?
PDMX is public domain, so its melodies can be shown next to whatever a model produces. I split it by twin family, so no test melody has a relative in training, and trained the same transformers on two versions of the training pool: raw, with every duplicate kept, and deduplicated, one melody per family. Both versions also hold 360 synthetic “canary” melodies inserted between 1 and 32 times.
Up to 7.4M parameters, no model reproduces a canary from its opening, even one inserted 32 times, and no model reproduces a real training melody. The reason is the transposition augmentation almost every melody model uses: it shows each melody in a new key every pass. Switch it off and the same model clearly prefers the canaries it has seen, more strongly the more often it saw them.
Free samples do share long runs with training melodies, 22% of samples at 1.8M parameters, but those runs are repeated notes, trills and looped arpeggios that occur in thousands of scores. The longest real melodic passage any sample shares with a training tune is 15 to 20 notes. So the leak in section 02 inflates scores with no memorization you could detect by looking for copies. Pick a sample and hear its longest borrowed passage next to the melody it comes from.
Limits
- The Session and IrishMAN are not analysed: The Session's licence forbids processing its tunes with language-model tools.
- Essen melodies are counted but never published here, because its licence forbids redistribution.
- The audit measures melodic identity, not musical quality; a twin can be a deliberate variant rather than an error.