INDEPENDENT RESEARCH · 2026 · PAPER IN PREPARATION FOR DAFX27

Building
Lacquer.

Give Lacquer a finished stereo mix and it rebuilds clipped peaks, removes echo and excess reverb, checks the instrument balance, evens out level problems, and masters the result against the measured norms of released music or a reference track. It reports every decision with the reading behind it, and leaves alone whatever is already fine.

Animation of a damaged mix being restored: a sweep reveals the cleaned spectrogram while the list of decisions appears below it
One real run on a held-out clip damaged with room reverb and a 163 ms echo. Left of the white line is the restored audio, right of it the damaged input. The echo is found and inverted by DSP, the reverb is removed by the network, the vocal reverb and the dynamics are measured and judged normal, and loudness is set for release.
01

Each problem goes to the tool that suits it

Problems with an exact answer use DSP. Hard clipping leaves the samples under the ceiling intact, so the peaks are rebuilt as the sparsest spectrum consistent with them. A discrete echo is one spike in the cepstrum, read off and inverted with an exact filter. Problems that need a prior about what music sounds like use a network: a 51 M parameter band-split transformer predicts a complex mask on the stereo spectrogram and removes room reverb. Aesthetic choices such as tone, stereo image and dynamics are measured against the range of released music and left alone inside it.

The network never generates audio. It multiplies the input spectrogram by a mask that starts as all ones, so anything it does not touch passes through unchanged. There is no codec or vocoder in the path.

The nine decision stages drawn from one clip: each card shows the quantity the stage measures and the decision it reached, and two before-and-after spectrogram pairs mark the echo copy and the room tail
The nine stages drawn from one 4 s clip with room reverb and a 210 ms echo added: each card shows what the stage measures and what it decided. Below, the echo stage removes the copy of each hit 210 ms later, and the network removes the room's tail. Green cards are deterministic DSP; orange cards use the learned model.
02

Nine decisions, each with a reading

The pipeline runs the stages in order and returns a report. A stage that finds nothing outside its range does nothing, so a track that is already fine comes out unchanged apart from the loudness target.

  1. Clipping. A channel whose samples pile up at a flat ceiling on both polarities was hard-clipped. The peaks are rebuilt by consistent sparse reconstruction. No ceiling, no action.
  2. Echo. A delayed copy is a sharp peak in the cepstrum. The delay and echo path are read off and inverted. A delay on the tempo grid is a delay effect or a loop, so it is reported and left alone.
  3. Room reverb. The network’s proposed change is measured first. Clean productions read about −42 dB and 95% of released tracks are under −20 dB. Above the −18 dB gate the reverb is removed, scaled by how far above; below it the network is bypassed.
  4. Vocal reverb. The separated vocal gets an absolute wetness reading. Produced vocals span a wide range, so the vocal is reduced to the edge of that range only when it is clearly past it. Stems are remixed only when the vocal changed.
  5. Instrument balance. Each stem’s loudness relative to the mix is compared with its range in professional mixes. Readings outside it are reported. On request a vocal is moved to the edge, applied as a stem difference so untouched stems contribute nothing.
  6. Level. A small controller network predicts a gain trajectory for the whole track, applied as a fader move when it exceeds 3 dB. A steady track is left alone.
  7. Tone and stereo image. Third-octave bands and band-wise side level outside the genre’s range are moved to its edge, by at most 1.5 dB unless a reference track is given.
  8. Dynamics. Band crest factors and the peak-to-loudness ratio are compared with the genre’s range. Peaky mixes get 2:1 compression on the excess; mixes that are already dense are flagged and protected.
  9. Loudness and peaks. Gain to the delivery target, then a two-stage true-peak limiter inside a 3 dB budget. A target that needs more is not reached, and the report says by how much.
MEASURE, THEN DECIDEPick a stage, set the reading, read the decision.

3. Room reverb

Network as a meter

The restoration network's proposed change, measured on a few windows before anything is applied.

Normal range. Clean productions read about −42 dB; 95% of released tracks are under −20 dB. The gate is −18 dB.

DECISION

Bypass the network.

−30.0 dB is below the −18 dB gate. The proposed change is discarded and the mix passes through untouched, which is what keeps finished productions safe.

Room reverb on 24 MUSDB18-HQ test songs: 4.7 to 9.7 dB SI-SDR. 3 of 98 clean released tracks trip this gate.

03

What the numbers say

Held-out music with synthetic damage. SI-SDR is closeness to the clean track in dB; higher is better. The last two rows are error measures, where lower is better.

TESTDAMAGEDAFTER LACQUER
Echo, 79 clips (DSP stage alone, no training)8.8 dB25.2 dB
Hard clipping, 29 clips (DSP stage alone, no training)20.1 dB26.7 dB
Room reverb on the whole mix, 24 MUSDB18-HQ test songs4.7 dB9.7 dB
Reverb on the vocal only, same songs6.7 dB7.8 dB
Reverb from a plug-in style the model never trained on, 40 clips3.6 dB8.7 dB
Reverb from six real rooms the model never trained on, 40 clips3.1 dB5.7 dB
Level problems inside a track, 20 clips (envelope error)4.5 dB2.4 dB
Tonal fault with a reference track, 200 tracks (tone error)3.1 dB0.19 dB
A separate judge that never sees the clean reference (Audiobox Aesthetics, production quality on a 1 to 10 scale) rates reverberant clips 6.37, the restored versions 6.63, and the clean originals 6.66.

Against other tools on the same reverberant clips, classical WPE and four released community dereverb models land between −0.5 and 3.9 dB where Lacquer reaches 9.1 dB from an input of 3.2 dB. On real rooms none of them trained on, none of them improves the average; Lacquer gained 1.7 dB there, and 2.6 dB after 6,500 further training steps with simulated rooms added.

With one stage removed at a time, each carries its own fault, and on clean clips 26 of 30 pass through bit-identical. Restored clips keep their notes and get their rhythm back: onset-envelope correlation with the clean track goes from 0.85 to 0.96 under reverb plus echo. Every number has its experiment, data and caveats in the experiment log.

04

Inside the model

Sound becomes a spectrogram: a 46 ms window slid forward every 11.6 ms, both channels, magnitude and phase kept. Frequency is cut into 62 bands, two bins wide at the bottom and up to 129 bins wide at the top, and each band becomes a 256-number token per frame. Twelve layers alternate attention along time and attention across bands. A small network per band then turns each token into a complex gain for every bin it covers, and an inverse STFT returns audio.

Overview of the band-split transformer with tensor shapes at every stage, from a stereo clip to a complex mask
The band-split transformer, with tensor shapes from one real forward pass on a 4 s clip.

The most direct evidence of what the network learned is in its attention. The traced clip carries an echo 210 ms late, which is 18 frames. In layer 12 the time attention has a second stripe 18 frames below the diagonal, and its average weight against lag peaks at exactly 210 ms. Over 30 held-out clips with random delays from 90 to 500 ms, layer 6 puts its largest relative increase in attention within one frame of the true delay in 30 of 30 clips. The delay was never an input; the network finds the earlier copy of each sound and uses it.

Time attention maps with a stripe at the echo delay, attention weight against lag peaking at 210 ms, a per-head view, and band attention
Time attention in layer 12 locks onto the echo delay without being told it. One of the eight heads carries most of it.

The warm start mattered more than architecture or size. With the architecture and parameter count held fixed, starting from a public vocal dereverb checkpoint was worth 2.9 dB on reverb after 1,000 steps and 3.6 dB after 8,000, against the same 51 M network from random weights. From scratch, the large model did no better than a 7.7 M one or a spectrogram U-Net. The released vocal checkpoint on its own wrecks full mixes, which is why the fine-tune exists.

05

Mastering as measurement

Mastering here means what a mastering engineer does to a mix that is already good: measure it, compare it with released music of its kind, and change as little as the readings call for. The norms come from 103,838 released tracks across 14 genres and 150 unmastered professional mixes, 60 features each.

Loudness, dynamics, tone, stereo image and instrument balance of released music by genre, against unmastered professional mixes
Unmastered professional mixes have the same tone and stereo image as released music and 4 dB more peak-to-loudness ratio. Mastering a good mix is mostly a dynamics and loudness job.

Tone cannot be corrected blind, and the bound is measurable. A track’s own tone is 5.2 dB away from the average of all music, 4.6 dB from its genre and 4.3 dB from its nearest neighbours, while a typical tonal mistake is 2.3 dB. Population norms therefore take back only a few percent of a tonal fault, and no estimator tried does better. So blind tone moves are capped at 1.5 dB and anything larger is reported as a mix problem. A reference track lifts the cap: on 200 held-out tracks with a deliberate tilt and the original as reference, 0.19 dB of tone error remains, where Matchering 2.0 leaves 0.72 dB.

How far a track's tone sits from population, genre and neighbour averages, and the tone error left by blind and reference mastering
Natural variation between tracks is twice a typical fault, so norms recover 1 to 3% of it and a reference recovers 94%.
Distortion against gain movement for limiters driven to the same loudness and true peak
Fifty unmastered mixes driven to −9 LUFS under a −1 dBTP ceiling by each limiter. Lower and further left is better.

The limiter is where an open tool can beat the baselines. At equal loudness and equal true peak the two-stage limiter leaves 5 to 10 dB less distortion than Matchering’s with less gain movement, and matches ffmpeg’s limiter on distortion with 2.5 dB less gain movement. Against iZotope Ozone 9 on the same songs it leaves 5 to 8 dB less distortion at equal or lower gain movement, while Ozone changes the long-term spectrum least. These are signal measures. Which limiter sounds better is a listening-test question.

The same two seconds through a hard clipper, Matchering, Ozone 9 and the two-stage limiter: waveform at a peak, applied gain, and the distortion a smooth gain cannot explain
The same two seconds through four limiters: the waveform at a peak, the applied gain, and the distortion a smooth gain cannot explain.
06

What does not work yet

  • Unseen real rooms. On six rooms outside the training set the gain drops from 5.9 dB to 1.7 dB. The network learned its 270 training rooms much better than room reverb in general. Adding 4,000 simulated rooms to training raised it to 2.6 dB within 6,500 steps, and that run continues.
  • Blind tone correction has an information bound. Track-to-track variation (5.2 dB) is twice a typical fault (2.3 dB). Without a reference, the system moves at most 1.5 dB and reports the rest.
  • Clean music is the hardest test. Run on 98 untouched released tracks, the first version of the repair half changed every one of them: tempo-synced delays were removed as echoes, quiet vocal samples were boosted, and the level controller nudged everything. After recalibrating on that audit, 80 of 98 come out bit-identical and the rest are listed proposals, each with its reading and each switchable.
  • Noise and soft saturation are not improved. Undoing compression or limiting did not train, and a published de-limiter only partly transfers to a different limiter.
  • The careful thresholds have a cost. Light reverb, tempo-synced echoes and moderately wet vocals are left alone on purpose.
  • Everything above is synthetic damage on real music. No recordings damaged in the field have been evaluated yet.
07

Where this came from, and where it goes

My engineering honors thesis took the obvious route first: encode the track with a neural codec, predict clean tokens with a large network (a 1.08 B parameter U-Net over EnCodec tokens) and decode. Measured properly, it could not beat its own input, for three reasons that belong to the route and to no particular network: level is invisible in the tokens, the audio losses sit behind an argmax, and the decoder caps quality at the codec’s own reconstruction. Lacquer keeps the question and changes the method: a mask on the spectrogram, DSP wherever the answer is exact, and a measurement before every action.

A paper is in preparation for DAFx27, the 30th International Conference on Digital Audio Effects in Cremona, August 2027. The call is not published yet. The current draft is here as a PDF, eight pages in the DAFx format, dated October 2026 and not yet peer reviewed. Before submission the work needs a blind listening test (the MUSHRA-style tooling is built and needs 15 or more listeners), a comparison against SonicMaster, the closest published system, and recordings damaged in the field.