How Vocals, Drums and Guitars Are Recovered from a Finished Song: Music Source Restoration in 2026
A finished song may sound like a single continuous recording, but it began as a collection of individual musical sources: lead and backing vocals, drums, bass, guitars, keyboards, synthesizers and other instruments. Once these parts have been mixed, processed and mastered, separating them again becomes difficult because their frequencies overlap and studio effects change the original signals. Music Source Restoration, usually shortened to MSR, addresses a more ambitious version of this problem. Instead of merely extracting the vocal or drum sound that can still be heard in the master, MSR aims to estimate the cleaner, less processed source that existed before much of the mixing and mastering chain was applied. Research published in 2025 established MSR as a distinct audio task, while work presented around ICASSP 2026 has turned it into an active field with dedicated benchmarks, competing systems and publicly available research tools.
From a Finished Master to Separate Musical Sources
Traditional music source separation already makes it possible to take a stereo recording and estimate several individual parts. A separation system might receive one WAV or FLAC file and produce separate files for vocals, drums, bass and other instruments. This is useful, but there is an important limitation. The extracted vocal normally retains characteristics that were added during production. If the singer was compressed, equalised and covered with reverb before the final master was created, an ordinary separation model generally extracts a vocal containing at least some of those changes. It separates what exists in the finished recording rather than attempting to reconstruct the earlier studio source.
Music Source Restoration adds another objective. The system has to identify which sounds belong to each musical source while also estimating how those sounds may have differed before processing. Consider a guitar recorded through several effects. Its frequencies may have been shaped with EQ, its loud and quiet notes may have been levelled by compression, and reverb may have placed it inside an artificial acoustic space. The final master can then receive further limiting, stereo processing and lossy encoding. MSR treats these transformations as part of the restoration problem rather than assuming that the mastered mix is simply a clean sum of several instruments.
This distinction became formalised with research led by Yongyi Zang, Zheqi Dai, Mark D. Plumbley and Qiuqiang Kong. Their original MSR work introduced RawStems, an annotated collection covering 578 songs and 354.13 hours of unprocessed source material, organised into eight main and 17 secondary instrument groups. The purpose was to give models examples of musical sources before the type of processing normally heard in a commercial master. By 2026, the field had progressed to a dedicated ICASSP challenge and a more structured evaluation process, making it possible to compare different methods on the same restoration problem rather than relying on demonstrations chosen by individual research teams.
What the System Listens for in Vocals, Drums and Guitars
A vocal has a recognisable combination of pitch movement, formants, consonants, breaths and phrasing. Even when a singer overlaps with guitars, keyboards or cymbals, these patterns provide clues that a separation model can use. The difficult part is that processing changes those clues. Compression can alter the relationship between quiet syllables and strong notes, reverb produces delayed copies of the voice, and EQ can remove or emphasise parts of its frequency range. An MSR system therefore does more than locate vocal energy. It also uses patterns learned from unprocessed recordings to estimate how a cleaner version of that vocal could sound.
Drums present a different problem. A kick drum produces a strong low-frequency impact, a snare contains a short transient followed by a noisier decay, and cymbals can occupy a broad high-frequency range. These sounds are often easier to recognise than to reconstruct accurately. Heavy bus compression can reshape drum transients, saturation can alter their harmonic content, and room or artificial reverb can spread a single hit across a much longer period. Percussion is especially challenging because short sounds can be masked by other instruments and may share similar frequency ranges with hi-hats, cymbals, shakers or effects.
Guitars vary even more widely. An acoustic guitar, a clean electric guitar and a heavily distorted electric guitar may have very different acoustic characteristics even though all belong to the same broad instrument family. Chords also occupy several frequencies simultaneously and can overlap with vocals, keyboards and synthesizers. A restoration model therefore has to use both the sound at a particular moment and the way musical information develops over time. It is not simply searching for a fixed frequency band labelled “guitar”. It is estimating which changing patterns are most consistent with guitar performance and then reconstructing that source while limiting material that belongs to other instruments.
Why Music Source Restoration Goes Beyond Ordinary Stem Separation
The simplest way to understand the difference is to imagine an old finished mix containing a vocal with strong reverb. Conventional stem separation may produce a useful isolated vocal, but that vocal can still contain the reverberant tail heard on the record. MSR attempts to go further by producing an estimate closer to the vocal before that effect was added. The same principle applies to drums shaped by compression, guitars altered by tonal processing and instruments affected by mastering. This does not mean that every studio decision can be perfectly undone. The system is estimating missing information from patterns learned during training, not travelling backwards through an exact record of the original production session.
Many of the strongest 2026 approaches divide the task into stages. First, the finished recording is separated into approximate musical sources. A second restoration stage then works on those estimates and attempts to correct the effects of production processing or other degradation. One system presented in 2026 used a BandSplit-RoFormer separator to estimate eight instrument stems and an additional “other” output, followed by specialised restoration models. Another competition system combined several separation models before applying targeted reconstruction to individual sources. The architectures differ, but both approaches reflect the same practical idea: identifying an instrument and restoring its earlier sound are related yet separate problems.
Researchers are also testing different orders for those operations. Some systems separate first and restore afterwards, while other work has investigated enhancing the mixed signal before performing separation. This matters because damage introduced before the model receives the recording can make instrument recognition harder. A heavily compressed or poorly encoded file may contain fewer useful details than a high-quality master. There is therefore no single universal MSR recipe in 2026. Current research is comparing several methods to determine which order and combination of separation, enhancement and reconstruction gives the most reliable results across different instruments and recording conditions.
What the 2026 Research Results Show in Practice
One important step towards realistic testing is MSRBench. The benchmark was created specifically to compare restored sources with their corresponding unprocessed references. It covers eight broad classes: vocals, guitar, keyboard, synthesizer, bass, drums, percussion and orchestra. Professional mixing engineers created mixtures while keeping access to the earlier source material, allowing researchers to measure how closely a restored output resembles an actual reference rather than judging only whether an instrument sounds isolated. The benchmark also includes twelve types of additional degradation representing problems such as acoustic interference, analogue-style artefacts and lossy audio coding.
The inaugural MSR Challenge reported in 2026 illustrates both the progress and the remaining difficulty. Five teams participated. According to the official challenge summary, the leading system reached 4.46 dB on the Multi-Mel-SNR objective measure and received a 3.47 overall mean opinion score in subjective evaluation. Performance was far from equal across instruments. Across submitted systems, bass averaged 4.59 dB on the reported Multi-Mel-SNR measure, while percussion averaged only 0.29 dB. This difference is useful because it shows that a single headline score does not describe restoration quality equally well for every type of musical material.
Those results also explain why MSR should not be treated as a finished replacement for original multitrack sessions. A convincing bass estimate does not mean that every cymbal hit, guitar resonance or backing-vocal detail can be reconstructed with the same accuracy. Some musical information becomes obscured when several sounds occupy the same frequencies, while other details may be permanently reduced by aggressive processing or lossy encoding. The technology is progressing quickly, but the strongest evidence available in 2026 still describes an active research problem. Listening tests and reference-based measurements remain important because a stem that appears correctly separated can still contain unnatural texture, missing transients or traces of neighbouring instruments.

Where Restored Music Stems Can Be Useful
One of the most practical uses for MSR is work with recordings whose original multitrack sessions are unavailable. An archive may possess a stereo master but not the separate tapes or digital session files from which it was created. A restoration system could provide estimated vocal, drum, guitar and other stems that allow engineers to inspect specific parts of the recording independently. This can support repair work, new mixes or research into older production techniques. The important distinction is that these files remain reconstructions. When an original multitrack exists, it is still the more reliable source because it contains information that no model has to infer.
Separate reconstructed sources can also help musicologists, teachers and musicians analyse arrangements. A student studying rhythm can focus on the drum part, while someone transcribing harmony may benefit from reducing vocals and percussion. Producers can use estimated stems as references when examining how instruments interact inside a dense mix. Similar techniques may assist with accessibility, rehearsal material and certain forms of music transcription. These uses do not require the reconstructed stem to be identical to the studio recording; they require enough separation and musical consistency for a person to hear a part more clearly than in the full master.
There are also potential production uses, particularly when only a finished recording survives. An engineer may want to adjust the balance between a voice and backing instruments, reduce an unwanted element or prepare material for a new authorised arrangement. MSR could provide cleaner starting material than ordinary stem extraction when the production effects embedded in the master interfere with the intended work. However, professional use requires careful listening. Restored stems may contain small artefacts that become more obvious after further EQ, compression or gain changes, so an output that sounds acceptable on its own may still need manual treatment before it can be used in a new mix.
Limits, Rights and Realistic Expectations
The central limitation of Music Source Restoration is that the finished song does not contain a complete hidden copy of every original track. Mixing combines signals, and some production operations remove or transform information. Limiting can flatten peaks, filtering can reduce frequencies, lossy codecs can discard data and several instruments can mask one another. A model can make an informed estimate based on patterns learned from other recordings, but it cannot prove that the reconstructed waveform is exactly what came from the microphone or instrument output years earlier. Calling the result an estimated restored stem is therefore more accurate than describing it as the recovered original studio file.
Typical errors include leakage, where part of one instrument remains audible in another stem; missing detail, where the model removes material that actually belongs to the target instrument; and generated texture that sounds plausible but does not precisely match the source. Reverb is particularly difficult because its reflections overlap with later sounds, while distorted guitars and dense percussion can occupy wide frequency ranges shared by other instruments. Results also depend on the recording supplied to the system. A clean lossless master normally preserves more useful evidence than a low-bitrate copy that has already passed through additional encoding or acoustic recording.
Technical ability also has to be separated from permission to use a recording. Creating stems from copyrighted music does not give the user ownership of the composition, master recording or performances contained in it. Commercial remixing, redistribution, model training, sampling and publication can involve rights that vary according to jurisdiction, licence and intended use. For archival or professional production work, provenance and authorisation should therefore be checked independently of the software used to create the stems. In 2026, Music Source Restoration is best understood as an increasingly capable audio reconstruction method: it can provide useful estimates of vocals, drums, guitars and other sources from a finished song, but the results still require technical judgement, listening checks and realistic expectations about information that may already have been lost.