Vocal Remover Beta
Separate vocals from the backing track using AI stem separation.
How to vocal remover
- Add a music track. The model downloads on first use — this is a one-time wait of around 30–60 seconds on a typical connection.
- Choose whether you want the instrumental, the vocal stem, or both.
- Wait for processing. A three-minute song takes two to four minutes. Download the result.
About this tool
Audio source separation — pulling individual instruments or vocals out of a mixed recording — was computationally intractable outside research labs until transformer-based models demonstrated that the problem was learnable. The models are still large, still slow, and still imperfect, but for the most common use case (separating a clear lead vocal from an instrumental backing) they produce results that are genuinely useful.
The separation artefacts are worth understanding before you start. The model works by predicting which frequency content belongs to the vocal mask and which belongs to the instrumental. On a well-produced track with clear spectral separation between voice and instruments, this works well. On a dense mix where the instruments occupy the same frequency range as the voice, the model has to make harder guesses and makes more mistakes. The "both stems" option lets you check both outputs and decide which is useful for your purpose.