Remove voice from video, keep the music
Most tools strip the voice but hand you an audio file. Learn how to remove the voice from a video and get the music-only result back as a full video.
Search for a way to remove voice from a video and you will find two completely different problems hiding behind the same words. Sometimes people want to strip the music and keep the talking — interviewers cleaning up a noisy session, students isolating a lecture. But just as often, the problem is the opposite: the voice is the problem. A vlog where the talking buries a beautiful street performance. A travel clip where your commentary ruins the ambient sound of the market. A gaming recording where the team chat drowns out the soundtrack you spent an hour picking.
This guide is for the second group. You want the voice gone and the music (or ambient sound) kept — and, crucially, you want the result back as a video, not as a lone audio file you now have to stitch back onto the picture yourself.
What "removing the voice" actually means
Let's get the vocabulary straight, because tool marketing makes this murky. Every modern AI separator works on stems. A stem is one ingredient of a mix, isolated into its own track. For music with vocals, the two most common stems are:
- the vocal stem — everything that sounds like a singing or speaking voice
- the instrumental stem — everything else: drums, bass, guitars, synths, ambient noise
"Removing the voice" therefore means keeping the instrumental stem and discarding the vocal stem. In the karaoke world this instrumental version has a traditional name — it is essentially what the karaoke track has always been: the song with the lead line taken out. The difference is that karaoke historically required either official instrumental releases or manual mixing skills. AI separation now produces the same result from any ordinary video file, in one pass.
One important nuance: in a vlog or travel clip, "the music" is often not a clean studio mix. It is background music layered over footsteps, wind, traffic, and room tone. Good separators treat everything that is not a voice as the instrumental side, which is exactly what you want here — the atmosphere stays, the talking disappears.
Why most tools hand you an audio file
Here is the frustrating part that every video creator eventually discovers. The vast majority of vocal removers — including the big free websites — accept your video as input, run the separation, and then give you back two audio files. Your video goes in; a WAV of the voice and a WAV of the instrumental come out. The picture never comes back.
That workflow made sense when these tools were built for musicians, who live in audio editors anyway. But if your source material is a video and your destination is a video platform, those two files are only half the job. You now have to:
- open a video editor and import the original clip and the instrumental WAV
- mute the original audio, lay the instrumental track underneath, and align it to the frame
- export a re-encode — waiting through a full render and losing a little quality along the way
It is maybe twenty minutes of fiddly work per clip, and alignment errors are easy to introduce and annoying to notice. If you process several clips, the busywork multiplies. And if you simply want to re-upload a cleaned version before some deadline, that timeline work is pure friction between you and the finish line.
How to remove the voice and keep the video
This is the gap our tool was built to close. BGM Remover runs the same class of AI separation — but it returns the result as videos, with the original picture untouched. The workflow looks like this:
- Drop your file on the home page. MP4, MOV, MKV, WEBM are accepted, up to 500 MB and 30 minutes.
- Sign in and press separate. The audio track is extracted inside your browser with WebAssembly and sent for AI separation. The video itself never leaves your device.
- Download four files. You get a music-only video (the picture with only the instrumental kept — this is your voice-removed video), a voice-only video, plus both stems as standalone audio if you want to work in an editor.
The music-only video is the deliverable most "voice remover" sites simply cannot give you. And because the picture is never re-encoded — the video stream is copied as-is, only the audio track is swapped — the footage is pixel-identical to what you uploaded.
New accounts also get a one-time 5-minute free trial with no watermark, so you can test the whole pipeline on a short clip before deciding anything.
When you actually need this
Removing the voice while keeping the music shows up in more workflows than you might expect:
- Atmosphere-first vlogs. Your montage works because of the music bed and ambient sound; the talking was scratch narration you planned to replace anyway. Delete the voice, keep the vibe.
- Gaming recordings. Discord and team chat contaminate hundreds of hours of footage. Strip the voices, keep the game audio, and your highlights become publishable.
- Reaction-style videos. Keep the original soundtrack of a clip while you re-record commentary over it, without the old voice bleeding through.
- Footage handoffs. A client sends talking-head footage but only wants the b-roll with its music. Hand back a finished music-only video instead of a confusing stem package.
- Re-dubbing prep. Remove the original voice to create a clean bed for a new voiceover in another language — useful for localizing your own content.
What quality should you expect
Honesty matters here, because every AI separator has limits. Separation is a prediction problem: the model listens and decides which parts of the mix belong to a voice. With clean recordings — a voice over background music — the results are usually very good. With extreme cases, you may hear slight artifacts: a breath that leaked into the instrumental, or a touch of music in the voice stem. The widely accepted rule of thumb is to test with a short excerpt before committing a long project, which is exactly why the free trial exists.
Two practical boundaries to know: files are accepted up to 500 MB and 30 minutes, and processing happens on the audio only, so a 10-minute video takes roughly the same separation time whether it is 4K or 1080p — the picture is copied, not recomputed.
Frequently asked questions
Does the picture lose quality?
No. The video stream is never re-encoded. Only the audio track is replaced, using stream copy — the lossless way to swap a soundtrack. What you upload is what comes back, frame for frame.
Can I do the opposite — remove the music and keep the voice?
Yes, and you do not have to choose in advance. The separation produces both stems at once, so you get a voice-only video and a music-only video in the same run. See remove vocals from video for that direction.
Is it free?
New accounts get a one-time 5-minute free trial — enough to fully process one or two short clips and judge the quality. After the trial, minute packs start at 4.99 dollars for 50 minutes and the minutes never expire. There is no subscription and no watermark.
What input formats are supported?
MP4, MOV, MKV, and WEBM, up to 500 MB and 30 minutes per file. Since the video track is copied rather than decoded, unusual codecs inside those containers still work as long as the audio track is readable.
Ready to try it
Drop a clip on the home page, sign in, and one press gives you the voice-removed video with the music intact — no timeline, no alignment, no re-encode. Remove the voice from your video and compare both outputs side by side.