Browser-local vs cloud vocal separation
Local in-browser models are free and private but capped. Cloud separation is fast and consistent but per-minute. An honest look at where each one wins.
A quiet shift happened in the last two years: vocal separation stopped being something you install. A growing family of tools now runs the entire AI model inside your browser — the model file downloads once, the separation runs on your own CPU or GPU, and nothing is ever uploaded. Meanwhile, cloud-based services kept getting faster and more accurate, at the cost of a per-minute or per-pack fee.
If you are choosing between the two architectures for your next project, this guide walks through the honest trade-offs — including the ones neither side likes to advertise. The open-source ecosystem behind browser inference is real and maturing fast; projects like Facebook's Demucs have published models that the whole industry builds on, and quantized versions of them now fit in a browser tab.
What a local model actually is
When a site says "AI runs locally in your browser," here is the mechanics behind it:
- On your first visit, the tool downloads a compressed neural network file — typically 30 to 80 MB depending on the model generation.
- The file is cached by your browser, so it downloads once.
- When you drop a file, the page runs the model in a WebAssembly or WebGPU runtime, using your processor.
- The result never leaves your machine.
The genuine advantages are real: it is free to operate at any volume (your hardware pays the bill), files with privacy constraints never travel, and it works offline. For a musician with a capable laptop, this is a legitimate setup.
What the model size quietly tells you
Here is the detail that product pages gloss over: the model file size is a proxy for its capability. State-of-the-art separation models are large — the full-quality versions of popular architectures run well over 100 MB uncompressed. Browser-local tools compress heavily to fit a reasonable download, usually landing at a two-stem (voice vs. everything else) model of moderate quality.
The practical consequences show up on hard material:
- clean spoken-word recordings over light background music: local models do genuinely well
- dense mixes where vocals and instruments share frequencies: small models leave audible residue
- live recordings, heavy effects, layered sound design: this is where the quality ceiling becomes visible
Reputable local tools say this themselves — their FAQs warn about residue on demanding mixes. It is not a flaw of any particular product; it is physics. A 40 MB model cannot carry the same information as a datacenter-grade one.
What the cloud side does better
Cloud separation inverts every trade-off:
- Model quality. Datacenter models are larger, newer, and trained harder. On difficult material — the footage where you actually need help — the difference is usually audible.
- Consistency. The same job produces the same result whether you are on a gaming PC, an office laptop, or a phone. Browser inference, by contrast, ranges from smooth (modern Chrome with a discrete GPU) to painful (older hardware, Safari's WebGPU gaps).
- Duration and size limits. Local tools commonly cap at around 10 minutes per file, because a user's browser tab is a hostile place for long, heavy computations. Cloud pipelines handle 30-minute files routinely.
- Deliverables. Cloud workflows can afford extra post-processing steps — like remuxing the separated audio back onto your original video stream — that are awkward to do reliably inside a browser session.
What the cloud side will not tell you
Fairness requires the other column too:
- Recurring cost. Cloud quality is paid per minute. If you process a lot, the bill is real — which is exactly why per-minute pricing should come with a meaningful free trial rather than a demo.
- Upload. Your audio travels to a server. For published or sensitive footage, check that the tool extracts audio locally and sends only the small audio track — or that audio is deleted on a fixed schedule.
- Dependency. A cloud tool that shuts down takes its quality with it. Local files keep working forever.
So which one should you use
A decision order that works in practice:
- Is the material short, clean, and non-commercial? A local browser tool will do the job free, and the quality will satisfy you.
- Is the footage long, dense, or destined for publication? Cloud separation pays for itself in quality and time — a 30-minute video processed in minutes beats an hour of your laptop grinding through the same task.
- Does the deliverable need to be a video? Check which tools return video with the original picture; most local tools return audio only.
- Is privacy a hard requirement? Both architectures can honor it — local by design, cloud by choosing a service that extracts audio in-browser and deletes it quickly.
Many creators end up using both: a local tool for quick throwaway jobs, a cloud service for the work that ships. That is not indecision — the two architectures genuinely optimize for different moments.
Frequently asked questions
Is browser-local separation actually good?
For clean, short, spoken-word material — yes, surprisingly good. The limitation is the ceiling: compressed two-stem models struggle with dense or live material, and there is no room in a 40 MB download for the nuance a larger model captures.
Why do local tools limit files to 10 minutes?
Long files mean long inference inside a browser tab: memory pressure, throttled background tabs, and users who close the window mid-job. The cap protects the experience. Cloud pipelines do not have this constraint.
Does cloud separation damage the video picture?
Not if the service remuxes instead of re-encoding. Good cloud tools copy the video stream as-is and swap only the audio track — the picture is pixel-identical to the upload.
Is my audio safe in the cloud?
Depends on the service. The privacy-conscious pattern is: extract the audio track in your browser, upload only that small track, process it, and delete it on a fixed schedule (24 hours, for example) — while the video never leaves your device. Verify this before uploading unreleased material.
The honest bottom line
Local models won the argument on privacy, price, and convenience for light jobs. Cloud models win on quality, duration, consistency, and deliverables for work that matters. Choose by the job in front of you — and if that job is "rescue a 20-minute video with dense audio and hand me back the original picture," run it through the cloud pipeline and compare the results yourself.