Audio to stems

Drop a song to split into stems, or select a file — it opens in the studio, ready to export.

WAV · AIF · AIFF · MP3 · M4A · FLAC · OGG

How to split a song into stems

Three steps in the browser, nothing to install — here is how the audio to stems converter works:

  1. Add your song: Drop a WAV, AIFF, MP3, M4A, FLAC or OGG — a finished mix, a rough bounce or a phone recording.
  2. It comes apart: BS-RoFormer separates vocals, drums, bass, guitar, keys and the rest; the kit can go further into kick, snare, hi-hat, toms, ride and crash.
  3. Take the stems: Full-length files, one per part — download them, or carry straight on into transcription or a DAW session.

Questions about the audio to stems converter?

We have answers.

How many stems do I get?

Vocals, drums, bass and the remaining melodic material, with further splits available depending on the source. A sparse arrangement separates further and more cleanly than a dense one.

Does it work on a live recording?

Yes, though bleed between sources makes it harder — a room recording where every mic hears every instrument is the worst case for any separator, including this one.

Can I get the stems as MIDI instead of audio?

Yes. Transcription runs per-stem, so you can take the separated parts through to notes rather than audio. That is a different route through the same pipeline.

Try our free audio to stems converter

  • What you get back

    A separated mix comes back as individual audio files, one per source. Vocals and drums are the cleanest — they are the most distinct in a mix and the model has the most to work with. Bass is reliable. Guitars and keys are harder, because in a dense arrangement they overlap in both frequency and time, and what you get is a best effort rather than a surgical extraction.

  • Why the separator is not demucs

    A lot of tools in this space wrap the same open-source model. This one runs a RoFormer-based separator, which is a transformer operating on the spectrogram rather than a convolutional network on the waveform. In practice the difference shows up on the hard material: dense mixes, heavy compression, and vocals that sit inside a wall of guitars.

  • From stems to a finished session

    Because separation here is a stage in a pipeline rather than the product, the same upload can come back as a finished session file instead — tracks laid out, tempo detected, each stem on its own named track. That is the audio-to-session route, and it is the one most people actually want.