Deep Dive

Stem Separation — Splitting Audio into Layers

GRAB / SMPLR / LIBRARY
GRAB — STEMS
Splitting stems… · AI · 12.4s · sample_01.wav
VOCALS
80% Pad
DRUMS
100% Pad
BASS
90% Pad
OTHER
70% Pad
A/B SAVE ALL

1. What Stem Separation Is

Stem separation takes a mixed audio recording and splits it into isolated layers — called "stems." A full song becomes four separate audio files: vocals, drums, bass, and everything else (other instruments). The technology uses AI (a machine learning model called htdemucs) that runs entirely on your device — no internet needed, your audio never leaves your phone.

Think of it like being handed the master tapes. You can mute the vocals and hear just the instrumental. You can take the drum break and build a new beat around it. You can isolate a bassline to study the groove. You can extract an acapella for a remix.

2. Where You Can Separate Stems

GRAB screen: After capturing or importing audio, tap the SPLIT button. This is the most common path.

Sampler screen: Long-press a pad with a loaded sample, open the overflow menu, choose Stem Separate.

Library screen: Select a sample (or multiple samples) in multi-select mode, then tap Stem Separation in the toolbar.

Note Editor screen: After capturing audio into a note, tap Stem Split & Transcribe to separate AND transcribe each stem to notation automatically.

3. The Separation Process

When you tap SPLIT, ESSNCE loads the htdemucs model into memory (about 289 MB). The model analyses your audio — looking at the frequency spectrum, identifying patterns that match "vocals" vs "drums" vs "bass."

Processing time depends on the audio length and your device speed. A 30-second clip on a recent phone takes 10–30 seconds. A 2-minute clip might take a minute or more.

A progress bar shows the current step (loading model → analysing → separating) with percentage. The app fires a subtle haptic at each milestone (25%, 50%, 75%, 100%).

Length limit: There's a length limit based on your device's RAM: 30 seconds on low-end devices, 60 seconds on mid-range, 120 seconds on high-end. A warning dialog appears if your audio exceeds the limit — you can truncate or cancel.

4. The Stem Tray (What You See After Separation)

When separation completes, a tray slides up from the bottom with four stem channels:

VOCALS (red)

The lead vocal, backing vocals, and any vocal-like sounds (spoken word, ad-libs, vocal chops).

DRUMS (orange)

Kick, snare, hi-hats, cymbals, percussion — anything percussive. The model separates the full drum bus, not individual drum sounds.

BASS (blue)

Bass guitar, synth bass, 808s, sub-bass — the low-end melodic instruments.

OTHER (purple)

Everything else: guitars, keyboards, strings, horns, pads, FX, and any sounds that don't fit the other categories.

Each stem channel has: a play button (hear just this stem), a colour dot, the stem name, a volume slider (0–150%), a level percentage readout, and a "Pad" chip to save this stem to a sampler pad.

Volume slider

0–150%. At 100%, the stem plays at its original detected level. Push above 100% to boost quiet stems. Pull below 100% to reduce dominant ones.

A/B toggle

Flips between the original mixed audio and the separated stems. Use this to check: did the separation lose anything important? Are the stems clean?

AI badge

Shows processing time (e.g., "AI · 12.4s") or "Basic" for non-AI processing.

Save All to Pads

Saves all four stems to consecutive sampler pads at once. Choose the starting pad (1, 5, 9, or 13). Vocals go to pad N, Drums to N+1, Bass to N+2, Other to N+3.

5. Quality and Limitations

Stem separation is impressive but not perfect. Know what to expect:

6. Creative Uses for Stems

Remix / Bootleg

Capture a song. Split stems. Load the vocals and drums onto pads — you now have the acapella and the drum break. Build a new instrumental around them. Classic remix workflow.

Drum Break Extraction

Capture a section of a song with a drum break. Split stems. Load just the DRUMS stem onto a pad. Open the waveform editor, find the best bar, trim to it, and chop it across pads. You've extracted a clean drum break from a full mix.

Bassline Study

Capture a song with a bassline you love. Split stems. Isolate the BASS stem. Load it onto a pad and slow it down (time stretch to 0.5×). Now you can hear every note clearly — transcribe it, learn it, or sample it.

Vocal Chop Instrument

Split stems from a song. Load the VOCALS stem onto a pad. Auto-chop it. Now each pad has a different vocal slice — a syllable, a breath, a phrase. Play them in a new order and you've built a vocal chop instrument.

Karaoke / Instrumental

Split stems. Mute the VOCALS channel (drag the slider to 0%). Save All to Pads. The DRUMS, BASS, and OTHER stems together are the instrumental. Export them or mix them with your own vocal recording.

7. Stem-to-Notation (in the Notes Screen)

After you've separated stems, you can go further — transcribe each stem to musical notation automatically. This is available from a note's Captured Sample section.

Tap Stem Split & Transcribe. ESSNCE separates the audio, then runs different AI models on each stem: CREPE pitch detection for vocals and bass (monophonic, precise), BasicPitch for the OTHER stem (polyphonic, chord-aware).

The result is a "Song" part on the staff with all stems merged into one multi-instrument notation. You can edit individual notes, change instruments, or export as MIDI.

The Refine button lets you adjust sensitivity and timing parameters and re-run the transcription if the first pass wasn't perfect.

Note: Not all stem types produce useful notation. Drums transcribe as percussion notation on a single staff — each drum sound maps to a different staff line or space. Vocals and bass transcribe best because they're monophonic (one note at a time). The OTHER stem polyphonic transcription is more experimental — chords may need manual correction.

8. Suggestions to Try

  1. Grab the intro of a favourite song — the first 30 seconds before the vocals come in. Split stems. The OTHER stem is the instrumental bed. Load it onto a pad, loop 4 bars, and build a new beat under it.
  2. Capture a drum fill (just 2–4 seconds). Split stems. Load the DRUMS stem. Auto-chop it. Now every drum hit from that fill is on its own pad. Re-arrange them for a completely different fill.
  3. Record yourself singing a melody with your phone mic. Capture it in GRAB. Split stems — since it's just a vocal, the VOCALS stem will be nearly perfect. Send it to Compose to see the notation.
  4. Capture a track with a distinctive bassline. Split stems. Load the BASS stem. Tempo-match it to your project BPM. Now the bassline plays in time with your beat — instant sample flip.
Stems unlocked

You can now split any audio into its component parts, extract the elements you need, and build entirely new creations from the pieces. Stem separation turns a finished song into a box of raw materials — and you know how to open it.

← Back to What's Next