every one of these systems makes you watch.
the last two years produced systems that write a finished song from a sentence. for a listener who wants a song, that’s magic. for the tens of millions of people who play an instrument — who have a guitar take recorded, a progression looping, an unfinished track open — it’s useless. it replaces exactly the thing they wanted to do themselves, and does nothing about the thing they can’t.
a guitarist can’t program drums. a beatmaker can’t write basslines. their sessions are full of their own playing and silent everywhere else. what they’re offered is loops and sample packs that don’t listen, virtual instruments that demand the production skill that’s missing, and now a text box that removes them from the music entirely.
the valuable seat at this table is the bandmate’s.
a bandmate doesn’t write your song. they hear what you’re playing and answer it — in your key, on your grid, under your playing. that is a different problem from writing whole songs, and it wants a different model: one trained to play the missing part of a mix when every other part is already there.
the consequence is structural. the more you’ve already played, the closer stems is to what it was trained on. a full session is its easiest case; a bare prompt is its hardest. that’s the exact inverse of text-to-song, and it’s why the instrument belongs inside the daw, where the context lives. (that’s how the model was trained; how large the effect is hasn’t been measured yet, and we’ll say so until it has.)
there is no text box for the music.
key, tempo, groove and arrangement are read from what’s already on the track. the few words you type describe the part — a kit, a room, a feel — not the song. that removes the entire prompt-engineering surface that makes generative tools feel like operating software instead of playing music.
and the output isn’t the machine’s song. it’s your song, finished. the pride of authorship is the whole reason anyone opens a daw; a tool that erases it has misread the room.
we tell you what it’s bad at, first.
generation isn’t real time yet: a take arrives after a wait you’ll notice, on one gpu, and you’re told so before you try it. drums and bass are strong; melody and vocals are still rough. no fake progress, no spinners, no claim we haven’t measured. in a field whose leaders are litigating over their training data, the tool that never once lied to a musician is the tool musicians keep.
stems ships today as an audio unit. it sits on a track like any other device, and takes arrive as plain audio in your project. just play.
get stems →