SepGen

Multi-Stem Audio-Video Separation and Generation in a Single Model

Aviad Dahan1, Rajaei Khatib1, Yonatan Bitton2, Idan Szpektor2, Lior Wolf1, Raja Giryes1

1Tel Aviv University   2Google

One model that creates the sounds of a video from text or pulls them out of an existing video, with one track per described sound.

SepGen overview and motivation
SepGen overview and motivation. (a) Joint audio-video generation emits a single soundtrack in which no source can be addressed. The video is lifted to 4D and rerendered from a new viewpoint while the same mixed signal reaches both ears. (b) SepGen emits one waveform per captioned source. Placed at its source and carried into the lifted scene, each waveform can be rendered as spatial sound from the new viewpoint (illustration).

Each sound on its own track

SepGen turns an existing audio-video generator into one that also outputs each described sound separately, in the same run that makes the video and its soundtrack.

The original soundtrack stays the same

The added weights never change the soundtrack the generator already makes. Each sound reads only its own description and is steered away from the other one.

One model for both tasks

The same weights create sounds from text or pull them out of an existing video. A short extra training round improves generation without hurting separation.

Abstract

How it works

1

Describe each sound

Write one description for the whole scene and one for each sound you want, for example “a jeep engine” and “a taxi honking”.

2

One model writes every track

SepGen adds two audio tracks to a pretrained video generator. In one run it makes the video, the full soundtrack and one track per description, and each track reads only its own description.

3

Separation and generation are both supported

Give it an existing video instead. It keeps the real soundtrack fixed, so the two extra tracks become the separated sounds.

How the noise is scheduled

Each line is one audio track as the model cleans it up, from pure noise at the top to the finished sound at the bottom.

Separation

Separation

The real audio-mix is given clean from the start, and the two tracks listen to it at every step.

Generating all tracks together

Generating all tracks together

All three tracks start as noise and are cleaned up together, so early on the two tracks listen to an audio-mix that is still mostly noise. The finished tracks come out unaligned with the audio-mix.

Generation in SepGen

Generation in SepGen

Once the audio-mix is clear enough, the tracks listen to the model’s current guess of the finished audio-mix instead of its noisy version. We call this Estimated Separation.

a track being madethe audio-mixwhat the tracks listen tothe model’s guess of the finished audio-mix

Examples

Stored benchmark outputs of SepGen next to the cascade baselines, grouped as same-category sources, speech, and sounds. Click any cell to hear that stem over the video.

Quantitative results

How to read the tables

Means over three seeds (Table 3: five seeds); best in bold, second underlined. In every table, each method is handed the identical audio-mix and captions. Column names are spelled out in the headers.

3.82 vs 3.35

Generation: how well each generated sound matches its description (SAM Audio Judge), compared with the best other method

0.05 vs 0.74

Separation: word error rate for each speaker on the Veo benchmark, compared with the best other method

63% / 58%

of listeners preferred SepGen in the generation and separation studies (the other method got 28% and 25%)

User study

We ran two listening tests, one for each mode, comparing SepGen against the top competitors in an A/B test. Listeners heard two unlabeled versions of the same track and picked the better one, or called it a tie. Each bar shows how often they chose SepGen, a tie, or the other method.

Separation user study results
(a) Separation, 48 listeners. Ten clips from the Veo benchmark and SAM-Audio-bench, against SAM Audio.
Generation user study results
(b) Generation, 51 listeners. Ten clips of the generation benchmark, against AudioSep.

BibTeX

@article{dahan2026sepgen,
  title   = {SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model},
  author  = {Dahan, Aviad and Khatib, Rajaei and Bitton, Yonatan and Szpektor, Idan and Wolf, Lior and Giryes, Raja},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}