Video generation with layered and controllable audio.
Demo film 10:48 Best with sound
One scene. Separate voices, music, and sound effects.
Hear the whole scene.
Then make it your own.
Press Play. Solo a voice, turn down the music, or listen to a sound on its own.
These illustrative examples include the high-resolution airship and study outputs. The complete benchmark and its repeated seeds are available in the comparison browser.
Compare this scene ↗One scene.
Separate sound layers.
Each source has its own track, ready to listen to or edit.
Change one sound.
Keep the rest.
Move an event or generate a new source track, then refine the video around the revised audio.
Play the original and edited versions, then mute or solo individual sources. Each edited video is refined around its revised, fixed audio.
The Unrented Room
Three voices, two instruments, and the sounds of a room. Explore each source in this Soundwich-on-H3 scene.
Play uses the final refined soundtrack. Mute and Solo audition the nine source stems from before the final refinement. Reset mix restores the final soundtrack.
Control you can
see and hear.
Compare the same scene across methods, listen to individual sources, and inspect what changes when a component is removed.
Scope and current limitations
The main benchmark covers roughly ten-second clips with up to four output stems. Short events, weak sounds, persistent music, and difficult-to-track entities remain challenging. Editing refinement can also change video details.