📷 Your scene, blocked out in 3D
Grey shapes, real camera📷 The generated shot
Same framing, fully renderedWrite it in plain English — "a small cafe, late winter afternoon, a woman comes in with her boyfriend" — or upload a reference photo of the location. The room, the furniture and the people are laid out for you in seconds. Drag anything that is not where you want it.
Fly the camera to the framing you want and press New shot. Move the playhead, frame the next angle, press it again. Each one is a locked-off setup with a hard cut between them. Key the action on a timeline: who moves, who turns, who is looking at whom.
Write how it is played — the register for the whole scene, and a note per beat describing what the faces do. Attach reference photos for the faces and the location. Hit generate: your camera move goes across as a motion reference with a written shot list, so the cuts and the framing survive.
Put the camera exactly where you want it, in 3D
Hard cuts between shots — not one long drifting move
Timed beats: say what happens at 0-4s and have it honoured
Direct facial performance in words, beat by beat
Characters hold their positions and eyelines across every cut
Reference photos for faces and locations carry through
Refine by asking — "move the counter left, add more customers"
No 3D modelling, rigging or camera-rig software to learn
📷 Wide — the room
📷 Cut — close on her
📷 Cut — the reverse angle
Yes. You place the camera in a 3D view of your scene, frame the shot, and lock it. The camera move is recorded and handed to the video model as a motion reference, so the generated video follows the framing you composed instead of a framing the model invented.
You can cut. Each shot is a locked-off camera setup, and the generator is told where the cuts fall — for example a wide at 0s, a close-up at 4s, the reverse angle at 7.5s. Without this, AI video models tend to smooth everything into one long drifting move.
No. You describe the scene in plain English — or upload a reference photo — and the layout is built for you: room, furniture, and the people in it. You move things by dragging. There is no modelling, texturing, or rigging.
You write direction in words, and it travels with the shot. A register for the whole scene ("everything is held back, nothing is cold") plus a note per beat describing what the face does. Physical description works far better than emotion labels — "the eyes drop to the table a beat too long" gives the model something to render, where "devastated" does not.
The characters are real objects in a 3D scene, so their positions, facing and eyelines are the same fact in every shot. Attach a reference photo for a face and it is used for that character throughout. Two people looking at each other across a cut actually look at each other, because the geometry says so.
You write timed beats — "0-4s: she comes in through the door", "4-7.5s: close on her face" — and those exact spans are used. If you would rather not, describe the scene and the beats are worked out for you.
Up to 15 seconds per generated clip, which is the limit of the underlying video models. A 15 second scene can hold four or five separate shots, which is enough for a complete beat of a story.