Video Generator
Add Node → Generate → Video Generator
This node makes video — from a text prompt on its own, from one or two still images, from a set of reference images, or from a clip you feed it.
It arrives already wired to an empty video slot, ready to fill once you generate.
What’s on the node

| Area | What goes in it |
|---|---|
| Prompt box | Where you describe the shot. Type into it directly, or wire something into the prompt handle and it turns read-only, showing the upstream text live. |
| Model and settings | The model picker, and under it whatever that model exposes — resolution, clip length, audio. The controls change with the model, and switching resets them. |
| P | One only. Takes a Prompt Generator, a Text Area, a Prompt Split, or a finished result’s own prompt handle. |
| I | Reference images, numbered in the order they’re sent to the model. Each takes an image, a Media Group or a Media Placeholder. |
| FL | First and last frame, on the models that support it: a starting image and an ending one, with the motion generated between them. |
| V | A finished clip, on the models that restyle existing footage. |
Which of these you actually see depends on the model — the slots change shape when you switch. A ring is hollow until something is connected and fills with colour once it is, so you can see at a glance what the node still needs. There’s more on all of this in How nodes connect.
Text-to-video, or image-to-video?
Type a prompt and generate, and the model invents the whole shot from words alone.
Connect an image instead, and it animates that instead of guessing: composition, subject and framing are already decided, and the model only has to supply the motion. This is almost always the better way in — with nothing to anchor it, a model has very little to go on, and you can burn through a lot of generations before landing near what you had in mind. An image made with the Image Generator works fine as a starting point.
Pick the model before you wire anything up. The node’s input slots change shape depending on which one you choose.
Which model should you use?
| Type | Model | Resolution | Duration | Price range | What it’s good at | |
|---|---|---|---|---|---|---|
| Pro | text-to-video image-to-video video-to-video |
Seedance 2.0 | 480p · 720p · 1080p · 4K | 4–15s | $0.60–18.00 | The most flexible of the six — every mode, first and last frame included — and among the best of those that reach 4K. |
| Pro | text-to-video image-to-video video-to-video |
Kling O3 | 720p · 1080p · 4K | 3–15s | $0.42–6.30 | The one for reference images, and for restyling a clip you already have. No first and last frame — that’s V3’s job. |
| Pro | text-to-video image-to-video |
Kling V3 | 720p · 1080p · 4K | 3–15s | $0.42–6.30 | The first-and-last-frame one: give it a start and an end, get the motion between. Every bit as good as O3, just for a different job. |
| Good | text-to-video image-to-video |
Veo 3.1 | 720p · 1080p · 4K | 4 · 6 · 8s | $0.80–3.20 | A solid general-purpose pick, and the one to reach for when the clip needs audio. Fixed lengths, and audio doubles the price. |
| Good | text-to-video image-to-video video-to-video |
Gemini Omni | 720p | 3–10s | $0.70–1.40 | Works from several reference images, but no first and last frame. Priced by length alone, with no resolution or audio surcharge. |
| Creative | text-to-video image-to-video |
LTX 2.3 | 480p · 720p · 1080p | 5–20s | $0.10–0.60 | An open model: cheap by a wide margin, good results, and the longest clips here. |
Prices are indicative and set by the providers, who revise them: see wavespeed.ai or runware.ai for current rates.
Prices are per clip, charged straight to whichever provider you’ve connected. The figures here are WaveSpeed’s rates; Runware prices its own catalogue. The low end is a five-second clip at the smallest size with audio off; the high end is fifteen seconds at the largest size with audio on — except for Veo 3.1 and Gemini Omni, which cap out at eight and ten seconds and are priced there. Both resolution and length move the figure a lot, so a cheap model run long and large can cost more than an expensive one kept short.
First and last frame, or reference images?
These are two different ways of feeding a model stills, and they don’t do the same thing.
With first and last frame you give a starting image and an ending image, and the model generates the motion that connects the two. Use it when you need a precise, controlled transition between two exact shots. Seedance 2.0, Kling V3 and Veo 3.1 support it; Kling O3, Gemini Omni and LTX 2.3 don’t — O3 works from reference images instead.
With reference images you drop in one or more stills as inspiration — a subject, a location, a style — and the model draws on them without locking any exact frame. More creative freedom, less control over where the shot ends up.
Careful — this one depends on your provider. Seedance 2.0 can run on reference images alone, with no first and last frame and no video wired in. That works through Runware, but not through WaveSpeed: there the model always wants a starting frame or a clip to work from. Same model, same node, different rules depending on which key you’ve connected.
Some models take a whole video as input instead, letting you restyle existing footage or change parts of it while keeping the original motion.
To connect anything, drag it from the Media Library onto the canvas, then link it to the matching input on the Video Generator node.
Switching to a different model resets your settings, so nothing carries over by mistake.