Node reference

Avatar Generator

Add Node → Generate → Avatar Generator

Give it a portrait and a voice recording, and it gives you back a talking video: the mouth moves in sync with the audio, and the face comes to life around it.

It arrives already wired to an empty video slot, ready to fill once you generate.

What’s on the node

The Avatar Generator node with its three input handles and an image and audio node connected

Area What goes in it
Prompt box Optional. A line or two describing the performance — see below. Leave it empty and the model works fine.
Model and settings The model picker. There’s little else to set: everything about the clip follows the audio you connect.
P Optional. Takes a Prompt Generator, a Text Area, a Prompt Split, or a finished result’s own prompt handle.
I Required. The portrait of whoever’s talking.
A Required. The voice recording. Its length decides both the clip’s length and what it costs.

A ring is hollow until something is connected and fills with colour once it is, so you can see at a glance what the node still needs. There’s more on all of this in How nodes connect.

What it needs

Two things, both required: an image of the person (or character) you want talking, and an audio clip of what they should say. Drag the image onto the I handle and the audio onto the A handle.

For the image, use a clear, front-facing shot with the whole face visible — no sunglasses, no hands near the mouth, no extreme angles. The model needs to read the face to sync the lips well.

For the audio, keep it clean: just the voice, no music or background noise, no people talking over each other. Anything extra in the recording makes the lip-sync less convincing.

There’s also an optional prompt, if you want to nudge the performance — more on that below.

Which model should you use?

Model Max length Price range What it’s good at
Pro Kling AI Avatar 2.0 Pro 5 min $0.56–1.68 The safest default and the most natural result — subtle expressions, believable movement. The one for real, photorealistic people.
Good Kling AI Avatar 2.0 Standard 5 min $0.28–0.84 The same idea at half the price and lower quality. Good for a quick test before committing to a Pro generation.
Creative HeyGen Avatar IV 2 min Runware only. Reads tone and rhythm in the audio and drives head movement and expression from it. The best choice for anything that isn’t a real photo — illustrated characters, mascots, 3D models.

Prices are indicative and set by the providers, who revise them: see wavespeed.ai or runware.ai for current rates.

Price follows the length of your audio, not any setting you pick: the figures above are for a five-second clip and a fifteen-second one, and the two Kling models bill in five-second blocks, so a six-second recording costs the same as a ten-second one. Those rates are WaveSpeed’s, and HeyGen has no figure at all because it only runs through Runware.

Prompting the avatar

You don’t need a prompt at all — every model works fine without one. But a short one helps: one or two sentences describing the performance, not the lip-sync itself (that part is automatic). Say who the avatar is and their tone, how they should look while speaking, and how animated they should be. Keep it simple and avoid contradictory or vague instructions — a few clear directions beat a long paragraph.

Treat your first generation as a test: check the lip-sync and expression on a short clip before committing to a longer one.

← Back to all articlesNeed help? Go to Support →