How to make AI talking head videos
September 26, 2026
Talking heads are one shot, one person, one line. Here is how to script, frame, and generate them — including Seedance and Kling on IQON.
A talking head is the simplest video format in marketing: one person speaks directly to the camera. AI talking heads add a generated person and synthesized speech so you can ship a pitch without booking talent.
Anatomy of a talking head
| Element | What to decide |
|---|---|
| Script | One spoken line, conversational, timed to length |
| Face | Fictional or stock — not a celebrity likeness |
| Frame | Vertical, front camera, minimal camera move |
| Audio | Native model speech or external voice (later pass) |
| Length | ~10–12 seconds for social hooks |
No scene changes. The power is intimacy.
Script rules that actually matter
Models rush if you give them too many words.
- Budget ~2 words per second for a natural pace.
- Direct delivery: "warm, unhurried, breath between sentences" — not "announcer voice."
- Quote the line in the video prompt so the model knows exactly what to say.
- Avoid the phrase "AI presenter" in speech prompts; it pushes a synthetic read.
IQON clamps over-long scripts to the last complete sentence so you do not get a clip that ends mid-thought.
Image-to-video vs text-to-video
Most talking heads start from a still frame (image-to-video). The still must pass the provider's likeness check.
Seedance runs a privacy precheck on input images. Photoreal faces — even AI-generated ones — often fail with:
The images or videos provided may contain likenesses of real people…
That is not a prompt bug. Cropping, blurring, or grid tricks are unreliable and against provider policy. Legitimate paths:
- Fictional face in an ordinary room — described as a specific character, not "a real person."
- Trusted pipeline output — Seedream still → Seedance I2V with BytePlus trusted URLs (IQON's Seedance workflow).
- Text-to-video fallback — same framing described in prose, no uploaded face.
- Kling I2V — IQON's current influencer test uses Nano Banana Pro for character + first frame, then Kling 2.6 Pro for ~10s with audio.
Read the full breakdown: Seedance likeness filter — what works.
Art direction for the still
Before video, nail the first frame:
- iPhone front-camera language (not "cinematic 8K portrait").
- Ordinary environment — desk, couch, window light.
- Imperfect skin texture reads more believable than beauty-filter smoothness.
IQON generates character portrait → first frame in room → video in one onboarding pass when you pick Influencer video.
When a talking head is enough
Use it for hooks, founder-style updates, product explainers, and UGC-style ads. Add b-roll, captions, and music in your editor if the campaign needs more — the AI clip is the talking piece, not the full ad.