Photo to Video

How to Turn a Photo Into a Talking Video

One good photo plus a script is enough to produce a video of someone speaking. Here is how the process works and how to get a natural result.

By Matt Hansen3 min read
A still portrait photo being animated into a talking video

Turning a still photo into a talking video used to be a visual-effects job. Now it is a two-input process: an image and an audio track. The quality gap between a bad result and a convincing one comes down to a handful of choices.

How photo-to-video actually works

Three stages run behind the scenes:

  1. Face analysis — the model locates facial landmarks: mouth corners, jawline, eyes, head pose.
  2. Audio analysis — the speech track is broken into phonemes, the sound units that map to mouth shapes.
  3. Animation — the still is warped frame by frame to match those mouth shapes, with small head and eye movements added so the result does not look frozen.

Everything the model has to invent is a chance to look wrong. That is why the input photo matters so much.

Choosing the right photo

Do:

  • Use a front-facing or slightly angled portrait.
  • Frame chest-up with visible shoulders — the shoulders sell the head motion.
  • Choose neutral or lightly smiling expressions with the mouth closed or barely open.
  • Use even lighting and a clean, uncluttered background.

Avoid:

  • Extreme angles or profile shots.
  • Wide grins showing full teeth — the model has to guess what is behind them.
  • Hands near the face.
  • Motion blur or low resolution.
  • Sunglasses and anything covering the mouth.

Getting the audio right

You have three options, in ascending order of quality:

  • Text-to-speech with a stock voice. Fast, free-feeling, generic.
  • A recorded voiceover. Authentic, but you are back to recording.
  • A cloned copy of your own voice. Type the script, hear yourself. This is the option that makes the result feel like footage rather than a demo.

Whatever the source, record or generate in a quiet environment. Background hiss makes lipsync models produce odd mouth movements between words.

Making it look natural

A few practical rules separate believable output from uncanny output:

  • Keep clips short. Under 60 seconds per shot. Long single takes accumulate small errors.
  • Do not over-animate. Subtle head movement reads as real; exaggerated motion reads as fake.
  • Match energy to the photo. An excited script over a deadpan portrait feels mismatched.
  • Add captions. They anchor the viewer''s attention and mask minor sync imperfections.
  • Cut away occasionally. A B-roll insert every few seconds keeps attention on the message.

When to use a still photo vs a full clone

Animating a single photo is perfect for one-off videos, testing an idea, or animating an image you already have.

A full clone is better once you are producing regularly: it gives you consistent framing, multiple outfits and backgrounds, and a face that stays recognisably the same across every video. If you plan to publish weekly, build the clone.

Animate photos of yourself, or of people who have given clear permission. Do not animate public figures or images scraped from the internet, and disclose synthetic presenters where a viewer might otherwise be misled. This is not just etiquette — platform policies increasingly require it.

A realistic quality checklist

Before you publish, watch your clip on a phone at normal speed and ask:

  • Does the mouth open and close on the right syllables?
  • Does the head move at all, and does that motion feel human?
  • Do the eyes blink?
  • Does the voice match the face''s apparent age and energy?
  • Would you scroll past it, or stop?

If the answer to the last question is "stop", the technical imperfections almost never matter.

Frequently asked questions

About Matt Hansen

Matt Hansen writes about AI video creation at Lyler, where creators turn a single selfie into talking-head videos.

Related reading