Lip Sync
Fitting mouth shapes to the beat of a sound
- Lip sync is the job of shaping and attaching mouth movement to match a voice. The sound comes first, and the picture follows it.
- When a sound happens matters far more than which word it is.
- A mouth shape gets pulled by the sounds around it. Swap in one shape per sound and it looks stiff.
- Human eyes are extremely sensitive to mismatch — a shift as small as a single blink gets caught right away.
- Rather than rebuild an entire face, it's common to redraw just the area around the mouth and lay it over the original footage.
Contents
1The analogy
When you jump rope, your feet aren't watching the rope. They're listening for the rope hitting the ground and timing the jump to that beat, rising before it lands. React after seeing it, and it's already too late.
That's the timing job lip sync does. The voice comes first, and mouth shapes get built and attached to match that sound. Sound sets the beat, and the picture follows.
They're alike in getting caught the moment they slip, too. In jump rope, being off by even one beat catches the rope on your foot; in lip sync, being off by even one instant makes a mismatch a viewer notices immediately. Nobody pays attention when it's synced well, but the instant it slips, that's all anyone sees.
2In detail
It reads the texture of the sound, not the letters
Building mouth shapes by reading a script isn't the usual approach. The same sentence gets spoken at a different pace by every person, and gets drawn out, broken up by breaths, or mixed with laughter along the way. A mouth built from letters alone keeps drifting away from the actual sound.
So the sound itself is the raw material. Cut the voice into short pieces and pull out how strongly each pitch shows up, and the moments the mouth opens wide and the moments it closes are already sitting right there in that measurement. It closes during quiet stretches and opens where sound bursts out.
Silent stretches matter too. How the mouth moves during an intake of breath, or a brief pause, has a big effect on how natural it looks. Freeze the mouth completely just because there's no sound, and it starts looking like a puppet.
A mouth shape gets pulled by the sounds around it
Fix one mouth shape per sound and swap them in order, and it comes out stiff and strange. That's because a human mouth pre-shapes itself for the sound coming next. If a rounded sound is about to happen, the mouth already starts rounding slightly during the sound before it.
So the shape at any instant is decided by looking at several moments before and after together, not just the sound happening right now. That's exactly why building this in real time requires grabbing a little bit of sound ahead of time, which introduces a short delay. Working from a sound that's already fully recorded avoids this problem, since the whole thing can be looked at at once.
For the same reason, when speech speeds up, the mouth doesn't finish opening fully before moving to the next shape. Imitating that blurring is part of what makes it look like a person actually talking.
Redrawing just the mouth
Rebuilding an entire face tends to make the eyes and head movement look off. So a common approach leaves the original footage alone and redraws only the region around the mouth and jaw, laying it over the original. Since the original lighting and skin texture stay in place, it comes out far more natural.
Smoothly stitching the boundary between the changed area and the untouched area is the tricky part. Any color mismatch at that seam shows up as a ring-like blur around the mouth. It shows especially clearly when the head turns sharply or a hand covers part of the mouth.
The eye catches a mismatch first
How much a sound and picture can drift apart before it's noticed is much narrower than it seems. People are fairly forgiving of sound arriving later than the picture — a sound from far away naturally arrives late anyway. The other way around, sound arriving before the picture, feels wrong much faster.
So when fine-tuning lip sync, it's safer to push the picture very slightly ahead of the sound. It's also known that a mismatch that jumps around from moment to moment bothers people far more than one that's off by a steady amount throughout.
Different voice, different language
Dubbing is where lip sync gets used most. Rebuild mouth shapes to match audio freshly recorded in another language, and it can look like the speaker actually said it in that language, no subtitles needed. It's used to release lectures or informational videos in several languages.
The same technique gets applied to animated characters getting a voice. Work that used to mean drawing the mouth frame by frame by hand gets generated straight from the sound instead, cutting the workload sharply.
But it also means words a real person never said can be laid onto their actual face. Combined with technology that imitates a voice, it can produce footage that looks like someone said something they never said. That's why marking a video as synthetic and obtaining consent are handled alongside this technology.
3More precisely
A lip sync model takes in features pulled from the sound together with a picture of the face, and redraws the area around the mouth to match the sound. Some approaches generate pixels straight from the sound; others first fix the position of reference points on the lips and jaw, then draw the picture to match those positions. Whether it turned out well gets checked with a separate judging model that measures how well the resulting mouth shape and the sound actually fit together.
Sound and mouth shape aren't a one-to-one match. Very different sounds often produce almost the same mouth shape, so sounds that share a shape get grouped together and handled as one. That's exactly why it's hard to tell what was said from the mouth shape alone.
The analogy breaks down in one place. Jump rope only needs the beat to land right; lip sync has to get timing and shape right together. Open right on beat with the wrong shape, and it still looks off. And a person jumping rope reacts to the sound in the moment, but lip sync is usually built having already received the whole sound in advance, looking ahead and behind as it goes.
4Try it yourself
5Common misconceptions
It's easy to think mouth shapes are built by reading the script, but actually the opening and closing moments are pulled from the sound itself, since the same line is spoken differently every time.
It's easy to think you can tell what was said just by watching the mouth, but actually very different sounds come out looking almost the same, which makes reversing it hard.
It's easy to think a small mismatch goes unnoticed, but actually even a very brief shift gets caught quickly, especially when the sound arrives before the picture.
7One-line summary
In shortLip sync builds and attaches mouth shapes to match the texture of a voice, and because sound leads and the picture follows, even the smallest mismatch shows up immediately.
Spotted an error or have a better analogy? Suggest an edit · Last updated2026-09-02