Phoneme-Driven Lip Synchronization via Viseme Action Units

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current facial animation tools, especially for lip synchronization, struggle to produce high-quality, nuanced, and editable results, often falling short of the quality achieved by performance capture and requiring significant animator effort, with procedural methods failing to capture the complexity of modern facial models and performance capture being limited by human capabilities and difficult to refine.

Innovation Solution

A method and system for animated lip synchronization that maps phonemes to visemes, synchronizes jaw and lip contributions, and outputs viseme action units, allowing for independent synchronization with facial muscles and dynamic simulation, enabling realistic and expressive facial animation that can be easily edited and refined.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If professional animators use laborious workflow with animation software to animate 3D facial rig, then high-quality facial animation is achieved, but significant time and effort are required

Engineering Contradiction:
Improvefacial animation qualityVSAvoidanimation production time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The facial animation system segments the complex animation task into distinct phoneme-based units. Each phoneme is mapped to corresponding viseme action units that control specific facial muscle groups (jaw, lips, tongue), allowing procedural generation of facial animation from speech input without manual keyframing of each facial element

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces phonemes as an intermediary layer between speech input and facial animation output. Speech is converted to phonemes, which then drive viseme action units that control facial rig, creating an automated pipeline that bridges audio and visual domains without requiring direct manual animation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If procedural approaches are used to automate lip-synchronization, then animation production is simplified, but quality does not keep pace with modern facial model complexity

Engineering Contradiction:
Improveautomation levelVSAvoidlip-synchronization quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system uses dynamic viseme action units that adapt their intensity and timing based on phoneme context, coarticulation effects, and speech characteristics. The jaw and lip contributions are dynamically blended rather than applied statically, allowing the procedural system to respond naturally to varying speech conditions and maintain quality across different phoneme sequences

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system modifies multiple parameters including viseme intensity, timing offsets, decay rates, and blending weights to achieve realistic lip synchronization. By adjusting these parameters based on phoneme properties and speech context, the procedural approach achieves quality comparable to manual animation while maintaining automation

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If performance capture is used to achieve high-quality facial animation, then realistic results are obtained, but the animation is limited by human performer capabilities and subsequent refinement is difficult

Engineering Contradiction:
Improvefacial animation qualityVSAvoidediting flexibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

Instead of capturing real human performance, the system creates a procedural copy of speech-driven facial motion through phoneme-to-viseme mapping. This synthetic approach generates animation that mirrors natural speech patterns without being constrained by actual performer limitations, and the procedural nature allows easy regeneration and adjustment

Inventive Principle:
Principle #26Copying

4Ease of operation

If simple interpolation between blend shapes is used in facial rig, then ease of operation is maintained, but anatomical accuracy and physical plausibility are compromised

Engineering Contradiction:
Improverig complexityVSAvoidanatomical accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system segments facial animation into anatomically distinct muscle groups (jaw, lips, tongue) controlled by separate viseme action units. This segmentation allows each facial element to be controlled independently based on phoneme requirements, maintaining anatomical plausibility while keeping the rig manageable through modular organization

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses parameter-based control of viseme intensity and timing to drive anatomically inspired facial muscles. By adjusting parameters like viseme strength, duration, and overlap, the system achieves realistic facial motion without complex rigid transformations, maintaining both anatomical accuracy and operational simplicity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10839825B2System and method for animated lip synchronization
Publication Date: 2020.11.17 JALI INC
  • US10839825B2 patent drawing
  • US10839825B2 patent drawing
  • US10839825B2 patent drawing

AI summary

A system and method for animated lip synchronization. The method includes: capturing speech input; parsing the speech input into phenomes; aligning the phonemes to the corresponding portions of the speech input; mapping the phonemes to visemes; synchronizing the visemes into viseme action units, the viseme action units comprising jaw and lip contributions for each of the phonemes; and outputting the viseme action units.