Lip-Syncing Face to Speech Using Machine Learning Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems fail to generate accurate lip motion for dynamic, unconstrained videos and struggle with temporal consistency in lip-synced animations, especially when the viseme for a person is not available in the look-up table, limiting their ability to handle generic identities and speech inputs.

Innovation Solution

A processor-implemented method using a machine learning model and a pre-trained lip-sync model that determines visual and audio representations of a face, modifies face crops, combines them with reference frames, and optimizes lip-synced frames to achieve improved visual quality by training with historical data and loss functions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing systems use look-up tables with visemes for lip-sync animation, then the animation can be generated from textual inputs, but the systems fail when the desired viseme for a person is not available in the look-up table

Engineering Contradiction:
Improveability to handle generic identitiesVSAvoidaccuracy of lip motion generation
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system creates a digital twin or virtual avatar that replicates the physical person's facial characteristics and lip motion patterns. This virtual copy can then be animated with synthesized speech, allowing the system to work with generic identities without requiring pre-recorded visemes for each individual.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transitions from discrete viseme selection to continuous facial parameter control. By representing lip motions as continuous parameters rather than discrete look-up table entries, the system can generate accurate lip-sync for any identity without being constrained by predefined viseme sets.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If existing systems are trained on single speaker data, then they generate good quality lip-sync for specific speakers, but they fail to work for generic identities and dynamic videos

Engineering Contradiction:
Improvequality of lip-sync for specific speakersVSAvoidability to handle generic identities and dynamic videos
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system creates a universal lip-sync model that can handle multiple speakers and video types. The virtual avatar approach allows the same system to work across different identities and video conditions, eliminating the need for separate training for each speaker while maintaining high quality output.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transitions from static image processing to dynamic video processing. By capturing and animating facial movements in real-time video sequences, the system can handle dynamic, unconstrained videos rather than just static images, improving both versatility and temporal consistency.

Inventive Principle:
Principle #15Dynamics

3Productivity

If existing systems use conventional lip-sync methods, then they can generate lip animations, but they fail to maintain temporal consistency in generated lip movements

Engineering Contradiction:
Improveability to generate lip animationsVSAvoidtemporal consistency of lip movements
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The system ensures continuous and smooth lip movements by processing video frames in temporal sequence and maintaining consistent facial states across frames. This approach preserves the natural flow of speech-related facial expressions, eliminating temporal inconsistencies and flickering artifacts.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12154548B2System and method for lip-syncing a face to target speech using a machine learning model
Publication Date: 2024.11.26 INT INST OF INFORMATION THCHNOLOGY HYDERABAD
  • US12154548B2 patent drawing
  • US12154548B2 patent drawing
  • US12154548B2 patent drawing

AI summary

A processor-implemented method for generating a lip-sync for a face to a target speech of a live session to a speech in one or more languages in-sync with improved visual quality using a machine learning model and a pre-trained lip-sync model is provided. The method includes (i) determining a visual representation of the face and an audio representation, the visual representation includes crops of the face; (ii) modifying the crops of the face to obtain masked crops; (iii) obtaining a reference frame from the visual representation at a second timestamp; (iv) combining the masked crops at the first timestamp with the reference to obtain lower half crops; (v) training the machine learning model by providing historical lower half crops and historical audio representations as training data; (vi) generating lip-synced frames for the face to the target speech, and (vii) generating an in-sync lip-synced frames by the pre-trained lip-sync model.