Lip-Syncing Face to Speech Using Machine Learning Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to generate accurate lip motion for dynamic, unconstrained videos and struggle with temporal consistency in lip-synced animations, especially when the viseme for a person is not available in the look-up table, limiting their ability to handle generic identities and speech inputs.
Innovation Solution
A processor-implemented method using a machine learning model and a pre-trained lip-sync model that determines visual and audio representations of a face, modifies face crops, combines them with reference frames, and optimizes lip-synced frames to achieve improved visual quality by training with historical data and loss functions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing systems use look-up tables with visemes for lip-sync animation, then the animation can be generated from textual inputs, but the systems fail when the desired viseme for a person is not available in the look-up table
Solution Approach 1:
The system creates a digital twin or virtual avatar that replicates the physical person's facial characteristics and lip motion patterns. This virtual copy can then be animated with synthesized speech, allowing the system to work with generic identities without requiring pre-recorded visemes for each individual.
Solution Approach 2:
The system transitions from discrete viseme selection to continuous facial parameter control. By representing lip motions as continuous parameters rather than discrete look-up table entries, the system can generate accurate lip-sync for any identity without being constrained by predefined viseme sets.
2Manufacturing precision
If existing systems are trained on single speaker data, then they generate good quality lip-sync for specific speakers, but they fail to work for generic identities and dynamic videos
Solution Approach 1:
The system creates a universal lip-sync model that can handle multiple speakers and video types. The virtual avatar approach allows the same system to work across different identities and video conditions, eliminating the need for separate training for each speaker while maintaining high quality output.
Solution Approach 2:
The system transitions from static image processing to dynamic video processing. By capturing and animating facial movements in real-time video sequences, the system can handle dynamic, unconstrained videos rather than just static images, improving both versatility and temporal consistency.
3Productivity
If existing systems use conventional lip-sync methods, then they can generate lip animations, but they fail to maintain temporal consistency in generated lip movements
Solution Approach 1:
The system ensures continuous and smooth lip movements by processing video frames in temporal sequence and maintaining consistent facial states across frames. This approach preserves the natural flow of speech-related facial expressions, eliminating temporal inconsistencies and flickering artifacts.
Data Source
AI summary
A processor-implemented method for generating a lip-sync for a face to a target speech of a live session to a speech in one or more languages in-sync with improved visual quality using a machine learning model and a pre-trained lip-sync model is provided. The method includes (i) determining a visual representation of the face and an audio representation, the visual representation includes crops of the face; (ii) modifying the crops of the face to obtain masked crops; (iii) obtaining a reference frame from the visual representation at a second timestamp; (iv) combining the masked crops at the first timestamp with the reference to obtain lower half crops; (v) training the machine learning model by providing historical lower half crops and historical audio representations as training data; (vi) generating lip-synced frames for the face to the target speech, and (vii) generating an in-sync lip-synced frames by the pre-trained lip-sync model.


