Lip Sync Image Synthesis Using Adversarial Loss and Discrimination Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current lip sync image generation technologies face challenges in accurately synthesizing lip sync images when the timing and identity of the voice and image do not match, leading to reduced accuracy in mouth shape prediction and lip sync image synthesis, especially when dealing with different speakers or languages.
Innovation Solution
A lip sync image generation device using machine learning, comprising an image synthesis model and a lip sync discrimination model, which employs different total loss functions for matching and non-matching voice and image pairs, incorporating reconstruction loss, sync loss, adversarial loss, and identity loss to improve synthesis accuracy and identity preservation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a lip sync model is trained using only matching voice and image pairs, then the model can accurately synthesize lip sync images when voice and image match, but the model fails to learn when voice and image timing or identity do not match
Solution Approach 1:
The patent changes the training parameters by introducing mismatched voice-image pairs with controlled timing offsets and identity variations. The training data includes pairs where voice and image are deliberately misaligned in time or belong to different speakers, allowing the model to learn robust lip sync synthesis under various conditions while maintaining accuracy for matched pairs.
2Adaptability or versatility
If a lip sync model is trained with multiple speakers and languages, then the model can handle different identities, but the accuracy of mouth shape prediction decreases due to unique utterance habits of each person
Solution Approach 1:
The patent segments the training process into different components: identity-agnostic lip movement patterns are learned separately from speaker-specific characteristics. The loss function is designed to focus on universal lip sync patterns that apply across different speakers and languages, while speaker-specific details are handled through data augmentation and controlled training examples.
Solution Approach 2:
The patent applies homogeneity by standardizing the representation of lip movements across different speakers and languages. The model learns a unified lip sync synthesis approach that works consistently across diverse inputs, reducing the impact of individual utterance habits through normalized training data and consistent evaluation metrics.
3Adaptability or versatility
If the timing discrepancy between voice and image is large, then the model can handle more challenging cases, but the accuracy of lip sync image synthesis is greatly reduced
Solution Approach 1:
The patent performs preliminary alignment and analysis of voice-image timing relationships during the training phase. The model is pre-trained with progressively increasing timing offsets, starting from small discrepancies and gradually learning to handle larger mismatches. This staged approach allows the model to build robust timing compensation capabilities while maintaining synthesis accuracy.
Data Source
AI summary
A lip sync image generation device based on machine learning according to a disclosed embodiment includes an image synthesis model, which is an artificial neural network model, and which uses a person background image and an utterance audio signal as an input to generate a lip sync image, and a lip sync discrimination model, which is an artificial neural network model, and which discriminates the degree of match between the lip sync image generated by the image synthesis model and the utterance audio signal input to the image synthesis model.


