Lip Sync Image Synthesis Using Adversarial Loss and Discrimination Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current lip sync image generation technologies face challenges in accurately synthesizing lip sync images when the timing and identity of the voice and image do not match, leading to reduced accuracy in mouth shape prediction and lip sync image synthesis, especially when dealing with different speakers or languages.

Innovation Solution

A lip sync image generation device using machine learning, comprising an image synthesis model and a lip sync discrimination model, which employs different total loss functions for matching and non-matching voice and image pairs, incorporating reconstruction loss, sync loss, adversarial loss, and identity loss to improve synthesis accuracy and identity preservation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a lip sync model is trained using only matching voice and image pairs, then the model can accurately synthesize lip sync images when voice and image match, but the model fails to learn when voice and image timing or identity do not match

Engineering Contradiction:
Improvelip sync image synthesis accuracyVSAvoidmodel adaptability to mismatched voice and image
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the training parameters by introducing mismatched voice-image pairs with controlled timing offsets and identity variations. The training data includes pairs where voice and image are deliberately misaligned in time or belong to different speakers, allowing the model to learn robust lip sync synthesis under various conditions while maintaining accuracy for matched pairs.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If a lip sync model is trained with multiple speakers and languages, then the model can handle different identities, but the accuracy of mouth shape prediction decreases due to unique utterance habits of each person

Engineering Contradiction:
Improvemodel capability to handle multiple speakers and languagesVSAvoidmouth shape prediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the training process into different components: identity-agnostic lip movement patterns are learned separately from speaker-specific characteristics. The loss function is designed to focus on universal lip sync patterns that apply across different speakers and languages, while speaker-specific details are handled through data augmentation and controlled training examples.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies homogeneity by standardizing the representation of lip movements across different speakers and languages. The model learns a unified lip sync synthesis approach that works consistently across diverse inputs, reducing the impact of individual utterance habits through normalized training data and consistent evaluation metrics.

Inventive Principle:
Principle #33Homogeneity

3Adaptability or versatility

If the timing discrepancy between voice and image is large, then the model can handle more challenging cases, but the accuracy of lip sync image synthesis is greatly reduced

Engineering Contradiction:
Improvemodel ability to handle timing discrepanciesVSAvoidlip sync image synthesis accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary alignment and analysis of voice-image timing relationships during the training phase. The model is pre-trained with progressively increasing timing offsets, starting from small discrepancies and gradually learning to handle larger mismatches. This staged approach allows the model to build robust timing compensation capabilities while maintaining synthesis accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12198713B2Learning method for generating lip sync image based on machine learning and lip sync image generation device for performing same
Publication Date: 2025.01.14 DEEPBRAIN AI INC
  • US12198713B2 patent drawing
  • US12198713B2 patent drawing
  • US12198713B2 patent drawing

AI summary

A lip sync image generation device based on machine learning according to a disclosed embodiment includes an image synthesis model, which is an artificial neural network model, and which uses a person background image and an utterance audio signal as an input to generate a lip sync image, and a lip sync discrimination model, which is an artificial neural network model, and which discriminates the degree of match between the lip sync image generated by the image synthesis model and the utterance audio signal input to the image synthesis model.