Lip Sync Image Generation Using Dual Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing lip sync image generation technologies fail to prevent the person background image from controlling the utterance movement, leading to unnatural lip sync when generating images without audio signals.

Innovation Solution

A dual artificial neural network model system is employed, where the first model generates utterance synthesis images using both person background images and audio signals, and silence synthesis images using only background images, while the second model classifies these images to ensure utterance movement is controlled solely by the audio signal, preventing background image influence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single neural network model is used to generate lip sync images, then the generation process is simple, but the person background image controls the utterance movement leading to unnatural lip sync

Engineering Contradiction:
Improvemodel structureVSAvoidlip sync accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The single neural network is divided into two separate networks: a first neural network that generates utterance synthesis images using both background image and audio signal, and a second neural network that generates silence synthesis images using only the background image. This segmentation allows independent control of utterance movement (controlled by audio signal) and background features, preventing the background image from inappropriately controlling utterance movement in silence conditions.

Inventive Principle:
Principle #1Segmentation

2Stability of the object's composition

If the person background image is used as input for generating silence synthesis images, then the background features are preserved, but the utterance movement becomes unnatural due to background image control

Engineering Contradiction:
Improvebackground consistencyVSAvoidutterance movement naturalness
Core Design Contradiction:
Stability of the object's compositionVSManufacturing precision

Solution Approach 1:

The system extracts and separates the utterance movement control from the background image input. The second neural network is specifically designed to generate silence synthesis images by using only the background image as input, excluding the audio signal. This extraction ensures that background features are preserved while preventing the background image from controlling utterance movement, as the audio signal (which carries utterance information) is completely excluded from the input.

Inventive Principle:
Principle #2Taking out (Extraction)

3Stability of the object's composition

If the audio signal is completely excluded from silence synthesis generation, then the background image stability is maintained, but the utterance movement control is lost

Engineering Contradiction:
Improveimage stabilityVSAvoidutterance control accuracy
Core Design Contradiction:
Stability of the object's compositionVSReliability

Solution Approach 1:

The system uses a feedback mechanism where the second neural network generates silence synthesis images based on background image input, and this generated image is then used as a reference to guide the first neural network's generation process. The feedback loop ensures that when audio signal is present, the utterance movement is properly controlled by the audio signal while maintaining consistency with the background image features, thus maintaining both stability and reliability.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12190903B2Apparatus and method for generating lip sync image
Publication Date: 2025.01.07 DEEPBRAIN AI INC
  • US12190903B2 patent drawing
  • US12190903B2 patent drawing
  • US12190903B2 patent drawing

AI summary

An apparatus for generating a lip sync image according to a disclosed embodiment has one or more processors and a memory which stores one or more programs executed by the one or more processors. The apparatus includes a first artificial neural network model configured to generate an utterance synthesis image by using a person background image and an utterance audio signal corresponding to the person background image as an input, and generate a silence synthesis image by using only the person background image as an input, and a second artificial neural network model configured to output, from a preset utterance maintenance image and the first artificial neural network model, classification values for the preset utterance maintenance image and the silence synthesis image by using the silence synthesis image as an input.