Neural Network Face Expression Synthesis for Voice-Driven Lip Sync

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies cannot control face expressions, including lip movements, in synthesized face moving images based on voice signals, leading to difficulties in maintaining identity and ensuring suitable facial expressions.

Innovation Solution

An image processing device and method that utilize two neural networks to predict action units from voice signals and transform face still images into corresponding expressions, ensuring accurate and natural lip movement synchronization with voice content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If the entire face area is used as the masking region for lip synthesis, then the lip movement can be synthesized, but it becomes difficult to hold the identity of the person

Engineering Contradiction:
Improvelip movement synthesis accuracyVSAvoididentity preservation
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent divides the face region into multiple areas: a first region (mouth and surrounding area) for synthesis and a second region (other face areas) for identity preservation. This segmentation allows selective processing where only the necessary region is synthesized while preserving the rest of the face identity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities to different regions: the first region undergoes synthesis processing to achieve accurate lip movement, while the second region maintains original characteristics to preserve identity. This local differentiation resolves the contradiction between synthesis accuracy and identity preservation.

Inventive Principle:
Principle #3Local quality

2Productivity

If neural network restoration is used for lip synthesis, then lip movement can be generated, but there is no guarantee that a face moving image with an expression that suits the voice can be generated

Engineering Contradiction:
Improvelip synthesis efficiencyVSAvoidexpression accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces a feedback mechanism where the synthesized first region is combined with the original second region, and the combined image is used to update the neural network model. This feedback loop ensures that the synthesized expressions accurately match the voice while maintaining overall facial expression consistency.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent merges the synthesized first region with the original second region to create a complete face moving image. This combination ensures that the generated expression not only matches the voice but also maintains consistency with the original facial characteristics and expressions.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240378783A1Image processing device, learning device, image processing method, learning method, image processing program, and learning program
Publication Date: 2024.11.14 NT T INC
  • US20240378783A1 patent drawing
  • US20240378783A1 patent drawing
  • US20240378783A1 patent drawing

AI summary

Provided is an image processing device 10 including: an action unit acquisition unit 101 configured to input a voice signal to a first neural network to obtain an action unit representing movement of a mimic muscle corresponding to the voice signal from the first neural network; and a face image generation unit 102 configured to input the action unit and a face still image to a second neural network to obtain a sequence of a generated image obtained by transforming an expression of the face still image into an expression corresponding to the voice signal from the second neural network.