Deep Learning Mouth Shape Generation with HD Transform

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating mouth shapes in synchronization with speech using deep learning networks produce inaccurate results due to low-definition training and lack of standardization in sync discrimination, resulting in poor mouth shape quality.

Innovation Solution

A method employing a mouth shape generation preprocessing deep learning model and a high-definition transform postprocessing deep learning model to transform a preset mouth shape into a synchronized mouth shape, using loss functions and feature vector combinations to enhance accuracy and quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If low-definition deep learning training is performed for learning stability, then training stability is improved, but mouth shape quality deteriorates

Engineering Contradiction:
Improvelearning stabilityVSAvoidmouth shape quality
Core Design Contradiction:
Stability of the object's compositionVSManufacturing precision

Solution Approach 1:

The patent divides the deep learning training process into two distinct stages: a first training stage using low-definition data for stability, and a second training stage using high-definition data for quality. This segmentation allows each stage to optimize for its specific goal without compromise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The first low-definition training stage serves as a preliminary action that establishes stable feature extraction capabilities before the second high-definition training stage refines the mouth shape generation quality. This preliminary training provides a solid foundation for subsequent high-quality generation.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If mouth shape is generated on the basis of initial mouth shape of input face image, then generation process is simplified, but mouth shape accuracy deteriorates

Engineering Contradiction:
Improvegeneration process complexityVSAvoidmouth shape accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent extracts and removes the constraint of the initial mouth shape from the generation process. By using a preset mouth shape template instead of relying on the input face image's mouth shape, the system achieves more accurate synchronization with speech while maintaining process simplicity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter basis for mouth shape generation from the variable initial mouth shape of input images to a standardized preset mouth shape template. This parameter change ensures consistency and accuracy in synchronization with speech across different input images.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If there is no standard for sync discrimination, then discrimination process is simplified, but mouth shape synchronization quality deteriorates

Engineering Contradiction:
Improvediscrimination process complexityVSAvoidsynchronization quality
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces a feedback mechanism using a pre-trained lip sync discrimination deep learning model. This model provides standardized discrimination of mouth shape synchronization quality, enabling effective training optimization while maintaining a relatively simple overall process structure.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12182920B2Method and apparatus for generating mouth shape by using deep learning network
Publication Date: 2024.12.31 KLLEON INC
  • US12182920B2 patent drawing
  • US12182920B2 patent drawing
  • US12182920B2 patent drawing

AI summary

The present invention relates to a method for generating a mouth shape by using a deep learning network, and comprises the steps of: receiving a face image as an input; generating a mouth shape of the face image into a preset mouth shape by using a first mouth shape generation preprocessing deep learning model; receiving speech information as an input; generating the preset mouth shape into a mouth shape synchronizing with the speech information by using a second mouth shape generation preprocessing deep learning model; and transforming the mouth shape synchronizing with the speech information into high definition by using a high definition transform postprocessing deep learning model.