Deep Learning Mouth Shape Generation with HD Transform
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating mouth shapes in synchronization with speech using deep learning networks produce inaccurate results due to low-definition training and lack of standardization in sync discrimination, resulting in poor mouth shape quality.
Innovation Solution
A method employing a mouth shape generation preprocessing deep learning model and a high-definition transform postprocessing deep learning model to transform a preset mouth shape into a synchronized mouth shape, using loss functions and feature vector combinations to enhance accuracy and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If low-definition deep learning training is performed for learning stability, then training stability is improved, but mouth shape quality deteriorates
Solution Approach 1:
The patent divides the deep learning training process into two distinct stages: a first training stage using low-definition data for stability, and a second training stage using high-definition data for quality. This segmentation allows each stage to optimize for its specific goal without compromise.
Solution Approach 2:
The first low-definition training stage serves as a preliminary action that establishes stable feature extraction capabilities before the second high-definition training stage refines the mouth shape generation quality. This preliminary training provides a solid foundation for subsequent high-quality generation.
2Device complexity
If mouth shape is generated on the basis of initial mouth shape of input face image, then generation process is simplified, but mouth shape accuracy deteriorates
Solution Approach 1:
The patent extracts and removes the constraint of the initial mouth shape from the generation process. By using a preset mouth shape template instead of relying on the input face image's mouth shape, the system achieves more accurate synchronization with speech while maintaining process simplicity.
Solution Approach 2:
The patent changes the parameter basis for mouth shape generation from the variable initial mouth shape of input images to a standardized preset mouth shape template. This parameter change ensures consistency and accuracy in synchronization with speech across different input images.
3Device complexity
If there is no standard for sync discrimination, then discrimination process is simplified, but mouth shape synchronization quality deteriorates
Solution Approach 1:
The patent introduces a feedback mechanism using a pre-trained lip sync discrimination deep learning model. This model provides standardized discrimination of mouth shape synchronization quality, enabling effective training optimization while maintaining a relatively simple overall process structure.
Data Source
AI summary
The present invention relates to a method for generating a mouth shape by using a deep learning network, and comprises the steps of: receiving a face image as an input; generating a mouth shape of the face image into a preset mouth shape by using a first mouth shape generation preprocessing deep learning model; receiving speech information as an input; generating the preset mouth shape into a mouth shape synchronizing with the speech information by using a second mouth shape generation preprocessing deep learning model; and transforming the mouth shape synchronizing with the speech information into high definition by using a high definition transform postprocessing deep learning model.


