Neural Network Key Point Extraction for Stable Lip Sync
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing key point training technologies face challenges in maintaining continuity and stability over time due to annotation noise in facial key point data, leading to deteriorated image quality in lip sync facial image synthesis.
Innovation Solution
A neural network-based key point training apparatus and method that includes a key point model trained to extract stable key points and an image reconstruction model trained to reconstruct images using these key points, minimizing differences between extracted and labeled key points, and between reconstructed and input images, using optimized parameters calculated through specific objective functions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If facial key points are predicted using correct answer values obtained by labelling, then key point extraction can be performed, but annotation noise causes instability over time and deteriorates image quality
Solution Approach 1:
The training process is segmented into two distinct stages: first training the key point model using labeled key point data, then training the image reconstruction model using the extracted key points. This segmentation allows the system to handle annotation noise by separating the key point extraction function from the image reconstruction function, with each model optimized for its specific task rather than attempting to solve both problems simultaneously
Solution Approach 2:
The key point model serves as an intermediary between the input image and the image reconstruction model. By introducing this intermediate component that extracts key points separately, the system can filter out annotation noise before passing information to the reconstruction model, thereby improving both measurement precision and temporal reliability
2Productivity
If facial key points are extracted and used for image synthesis, then lip sync facial images can be generated, but annotation noise in correct answer values causes image quality deterioration
Solution Approach 1:
The patent extracts only the essential key point information from images through dedicated training, separating this extraction function from the overall image synthesis process. By taking out the key point extraction task and training it independently with its own optimization criteria, the system removes the negative impact of annotation noise on final image quality while preserving the productivity of automated image synthesis
Solution Approach 2:
The patent changes the training parameters and objective functions for different models: the key point model uses parameters optimized for key point accuracy, while the image reconstruction model uses parameters optimized for image quality. This parameter differentiation allows each model to operate at optimal precision for its specific function, resolving the contradiction between productivity and manufacturing precision
Data Source
AI summary
A neural network-based key point training apparatus according to an embodiment includes a key point model trained to extract key points from an input image, and an image reconstruction model trained to reconstruct the input image with the key points output by the key point model as the input. The optimized parameters of the key point model and the image reconstruction model can be calculated.


