Neural Network Key Point Extraction for Stable Lip Sync

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing key point training technologies face challenges in maintaining continuity and stability over time due to annotation noise in facial key point data, leading to deteriorated image quality in lip sync facial image synthesis.

Innovation Solution

A neural network-based key point training apparatus and method that includes a key point model trained to extract stable key points and an image reconstruction model trained to reconstruct images using these key points, minimizing differences between extracted and labeled key points, and between reconstructed and input images, using optimized parameters calculated through specific objective functions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If facial key points are predicted using correct answer values obtained by labelling, then key point extraction can be performed, but annotation noise causes instability over time and deteriorates image quality

Engineering Contradiction:
Improvekey point extraction accuracyVSAvoidcontinuity over time
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The training process is segmented into two distinct stages: first training the key point model using labeled key point data, then training the image reconstruction model using the extracted key points. This segmentation allows the system to handle annotation noise by separating the key point extraction function from the image reconstruction function, with each model optimized for its specific task rather than attempting to solve both problems simultaneously

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The key point model serves as an intermediary between the input image and the image reconstruction model. By introducing this intermediate component that extracts key points separately, the system can filter out annotation noise before passing information to the reconstruction model, thereby improving both measurement precision and temporal reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If facial key points are extracted and used for image synthesis, then lip sync facial images can be generated, but annotation noise in correct answer values causes image quality deterioration

Engineering Contradiction:
Improveimage synthesis capabilityVSAvoidimage quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent extracts only the essential key point information from images through dedicated training, separating this extraction function from the overall image synthesis process. By taking out the key point extraction task and training it independently with its own optimization criteria, the system removes the negative impact of annotation noise on final image quality while preserving the productivity of automated image synthesis

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the training parameters and objective functions for different models: the key point model uses parameters optimized for key point accuracy, while the image reconstruction model uses parameters optimized for image quality. This parameter differentiation allows each model to operate at optimal precision for its specific function, resolving the contradiction between productivity and manufacturing precision

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240428615A1Method and apparatus for learning key point of based neural network
Publication Date: 2024.12.26 DEEPBRAIN AI INC
  • US20240428615A1 patent drawing
  • US20240428615A1 patent drawing
  • US20240428615A1 patent drawing

AI summary

A neural network-based key point training apparatus according to an embodiment includes a key point model trained to extract key points from an input image, and an image reconstruction model trained to reconstruct the input image with the key points output by the key point model as the input. The optimized parameters of the key point model and the image reconstruction model can be calculated.