Speech Video Generation With Fused Features for Landmark Continuity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech video generation technologies face challenges in synthesizing high-quality lip sync face images due to annotation noise in face landmark data, leading to unstable continuity over time and image quality deterioration.
Innovation Solution
A speech video generation device that combines image feature vectors from person background images with voice feature vectors from speech audio signals to generate a combined vector, which is then used to reconstruct the speech video and predict the landmark, thereby improving prediction accuracy and image quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If face landmark data with annotation noise is used for synthesis, then the synthesis process can proceed, but image quality deteriorates and continuity becomes unstable
Solution Approach 1:
The patent extracts and removes annotation noise from face landmark data through a learning model that distinguishes between valid landmark information and noise. The system separates the noise component from the original landmark data to produce cleaned landmark sequences that maintain continuity and accuracy for synthesis.
Solution Approach 2:
The patent implements a feedback mechanism where the learning model continuously refines landmark prediction by comparing generated results with ground truth data. The model uses error feedback to adjust its predictions, reducing annotation noise accumulation and improving temporal consistency across video frames.
2Ease of operation
If face landmark is aligned in standard space with simplifying conversion, then processing becomes easier, but information loss and distortion occur
Solution Approach 1:
The patent transitions from 2D landmark alignment to 3D spatial reasoning by introducing depth information and spatial relationships. The system models face landmarks in three-dimensional space, preserving spatial accuracy and avoiding the information loss inherent in 2D simplifying conversions while maintaining computational feasibility.
3Device complexity
If reference point is located at virtual position, then calculation is simplified, but control of speaking part movement becomes difficult
Solution Approach 1:
The patent segments the face into multiple functional regions with dedicated reference points, including a specific reference point for the speaking part (mouth area) separate from the overall face reference point. This segmentation enables independent control of lip movements while maintaining overall face structure, achieving both computational efficiency and precise control.
Data Source
AI summary
A speech video generation device according to an embodiment includes a first encoder, which receives an input of a person background image that is a video part in a speech video of a predetermined person, and extracts an image feature vector from the person background image, a second encoder, which receives an input of a speech audio signal that is an audio part in the speech video, and extracts a voice feature vector from the speech audio signal, a combining unit, which generates a combined vector by combining the image feature vector output from the first encoder and the voice feature vector output from the second encoder, a first decoder, which reconstructs the speech video of the person using the combined vector as an input, and a second decoder, which predicts a landmark of the speech video using the combined vector as an input.


