Speech Video Generation With Fused Features for Landmark Continuity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech video generation technologies face challenges in synthesizing high-quality lip sync face images due to annotation noise in face landmark data, leading to unstable continuity over time and image quality deterioration.

Innovation Solution

A speech video generation device that combines image feature vectors from person background images with voice feature vectors from speech audio signals to generate a combined vector, which is then used to reconstruct the speech video and predict the landmark, thereby improving prediction accuracy and image quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If face landmark data with annotation noise is used for synthesis, then the synthesis process can proceed, but image quality deteriorates and continuity becomes unstable

Engineering Contradiction:
Improvesynthesis process efficiencyVSAvoidimage quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent extracts and removes annotation noise from face landmark data through a learning model that distinguishes between valid landmark information and noise. The system separates the noise component from the original landmark data to produce cleaned landmark sequences that maintain continuity and accuracy for synthesis.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements a feedback mechanism where the learning model continuously refines landmark prediction by comparing generated results with ground truth data. The model uses error feedback to adjust its predictions, reducing annotation noise accumulation and improving temporal consistency across video frames.

Inventive Principle:
Principle #23Feedback

2Ease of operation

If face landmark is aligned in standard space with simplifying conversion, then processing becomes easier, but information loss and distortion occur

Engineering Contradiction:
Improveprocessing simplicityVSAvoidlandmark accuracy
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent transitions from 2D landmark alignment to 3D spatial reasoning by introducing depth information and spatial relationships. The system models face landmarks in three-dimensional space, preserving spatial accuracy and avoiding the information loss inherent in 2D simplifying conversions while maintaining computational feasibility.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Device complexity

If reference point is located at virtual position, then calculation is simplified, but control of speaking part movement becomes difficult

Engineering Contradiction:
Improvecalculation complexityVSAvoidcontrol precision
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The patent segments the face into multiple functional regions with dedicated reference points, including a specific reference point for the speaking part (mouth area) separate from the overall face reference point. This segmentation enables independent control of lip movements while maintaining overall face structure, achieving both computational efficiency and precise control.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12347197B2Device and method for generating speech video along with landmark
Publication Date: 2025.07.01 DEEPBRAIN AI INC
  • US12347197B2 patent drawing
  • US12347197B2 patent drawing
  • US12347197B2 patent drawing

AI summary

A speech video generation device according to an embodiment includes a first encoder, which receives an input of a person background image that is a video part in a speech video of a predetermined person, and extracts an image feature vector from the person background image, a second encoder, which receives an input of a speech audio signal that is an audio part in the speech video, and extracts a voice feature vector from the speech audio signal, a combining unit, which generates a combined vector by combining the image feature vector output from the first encoder and the voice feature vector output from the second encoder, a first decoder, which reconstructs the speech video of the person using the combined vector as an input, and a second decoder, which predicts a landmark of the speech video using the combined vector as an input.