Speech Moving Image Generation via Dual Encoder Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating speech moving images are complex and often result in unnatural images due to the need for key point extraction, alignment, and separate synthesis processes, and may degrade in quality if mask processing is not properly performed.

Innovation Solution

A method and device for generating speech moving images using a computing device with one or more processors and a memory, which includes a first encoder for compressing image feature vectors from a person background image, a second encoder for compressing voice feature vectors from a speech audio signal, and an image reconstruction unit that uses a combination vector to reconstruct the speech moving image, thereby simplifying the neural network structure and improving image quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If key point extraction, alignment, and separate synthesis processes are used to generate speech moving images, then the image can be synthesized to match input voice, but the procedure becomes complicated

Engineering Contradiction:
Improveimage synthesis accuracyVSAvoidprocedure complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges the key point extraction, alignment, and image synthesis processes into a single integrated neural network model. The encoder processes the input image to extract features, the transformer model performs alignment and synthesis in one unified operation, and the decoder reconstructs the output image, eliminating the need for separate processing steps while maintaining synthesis accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The transformer model serves multiple functions simultaneously: it performs key point alignment, speech synthesis, and temporal transformation within a single architectural framework. This multi-functional approach replaces the need for separate specialized processes, simplifying the overall procedure while maintaining the ability to generate accurate speech-moving images.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If only face portion is cut off and alignment is made according to size and position to generate speech moving image, then the procedure is simplified, but the result becomes unnatural

Engineering Contradiction:
Improveprocedure simplicityVSAvoidnaturalness of speech movement
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent replaces the mechanical approach of simple face cropping and geometric alignment with a learning-based system. The neural network automatically learns the complex transformations and alignments needed to generate natural speech movements, substituting manual geometric operations with intelligent data-driven processing that captures nuanced facial dynamics.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The transformer model dynamically adjusts multiple parameters including key point positions, facial feature coordinates, and temporal transformations based on the input speech signal. This allows the system to maintain procedural simplicity while achieving natural results through learned parameter transformations rather than fixed geometric rules.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If mask processing is not properly performed on person background image, then the processing is simpler, but the speech moving image quality degrades

Engineering Contradiction:
Improvemask processing simplicityVSAvoidspeech moving image quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The neural network performs preliminary feature extraction and separation during the encoding phase, preparing the image data in advance to handle potential mask processing issues. By pre-processing the image features and separating relevant from irrelevant information before the main synthesis operation, the system can maintain quality even when mask processing is not optimal.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The decoder provides feedback to the encoder by comparing the reconstructed image with the target, allowing the system to automatically adjust and compensate for inadequate mask processing. This feedback mechanism enables the network to learn from errors and improve the separation of speech-related features from background elements, maintaining image quality without requiring perfect mask processing.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12205212B2Method and device for generating speech moving image
Publication Date: 2025.01.21 DEEPBRAIN AI INC
  • US12205212B2 patent drawing
  • US12205212B2 patent drawing
  • US12205212B2 patent drawing

AI summary

A device which generates a speech moving image includes a first encoder, a second encoder, a combination unit, and an image reconstruction unit. The first encoder receives a person background image in which a portion related to speech of a person that is a video part of the speech moving image of the person is covered with a mask, extracts an image feature vector from the person background image, and compresses the extracted image feature vector. The second encoder receives a speech audio signal that is an audio part of the speech moving image, extracts a voice feature vector from the speech audio signal, and compresses the extracted voice feature vector. The combination unit generates a combination vector of the compressed image feature vector and the compressed voice feature vector. The image reconstruction unit reconstructs the speech moving image of the person with the combination as an input.