Speech Image Synthesis via Masked Region Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating speech moving images are complex and often produce unnatural results due to the need for key point extraction and alignment, and may not accurately reflect the natural movement of a person during speech.

Innovation Solution

A learning device with a first machine learning model that generates a mask for the speech portion of a person's image, allowing for the synthesis of a background image that includes the face and upper body, enabling the generation of a more natural speech moving image by reflecting unique gestures and movements without a separate key point prediction process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If key point extraction and alignment processes are used to generate speech moving images, then the accuracy of matching voice and image is improved, but the complexity of the procedure increases

Engineering Contradiction:
Improveaccuracy of matching voice and imageVSAvoidcomplexity of the procedure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the necessary facial region from the original image using a mask generated by a machine learning model, removing the need for complex key point extraction and alignment processes. The mask isolates the face portion that needs to be synthesized, simplifying the overall procedure while maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses a pre-generated mask as a template to guide the synthesis process. Instead of performing complex transformations on the entire image, the system copies and synthesizes only the masked facial region, significantly reducing procedural complexity while preserving matching accuracy.

Inventive Principle:
Principle #26Copying

2Device complexity

If only the face portion is cut off and aligned for synthesis, then the procedure is simplified, but the natural movement of the person is not reflected

Engineering Contradiction:
Improvesimplicity of the procedureVSAvoidnatural movement of the person
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies different processing qualities to different regions: the masked face portion undergoes detailed synthesis to capture natural movements and expressions, while the rest of the image remains unchanged. This localized approach ensures natural movement representation without requiring complex global processing.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The machine learning model acts as an intermediary that generates masks capturing the essential facial features and movements. This intermediary structure enables the system to preserve natural movements without requiring direct complex transformation of the entire image, bridging the gap between simplicity and realism.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If a mask is generated to mask the speech-related portion, then the synthesis accuracy is improved, but the neural network structure becomes more complex

Engineering Contradiction:
Improvesynthesis accuracyVSAvoidneural network structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the image processing task into two distinct components: mask generation and image synthesis. The mask generation module creates a binary mask identifying the face region, while the synthesis module processes only this masked region. This segmentation improves synthesis accuracy by focusing computational resources on the relevant area while managing network complexity through functional separation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12131441B2Learning device and method for generating image
Publication Date: 2024.10.29 DEEPBRAIN AI INC
  • US12131441B2 patent drawing
  • US12131441B2 patent drawing
  • US12131441B2 patent drawing

AI summary

A learning device for generating an image according to an embodiment disclosed is a computing device including one or more processors and a memory storing one or more programs executed by the one or more processors. The learning device includes a first machine learning model that generates a mask for masking a portion related to speech in a person basic image with the person basic image as an input, and generates a person background image by synthesizing the person basic image and the mask.