Speech Image Synthesis via Masked Region Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating speech moving images are complex and often produce unnatural results due to the need for key point extraction and alignment, and may not accurately reflect the natural movement of a person during speech.
Innovation Solution
A learning device with a first machine learning model that generates a mask for the speech portion of a person's image, allowing for the synthesis of a background image that includes the face and upper body, enabling the generation of a more natural speech moving image by reflecting unique gestures and movements without a separate key point prediction process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If key point extraction and alignment processes are used to generate speech moving images, then the accuracy of matching voice and image is improved, but the complexity of the procedure increases
Solution Approach 1:
The patent extracts only the necessary facial region from the original image using a mask generated by a machine learning model, removing the need for complex key point extraction and alignment processes. The mask isolates the face portion that needs to be synthesized, simplifying the overall procedure while maintaining accuracy.
Solution Approach 2:
The patent uses a pre-generated mask as a template to guide the synthesis process. Instead of performing complex transformations on the entire image, the system copies and synthesizes only the masked facial region, significantly reducing procedural complexity while preserving matching accuracy.
2Device complexity
If only the face portion is cut off and aligned for synthesis, then the procedure is simplified, but the natural movement of the person is not reflected
Solution Approach 1:
The patent applies different processing qualities to different regions: the masked face portion undergoes detailed synthesis to capture natural movements and expressions, while the rest of the image remains unchanged. This localized approach ensures natural movement representation without requiring complex global processing.
Solution Approach 2:
The machine learning model acts as an intermediary that generates masks capturing the essential facial features and movements. This intermediary structure enables the system to preserve natural movements without requiring direct complex transformation of the entire image, bridging the gap between simplicity and realism.
3Measurement precision
If a mask is generated to mask the speech-related portion, then the synthesis accuracy is improved, but the neural network structure becomes more complex
Solution Approach 1:
The patent segments the image processing task into two distinct components: mask generation and image synthesis. The mask generation module creates a binary mask identifying the face region, while the synthesis module processes only this masked region. This segmentation improves synthesis accuracy by focusing computational resources on the relevant area while managing network complexity through functional separation.
Data Source
AI summary
A learning device for generating an image according to an embodiment disclosed is a computing device including one or more processors and a memory storing one or more programs executed by the one or more processors. The learning device includes a first machine learning model that generates a mask for masking a portion related to speech in a person basic image with the person basic image as an input, and generates a person background image by synthesizing the person basic image and the mask.


