Latent Structural Diffusion Model for Anatomical Coherence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated image generation systems, particularly those using text-to-image diffusion models, struggle to produce high-fidelity and anatomically coherent human images due to the complexity of human anatomy and the difficulty in conveying structural information through text prompts.
Innovation Solution
A latent structural diffusion model is introduced, which conditions image generation on both explicit appearance and structural aspects such as depth information and surface normal maps, using a pose map to provide structural guidance and ensure anatomical coherence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If text-to-image diffusion models are used for automated image generation, then the system can automatically generate images from natural language descriptions, but the generated images lack anatomical coherence and structural integrity
Solution Approach 1:
The patent introduces pose maps as an intermediary between the text prompt and the image generation process. The pose map encodes structural information about human anatomy and body configuration, serving as a mediator that guides the diffusion model to generate anatomically coherent images while maintaining automatic generation capabilities
Solution Approach 2:
The patent performs preliminary encoding of structural information into pose maps before the actual image generation process. By pre-processing and encoding anatomical structures into the pose map, the system prepares structural guidance in advance, ensuring that the diffusion model receives both textual and structural constraints simultaneously
2Ease of operation
If text prompts alone are used to convey structural information, then the input process remains simple, but the structural information is insufficient for generating coherent human images
Solution Approach 1:
The patent segments the input requirements into two distinct components: text prompts for semantic content and pose maps for structural information. This segmentation allows the system to maintain simple text-based user input while internally utilizing detailed pose map representations to preserve complete structural information during generation
3Productivity
If existing diffusion models generate human images, then the generation process is fast, but the output images contain disjointed body parts and unnatural poses
Solution Approach 1:
The pose map acts as a continuous intermediary throughout the generation process, providing ongoing structural constraints that prevent disjointed body parts and unnatural poses while maintaining the efficiency of the diffusion model's generation speed
Solution Approach 2:
The patent incorporates structural feedback through the pose map, which provides continuous guidance during the diffusion process. The pose map serves as a feedback mechanism that constantly reinforces anatomical coherence, ensuring that generated images maintain structural integrity without significantly slowing down the generation process
Data Source
AI summary
Examples described herein relate to automatic image generation. A plurality of inputs is accessed. The inputs include first input data and second input data. The first input data includes a text prompt describing a desired image and the second input data is indicative of one or more structural features of the desired image. One or more intermediate outputs are generated via a first generative machine learning model that uses the plurality of inputs as first control signals. An output image is generated via a second generative machine learning model that uses at least a subset of the plurality of inputs and at least a subset of the one or more intermediate outputs as second control signals. The output image is presented at a user device of a user.


