Diffusion Scene Generation for High-Resolution Human Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-image conversion models struggle with limited resolution and text-image correspondence in complex scenes, leading to distorted and unnatural images, especially when enlarging images to high resolution and due to a limited number of tokens in text encoders.
Innovation Solution
The method employs high-frequency noise injection and adaptive joint diffusion techniques, using window-based latent vector reconstruction to maintain high resolution and natural text-image correspondence, even in complex scenes with multiple instances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If existing text-to-image conversion models are used, then simple scenes can be generated, but high-resolution complex scenes with multiple human instances cannot be generated naturally
Solution Approach 1:
The patent divides the complex scene generation task into multiple stages: first generating a base image at lower resolution, then progressively refining it through multiple diffusion steps. The image is processed in segments through window-based operations and latent vector manipulations at different resolution levels, allowing detailed generation of multiple human instances while maintaining overall scene coherence and accurate text-image correspondence.
2Manufacturing precision
If image is enlarged to high resolution, then detail is improved, but distortion and unnatural appearance occur
Solution Approach 1:
The patent performs preliminary diffusion processing at lower resolution to establish the overall scene structure and object arrangements before upsampling. By pre-establishing the compositional framework and then progressively refining details through controlled diffusion steps, the method avoids distortion that would occur from direct upsampling, maintaining natural appearance while achieving high resolution.
3Productivity
If number of tokens in text encoder is limited, then processing speed is maintained, but complex scenes with multiple instances cannot be accurately represented
Solution Approach 1:
The patent moves the detailed representation task from the text encoder's token space to the image generation space. Instead of attempting to encode all scene details into limited text tokens, the method uses the diffusion model's latent space to represent complex scene details, allowing rich scene representation while keeping text encoder processing efficient. The text prompt provides high-level guidance while the diffusion process generates detailed scene content.
Data Source
AI summary
The present disclosure relates to technology that generates a high-resolution image from a text prompt by applying an adaptive joint diffusion technique to a latent vector injected with high-frequency noise to generate a high-resolution human-centric scene.


