Patch-Wise Diffusion Image Generation to Reduce High-Resolution Compute
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional diffusion models are limited to generating low-resolution images and require extensive computational resources and multiple stages to produce high-resolution images, making them inefficient and costly.
Innovation Solution
A diffusion model trained with mixed-resolution datasets to generate high-resolution images by creating and combining image patches, using a global image code and local diffusion processes to ensure consistency, allowing single-stage training and efficient high-resolution image synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If conventional diffusion models are used to generate high-resolution images, then image resolution is improved, but computational resources and time required increase significantly
Solution Approach 1:
The patent divides the image generation process into multiple stages, starting with low-resolution patches and progressively refining to high-resolution images. The diffusion model generates images at different resolutions in sequence, combining multiple low-resolution generations into a final high-resolution output, thereby reducing the computational burden of direct high-resolution generation
Solution Approach 2:
The patent performs preliminary generation at low resolution before final high-resolution output. By first generating low-resolution images and using them as conditions for subsequent high-resolution generation, the system prepares intermediate results that guide the final high-resolution synthesis, reducing overall computational requirements
2Manufacturing precision
If multiple stages are used to generate high-resolution images, then image quality is improved, but process complexity increases
Solution Approach 1:
The patent employs a single diffusion model that serves multiple functions across different resolution stages. The same model architecture is used for both low-resolution and high-resolution generation, with resolution being controlled by input parameters rather than requiring separate specialized models for each stage
Solution Approach 2:
The patent combines multiple low-resolution image generations into a single high-resolution output by using one generated image as a condition for the next generation stage. This merging approach integrates results from multiple stages while maintaining a unified process framework
3Manufacturing precision
If extensive computational resources are allocated to high-resolution image generation, then image resolution is improved, but cost increases
Solution Approach 1:
By segmenting the computation into progressive resolution stages, the patent avoids the exponential computational cost of direct high-resolution generation. Each stage operates at a manageable resolution level, accumulating computational effort incrementally rather than requiring massive resources upfront
Solution Approach 2:
The patent uses partial action by generating images at intermediate resolutions that are sufficient for guiding the final high-resolution synthesis. Rather than performing full high-resolution computation at every stage, the method uses lower-resolution intermediates that provide adequate guidance while consuming fewer resources
Data Source
AI summary
Aspects of the methods, apparatus, non-transitory computer readable medium, and systems include obtaining a noise map and a global image code encoded from an original image and representing semantic content of the original image; generating a plurality of image patches based on the noise map and the global image code using a diffusion model; and combining the plurality of image patches to produce an output image including the semantic content.


