Generative image super-resolution method based on diffusion model
Through the consistent path matching super-resolution method, the mapping function from low-resolution images to high-resolution images is directly learned, solving the problem of high computing costs in the prior art, and generating high-resolution images with high realistic sense of high-resolution images under limited computing resources.
Patent Information
- Application Number
- CN202510348172.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
Existing super-resolution methods based on diffusion models are computationally costly in the inference process, making it difficult to generate high-resolution images with high realistic sense of high resolution under limited computing resources.
A consistent path matching super-resolution (CTMSR) method is proposed. By constructing a diffusion model and leveraging consistency loss and distributed path matching loss, we can directly learn the mapping function from low-resolution images to high-resolution images to achieve single-step inference.
This method can generate high-resolution images with only one-step inference, avoiding the limitations of distillation and fixed backbone networks, and significantly reducing computational costs.
Smart Images

Figure CN120198291A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image super-resolution, and particularly relates to a generative image super-resolution method based on a diffusion model. Background Art
[0002] The image super-resolution task aims to generate a corresponding high-resolution image based on the input low-resolution image. As is well known, the image super-resolution task is a typical ill-posed problem in the field of low-level vision because each low-resolution image may correspond to multiple potential high-resolution images.
[0003] Early classical super-resolution methods mainly restored high-resolution images by optimizing the root mean square error loss function. This supervised learning method forced the model to learn the expected value of all possible high-resolution images, ultimately resulting in blurred super-resolution results. In contrast, generative super-resolution methods aim to generate a high-resolution image estimation result that best conforms to the natural image distribution, thereby obtaining a more realistic high-resolution image.
[0004] In recent years, diffusion models have demonstrated powerful capabilities in modeling complex distributions (such as natural image distributions), and thus have great potential in generative super-resolution tasks. Early research on super-resolution methods based on diffusion models mainly adopted two approaches: 1. Conditioning on the low-resolution image and training as a conventional diffusion model (such as DDPM); 2. Using a pre-trained diffusion model as a prior and adjusting the reverse process under the guidance of the low-resolution image. Although these methods can achieve good results, their inference processes usually require hundreds of steps of calculation, resulting in extremely high computational costs.
[0005] Therefore, researchers have proposed various methods to accelerate the inference of diffusion model-based super-resolution: some studies (such as DDIM, DPMSolver) have explored more advanced inference strategies to reduce the sampling steps; other studies (such as ResShift) have modeled the initial state of the diffusion process, setting it as a low-quality image with a small amount of noise instead of pure noise, thereby significantly reducing the number of inference steps required by generative super-resolution methods. SinSR further reformulated the inference process of ResShift as an ordinary differential equation (ODE) and directly distilled it into one step. However, as pointed out by Rectified Flow, the performance of the model distilled into one step is limited by the teacher model; if the ODE path is not optimized to be close to a straight line during the training process of the teacher model, direct distillation may only yield sub-optimal results. In addition, the distillation process of the teacher model involves multi-step sampling to generate training data pairs, thus significantly increasing the training cost.
[0006] In addition to the above methods, some SR methods based on Stable Diffusion utilize the powerful generative ability of the pre-trained StableDiffusion model to achieve impressive super-resolution results in single-step inference. However, due to the dependence of these methods on a fixed backbone network and the difficulty in scaling to smaller models, there are certain limitations in practical applications. How to design a generative super-resolution method for single-step inference to generate high-fidelity restoration results at a limited computational cost without relying on distillation and a fixed backbone network remains an unsolved challenge. Summary of the Invention
[0007] To solve the above problems, the present invention proposes an efficient generative super-resolution method: Consistency Path Matching Super-Resolution (CTMSR), which can generate high-resolution images with high perceptual quality in only one-step inference. The technical solution of the present invention is as follows:
[0008] A generative image super-resolution method based on a diffusion model, comprising the following steps:
[0009] Step 1: Obtain an image dataset, including the low-resolution image to be processed;
[0010] Step 2: Use a U-Net network to generate a high-resolution image;
[0011] Step 3: Construct a diffusion model, specifically perform forward diffusion on the high-resolution image obtained in Step 2, and obtain the super-resolution result through consistency loss and forward diffusion;
[0012] Step 4: First, construct a forward diffusion process based on the super-resolution result obtained in Step 3 as the diffusion model of the false distribution path; minimize the distance loss between the diffusion model obtained in Step 3 and the diffusion model of the false distribution path to obtain a trained consistency-enhanced diffusion model for the image generative super-resolution task.
[0013] The specific implementation of Step 3 is as follows:
[0014] First, formulate the forward diffusion process:
[0015]
[0016] where x t is the noise map obtained from the forward diffusion process, e0 = y0 - x0 is the residual term, ∈ is Gaussian white noise; y0 represents the original low-resolution image, x0 represents the high-resolution image obtained after passing through the U-Net network, and α(t) and σ(t) are preset coefficient functions;
[0017] Then, train the diffusion model f through consistency lossθ Model the distribution conversion path, where θ is the model parameter, and the consistency loss formula is as follows:
[0018] L CT = [d(f θ (x t , y0, t), f θ- (x t-1 , y0, t - 1))]
[0019] where d(·, ·) represents a predefined function for measuring the distance between samples, f θ- represents the diffusion model that stops the propagation of gradients, t represents the number of forward diffusion steps; after training is completed, use f θ′ (x T , y0, T) to directly estimate the super-resolution result T represents the maximum number of steps in the forward diffusion process.
[0020] The specific steps of step 4 are as follows:
[0021] Construct a diffusion model based on the false distribution path, and the forward diffusion process is expressed as follows:
[0022]
[0023] where x t is the noise map obtained from the forward diffusion process based on the false distribution path, represents the residual corresponding to the super-resolved image;
[0024] Then construct a loss function to align the path from t to the false distribution and the path from x
[0025]
[0026] represents the mean square error, and keep the parameter θ ′ unchanged by only updating the parameter θ, and finally obtain the optimal solution:
[0027] θ’ = argmin θ L DTD
[0028] Finally, obtain the trained consistency-enhanced diffusion model for the image generation-based super-resolution task.
[0029] The U-Net network specifically includes a mapping module, an encoder, and a decoder;
[0030] The mapping module is responsible for processing noise and conditional information, mapping the noise conditions to the embedding space through the position encoding layer and the linear layer to obtain the noise condition embedding vector;
[0031] The encoder fuses the image features and the noise condition embedding vector and gradually performs feature extraction and downsampling through the UNetBlock module;
[0032] The decoder gradually restores the image resolution in a symmetric manner to the encoder, and uses skip connections to transfer the low-level features extracted by the encoder to the decoder, finally obtaining the super-resolution result.
[0033] The UNetBlock module is specifically as follows:
[0034] First, the UNetBlock module normalizes the image features and passes through the convolutional layer;
[0035] Then, a linear transformation is performed on the noise condition embedding vector to map it to the same dimension as the image features, and the two are added to obtain the image features after noise embedding;
[0036] The embedded image features pass through the activation function, convolutional layer and dropout layer, then perform skip connection with the embedded image features, enter the self-attention mechanism module for feature aggregation, and finally perform residual connection with the input features through the convolutional layer, and are scaled as the output of the UNetBlock module.
[0037] The features extracted by each UNetBlock module of the encoder will be stored. In each decoding stage, the decoder first performs UNetBlock module processing, then extracts the features of the corresponding layer of the encoder from the stored variables, and fuses them with the current decoding layer through skip connections.
[0038] The beneficial effects of the present invention are as follows:
[0039] Different from the model distilled into one step from a pre-trained generative super-resolution model, the method of the present invention utilizes the latest consistency training technology to directly learn the mapping function from a noisy low-resolution image to a high-resolution image. The proposed consistency training strategy enables the model to directly learn the probability flow ordinary differential equation (PF-ODE) path, thereby eliminating the limitation of relying on a pre-trained multi-step diffusion model. In addition, based on the learned PF-ODE path, the noisy low-resolution image distribution can be transformed into a natural image distribution, and a distribution path matching (DTM) loss is proposed to further improve the perceptual quality of the super-resolution results. The distribution path matching loss minimizes the distribution difference between them by matching the paths from the noisy low-resolution image distribution to the super-resolution results and the natural image distribution at the distribution path level, thereby improving the perceptual quality of the super-resolution results. Description of the Drawings
[0040] Figure 1 It is a schematic diagram of the consistency training and distribution path matching algorithms.
[0041] Figure 2 It is a flowchart of the consistency path matching super-resolution algorithm. Detailed Description of the Invention
[0042] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a generative image super-resolution method based on a diffusion model of the present invention with reference to the drawings.
[0043] The specific steps of this method are as follows:
[0044] Step 1: Obtain an image dataset for super-resolution training. To achieve generative image super-resolution, the training set of the ImageNet open-source dataset is used for training. This dataset includes 1.28 million pictures, which are randomly cropped into high-resolution pictures with a resolution of 256×256 during training, and then low-resolution pictures with a resolution of 64×64 are synthesized according to the image degradation process of RealESRGAN.
[0045] Step 2: Network structure design, using a U-Net network to generate high-resolution images. This method adopts a U-Net type network architecture, aiming to improve the performance in image processing tasks, especially image super-resolution and generation tasks. The network includes a mapping module, an encoder, and a decoder;
[0046] The mapping module of the network is responsible for processing noise and conditional information. First, the noise condition is processed through a positional encoding layer (PositionalEmbedding) to ensure the full utilization of spatial information; then it is further mapped to the embedding space through linear layer 0 and linear layer 1 to prepare for subsequent feature extraction.
[0047] In the encoder part, through multiple UNetBlock modules, the network can capture both the low-level features and global information of the image simultaneously. Specifically, the input image features and the noise condition embedding vector pass through convolutional layers and are gradually subjected to feature extraction and downsampling through the UNetBlock modules. The encoder consists of multiple resolution levels, and the number of channels at each level is adjusted according to the number of channels pre-set for the current resolution. Features at certain specific resolutions also pass through self-attention layers to enhance the global feature capture ability. The features extracted at each stage of the encoder are stored for later use in the decoding process.
[0048] Furthermore, UNetBlock is the core computational unit of this network structure, responsible for feature transformation, skip connections, and an optional self-attention mechanism. Its main function is to perform feature extraction and fusion at different scales and ensure the effective transmission of information through skip connections. The input features of this module include image features and the noise condition embedding vector, and through a series of normalization, convolution, affine transformation, and attention mechanisms, the transformed features are finally output. The specific structure of UNetBlock is as follows:
[0049] First, UNetBlock normalizes the image features and performs the first 3×3 convolution using a convolutional layer; then, it performs a linear transformation on the noise condition embedding vector to map it to the same dimension as the image features, and adds the two to obtain the image features after noise embedding. Next, after passing through the silu activation function, the embedded features pass through a convolutional layer for the second 3×3 convolution, and at the same time, a dropout layer is applied to increase the regularization ability. Subsequently, the features here are subjected to a skip connection with the input features to retain the input information. Then it enters the self-attention mechanism module, where the features after the skip connection are normalized and query (q), key (k), and value (v) are generated through a linear layer, then the attention weights are calculated and feature aggregation is performed, and finally, a transformation is carried out through convolution to make the output have a residual connection with the input features. Finally, the features are scaled and used as the output of UNetBlock to complete the entire feature processing process. The design of this module combines residual connections and the attention mechanism, ensuring effective information propagation and enhanced feature expression ability under multi-scale conditions.
[0050] The decoder part gradually restores the image resolution in a symmetric way to the encoder, and uses skip connections to transfer the low-level features extracted by the encoder to the decoder, ensuring the retention of details. Specifically, in each decoding stage, the decoder first performs UNetBlock processing, then extracts the features of the corresponding layer of the encoder from the stored variables, and fuses them with the current decoding layer through skip connections. This ensures that the high-resolution detail information can be fully utilized during the decoding process. The final output is processed through normalization and convolution to obtain the super-resolution result.
[0051] Overall, the network design ensures both fine detail restoration and understanding of the overall image structure in image processing tasks through hierarchical feature extraction and global information modeling. Combining the local feature extraction ability of U-Net and the global modeling advantages of the self-attention mechanism, this network can effectively improve the image quality and restore more details in image super-resolution and restoration tasks.
[0052] Step 3: Construct a diffusion model, specifically perform forward diffusion on the high-resolution image obtained in Step 2, and obtain the super-resolution result through consistency loss and forward diffusion; in order to better utilize the prior information of the low-resolution image, a specific forward diffusion process is first formulated for the super-resolution task:
[0053]
[0054] Here, x t is the noise map obtained from the forward diffusion process, e0 = y0 - x0 is the residual term, and ∈ is Gaussian white noise. y0 represents the low-resolution image, i.e., the original image, x0 represents the high-resolution image obtained after passing through the U-Net network, and α(t) and σ(t) are preset coefficient functions that control the growth rates of the residual and noise respectively. The above diffusion process defines the distribution conversion path from y0 to x0. Classical diffusion methods train neural networks through multiple iterative samplings to achieve generative image super-resolution. To achieve single-step diffusion super-resolution, a neural network-based diffusion model f θ is trained through a consistency training method to model the distribution conversion path, where θ is the model parameter. During the training process, as Figure 1 and Figure 2 shown, two adjacent points (i.e., x t and x t-1 ) are randomly selected, and the following formula is used to minimize the consistency diffusion enhancement loss between them:
[0055] L CT = [d(f θ (x t , y0, t), f θ- (x t-1 , y0, t - 1))]
[0056] Among them, d(·,·) represents a predefined function for measuring the distance between samples, and f θ- represents a diffusion model that stops the propagation of gradients, and t represents the number of forward diffusion steps. The above training process constrains the consistency of different sampling points in the distribution mapping path to the final distribution mapping, facilitating the neural network to directly learn the single-step mapping process. After the training is completed, f θ′ (x T , y0, T) can be directly used to estimate the super-resolution result Here, T represents the maximum number of steps in the forward diffusion process, is similar to the high-quality representation content of the original high-resolution image x0, and at the same time conforms to the distribution characteristics of the high-quality representation, which helps to generate super-resolution results that conform to the human visual perception preference.
[0057] Step 4: Enhance the perceptual quality of the super-resolution result based on the diffusion path matching method; first, construct a forward diffusion process based on the super-resolution result obtained in Step 3 as the diffusion model of the false distribution path; minimize the distance loss between the diffusion model obtained in Step 3 and the diffusion model of the false distribution path to obtain a trained consistency-enhanced diffusion model for the image generation-based super-resolution task. Specifically, based on the distribution path constructed by the consistency training method, the diffusion path matching (Trajectory Matching) method is further proposed, and the perceptual quality of the reconstructed features is further enhanced by designing the loss function L DTD First, the model f θ′ obtained by the consistency training method can represent the distribution path from the low-resolution image to the high-resolution image, such as Figure 2 shown as f θ′ (x t , y0, t). Similarly, a function can also be constructed to represent the false distribution path to the enhanced pseudo-real distribution: Among them and x t follow the same forward diffusion process:
[0058]
[0059] Among them, x t is the noise map obtained from the forward diffusion process based on the false distribution path, represents the result of the diffusion model super-resolution, represents the residual corresponding to the super-resolved image. In order to make the result restored by the model closer to the natural image distribution at the distribution level, a loss function is constructed to the path to the false distribution and x tAlign the path to the true distribution. Specifically, the present invention aims to minimize f θ’ (x t , y0, t) and The distribution path distance loss function between them is as follows, as Figure 2 shown, and the specific expression is as follows:
[0060]
[0061] represents the mean square error, and by only updating the parameter θ, the parameter θ ′ is kept unchanged, and finally the optimal solution is obtained:
[0062] θ’ = argmin θ L DTD
[0063] This loss function can further enhance the super-resolution result on the basis of the consistency diffusion model, making it more in line with the natural image distribution and more in line with the human visual perception in terms of visual appearance. Finally, a trained consistency-enhanced diffusion model is obtained for the image generation-based super-resolution task.
[0064] During the training process of the model, the Adam optimizer is used for optimization. The entire training process is carried out at a constant learning rate of 5e-5, and the batch size is set to 32. First, the consistency training strategy is adopted to train for 500,000 iteration cycles, and then the consistency loss and the path matching loss function are jointly optimized, and training is continued for 10,000 iteration times.
[0065] A large number of experimental results on the synthetic dataset and the real dataset clearly prove the superiority of this method. With less inference calculation amount, the method proposed by the present invention can generate state-of-the-art high-fidelity super-resolution results.
[0066] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A generative image super-resolution method based on a diffusion model, characterized in that: The following steps are involved: Step 1: Obtain an image dataset, including the low-resolution image to be processed; Step 2: Generate high-resolution images using the U-Net network; Step 3: Construct a diffusion model, specifically perform forward diffusion on the high-resolution image obtained in step 2, and obtain a super-resolution result through consistency loss and forward diffusion; Step 4: First, a forward diffusion process is constructed based on the super-resolution result obtained in step 3 as the diffusion model of the false distribution path; the distance loss between the diffusion model obtained in step 3 and the diffusion model of the false distribution path is minimized to obtain a trained consistency-enhanced diffusion model for the image generation super-resolution task.
2. The generative image super-resolution method based on a diffusion model according to claim 1, characterized in that: The step 3 is as follows: First formulate the forward diffusion process: Among them, x t is the noise map obtained by the forward diffusion process, e0=y0-x0 is the residual term, ∈ is Gaussian white noise; y0 represents the original low-resolution image, x0 represents the high-resolution image obtained after passing through the U-Net network, α(t) and σ(t) are pre-set coefficient functions; Then the diffusion model f is trained by the consistency loss θ To model the distribution transformation path, where θ is the model parameter, the consistency loss formula is as follows: L CVT =[d(f θ (x t ,y0,t),f θ- (x t-1 ,y0,t-1))] Where d(·,·) represents a predefined function used to measure the distance between samples, and f θ- Represents a diffusion model that stops propagating gradients, and t represents the number of forward diffusion steps; after training, use f θ′ (x T ,y0,T) directly estimates the super-resolution result T represents the maximum number of forward diffusion process steps.
3. The generative image super-resolution method based on a diffusion model according to claim 2, characterized in that: The step 4 is specifically as follows: A diffusion model based on the false distribution path is constructed, and the forward diffusion process is expressed as follows: Among them, x t is the noise map obtained by the forward diffusion process based on the false distribution path, Represents the residual corresponding to the image after super resolution; Then construct the loss function The path to the spurious distribution and x t The path to the true distribution is aligned, and the loss function is specifically expressed as follows: represents the mean square error, by updating only the parameter θ to keep the parameter θ ′ The optimal solution is finally obtained: <h2 style=";text-align:left;direction:ltr">θ' = argmin<h2 style=";text-align:left;direction:ltr"> θ <h2 style=";text-align:left;direction:ltr"> L<h2 style=";text-align:left;direction:ltr"> DTD Finally, a trained consistency-enhanced diffusion model is obtained for image generative super-resolution tasks.
4. The generative image super-resolution method based on a diffusion model according to claim 3, characterized in that: The U-Net network specifically includes a mapping module, an encoder, and a decoder; The mapping module is responsible for processing noise and condition information, mapping the noise condition to the embedding space through the coding layer and the linear layer to obtain the noise condition embedding vector; The encoder fuses the image features and the noise condition embedding vector and gradually performs feature extraction and downsampling through the UNetBlock module; The decoder gradually restores the image resolution in a symmetrical manner to the encoder, and uses skip connections to pass the low-level features extracted by the encoder to the decoder, ultimately obtaining a super-resolution result.
5. The generative image super-resolution method based on a diffusion model according to claim 4, characterized in that: The UNetBlock module is as follows: First, the UNetBlock module normalizes the image features and passes through the convolution layer; Then, the noise condition embedding vector is linearly transformed to map it to the same dimension as the image feature, and the two are added to obtain the image feature after noise embedding; The embedded image features pass through the activation function, convolution layer and dropout layer, and then make a jump connection with the embedded image features, enter the self-attention mechanism module for feature aggregation, and finally make a residual connection with the input features through the convolution layer, and after scaling, serve as the output of the UNetBlock module.
6. The generative image super-resolution method based on a diffusion model according to claim 5, characterized in that: The features extracted by each UNetBlock module of the encoder are stored. At each decoding stage, the decoder first performs UNetBlock module processing, then extracts the features of the corresponding layer of the encoder from the stored variables, and fuses them with the current decoding layer through skip connections.
Citation Information
Cited By
Image decomposition method, device and computer program product
CN121053146A
MRI image super-resolution method and system based on single-step diffusion
CN121258797A
A single-step diffusion-based MRI image super-resolution method and system
CN121258797B