Real image super-resolution method based on uncertainty diffusion model
By constructing a real image super-resolution method based on the uncertainty diffusion model, using the uncertainty prediction module to add different noise intensities to different regions, the problem of insufficient information utilization in the existing methods is solved, the image super-resolution quality is improved and the calculation cost is reduced.
Patent Information
- Application Number
- CN202510341650.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing real image super-resolution method based on diffusion models has limited ability to utilize low-resolution image information, especially in flat areas, insufficient information utilization and high computational cost, and large amount of calculation in the training process, which affects training efficiency.
The end-to-end depth real image super-resolution model is constructed, including uncertainty prediction module and uncertain diffusion process. Through pixel de-shuffling, denoising network, nearest neighbor interpolation and convolutional layer, the uncertainty prediction module is used to add noise of different intensities to different regions, the flat area retains more original image information, and the edge area increases noise intensity, thereby improving the information utilization efficiency of the diffusion process.
Improves the quality and performance of image super resolution, reduces computing requirements, and is suitable for practical applications.
Smart Images

Figure CN120278882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of real image super-resolution, and in particular to a real image super-resolution method based on an uncertainty diffusion model. Background Art
[0002] Single Image Super-Resolution (SR) aims to recover a clear High-Resolution (HR) image from a degraded and contaminated Low-Resolution (LR) image, which is a classical problem in the field of computer vision. Existing super-resolution methods have explored a variety of advanced network architectures and complex degradation models to improve performance in classical super-resolution and real-world super-resolution tasks respectively.
[0003] Early real image super-resolution method SRGAN (Super-Resolution Generative Adversarial Network) trains two models, a generator and a discriminator, simultaneously through generative adversarial training. The discriminator is used to make the distribution of the textures generated by the generator closer to the distribution of real images, so as to solve the problem that the details of the images generated by general image super-resolution models are too smooth. BSRGAN and RealESRGAN further improve the performance of real super-resolution methods by designing more complex training processes and constructing degradations closer to real images. In recent years, Diffusion Models have demonstrated impressive capabilities in image synthesis, providing a promising new approach for real-world image super-resolution tasks. These models transform pure Gaussian noise into high-quality images through a predefined Markov chain, and their theoretical basis endows them with powerful modeling capabilities to bridge different data distributions. To utilize the modeling capabilities of diffusion models to recover the missing details in LR images, some methods start the super-resolution process by sampling from the standard Gaussian distribution and gradually refine the noise input into a high-quality output. However, the methods starting from pure noise were originally designed for image synthesis tasks, resulting in poor performance in super-resolution tasks. In addition, these methods usually require a long sampling process, limiting their practicality in actual applications.
[0004] The super-resolution diffusion model SRDiff and the iterative refinement super-resolution diffusion model SR3 are the first methods to apply diffusion models to the image super-resolution task. They use the low-resolution input as conditional information, demonstrating the effectiveness of diffusion models in generating perceptually high-quality super-resolution images. Despite their superior performance, SR3 still has problems such as bias and a high computational cost during the sampling process. The latent space super-resolution diffusion model LDM-SR trains an autoencoder and performs the diffusion process in its low-dimensional latent space. This method enables the model to focus on perceptually relevant details and significantly improves computational efficiency. The residual transformation super-resolution diffusion model ResShift greatly simplifies the diffusion process by embedding the LR image into the initial noise map and gradually recovering the residual between the LR and HR images. It no longer needs to model the entire HR image from noise but only needs to estimate the LR-HR residual, thus shortening the sampling process and significantly enhancing the super-resolution effect. Although the performance has been improved, there are still challenges, especially when most details are severely masked by noise, which brings additional challenges because in the super-resolution task, using information from the surrounding area is crucial for recovering missing details.
[0005] Despite many advancements in enhancing the image super-resolution modeling ability based on diffusion models, in recent research, few methods have focused on the inherent information in LR images, that is, the flat areas are already close to the target, while the edge and texture areas are farther away. At the same time, recent diffusion model-based methods have all used the vector quantization generative adversarial network VQGAN to transform the diffusion process into the latent space to improve the ability to recover image details. However, when used simultaneously with the feature loss, i.e., the LPIPS loss (used to enhance perceptual quality), this method will cause a significant increase in the computational amount during the backpropagation process of network training, seriously affecting the training efficiency. Summary of the Invention
[0006] To solve the problem of the limited ability of existing diffusion model-based super-resolution methods to utilize LR image information, the present invention proposes a real image super-resolution method based on an uncertainty diffusion model, and the method includes the following steps:
[0007] Step S1, construct an end-to-end deep real image super-resolution model, the model includes an uncertainty prediction module and an uncertainty diffusion process, and the uncertainty diffusion process includes a pixel de-shuffle layer, a denoising network, a nearest neighbor interpolation, and a convolutional layer;
[0008] Step S2, construct a training dataset for the end-to-end deep real image super-resolution model, use the ImageNet open-source dataset as the training dataset, and perform supervised training on the denoising module of the diffusion process composed of the denoising network, the nearest neighbor interpolation, and the convolutional layer;
[0009] In step S3, the trained end-to-end deep real image super-resolution model is used for image super-resolution processing. The uncertainty prediction module predicts the uncertainty of each point in the image, and different intensities of noise are added according to the high or low uncertainty of each point. In the area with low uncertainty, that is, the smooth area in the image, a lower noise intensity, that is, a smaller noise weight, is adopted to retain more original image information. In the area with high uncertainty, that is, the edge and texture areas in the image, a higher noise intensity, that is, a larger noise weight, is adopted so that the diffusion process can better utilize the information in the low-resolution image and improve the super-resolution effect; the uncertainty diffusion process obtains the denoised high-resolution image through pixel de-shuffling, denoising, nearest neighbor interpolation, and convolutional upsampling.
[0010] Further, in step S3, the uncertainty prediction module first inputs the low-resolution image y0 into the image super-resolution network g to obtain a rough high-resolution image g(y0), and then calculates the difference between the high-resolution image g(y0) and the low-resolution image y0 upsampled by bicubic interpolation to obtain the uncertainty prediction value. The uncertainty prediction value is further transformed through a monotonically increasing function u to obtain the noise intensity weight w based on uncertainty u (y0) = u(ψ est (y0)), where u is expressed as:
[0011]
[0012] where ψ max and b u are hyperparameters. For the region where the uncertainty ψ > ψ max , the noise intensity weight is set to u(ψ) = 1. Finally, a noise intensity weight map with the same size as the input image and the weight values in the interval [b u , 1] is obtained.
[0013] Further, the image super-resolution network includes multiple feature extraction modules and an upsampling module. Each feature extraction module includes six self-attention layers and a convolutional layer for extracting image features. The upsampling module includes a convolutional layer and a pixel shuffling layer for improving the image resolution.
[0014] Further, in step S3, the uncertainty diffusion process is specifically as follows:
[0015] Multiply the noise intensity weight by Gaussian noise with a mean of 0 and a variance of k 2 η T I to obtain anisotropic noise with different variances in different regions, and add it to the low-resolution image y0 to obtain the noisy image x at the starting time T of the diffusion process T :
[0016]
[0017] Among them, T is a hyperparameter representing the number of steps of the diffusion model, and η T , κ are hyperparameters for controlling the noise addition and denoising rates of the diffusion process, and I is the identity matrix;
[0018] Starting from the initial time T, perform backward iteration. During each backward iteration of the diffusion process, the noisy image x t First, convert the image spatial dimension to the channel dimension through the pixel unshuffle layer, and then extract the noise-free image at time 0 through the denoising network f, where 1 ≤ t ≤ T:
[0019]
[0020] And obtain the probability distribution of the noisy image at time x t-1 through interpolation:
[0021]
[0022] Among them, η t , η t-1 , α t and κ are hyperparameters for controlling the noise addition and denoising rates of the diffusion process, I is the identity matrix, and then the feature map obtained by the denoising network is upsampled to the original size through nearest neighbor interpolation and convolutional layers to obtain the denoised image; finally, after T iterations, the final denoised image x0 is obtained.
[0023] Furthermore, the denoising network includes three layers of downsampling feature extraction modules, one layer of intermediate layer feature extraction module, and three layers of upsampling feature extraction modules. There are residual connections between the feature extraction modules corresponding to different scales. Among them, the downsampling feature extraction modules obtain feature maps of different scales, and then the intermediate layer feature extraction module refines the features of the smallest scale, and the upsampling feature extraction modules combine the downsampled depth image features with the feature maps of different scales to obtain a multi-scale fused feature map.
[0024] Furthermore, the training strategy in step S2 is: use the combination of pixel loss, that is, L1 loss, and feature loss, that is, LPIPS loss function, as the loss function of the denoising module in the diffusion process:
[0025]
[0026] Among them, the L1 loss is used to ensure the fidelity of the network output image, the LPIPS loss is used to ensure the perceptual quality of the network output image, and λ is a hyperparameter used to control the trade-off between fidelity and perceptual quality. During the training process of the model, the Adam optimizer is used for optimization.
[0027] By considering the different characteristics of different regions (parts with different uncertainties) in the image, and thus the different noises required in the actual diffusion process, the present invention proposes a real-image super-resolution method based on an uncertainty-guided anisotropic diffusion model to achieve higher-quality deep-learning real-image super-resolution. First, a simple image super-resolution network is used to estimate the uncertainties of different regions in the input image, and different intensities of noise are added according to the uncertainty level of each point. By introducing the uncertainty method, the division of different regions (smooth / edge regions) in the image is obtained, and the information in the low-resolution image is better applied to the diffusion process. Second, in order to use this uncertainty information in the diffusion process, the uncertainty value of each point is transformed into the corresponding noise intensity weight. A lower noise intensity, that is, a smaller noise weight, is used in the region with low uncertainty, namely the smooth region in the image, to retain more original image information. A higher noise intensity, that is, a larger noise weight, is used in the region with high uncertainty, namely the edge and texture regions in the image, so that the diffusion process can better utilize the information in the low-resolution image and improve the super-resolution effect. Finally, a denoised high-resolution image is obtained through the uncertainty-guided anisotropic diffusion process, which includes pixel de-shuffling, a denoising network, nearest-neighbor interpolation, and convolutional layers. Through the solution of the present invention, the diffusion process is more suitable for the super-resolution task, improving the performance and quality of image super-resolution and reducing the computing power requirements. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0029] Figure 1 It is a schematic flowchart of a real-image super-resolution method based on an uncertainty diffusion model provided by an embodiment of the present invention;
[0030] Figure 2 It is a schematic structural diagram of an image super-resolution network provided by an embodiment of the present invention;
[0031] Figure 3 It is a schematic structural diagram of a denoising network provided by an embodiment of the present invention. Detailed Embodiments
[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0033] The flow of a real image super-resolution method based on an uncertainty diffusion model provided in this embodiment is as Figure 1 shown, and the method includes the following steps:
[0034] Step 1: Construct an end-to-end deep real image super-resolution model, which includes an uncertainty prediction module and an uncertainty diffusion process. The uncertainty diffusion process includes a pixel de-shuffle layer, a denoising network, a nearest neighbor interpolation, and a convolutional layer;
[0035] Step 2: Construct a training dataset for the end-to-end deep real image super-resolution model. Use the ImageNet open-source dataset as the training dataset, and perform supervised training on the denoising module of the diffusion process composed of the denoising network, nearest neighbor interpolation, and convolutional layer. This dataset contains 1.25 million images. During training, randomly crop them into high-resolution images with a resolution of 256×256, and use the degradation process in RealESRGAN to generate low-resolution images with a size of 64×64 corresponding to each image, so as to form the output (256) and input (64) pairs of the training network. The degradation process includes blurring, up / downsampling, noise, and JPEG compression processes;
[0036] Step 3: Use the trained end-to-end deep real image super-resolution model for image super-resolution processing. The uncertainty prediction module predicts the uncertainty of each point in the image, and adds different intensities of noise according to the high and low uncertainty of each point. In the area with low uncertainty, that is, the smooth area in the image, a lower noise intensity, that is, a smaller noise weight, is used to retain more original image information. In the area with high uncertainty, that is, the edge and texture areas in the image, a higher noise intensity, that is, a larger noise weight, is used to enable the diffusion process to better utilize the information in the low-resolution image and improve the super-resolution effect; the uncertainty diffusion process obtains a denoised high-resolution image through pixel de-shuffle, denoising, nearest neighbor interpolation, and convolutional upsampling.
[0037] Specifically, the uncertainty prediction module first inputs the low-resolution image y0 into the image super-resolution network g to obtain a rough high-resolution image g(y0). The structure of the image super-resolution network g is as Figure 2As shown in the figure, the image super-resolution network includes multiple feature extraction modules and an upsampling module. Each feature extraction module includes six self-attention layers and a convolutional layer for extracting image features. The upsampling module includes a convolutional layer and a pixel shuffling layer for enhancing the image resolution. To calculate the corresponding noise intensity weight w u (y0), the high-resolution image g(y0) is used to calculate the difference with the low-resolution image y0 upsampled by bicubic interpolation to obtain the uncertainty prediction value. The uncertainty prediction value and is further transformed through a monotonically increasing function u to obtain the noise intensity weight w based on uncertainty u (y0) = u(ψ est (y0)), where u is expressed as:
[0038]
[0039] where ψ max and b u are hyperparameters. For the region where the uncertainty ψ > ψ max , the noise intensity weight is set to u(ψ) = 1. Finally, a noise intensity weight map with the same size as the input image and the weight values in the interval [b u , 1] is obtained.
[0040] After obtaining the noise intensity weight, an uncertainty diffusion process is carried out. For the anisotropic diffusion process guided by uncertainty, it is based on the framework of the general isotropic diffusion process and applies the noise intensity weight based on uncertainty to the noise. The specific process is as follows: Multiply the noise intensity weight by a Gaussian noise with a mean of 0 and a variance of κ 2 η T I to obtain anisotropic noise with different variances in different regions, and add it to the low-resolution image y0 to obtain the noisy image x at the starting time T of the diffusion process T :
[0041]
[0042] where T is a hyperparameter representing the number of steps of the diffusion model, η T , κ are hyperparameters controlling the noise addition and denoising rates of the diffusion process, and I is the identity matrix;
[0043] Starting from the starting time T, a reverse iteration is carried out. In each reverse iteration of the diffusion process, the noisy image x t first converts the image spatial dimension to the channel dimension through a pixel unshuffling layer to reduce the computational power requirement for the denoising process, and then extracts the noise-free image at time 0 through the denoising network f, where 1 ≤ t ≤ T:
[0044]
[0045] and obtain x through interpolation t-1 Probability distribution of the noisy image at time
[0046]
[0047] where η t 、η t-1 、α t and κ are hyperparameters that control the noise addition and denoising rates of the diffusion process, I is the identity matrix. Subsequently, the feature map obtained by the denoising network is upsampled to the original size through nearest neighbor interpolation and a convolutional layer to obtain the denoised image; the final denoised image x0 is obtained after T iterations.
[0048] The structure of the denoising network is as shown in Figure 3 . The denoising network recovers a clear image with rich details from the noisy image through deep feature extraction. The denoising network includes three layers of downsampling feature extraction modules, one middle layer feature extraction module, and three layers of upsampling feature extraction modules. There are residual connections between the feature extraction modules corresponding to different scales. Among them, the downsampling feature extraction modules obtain feature maps of different scales to extract depth image features in a larger range. Subsequently, the middle layer feature extraction module refines the features of the smallest scale, and the upsampling feature extraction modules combine the downsampled depth image features with feature maps of different scales to obtain a multi-scale fused feature map.
[0049] Train the diffusion process denoising module composed of the denoising network, nearest neighbor interpolation, and convolutional layers. The training strategy is as follows: Since real image super-resolution methods require considering both the fidelity and perceptual quality of the image simultaneously, to achieve multi-objective optimization, the combination of the pixel loss (L1 loss) and the feature loss (LPIPS loss) function is used as the loss function of the diffusion process denoising module:
[0050]
[0051] where the L1 loss is used to ensure the fidelity of the network output image, the LPIPS loss is used to ensure the perceptual quality of the network output image, and λ is a hyperparameter used to control the trade-off between fidelity and perceptual quality.
[0052] During the training process of the model, the Adam optimizer is used for optimization. A total of 200,000 iterations are trained. The learning rate linearly increases from 0 to 2*10 -4 in the initial 5,000 iterations, and then the cosine annealing method is used to gradually reduce the learning rate to 2*10 -5 until the network converges. The batch size is set to 32.
[0053] A real - image super - resolution method based on an uncertainty - guided anisotropic diffusion model proposed by the present invention regards different regions in the LR image as different time steps in an isotropic diffusion process, leading to an anisotropic diffusion process. In this process, flat regions are assigned a lower noise level (such as when t is close to the 0 - moment), while edge and texture regions are assigned a larger noise (such as when t is close to the T - moment). To achieve region - specific noise control, the idea of uncertainty is adopted, and a simple super - resolution network is used to estimate the uncertainty (variance) of different regions in the input LR image. This uncertainty reflects the difficulty of recovering HR details and is closely related to the distribution difference between LR and HR, indicating the required amount of noise, so that the diffusion process is more suitable for the super - resolution task to obtain better super - resolution quality.
[0054] Finally, it should be noted that: the above - mentioned embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real image super-resolution method based on an uncertainty diffusion model, characterized in that The method includes the following steps: Step S1, construct an end-to-end deep real image super-resolution model, the model includes an uncertainty prediction module and an uncertainty diffusion process, and the uncertainty diffusion process includes a pixel de-shuffle layer, a denoising network, nearest neighbor interpolation, and a convolutional layer; Step S2, construct a training dataset for the end-to-end deep real image super-resolution model, use the ImageNet open-source dataset as the training dataset, and perform supervised training on the denoising module of the diffusion process composed of the denoising network, nearest neighbor interpolation, and convolutional layer; Step S3, use the trained end-to-end deep real image super-resolution model for image super-resolution processing. The uncertainty prediction module predicts the uncertainty of each point in the image, and different intensities of noise are added according to the high or low uncertainty of each point. In the area with low uncertainty, that is, the smooth area in the image, a lower noise intensity, that is, a smaller noise weight, is adopted to retain more original image information. In the area with high uncertainty, that is, the edge and texture areas in the image, a higher noise intensity, that is, a larger noise weight, is adopted to enable the diffusion process to better utilize the information in the low-resolution image and improve the super-resolution effect; the uncertainty diffusion process obtains the denoised high-resolution image through pixel de-shuffle, denoising, nearest neighbor interpolation, and convolutional upsampling.
2. The method according to claim 1, wherein In step S3, the uncertainty prediction module first inputs the low-resolution image y0 into the image super-resolution network g to obtain a rough high-resolution image g(y0), and then calculates the difference between the high-resolution image g(y0) and the low-resolution image y0 upsampled by bicubic interpolation to obtain an uncertainty prediction value. The uncertainty prediction value is further transformed by a monotonically increasing function u to obtain a noise intensity weight w based on uncertainty u (y0) = u(ψ est (y0)), where u is expressed as: Among them, ψ max and b u are hyperparameters. For the region where the uncertainty ψ > ψ max , the noise intensity weight is set to u(ψ) = 1. Finally, a noise intensity weight map with the same size as the input image and the weight values in the interval [b u , 1] is obtained.
3. The method according to claim 2, wherein The image super-resolution network includes multiple feature extraction modules and an upsampling module. Each feature extraction module includes six self-attention layers and a convolutional layer for extracting image features, and the upsampling module includes a convolutional layer and a pixel shuffle layer for improving the image resolution.
4. The method according to claim 1, characterized in that In step S3, the uncertainty diffusion process is specifically as follows: Multiply the noise intensity weight by Gaussian noise with a mean of 0 and a variance of κ that is randomly sampled 2 η T I to obtain anisotropic noise with different variances in different regions, and add it to the low-resolution image y0 to obtain the noisy image x at the starting time T of the diffusion process T : Among them, T is a hyperparameter representing the number of steps of the diffusion model, and η T , κ are hyperparameters that control the noise addition and denoising rates of the diffusion process, and I is the identity matrix; Perform reverse iteration starting from the initial time T. During the reverse iteration of each diffusion process, the noisy image x t first converts the spatial dimension of the image to the channel dimension through the pixel unshuffle layer, and then extracts the denoised image at time 0 through the denoising network f, where 1 ≤ t ≤ T: and obtain x through interpolation t-1 Probability distribution of the noisy image at Among them, η t 、η t-1 、α t and κ are hyperparameters that control the noise addition and denoising rates of the diffusion process. I is the identity matrix. Subsequently, the feature map obtained by the denoising network is upsampled to the original size through nearest-neighbor interpolation and convolutional layers to obtain the denoised image. Finally, after T iterations, the final denoised image x0 is obtained.
5. The method according to claim 1, characterized in that The denoising network includes three downsampling feature extraction modules, one intermediate layer feature extraction module, and three upsampling feature extraction modules. There are residual connections between the feature extraction modules corresponding to different scales. Among them, the downsampling feature extraction modules obtain feature maps of different scales, and then the intermediate layer feature extraction module refines the features of the smallest scale, and the upsampling feature extraction modules combine the downsampled depth image features with the feature maps of different scales to obtain a multi-scale fused feature map.
6. The method according to claim 1, characterized in that, The training strategy in step S2 is: use the combination of pixel loss, that is, L1 loss, and feature loss, that is, LPIPS loss function, as the loss function of the diffusion process denoising module: Among them, the L1 loss is used to ensure the fidelity of the network output image, and the LPIPS loss is used to ensure the perceptual quality of the network output image. λ is a hyperparameter used to control the trade-off between fidelity and perceptual quality. During the training process of the model, the Adam optimizer is used for optimization.
Citation Information
Cited By
Defect detection method and system for gantry machine tool workbench casting part
CN120927812A