Residual-guided progressive diffusion image super-resolution method, system, electronic device or computer readable storage medium
Patent Information
- Application Number
- CN202610919160.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-25
AI Technical Summary
[0011]本发明的目的是提供一种残差引导的渐进式扩散图像超分辨率方法、系统、电子设备或计算机可读存储介质,用于解决现有固定分辨率扩散超分辨率方法中退化路径失配、噪声强度缺乏空间自适应、纹理区域细节不足、平滑区域随机伪影、单一网络任务耦合、采样步数较多以及固定高分辨率推理开销较大的问题
[0031](1)由于本发明中渐进式扩散框架的图像退化过程考虑到了像素空间分辨率的退化,这与实际高分辨率图像到低分辨率图像的物理退化规律具有一致性,能够降低固定尺度采样带来的结构错位和颜色不一致。
Smart Images

Figure CN122820441A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing, computer vision, and deep learning, and in particular to a residual-guided progressive diffusion image super-resolution method, system, electronic device, or computer-readable storage medium. Background Technology
[0002] Single-Image Super-Resolution (SISR) aims to recover a corresponding high-resolution image from a low-resolution image. It is widely used in scenarios such as mobile image enhancement, security monitoring, medical image-assisted analysis, remote sensing image processing, digital document restoration, and video content production. This problem is essentially an ill-conditioned inverse problem: the same low-resolution input may correspond to multiple high-resolution solutions, thus requiring a balance between visual fidelity and subjectively perceived quality.
[0003] Traditional methods typically rely on manual priors such as adaptive sparse domain selection and regularization, nonlocal centralized sparse representation, and weighted kernel norm minimization. With the advent of deep learning methods, SRCNN, SRGAN, and SwinIR have been used to learn end-to-end mappings from low to high resolution. In recent years, Denoising Diffusion Probabilistic Models (DDPM) have been introduced into super-resolution tasks due to their strong generative priors. Typical approaches include conditional denoising diffusion methods such as SR3 and SRDiff, as well as methods like ResShift and SinSR to improve sampling efficiency.
[0004] The existing technologies closest to this invention mainly fall into two categories: one is the fixed-resolution conditional diffusion super-resolution method based on DDPM, such as SR3 and SRDiff, which performs iterative denoising at the target high-resolution scale, with the low-resolution image only serving as conditional input; the other is residual transfer diffusion methods such as ResShift, which reduce the number of sampling steps by designing Markov residual transfer paths between the high-resolution image and the upsampled low-resolution image. Specifically, the typical process of a fixed-resolution conditional diffusion super-resolution scheme (such as SR3 and SRDiff) is as follows: first, the low-resolution image is mapped to the target resolution through bicubic interpolation or an encoder; then, a diffusion sampling chain from Gaussian noise to the high-resolution image is constructed in the target resolution space; at each time step, the neural network uses the current noise sample, the time step encoding, and the low-resolution condition as input to predict the noise or predict the previous state; after dozens or even hundreds of iterations, a high-resolution output is obtained. ResShift-like schemes further assume that there is a residual between the high-resolution image and the upsampled low-resolution image, and design a residual transfer process at a fixed target scale to complete the reconstruction with fewer sampling steps. However, the state transition still mainly occurs within the same spatial scale, and the stepwise spatial evolution from low resolution to high resolution is not explicitly embedded in the diffusion time step.
[0005] However, the aforementioned existing technologies still have the following shortcomings:
[0006] (1) Existing fixed-resolution or fixed-target-scale diffusion super-resolution methods such as SR3, SRDiff, and ResShift treat super-resolution mainly as a conditional denoising or residual transfer task, and do not incorporate the progressive spatial degradation from high resolution to low resolution into the forward diffusion chain. This results in a lack of correspondence between forward degradation and backsampling at the resolution level, which can easily lead to structural misalignment and color inconsistency.
[0007] (2) Existing conditional diffusion super-resolution methods such as SR3 and SRDiff use the same or approximately the same noise intensity in all regions without distinguishing between textured regions and smooth regions. If the noise is too strong, artifacts and color shifts are likely to occur in smooth regions. If the noise is too weak, the edges and textures are difficult to refine sufficiently.
[0008] (3) Fixed-scale diffusion sampling such as DDPM, SR3, and SRDiff usually rely solely on changes in noise intensity to complete forward degradation, requiring a long diffusion chain to make the image close to pure noise. Although methods such as ResShift and SinSR attempt to reduce the number of sampling steps, the reverse inference process still does not fully utilize the small computational scale of the low-resolution stage, making it difficult to balance high-quality reconstruction with the requirements of real-time or lightweight deployment.
[0009] (4) Although DDPM and ResShift provide noise scheduling or residual migration ideas respectively, the existing schemes lack a noise scheduling mechanism that matches the resolution level. If the noise changes too fast or too slow, it will cause insufficient contribution of some time steps, thus affecting sampling efficiency and generation quality.
[0010] (5) Although the diffusion super-resolution method based on latent space can reduce the computational load by compressing the image to a low-dimensional latent space, it requires an additional encoder-decoder structure, which may introduce compression loss. Moreover, the diffusion process in the latent space lacks the interpretability of the pixel domain. Its residual is represented as the difference of the encoded feature vector, which cannot directly correspond to the high-frequency structural information of the pixel domain, and it is difficult to adaptively adjust the noise intensity according to the local content of the image. Summary of the Invention
[0011] The purpose of this invention is to provide a residual-guided progressive diffusion image super-resolution method, system, electronic device, or computer-readable storage medium to solve the problems of degradation path mismatch, lack of spatial adaptation of noise intensity, insufficient texture details, random artifacts in smooth regions, single network task coupling, large number of sampling steps, and large inference overhead at fixed high resolution in existing fixed resolution diffusion super-resolution methods.
[0012] To achieve the above objectives, the technical solution adopted by this invention is: a residual-guided progressive diffusion image super-resolution method, comprising the following steps:
[0013] S1. Construct a residual-guided progressive diffusion image super-resolution model, which includes a low-resolution image feature extraction network based on residual nested residual dense blocks RRDB, a denoising U-Net network, and an upsampling decoder based on implicit neural representation INR.
[0014] S2. Obtain the high-resolution training image x0 and generate the corresponding low-resolution image y0 for training;
[0015] S3. Construct a forward diffusion process in pixel space that simultaneously incorporates resolution degradation and noise perturbation: Downsample the high-resolution image step-by-step according to a preset time step and resolution step size to obtain the noise-free intermediate image I corresponding to time step t. t Simultaneously, Gaussian noise is added according to the designed noise scheduling to obtain the diffusion intermediate state x. t ;
[0016] S4. Perform conditional encoding on the low-resolution image y0 to obtain low-resolution image features; simultaneously encode the time step t to obtain the temporal conditional embedding.
[0017] S5. Pre-trained denoising U-Net network: Freeze the parameters of the low-resolution image feature extraction network based on residual nested residual dense block RRDB, and perform independent pre-training of the denoising U-Net network in pixel space.
[0018] S6. Training the upsampling network: Freeze all parameters of the pre-trained denoising U-Net network, and use the pre-features of the output layer of the denoising U-Net network as implicit neural representations, inputting them into the upsampling decoder for training;
[0019] S7. Perform reverse progressive sampling: During inference, the low-resolution image y0 to be processed is input into the residual-guided progressive diffusion image super-resolution model, and the initial diffusion state x is obtained by superimposing it with Gaussian noise according to a preset noise strategy. T Starting from t=T, denoising and upsampling are performed step by step until t=0 to obtain a super-resolution image of the target resolution.
[0020] As a further improvement of the present invention, spatial resolution evolution and noise diffusion evolution are jointly embedded in the single-image super-resolution diffusion process, so that each diffusion time step corresponds to a controlled resolution transition, and the diffusion sampling chain is shortened by forward resolution degradation and reverse progressive resolution enhancement.
[0021] As a further improvement of the present invention, the Gaussian noise guided by residual modulation is spatially adaptively modulated using adjacent resolution residual maps, and the modulation operator is: ,in, Indicates standard Gaussian noise. This is a cross-resolution residual map between noise-free intermediate images of adjacent resolutions. and These are the scaling parameter and the offset parameter, respectively. This indicates element-wise multiplication; the modulation operator is used to enhance the noise intensity in high-frequency textured regions and suppress the noise intensity in smooth regions.
[0022] As a further improvement of the present invention, the serial recovery network includes a denoising network and a residual upsampling decoder; the denoising network is used to recover a structurally consistent noise-free intermediate image from the current diffusion state, and the residual upsampling decoder is used to predict cross-resolution residuals based on the intermediate features or output features of the denoising network, and add the cross-resolution residuals to the interpolation upsampling result to compensate for high-frequency details and complete the upsampling; wherein, the denoising network may be a denoising U-Net network, and the residual upsampling decoder may be an upsampling decoder based on the implicit neural representation INR.
[0023] As a further improvement of the present invention, the noise scheduling adopts a power-law form, assuming that the noise sequence satisfies Where p is the power exponent. Initial noise, The noise is the terminal noise; the noise intensity transfer rate is controlled by adjusting the parameter p, so that the noise evolution rate matches the resolution degradation or recovery rate.
[0024] As a further improvement of the present invention, in the reverse sampling process, network inference is performed in the early stage of sampling at a low resolution. As the time step decreases, the sample space size is gradually expanded according to the preset resolution step size until the target resolution is reached. By introducing resolution degradation into the forward diffusion process and performing progressive resolution increase in the reverse process, the diffusion sampling chain is shortened and the inference computation in the early stage of sampling is reduced.
[0025] This invention also provides a residual-guided progressive diffusion image super-resolution system, including an image input interface, a model parameter storage module, an inference scheduling module, a processor, a memory, and an image output interface; the model parameter storage module is used to store model parameters for implementing the method described above; the inference scheduling module is used to perform residual-guided noise modulation, denoising, and residual upsampling according to the diffusion time step; after the processor executes the program in the memory, it outputs a super-resolution image of the target resolution through the image output interface.
[0026] The present invention also provides an electronic device or a computer-readable storage medium, the electronic device including a processor and a memory, the memory storing a program that, when executed by the processor, implements the residual-guided progressive diffusion image super-resolution method as described above; the computer-readable storage medium storing program instructions that, when executed by the processor, implement the method as described above.
[0027] This invention represents single-image super-resolution reconstruction as a progressive diffusion trajectory across multiple resolution levels: the forward process simultaneously performs resolution degradation and noise degradation, while the reverse process simultaneously performs denoising and resolution upsampling in the opposite direction; each diffusion time step corresponds to a controlled resolution transition, and the noise is spatially adaptively modulated based on the cross-resolution residual.
[0028] Furthermore, this invention uses the progressive downsampling and noise addition of high-resolution images as the forward diffusion process, and the progressive denoising and upsampling of low-resolution states as the corresponding inverse sampling process, thus unifying the noise evolution and resolution evolution within the same diffusion chain. Simultaneously, this invention utilizes the cross-resolution residual between noise-free intermediate images of adjacent resolutions to characterize local high-frequency structures and texture complexity, and accordingly performs spatial adaptive modulation on the diffused noise, enabling edges and textured regions to retain stronger random refinement capabilities and reducing noise disturbances in smooth regions. This method addresses the problems of degradation path mismatch, lack of spatial adaptation in noise intensity, insufficient detail in textured regions, random artifacts in smooth regions, single-network task coupling, large number of sampling steps, and high inference overhead at fixed high resolution in existing fixed-resolution diffusion super-resolution methods.
[0029] This invention explicitly incorporates the spatial resolution decrease / increase process into the diffusion time step in pixel space, ensuring that the diffusion chain includes both noise evolution and resolution evolution. Spatial degradation assists the forward state in more quickly approaching the terminal noise distribution. The cross-resolution residual map between noise-free intermediate images at adjacent resolution levels in pixel space (i.e., the difference between noise-free intermediate images at adjacent resolutions, characterizing local high-frequency structure and texture complexity) serves as the basis for noise modulation, allowing textured and smooth regions to use different local noise intensities. A dual-network structure combining a denoising U-Net and a residual upsampling decoder in series clearly defines the division of labor between structural denoising, upsampling, and high-frequency detail compensation. Power-law noise scheduling adapted to progressive diffusion avoids premature image degradation into pure noise or inefficiency caused by slow early diffusion. Forward resolution degradation shortens the diffusion chain, and progressive inference from low to high resolution in reverse reduces early computation, achieving high-quality super-resolution reconstruction with fewer sampling steps; 4x super-resolution inference can be completed in just 8 sampling steps.
[0030] The beneficial effects of this invention are:
[0031] (1) Since the image degradation process of the progressive diffusion framework in this invention takes into account the degradation of pixel spatial resolution, it is consistent with the physical degradation law of actual high-resolution images to low-resolution images, which can reduce the structural misalignment and color inconsistency caused by fixed-scale sampling.
[0032] (2) Since the noise intensity is adaptively modulated by the cross-resolution residual, it can enhance the generation of details in high-frequency areas such as text, edges, and wall textures, while suppressing random artifacts in smooth areas such as sky and solid color background.
[0033] (3) Due to the use of a serial dual-network structure, the denoising network focuses on structural recovery, and the upsampling decoder adopts a residual design and focuses on high-frequency detail compensation and upsampling, which can reduce the learning burden of a single network and improve reconstruction stability.
[0034] (4) Since the forward diffusion chain synchronously includes resolution degradation and noise perturbation, and the reverse sampling chain denoises and improves resolution step by step in the corresponding direction, the image can be restored from an approximately noisy state to a high-resolution result in a shorter time step; at the same time, the network runs at a lower resolution in the early stage of reverse sampling, which can significantly reduce pixel-level computation and memory usage, and improve inference efficiency. Attached Figure Description
[0035] Figure 1 This is a diagram illustrating the overall architecture of the RGP-Diff progressive upsampling diffusion framework in an embodiment of the present invention.
[0036] Figure 2 This is a schematic diagram of the overall process and progressive resolution path of residual-guided progressive diffusion image super-resolution in an embodiment of the present invention;
[0037] Figure 3 This is a schematic diagram of the single-time-step serial dual-network recovery and residual-guided sampling update structure in an embodiment of the present invention;
[0038] Figure 4 This is a schematic diagram of the spatial adaptive noise modulation process guided by high-frequency residuals in an embodiment of the present invention. Detailed Implementation
[0039] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0040] Example 1
[0041] A residual-guided progressive diffusion image super-resolution method includes:
[0042] S1: Construct a residual-guided progressive diffusion image super-resolution model. The model consists of three sub-networks: (1) a low-resolution image feature extraction network based on RRDB (Residual-in-Residual Dense Block), used to encode low-resolution conditional images into deep feature representations; (2) a denoising U-Net network, which takes the current diffusion state, time step embedding, and low-resolution image features as input, and is responsible for recovering noise-free intermediate images from noisy images; (3) an upsampling decoder based on INR (Implicit Neural Representation), which takes the output layer pre-features of the denoising U-Net as input, achieves arbitrary-scale upsampling through multi-scale hierarchical position encoding, and compensates for high-frequency details through residual prediction. Among them, the denoising U-Net and INR upsampling decoder adopt a serial dual-network architecture to decouple the denoising task from the upsampling task.
[0043] S2: Obtain high-resolution training images And generate the corresponding low-resolution image. As one implementation method, the DIV2K and Flickr2K datasets can be used, with randomly cropped 256×256 image patches for 4x super-resolution training. Low-resolution images By analyzing high-resolution images The image is obtained by performing bicubic downsampling, with the downsampling factor determined based on the target task (e.g., 4x super-resolution means downsampling 256×256 to 64×64). During the training phase, high-resolution images... As the starting point of the forward diffusion process, low-resolution images As a conditional input to the denoising network, it provides structural prior information to the model.
[0044] S3: Construct a forward diffusion process in pixel space that simultaneously incorporates resolution degradation and noise perturbation. Set the total number of diffusion time steps to T (e.g., T=8), and the resolution step size between adjacent time steps to... High-resolution image After stepwise downsampling, the noise-free intermediate image at time step t The space dimensions are , where d is The side length, i.e. , This represents the downsampling operator. Noise scheduling uses a power-law form: Let the noise sequence... satisfy , where p is the power exponent (preferably p=2). =0.001 makes the initial noise approach zero. =0.999 makes the terminal noise approach 1. Diffusion intermediate state. Obtained through the following formula: , that is ,in Standard Gaussian noise, The residual-guided noise modulation operator is defined as follows: , This is a cross-resolution residual map (i.e., high-frequency residual) between noise-free intermediate images of adjacent resolutions. and These are the scaling parameter and the offset parameter (preferred). =20, =0.6), This indicates element-wise multiplication. This modulation mechanism imparts stronger noise to textured regions with larger residuals to enhance generated details, while imparting weaker noise to smooth regions with smaller residuals to suppress artifacts.
[0045] S4: For low-resolution images Conditional encoding is performed, and time step t is also encoded. Low-resolution conditional encoding employs a deep feature extraction network based on RRDB (Residual-in-Residual Dense Block): the low-resolution image... The input is an RRDB encoder, which extracts deep semantic features and structural information through multiple layers of residual dense blocks, and outputs a low-resolution image feature map. The feature map is injected into the feature layers of the denoising U-Net at various scales through cross-layer connections, providing structural prior guidance for the denoising process. The time step encoding uses sinusoidal position encoding or learned embedding to map the discrete time step t into a continuous vector representation, resulting in a time conditional embedding. This embedding is injected into the denoising network through adaptive normalization (AdaGN), enabling the network to perceive the current diffusion stage and adjust the denoising intensity accordingly.
[0046] S5: Pre-trained Denoising Network. In the first stage of training, the parameters of the RRDB low-resolution encoder are frozen, and the denoising U-Net is pre-trained independently in pixel space. Specifically, high-resolution images are randomly sampled from the training set. Randomly select time steps Calculate the diffusion state based on the forward diffusion formula. Noise-free intermediate image of the target .Will The temporal conditional embedding t and the low-resolution image features extracted by the RRDB encoder are jointly input into the denoising U-Net. The U-Net adopts an encoder-decoder structure, which includes multi-scale residual connections and attention modules, and is responsible for predicting the denoised image. Calculate the denoising loss. Using the AdamW optimizer (initial learning rate) The process continues until the loss converges. This stage enables U-Net to learn to recover a clear image structure at the current resolution from a noisy state, given low-resolution conditions and time steps.
[0047] S6: Training the Upsampling Network. In the second training stage, all parameters of the pre-trained denoising U-Net (including the RRDB encoder) are frozen. The pre-features of the denoising U-Net's output layer are used as implicit neural representations (INR) and input into the upsampling decoder. The upsampling decoder adopts the multi-scale hierarchical location decoding network structure in HIIF (Hierarchical encoding based Implicit Image Function). The INR features output by U-Net are interpolated in the feature space and then mapped to pixel values by a multilayer perceptron (MLP), thus achieving upsampling at arbitrary scales through continuous coordinate decoding. Simultaneously, the original output layer of U-Net is retained to generate a preliminary denoised image. ,Will After being upsampled to the target resolution via bilinear interpolation, the result is added to and fused with the output of the INR decoder via a skip connection to obtain the final output. ,in Indicates interpolation upsampling. This represents the cross-resolution residuals predicted by the INR decoder. The training objective is... Compared with a real noise-free intermediate image The L1 loss between them. Because the INR decoder operates in a residual manner (i.e., focuses on predicting high-frequency residuals). Instead of a complete image, the learning burden is significantly reduced, and skip connections ensure the integrity of the basic structure. Furthermore, the INR decoder can flexibly align the output to any higher-resolution, noise-free intermediate image. This enables super-resolution reconstruction at arbitrary magnification and non-integer scales, thereby enhancing model generalization.
[0048] S7: Perform inverse progressive sampling. During inference, the low-resolution image to be processed is... The input RRDB encoder extracts conditional features and performs forward relation processing. Will The initial diffusion state is obtained by superimposing the residual-modulated Gaussian noise. The reverse process is performed using the Cold Diffusion generalized framework, defining the objective of each time step as regressing from the current degenerate state. Directly recover the noise-free intermediate image of the next lower resolution. ,in This is a serial recovery network composed of a denoising U-Net and an INR decoder. The specific recursive process is as follows: Starting from t=T, (1) the current sample With time step t, low-resolution conditional features Input a denoising U-Net, output the denoising result (2) Input the INR features from the output layer into the upsampled decoder to predict the cross-resolution residuals. After skip connection and interpolation upsampling Fusion (3) Using the residual predicted by the upsampling decoder Constructing a noise modulation operator (4) Inject noise according to the forward relationship to obtain the next diffusion intermediate state: ,in The iterative process recursively executes from t=T to t=1, and directly outputs the result at t=0. High-resolution images are used as the target resolution. Since the upsampling network employs a residual design, its residual prediction results can be directly utilized during inference. Guided noise modulation eliminates the need for additional residual map calculations, achieving seamless coupling between residual information and noise modulation.
[0049] The above steps can be deployed as a software algorithm or implemented as an image processing system. Combined with... Figures 1 to 4 The data stream shown indicates that the system may include: an image input and preprocessing module, a progressive resolution path construction module, a forward degradation and noise scheduling module, a low-resolution conditional and time-step encoding module, a denoising structure restoration module, a residual upsampling decoding module, a jump-join fusion module, a residual-guided noise modulation module, and a reverse recursive sampling output module. The training phase may also include a training sample construction and L1 loss optimization module. Each module can be implemented by a processor executing program instructions stored in memory, or it can be merged or split into equivalent software functional units as needed for deployment.
[0050] Figure 2 In the diagram, 101 is the low-resolution image input y0, 102 is the near-pure noise image xT, 103 is the progressive backsampling module, and 104 is the high-resolution image output x0. The low-resolution image is also used for conditional feature extraction, and Gaussian noise εt is used to construct the diffusion state. The progressive path with the resolution gradually increasing from 64 to 256 is shown below. Figure 3 In the diagram, 201 represents the current sample xt, 202 represents the denoising network U-Net, 203 represents the INR decoder, 204 represents the noise-free intermediate image It-1, 205 represents the next time step sample xt-1, 206 represents the low-resolution conditional image y0, 207 represents the feature extraction module, 208 represents the time condition t, 209 represents the Gaussian noise εt, and 210 represents the residual-guided noise modulation module. The low-resolution conditional features and the time condition are jointly injected into the denoising network. The denoising network and the INR decoder sequentially complete the structure restoration, residual prediction, and upsampling, and combine the modulated noise with the noise-free intermediate image to obtain the next time step sample. Figure 4 In the diagram, 301 represents the high-frequency residual plot r. t 302 represents Gaussian noise ε t 303 represents the modulated noise M(ε) t ), 304 enhances noise intensity in high-frequency textured regions, and 305 suppresses noise intensity in smooth regions; s t m is the residual scaling factor. t As an offset parameter, this modulation process allows textures or edge regions with larger residuals to retain stronger random refinement capabilities, while reducing noise perturbations in smooth regions with smaller residuals.
[0051] On the DIV2K 4x super-resolution validation set, our method achieved a PSNR of 28.50, an SSIM of 0.7902, and an LPIPS of 0.1104, outperforming representative methods such as ESRGAN, SRFlow, SRDiff, and ResShift in fidelity (see Table 1). Stable reconstruction results were also achieved on multiple test sets, including Set5, Set14, BSD100, Urban100, and Manga109 (see Table 3).
[0052] In terms of parameter count and inference efficiency, this method has 80.00M parameters and an inference time of 0.2606 seconds per image, which is about twice as fast as ResShift and about 6.6 times faster than SRDiff (see Table 4). This efficiency advantage stems from the asymptotic resolution path, which allows the initial sampling to be performed at a low resolution. Although the latent space diffusion method can reduce the size of the diffused image through encoder compression, it requires the introduction of an additional variational autoencoder, resulting in a higher overall network parameter count and memory usage. This method does not require such a compression network and can directly reduce the computational load by leveraging the advantage of the asymptotic resolution path running at a low resolution in the early stages, with a smaller parameter count and lower memory usage.
[0053] Ablation experiments (see Table 2) show that removing residual guided noise modulation reduces PSNR by approximately 0.18 dB, indicating its crucial role in balancing structural fidelity and overall performance. Removing the serial dual network reduces PSNR by approximately 0.67 dB while significantly increasing LPIPS, demonstrating the critical importance of dual network specialization for texture detail synthesis. Removing the progressive resolution path reduces PSNR by approximately 0.49 dB and decreases inference efficiency, indicating that this path is fundamental for balancing reconstruction quality and efficiency. These three technical features make irreplaceable synergistic contributions to the overall effectiveness of this invention.
[0054] Table 1 presents a quantitative comparison of our method with representative super-resolution methods on the DIV2K 4x validation set, illustrating the overall effectiveness of our method compared to fixed-resolution diffusion methods and residual migration methods.
[0055]
[0056] Table 1
[0057] Table 2 shows the ablation experiment results of the key modules of this method, which are used to illustrate the contributions of residual-guided noise modulation, serial dual network and progressive resolution path to reconstruction quality.
[0058]
[0059] Table 2
[0060] Table 3 shows the results of this method at 4x single-image super-resolution on multiple public test sets, which are used to supplement the explanation of its cross-scene reconstruction effect.
[0061]
[0062] Table 3
[0063] Table 4 compares the parameter quantities and inference time per image of the proposed method with those of a representative diffusion super-resolution method under the same test settings, serving to further illustrate the inference efficiency of the present invention.
[0064]
[0065] Table 4
[0066] Table 5 provides explanations of English abbreviations and codes with special meanings.
[0067]
[0068] Table 5
[0069] Example 2
[0070] 4x Natural Image Super-Resolution Reconstruction: In this embodiment, the objective is to reconstruct a 64×64 low-resolution image into a 256×256 high-resolution image. The diffusion time step T=8 is set, with the resolution increasing progressively between adjacent steps, for example, 64, 88, 112, 136, 160, 184, 208, 232, 256. During training, starting from the high-resolution image x0, a forward degradation chain of progressive downsampling and noise addition is constructed in the opposite direction.
[0071] At each training time step t, a noise-free intermediate image I is first obtained by the downsampling operator P. t Then add the residual plot r t Modulated noise. The residual map can be obtained by prediction using an INR decoder and can be normalized or truncated to ensure that the noise modulation weights fall within a stable range.
[0072] Recovery model G θ Receive the current degradation state x t The time-step encoding t and the low-resolution conditional image y0 are used. The denoising U-Net extracts structural features and outputs a preliminary denoised representation. The upsampling decoder compensates for high-frequency details across resolutions using multi-scale positional encoding and predictive residuals. Finally, I is obtained through jump-join fusion. t-1 The estimated value. The training objective is to find the approximate value of this estimated value and the true I. t-1 L1 loss between.
[0073] During inference, given any low-resolution image y0 to be enhanced, start from the lowest resolution state x TInitially, denoising, upsampling, and residual-guided noise modulation are performed sequentially from t=8 to 1. A high-resolution image of the target is output at the last time step. This embodiment is implemented using PyTorch in the paper's experiments, trained on a single NVIDIA GeForce RTX 4090 GPU, with AdamW optimizer and an initial learning rate of 1×10⁻⁶. -4 .
[0074] Example 3
[0075] Parameter range and alternative implementations: The diffusion time step T can be set from 4 to 24 depending on the device's computing power and image quality requirements; since forward resolution degradation can help the image approach the terminal noise state more quickly, it can be set to 8 to balance inference efficiency and detail quality.
[0076] The noise scheduling exponent p can be set from 0.5 to 8, with p=2 being preferred. If p is too small, the image degradation speed will be too fast and not smooth enough; if p is too large, the degradation speed will be too slow and the efficiency will be low, both of which may affect the reconstruction quality. The residual modulation parameter st can be set from 10 to 40, and mt can be set from 0 to 1; st=20 and mt=0.6 are preferred.
[0077] The denoising network is not limited to U-Net, but can also be replaced by Transformer, convolutional residual network or hybrid structure; the low-resolution conditional encoder is not limited to RRDB, but can also be replaced by CNN, Swing Transformer or other feature extraction networks; the upsampling decoder can adopt implicit neural representation based on multi-scale position encoding or continuous coordinate decoding structure, or can be replaced by other upsampling structures that support reconstruction at different magnifications and non-integer scales.
[0078] In addition to 4x super-resolution of natural images, this method can also be applied to image restoration tasks such as document images, remote sensing images, medical images, video frame enhancement, and low-definition surveillance image enhancement.
[0079] Example 4
[0080] System and Device Implementation: This invention can also be implemented as an image super-resolution system. The system includes a processor, a memory, and a program stored in the memory. When the program is executed by the processor, it implements the above-described method steps. The system may also include an image input interface, a model parameter storage module, an inference scheduling module, and an image output interface. For mobile or edge deployments, computational load can be reduced through model pruning, quantization, or distillation; for cloud deployments, multiple images can be processed in batches and super-resolution results can be output.
[0081] In one specific deployment method, model training is completed on the server side, and the inference model is loaded in the form of a weight file. After the user inputs a low-resolution image, the system first standardizes the image size and pixel range, and can divide the input image into several low-resolution small image blocks; then, for each small image block, a progressive diffusion sampling process that expands from low resolution to high resolution is called to generate corresponding high-resolution small image blocks, so that the initial sampling is completed on a smaller size to reduce computational overhead; after each small image block completes super-resolution reconstruction, the system stitches them together into a complete high-resolution image according to the spatial position at the time of division, and finally performs post-processing operations such as pixel cropping, color space conversion, and format saving.
[0082] The embodiments described above are merely illustrative of specific implementations of the present invention, and while the descriptions are detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A residual-guided progressive diffusion image super-resolution method, characterized in that, Includes the following steps: S1. Construct a residual-guided progressive diffusion image super-resolution model, which includes a low-resolution image feature extraction network based on residual nested residual dense blocks RRDB, a denoising U-Net network, and an upsampling decoder based on implicit neural representation INR. S2. Obtain the high-resolution training image x0 and generate the corresponding low-resolution image y0 for training; S3. Construct a forward diffusion process in pixel space that simultaneously incorporates resolution degradation and noise perturbation: Downsample the high-resolution image step-by-step according to a preset time step and resolution step size to obtain the noise-free intermediate image I corresponding to time step t. t Simultaneously, Gaussian noise is added according to the designed noise scheduling to obtain the diffusion intermediate state x. t ; S4. Perform conditional encoding on the low-resolution image y0 to obtain low-resolution image features; simultaneously encode the time step t to obtain the temporal conditional embedding. S5. Pre-trained denoising U-Net network: Freeze the parameters of the low-resolution image feature extraction network based on residual nested residual dense block RRDB, and perform independent pre-training of the denoising U-Net network in pixel space. S6. Training the upsampling network: Freeze all parameters of the pre-trained denoising U-Net network, and use the pre-features of the output layer of the denoising U-Net network as implicit neural representations, inputting them into the upsampling decoder for training; S7. Perform reverse progressive sampling: During inference, the low-resolution image y0 to be processed is input into the residual-guided progressive diffusion image super-resolution model, and the initial diffusion state x is obtained by superimposing it with Gaussian noise according to a preset noise strategy. T Starting from t=T, denoising and upsampling are performed step by step until t=0 to obtain a super-resolution image of the target resolution.
2. The residual-guided progressive diffusion image super-resolution method according to claim 1, characterized in that, Spatial resolution evolution and noise diffusion evolution are jointly embedded in the single-image super-resolution diffusion process, so that each diffusion time step corresponds to a controlled resolution transition, and the diffusion sampling chain is shortened by forward resolution degradation and reverse progressive resolution enhancement.
3. The residual-guided progressive diffusion image super-resolution method according to claim 2, characterized in that, The Gaussian noise modulated by residual guidance is spatially adaptively modulated using adjacent resolution residual maps, and the modulation operator is: ,in, Indicates standard Gaussian noise. This is a cross-resolution residual map between noise-free intermediate images of adjacent resolutions. and These are the scaling parameter and the offset parameter, respectively. This indicates element-wise multiplication; the modulation operator is used to enhance the noise intensity in high-frequency textured regions and suppress the noise intensity in smooth regions.
4. The residual-guided progressive diffusion image super-resolution method according to claim 3, characterized in that, The serial recovery network includes a denoising network and a residual upsampling decoder. The denoising network is used to recover a structurally consistent, noise-free intermediate image from the current diffusion state. The residual upsampling decoder is used to predict cross-resolution residuals based on the intermediate or output features of the denoising network, and add the cross-resolution residuals to the interpolation upsampling result to compensate for high-frequency details and complete the upsampling. The denoising network can be a denoising U-Net network, and the residual upsampling decoder can be an upsampling decoder based on the implicit neural representation INR.
5. The residual-guided progressive diffusion image super-resolution method according to claim 4, characterized in that, The noise scheduling adopts a power-law form, assuming the noise sequence satisfies... Where p is the power exponent. Initial noise, The noise is the terminal noise; the noise intensity transfer rate is controlled by adjusting the parameter p, so that the noise evolution rate matches the resolution degradation or recovery rate.
6. The residual-guided progressive diffusion image super-resolution method according to claim 5, characterized in that, In the reverse sampling process, network inference is performed at a low resolution in the early stage of sampling. As the time step decreases, the sample space size is gradually expanded according to the preset resolution step size until the target resolution is reached. By introducing resolution degradation into the forward diffusion process and performing progressive resolution increase in the reverse process, the diffusion sampling chain is shortened and the inference computation in the early stage of sampling is reduced.
7. A residual-guided progressive diffusion image super-resolution system, characterized in that, It includes an image input interface, a model parameter storage module, an inference scheduling module, a processor, a memory, and an image output interface; the model parameter storage module is used to store model parameters for implementing the method according to any one of claims 1 to 6; the inference scheduling module is used to perform residual guided noise modulation, denoising, and residual upsampling according to the diffusion time step; After the processor executes the program in the memory, it outputs a super-resolution image of the target resolution through the image output interface.
8. An electronic device or computer-readable storage medium, characterized in that, The electronic device includes a processor and a memory, the memory storing a program that, when executed by the processor, implements the residual-guided progressive diffusion image super-resolution method according to any one of claims 1 to 6; the computer-readable storage medium stores program instructions that, when executed by the processor, implement the method according to any one of claims 1 to 6.