Real-world Image Super-resolution Reconstruction Method Based on Superpixel Attention Mechanism
By introducing a superpixel attention mechanism and a GAN network architecture SPGAN model, the problem of high-frequency detail recovery of real-world low-resolution images is solved, and a higher-quality image super-segment reconstruction effect is achieved.
Patent Information
- Application Number
- CN202411836612.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-12-13
AI Technical Summary
The prior art is difficult to effectively restore high-frequency details of real-world low-resolution images, and supervised learning methods based on image pairs have poor generalization effect in the real world.
The image super-segment reconstruction method based on the superpixel attention mechanism is adopted, combined with the GAN network architecture, a two-stage gating degradation model and superpixel module are introduced, and the SPGAN model is trained through L1 loss, perceived loss, full variation regularization and GAN loss to improve the attention ability of pixels at the edge of the image.
It significantly improves the real-world image super-scoring effect, reduces blur artifacts, improves the fineness and reconstruction quality of the image, and significantly improves the PSNR and SSIM indicators.
Smart Images

Figure CN119784591B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a real-world image super-resolution reconstruction method based on a superpixel attention mechanism. Background Art
[0002] Real-world image super-resolution (RSISR) is an important research topic in computer vision. In real life, there are a large number of devices for collecting real-world images. However, due to factors such as device performance limitations, the real-world images collected are often low-resolution images (LR) with insufficient quality, suffering from problems such as detail loss, inaccurate color, and compression artifacts, which seriously affect the effects of computer vision tasks such as medical imaging, satellite remote sensing, video surveillance, and traffic detection. Therefore, it is of great practical significance to study how to restore high-resolution images (HR) from low-resolution images (LR) obtained from the real world.
[0003] Real-world image super-resolution (RSISR) is different from ordinary single-image super-resolution (SISR). The degradation process of images in the real world is very complex, including factors such as blur, noise, and distortion. Moreover, the degraded low-resolution images may not retain detailed features such as textures of real-world pictures. At the same time, it is very difficult to collect well-aligned low-resolution and high-resolution image pairs (LR-HR) in real-world image super-resolution tasks, which makes it difficult for some supervised learning methods based on image pairs to achieve good super-resolution effects and the generalization effect is also relatively poor.
[0004] Recently, methods based on degradation modeling have made great progress in some real-world super-resolution tasks. An artificial-designed degradation model is used to degrade the HR image to generate the corresponding LR image, and then the HR-LR image pair is used for model training. The core of this method lies in simulating the degradation process of images from high resolution to low resolution in the real world, including the influence of factors such as blur, compression, and noise. In this way, a large number of synthetic data sets can be generated for training and optimizing the super-resolution model. These synthetic data sets can not only provide sufficient training samples but also simulate different degradation scenarios to enhance the generalization ability of the model.
[0005] In practical applications, the degradation modeling method can significantly improve the performance of super-resolution reconstruction. For example, by mixing degradation operations and setting different degradation parameters, LR images with a wide range of degradation effects can be generated. The advantage of this method is that it can easily obtain a very large number of paired degraded LR images without the need to collect them laboriously or endure the misalignment problem of unpaired training data. In addition, by changing the degradation parameters, various degradation models can be easily obtained, thereby improving the adaptability and robustness of the model in different scenarios.
[0006] With the development of deep learning technology, real-world super-resolution technology based on deep learning has also been actively explored and developed. Currently, various real-world image super-resolution methods based on deep learning have been proposed, and good reconstruction results have been achieved on public datasets. These methods can effectively learn the high-frequency information in images, including structures, textures, etc., by learning the mapping relationship of images, so as to restore rich details and structures in low-resolution images. Therefore, super-resolution technology based on degradation modeling has broad prospects in practical applications, especially in the fields of medical diagnosis, remote sensing images, computer vision research, etc. Summary of the Invention
[0007] To address the above technical problems, the present invention provides a real-world image super-resolution reconstruction method based on a superpixel attention mechanism, which can generate a low-resolution degraded image closer to the characteristics of real-world degradation, thereby obtaining more excellent super-resolution performance in real-world scenarios. At the same time, a real-world image super-resolution method based on superpixels is proposed, which integrates the network architecture of GAN and effectively strengthens the model's attention to image edge pixels using the superpixel module, further improving the fineness of the super-resolution result.
[0008] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0009] A real-world image super-resolution reconstruction method based on a superpixel attention mechanism includes the following steps:
[0010] S1. Taking GAN as the basic structure, integrating the superpixel attention mechanism into the generator network to construct a real-world image super-resolution model SPGAN;
[0011] S2. Introducing a two-stage gated degradation model to obtain a real-world degraded image by adjusting the blur kernel, size, noise type, and performing JPEG compression;
[0012] S3. Designing a superpixel module in combination with superpixel attention to group similar pixels perceptually;
[0013] S4. Training SPGAN by combining L1 loss, perceptual loss, total variation regularization, and GAN loss, and adopting exponential moving average EMA to obtain more stable training and better performance.
[0014] In the S1, the main framework of the real-world image super-resolution model SPGAN is a GAN structure. The generator is used for the super-resolution process of the degraded low-resolution image LR, and the discriminator is responsible for judging the relative authenticity of the super-resolution SR image and the high-resolution HR image.
[0015] The method for super-resolution reconstruction using the real-world image super-resolution model SPGAN is as follows: The low-resolution image after two-stage gated degradation is input into the generator. The image captures features through N superpixel blocks SPI and improves the reconstruction quality through PixelShuffle upsampling operation. At the same time, the overall residual structure is combined to obtain the final SR image. Finally, this SR image and its corresponding HR image are jointly input into the discriminator to determine the authenticity of the image. Through the adversarial training of the two, the super-resolution reconstruction effect of the image is improved.
[0016] The method for introducing the two-stage gated degradation model in S2 is as follows: The two-stage gated degradation model includes two identical stages, and each stage includes four parts: blur kernel, resizing, adding noise, and JPEG compression. The mathematical model can be expressed as:
[0017]
[0018] Where represents the second-order gated degradation model, where D 1 and D 2 respectively represent the degradation operations of the first stage and the second stage. D k represents blur kernel degradation, D s represents downsampling degradation, D n represents noise degradation, D j represents JPEG compression degradation; σ represents the gating probability, which is used to control whether the corresponding degradation operation takes effect. Through the above four parts, the gating operation mechanism is incorporated to achieve flexible regulation of the degradation operation intensity, so as to more accurately simulate the actual degradation phenomenon.
[0019] The method for designing the superpixel module by combining superpixel attention in S3 is as follows: It includes superpixel segmentation, superpixel cross-attention, superpixel internal attention, and local attention module. The overall uses a residual structure to simplify training. The superpixel internal attention uses two cascaded linear layers to replace the two matrix multiplication operations in the self-attention operation, making the model have lower complexity and better performance. Connecting multiple superpixel attention modules in parallel can better strengthen the model's attention to local and global features.
[0020] The method for training SPGAN by combining L1 loss, perceptual loss, total variation regularization, and GAN loss in S4 is as follows:
[0021] First, the L1 loss is used to calculate the average error between the super-resolved image and the real image. The loss formula is defined as follows:
[0022]
[0023] Where n is the total number of samples, f(x) is the super-resolved image, and yi For real-world images, the L1 loss function iteratively optimizes the model by calculating the absolute value of the difference between f(x) and y i to optimize the model iteratively;
[0024] Train a model for PSNR, and name the trained model SPNet;
[0025] Then, we use the trained PSNR-oriented model as the initialization of the generator, and train SPGAN by combining L1 loss, perceptual loss, total variation regularization, and GAN loss.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0027] The present invention proposes a real-world image super-resolution reconstruction method SPGAN based on a superpixel attention mechanism. A two-stage gated degradation model is proposed, which makes the low-resolution image after model degradation closer to the real-world degraded image, reducing to a certain extent the unsatisfactory super-resolution effect caused by the large gap between the image after model degradation and the real-world degraded image, and further improving the super-resolution effect of real-world images. At the same time, combined with the GAN network structure, the superpixel module is applied to the generator network, enabling the model to pay more attention to the edge pixels of the image, and solving to a certain extent the problems such as image blur artifacts caused by the GAN network, and improving the fineness of the super-resolution image. The pre-trained model SPNet network in the present invention is quantitatively compared on the Set14 dataset. When the noise level is 15, a PSNR value of 23.74 and an SSIM value of 0.6165 are obtained; when the noise level is 25, a PSNR value of 23.03 and an SSIM value of 0.578 are obtained; on the BSD100 dataset for quantitative comparison, when the noise level is 15, a PSNR value of 23.83 and an SSIM value of 0.6176 are obtained; when the noise level is 25, a PSNR value of 23.31 and an SSIM value of 0.5864 are obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, other implementation drawings can be obtained according to the provided drawings without creative efforts.
[0029] The structures, proportions, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the implementation conditions of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the efficacy that the present invention can produce and the purpose that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0030] Figure 1 It is the overall architecture diagram of the SPGAN of the present invention;
[0031] Figure 2 It is the structure diagram of the superpixel block in the SPGAN model of the present invention;
[0032] Figure 3 It is the comparison diagram of the experimental results of the SPGANet of the present invention and excellent real-world super-resolution methods for synthetic images;
[0033] Figure 4 It is the comparison diagram of the experimental results of the SPGAN of the present invention and excellent real-world super-resolution methods for real-world images. Detailed implementation manners
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than a limitation on the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0035] The following will further describe in detail the specific implementation manners of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but not to limit the scope of the present invention.
[0036] The present invention is implemented under the pytorch deep learning framework. The present invention provides a real-world image super-resolution reconstruction method based on a superpixel attention mechanism, which specifically includes the following steps:
[0037] 1. Data preparation
[0038] The data samples of the present invention are the DF2K image super-resolution dataset:
[0039] The DF2K dataset is obtained by merging the DIV2K dataset and the Flickr2K dataset. The DIV2K dataset contains 1000 low-resolution images of various different degradation types, including 800 training images, 100 validation images, and 100 test images; the Flickr2K dataset contains 2650 images with a resolution of 2K, and the image content includes various types such as people, animals, and landscapes, covering a rich variety of natural scenes and daily objects.
[0040] During the training process, the images are randomly cropped to a size of 256×256.
[0041] 2. Model Construction
[0042] The main framework of the constructed SPGAN model is a GAN structure, and the specific network structure is as Figure 1 shown. The SPGAN model contains two major modules: a generator and a discriminator. The generator first uses a 3×3 convolutional layer to extract shallow features of the low-resolution image I LR , and then uses n superpixel modules to form an encoder. As Figure 2 shown, each superpixel module consists of four parts: a superpixel segmentation module, a superpixel cross-attention module, a superpixel intra-attention module, and a local attention module. In terms of the details of constructing the superpixel block, we first pass the pixel information to the superpixel layer through a linear layer to generate the query vector Q s , key vector K I , and value vector V I required for the self-attention mechanism. These vectors are obtained by multiplying with the corresponding weight matrices and . Then, we calculate the similarity between the query vector Q s and the key vector K I , and use this similarity as the attention weight to multiply with the value vector V I to obtain the updated superpixel features. This process can be expressed as:
[0043]
[0044] where is the scaling factor used to prevent gradient disappearance. The updated superpixel features are then mapped back to the pixel level through the cross-attention mechanism, which solves the problem of long-range information transmission between superpixels and ensures the effective flow of information between pixels. After passing through n superpixel blocks, the extracted features pass through a 3×3 convolutional layer and an upsampling Pixel Shuffle layer, and finally the result is combined with the low-resolution image I LRPerform residual fusion; the discriminator is designed with a U-Net structure with stronger discrimination ability, which can provide more detailed discrimination feedback. At the same time, in order to increase the training stability, the spectral normalization regularization method is used to stabilize the training process.
[0045] 3. Model Training
[0046] The training process adopts two stages.
[0047] First, we use the L1 loss to calculate the average error between the super-resolution image and the real image. The loss formula is defined as follows:
[0048]
[0049] where n is the total number of samples, f(x) is the super-resolved image, and y i is the real-world image. The L1 loss function iteratively optimizes the model by calculating the absolute value of the difference between f(x) and y i .
[0050] Train a model for PSNR, and the trained model is named SPNet.
[0051] Then, we use the trained PSNR-oriented model as the initialization of the generator and train SPGAN in combination with the L1 loss, perceptual loss, total variation regularization, and GAN loss.
[0052] 4. Test Results
[0053] The method for super-resolution reconstruction of training from low-resolution images to obtain super-resolution reconstruction results is as follows: For the real-world image dataset, bicubic interpolation is used to adjust the resolution of the test images in the dataset to 256×256 as the HR image; the designed two-stage gated degradation model is used to degrade the HR image to an image with a resolution of 16×16 as the LR image. First, use the trained SPNet model to reconstruct the LR image to obtain the SR image, and use SSIM and PNSR as performance evaluation indicators to evaluate the reconstruction effect; then, use the SPGAN model to super-resolution reconstruct the LR image to obtain the SR image, and use NIQE as the performance evaluation indicator to evaluate the reconstruction effect.
[0054] 5. Model Evaluation
[0055] Use the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) evaluation indicators calculated from the super-resolution results and real images to evaluate the performance of the SPNet model.
[0056] Table 1 Comparison results of different models on the Set14 and BSD100 datasets
[0057]
[0058] The experiment first compared SPNet with four advanced real-world image super-resolution algorithms, namely BSRNet, DASR, Real-SwinIR-L, and Real-ESRNet. The quantitative comparison of the super-resolution image effect evaluation was carried out on the Set14 dataset and the BSD100 dataset. The best indicators in the table are in bold font.
[0059] It can be seen from Table 1 that under the same experimental environment, compared with the comparison algorithms, SPNet has excellent PSMR and SSIM results in the Set14 dataset and the BSD100 dataset. In the Set14 dataset, under two noise levels, the PSNR has a significant improvement compared with the BSRNet and DASR methods, and also has a slight improvement compared with Real-SwinIR-L. The performance is similar to that of the Real-ESRNet model; the SSIM is also very close to that of the Real-ESRNet. In the BSD100 dataset, it also performs well under two noise levels. The PSNR still has a large improvement compared with the BSRNet and DASR methods, and also has a slight improvement compared with Real-SwinIR-L; compared with the Real-ESRNet, both the PSNR and SSIM are very close. Through Figure 3 the comparison results, it can be found that SPNet has a better reconstruction effect on severely degraded images compared with the current excellent real-world super-resolution models.
[0060] The NIQE (Natural Image Quality Evaluator) score was calculated using the super-resolution results and the real images to evaluate the performance of the SPGAN model.
[0061] Table 2 Comparison results of different models on multiple real-world image datasets
[0062]
[0063] Table 2 compared the SPGAN model with four advanced real-world image super-resolution methods, namely Bicubic, ESRGAN, CDC, and BSRGAN. The quantitative comparison of the super-resolution image effect evaluation was carried out on the RealSR-Canon dataset, the OST300 dataset, and the ADE20Kval dataset. The best indicators in the table are in bold font.
[0064] It can be found from the comparison results in Table 2 that SPGAN outperforms most real-world super-resolution methods on all three datasets. On the RealSR-Canon dataset and the OST300 dataset, the NIQE scores of the SPGAN model are 5.6214 and 3.2413 respectively, and the reconstruction effect is significantly better than other real-world super-resolution methods. On the ADE20Kval dataset, the NIQE score of the SPGAN model is also slightly better than that of other excellent super-resolution models. This shows that SPGAN can better handle various degradation situations in the real world and reconstruct LR images into high-quality SR images to the greatest extent. Through Figure 4 comparison, it can be found that SPGAN not only has a great improvement in enhancing image resolution, but also can well retain detailed information such as edge contours and textures, which indicates that SPGAN has superior performance in real-world image super-resolution tasks.
[0065] The above only elaborates on the preferred embodiments of the present invention in detail. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and all such changes should be included within the protection scope of the present invention.
Claims
1. A real-world image super-resolution reconstruction method based on a superpixel attention mechanism, characterized in that It includes the following steps: S1. Based on GAN, incorporate the superpixel attention mechanism into the generator network to construct the real-world image super-resolution model SPGAN. The method for super-resolution reconstruction by the real-world image super-resolution model SPGAN is as follows: The low-resolution image after two-stage gated degradation is input into the generator. The image captures features through N superpixel blocks SPI and improves the reconstruction quality through the Pixel Shuffle upsampling operation. At the same time, the overall residual structure is combined to obtain the final SR image. Finally, this SR image and its corresponding HR image are input into the discriminator to judge the authenticity of the image. Through the adversarial training of the two, the super-resolution reconstruction effect of the image is improved; S2. Introduce a two-stage gated degradation model to obtain the real-world degraded image by adjusting the blur kernel, size, noise type, and performing JPEG compression; S3. Design a superpixel module in combination with superpixel attention to group similar pixels perceptually. The method for designing the superpixel module in combination with superpixel attention in S3 is as follows: It includes superpixel segmentation, superpixel cross-attention, superpixel internal attention, and a local attention module. The overall uses a residual structure to simplify training. The superpixel internal attention replaces the two matrix multiplication operations in the self-attention operation with two cascaded linear layers, making the model have lower complexity and better performance. Connecting multiple superpixel attention modules in parallel can better enhance the model's attention to local and global features; S4. Train SPGAN by combining L1 loss, perceptual loss, total variation regularization, and GAN loss, and use exponential moving average EMA to obtain more stable training and better performance.
2. The real-world image super-resolution reconstruction method based on the superpixel attention mechanism according to claim 1, wherein The main framework of the real-world image super-resolution model SPGAN in S1 is a GAN structure. The generator is used for the super-resolution process of the degraded low-resolution image LR, and the discriminator is responsible for judging the relative authenticity of the super-resolution SR image and the high-resolution HR image.
3. The real-world image super-resolution reconstruction method based on the superpixel attention mechanism according to claim 1, characterized in that The method of introducing the two-stage gated degradation model in S2 is as follows: The two-stage gated degradation model consists of two identical stages, and each stage includes four parts: a blur kernel, resizing, adding noise, and JPEG compression; the mathematical model is expressed as: where represents the second-order gated degradation model, and the and represent the degradation operations of the first stage and the second stage respectively, represents the blur kernel degradation, represents the downsampling degradation, represents the noise degradation, represents the JPEG compression degradation; represents the gating probability, which is used to control whether the corresponding degradation operation takes effect; through the above four parts, the gating operation mechanism is incorporated to achieve flexible control of the degradation operation intensity, so as to more accurately simulate the actual degradation phenomenon.
4. The real-world image super-resolution reconstruction method based on the superpixel attention mechanism according to claim 1, wherein The method for training SPGAN by combining L1 loss, perceptual loss, total variation regularization, and GAN loss in S4 is as follows: First, use L1 loss to calculate the average error between the super-resolved image and the real image. The loss formula is defined as follows: Among them n is the total number of samples, f(x) is the image after super-resolution, y i is the real-world image. The L1 loss function iteratively optimizes the model by calculating f(x) and y i the absolute value of the difference between them; Train a model for PSNR, and the trained model is named SPNet; Then, use the trained model for PSNR as the initialization of the generator, and train SPGAN by combining L1 loss, perceptual loss, total variation regularization, and GAN loss.
Citation Information
Patent Citations
Unsupervised semantic segmentation algorithm based on adversarial network and self-attention mechanism
CN115346045A
Remote sensing image blind super-resolution model training method and system
CN118172248A