Image super-resolution reconstruction method based on generalized window integration strategy

By adopting a super-segment reconstruction network based on a generalized lightweight window integration strategy and a hybrid attention module in image super-resolution reconstruction, the problem of insufficient global feature extraction capability under window size limitation is solved, and efficient image reconstruction and optimized utilization of computing resources are achieved.

CN120219170APending Publication Date: 2025-06-27XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510297273.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the image super-resolution reconstruction, due to window size limitations, long-distance dependencies are difficult to capture, resulting in insufficient global feature extraction capabilities, low computing efficiency, and easy artifacts or blurs to appear near the edges of the image.

Method used

A super-segment reconstruction network based on a generalized lightweight window integration strategy and a hybrid attention module consists of a feature map subnet, a residual group subnet and a high-resolution reconstruction subnet cascade. The hybrid attention module uses the window attention integration module and the channel self-attention module, combined with the adaptive fusion module, to effectively capture long-distance dependencies and reduce computational complexity.

Benefits of technology

It improves the global feature extraction capability of image super-resolution reconstruction, improves the computing efficiency, reduces edge artifacts and blurring, and significantly improves the quality of image reconstruction and the generalization capability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219170A_ABST
    Figure CN120219170A_ABST
Patent Text Reader

Abstract

The invention provides an image super-resolution reconstruction method based on a generalized window integration strategy. The generalized window integration strategy is realized by a mixed attention module in a residual group sub-network. The window attention integration module passes through a plurality of window integration modules which are connected in parallel and then is connected with the channel self-attention module in series, space and channel dimension information is extracted, and fusion is carried out through the self-adaptive fusion module. Therefore, the method not only can enhance the capability of capturing boundary pixels and information outside a window and relieve the problem of boundary artifacts, but also can reduce the calculation complexity, effectively improve the efficiency and performance of a network model, and balance the performance and calculation power consumption, thereby improving the feasibility and universality of a super-resolution task in practical application, and improving the robustness of the super-resolution task. And the super-resolution reconstruction effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and further relates to an image super-resolution reconstruction method based on a generalized window integration strategy in the field of image super-resolution reconstruction. The present invention can be used for image super-resolution reconstruction in medical diagnosis, remote sensing technology, video surveillance, and smartphones. Background Art

[0002] In computer vision tasks, the receptive field refers to the size of the "visible" area of a neural network on an image. Image super-resolution algorithms based on convolutional neural networks can increase the receptive field by stacking convolutional layers and pooling layers. A larger receptive field enables the algorithm to more accurately understand the image context and extract richer features, which can be used to reconstruct the details and textures of high-resolution images, improving the algorithm performance and reconstruction quality. In addition, increasing the receptive field can also help the network better handle image scale changes and deformations, enhancing the algorithm's robustness and generalization ability. To better utilize the receptive field, researchers proposed the SwinIR super-resolution method for image restoration. SwinIR divides the image into non-overlapping windows and restricts the self-attention mechanism to calculate within the windows. This method has made breakthrough progress in low-level vision fields such as image super-resolution, image denoising, and image deblurring, and is superior to the current state-of-the-art methods in terms of performance. However, it uses sliding windows for communication, which limits the receptive field of the model. In addition, to reduce the computational burden, SwinIR adopts the strategy of performing self-attention calculations in small windows, effectively integrating features by capturing local features within the windows and using window translation technology to integrate information from different windows. However, this method shows certain limitations in constructing long-range feature dependencies, resulting in the pixel range utilized by SwinIR being much smaller than that of conventional methods.

[0003] Fuzhou University proposed a method for a super-resolution reconstruction model based on a residual hybrid attention network in its patent document "Image Super-Resolution Reconstruction Model and Method Based on Residual Hybrid Attention Network" (Patent Application No.: 202210940743.1, Publication No.: CN 115222601 A). This method constructs a network that can extract local and global features from the input low-resolution image, obtain rich high- and low-frequency information, and achieve high-precision super-resolution image reconstruction. The specific implementation steps of this method are to input the low-resolution image into the shallow feature extraction module to extract shallow features; then input the shallow features into the deep feature extraction module composed of multiple cascaded residual separation hybrid attention groups and global residual connections to extract deep features, and finally use sub-pixel convolution in the reconstruction module to upsample the deep features to obtain a high-resolution image. Among them, the residual separation hybrid attention module uses channel separation technology to split the feature map and send it to two branches for processing, and then fuses the local features extracted by the residual triple attention module and the global features extracted by the efficient Swin Transformer module to obtain rich high- and low-frequency information. The disadvantages of this method are that since this method is based on the Transformer method, it requires a large amount of computing resources and a long inference time; and the pixel range utilized by the existing super-resolution method based on sliding windows is much smaller than that of the conventional method, weakening the long-distance modeling function of the Transformer, resulting in a large limitation in the extraction of effective information and not particularly good image generation effects. Summary of the Invention

[0004] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies, and propose an image super-resolution reconstruction method based on a generalized window integration strategy to solve the problems in the existing technologies that the window size limits the model's ability to extract global features in capturing long-distance dependency relationships, low computational efficiency, and the appearance of artifacts or blurring near the image edges.

[0005] To achieve the above object, the technical idea of the present invention is to propose a super-resolution reconstruction network based on a generalized lightweight window integration strategy and a hybrid attention module. The network is composed of a feature mapping sub-network, a residual group sub-network, and a high-resolution reconstruction sub-network connected in cascade. The residual group sub-network is formed by cascading multiple hybrid attention modules with the same structure and then connecting them with one convolutional layer. The generalized lightweight window integration strategy is implemented by the hybrid attention module in the residual group sub-network. The window attention integration module in the hybrid attention module extracts the main features of each sub-window in an innovative way and uses these features as labels to establish the attention relationship between windows. In the specific implementation process, the pooling layer extracts features for each window, then calculates the similarity through normalization and activation functions (such as GELU), generates a similarity matrix, and fuses information by selecting windows with high similarity, thereby realizing the extraction of global features. This method can effectively capture long-range dependencies and solve the problem that local windows in traditional sliding window methods are difficult to effectively obtain global information. The hybrid attention module extracts information from the spatial and channel dimensions respectively by connecting the window integration module and the channel self-attention module in parallel, and then fuses them through the adaptive fusion module, reducing the computational complexity. The combination of the window integration module and the adaptive fusion module enables the model to effectively fuse information from different dimensions while reducing the computational overhead, thus balancing performance and computing power consumption. While ensuring high performance, it reduces the consumption of computing resources through parallel processing and effectively improves the computing efficiency. This technical solution can not only enhance the ability to capture boundary pixels and information outside the window, alleviate the boundary artifact problem, but also significantly improve the utilization efficiency of computing resources, thereby improving the feasibility and generality of super-resolution tasks in practical applications.

[0006] To achieve the above object, the technical solution adopted by the invention includes the following:

[0007] Step 1, construct a hybrid attention module composed of a window attention integration module, a channel self-attention module, and an adaptive fusion module;

[0008] Step 2, construct a super-resolution reconstruction network composed of a feature mapping sub-network, a residual group sub-network, and a high-resolution reconstruction sub-network connected in cascade;

[0009] Step 3, generate a training set;

[0010] Step 4, perform rough structure and detailed fine-tuning training on the network model respectively;

[0011] Step 5, input the low-resolution natural image to be reconstructed into the trained super-resolution reconstruction network, and output the high-resolution image of the image.

[0012] The structure of the described window attention integration module is successively composed of a normalization layer, a pooling layer, and an activation layer in cascade.

[0013] Furthermore, the structure of the described channel self-attention module is successively: a normalization layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and an activation layer; among them, the normalization layer is cascaded with the three convolutional layers respectively, the outputs of the first convolutional layer and the second convolutional layer are cascaded and then cascaded with the activation layer, and its output is cascaded with the third convolutional layer, and finally a residual connection is made with the output of the window attention integration module.

[0014] Furthermore, the structure of the described adaptive fusion module is successively: a pooling layer, a first fully connected layer, a second fully connected layer, a first activation layer, a second activation layer, and its connection method is that the pooling layer is cascaded with the first fully connected layer, connected to the second fully connected layer through the first activation layer, and then through the second activation layer to obtain the output; its input is the outputs of the window attention integration module and the channel self-attention module.

[0015] Furthermore, the parameter settings in the described hybrid attention module are as follows: the sizes of the convolutional layers are all 1*1, the numbers of convolutional kernels are all set to 12, the convolutional strides are all set to 1, the activation layer in the window attention integration module is implemented by the ReLU activation function, and the activation layer in the channel self-attention module is implemented by the Softmax activation function; the first activation layer in the adaptive fusion module is implemented by the ReLU activation function, and the second activation layer is implemented by the Softmax activation function.

[0016] Furthermore, the described feature mapping sub-network includes a feature extraction module and a feature conversion module; among them:

[0017] The feature extraction module is composed of a first convolutional layer, an activation layer, a normalization layer, and a second convolutional layer in cascade.

[0018] The structure of the feature conversion module is successively: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and the connection method of the four convolutional layers is cascade, and the output of the first convolutional layer is connected to the outputs of the remaining convolutional layers as a residual connection;

[0019] The parameters in the feature extraction module and the feature conversion module are set as follows:

[0020] The sizes of the first convolutional layers in both modules are 1*1, the sizes of the remaining convolutional layers are 3*3, the numbers of convolutional kernels are all set to 12, the convolutional strides are all set to 1, and the activation layers are all implemented by the ReLU function.

[0021] Further, the residual group sub-network is composed of t hybrid attention modules with the same structure cascaded and then cascaded with 1 convolutional layer, where the value of t is limited by the fitting degree of the network model. The more the number of t, the better the effect of the network model, but problems such as model overfitting need to be considered and selected according to the actual situation.

[0022] Further, the high-resolution reconstruction sub-network is composed of a first convolutional layer, an upsampling layer, and a final convolutional layer; among them, the structure of the upsampling layer is in turn a nearest neighbor interpolation magnification layer, a first convolutional layer, an activation layer, and a second convolutional layer; the parameter settings in the high-resolution reconstruction sub-network are as follows: the size of the convolutional layers is all 1*1, the size of the final convolutional layer is 3*3, the number of all convolutional kernels is set to 12, the convolutional stride is set to 1, and all activation layers are implemented by the ReLU function.

[0023] Further, the steps for generating the training set are as follows:

[0024] First step, select at least 3400 high-resolution natural images with 2K resolution to form a data set;

[0025] Second step, preprocess each high-resolution natural image to obtain the corresponding low-resolution image;

[0026] Third step, crop each high-resolution image into image patches of 64*64 size, and crop each low-resolution image into low-resolution image patches of 32*32 size;

[0027] Fourth step, form a training set from all the cropped high-resolution image patches and low-resolution image patches.

[0028] Further, the steps for respectively performing rough structure and detailed fine-tuning training on the network model are as follows:

[0029] First step, set the training parameters: The training stage is divided into two stages: rough structure and detailed fine-tuning. The batch size in the first stage is 8, and the initial learning rate is 2*10 -4 ; the initial learning rate in the second stage is 10 -4 , and the Adam optimizer is used in both stages, and the momentum parameters β1 = 0.9 and β2 = 0.999 are set;

[0030] Second step, input all low-resolution images into the super-resolution reconstruction network batch by batch, and use the gradient descent method to update the network parameters until the L1 loss function converges, and obtain a preliminarily trained super-resolution reconstruction network;

[0031] Third step, input all low-resolution images into the super-resolution reconstruction network batch by batch, and use the gradient descent method to update the network parameters until the MSE loss function converges, and obtain a trained super-resolution reconstruction network.

[0032] The described L1 loss function is as follows:

[0033]

[0034] Among them, L1 represents the absolute error between each pixel when initially training the super-resolution reconstruction network, N represents the total number of training samples, y i represents the output of the super-resolution reconstruction network, and GT i represents the real image corresponding to the output image.

[0035] The described MSE loss function is as follows:

[0036]

[0037] Among them, L2 represents the square of the difference between each pixel when secondarily training the super-resolution reconstruction network.

[0038] Compared with the prior art, the present invention has the following advantages:

[0039] First, through the window attention integration module, the present invention uses the extracted window features as labels to establish the attention relationship between windows, calculates the similarity through normalization and activation functions (such as GELU), generates a similarity matrix, and then fuses the window information with high similarity to obtain the extraction of global features. Accordingly, it can effectively capture long-range dependence relationships, and overcomes the problem of low efficiency in aggregating global features caused by the difficulty of local windows in the traditional sliding window method to effectively obtain global information, enabling the present invention to extract and aggregate features within the entire global range and improving the efficiency of aggregating global features.

[0040] Second, the channel self-attention module of the present invention extracts spatial and channel dimension information respectively, and then fuses them through the adaptive fusion module, reducing the computational complexity. It overcomes the problem of high computational cost existing in the existing Transformer-based models, enabling the present invention to effectively improve the efficiency and performance of the network model, balance performance and computing power consumption, and be more feasible and versatile in practical applications.

[0041] Thirdly, in the hybrid attention module of the present invention, after multiple window integration modules connected in parallel are integrated and then cascaded with the channel self-attention module, the combination of the window integration module and the adaptive fusion module enables the model to effectively fuse information from different dimensions while reducing the computational overhead, thereby balancing performance and computing power consumption. While ensuring high performance, the consumption of computing resources is reduced through parallel processing, overcoming the artifacts and blurring deficiencies existing in existing models, and effectively improving the computing efficiency. This enables the present invention not only to enhance the ability to capture boundary pixels and information outside the window, alleviate the boundary artifact problem, but also to significantly improve the utilization efficiency of computing resources, thereby improving the feasibility and universality of the super-resolution task in practical applications and enhancing the effect of super-resolution reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flowchart of the implementation of the embodiments of the present invention;

[0043] Figure 2 is a schematic structural diagram of the super-resolution reconstruction network of the present invention;

[0044] Figure 3 is a schematic structural diagram of the hybrid attention module of the present invention;

[0045] Figure 4 is a schematic structural diagram of the window attention integration module of the present invention;

[0046] Figure 5 is a schematic structural diagram of the channel self-attention module of the present invention;

[0047] Figure 6 is a simulation diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The present invention will be further described below with reference to the accompanying drawings.

[0049] Refer to Figure 1 for further description of the specific steps for implementing the present invention.

[0050] The super-resolution reconstruction network of the present invention based on the generalized lightweight window integration strategy and the hybrid attention module is composed of a feature mapping sub-network, a residual group sub-network, and a high-resolution reconstruction sub-network.

[0051] To more vividly describe the structural characteristics of the super-resolution reconstruction network, refer to Figure 2 for further description of the feature mapping sub-network constructed in step 1, the residual group sub-network constructed in step 2, and the high-resolution reconstruction sub-network constructed in step 3 of the present invention.

[0052] Step 1. Construct a feature mapping sub-network.

[0053] The feature mapping sub-network includes a feature extraction module and a feature transformation module, and its structure is as shown in Figure 2 the first part in

[0054] The structure of the feature extraction module is successively: the first convolutional layer, the activation layer, the normalization layer, and the second convolutional layer. The connection method is cascading.

[0055] The structure of the feature transformation module is successively: the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer. The connection method of the four convolutional layers is cascading, and the output of the first convolutional layer is connected with the outputs of the remaining convolutional layers by residual connection.

[0056] Set the parameters in the feature extraction module and the feature transformation module as follows:

[0057] The size of the first convolutional layer of both modules is 1*1, the size of the remaining convolutional layers is 3*3, the number of convolutional kernels is set to 12, the convolutional stride is set to 1, and the activation layer is implemented by the ReLU function.

[0058] Step 2. Construct the residual group sub-network.

[0059] In the embodiment of the present invention, the residual group sub-network is composed of four hybrid attention modules with the same structure (t = 4) and one convolutional layer. The residual group sub-network is the second part of the overall network, and its structure is as shown in Figure 2 the second part in

[0060] The first hybrid attention module, the second hybrid attention module, the third hybrid attention module, and the fourth hybrid attention module are cascaded and then cascaded with the convolutional layer, and its output is connected with the output of the feature mapping module by residual connection.

[0061] Each hybrid attention module is as shown in Figure 3 shown.

[0062] The structure of the hybrid attention module is successively: the window attention integration module, the channel self-attention module, and the adaptive fusion module. The window attention module and the channel self-attention module are connected in parallel, and the information extracted from the spatial and channel dimensions is adaptively weighted and summed respectively, and then the channel self-attention weight is extracted by the adaptive fusion module to obtain the features that adaptively fuse the two branches from the spatial and channel dimensions and aggregate the long-distance information between different windows. Better obtain the information outside the window from the pixels in the boundary region.

[0063] The window attention integration module is as shown in Figure 4As shown in the figure. Its structure is as follows: the normalization layer, the pooling layer, and the activation layer are cascaded. The pooling layer in the window attention module divides the feature map into multiple windows, extracts the features of each window through the pooling operation, generates window tokens, calculates the similarity of the feature map between different windows, generates the similarity matrix of each feature map, and based on the similarity matrix, selects the windows with high similarity for information fusion to generate new window features.

[0064] The channel self-attention module is as Figure 5 shown in the figure. Its structure is as follows: the normalization layer, the first convolutional layer, the second convolutional layer, the third convolutional layer, and the activation layer. The connection method is that the normalization layer is cascaded with the three convolutional layers respectively, the outputs of the first convolutional layer and the second convolutional layer are cascaded and then cascaded with the activation layer, and its output is cascaded with the third convolutional layer, and finally a residual connection is made with the output of the window attention integration module.

[0065] The structure of the adaptive fusion module is as follows: the pooling layer, the first fully connected layer, the second fully connected layer, the first activation layer, and the second activation layer. The connection method is that the pooling layer is cascaded with the first fully connected layer, connected to the second fully connected layer through the first activation layer, and then passed through the second activation layer to obtain the output. Its input is the outputs of the window attention integration module and the channel self-attention module.

[0066] The output of the hybrid attention module is the residual connection between the output of the feature mapping module and the output of the adaptive fusion module.

[0067] The parameter settings in the hybrid attention module are as follows:

[0068] The size of the convolutional layer is 1*1, the number of convolutional kernels is set to 12, the convolutional stride is set to 1. The activation layer in the window attention integration module is implemented by the ReLU activation function, and the activation layer in the channel self-attention module is implemented by the Softmax activation function. The first activation layer in the adaptive fusion module is implemented by the ReLU activation function, and the second activation layer is implemented by the Softmax activation function.

[0069] Step 3. Construct a high-resolution reconstruction sub-network.

[0070] The high-resolution reconstruction sub-network consists of the first convolutional layer, the upsampling layer, and the final convolutional layer. Its structure is as Figure 2 shown in the third part in the figure.

[0071] The structure of the upsampling layer is as follows: the nearest neighbor interpolation magnification layer, the first convolutional layer, the activation layer, and the second convolutional layer.

[0072] The parameter settings in the high-resolution reconstruction sub-network are as follows:

[0073] The sizes of all convolutional layers are 1*1, the size of the final convolutional layer is 3*3, the number of all convolutional kernels is set to 12, the convolutional stride is set to 1, and all activation layers are implemented by the ReLU function.

[0074] Step 4. Generate the training set.

[0075] The image dataset of the embodiment of the present invention selects at least 800 high-resolution natural images from the DIV2K dataset and 2,650 high-resolution natural images from the Flickr2K dataset, and combines the above images into a high-resolution image dataset. Selecting from the two datasets respectively is to enrich the types of images, improve the generalization ability of the model in different real scenarios, and help the deep learning model better learn image details and texture information.

[0076] For each high-resolution natural image, preprocessing is performed, that is, 1 / 2 times downsampling, bicubic interpolation are performed in sequence, and then processing is performed using an anisotropic Gaussian blur kernel to obtain the corresponding low-resolution image.

[0077] The expression of the anisotropic Gaussian blur kernel is as follows:

[0078]

[0079] Among them, g represents the Gaussian blur kernel, π represents the pi, σ represents the variance of the Gaussian distribution, exp represents the exponential function with the natural constant e as the base, m represents the position coordinates of the blur kernel, m = [x, y] T , [x, y] represents the position of a point relative to the center coordinates of the blur kernel, the superscript T represents the transpose operation, and R represents the covariance matrix, which are obtained by the following formulas respectively:

[0080]

[0081] Among them, θ represents the rotation angle of the main axis of the Gaussian distribution in the Gaussian blur kernel, specifically the angle used to rotate the Gaussian distribution (i.e., the blur kernel), satisfying the uniform probability distribution of 0 to π, and λ1, λ2 are random eigenvalues, satisfying the uniform probability distribution of 0.2 to 0.5 and 0.2 to λ1 respectively.

[0082] Each high-resolution image is cropped into image patches of size 64*64, and each low-resolution image is cropped into low-resolution image patches of size 32*32.

[0083] All the cropped high-resolution image patches and low-resolution image patches are combined to form the training set.

[0084] Step 5. Initially train the super-resolution reconstruction network.

[0085] All low-resolution images are input into the super-resolution reconstruction network in batches. The L1 loss function is used to calculate the absolute error between the output image and the real high-resolution image, and the loss value is obtained. This loss value is input into the Adam optimizer to calculate the gradients of each module in the network, and the network parameters are updated according to the gradients.

[0086] The initial training uses a batch size of 8 and an initial learning rate of 2*10 -4 , and the learning rate is halved after every 10,000 iterations until 100,000 iterations are completed. The rough structure training of the network model is completed to obtain a preliminarily trained super-resolution reconstruction network.

[0087] The L1 loss function is defined as follows:

[0088]

[0089] where L1 represents the absolute error between each pixel during the preliminary training of the super-resolution reconstruction network, N represents the total number of training samples, y i represents the output of the super-resolution reconstruction network, and GT i represents the real image corresponding to the output image.

[0090] Step 6. Secondarily train the super-resolution reconstruction network.

[0091] All low-resolution images are input into the preliminarily trained super-resolution reconstruction network to output high-resolution image blocks; the mean square error between this output and the corresponding real high-resolution image blocks is calculated to obtain the loss value, and the mean square error is represented by L2. The training in this stage uses an initial learning rate of 1*10 -4 , iterates 200,000 times, and the learning rate is decayed after every 10,000 iterations. The detailed fine-tuning of the model is completed to obtain a trained super-resolution reconstruction network.

[0092] The L2 loss function is defined as follows:

[0093]

[0094] where L2 represents the square of the difference between each pixel during the secondary training of the super-resolution reconstruction network.

[0095] Step 7. Perform super-resolution reconstruction on the image.

[0096] The low-resolution natural image to be reconstructed is input into the trained super-resolution reconstruction network to output the high-resolution image of this image.

[0097] The following further illustrates the effect of the present invention in combination with simulation experiments:

[0098] 1. Simulation experiment conditions.

[0099] The simulation hardware platform of the present invention is: an Intel Core i9-13900K CPU with a main frequency of 3.00 GHz, 64 GB of memory, and an NVIDIA RTX 4090 graphics card.

[0100] The simulation software platform of the present invention is: the Ubuntu22.04 operating system, Pytorch 2.0.1 and Python 3.11.

[0101] The test samples used in the simulation experiment of the present invention are selected from the DIV2K (800 2K images) and Flickr2K (2650 images) to construct a training set, with a total of 3450 high-resolution images. For each image, bicubic downsampling is used to generate the corresponding low-resolution image, and then image patches with a size of 64×64 are extracted through random rotation, flipping, and cropping operations to provide diverse samples for the model. At the same time, Set5, Set14, BSD100, and Urban100 are selected as the test data sets to objectively evaluate the reconstruction effect.

[0102] 2. Simulation content and result analysis.

[0103] In the simulation experiment of the present invention, the present invention and four existing technologies (bicubic method, RCAN method, SwinIR method, DAT method) are respectively used to perform super-resolution reconstruction on the input DIV2K images and Flickr2K images by magnifying 2 times, magnifying 3 times, and magnifying 4 times to obtain super-resolution result images.

[0104] To verify the simulation experiment effect of the present invention, all the low-resolution images in the DIV2K validation set and the Flickr2K validation set are input into the trained generation network for super-resolution reconstruction to obtain the super-resolution result images of all the images in the DIV2K validation set and the Flickr2K validation set.

[0105] The four methods in the existing technologies used in the simulation experiment of the present invention refer to:

[0106] The bicubic method of the existing technology is a traditional image super-resolution method, specifically referring to bicubic interpolation.

[0107] The existing RCAN method refers to a deep learning model for image super-resolution publicly disclosed by Yulun Zhang et al. in their paper "Image Super-Resolution Using Very Deep Residual Channel Attention Networks" (2018, ECCV - European Conference on Computer Vision). Its core lies in combining the residual structure and the channel attention mechanism, deepening the network structure through the characteristics of the residual network, and using the channel attention mechanism to adaptively adjust the weights of feature channels, thereby improving the effect of image super-resolution.

[0108] The existing SwinIR method refers to an image restoration method based on the Swin Transformer publicly disclosed by Jingyun Liang et al. in their paper "SwinIR: Image Restoration Using Swin Transformer" (2021, IEEE International Conference on Computer Vision), aiming to improve the performance of the image restoration task and reduce the number of parameters through the Transformer architecture.

[0109] The existing DAT method refers to a Transformer model for image super-resolution publicly disclosed by Zheng Chen et al. in their paper "Dual Aggregation Transformer for Image Super-Resolution" (2023, IEEE International Conference on Computer Vision). This method aggregates the spatial and channel features of the image through a dual aggregation (between blocks and within blocks) method, thereby enhancing the representation ability of the model.

[0110] The following combines Figure 6 to further describe the effects of the present invention.

[0111] To prove the effects of the present invention, take the "033" image selected from the input Urban100 images as an example, and combine Figure 6 to further describe the simulation effects of the present invention.

[0112] Figure 6 (a) is the result image after the low-resolution "033" image in the Urban100 validation set is super-resolution reconstructed and magnified 2 times using the existing bicubic method. Figure 6(b) is the result image after the low-resolution "0033" image in the Urban100 validation set is super-resolution reconstructed and magnified by 2 times using the existing RCAN method. Figure 6 (c) is the result image after the low-resolution "033" image in the Urban100 validation set is super-resolution reconstructed and magnified by 2 times using the existing SwinIR method. Figure 6 (d) is the result image after the low-resolution "033" image in the Urban100 validation set is super-resolution reconstructed and magnified by 2 times using the existing DAT method. Figure 6 (e) is the result image after the low-resolution "033" image in the Urban100 validation set is super-resolution reconstructed and magnified by 2 times using the super-resolution technology based on the window integration strategy of the present invention.

[0113] Comparison Figure 6 It can be seen that there is a large amount of repeated texture between the windows of the building, and the method of the present invention can obtain the information of windows with similar texture within a certain distance. Therefore, compared with the methods of the other four existing technologies, the method of the present invention restores an image with a correct structure and clear texture.

[0114] Next, through two evaluation indicators, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM), the reconstruction effect of the method in the simulation experiment of the present invention is objectively evaluated.

[0115] The calculation formulas of the two evaluation indicators are as follows:

[0116]

[0117]

[0118] Among them, n represents the total number of pixels in the test result image, MSE represents the mean square error between the test result image and the true high-resolution image, m and n respectively represent the test result image and the true high-resolution image, μ m represents the mean value of the pixel values of m, μ n represents the mean value of the pixel values of n, represents the variance of m, represents the variance of n, σ mn represents the covariance of m and n. The higher the PNSR and SSIM indicators, the higher the structural fidelity of the reconstructed image.

[0119] The comparison results of the average values of the objective evaluation indicators for reconstructing the Urban100 validation set images by the present invention and the above four methods in the prior art are shown in Table 1.

[0120] Table 1 Objective index table for reconstructing images by the present invention and the comparison methods

[0121] Algorithm PSNR↑ SSIM↑ bicubic 25.63 0.8330 RCAN 33.34 0.9384 SwinIR 33.81 0.9427 DAT 34.37 0.9443 The present invention 34.37 0.9458

[0122] As can be seen from Table 1, the average values of PSNR and SSIM after testing the Urban100 validation set samples in the content reconstruction stage of the present invention are higher than the four comparative methods in the above-mentioned prior art, indicating that the images reconstructed by the present invention have better fidelity.

Claims

1. An image super-resolution reconstruction method based on a generalized window integration strategy, characterized in that: Construct a hybrid attention module consisting of a window attention integration module, a channel self-attention module, and an adaptive fusion module; the specific steps of the reconstruction method include the following: Step 1, construct a hybrid attention module consisting of a window attention integration module, a channel self-attention module, and an adaptive fusion module; Step 2, construct a super-resolution reconstruction network consisting of a cascade of a feature mapping sub-network, a residual group sub-network, and a high-resolution reconstruction sub-network; Step 3, generated training set; Step 4, respectively perform structural rough adjustment and detail fine-tuning training on the network model; Step 5: Input the low-resolution natural image to be reconstructed into the trained super-resolution reconstruction network, and output a high-resolution image of the image.

2. The image super-resolution reconstruction method according to claim 1, characterized in that: The structure of the window attention integration module described in step 1 is: a normalization layer, a pooling layer, and an activation layer cascaded.

3. The image super-resolution reconstruction method according to claim 1, characterized in that: The structure of the channel self-attention module described in step 1 is: normalization layer, the first convolution layer, the second convolution layer, the third convolution layer, and the activation layer; the normalization layer is cascaded with the three convolution layers respectively, the output of the first convolution layer is cascaded with the second convolution layer, and then cascaded with the activation layer, and its output is cascaded with the third convolution layer, and finally residually connected with the output of the window attention integration module.

4. The image super-resolution reconstruction method according to claim 1, characterized in that: The structure of the adaptive fusion module described in step 1 is: pooling layer, first fully connected layer, second fully connected layer, first activation layer, second activation layer, and its connection method is that the pooling layer is cascaded with the first fully connected layer, connected with the second fully connected layer through the first activation layer, and then through the second activation layer to obtain the output; its input is the output of the window attention integration module and the channel self-attention module.

5. The image super-resolution reconstruction method according to claims 1-4, characterized in that: The parameters in the hybrid attention module are set as follows: the size of the convolution layer is 1*1, the number of convolution kernels is set to 12, the convolution step size is set to 1, the activation layer in the window attention integration module is implemented by the ReLU activation function, and the activation layer in the channel self-attention module is implemented by the Softmax activation function; The first activation layer in the adaptive fusion module is implemented by the ReLU activation function, and the second activation layer is implemented by the Softmax activation function.

6. The image super-resolution reconstruction method according to claim 1, characterized in that: The feature mapping subnetwork described in step 2 includes a feature extraction module and a feature conversion module; wherein: The feature extraction module is composed of a first convolution layer, an activation layer, a normalization layer, and a second convolution layer in cascade; The structure of the feature conversion module is: the first convolution layer, the second convolution layer, the third convolution layer, and the fourth convolution layer. The four convolution layers are connected in cascade, and the output of the first convolution layer is residually connected with the outputs of the remaining convolution layers. Set the parameters in the feature extraction module and feature conversion module as follows: The size of the first convolutional layer of both modules is 1*1, and the size of the remaining convolutional layers is 3*3. The number of convolution kernels is set to 12, the convolution step size is set to 1, and the activation layers are implemented by ReLU function.

7. The image super-resolution reconstruction method according to claim 1, characterized in that: The residual group subnetwork described in step 2 is composed of t hybrid attention modules with the same structure cascaded and then cascaded with one convolutional layer, where the value of t is limited by the degree of fit of the network model.

8. The image super-resolution reconstruction method according to claim 1, characterized in that: The high-resolution reconstruction subnetwork described in step 2 consists of the first convolutional layer, the upsampling layer and the final convolutional layer; wherein the structure of the upsampling layer is the nearest neighbor interpolation amplification layer, the first convolutional layer, the activation layer, and the second convolutional layer in sequence; the parameters in the high-resolution reconstruction subnetwork are set as follows: the convolutional layer size is 1*1, the final convolutional layer size is 3*3, the number of all convolution kernels is set to 12, the convolution step size is set to 1, and all activation layers are implemented by ReLU function.

9. The image super-resolution reconstruction method according to claim 1, characterized in that: The steps for generating the training set described in step 3 are as follows: The first step is to select at least 3,400 high-resolution natural images with a resolution of 2K to form a dataset; The second step is to preprocess each high-resolution natural image to obtain the corresponding low-resolution image; The third step is to crop each high-resolution image into a 64*64 image block, and crop each low-resolution image into a 32*32 low-resolution image block; The fourth step is to combine all the cropped high-resolution image patches and low-resolution image patches into a training set.

10. The image super-resolution reconstruction method according to claim 1, characterized in that: The steps for performing coarse structural and detailed fine-tuning training on the network model described in step 4 are as follows: The first step is to set the training parameters: the training phase is divided into two stages: coarse structure adjustment and detail fine-tuning. The batch size of the first stage is 8, and the initial learning rate is 2*10 -4 ; The initial learning rate of the second stage is 10 -4 , Adam optimizer is used in both stages, and momentum parameters β1 = 0.9, β2 = 0.999 are set; In the second step, all low-resolution images are input into the super-resolution reconstruction network in batches, and the network parameters are updated using the gradient descent method until the L1 loss function converges, thus obtaining a preliminarily trained super-resolution reconstruction network. In the third step, all low-resolution images are input into the super-resolution reconstruction network in batches, and the network parameters are updated using the gradient descent method until the MSE loss function converges to obtain a trained super-resolution reconstruction network.

11. The image super-resolution reconstruction method according to claim 10, characterized in that: The L1 loss function is as follows: Where L1 represents the absolute error between each pixel when initially training the super-resolution reconstruction network, N represents the total number of training samples, and y i Represents the output of the super-resolution reconstruction network, GT i Indicates the real image corresponding to the output image.

12. The image super-resolution reconstruction method according to claim 11, characterized in that: The MSE loss function is as follows: Among them, L2 represents the square of the difference between each pixel when retraining the super-resolution reconstruction network.

Citation Information

Patent Citations

  • Image super-resolution reconstruction model and method based on residual mixed attention network

    CN115222601A