Image super-division reconstruction Transform method based on cross-window information aggregation and big kernel attention
By introducing cross-window information aggregation and large-core attention modules in the image super-resolution reconstruction Transformer model, the limited scope of information utilization and artifact problems are solved, and higher quality image reconstruction and excellent quantitative performance are achieved.
Patent Information
- Application Number
- CN202510172107.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has problems in image super-resolution reconstruction with limited information utilization range and artifacts appearing in reconstruction images, especially the Transformer model performs poorly in cross-window information interaction.
A Transformer method based on cross-window information aggregation and large-core attention is proposed. Through the parallel structure of mixed attention module and large-core attention module, the effective integration and utilization of cross-window information is realized.
This method can significantly improve the quality of image super-resolution reconstruction, reduce artifacts, and perform superiorly in quantitative performance such as PSNR and SSIM.
Smart Images

Figure CN120125438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image super-resolution reconstruction, and specifically to an image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention. Background Art
[0002] The progress of the times has made people's requirements for the quality of digital images increasingly stringent. However, the image quality of the pictures obtained in actual scenarios often leaves much to be desired and cannot contain all the information of the image scene. The reasons are as follows. On the one hand, it is limited by the current image acquisition method, and the digitized image information can only approximate the real scene as much as possible. On the other hand, it is restricted by the current image acquisition equipment, and the image acquisition process will be interfered by many external factors. Generally speaking, the obtained pictures will first undergo translation and rotation, during which various blurs will occur, such as motion blur, aperture focus blur, etc. In addition, the whole process is accompanied by the introduction of noise in the external environment, resulting in the loss of high-frequency details in the image. In this case, if storage compression is required, downsampling operations will be carried out, which will further reduce the information carried by the image. Therefore, people need a method to restore high-quality pictures from images with limited information, and the emergence of super-resolution reconstruction technology is to solve this problem.
[0003] Super-resolution (SR) reconstruction processes the existing low-resolution images through software technology, restores the lost high-frequency details while improving the image resolution, and has the characteristics of low cost and strong practicability, becoming a research hotspot in the field of image processing.
[0004] The current deep neural networks have achieved good results in the task of image super-resolution. Dong C et al. first proposed a three-layer CNN model SRCNN to learn the end-to-end mapping from LR images to HR images. Since then, a large number of deep methods based on CNNs have been proposed in recent years and have made significant progress. The superior reconstruction performance of methods based on CNNs mainly comes from the deep structure and residual learning. However, since the convolutional neural network adopts a local mechanism, it hinders the establishment of global dependencies and limits the performance of the model. A major breakthrough in the field in recent years was the proposal of the VIT model in 2020, which applied the Transformer mechanism to the visual field and provided a new idea for the design of the super-resolution reconstruction model structure. On the basis of this architecture, Yang et al. designed a new TTSR model with a Ref-Transfermor structure, which better restored the texture information of the image.
[0005] In 2021, Chen et al. proposed a new pre-trained model IPT (Image Processing Transformer) for completing low-level vision tasks such as super-resolution reconstruction. In addition, Liu et al. proposed the Swin Transformer model, which adopts a hierarchical architecture and has the flexibility to model at various scales, and has made improvements in the computational complexity of the Transformer. In the same year, J Liang et al. introduced the Swin Transformer into the field of super-resolution reconstruction. The Swin Transformer combines the advantages of CNN and Transformer, and can not only obtain a larger receptive field through a deep CNN, but also obtain long-range dependencies through the Transformer.
[0006] Despite the success, it has not been fully revealed in the prior art that the Transformer is better than the CNN. An intuitive explanation is that such a network can benefit from the self-attention mechanism and utilize long-range dependency information. In addition, although SwinIR generally obtains higher quantitative performance, due to the limited range of information utilization, it produces results inferior to RCAN on some samples. These phenomena indicate that the Transformer has a stronger local modeling ability for information, but the range of its information utilization can be further expanded. In addition, artifacts will appear in the reconstructed images of SwinIR, which indicates that the shift window mechanism cannot fully achieve cross-window information interaction. Summary of the Invention
[0007] Aiming at the deficiencies of the prior art, the present invention discloses an image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention to solve the problems proposed in the above background technology.
[0008] To achieve the above object, the present invention provides the following technical solutions: An image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention, comprising the following steps:
[0009] S1. Prepare the input training set data: Degrade the high-resolution (HR) image to obtain the corresponding low-resolution (LR) image, and thus construct the training set {L i , H i}, where L i is the low-resolution image, and H i is the corresponding high-resolution image, and the subscript i represents the i-th in the low-resolution or high-resolution image;
[0010] S2. Construct and train an image super-resolution reconstruction Transformer method network with cross-window information aggregation and large kernel attention, where the network structure includes a shallow feature extraction layer, a deep feature extraction layer, and an image reconstruction part. The specific steps are as follows:
[0011] 1) Input the low-resolution image into the shallow feature extraction layer for feature extraction to obtain shallow features.
[0012] 2) Then input the shallow features into the deep feature extraction layer for feature extraction to obtain deep features.
[0013] Among them, the deep feature extraction layer is composed of several residual group (RG) modules and a convolutional layer in series. A residual group (RG) module is composed of several mixed attention blocks (MAB) and a convolutional layer in series.
[0014] Among them, the mixed attention block (MAB) is composed of dual attention TFA and a large kernel attention block (LKAB) in parallel.
[0015] 3) Then input the finally extracted deep features into the upsampling reconstruction module, and use the sub-pixel convolutional layer to upsample the final feature map F D and use the convolutional layer to obtain the final reconstructed image I HQ ;
[0016] 4) And use the loss function to constrain the difference between the predicted reconstructed image and the original high-resolution image, continuously optimize and adjust the network model parameters until the preset conditions are met to obtain a trained image super-resolution reconstruction model.
[0017] S3. Input a low-resolution image to be super-resolved into the trained model to obtain the finally reconstructed high-resolution image.
[0018] Preferably, in step S2:
[0019] First, input the low-resolution image L i ∈R H×W×3 into a shallow feature extraction layer to obtain the shallow features F 0 ∈R H×W×C of the low-resolution image, where H and W respectively represent the height and width of the input image, and C represents the number of feature channels. The formula is as follows:
[0020] F 0 =H Conv3×3 (L i)
[0021] Among them, H Conv3×3 (·) represents a shallow feature extraction operator, which is a 3×3 convolution operation, and then F 0 is used as the input to the deep feature extraction layer.
[0022] Preferably, the shallow feature F 0 is used as the input to the deep feature extraction layer. The structure of the deep feature extraction layer is to first connect k residual group modules in series, then connect a 3×3 convolution in series, and finally add the output of this convolution to F 0 to obtain the final output F D of the deep feature extraction layer module. Among them, the structure of the residual group module is to first connect N hybrid attention modules in series, then connect a 3×3 convolution in series, and finally add the output of this convolution to the input of this module to obtain the output of a residual group module, which is formulated as:
[0023]
[0024] Among them, F RGn represents the output of the nth hybrid attention module; H SAB(.) represents a hybrid attention module; F RGn-1 represents the output of the previous residual group module, and F d represents the output of a residual group module;
[0025] Similarly, for the outputs of all k RG modules, that is, the final output of the deep feature extraction layer is expressed by the formula as:
[0026] F D = H Conv3×3 (F M ) + F 0
[0027] Among them, F M represents the output of the last residual group module, and F D represents the final output of the deep feature extraction layer module.
[0028] Preferably, in order to better perform cross-window information aggregation, the process of the hybrid attention module processing data includes the following:
[0029] Given the input feature map Y, it first passes through the first LayerNorm (LN) layer, and its output Y N is respectively input into the parallel branches of the LKAB and TFA modules. The results of the parallel branches are added to the input Y element by element, and the obtained result Y M then passes through the second LN and a multi-layer perceptron (MLP) and then is added to Y MThe outputs of them are added to obtain the output E of the hybrid attention module, and this process is expressed by the formula as follows:
[0030] Y N = LN(Y)
[0031] Y M = TFA(Y N ) + LKAB(Y N ) + Y
[0032] E = MLP(LN(Y M )) + Y M
[0033] TFA is used to find the window on the feature map that is most relevant to the semantic information of the current window. At the same time, position information is added for constraint, reducing the computational complexity of the model while adding position information for further information fusion. In addition, many works have shown that large-kernel attention can help the network obtain a better receptive field and utilize context information. Therefore, TFA is connected in parallel with an LKAB to integrate the advantages of the large-kernel self-attention mechanism in processing local context information and large receptive fields, improving the model's ability to process global and local information.
[0034] Preferably, in order to better perform cross-window information aggregation, the process of TFA processing the feature map includes the following:
[0035] Given the input feature map X ∈ R H×W×C , first, X is divided into M×M non-overlapping windows, so that each window contains HW / M 2 feature vectors, that is, X is reshaped into a tensor Query (Q), key (K), and value (V) are generated through linear projection: Q = X s W q , K = X s W k , where W q , W k , W v are the projection matrices of Q, K, and V;
[0036] Then, the elements corresponding to the HW / M 2 feature vectors in tensors Q and K are averaged to obtain the window-level query and key Q S , and the affinity matrix Y S representing the semantic relationship between windows is calculated, and it is expressed by the formula as follows:
[0037] Y S = Q S (K S ) T
[0038] Since the matrix Y S is used to measure the semantic relevance between two windows, in order to select the n windows that are most semantically relevant to the current window, only the first n elements of each row of the matrix Y S are retained, and their indices are recorded:
[0039] I S = topnIndex(Y s )
[0040] That is, the index matrix I S contains the indices of the n windows that are most semantically relevant to the window corresponding to each row;
[0041] At the same time, a constraint on the position information is added on the basis of the semantic information and recorded in the position matrix Z S ∈ M 2 × n, that is, Z S contains the Euclidean distances between the center pixels of the window corresponding to each row and the n windows that are most semantically relevant to it. Specifically, it is expressed by the formula:
[0042] Z S = Dist(I s )
[0043] Then, for the query token Q i of the i-th window, according to the indices in the semantic relationship graph I S , the keys and values of the n windows that are most semantically relevant corresponding to the indices are found in the K and V tensors, and are stored row by row in the i-th key matrix and the i-th value matrix respectively, that is
[0044] V i C = converge(V, I S )
[0045] All M 2 key matrices form a key tensor
[0046] The i-th row vector α S in the position matrix Z i ∈ R 1×n is expanded,
[0047] that is, the elements α i in α i,j , (j = 1,... n) are multiplied by the all-ones matrix to obtain
[0048] Let Vi C Multiply element-wise with W i to obtain all M 2 V's i CW to form the value tensor
[0049] Perform attention calculation using the key and value tensors to obtain the position semantic attention map O; which is expressed by the formula as follows:
[0050]
[0051] where d represents the dimension of the key and value.
[0052] Preferably, in order to consider local context information, large receptive fields, and linear complexity, use the large kernel attention module to process features; input the feature map X into the large kernel attention module LKAB, and successively pass through a 1×1 convolution, a GELU activation function, a large kernel attention LKA (Large Kernel Attention), and a 1×1 convolution, and add the obtained result to the input X as the output;
[0053] where the input S of the large kernel attention successively passes through a (2d - 1)×(2d - 1) depthwise convolution DW-Conv(·), a depthwise dilated convolution DW-D-Conv(·) with a dilation rate of d and a size of and a 1×1 convolution, and multiply the obtained attention map element-wise with the input S as the output of the large kernel attention; the large kernel attention is regarded as decomposing a K×K convolution into three parts: depthwise convolution, depthwise dilated convolution, and a 1×1 pointwise convolution, and capturing long-range relationships with a relatively small computational cost and parameters through this decomposition; after obtaining the long-range relationships, estimate the importance of a point and generate an attention map, and the formula of the large kernel attention module is as follows:
[0054] S = GELU(Conv 1×1 (X))
[0055] Attention = Conv 1×1 (DW-D-Conv(DW-Conv(S)))
[0056]
[0057] LKAB(X) = Conv 1×1 (Output) + X
[0058] Among them, Attention is the attention map, and LKAB(X) is the output of the large kernel attention module. The large kernel attention LKA utilizes local information, captures long-range dependencies, and is adaptive in both the channel and spatial dimensions.
[0059] Preferably, the upsampling reconstruction module includes an upsampling layer and a reconstruction layer;
[0060] Among them, the upsampling layer uses a sub-pixel convolutional layer, and H UP (·) is the upsampling operator, as shown in the following formula:
[0061] F UP =H UP (F D )
[0062] Among them, the input F D represents the output of the deep feature extraction layer; the reconstruction layer uses a simple 3×3 convolutional layer to perform the final adjustment and optimization on F UP , and H REC (·) is this process, as shown in the following formula:
[0063] I HQ =H REC (F UP )
[0064] Among them, I HQ is the output of the reconstruction layer, that is, the finally reconstructed high-resolution image.
[0065] Preferably, the degradation processing includes:
[0066] L = (Y * k)↓ s
[0067] Among them, L is the LR image, Y is the HR image, * is the convolution operation, k is the blur kernel, and ↓s represents s-fold downsampling.
[0068] Preferably, the L 1 loss between the high-resolution image and the super-resolution reconstructed image is used as the loss function of the model.
[0069] Compared with the prior art, the beneficial effects of the present invention are:
[0070] 1. A novel hybrid attention Transformer module is proposed in the present invention, which combines large kernel attention. Through the self-attention scheme based on position information and semantic information, a wider receptive field is provided, and the fusion of position and semantic information is added to perform cross-window information fusion, realizing the further utilization of image pixel information, and taking advantage of their complementary advantages of being able to utilize global statistics and powerful local fitting capabilities.
[0071] 2. In the present invention, the performance of the network is further explored by using a parallel large-kernel attention structure. The large-kernel attention can utilize local information, capture long-range dependencies, and is adaptive in both the channel and spatial dimensions. Compared with the highly representative SISR methods in recent years, the method proposed by the present invention can reconstruct higher-quality super-resolution images. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.
[0073] In the drawings:
[0074] Figure 1 is a schematic diagram of the network structure of the image super-resolution reconstruction method of the present invention;
[0075] Figure 2 is a schematic diagram of the network structure of the hybrid attention module in the embodiment of the present invention;
[0076] Figure 3 is a schematic diagram of the network structure of the large-kernel attention module in the embodiment of the present invention;
[0077] Figure 4 is a schematic diagram of the original high-resolution image in the embodiment of the present invention;
[0078] Figure 5 is a schematic diagram of the image obtained after being processed by the EDSR method in the embodiment of the present invention;
[0079] Figure 6 is a schematic diagram of the image obtained after being processed by CARN in the embodiment of the present invention;
[0080] Figure 7 is a schematic diagram of the image obtained after being processed by NGswin in the embodiment of the present invention;
[0081] Figure 8 is a schematic diagram of the image obtained after being processed by SwinIR in the embodiment of the present invention;
[0082] Figure 9 is a schematic diagram of the image obtained after being processed by the image super-resolution reconstruction method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0083] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0084] Current single - image super - resolution (SISR) methods based on Transformer networks still have some deficiencies: When ordinary window self - attention is used to reconstruct high - resolution images, due to the limited receptive field, more windows are ignored. Transformer - based methods have shown impressive performance in low - level vision tasks, such as image super - resolution. Therefore, in order to better utilize the information of different windows for better reconstruction, this method proposes a new hybrid attention model that combines large - kernel attention and a self - attention scheme based on position information and semantic information, thus providing a wider receptive field and leveraging the complementary advantages of their ability to utilize global statistics and powerful local fitting capabilities. In addition, the application of the self - attention mechanism based on position information and semantic information greatly improves the representation ability of the network.
[0085] Embodiment 1:
[0086] As Figure 1 shown, this embodiment details the Transformer method for image super - resolution reconstruction based on cross - window information aggregation and large - kernel attention. The image super - resolution reconstruction method includes the following steps:
[0087] S1. Prepare the input training set data: Degrade the high - resolution (HR) image to obtain the corresponding low - resolution (LR) image, and thus construct the training set {L i , H i}, where L i is the LR image, and H i is the corresponding HR image. The subscript i represents the i - th in the LR or HR image;
[0088] S2. Construct and train an image super-resolution reconstruction Transformer method network with cross-window information aggregation and large kernel attention. The network structure includes a shallow feature extraction layer, a deep feature extraction layer, and an image reconstruction part. The specific steps are as follows: Input a low-resolution image into the shallow feature extraction layer for feature extraction to obtain shallow features. Then input the shallow features into the deep feature extraction layer for feature extraction to obtain deep features. The deep feature extraction layer is composed of several ResidualGroup (RG) modules and a convolutional layer in series. An RG module is composed of several MixedAttentionBlock (MAB) modules and a convolutional layer in series. Among them, an MAB is composed of a Two-FoldAttention (TFA) and a LargeKernelAttentionBlock (LKAB) in parallel. Then input the finally extracted deep features into the upsampling reconstruction module, and use a sub-pixel convolutional layer to upsample the final feature map F D to obtain the final reconstructed image I by using a convolutional layer HQ ; and use a loss function to constrain the difference between the predicted reconstructed image and the original HR image, continuously optimize and adjust the network model parameters until the preset conditions are met to obtain a trained image super-resolution reconstruction model;
[0089] S3. Input an LR image to be super-resolved into the trained model to obtain the finally reconstructed high-resolution image.
[0090] Specifically, in step S2, first input a low-resolution (LR) image L i ∈R H×W×3 into a shallow feature extraction layer to obtain the shallow features F 0 ∈R H×W×C , where H and W respectively represent the height and width of the input image, and C represents the number of feature channels. The formula is as follows:
[0091] F 0 =H Conv3×3 (L i ) (1)
[0092] where H Conv3×3 (·) represents the shallow feature extraction module, which is a 3×3 convolutional operation, and then F 0 is used as the input of the deep feature extraction layer.
[0093] Take the shallow feature F 0 as the input of the deep feature extraction layer. The structure of the deep feature extraction layer is to first connect k RG modules in series, then connect a 3×3 convolution in series, and finally add the output of this convolution to F 0Add them up to get the final output F of the depth feature extraction layer module D , where the structure of the RG module is to first connect N hybrid attention modules in series, then connect a 3×3 convolution in series, and finally add the output of this convolution to the input of this module to obtain the output of one RG module, which is formulated as:
[0094]
[0095] where, F RGn represents the output of the nth hybrid attention module; H SAB(.) represents a hybrid attention module; F RGn-1 represents the output of the previous RG module, and F d represents the output of one RG module;
[0096] Similarly, for the outputs of all k RG modules, that is, the final output of the depth feature extraction layer is expressed by the formula as:
[0097] F D =H Conv3×3 (F M )+F 0 (5)
[0098] F M represents the output of the last RG module, and F D ∈R H×W×C represents the final output of the depth feature extraction layer module.
[0099] As Figure 2 shown, in order to better perform cross-window information aggregation, a hybrid attention module composed of a dual attention TFA and a large kernel attention module in parallel is set. The process of the hybrid attention module processing data is as follows: Given the input feature map Y, first pass through the first LayerNorm (LN) layer, and its output Y N is respectively input into the parallel branches of the LKAB and TFA modules. The results of the parallel branches are added to the input Y element by element, and the obtained result Y M is then passed through the second LN and a multi-layer perceptron (MLP) and then added to Y M to obtain the output E of the hybrid attention module. This process is expressed by the formula as:
[0100] Y N =LN(Y) (6)
[0101] Y M =TFA(Y N )+LKAB(Y N )+Y (7)
[0102] E = MLP(LN(YM )) + Y M (8)
[0103] Find the window on the feature map that is most semantically relevant to the current window through TFA, and at the same time add position information for constraint, reducing the computational complexity of the model while also adding position information for further information fusion; in addition, many works have shown that large-kernel attention can help the network obtain a better receptive field and utilize context information; therefore, TFA is connected in parallel with an LKAB, integrating the advantages of the large-kernel self-attention mechanism for processing local context information and large receptive fields, and improving the model's ability to process global and local information.
[0104] Specifically, to better perform cross-window information aggregation, the process of TFA processing the feature map is as follows: Given the input feature map X ∈ R H×W×C , first divide X into M×M non-overlapping windows, so that each window contains HW / M 2 feature vectors, that is, deform X into a tensor Generate queries (Q), keys (K), and values (V) through linear projection:
[0105]
[0106] where W q , W k , W v are the projection matrices of Q, K, and V.
[0107] Then, take the average of the elements corresponding to the HW / M 2 feature vectors in tensors Q and K to obtain window-level queries and keys Q S , and calculate the affinity matrix Y S representing the semantic relationship between windows. This step can be expressed by the formula:
[0108] Y S = Q S (K S ) T (10)
[0109] Since the matrix Y S is used to measure the semantic correlation between two windows, to select the n windows that are most semantically relevant to the current window, only keep the first n elements of each row of the matrix Y S , and record their indices:
[0110] I S = topnIndex(Y s ) (11)
[0111] That is, the index matrix I Scontains the indices of the n windows that are most semantically relevant to the window semantics corresponding to each row; in addition, this method also adds a constraint on the position information based on the semantic information, which is recorded in the position matrix Z S ∈M 2 ×n, that is, Z S contains the Euclidean distances between the center pixels of the n windows that are most semantically relevant to the window corresponding to each row. Specifically, it is expressed by the formula:
[0112] Z S = Dist(I s ) (12)
[0113] Then, for the query token Q i of the i-th window, according to the indices in the semantic relationship graph I S , find the keys and values of the n semantically most relevant windows corresponding to the indices in the K and V tensors, and store them row by row in the i-th key matrix and the i-th value matrix , that is
[0114]
[0115] All M 2 key matrices form the key tensor
[0116] Expand the i-th row vector α S in the position matrix Z i ∈R 1×n , that is, multiply the elements α i in α i,j , (j = 1,...n) by the all-ones matrix to get Let V i C be multiplied element-wise with W i to get All M 2 V i CW form the value tensor
[0117] Use the key and value tensors to calculate the attention, and obtain the position-semantic attention map O; it is expressed by the formula as follows:
[0118]
[0119] where d represents the dimension of the key and value.
[0120] Specifically, for example Figure 3As shown, in order to consider local context information, large receptive fields, and linear complexity, a large kernel attention module is used to process features; the feature map X is input into the LKAB, passing through a 1×1 convolution, a GELU activation function, a large kernel attention (LKA), and a 1×1 convolution in sequence, and the resulting output is added to the input X as the final output; among them, the input S of the large kernel attention passes through a (2d - 1)×(2d - 1) depthwise convolution DW-Conv(·), a depthwise dilated convolution DW-D-Conv(·) with an expansion rate of d and size of, and a 1×1 convolution, and the resulting attention map is element-wise multiplied by the input S as the output of the large kernel attention; the large kernel attention can be regarded as decomposing a K×K convolution into three parts: depthwise convolution, depthwise dilated convolution, and a 1×1 pointwise convolution. Through this decomposition, long-range relationships can be captured with relatively low computational cost and fewer parameters; after obtaining the long-range relationships, the importance of a point can be estimated and an attention map can be generated. The formula of the large kernel attention module is as follows:
[0121] S = GELU(Conv 1×1 (X)) (15)
[0122] Attention = Conv 1×1 (DW-D-Conv(DW-Conv(S))) (16)
[0123]
[0124] LKAB(X) = Conv 1×1 (Output)+X (18)
[0125] where Attention is the attention map, LKAB(X) is the output of the large kernel attention module, and LKA can utilize local information, capture long-range dependencies, and is adaptive in both channel and spatial dimensions.
[0126] Specifically, the upsampling and reconstruction module includes an upsampling layer and a reconstruction layer;
[0127] Among them, the upsampling layer uses a sub-pixel convolutional layer, and let H UP(·) be the upsampling operator, as shown in the following formula:
[0128] F UP = H UP (F D ) (19)
[0129] where the input F DRepresents the output of the last feature fusion layer; the reconstruction layer uses a simple convolutional layer to upsample the output F UP For the final adjustment and optimization, set H REC(·) For this process, as shown in the following equation:
[0130] I HQ = H REC (F UP ) (20)
[0131] where I HQ is the output of the reconstruction layer, that is, the finally reconstructed high-resolution image.
[0132] During the training process, the L 1 loss between the high-resolution image and the super-resolution reconstructed image is used as the loss function of the model.
[0133] Example 2:
[0134] This example conducts experimental illustration on the image super-resolution Transformer method based on cross-window information aggregation and large kernel attention of the present invention, and conducts performance comparative analysis with existing methods.
[0135] In the experimental setup section, the size of the window in this example is set to 8×8. This example follows the work of most SISR tasks to train and test the method of the present invention.
[0136] This example conducts experiments under a ×4 magnification factor. Two objective evaluation metrics, peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), are used to quantitatively evaluate the performance of the method proposed by the present invention.
[0137] Specifically, this example trains on a dataset DIV2K of 800 images and tests on a benchmark dataset Urban100. In the training stage, the size of the input LR image patches is set to 64×64, the batch size is set to 2, and the number of training iterations is 500,000 times. This example uses random rotation by 90 degrees, 180 degrees, 270 degrees, and horizontal flipping for data augmentation of the training dataset. The ADAM optimizer is used, with β 1 = 0.9, β 2 = 0.999. The initial learning rate is set to 2×10 -4 , and it is halved at 250,000 times, 300,000 times, 450,000 times, and 475,000 times of iteration respectively.
[0138] In this example, the images in the BSD100 dataset are used as test images. Table 1 lists the average PSNR and SSIM of the result images obtained by this example and other advanced image super-resolution methods.
[0139] As can be seen from Table 1, the results obtained by the method of this embodiment have been significantly improved compared with the results of other methods.
[0140]
[0141] It can be known from the comparison in Table 1 that the method of the present invention has obtained the highest average PSNR and the highest average SSIM compared with EDSR, CARN, NGswin, and SwinIR.
[0142] In addition to the objective evaluation indicators, in this embodiment, the method proposed by the present invention is also compared with the methods of EDSR, CARN, NGswin, and SwinIR on the subjective visual graph. The subjective visual graph is as Figures 4 to 9 shown. It can be seen from the figure that compared with other methods, the method proposed by the present invention can better restore the structural information and texture information and is closer to the real HR image.
[0143] Finally, it should be noted that the above are only preferred examples of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. Image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention, characterized by: The following steps are involved: S1. Prepare the input training set data: Degrade the high-resolution image to obtain the corresponding low-resolution image, and construct the training set {L i ,H i }, where L i is a low-resolution image, H i is the corresponding high-resolution image, and the subscript i represents the i-th low-resolution or high-resolution image; S2. Construct and train a network of image super-resolution reconstruction Transformer method with cross-window information aggregation and large kernel attention, where the network structure includes a shallow feature extraction layer, a deep feature extraction layer and an image reconstruction part; the specific steps are: 1) Input the low-resolution image into the shallow feature extraction layer for feature extraction to obtain shallow features; 2) Then the shallow features are input into the deep feature extraction layer for feature extraction to obtain deep features. Among them, the deep feature extraction layer is composed of several residual group modules and a convolutional layer in series; a residual group module is composed of several hybrid attention modules and a convolutional layer in series; Among them, the hybrid attention module is composed of dual attention TFA and a large core attention module LKAB in parallel; 3) Then the final extracted deep features are input into the upsampling reconstruction module, and the final feature map F is reconstructed using the sub-pixel convolution layer. D Upsample and use the convolution layer to get the final reconstructed image I HQ ; 4) The difference between the predicted reconstructed image and the original high-resolution image is constrained by the loss function, and the network model parameters are continuously optimized and adjusted until the preset conditions are met to obtain a trained image super-resolution reconstruction model; S3. Input a low-resolution image to be super-resolved into the trained model to obtain the final reconstructed high-resolution image.
2. The image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention according to claim 1 is characterized in that: In step S2: First, the low-resolution image L i ∈R H×W×3 Input a shallow feature extraction layer to obtain the shallow features F0∈R of the low-resolution image H×W×C , where H and W represent the height and width of the input image respectively, and C represents the number of feature channels. The formula is as follows: F0=H Conv3×3 (L i ) Among them, H Conv3×3 (·) represents the shallow feature extraction operator, which is a 3×3 convolution operation, and then F0 is used as the input of the deep feature extraction layer.
3. The image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention according to claim 2, characterized in that: The shallow feature F0 is used as the input of the deep feature extraction layer. The structure of the deep feature extraction layer is to first connect k residual group modules in series, then connect a 3×3 convolution in series, and finally add the output of the convolution to F0 to obtain the final output F of the deep feature extraction layer module. D , where the structure of the residual group module is to first connect N hybrid attention modules in series, then connect a 3×3 convolution in series, and finally add the output of the convolution to the module input to obtain the output of a residual group module, which is formulated as: Among them, F RGn represents the output of the nth hybrid attention module; H SAB(.) represents a hybrid attention module; F RGn-1 represents the output of the previous residual group module, F d Represents the output of a residual group module; Similarly, the output of all k RG modules, that is, the final output of the deep feature extraction layer, is expressed as: F D =H Conv3×3 (F M )+F0 Among them, F M represents the output of the last residual group module, F D Represents the final output of the deep feature extraction layer module.
4. The image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention according to claim 3 is characterized by: The process of data processing in the hybrid attention module includes the following: Given an input feature map Y, it first passes through the first LayerNorm layer, and its output Y N The output is input into the parallel branches of the large core attention module LKAB and the TFA module respectively. The output of the parallel branch is added element by element to the input Y. The result Y M After the second LayerNorm layer and a multi-layer perceptron, it is combined with Y M The output E of the hybrid attention module is obtained by adding them together. The process is expressed as follows: Y N =LN(Y) AND M =TFA(Y N )+LKAB(Y N )+Y E=MLP(LN(Y M ))+Y M The TFA module is used to find the window on the feature map that is most relevant to the semantic information of the current window, and position information is added for constraint, which reduces the amount of model calculation and further integrates the information while adding position information.
5. The image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention according to claim 4, characterized in that: The process of TFA module processing feature map includes the following: Given an input feature map X∈R H×W×C , first divide X into M×M non-overlapping windows, so that each window contains HW / M 2 feature vectors, transforming X into a tensor Generate query (Q), key (K), value (V) through linear projection: Q = X s W q , K = X s W k , Where W q , W k , W v is the projection matrix of Q, K, V; Then, for the HW / M in the tensors Q and K 2 The elements corresponding to the feature vectors are averaged to obtain the window-level query and key And calculate the affinity matrix Y representing the semantic relationship between windows S , expressed as: Y S =Q S (K S ) T Select the n windows that are most relevant to the current window semantics and retain the matrix Y S The first n elements of each row are recorded with their indexes: I S =topnIndex(Y s ) That is, the index matrix I S Contains the indices of the n windows that are most relevant to the window semantics corresponding to each row; At the same time, the constraints of position information are added on the basis of semantic information and recorded in the position matrix Z S ∈M 2 ×n, that is, Z S contains the Euclidean distance between the window corresponding to each row and the center pixels of the n windows that are most semantically relevant to it, which can be expressed as: Z S =Dist(I s ) Then, for the query token Q of the i-th window i , according to the semantic relationship graph I S The index in , finds the key and value of the n most semantically relevant windows corresponding to the index in the K, V tensors, and stores them row by row in the i-th key matrix and the i-th value matrix In All M 2 The key matrices form the key tensor For the position matrix Z S The i-th row vector α i ∈R 1×n To expand, Ready-to-use α i The element α in i,j ,(j=1,...n) and the all-one matrix Multiply to get make With W i Multiply element by element to get All M 2 V i CW Composition value tensor The key and value tensors are used to calculate the attention and obtain the position semantic attention map O; it can be expressed as follows: Where d represents the dimension of key and value.
6. The image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention according to claim 4, characterized in that: Use the large core attention module LKAB to process the features; input the feature map X into the large core attention module LKAB, and then go through a 1×1 convolution, a GELU activation function, a large core attention LKA and a 1×1 convolution, and add the result to the input X as the output; The input S of the large core attention is sequentially passed through a (2d-1)×(2d-1) deep convolution DW-Conv(·), an expansion rate of d, and a size of The deep hole convolution DW-D-Conv(·) and a 1×1 convolution are used to obtain the attention map, which is element-wise multiplied with the input S as the output of the large core attention. The large core attention is regarded as decomposing a K×K convolution into three parts: deep convolution, deep hole convolution and a 1×1 channel convolution. Through this decomposition, the long-range relationship is captured with a small computational cost and parameters. After obtaining the long-range relationship, the importance of a point is estimated and the attention map is generated. The formula of the large core attention module is as follows: S=GELU(Conv. 1×1 (X)) Attention=Conv 1×1 (DW-D-Conv(DW-Conv(S))) LKAB(X)=Conv 1×1 (Output)+X Among them, Attention is the attention map, LKAB(X) is the output of the large-core attention module, and the large-core attention LKA uses local information to capture long-range dependencies and is adaptive in both channel and spatial dimensions.
7. The image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention according to claim 1, characterized in that: The upsampling and reconstruction module includes an upsampling layer and a reconstruction layer; The upsampling layer uses a sub-pixel convolution layer, setting H UP (·) is the upsampling operator, as follows: F UP =H UP (F D ) Among them, input F D Represents the output of the deep feature extraction layer; the reconstruction layer uses a simple 3×3 convolutional layer to transform F UP Make final adjustments and optimizations, set H REC (·) is the process, as follows: I HQ =H REC (F UP ) Among them, I HQ is the output of the reconstruction layer, that is, the final reconstructed high-resolution image.
8. The image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention according to claim 1, characterized in that: Degradation treatment includes: L=(Y*k)↓ s Among them, L is the low-resolution image, Y is the high-resolution image, * is the convolution operation, k is the blur kernel, and ↓s means downsampling by s times.
9. The image super-resolution reconstruction Transformer method based on cross-window information aggregation and large kernel attention according to claim 1, characterized in that: The L1 loss of both the high-resolution image and the super-resolution reconstructed image is used as the loss function of the model.
Citation Information
Cited By
Unmanned aerial vehicle image super-resolution method based on multi-scale mixed attention
CN120655511A