Image Restoration Method Based on Cross-Asymmetric Convolution Multi-Aggregation Self-Attention Network

By using cross-asymmetric convolution and multi-aggregation self-attention networks in remote sensing image processing, the problems of limited nonlinear feature capture capabilities and high calculation cost of attention mechanisms in the prior art are solved, and more efficient feature extraction and image recovery effects are achieved.

CN120047336BActive Publication Date: 2025-06-20SHAANXI SILK ROAD DIGITAL INTELLIGENT NAVIGATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510517758.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-20
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing remote sensing image processing technology based on convolutional networks has problems such as limited nonlinear feature capture capabilities, high attention mechanism calculation cost, and low global information utilization rate. It is difficult to effectively extract features of complex image scenes, affecting the image recovery effect.

Method used

The image restoration method based on cross-asymmetric convolution multi-aggregation self-attention network is adopted, and deep feature extraction and image reconstruction are realized through shallow feature extraction module, cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module, decoder module, upsampling module and image reconstruction recovery module.

Benefits of technology

Through the cross-asymmetric convolution and multi-aggregation self-attention mechanism, the insufficient feature components extracted by single-dimensional convolution can be fully extracted, the calculation amount can be reduced, and the accuracy of feature extraction and image recovery quality can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047336B_ABST
    Figure CN120047336B_ABST
Patent Text Reader

Abstract

The present invention discloses an image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network. S1: Extract shallow information from a low-resolution remote sensing image to obtain shallow image feature information. S2: Input the shallow feature information into a cross-asymmetric convolution module to obtain cross-asymmetric convolution feature information. S3: Perform multi-aggregation self-attention extraction on the cross-asymmetric convolution feature information to obtain multi-aggregation self-attention features. By means of two designed cross-asymmetric convolution blocks and two cross operations, the information of the upper and lower two paths of the first and second paths can be not lost, and at the same time, the feature components that are not fully extracted by the single-dimensional convolution can be fully extracted, ensuring that the extracted feature information will not deteriorate. Through multi-aggregation self-attention extraction, multi-aggregation self-attention features are obtained, which can reduce the amount of calculation while giving more-dimensional binary weights to the target image to more accurately obtain the required features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing, and specifically relates to an image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network. Background Art

[0002] Remote sensing image processing technology has evolved from traditional probability solving to deep learning. With the development of neural networks, current methods centered on convolutional neural networks have made remarkable progress in tasks such as classification, detection, and segmentation. However, traditional symmetric convolutions have limited ability to capture non-linear features, and there are problems such as high computational cost and low utilization rate of global information in the attention mechanism. Existing remote sensing image processing technologies based on convolutional networks mainly optimize the convolutional network. However, simply stacking the number of convolutional layers, adjusting parameters, and changing channels cannot enable the convolutional network to fully extract complex features and high-frequency information. If too many convolutional layers are stacked, it will also lead to an increase in the number of parameters. Existing attention mechanisms mostly use simple concatenations of single attentions such as channel attention and spatial attention, and cannot effectively extract complex image scenes such as remote sensing images. At the same time, simply introducing attention will also lead to insufficient extraction of information of interest in specific data, affecting the final image restoration effect, so optimization is required.

[0003] Therefore, the present invention proposes an image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network. Summary of the Invention

[0004] The purpose of the present invention is to provide an image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: An image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network, including a shallow feature extraction module, a cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module, a decoder module, an upsampling module, and an image reconstruction and restoration module; the cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module includes a cross-asymmetric convolution module and a multi-aggregation self-attention module;

[0006] The specific steps include: S1: Extract shallow information from a low-resolution remote sensing image to obtain image shallow feature information;

[0007] S2: Transmit the shallow feature information into the cross-asymmetric convolution module to obtain cross-asymmetric convolution feature information;

[0008] S3: Perform multi-aggregation self-attention extraction on the cross-asymmetric convolution feature information to obtain multi-aggregation self-attention features, and perform the next cross-asymmetric convolution multi-aggregation self-attention extraction on the features after combining the cross-asymmetric convolution features with the multi-aggregation self-attention features;

[0009] S4: Repeat steps S2 and S3, stack multiple cross-asymmetric convolution modules and multi-aggregation self-attention modules to obtain the deep features;

[0010] S5: Input the deep features into the decoder module, stack them with the shallow features to obtain the total extracted features, and then obtain the restored high-resolution remote sensing image through the upsampling module and the image reconstruction and restoration module.

[0011] Preferably, step S1 specifically includes:

[0012] S11: Construct a remote sensing image dataset and obtain low-resolution remote sensing images from high-resolution remote sensing images;

[0013] S12: Perform shallow feature extraction on the low-resolution remote sensing images through a shallow extraction convolution block to obtain a shallow feature image.

[0014] Preferably, the shallow feature image undergoes two cross-asymmetric feature extractions, including: the shallow feature image undergoes the first cross-asymmetric feature extraction, two-way parallel asymmetric convolutions to obtain two sets of dimensional features. For the first-way asymmetric convolution layer and the second-way asymmetric convolution layer, horizontal single-dimensional asymmetric convolution and vertical single-dimensional asymmetric convolution are respectively used for extraction to obtain horizontal single-dimensional features and vertical single-dimensional features;

[0015] After one extraction, cross-connection is performed. The single-dimensional features extracted by the horizontal single-dimensional asymmetric convolution are input into the first way and the second way, and the single-dimensional features extracted by the vertical single-dimensional asymmetric convolution are input into the first way and the second way for the second extraction;

[0016] During the second extraction, the first way and the second way respectively use vertical single-dimensional asymmetric convolution and horizontal single-dimensional asymmetric convolution for extraction to obtain vertical single-dimensional features and horizontal single-dimensional features, and perform the third extraction;

[0017] After the second extraction, the vertical single-dimensional features extracted in the second extraction are input into the first way and the second way, and the horizontal single-dimensional features extracted in the second extraction are input into the first way and the second way for the third extraction;

[0018] During the third extraction, the first way and the second way respectively use horizontal single-dimensional asymmetric convolution and vertical single-dimensional asymmetric convolution for extraction to obtain horizontal single-dimensional features and vertical single-dimensional features, obtain the third extraction features, and aggregate the features of the two ways;

[0019] The feature image obtained by the first cross - asymmetric feature extraction is subjected to the second cross - asymmetric feature extraction. The extraction steps are as follows: the third - extraction feature of the first cross - asymmetric is passed through two - path parallel asymmetric convolutions to obtain two groups of dimensional features. For the first - path asymmetric convolution layer and the second - path asymmetric convolution layer, vertical single - dimensional asymmetric convolution and horizontal single - dimensional asymmetric convolution are respectively used for extraction to obtain vertical single - dimensional features and horizontal single - dimensional features;

[0020] After the above - mentioned first extraction of the second cross - asymmetric feature through two - path parallel asymmetric convolutions, for the obtained vertical single - dimensional features and horizontal single - dimensional features, cross - connection is performed. The single - dimensional features extracted by the vertical single - dimensional asymmetric convolution are passed into the first path and the second path, and the single - dimensional features extracted by the horizontal single - dimensional asymmetric convolution are passed into the first path and the second path for the second extraction;

[0021] During the second extraction, the first path and the second path respectively use horizontal single - dimensional asymmetric convolution and vertical single - dimensional asymmetric convolution for extraction to obtain horizontal single - dimensional features and vertical single - dimensional features, and then perform the third extraction;

[0022] After two extractions, the horizontally single - dimensional features obtained from the second extraction are passed into the first path and the second path, and the vertically single - dimensional features obtained from the second extraction are passed into the first path and the second path for the third extraction;

[0023] During the third extraction, the first path and the second path respectively use vertical single - dimensional asymmetric convolution and horizontal single - dimensional asymmetric convolution for extraction to obtain vertical single - dimensional features and horizontal single - dimensional features, and the features of the two paths are aggregated to obtain the third - extraction feature.

[0024] Preferably, multi - aggregation self - attention extraction is performed on the cross - asymmetric convolution:

[0025] The multi - aggregation self - attention includes spatial graph attention and channel graph attention. For the input features, first, they are set as query Query, key Key, and value Value matrices through linear projection 、 、 . Then, the product of the query and the key - value is calculated to obtain the channel attention map, and after Softmax normalization, it is multiplied by the value matrix to obtain the channel attention aggregation result . Then, the input query Query, key Key, and value Value matrices are transposed to obtain spatial query Query, key Key, and value Value matrices 、 、 , and through calculation Query the product with the key value to obtain the spatial attention map, and after normalizing it through Softmax, multiply it with the Value matrix to obtain the spatial attention aggregation result . Finally, perform an inverse transpose operation to output the result and convert it to the original dimension, and use an activation function to increase its non-linearity to obtain the final multi-aggregation result . Then, pass the final multi-aggregation input result to the output through a residual connection to ensure the stability of its extraction effect.

[0026] Preferably, multiple S2 and S3 are repeatedly concatenated, and a residual connection is introduced in S2 and S3 to obtain the cross-asymmetric convolutional multi-aggregation self-attention network.

[0027] Preferably, the network is iteratively trained, and the loss value for each weight parameter modification during the iterative process is calculated using the L1 loss function: where represents the cross-asymmetric convolutional multi-aggregation self-attention network model, represents the hyperparameters corresponding to the network, is the remotely sensed image after reconstruction and recovery, is the corresponding reference true remotely sensed image, represents the number of remotely sensed images.

[0028] Preferably, for the first cross-asymmetric feature extraction and the second cross-asymmetric feature extraction, a residual connection is used to splice the input feature to the output, and local aggregation features can be obtained.

[0029] Compared with the prior art, the beneficial effects of the present invention are:

[0030] By means of the two designed cross-asymmetric convolutional blocks, compared with using standard convolutions and simple stacking of asymmetric convolutions, the two cross operations can prevent the loss of information in the upper and lower two paths of the first path and the second path, and at the same time can fully extract the feature components that are not fully extracted by one-dimensional convolutions, ensuring that the extracted feature information will not deteriorate.

[0031] Through multi-aggregation self-attention extraction, multi-aggregation self-attention features are obtained. Compared with simply using channel attention, spatial attention, or a simple combination of the two, by introducing the idea of self-attention, while reducing the computational amount, more-dimensional binary weights can be given to the target image to more accurately obtain the required features.

[0032] The present invention also introduces residual connections in different modules to extract local feature information, multi-aggregation information, and global feature information, enhance the target receptive field, eliminate unnecessary receptive fields, and improve the restoration quality of images. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic flowchart of the method of the present invention;

[0034] Figure 2 is a diagram of the first cross-asymmetric convolution block and the second cross-asymmetric convolution block of the present invention;

[0035] Figure 3 is a diagram of the multi-aggregation self-attention network of the present invention;

[0036] Figure 4 is a schematic flowchart of the principle of the cross-asymmetric convolution multi-aggregation self-attention of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Please refer to Figure 1 and Figure 4 , the present invention provides a technical solution: a remote sensing image processing method based on a cross-asymmetric convolution multi-aggregation self-attention network. S1: Extract shallow information from a low-resolution remote sensing image to obtain shallow feature information of the image;

[0039] In this embodiment, specifically, first, the input low-resolution image is passed through a and convolution to perform shallow feature extraction and obtain a shallow feature information of the remote sensing image;

[0040] S2: Input the shallow feature information into the cross-asymmetric convolution module to obtain cross-asymmetric convolution feature information;

[0041] In this embodiment, specifically, further extract the shallow feature information through the cross-asymmetric convolution module. For example, Figure 2 , the cross-asymmetric convolution module is divided into a first cross-asymmetric convolution block and a second cross-asymmetric convolution block, including:

[0042] Copy the shallow feature information and input it into two paths for three extractions. For the first path, use convolution, convolution, Feature extraction is performed by convolution. For the second path, convolution, convolution, and convolution are used for feature extraction. For the first and second feature extractions, cross operations are performed for each path, that is, the features extracted from the first path are passed into the first and second paths of the next extraction, and the features extracted from the second path are passed into the first and second paths of the next extraction. After the third extraction, the features of the first and second paths are aggregated to obtain the first cross-asymmetric convolution feature;

[0043] The first asymmetric convolution feature information is copied and then passed into a dual path for three extractions. For the first path, convolution, convolution, and convolution are used for feature extraction. For the second path, convolution, convolution, and convolution are used for feature extraction. For the first and second feature extractions, cross operations are performed for each path, that is, the features extracted from the first path are passed into the first and second paths of the next extraction, and the features extracted from the second path are passed into the first and second paths of the next extraction. After the third extraction, the features of the first and second paths are aggregated to obtain the second cross-asymmetric convolution feature;

[0044] Meanwhile, residual connections are introduced during both the first cross-asymmetric feature extraction and the second cross-asymmetric feature extraction, that is, the corresponding input features are directly passed into the output to obtain local aggregated feature information;

[0045] S3: Perform multi-aggregation self-attention extraction on the cross-asymmetric convolution feature information to obtain multi-aggregation self-attention features, and perform the next cross-asymmetric convolution multi-aggregation self-attention extraction on the features after combining the cross-asymmetric convolution features and the multi-aggregation self-attention features;

[0046] S4: Repeat steps S2 and S3, stack multiple cross-asymmetric convolution modules and multi-aggregation self-attention modules to obtain the deep features;

[0047] Performing multi-aggregation self-attention extraction on the features extracted by the cross-asymmetric convolution module includes:

[0048] Performing two multi-aggregation self-attention extractions of spatial graph attention and channel graph attention. For the input features, first project them into query (Query), key (Key), and value (Value) matrices through linear projection 、 、 , and then calculate the product of the query and the key-value to obtain the channel attention map, specifically as follows:

[0049] , , Among them, represents the input feature, , , respectively represent the linear projection matrices corresponding to the query, key, and value, and after multiplying with the value matrix after Softmax normalization, the channel attention aggregation result is obtained , Among them, represents the transpose matrix of, and then, in order to perform spatial attention aggregation, for , , , they are respectively rotated , , , among which, represents the rotation operation. By calculating the product of the query and the key-value, the spatial attention map is obtained, and after Softmax normalization, it is multiplied with the value matrix to obtain the spatial attention aggregation result , Finally, the output result is converted to the original dimension by an inverse transpose operation, and its nonlinearity is increased by an activation function to obtain the final multi-aggregation result , , and then the final multi-aggregation input result is given to the output through a residual connection to ensure the stability of its extraction effect. Through this self-attention aggregation method, pixels in the channel dimension and the spatial dimension can be aggregated, so as to realize the all-directional multi-aggregation self-attention extraction; in the process of multi-aggregation self-attention extraction, residual connections or attention local information are also introduced to make the extracted features more abundant;

[0050] Stack the cross-asymmetric convolution multi-aggregation self-attention modules multiple times to obtain deep feature information;

[0051] S5: Input the deep features into the decoder module, stack them with the shallow features to obtain the total extracted features, and then obtain the restored high-resolution remote sensing image through the upsampling module and the image reconstruction and restoration module;

[0052] In this embodiment, specifically, a convolution with a convolution kernel is used as the decoder for feature decoding, and at the same time, a residual connection is used to fuse the shallow features and the cross-asymmetric convolution multi-aggregation self-attention deep features to obtain the fusion information with texture structure detail information and image low-frequency information;

[0053] Input the fusion information into the upsampling module and the image reconstruction and restoration module to obtain the restored high-resolution remote sensing image. The upsampling uses a sub-pixel convolutional layer, which is a Convolution, with 64 input channels and 256 output channels. The sub-pixel convolution layer can rearrange the feature map with a size of into a feature map of . The reconstruction recovery output consists of a convolution with a kernel size of , and its function is to convert the 64 output channels into 3 channels.

[0054] Except for the input module, the final upsampling module, and the image reconstruction recovery module, other modules are performed in the channel dimension of 64.

[0055] Although the embodiments of the present invention have been shown and described, see the above detailed description. For those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image restoration method based on a cross-asymmetric convolutional multi-aggregation self-attention network, characterized by: It includes a shallow feature extraction module, a cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module, a decoder module, an upsampling module and an image reconstruction and restoration module; the cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module includes a cross-asymmetric convolution module and a multi-aggregation self-attention module; The specific steps include: S1: extracting shallow information from low-resolution remote sensing images to obtain shallow feature information of the images; S2: passing the shallow feature information into the cross asymmetric convolution module to obtain cross asymmetric convolution feature information; The shallow feature image is subjected to two cross asymmetric feature extractions, including: the shallow feature image is subjected to a first cross asymmetric feature extraction, two parallel asymmetric convolutions, to obtain two sets of dimensional features, and for the first asymmetric convolution layer and the second asymmetric convolution layer, horizontal single-dimensional asymmetric convolutions and vertical single-dimensional asymmetric convolutions are respectively used to extract, to obtain horizontal single-dimensional features and vertical single-dimensional features; S3: Perform multi-aggregate self-attention extraction on the cross asymmetric convolution feature information to obtain multi-aggregate self-attention features, combine the cross asymmetric convolution features with the features after the multi-aggregate self-attention features, and perform the next cross asymmetric convolution multi-aggregate self-attention extraction; Perform multi-aggregation self-attention extraction on the cross asymmetric convolution: The multi-aggregate self-attention includes spatial graph attention and channel graph attention. For the input features, they are first set as query, key, and value matrices through linear projection. , , , then calculate the query With key value The product of , get the channel attention map, and then normalize it with the value matrix after Softmax The product of , we get the channel attention aggregation result. , then transpose the input query, key, and value matrices to obtain the spatial query, key, and value matrices , , , by calculating Query and key value The spatial attention map is obtained by multiplying the value matrix by Softmax normalization. Multiply them together to get the spatial attention aggregation result Finally, the inverse permutation operation is used to convert the output result into the original dimension, and the activation function is used to increase its nonlinearity to obtain the final multi-aggregation result. , and then the final multi-aggregation input result is given to the output through residual connection to ensure the stability of its extraction effect; S4: repeat steps S2 and S3, stack multiple cross asymmetric convolution modules and multi-aggregation self-attention modules to obtain the deep features; S5: The deep features are passed into the decoder module, and after being superimposed with the shallow features, the total extracted features are obtained, and then the restored high-resolution remote sensing image is obtained through the upsampling module and the image reconstruction and restoration module.

2. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 1, characterized in that: Step S1 specifically includes: S11: Construct remote sensing image datasets and obtain low-resolution remote sensing images from high-resolution remote sensing images; S12: Performing shallow feature extraction on the low-resolution remote sensing image through a shallow extraction convolution block to obtain a shallow feature image.

3. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 2 is characterized in that: After one extraction, a cross connection is performed, and the single-dimensional features extracted by the horizontal single-dimensional asymmetric convolution are passed into the first path and the second path, and the single-dimensional features extracted by the vertical single-dimensional asymmetric convolution are passed into the first path and the second path for a second extraction; During the second extraction, the first path and the second path respectively use vertical single-dimensional asymmetric convolution and horizontal single-dimensional asymmetric convolution to extract, obtain vertical single-dimensional features and horizontal single-dimensional features, and perform a third extraction; After the second extraction, the vertical single-dimensional features extracted for the second time are passed into the first and second paths, and the horizontal single-dimensional features extracted for the second time are passed into the first and second paths for the third extraction; During the third extraction, the first path and the second path are extracted by horizontal single-dimensional asymmetric convolution and vertical single-dimensional asymmetric convolution respectively to obtain horizontal single-dimensional features and vertical single-dimensional features, obtain the third extraction features, and aggregate the two-path features; The feature image obtained by the first cross asymmetric feature extraction is subjected to the second cross asymmetric feature extraction. The extraction step is to obtain two sets of dimensional features by subjecting the third extraction features of the first cross asymmetric to two parallel asymmetric convolutions. For the first asymmetric convolution layer and the second asymmetric convolution layer, vertical single-dimensional asymmetric convolution and horizontal single-dimensional asymmetric convolution are respectively used to extract, so as to obtain vertical single-dimensional features and horizontal single-dimensional features. After the second cross asymmetric feature is extracted for the first time through two-way parallel asymmetric convolution, the vertical single-dimensional feature and the horizontal single-dimensional feature are cross-connected, the single-dimensional feature extracted by the vertical single-dimensional asymmetric convolution is passed to the first path and the second path, and the single-dimensional feature extracted by the horizontal single-dimensional asymmetric convolution is passed to the first path and the second path for a second extraction; During the second extraction, the first and second paths are extracted using horizontal single-dimensional asymmetric convolution and vertical single-dimensional asymmetric convolution respectively to obtain horizontal single-dimensional features and vertical single-dimensional features for the third extraction; After the second extraction, the horizontal single-dimensional features extracted for the second time are passed into the first path and the second path, and the vertical single-dimensional features extracted for the second time are passed into the first path and the second path for the third extraction; During the third extraction, the first and second paths are extracted using vertical single-dimensional asymmetric convolution and horizontal single-dimensional asymmetric convolution respectively to obtain vertical single-dimensional features and horizontal single-dimensional features, and the two-path features are aggregated to obtain the third extraction features.

4. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 3, characterized in that: Repeatedly splice multiple S2 and S3, and introduce residual connections in S2 and S3 to obtain the cross asymmetric convolutional multi-aggregation self-attention network.

5. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 1, characterized in that: The network is iteratively trained, and the loss value of each weight parameter modification during the iteration process is calculated using the L1 loss function: in, represents the cross asymmetric convolutional multi-aggregation self-attention network model, represents the hyperparameters corresponding to the network, is the reconstructed and restored remote sensing image. is with The corresponding benchmark real remote sensing image, Represents the number of remote sensing images.

6. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 4, characterized in that: For the first cross asymmetric feature extraction and the second cross asymmetric feature extraction, residual connection is used to splice the input features to the output, wherein the second cross asymmetric feature extraction takes the aggregated features of the first cross asymmetric feature extraction as input to obtain local aggregated features.

Citation Information

Patent Citations

  • Face super-resolution method and device based on cross convolution attention adversarial learning

    CN114757832A

  • KR20230141376A