Image restoration method based on cross asymmetric convolution multi-aggregation self-attention network
By using cross-asymmetric convolution and multi-aggregation self-attention networks in remote sensing image processing, the problem of difficulty in extracting complex features and high-frequency information in the prior art is solved, and a more efficient image recovery effect is achieved.
Patent Information
- Application Number
- CN202510517758.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing remote sensing image processing technology based on convolutional networks is difficult to fully extract complex features and high-frequency information, and the attention mechanism calculation cost is high and the global information utilization rate is low, which affects the image recovery effect.
The image restoration method based on cross-asymmetric convolution multi-aggregation self-attention network is adopted, and deep feature extraction and image reconstruction are realized through shallow feature extraction module, cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module, decoder module, upsampling module and image reconstruction recovery module.
Through cross-asymmetric convolution and multi-aggregation self-attention mechanisms, feature components can be fully extracted, computational amounts can be reduced, image recovery quality can be improved, and target receptive fields can be enhanced.
Smart Images

Figure CN120047336A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing technology, and particularly relates to an image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network. Background Technique
[0002] Remote sensing image processing technology has evolved from traditional probability solving to deep learning. With the development of neural networks, current methods centered on convolutional neural networks have made remarkable progress in tasks such as classification, detection, and segmentation. However, traditional symmetric convolutions have limited ability to capture non-linear features, and attention mechanisms have problems such as high computational cost and low utilization rate of global information. Existing remote sensing image processing technologies based on convolutional networks mainly optimize the convolutional network. However, simply stacking convolutional layers, adjusting parameters, and changing channels cannot enable the convolutional network to fully extract complex features and high-frequency information. If too many convolutional layers are stacked, it will also lead to an increase in the number of parameters. Existing attention mechanisms mostly use simple concatenations of single attentions such as channel attention and spatial attention, and cannot effectively extract complex image scenes such as remote sensing images. At the same time, simply introducing attention will also lead to insufficient extraction of information of interest in specific data, affecting the final image restoration effect, so optimization is required.
[0003] Therefore, the present invention proposes an image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network. Summary of the Invention
[0004] The purpose of the present invention is to provide an image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network to solve the problems raised in the above background technique.
[0005] To achieve the above purpose, the present invention provides the following technical solution: An image restoration method based on a cross-asymmetric convolution multi-aggregation self-attention network, including a shallow feature extraction module, a cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module, a decoder module, an upsampling module, and an image reconstruction and restoration module; the cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module includes a cross-asymmetric convolution module and a multi-aggregation self-attention module; The specific steps include: S1: Extract shallow information from a low-resolution remote sensing image to obtain shallow feature information of the image; S2: Input the shallow feature information into the cross-asymmetric convolution module to obtain cross-asymmetric convolution feature information; S3: Perform multi-aggregation self-attention extraction on the cross-asymmetric convolution feature information to obtain multi-aggregation self-attention features, and perform the next cross-asymmetric convolution multi-aggregation self-attention extraction on the features after combining the cross-asymmetric convolution features and the multi-aggregation self-attention features; S4: Repeat steps S2 and S3, stack multiple cross - asymmetric convolution modules and multi - aggregation self - attention modules to obtain the deep features; S5: Input the deep features into the decoder module. After superimposing them with the shallow features, obtain the overall extracted features, and then through the up - sampling module and the image reconstruction and restoration module, obtain the restored high - resolution remote sensing image.
[0006] Preferably, step S1 specifically includes: S11: Construct a remote sensing image dataset, and obtain low - resolution remote sensing images from high - resolution remote sensing images; S12: Perform shallow feature extraction on the low - resolution remote sensing images through a shallow extraction convolution block to obtain a shallow feature image.
[0007] Preferably, the shallow feature image undergoes two cross - asymmetric feature extractions, including: The shallow feature image passes through the first cross - asymmetric feature extraction, with two - path parallel asymmetric convolutions to obtain two sets of dimensional features. For the first - path asymmetric convolution layer and the second - path asymmetric convolution layer, horizontal single - dimensional asymmetric convolution and vertical single - dimensional asymmetric convolution are respectively used for extraction to obtain horizontal single - dimensional features and vertical single - dimensional features; After one extraction, perform cross - connection. Input the single - dimensional features extracted by the horizontal single - dimensional asymmetric convolution into the first path and the second path, and input the single - dimensional features extracted by the vertical single - dimensional asymmetric convolution into the first path and the second path for the second extraction; During the second extraction, the first path and the second path respectively use vertical single - dimensional asymmetric convolution and horizontal single - dimensional asymmetric convolution for extraction to obtain vertical single - dimensional features and horizontal single - dimensional features, and perform the third extraction; After the second extraction, input the vertical single - dimensional features extracted in the second extraction into the first path and the second path, and input the horizontal single - dimensional features extracted in the second extraction into the first path and the second path for the third extraction; During the third extraction, the first path and the second path respectively use horizontal single - dimensional asymmetric convolution and vertical single - dimensional asymmetric convolution for extraction to obtain horizontal single - dimensional features and vertical single - dimensional features, obtain the third - extraction features, and aggregate the two - path features; For the feature image obtained from the first cross - asymmetric feature extraction, perform the second cross - asymmetric feature extraction. The extraction step is through two - path parallel asymmetric convolutions to obtain two sets of dimensional features. For the first - path asymmetric convolution layer and the second - path asymmetric convolution layer, vertical single - dimensional asymmetric convolution and horizontal single - dimensional asymmetric convolution are respectively used for extraction to obtain vertical single - dimensional features and horizontal single - dimensional features; After the first extraction, cross-connection is performed. The one-dimensional features extracted by the vertical one-dimensional asymmetric convolution are input into the first path and the second path, and the one-dimensional features extracted by the horizontal one-dimensional asymmetric convolution are input into the first path and the second path for the second extraction; During the second extraction, the first path and the second path respectively use horizontal one-dimensional asymmetric convolution and vertical one-dimensional asymmetric convolution for extraction to obtain horizontal one-dimensional features and vertical one-dimensional features, and then the third extraction is performed; After the second extraction, the horizontal one-dimensional features extracted in the second extraction are input into the first path and the second path, and the vertical one-dimensional features extracted in the second extraction are input into the first path and the second path for the third extraction; During the third extraction, the first path and the second path respectively use vertical one-dimensional asymmetric convolution and horizontal one-dimensional asymmetric convolution for extraction to obtain vertical one-dimensional features and horizontal one-dimensional features, and the features of the two paths are aggregated to obtain the features extracted in the third extraction.
[0008] Preferably, multi-aggregation self-attention extraction is performed on the cross asymmetric convolution: The multi-aggregation self-attention includes spatial graph attention and channel graph attention. For the input features, they are first set as query Query, key Key, and value Value matrices through linear projection 、 、 . Then, the product of the query and the key-value is calculated to obtain the channel attention map, and after being normalized by Softmax, it is multiplied by the value matrix to obtain the channel attention aggregation result . Then, the input query Query, key Key, and value Value matrices are transposed to obtain the spatial query Query, key Key, and value Value matrices 、 、 . By calculating the product of the query and the key-value, the spatial attention map is obtained, and after being normalized by Softmax, it is multiplied by the value matrix to obtain the spatial attention aggregation result . Finally, an inverse transpose operation is used to output the result and convert it to the original dimension, and an activation function is used to increase its non-linearity to obtain the final multi-aggregation result . Then, the final multi-aggregation input result is given to the output through a residual connection to ensure the stability of its extraction effect.
[0009] Preferably, multiple S2 and S3 are repeatedly spliced, and residual connections are introduced into S2 and S3 to obtain the cross asymmetric convolution multi-aggregation self-attention network.
[0010] Preferably, the network is iteratively trained, and the loss value of each weight parameter modification during the iterative process is calculated using the L1 loss function: Among them, Denote the cross - asymmetric convolution multi - aggregation self - attention network model Denote the hyperparameters corresponding to the network is the remotely sensed image after reconstruction and recovery is related to the corresponding benchmark true remotely sensed image Denote the number of remotely sensed images
[0011] Preferably, for the first cross - asymmetric feature extraction and the second cross - asymmetric feature extraction, residual connections are used to splice the input features to the output, and local aggregation features can be obtained
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows By means of two designed cross - asymmetric convolution blocks, compared with using standard convolution and simple stacking of asymmetric convolutions, two cross - operations can ensure that the information of the upper and lower two paths of the first path and the second path is not lost, and at the same time, it can fully extract the feature components that are not fully extracted by one - dimensional convolution, ensuring that the extracted feature information will not deteriorate
[0013] Through multi - aggregation self - attention extraction, multi - aggregation self - attention features are obtained. Compared with simply using channel attention, spatial attention or a simple combination of the two, by introducing the idea of self - attention, while reducing the computational amount, more - dimensional binary weights can be given to the target image to more accurately obtain the required features
[0014] The present invention also introduces residual connections in different modules to extract local feature information, multi - aggregation information and global feature information, enhance the target receptive field, hide the unnecessary receptive field, and enhance the recovery quality of the image BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is the schematic flow chart of the method of the present invention Figure 2 is the diagram of the first cross - asymmetric convolution block and the second cross - asymmetric convolution block of the present invention Figure 3 is the multi - aggregation self - attention network diagram of the present invention Figure 4 is the schematic principle flow chart of cross - asymmetric convolution multi - aggregation self - attention of the present invention DETAILED DESCRIPTION OF THE INVENTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] Please refer to Figure 1 and Figure 4 , the present invention provides a technical solution: a remote sensing image processing method based on a cross-asymmetric convolution multi-aggregation self-attention network. S1: Extract shallow information from the low-resolution remote sensing image to obtain the shallow feature information of the image. In this embodiment, specifically, first, the input low-resolution image is passed through a and convolution to perform shallow feature extraction and obtain a shallow feature information of the remote sensing image. S2: Input the shallow feature information into the cross-asymmetric convolution module to obtain the cross-asymmetric convolution feature information. In this embodiment, specifically, further extract the shallow feature information through the cross-asymmetric convolution module. For example, Figure 2 , the cross-asymmetric convolution module is divided into a first cross-asymmetric convolution block and a second cross-asymmetric convolution block, including: Copy the shallow feature information and input it into two paths for three extractions. For the first path, use convolution, convolution, convolution for feature extraction. For the second path, use convolution, convolution, convolution for feature extraction. For the first and second feature extractions, cross operations are performed on each path, that is, the feature extracted from the first path is input into the first and second paths of the next extraction, and the feature extracted from the second path is input into the first and second paths of the next extraction. After the third extraction, the features of the first and second paths are aggregated to obtain the first cross-asymmetric convolution feature. Copy the first asymmetric convolution feature information and input it into two paths for three extractions. For the first path, use convolution, convolution, convolution for feature extraction. For the second path, use convolution, convolution, Feature extraction is performed by convolution. For the first and second feature extractions, cross operations are adopted for each path, that is, the features extracted from the first path are passed into the first and second paths of the next extraction, and the features extracted from the second path are passed into the first and second paths of the next extraction. After the third extraction, the features of the first and second paths are aggregated to obtain the second cross-asymmetric convolution features; Meanwhile, residual connections are introduced during the first cross-asymmetric feature extraction and the second cross-asymmetric feature extraction, that is, the corresponding input features are directly passed into the output to obtain local aggregated feature information; S3: Perform multi-aggregation self-attention extraction on the cross-asymmetric convolution feature information to obtain multi-aggregation self-attention features, and perform the next cross-asymmetric convolution multi-aggregation self-attention extraction on the features after combining the cross-asymmetric convolution features with the multi-aggregation self-attention features; S4: Repeat steps S2 and S3, stack multiple cross-asymmetric convolution modules and multi-aggregation self-attention modules to obtain the deep features; Performing multi-aggregation self-attention extraction on the features extracted by the cross-asymmetric convolution module includes: Two multi-aggregation self-attention extractions of spatial graph attention and channel graph attention. For the input features, first set them as query (Query), key (Key), and value (Value) matrices through linear projection , , . Then calculate the product of the query and the key-value to obtain the channel attention map, specifically as follows: , , where, represents the input features, , , respectively represent the linear projection matrices corresponding to the query, key, and value, and after multiplying with the value matrix through Softmax normalization, the channel attention aggregation result , , where, represents 's transposed matrix. Then, for the purpose of spatial attention aggregation, for , , , rotation is performed respectively, , , , where, represents the rotation operation, By calculating the product of the query and the key value, a spatial attention map is obtained, and after Softmax normalization, it is multiplied by the value matrix to obtain the spatial attention aggregation result ,
[0018] Finally, an inverse transpose operation is adopted to convert the output result to the original dimension, and an activation function is used to increase its non-linearity to obtain the final multi-aggregation result , , and then the final multi-aggregation input result is given to the output through a residual connection to ensure the stability of its extraction effect. Through this self-attention aggregation method, pixels in the channel dimension and the spatial dimension can be aggregated, so as to realize the all-direction multi-aggregation self-attention extraction; during the multi-aggregation self-attention extraction process, a residual connection or attention local information is also introduced to make the extracted features more abundant; Stack the cross-asymmetric convolution multi-aggregation self-attention modules multiple times to obtain deep feature information; S5: Transmit the deep features into the decoder module, and after superimposing with the shallow features, the total extracted features are obtained. Then, through the upsampling module and the image reconstruction and restoration module, the restored high-resolution remote sensing image is obtained; In this embodiment, specifically, a convolution with a convolution kernel is used as the decoder for feature decoding. At the same time, a residual connection is adopted to fuse the shallow features and the cross-asymmetric convolution multi-aggregation self-attention deep features to obtain the fusion information with texture structure detail information and image low-frequency information; The fusion information is transmitted into the upsampling module and the image reconstruction and restoration module to obtain the restored high-resolution remote sensing image. The upsampling adopts a sub-pixel convolutional layer, which is a convolution, with an input channel of 64 and an output channel of 256. The sub-pixel convolutional layer can rearrange the feature map with a size of into a feature map of . The reconstruction and restoration output is a convolution with a convolution kernel size of , and its function is to convert the output 64 channels into 3 channels.
[0019] Except for the input module, the final upsampling module and the image reconstruction and restoration module, other modules are all performed in the channel dimension of 64.
[0020] Although the embodiments of the present invention have been shown and described, as detailed in the above description, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An image restoration method based on a cross-asymmetric convolutional multi-aggregation self-attention network, characterized by: It includes a shallow feature extraction module, a cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module, a decoder module, an upsampling module and an image reconstruction and restoration module; the cross-asymmetric convolution multi-aggregation self-attention network deep feature extraction module includes a cross-asymmetric convolution module and a multi-aggregation self-attention module; The specific steps include: S1: extracting shallow information from low-resolution remote sensing images to obtain shallow feature information of the images; S2: passing the shallow feature information into the cross asymmetric convolution module to obtain cross asymmetric convolution feature information; S3: Perform multi-aggregate self-attention extraction on the cross asymmetric convolution feature information to obtain multi-aggregate self-attention features, combine the cross asymmetric convolution features with the features after the multi-aggregate self-attention features, and perform the next cross asymmetric convolution multi-aggregate self-attention extraction; S4: repeat steps S2 and S3, stack multiple cross asymmetric convolution modules and multi-aggregation self-attention modules to obtain the deep features; S5: The deep features are passed into the decoder module, and after being superimposed with the shallow features, the total extracted features are obtained, and then the restored high-resolution remote sensing image is obtained through the upsampling module and the image reconstruction and restoration module.
2. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 1, characterized in that: Step S1 specifically includes: S11: Construct remote sensing image datasets and obtain low-resolution remote sensing images from high-resolution remote sensing images; S12: Performing shallow feature extraction on the low-resolution remote sensing image through a shallow extraction convolution block to obtain a shallow feature image.
3. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 2 is characterized in that: The shallow feature image is subjected to two cross asymmetric feature extractions, including: the shallow feature image is subjected to a first cross asymmetric feature extraction, two parallel asymmetric convolutions, to obtain two sets of dimensional features, and for the first asymmetric convolution layer and the second asymmetric convolution layer, horizontal single-dimensional asymmetric convolutions and vertical single-dimensional asymmetric convolutions are respectively used to extract, to obtain horizontal single-dimensional features and vertical single-dimensional features; After one extraction, a cross connection is performed, and the single-dimensional features extracted by the horizontal single-dimensional asymmetric convolution are passed into the first path and the second path, and the single-dimensional features extracted by the vertical single-dimensional asymmetric convolution are passed into the first path and the second path for a second extraction; During the second extraction, the first path and the second path respectively use vertical single-dimensional asymmetric convolution and horizontal single-dimensional asymmetric convolution to extract, obtain vertical single-dimensional features and horizontal single-dimensional features, and perform a third extraction; After the second extraction, the vertical single-dimensional features extracted for the second time are passed into the first and second paths, and the horizontal single-dimensional features extracted for the second time are passed into the first and second paths for the third extraction; During the third extraction, the first path and the second path are extracted by horizontal single-dimensional asymmetric convolution and vertical single-dimensional asymmetric convolution respectively to obtain horizontal single-dimensional features and vertical single-dimensional features, obtain the third extraction features, and aggregate the two-path features; The feature image obtained by the first cross asymmetric feature extraction is subjected to the second cross asymmetric feature extraction. The extraction step is to obtain two sets of dimensional features by subjecting the third extraction features of the first cross asymmetric to two parallel asymmetric convolutions. For the first asymmetric convolution layer and the second asymmetric convolution layer, vertical single-dimensional asymmetric convolution and horizontal single-dimensional asymmetric convolution are respectively used to extract, so as to obtain vertical single-dimensional features and horizontal single-dimensional features. After one extraction, a cross connection is performed, and the single-dimensional features extracted by the vertical single-dimensional asymmetric convolution are passed into the first path and the second path, and the single-dimensional features extracted by the horizontal single-dimensional asymmetric convolution are passed into the first path and the second path for a second extraction; During the second extraction, the first path and the second path respectively use horizontal single-dimensional asymmetric convolution and vertical single-dimensional asymmetric convolution to extract, obtain horizontal single-dimensional features and vertical single-dimensional features, and perform the third extraction; After the second extraction, the horizontal single-dimensional features extracted for the second time are passed into the first path and the second path, and the vertical single-dimensional features extracted for the second time are passed into the first path and the second path for the third extraction; During the third extraction, the first path and the second path respectively use vertical single-dimensional asymmetric convolution and horizontal single-dimensional asymmetric convolution to extract, to obtain vertical single-dimensional features and horizontal single-dimensional features, and the two-path features are aggregated to obtain the third extraction features.
4. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 1, characterized in that: Perform multi-aggregation self-attention extraction on the cross asymmetric convolution: The multi-aggregate self-attention includes spatial graph attention and channel graph attention. For the input features, they are first set as query, key, and value matrices through linear projection. , , , then calculate the product of the query and the key value to get the channel attention map, and then multiply it with the value matrix after Softmax normalization to get the channel attention aggregation result , then transpose the input query, key, and value matrices to obtain the spatial query, key, and value matrices , , , by calculating the product of the query and the key value, we get the spatial attention map, and then multiply it with the value matrix after Softmax normalization to get the spatial attention aggregation result Finally, the inverse permutation operation is used to convert the output result into the original dimension, and the activation function is used to increase its nonlinearity to obtain the final multi-aggregation result. , and then the final multi-aggregation input result is given to the output through residual connection to ensure the stability of its extraction effect.
5. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 3, characterized in that: Repeatedly splice multiple S2 and S3, and introduce residual connections in S2 and S3 to obtain the cross asymmetric convolutional multi-aggregation self-attention network.
6. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 1, characterized in that: The network is iteratively trained, and the loss value of each weight parameter modification during the iteration process is calculated using the L1 loss function: in, represents the cross asymmetric convolutional multi-aggregation self-attention network model, represents the hyperparameters corresponding to the network, is the reconstructed and restored remote sensing image. is with The corresponding benchmark real remote sensing image, Represents the number of remote sensing images.
7. The image restoration method based on cross asymmetric convolutional multi-aggregation self-attention network according to claim 5, characterized in that: For the first cross asymmetric feature extraction and the second cross asymmetric feature extraction, residual connection is used to splice the input features to the output, wherein the second cross asymmetric feature extraction takes the aggregated features of the first cross asymmetric feature extraction as input to obtain local aggregated features.
Citation Information
Patent Citations
Two-way cross-connected convolutional neural network for image segmentation
CN113192089A
Face super-resolution method and device based on cross convolution attention adversarial learning
CN114757832A
Food material image classification model establishment method based on attention and depth feature fusion
CN114898360A
Depth map super-resolution reconstruction method and system based on asymmetric cross attention
CN116402692A
Image processing method, apparatus including image processing model, image processing apparatus, device, storage medium, and program product
CN119579994A
Cited By
Knowledge editing method based on multi-language semantic retrieval
CN120892551A
Ultrahigh-definition image restoration method, device and equipment based on clustering center feature scanning
CN121391673A