Adaptive Attention-Based Multi-Scale Super-Resolution Reconstruction Method for Agricultural Remote Sensing Images
By introducing adaptive attention and multi-scale processing technologies in agricultural remote sensing image processing, the problems of poor generalization and loss of details in agricultural remote sensing image processing are solved, and higher quality image super-resolution reconstruction is achieved.
Patent Information
- Application Number
- CN202411449749.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-10-17
AI Technical Summary
When processing agricultural remote sensing images, existing deep learning models have poor generalization, which are prone to reconstruction distortion and details loss, and fail to fully consider the characteristics and needs of agricultural remote sensing images.
A multi-scale agricultural remote sensing image super-resolution reconstruction method is proposed based on adaptive attention. The super-resolution reconstruction quality of the image is improved through shallow feature extraction, multi-scale adaptive attention processing, aggregation processing and upsampling.
Through multi-scale convolution and adaptive attention mechanisms, we can improve the diversity of image feature scales and reconstruction quality, optimize the application of agricultural scenarios, and improve the generalization ability of model.
Smart Images

Figure CN119205512B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a multi-scale agricultural remote sensing image super-resolution reconstruction method based on adaptive attention. Background Art
[0002] Agricultural remote sensing technology is an important part of the development of modern agriculture. Remote sensing images obtained by satellites or drones can effectively monitor various information such as the growth status of crops, the distribution of pests and diseases, and soil humidity, thereby realizing fine agricultural management. However, the resolution of remote sensing images directly affects their application effects. High-resolution images can provide clearer and more detailed surface information, which helps to more accurately analyze various characteristics of agricultural plots. However, due to the technical limitations of remote sensing sensors and the influence of imaging conditions, directly obtaining high-resolution images faces high technical and cost challenges.
[0003] Currently, super-resolution technology has been widely studied and applied in improving the resolution of remote sensing images. Traditional super-resolution methods mainly include interpolation methods and reconstruction-based methods. Interpolation methods such as bilinear interpolation and bicubic interpolation, although simple to calculate, often cannot effectively improve the details and quality of images. Reconstruction-based methods reconstruct high-resolution images by fusing multiple low-resolution images. Although they can improve the image quality to a certain extent, they have high requirements for the consistency of the image sequence and a large computational complexity. In recent years, deep learning technology has shown great potential in super-resolution image reconstruction. The convolutional neural network CNN can capture the high-frequency details of images through large-scale data training, thereby generating high-quality high-resolution images. For example, models such as SRCNN, VDSR, and ESRGAN have shown excellent performance on various datasets. However, when these deep learning models process agricultural remote sensing images, there are still some significant deficiencies:
[0004] Poor model generalization: Existing deep learning models perform inconsistently on different types of agricultural remote sensing images. Especially in complex agricultural scenarios, problems such as reconstruction distortion and detail loss are likely to occur.
[0005] Lack of agricultural scenario optimization: Most super-resolution models fail to fully consider the characteristics and requirements of agricultural remote sensing images and cannot effectively improve the image quality in specific agricultural application scenarios. Summary of the Invention
[0006] The object of the present invention is to propose a multi-scale agricultural remote sensing image super-resolution reconstruction method based on adaptive attention to solve the problems existing in the above-mentioned prior art. By using a deep learning model and a variety of advanced image processing techniques, it aims to improve the quality of low-resolution agricultural remote sensing images, realize high-resolution image reconstruction, and meet the needs of high-resolution remote sensing images in the fields of agricultural monitoring, precision agriculture management, etc.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A multi-scale agricultural remote sensing image super-resolution reconstruction method based on adaptive attention, comprising:
[0009] S1. Extract shallow features from a low-resolution RGB agricultural remote sensing image to obtain shallow features of the low-resolution image;
[0010] S2. Perform multi-scale adaptive attention processing on the shallow features to obtain adaptive attention multi-scale features;
[0011] S3. Aggregate the adaptive attention multi-scale features to obtain global aggregated features. After adding the global aggregated features to the adaptive attention multi-scale features, perform deep feature extraction to obtain current deep features;
[0012] S4. Repeat S2-S3 for the current deep features until the target number of times to obtain the final deep features;
[0013] S5. After adding the final deep features to the shallow features, perform upsampling to obtain a super-resolution agricultural remote sensing image.
[0014] Optionally, extracting shallow features from a low-resolution RGB agricultural remote sensing image includes:
[0015] Extract a low-resolution RGB agricultural remote sensing image from the agricultural remote sensing image;
[0016] Perform shallow feature extraction on the low-resolution RGB agricultural remote sensing image.
[0017] Optionally, performing multi-scale adaptive attention processing on the shallow features includes:
[0018] The shallow features are first subjected to convolution processing and then replicated, and then respectively pass through a first masked convolutional layer and a second masked convolutional layer to obtain two sets of scale features;
[0019] Add and fuse the two sets of scale features and then perform global average pooling to obtain channel information features;
[0020] Fuse the channel information features, then copy and input them into a fully connected layer respectively, and perform a softmax operation in the channel dimension to obtain two channel attention weights;
[0021] Multiply the two channel attention weights by the corresponding two sets of scale features respectively, and then perform feature addition to obtain adaptive attention features;
[0022] Add the adaptive attention features and the shallow features to obtain adaptive attention multi-scale features.
[0023] Optionally, the aggregation processing of the adaptive attention multi-scale features includes:
[0024] Perform local aggregation processing on the adaptive attention multi-scale features to obtain local aggregation features;
[0025] Perform global aggregation processing on the local aggregation features to obtain global aggregation features.
[0026] Optionally, the local aggregation processing of the adaptive attention multi-scale features includes:
[0027] Perform spatial dimension processing on the adaptive attention multi-scale features to obtain spatial attention features;
[0028] Perform feature information suppression processing on the spatial attention features to obtain X GDFN features;
[0029] For the X GDFN features, perform channel self-attention processing to obtain channel self-attention features;
[0030] Add the channel self-attention features to the X GDFN features to obtain the local aggregation features.
[0031] Optionally, obtaining the spatial attention features includes:
[0032] Divide the adaptive attention multi-scale features into non-overlapping local blocks of size P×P;
[0033] Perform dimensional transformation on the non-overlapping local blocks in the spatial dimension to obtain feature X;
[0034] Perform dimensional transformation on the feature X, and convert the feature X into a query Q through linear projection S , key K S and value matrix V S ;
[0035] For the query Q S , key K S and value matrix VS Perform multi-head attention processing to obtain
[0036] Calculate the product of the query and the key , and perform softmax processing to obtain the spatial attention weight;
[0037] Multiply the spatial attention weight by the value matrix to obtain a preset feature;
[0038] Perform dimensional transformation on the preset feature, and perform fully connected processing. Add the feature after fully connected processing to the feature X to obtain the spatial attention feature.
[0039] Optionally, obtain the X GDFN The feature includes:
[0040] Perform dimensional transformation on the spatial attention feature, denoted as X GDFNIN ;
[0041] After X GDFNIN undergoes dimensional transformation, it passes through the first per-channel convolution and is evenly split into two features X GDFN1 , X GDFN2 ;
[0042] After X GDFN1 passes through the GELU activation function and is multiplied by X GDFN2 , and then convolution processing is performed, and added to X GDFNIN to obtain the X GDFN feature.
[0043] Optionally, obtaining the channel self-attention feature includes:
[0044] After performing per-channel convolution on the X GDFN feature, it is evenly divided to obtain the query Q C , the key K C , and the value matrix V C ;
[0045] After multiplying the Q C and K C matrices, calculate softmax and then multiply by V C to obtain a preset dimensional feature;
[0046] After transforming the preset dimensional feature, perform a convolution operation to obtain the channel self-attention feature.
[0047] Optionally, obtaining the preset deep feature includes:
[0048] After accumulating the global aggregation feature and the adaptive attention multi-scale feature, feature X is obtained ESAIN ;
[0049] Compress the feature X ESAIN to obtain feature X ESA1 ;
[0050] Perform receptive field expansion processing on the feature X ESA1 to obtain feature X ESA2 ;
[0051] Perform pooling processing on the feature X ESA2 and then perform convolution processing to obtain the spatial dimension correlation feature;
[0052] Use bilinear interpolation to perform dimension restoration on the spatial dimension correlation feature to obtain feature X ESA3 ;
[0053] Perform the first convolution processing on the feature X ESA1 to obtain feature X ESA4 ;
[0054] Add the feature X ESA3 and the feature X ESA4 , perform the second convolution processing, restore the number of channels, and generate an attention mask through the sigmoid activation function;
[0055] Perform dot product operation on the attention mask and X ESAIN to generate a preset deep feature with long-range dependence.
[0056] The beneficial effects of the present invention are as follows:
[0057] Due to the limitations of remote sensing image acquisition hardware devices such as drones, the resolution of the obtained remote sensing images is low, resulting in the loss of image details, blurred features, and difficult information acquisition, which limits the accurate identification and monitoring of fine features or small-scale objects in the agricultural application field. The method proposed by the present invention enhances the diversity of image feature scales through multi-scale convolution and improves the quality of image super-resolution reconstruction. Through neural architecture search, the optimal network structure is searched in a preset space to obtain adaptive attention, which can autonomously select the convolution kernel size according to the characteristics of agricultural image data, achieve a more reasonable receptive field, optimize the application of agricultural scenarios, and improve the generalization ability of the model. Brief Description of the Drawings
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0059] Figure 1 It is a diagrammatic illustration of the input feature block method for the local aggregation module and the global aggregation module in the embodiments of the present invention;
[0060] Figure 2 It is a flowchart of the multi-scale agricultural remote sensing image super-resolution reconstruction method based on adaptive attention in the embodiments of the present invention;
[0061] Figure 3 It is an architecture diagram of the multi-scale agricultural remote sensing image super-resolution reconstruction method based on adaptive attention in the embodiments of the present invention;
[0062] Figure 4 It is the adaptive attention multi-scale network module in the architecture diagram of the embodiments of the present invention;
[0063] Figure 5 It is the masked adaptive convolution module in the architecture diagram of the embodiments of the present invention. Detailed implementation manners
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0065] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0066] As Figure 2 - Figure 3 shown, this embodiment proposes a multi-scale agricultural remote sensing image super-resolution reconstruction method based on adaptive attention, including:
[0067] S1. Perform shallow feature extraction on the low-resolution RGB agricultural remote sensing image to obtain the shallow features of the low-resolution image;
[0068] Specifically, in this embodiment, the low-resolution agricultural remote sensing image is mapped to the RGB space, and the shallow features of the low-resolution image are extracted through the shallow feature extraction network;
[0069] S2. Perform multi-scale adaptive attention processing on the shallow features to obtain adaptive attention multi-scale features;
[0070] Specifically, in this embodiment, the features of the previous module are fed into the adaptive attention multi-scale network module to obtain adaptive attention multi-scale features;
[0071] S3. Perform aggregation processing on the adaptive attention multi-scale features to obtain global aggregation features. After adding the global aggregation features and the adaptive attention multi-scale features, perform deep feature extraction to obtain the preset deep features;
[0072] Specifically, in this embodiment, the adaptive attention features are fed into the local aggregation module to obtain local aggregation features. The local aggregation features are fed into the global aggregation module and added to the input of S2, and then fed into the ESA module to obtain deeper features;
[0073] S4. Repeat the processing of S2 - S3 several times for the preset deep features to obtain the final deep features;
[0074] S5. After adding the final deep features and the shallow features, perform upsampling processing to obtain the super-resolution agricultural remote sensing image;
[0075] Specifically, in this embodiment, the deep features are added to the shallow features in S1, and then fed into the PixelShuffle module for upsampling to obtain the final super-resolution agricultural remote sensing image.
[0076] Further, the shallow feature extraction of the low-resolution RGB agricultural remote sensing image includes:
[0077] Extract the low-resolution RGB agricultural remote sensing image from the agricultural remote sensing image; use a convolution with a convolution kernel size of 3×3 to perform shallow feature extraction on the low-resolution RGB agricultural remote sensing image.
[0078] Specifically, in this embodiment, the low-resolution RGB agricultural remote sensing image is fed into the shallow feature extraction module. In this example, the dimension of the low-resolution agricultural remote sensing image is 1×3×64×64. After passing through the shallow feature extraction network which includes a 3×3 convolution with a stride of 1 and a padding value of 1. The low-resolution RGB image is fed into the shallow feature extraction network to obtain the shallow features Y of the low-resolution image shallow , and the dimension is 1×64×64×64 at this time.
[0079] Further, the multi-scale adaptive attention processing of the shallow features includes:
[0080] The shallow features first go through a 1×1 convolution and then are copied. After that, they respectively go through a 5×5 masked convolution and a 7×7 masked convolution to obtain two sets of scale features;
[0081] The two sets of scale features are added and fused, and then global average pooling is performed to obtain channel information features;
[0082] Channel information fusion is performed on the channel information features, and then they are copied and respectively input into a fully connected layer, and a softmax operation is performed in the channel dimension to obtain two channel attention weights;
[0083] The two channel attention weights are respectively multiplied by the corresponding two sets of scale features, and then the features are added to obtain adaptive attention features;
[0084] The adaptive attention features and the shallow features are added to obtain adaptive attention multi-scale features.
[0085] Specifically, in this embodiment, the input feature is sent into the adaptive attention multi-scale network module, and the input feature is the shallow feature Y in step (1) shallow or the deep feature in step (4), as Figure 4 shown. The adaptive attention multi-scale network includes two convolutions with a kernel size of 1×1, a stride of 1, and a padding value of 0, a 5×5 masked convolution with a kernel size of 1×1, a stride of 1, and a padding value of 2, a 7×7 masked convolution with a kernel size of 1×1, a stride of 1, and a padding value of 3. Two GELU activation layers, two ReLU activation layers. A fully connected layer with an input channel of 64 and an output channel of 16. Two fully connected layers with an input channel of 16 and an output channel of 64. The input feature first goes through a 1×1 convolution and then is copied. After that, it respectively goes through a 5×5 masked adaptive convolution and a 7×7 masked adaptive convolution to obtain two sets of scale features and The 5×5 masked convolution is as Figure 5 shown. A convolution kernel of size 5×5 can be regarded as a 5×5 matrix, and a convolution kernel of size 3×3 can be regarded as a 3×3 matrix. It can be seen that there is a certain connection between the 5×5 convolution kernel and the 3×3 convolution, that is, a 5×5 convolution kernel can be regarded as a 3×3 convolution kernel wrapped with a circle of parameters. The formula is as follows:
[0086] w 5×5 =w 3×3 +w 5× 5\3×3 (1)
[0087] where w 5×5 represents a 5×5 convolution kernel, w 3×3 represents a 3×3 convolution kernel, w5×5\3×3 Represents the set of parameters in the outermost wrapping. After transforming the above formula once, we get the following formula:
[0088] w k = w 3×3 + αw 5×5\3×3 (2)
[0089] where w k is an adaptive convolutional kernel, and α is a neural architecture search operator. If α = 0, then this convolutional kernel can be regarded as a 3×3 convolution. If α = 1, then this convolutional kernel can be regarded as a 5×5 convolution. Similarly, a 7×7 convolutional kernel can be regarded as a 5×5 convolutional kernel with a set of parameters in the outermost wrapping. The formula for the neural architecture search operator α is as follows:
[0090] α = sigmoid(δt) (3)
[0091] where t is a trainable parameter, similar to a weight, obtained by training with a neural network. δ is a parameter that changes with the epoch (when all images in a complete dataset have been trained once through the neural network, this process is called one epoch). When δ → ∞, it can be seen from the image that the value of α can only be 0 or 1. δ → ∞ is certainly impossible to achieve in reality, but it can be considered that when δ takes a relatively large number, δ can be approximately regarded as infinity. The formula for δ changing with the epoch is as follows:
[0092]
[0093] where δ start is the initial value of δ, δ fin is the final value of δ, epoch_now represents the current epoch being executed during the training of the model, and num_epochs represents the total number of epochs required for training. By training with a neural network to determine whether α should take 0 or 1, adaptive convolution can be achieved. In this example, δ start takes 1, and δ fin takes 10000.
[0094] Sum the features and element-wise to obtain the feature U, then perform global average pooling to obtain the channel information feature S of 1×64×1×1. Through a fully connected layer with an input channel of 64 and an output channel of 16 for channel information fusion, the feature Z with a dimension of 1×16×1×1 is obtained. After replication, it passes through a fully connected layer with an input channel of 16 and an output channel of 64 respectively, and a softmax operation is performed on the channel dimension to obtain two channel attention weights a and b. a and b are respectively multiplied by the corresponding two sets of scale features After multiplication, the features are added to obtain the adaptive attention feature V. The adaptive attention feature and the input feature of this step are added to obtain the adaptive attention multi-scale feature V multi , whose dimension is 1×64×64×64.
[0095] Furthermore, the aggregation process of the adaptive attention multi-scale feature includes:
[0096] Perform local aggregation processing on the adaptive attention multi-scale feature to obtain the local aggregation feature;
[0097] Perform global aggregation processing on the local aggregation feature to obtain the global aggregation feature.
[0098] The local aggregation processing of the adaptive attention multi-scale feature includes:
[0099] Perform spatial dimension processing on the adaptive attention multi-scale feature to obtain the spatial attention feature;
[0100] Perform feature information suppression processing on the spatial attention feature to obtain X GDFN feature;
[0101] For X GDFN feature, perform channel self-attention processing to obtain the channel self-attention feature;
[0102] Add the channel self-attention feature to X GDFN feature to obtain the local aggregation feature.
[0103] Obtaining the spatial attention feature includes:
[0104] Divide the adaptive attention multi-scale feature into non-overlapping local blocks of size P×P;
[0105] Perform dimensional transformation on the non-overlapping local blocks in the spatial dimension to obtain feature X;
[0106] Perform dimensional transformation on feature X and convert feature X into query Q through linear projection S , key K S and value matrix V S ;
[0107] For query Q S , key K S and value matrix V S perform multi-head attention processing to obtain
[0108] Calculate the product of the query and the key and perform softmax processing to obtain the spatial attention weight;
[0109] Multiply the spatial attention weight with the value matrix to obtain a preset feature;
[0110] Perform dimensional transformation on the preset feature, and perform a fully connected process. Add the feature after the fully connected process to the feature X to obtain the spatial attention feature.
[0111] Obtain X GDFN The feature includes:
[0112] Perform dimensional transformation on the spatial attention feature, denoted as X GDFNIN ;
[0113] For X GDFNIN After performing dimensional transformation through a 1×1 convolution, perform a 3×3 per-channel convolution and evenly split it into two features X GDFN1 、X GDFN2 ;
[0114] For X GDFN1 Multiply it with the GELU activation function and X GDFN2 and then pass through a 1×1 convolution, and add it to X GDFNIN to obtain the X GDFN feature.
[0115] Obtaining the channel self-attention feature includes:
[0116] For X GDFN Send the feature into a 1×1 convolution, a 3×3 per-channel convolution in sequence, and then evenly divide it to obtain Q C 、K C 、V C ;
[0117] For Q C 、K C After matrix multiplication, calculate softmax and then multiply it with V C to obtain the preset dimensional feature;
[0118] After transforming the preset dimensional feature, send it into a 1×1 convolution to obtain the channel self-attention feature.
[0119] Specifically, in this embodiment, the obtained adaptive attention multi-scale feature is sent into the local aggregation module. This module first divides the feature into non-overlapping local blocks of size 8×8. The dimensions of these blocks are aggregated into the spatial dimension. The input feature is transformed from dimension 1×64×64×64 to 1×8×8×8×8×64 to obtain the feature X. Then it is transformed into 64×64×64 and the obtained feature is converted into query, key, and value matrices Q S 、K S 、V S, including 1 fully-connected layer, with 64 input channels and 192 output channels. Q S 、K S 、V S all have dimensions of 64×64×64. They are divided into 4 multi-head attention features to obtain with dimensions of 64×4×64×16. Calculate the product of the query and the key , and perform softmax to obtain the spatial attention weights with dimensions of 64×4×64×64. Then multiply by to obtain features with dimensions of 64×4×64×16, and then perform dimensional transformation to obtain features with dimensions of 64×8×8×64. After passing through a fully-connected layer with 64 input channels and 64 output channels. Then change the dimensions to 1×8×8×8×8×64, and finally add to the feature X to obtain the spatial attention feature. Change the dimensions of the spatial attention feature to 1×64×64×64, denoted as X GDFNIN . Feed it into the GDFN module to suppress features with less information and only allow useful information to pass through. This module consists of two convolutions with a kernel size of 1×1, a stride of 1, a padding value of 0, and a convolution with a kernel size of 3×3, a stride of 1, and a padding value of 1. The input feature first passes through a 1×1 convolution, and the dimensions become 1×128×64×64. After passing through a 3×3 depthwise convolution and splitting it into two on average, two features X GDFN1 、X GDFN2 are obtained. Multiply X GDFN1 after passing through the GELU activation function by X GDFN2 and then pass through a 1×1 convolution, and add it to X GDFNIN to obtain the GDFN feature X GDFN , with dimensions of 1×64×64×64. Subsequently, feed the X GDFN feature into the channel self-attention module. This module consists of two convolutions with a kernel size of 1×1, a stride of 1, a padding value of 0, and a convolution with a kernel size of 3×3, a stride of 1, and a padding value of 1. First, feed the X GDFN feature into a 1×1 convolution to obtain a feature with dimensions of 1×192×64×64. Then feed it into a 3×3 depthwise convolution to obtain a feature of 1×192×64×64. Divide it evenly to obtain Q C 、K C 、V C , with dimensions of 1×64×4×64×16. After multiplying the Q C and K C matrices and calculating softmax, then multiply by V CMultiply to obtain a dimension of 1×64×4×16×64. Then perform dimension conversion to get 1×64×64×64, and send it into a 1×1 convolution to obtain channel self-attention features with a dimension of 1×64×64×64. Add the channel self-attention features and X GDFN features to obtain local aggregation features.
[0120] The calculation process of the global aggregation features is similar to that of the local aggregation features. The only difference lies in the input feature chunking method. The local aggregation feature calculation method is continuous chunking, and the global aggregation feature calculation method is interval chunking, as Figure 1 shown. This module first divides the features into non-overlapping global blocks of size 8×8. The dimensions of these blocks are aggregated into the spatial dimension. The input features are converted from the dimension of 1×64×64×64 to 1×8×8×8×8×64 to obtain feature X'. Then it is converted to 64×64×64 and the obtained features are converted into query, key, and value matrices Q' S 、K' S 、V' S , including 1 fully connected layer with an input channel of 64 and an output channel of 192. The dimensions of Q' S 、K' S 、V' S are all 64×64×64. Divide them into 4 multi-head attention features to obtain features with a dimension of 64×4×64×16. Calculate the product of the query and the key , and perform softmax to obtain spatial attention weights with a dimension of 64×4×64×64. Then multiply with to obtain features with a dimension of 64×4×64×16, and then perform dimension conversion to obtain features with a dimension of 64×8×8×64. After passing through a fully connected layer with an input channel of 64 and an output channel of 64. Then change the dimension to 1×8×8×8×8×64, and finally add it to feature X' to obtain spatial attention features. Change the dimension of the spatial attention features to 1×64×64×64, denoted as X' GDFNIN . Send it into the GDFN module to suppress features with less information and only allow useful information to pass through. This module consists of two convolutions with a kernel size of 1×1, a stride of 1, a padding value of 0, and a convolution with a kernel size of 3×3, a stride of 1, and a padding value of 1. The input features first pass through a 1×1 convolution, and the dimension becomes 1×128×64×64. After passing through a 3×3 per-channel convolution and splitting it into two on average, two features X' GDFN1 、X' GDFN2 are obtained. Multiply X' GDFN1 after passing through the GELU activation function with X' GDFN2 and then pass through a 1×1 convolution, and add it to X'GDFNIN Add them to obtain the GDFN feature X'. GDFN , whose dimension is 1×64×64×64. Subsequently, X' GDFN The feature is fed into the channel self-attention module. This module consists of two convolutions with a kernel size of 1×1, a stride of 1, a padding value of 0, and a convolution with a kernel size of 3×3, a stride of 1, and a padding value of 1. First, the X' GDFN The feature is fed into a 1×1 convolution to obtain a feature with a dimension of 1×192×64×64. Then it is fed into a 3×3 per-channel convolution to obtain a feature of 1×192×64×64. It is evenly divided to obtain Q' C , K' C , V' C , all with a dimension of 1×64×4×64×16. After multiplying the Q' C , K' C matrices and calculating softmax, then multiplying with V' C , a feature with a dimension of 1×64×4×16×64 is obtained. Then, after dimension transformation to 1×64×64×64, it is fed into a 1×1 convolution to obtain the channel self-attention feature, whose dimension is 1×64×64×64. Add the local aggregation feature and the X' GDFN feature to obtain the global aggregation feature.
[0121] Furthermore, obtaining the preset deep features includes:
[0122] After adding the global aggregation feature and the adaptive attention multi-scale feature, obtain the feature X ESAIN ;
[0123] Use a 1×1 convolution to compress the feature X ESAIN to obtain the feature X ESA1 ;
[0124] Use a 3×3 convolution with a stride of 2 to perform receptive field expansion processing on the feature X ESA1 to obtain the feature X ESA2 ;
[0125] Use a 7×7 window and a max pooling layer with a stride of 3 to perform pooling processing on the feature X ESA2 , and then use a 3×3 convolution with a stride of 1 to process it to obtain the spatial dimension correlation feature;
[0126] Use bilinear interpolation to restore the dimension of the spatial dimension correlation feature to obtain the feature X ESA3 ;
[0127] Feed the feature X ESA1 into a 1×1 convolution to obtain the feature X ESA4 ;
[0128] Add feature X ESA3 and feature X ESA4 together and feed them into a 1×1 convolution. After restoring the number of channels, generate an attention mask through the sigmoid activation function;
[0129] Perform a dot product operation on the attention mask and X ESAIN to generate a preset deep feature with long-range dependencies.
[0130] Specifically, in this embodiment, the global aggregation feature and the input in S2 are accumulated to form feature X ESAIN and fed into the ESA module. The ESA module consists of three 1×1 convolutions with a stride of 1 and a padding value of 0, one 3×3 convolution with a stride of 1 and a padding value of 1, and one 3×3 convolution with a stride of 2 and a padding value of 0. First, use a 1×1 convolution to compress the feature from the dimension of 1×64×64×64 to 1×16×64×64, obtaining feature X ESA1 . Subsequently, use a 3×3 convolution with a stride of 2 to expand the receptive field, obtaining feature X ESA2 , with a dimension of 1×16×31×31. Perform pooling using a 7×7 window and a max pooling layer with a stride of 3, and then use a 3×3 convolution with a stride of 1 to obtain a spatial dimension correlation feature with a dimension of 1×16×8×8. Use bilinear interpolation to restore the feature dimension to 1×16×64×64, obtaining feature X ESA3 . Feed X ESA1 into a 1×1 convolution to obtain feature X ESA4 . Feed X ESA3 and X ESA4 together and feed them into a 1×1 convolution to restore the number of channels, obtaining a feature with a dimension of 1×64×64×64. Generate an attention mask (mask) through the sigmoid activation function and perform a dot product operation with X ESAIN to generate a deeper feature with long-range dependencies.
[0131] S5. After accumulating the final deep feature and the shallow feature, perform upsampling to obtain a super-resolution agricultural remote sensing image; including:
[0132] Feed the final deep feature into a 3×3 convolution with a stride of 1 and a padding value of 1, then add it to S1 and feed it into the PixelShuffle module for upsampling to obtain a super-resolution image with a specified magnification factor. For example, if the magnification factor is 2, the output image scale is 1×3×128×128. If the magnification factor is 4, the output image scale is 1×3×256×256, and so on.
[0133] The multi-scale agricultural remote sensing image super-resolution reconstruction method based on adaptive attention proposed in this embodiment utilizes the idea of automatically designing a neural network architecture, and automatically explores and optimizes the structure of the neural network through an algorithm to find the best remote sensing image super-resolution reconstruction network architecture for agricultural applications. For the method proposed by the present invention, first, the low-resolution agricultural remote sensing image is sent into a shallow feature extraction network to extract the shallow features of the image. Secondly, through the deep feature extraction module. The deep feature extraction module consists of three parts: an adaptive attention multi-scale network module, a local aggregation module, and a global aggregation module. The adaptive attention multi-scale network module is divided into an adaptive convolution module and an attention module. Through the adaptive convolution module, the algorithm automatically searches and optimizes the structure of the neural network to extract different scale features that are optimal for agriculture. The attention module strengthens the important scale features and suppresses the unimportant scale features. Adaptive attention multi-scale features are generated through the adaptive attention multi-scale network module. Subsequently, local aggregation features are obtained through the local aggregation module, and global aggregation features are obtained through the global aggregation module. The global aggregation features are added to the input features of the deep feature extraction module to obtain deeper features. The deep feature extraction module is repeated n times to obtain deep features. Finally, after the deep features and the shallow features are added and convolved, the PixelShuffle module is used for upsampling to generate a high-resolution agricultural remote sensing image. Compared with the traditional remote sensing image reconstruction method, through the adaptive attention multi-scale network module in S2, the neural network automatically selects the convolution kernel size with the help of a search operator, and the network structure is automatically designed. It is possible to automatically design the neural network architecture and optimize the network structure according to the characteristics of specific agricultural remote sensing image application scenarios. The present invention can also utilize the different information extracted from different receptive fields of agricultural remote sensing images, strengthen the important receptive fields, suppress the unimportant receptive fields, and improve the reconstruction quality.
[0134] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A multi-scale agricultural remote sensing image super-resolution reconstruction method based on adaptive attention, characterized in that: include: S1. Extract shallow features from low-resolution RGB agricultural remote sensing images to obtain shallow features of low-resolution images; S2. performing multi-scale adaptive attention processing on the shallow features to obtain adaptive attention multi-scale features; S3. Aggregate the adaptive attention multi-scale features to obtain global aggregate features, add the global aggregate features to the adaptive attention multi-scale features, perform deep feature extraction, and obtain current deep features; Aggregating the adaptive attention multi-scale features includes: Performing local aggregation processing on the adaptive attention multi-scale features to obtain local aggregation features; Performing global aggregation processing on the local aggregation features to obtain global aggregation features; The local aggregation processing of the adaptive attention multi-scale features includes: Performing spatial dimension processing on the adaptive attention multi-scale feature to obtain a spatial attention feature; The spatial attention feature is subjected to feature information suppression processing to obtain X GDFN feature; For the X GDFN The features are processed by channel self-attention to obtain channel self-attention features; The channel self-attention feature is combined with X GDFN Adding the features to obtain the local aggregated features; Acquiring the spatial attention feature includes: Divide the adaptive attention multi-scale features into non-overlapping local blocks of size P×P; The non-overlapping local blocks are transformed in spatial dimension to obtain feature X; Transform the dimension of the feature X and transform the feature X into a query Q through linear projection S , key K S Sum value matrix V S ; For the query Q S , key K S Sum value matrix V S Perform multi-head attention processing to obtain Calculation query and key The product of and softmax processing is performed to obtain the spatial attention weight; The spatial attention weights and value matrix Multiply to obtain preset features; Performing dimension conversion on the preset feature and performing full connection processing, adding the feature after the full connection processing to the feature X to obtain the spatial attention feature; S4. Repeat S2-S3 for the current deep feature until the target number of times to obtain the final deep feature; S5. After accumulating the final deep features and the shallow features, upsampling is performed to obtain a super-resolution agricultural remote sensing image.
2. The method for super-resolution reconstruction of multi-scale agricultural remote sensing images based on adaptive attention according to claim 1, characterized in that: Shallow feature extraction of low-resolution RGB agricultural remote sensing images includes: Extract low-resolution RGB agricultural remote sensing images from agricultural remote sensing images; Shallow feature extraction is performed on the low-resolution RGB agricultural remote sensing image.
3. The method for super-resolution reconstruction of multi-scale agricultural remote sensing images based on adaptive attention according to claim 1, characterized in that: Performing multi-scale adaptive attention processing on the shallow features includes: The shallow features are first convolved and then copied and passed through a first masked convolution layer and a second masked convolution layer respectively to obtain two sets of scale features; The two sets of scale features are added and fused, and then global average pooling is performed to obtain channel information features; Perform channel information fusion on the channel information features, copy them and input them into a fully connected layer respectively, and perform a softmax operation on the channel dimension to obtain the attention weights of two channels; The two channel attention weights are multiplied by the corresponding two sets of scale features, and then the features are added to obtain the adaptive attention features; The adaptive attention feature and the shallow feature are added to obtain an adaptive attention multi-scale feature.
4. The method for super-resolution reconstruction of multi-scale agricultural remote sensing images based on adaptive attention according to claim 1, characterized in that: Get the X GDFN Features include: The spatial attention feature is transformed into a dimension, denoted as X GDFNIN ; X GDFNIN After the dimension transformation, it is convolved channel by channel and evenly split into two features X GDFN1 , X GDFN2 ; X GDFN1 After GELU activation function and X GDFN2 Multiply and then convolve, and X GDFNIN Add together to get the X GDFN feature.
5. The method for super-resolution reconstruction of multi-scale agricultural remote sensing images based on adaptive attention according to claim 1, characterized in that: Obtaining the channel self-attention feature includes: For the X GDFN After the features are convolved channel by channel, they are evenly divided to obtain the query Q C , key K C , value matrix V C ; Q C , K C After matrix multiplication, calculate softmax and then add V C Multiply to obtain the preset dimension features; After converting the preset dimensional features, a convolution operation is performed to obtain channel self-attention features.
6. The method for super-resolution reconstruction of multi-scale agricultural remote sensing images based on adaptive attention according to claim 1, characterized in that: Obtaining the current deep features includes: After accumulating the global aggregate feature and the adaptive attention multi-scale feature, the feature X is obtained. ESAIN ; For the feature X ESAIN Compress and get feature X ESA1 ; For the feature X ESA1 Expand the receptive field and obtain feature X ESA2 ; For the feature X ESA2 Perform pooling processing and then convolution processing to obtain spatial dimension correlation features; Use bilinear interpolation to restore the dimension of the spatial dimension correlation feature to obtain feature X ESA3 ; The feature X ESA1 Perform the first convolution process to obtain feature X ESA4 ; The feature X ESA3 and the feature X ESA4 Add them together, perform the second convolution, restore the number of channels, and generate the attention mask through the sigmoid activation function; Combine the attention mask with X ESAIN Dot product operation generates current deep features with longer distance dependencies.
Citation Information
Patent Citations
Super-resolution method and device based on low-spatial-resolution remote sensing image
CN115409698A