Remote sensing image reconstruction method and device based on multi-scale window space channel attention

Through the remote sensing image reconstruction method based on the attention of multi-scale window space channel, the problem of insufficient details in the remote sensing image reconstruction process in the prior art is solved, high-quality image reconstruction is achieved, and the authenticity and reconstruction performance of the image are improved.

CN120013761APending Publication Date: 2025-05-16XIDIAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510087187.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is prone to artifacts, jagging and blurring in the process of remote sensing image reconstruction, especially in the recovery of high-frequency details, which leads to the reconstructed remote sensing image not rich enough and too smooth, which limits its application in many fields.

Method used

The remote sensing image reconstruction method based on the attention of multi-scale window space channel is adopted. The trained reconstruction model is used to process the low-resolution remote sensing image, extract and fuse multi-scale feature information to realize the reconstruction of high-resolution images.

Benefits of technology

This method can enrich the texture details of the reconstructed remote sensing image, making the reconstructed image more realistic and natural, and at the same time achieve optimal performance in peak signal-to-noise ratio and structural similarity, improving the reconstruction quality of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013761A_ABST
    Figure CN120013761A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image reconstruction method and device based on multi-scale window space channel attention, and relates to the technical field of image processing, and the method comprises the steps: obtaining a to-be-reconstructed low-resolution remote sensing image; processing a to-be-reconstructed low-resolution remote sensing image by adopting the trained reconstruction model to obtain a reconstructed high-resolution remote sensing image; wherein the trained reconstruction model is obtained by training an initial reconstruction model by taking data of a preset category as a training data set and aiming at extracting and fusing multi-scale feature information. The texture detail information of the reconstructed remote sensing image can be enriched, and the remote sensing image reconstruction quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a remote sensing image reconstruction method and device based on multi-scale window spatial channel attention. Background Art

[0002] With the continuous development of remote sensing technology, the number of remote sensing platforms such as drones and satellites has continued to expand, and there are more ways to obtain remote sensing images. However, the transmission of a large number of high-resolution images by platforms such as drones and satellites requires a large amount of transmission bandwidth resources. Therefore, platforms such as satellites may first convert high-resolution remote sensing images into low-resolution remote sensing images for transmission, and ground platforms need to reconstruct them through image super-resolution methods.

[0003] However, in the process of reconstructing images based on traditional interpolation and reconstruction methods, artifacts, aliasing and blurring are prone to occur, especially in the recovery of high-frequency details. In recent years, remote sensing image reconstruction methods based on convolutional neural networks and Transformers have greatly improved the above phenomena. However, remote sensing images have very complex spatial distribution and large cross-scale characteristics. The texture details of remote sensing images reconstructed by these methods are not rich enough and are too smooth, which limits the application of remote sensing images in urban planning, environmental monitoring, homeland security, land surveying, emergency rescue, vegetation monitoring, resource exploration, meteorological analysis and other fields. Therefore, there is an urgent need for a method to enrich the texture detail information of reconstructed remote sensing images. Summary of the invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a remote sensing image reconstruction method and device based on multi-scale window spatial channel attention. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0005] In a first aspect, the present invention provides a remote sensing image reconstruction method based on multi-scale window spatial channel attention, comprising:

[0006] Acquire low-resolution remote sensing images to be reconstructed;

[0007] The trained reconstruction model is used to process the low-resolution remote sensing image to be reconstructed to obtain a reconstructed high-resolution remote sensing image;

[0008] The trained reconstruction model uses data of preset categories as training data sets, and trains the initial reconstruction model for the purpose of extracting and fusing multi-scale feature information.

[0009] In a second aspect, the present invention also provides a remote sensing image reconstruction device based on multi-scale window spatial channel attention, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus;

[0010] Memory, used to store computer programs;

[0011] The processor is used to implement the above method when executing the program stored in the memory.

[0012] Beneficial effects of the present invention:

[0013] The present invention provides a remote sensing image reconstruction method and device based on multi-scale window spatial channel attention, which can enrich the texture detail information of low-resolution remote sensing images after reconstruction, making the reconstructed images more realistic and natural, while achieving optimal performance in peak signal-to-noise ratio and structural similarity, with higher precision.

[0014] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flow chart of a remote sensing image reconstruction method based on multi-scale window spatial channel attention provided by an embodiment of the present invention;

[0016] Figure 2 is a schematic diagram of a trained reconstruction model provided by an embodiment of the present invention;

[0017] Figure 3 is a schematic diagram of a trained multi-scale window spatial channel attention group module provided by an embodiment of the present invention;

[0018] Figure 4 is a schematic diagram of a trained multi-scale feature extraction module provided by an embodiment of the present invention;

[0019] Figure 5 is a schematic diagram of a trained window space channel attention module provided by an embodiment of the present invention;

[0020] Figure 6 is a schematic diagram of a window spatial attention module provided by an embodiment of the present invention;

[0021] Figure 7 is a schematic diagram of a channel attention module provided by an embodiment of the present invention;

[0022] Figure 8 It is a schematic diagram of a visualization result at a WHU-RS19 × 4 upsampling rate provided by an embodiment of the present invention;

[0023] Fig. 9 It is a schematic diagram of the visualization result under the RSSCN7×3 upsampling ratio provided by an embodiment of the present invention;

[0024] Fig.10 It is a schematic diagram of the visualization result under the RSSCN7×2 upsampling ratio provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0026] See also Figure 1 , Figure 1 1 is a flow chart of a remote sensing image reconstruction method based on multi-scale window spatial channel attention provided by an embodiment of the present invention. The remote sensing image reconstruction method based on multi-scale window spatial channel attention provided by the present invention comprises:

[0027] S101. Obtain a low-resolution remote sensing image to be reconstructed.

[0028] Specifically, in this embodiment, a remote sensing image is acquired, and a bicubic interpolation down-sampling method is used to process the remote sensing image into a low-resolution remote sensing image to be reconstructed.

[0029] S102, using the trained reconstruction model to process the low-resolution remote sensing image to be reconstructed to obtain a reconstructed high-resolution remote sensing image;

[0030] The trained reconstruction model uses data of preset categories as training data sets, and trains the initial reconstruction model for the purpose of extracting and fusing multi-scale feature information.

[0031] Specifically, in this embodiment, see Figure 2 , Figure 2 is a schematic diagram of a trained reconstruction model provided by an embodiment of the present invention, wherein the trained reconstruction model includes a trained first convolutional layer, a plurality of cascaded trained multi-scale window spatial channel attention group modules, a trained feature fusion layer, and a trained upsampling module; the trained reconstruction model is used to process a low-resolution remote sensing image to be reconstructed to obtain a high-resolution remote sensing image, including:

[0032] The trained first convolutional layer is used to extract features from the low-resolution remote sensing image to be reconstructed to obtain shallow features;

[0033] Using a cascade of multiple trained multi-scale window spatial channel attention group modules to process the shallow features to obtain multiple features;

[0034] Use the trained feature fusion layer to fuse multiple features and shallow features to obtain deep features;

[0035] The trained upsampling module is used to process the deep features to obtain the reconstructed high-resolution remote sensing image.

[0036] In this embodiment, see Figure 3 , Figure 3 : is a schematic diagram of a trained multi-scale window space channel attention group module provided by an embodiment of the present invention, wherein the trained multi-scale window space channel attention group module comprises a trained multi-scale feature extraction module, a trained second convolutional layer, a trained normalization layer, a cascade of multiple trained window space channel attention modules and a trained third convolutional layer; the shallow features are processed by using the cascade of multiple trained multi-scale window space channel attention group modules to obtain multiple features, including:

[0037] For the first-level trained multi-scale window spatial channel attention group module, the trained multi-scale feature extraction module is used to process the shallow features to obtain multi-scale features, and the multi-scale features are added to the shallow features to obtain the first fusion features;

[0038] The trained second convolutional layer is used to perform a convolution operation on the first fusion feature to obtain the first convolutional feature;

[0039] Use the trained normalization layer to normalize the first convolution feature to obtain the normalized feature;

[0040] Using a cascade of multiple trained window space channel attention modules to process the normalized features to obtain aggregated features;

[0041] The aggregated features are processed using the trained third convolution layer to obtain the second convolution feature, and the second convolution feature is added to the first fusion feature to obtain the first feature;

[0042] For the i-th level trained multi-scale window spatial channel attention group module, the output of the i-1-th level trained multi-scale window spatial channel attention group module is processed as input to obtain the i-th feature; for all cascaded trained multi-scale window spatial channel attention group modules, multiple features are obtained.

[0043] In this embodiment, six groups of trained multi-scale window spatial channel attention group modules are provided.

[0044] In this embodiment, see Figure 4 , Figure 4: is a schematic diagram of a trained multi-scale feature extraction module provided by an embodiment of the present invention, wherein the trained multi-scale feature extraction module includes multiple feature extraction paths, a first splicing layer and a convolutional layer 1 arranged in parallel; each feature extraction path includes a first module and a second module, the first module includes a convolutional layer 2 and an activation function 1, and the second module includes a second splicing layer, a convolutional layer 3 and an activation function 2, wherein in the same feature extraction path, the convolutional layer 1 and the convolutional layer 2 have the same size; the shallow features are processed using the trained multi-scale feature extraction module to obtain multi-scale features, including:

[0045] The first modules in different feature extraction paths are used to perform convolution and nonlinear activation function operations on shallow features in sequence to obtain first features of different scales;

[0046] The second modules in different feature extraction paths are used to perform concatenation, convolution and nonlinear activation function operations on all first features of different scales in sequence, so as to obtain second features of different scales respectively;

[0047] The first splicing layer is used to splice the second features of different scales to obtain spliced ​​multi-scale features;

[0048] A pair of convolutional layers is used to perform convolution operation on the concatenated multi-scale features to obtain multi-scale features.

[0049] In this embodiment, see Figure 5 , Figure 5 It is a schematic diagram of a trained window space channel attention module provided by an embodiment of the present invention, wherein the trained window space channel attention module includes a third module and a fourth module, wherein the third module includes a window space attention module, a first depth separation convolution module, a convolution layer four, an activation function three, a global average pooling layer one, a convolution layer five and an activation function four, and the fourth module includes a channel attention module, a second depth separation convolution module, a global average pooling layer two, a convolution layer six, an activation function five, a convolution layer seven and an activation function six; a plurality of the trained window space channel attention modules are cascaded to process the normalized features to obtain aggregated features, including:

[0050] For the first-level trained window space channel attention module, the normalized features are dimensionally changed, and the window space attention module is used to process the dimensionally changed normalized features to obtain the window space attention features;

[0051] Convolution layer 4 is used to perform convolution operation on the window spatial attention feature, and activation function 3 is used to perform nonlinear operation on the result processed by convolution layer 4 to obtain the first spatial weight matrix;

[0052] Using a first depth separation convolution module to process the normalized features to obtain a first depth separation feature;

[0053] A global average pooling layer 1 is used to perform a pooling operation on the first depth separation feature, a convolution layer 5 is used to perform a convolution operation on the result processed by the global average pooling layer 1, and an activation function 4 is used to perform a nonlinear operation on the result processed by the convolution layer 5 to obtain a first channel weight matrix;

[0054] Multiply the first channel weight matrix by the window spatial attention feature to obtain feature one; multiply the first spatial weight matrix by the first depth separation feature to obtain feature two; add feature one and feature two to obtain the window spatial attention enhanced feature;

[0055] The feature dimension of the window space attention enhancement is changed, and the channel attention module is used to process the window space attention enhancement feature after the dimension change to obtain the channel attention feature;

[0056] The global average pooling layer 2 is used to perform a pooling operation on the channel attention features, the convolution layer 6 is used to perform a convolution operation on the result processed by the global average pooling layer 2, and the activation function 5 is used to perform a nonlinear operation on the result processed by the convolution layer 6 to obtain the second channel weight matrix;

[0057] The second depth separation convolution module is used to process the features of the window space attention enhancement to obtain the second depth separation features;

[0058] The convolution layer 7 is used to perform a convolution operation on the second depth separation feature, and the activation function 6 is used to perform a nonlinear operation on the result processed by the convolution layer 7 to obtain a second spatial weight matrix;

[0059] Multiply the channel attention feature by the second spatial weight matrix to obtain feature three; multiply the second depth separation feature by the second channel weight matrix to obtain feature four; add feature three and feature four to obtain the first aggregate feature;

[0060] For the j-th level trained window space channel attention module, the output of the j-1-th level trained window space channel attention module is processed as input to obtain the j-th aggregated feature; for the output of the last level window space channel attention module, the aggregated feature is obtained.

[0061] In this embodiment, see Figure 6 , Figure 6 is a schematic diagram of a window spatial attention module provided by an embodiment of the present invention, the window spatial attention module includes activation function seven and a third splicing layer; the window spatial attention module is used to process the normalized features after the dimension change to obtain the window spatial attention features, including:

[0062] Obtain the query matrix, key matrix and value matrix through linear mapping, and change the dimensions of the query matrix, key matrix and value matrix; divide the query matrix, key matrix and value matrix after the dimension change into multiple parts evenly on the channel dimension;

[0063] For each query matrix, key matrix and value matrix after the dimension change, a first weight matrix is ​​calculated according to the query matrix and the key matrix;

[0064] A preset relative position encoding method is used to obtain a position encoding, and the position encoding is added to the first weight matrix to obtain a second weight matrix; if the current window spatial attention module is in an even position, the preset mask is added to the second weight matrix as an updated second weight matrix;

[0065] The second weight matrix is ​​normalized by using activation function seven to obtain a third weight matrix;

[0066] Multiply the third weight matrix by the value matrix to obtain the feature calculated by the window spatial attention, and transform the dimension of the feature to obtain a window spatial attention feature;

[0067] For the query matrix, key matrix and value matrix after all dimensions have changed, multiple window spatial attention features are obtained, and the third splicing layer is used to splice the multiple window spatial attention features to obtain the window spatial attention features.

[0068] In this embodiment, see Figure 7 , Figure 7 is a schematic diagram of a channel attention module provided by an embodiment of the present invention, the channel attention module includes an activation function eight and a fourth splicing layer; the channel attention module is used to process the features of the window space attention enhancement after the dimension change to obtain the channel attention features, including:

[0069] Obtain a query matrix, a key matrix, and a value matrix through linear mapping, and change the dimensions of the query matrix, the key matrix, and the value matrix; divide the query matrix, the key matrix, and the value matrix after the dimension change into multiple equal parts on the channel dimension;

[0070] For each query matrix, key matrix and value matrix after the dimension change, a first weight matrix is ​​calculated according to the query matrix and the key matrix;

[0071] The activation function eight is used to normalize the first weight matrix to obtain a second weight matrix;

[0072] Multiply the second weight matrix by the value matrix to obtain the feature calculated by the channel attention, and transform the dimension of the feature to obtain a channel attention feature;

[0073] For the query matrix, key matrix and value matrix after all dimensions have changed, multiple channel attention features are obtained, and the fourth splicing layer is used to splice the multiple channel attention features to obtain the channel attention features.

[0074] In this embodiment, the process of obtaining the trained reconstruction model includes:

[0075] Acquire data of multiple preset categories, pre-process the data of the preset categories as samples in a training data set, and obtain labels of the samples in the training data set;

[0076] The samples in the training data set are input into the reconstruction model in batches for training until the number of training times or the degree of convergence meets the preset conditions, and a trained reconstruction model is obtained.

[0077] Specifically, the trained reconstruction model is obtained through the following process, including:

[0078] S201. Generate low-resolution images from the existing remote sensing datasets AID, WHU-RS19 and RSSCN7 using bicubic interpolation downsampling. The images sampled from the remote sensing dataset AID are used as samples in the training dataset, with a total of 10,000 images, each with a resolution of 600×600. The images sampled from the remote sensing datasets WHU-RS19 and RSSCN7 are used as samples in the verification dataset. 1,005 images are sampled from the remote sensing dataset WHU-RS19, with a resolution of 600×600 for each image, and 2,800 images are sampled from the remote sensing dataset RSSCN7, with a resolution of 400×400 for each image.

[0079] It should be noted that the data in the remote sensing datasets AID, WHU-RS19 and RSSCN7 are high-resolution images.

[0080] S202, randomly crop a 48×48 image block I from the low-resolution image LR , and use flipping 90°, 180°, and 270° for data enhancement to expand the training data; at the same time, crop the high-resolution image block I at the corresponding magnification from the high-resolution image HR As training labels; normalize the pixels of the low-resolution image block and its corresponding high-resolution image block to [0, 1] to obtain a low-resolution image block I with a size of 48×48×3 LR and a high-resolution image patch I of size (48×r)×(48×r)×3 HR , r is the magnification.

[0081] S203, input the low-resolution image block into the convolution with a convolution kernel size of 3×3, a step size of 1, and an output dimension of 64 to extract shallow features F 0 , F 0 The dimensions are 48×48×64.

[0082] S204, the shallow feature F 0 Multiple features are obtained by inputting the cascaded multiple multi-scale window spatial channel attention group (MSWSCAG) modules, where the MSWSCAG module structure is as follows: Figure 3 shown.

[0083] S2041. For the first-level trained multi-scale window spatial channel attention group module, the trained multi-scale feature processing process includes:

[0084] a. The shallow feature F 0 Different features are extracted by convolution with kernel sizes of 3×3, 5×5 and 7×7, step sizes of 1, padding of 1, 2 and 3 respectively, and ReLU nonlinear activation function to obtain the first features of three scales respectively;

[0085] b. The first features of the three scales obtained in step a are concatenated in the channel dimension through different Contact operations to obtain concatenated features with a size of 48×48×192;

[0086] c. The concatenated features obtained in step b are extracted through convolution and ReLU nonlinear activation functions with kernel sizes of 3×3, 5×5 and 7×7, step sizes of 1, padding of 1, 2 and 3, respectively, to obtain the second features of three scales, with feature sizes of 48×48×64;

[0087] d. The second features of the three scales obtained in step c are fused by convolution with a convolution kernel size of 3×3 to obtain multi-scale features. At the same time, the multi-scale features are added to the shallow features to obtain the first fused features with a size of 48×48×64.

[0088] S2042. The first fusion feature is upgraded in the channel dimension through a convolution with a convolution kernel size of 3×3, a step size of 1, and a padding of 1 to obtain a first convolution feature.

[0089] S2043. Normalize the first convolutional feature through the LayerNorm layer to obtain a normalized feature with a size of 48×48×180.

[0090] S2044, the normalized features are passed through the window spatial attention (WSA) module and the depth separation convolution in parallel, and the parallel structure of WSA and depth separation convolution is as follows Figure 5As shown in the left part, the WSA module is as follows Figure 6 shown.

[0091] a. Reshape the normalized features to 2304×180, then obtain the query matrix query, key matrix key and value matrix value through linear mapping, and then reshape the query, key and value to 16×144×180 feature maps;

[0092] b. In order to use the multi-head attention mechanism within each window (window size is 6×24), the query matrix query, key matrix key and value matrix value are evenly divided into 6 parts in the channel dimension, that is, there are 6 heads in total, and the size of each query matrix query, key matrix key and value matrix value is 16×144×30;

[0093] c. Perform matrix multiplication on query and key, and then divide by Where d = C / h = 30, the first weight matrix is ​​obtained, and the relative position encoding method proposed in Swin Transformer is used to obtain the position encoding; then the position encoding is added to the first weight matrix to obtain the second weight matrix. If the current number of WSA blocks is an even number (for example, the 2nd, 4th, and 6th WSA modules in the cascade process), then the second weight matrix needs to be added to the mask in the current WSA module, which means that the window attention calculation is performed after the window slides to capture the global context information. The obtained matrix is ​​still called the second weight matrix; then the secondary weight matrix is ​​normalized by softmax to obtain the third weight matrix; finally, the third weight matrix is ​​matrix-multiplied with value to obtain the feature map after the window space attention calculation, and then the Reshape operation is used to resize it to 48×48×30 window space attention features;

[0094] d. Through the Contact operation, the six feature maps obtained using the multi-head attention mechanism are spliced ​​in the channel dimension to obtain the window space attention feature with a size of 48×48×180;

[0095] f. The normalized features are passed through a depth separation convolution with a convolution kernel size of 3×3, a step size of 1, a padding of 1, and a grouping of 180. Then, the BatchNorm normalization function is passed, and finally, the Relu nonlinear activation function is used to obtain the first depth separation feature with a size of 48×48×180.

[0096] g. The window spatial attention feature uses convolution with a kernel size of 1×1 and the Sigmoid function to extract the spatial weight matrix to obtain the first spatial weight matrix; the first depth separation feature uses global average pooling, convolution with a kernel size of 1×1 and the Sigmoid function to extract the channel weight matrix to obtain the first channel weight matrix; then the window spatial attention feature is element-wise multiplied with the first channel weight matrix, and the first spatial weight matrix is ​​element-wise multiplied with the first spatial weight matrix; finally, the results of the above two element multiplications are added to the matrix elements to obtain the feature map after the enhanced window spatial attention module, that is, the window spatial attention enhanced feature, whose size is 48×48×180. The purpose of doing this is to adaptively re-weight the two branch features in the spatial dimension and channel dimension to achieve the aggregation of global and local spatial channel features within the block;

[0097] h. The features enhanced by the window spatial attention are reshaped into features of size 2034×180 through the Reshape operation, and then the query matrix query, key matrix key and value matrix value are obtained through linear mapping;

[0098] i. In order to use the multi-head attention mechanism, the query matrix query, key matrix key and value matrix value are evenly divided into 6 parts in the channel dimension, that is, there are 6 heads in total, and the size of each query matrix query, key matrix key and value matrix value is 2034×30;

[0099] j. Perform matrix multiplication on query and key, and then divide by α (α is a learnable parameter) to adjust the inner product to obtain the first weight matrix; then perform softmax normalization on the first weight matrix to obtain the second weight matrix; finally, perform matrix multiplication on the second weight matrix and value to obtain the feature map after channel attention calculation, and then use the Reshape operation to resize it to 48×48×30;

[0100] k. Through the Contact operation, the six feature maps obtained using the multi-head attention mechanism are concatenated in the channel dimension to obtain the channel attention feature with a size of 48×48×180;

[0101] l. The window spatial attention enhanced features are passed through the depth separation convolution with a convolution kernel size of 3×3, a step size of 1, a padding of 1, and a grouping of 180. Then, the BatchNorm normalization function is passed, and finally the Relu nonlinear activation function is used to obtain the second depth separation feature with a size of 48×48×180.

[0102] m. The second depth separation feature uses convolution with a kernel size of 1×1 and a Sigmoid function to extract the second spatial weight matrix; the channel attention feature uses global average pooling, convolution with a kernel size of 1×1 and a Sigmoid function to extract the second channel weight matrix; then the second depth separation feature is element-wise multiplied with the second channel weight matrix, and the channel attention feature is element-wise multiplied with the second spatial weight matrix; finally, the results of the above two element multiplications are added together to obtain the feature map after the channel attention module, that is, the first aggregated feature, whose size is 48×48×180. The purpose of doing this is to adaptively re-weight the two branch features in the spatial dimension and the channel dimension to achieve the aggregation of global and local spatial channel features within the block, and after alternating cascading of the WSA module and the CA module, the aggregation of spatial channel features is enhanced between blocks;

[0103] S2045. The output of the last-level window spatial channel attention module is used as the aggregated feature. The aggregated feature is reduced in dimension through a convolution with a convolution kernel size of 3×3, a step size of 1, and a padding of 1, keeping the input and output of the feature map size of the MSWSCAG module consistent, and at the same time, the matrix elements are added to the first fusion feature to obtain the i-th feature.

[0104] S205, the present invention sets 6 window space channel attention modules, and obtains 6 features, namely F 1 、F 2 、F 3 、F 4 、F 5 and F 6 .

[0105] S206, multiple features F 0 、F 1 、F 2 、F 3 、F 4 、F 5 and F 6 Add the elements, and then perform multi-level feature fusion through convolution with a convolution kernel size of 1×1 and a step size of 1, while removing redundant information to obtain the final deep feature with a size of 48×48×64;

[0106] The fusion expression is:

[0107]

[0108] Among them, F represents the deep features, It represents a convolution operation with a kernel size of 1×1 and a stride of 1, and ⊕ represents the addition of matrix elements.

[0109] S207, extract the feature F' of the deep feature F through the convolution with a kernel size of 3×3, a step size of 1, a padding of 1 and the LeakyRelu non-linear activation function, and the size is 48×48×64.

[0110] S208, use sub-pixel convolution to upsample the feature F'. Sub-pixel convolution is essentially to use a convolution with a kernel size of 3×3, a step size of 1, and a padding of 1 to increase the channel dimension, that is, to change the image size from 48×48×64 to 48×48×(64×r 2 ), where r is the magnification factor, and then PixeShuffle is used to reduce the feature size from 48×48×(64×r 2 ) becomes (48×r)×(48×r)×64, which means r-fold upsampling is achieved, and its feature is named "F".

[0111] S209, the feature F" is reduced from 64 dimensions to 3 dimensions through a convolution with a convolution kernel size of 3×3, a step size of 1, and a padding of 1. The feature size is (48×r)×(48×r)×3, which is the final reconstructed remote sensing image.

[0112] S2010, the remote sensing image I SR With the corresponding high-resolution image I HR For comparison, the loss is calculated through the loss function;

[0113] The expression of the loss function is:

[0114]

[0115] Among them, G(·) is the model function, θ is the parameter in the model, is the input LR image feature, is the target HR image feature, Yes Input Generated SR image features

[0116] S2011. Derivative the loss function, use the derivative as the gradient of the output layer, and then back-propagate along the upsampling module, deep feature extraction module, and shallow feature extraction module in order to obtain the weight gradient value of each convolutional layer and fully connected layer; use the gradient descent method to update the weight gradient of each convolutional layer and fully connected layer, so as to continuously reduce the value of the loss function until the loss function remains basically unchanged, that is, the model converges, and finally save the network parameters of each layer, that is, save the entire model.

[0117] S2012. Select model parameters corresponding to the magnification (saved during the training process), automatically load the parameters of each layer of the network during the model loading process, input a low-resolution image, and the network automatically generates a high-resolution image corresponding to the magnification.

[0118] In order to measure the quality of reconstructed remote sensing images from an objective perspective, the peak signal-to-noise ratio (PNSR) and result similarity (SSIM) are selected as measurement indicators. The higher the PNSR and SSIM, the better the image quality. The peak signal-to-noise ratio (PNSR) is calculated at the pixel level, and the calculation formula is as follows:

[0119]

[0120] in, is the pixel value at (i, j) of the SR image generated by the model, is the pixel value of the HR image at (i, j), H is the height of the image, W is the width of the image, L is the maximum pixel value, and L is 255 for an 8-bit image;

[0121] Structural similarity (SSIM) starts from the perspective of image structural information, brightness and contrast. It uses the image mean as an estimate of brightness, the standard deviation as an estimate of contrast, and the covariance as a measure of structural similarity, reflecting the properties of the object structure.

[0122] Given two images x and y, the SSIM calculation formula for the two is as follows:

[0123] C 1 =(k 1 l) 2 ;

[0124] C 2 =(k 2 l) 2 ;

[0125]

[0126] Among them, μ x is the mean of x, μ y is the mean of y, is the variance of x, is the variance of y, σ xy is the covariance of x and y, C 1 and C 2 is a coefficient used to maintain stability, l is the dynamic range of pixel values, and k 1 is 0.01, k 2 It is 0.03, where the range of SSIM is [0,1]. The closer the SSIM value is to 1, the more similar the structures of the two images are, and the better the reconstructed image effect is.

[0127] See also Figure 8 to Figure 10 , respectively, are some low-resolution remote sensing image reconstruction results of the present invention on the WHU-RS19 and RSSCN7 datasets. Figure 8 As shown in the figure, the desert texture structure in the reconstructed area is relatively delicate. Although most networks can reconstruct the approximate structure, the image reconstructed by this method is more natural at the desert concave and convex lines than that reconstructed by other networks, and is more similar to the HR image. Fig. 9 As shown in , the reconstructed area restores the general structure of the HR image. Compared with other networks, the reconstructed image of this method is more similar to the HR image, and it achieves the best peak signal-to-noise ratio and structural similarity. Fig.10 As shown in the figure, the structure of the reconstructed area is very complex. Although most networks can reconstruct the corresponding SR image well, the white lines between the small solar panels in this method are clearer than those reconstructed by other networks, and the surface visual perception is more realistic.

[0128] In summary, the remote sensing image reconstruction method based on multi-scale window spatial channel attention provided by the present invention has the following beneficial effects:

[0129] 1. The multi-scale window channel attention module designed by the present invention can realize feature aggregation between modules and within modules, which helps the network pay more attention to important feature information and enhances the model's ability to extract and fuse multi-scale feature information.

[0130] 2. Considering that features at different depths in the network are helpful for reconstructing images, the present invention designs a multi-level feature fusion module to aggregate feature information at different levels in the network.

[0131] 3. The present invention was tested on the WHU-RS19 and RSSCN7 datasets. The experimental results show that it not only enriches the texture detail information of the reconstructed remote sensing image, but also achieves the best performance in peak signal-to-noise ratio (PNSR) and structural similarity (SSIM), thereby improving the image reconstruction quality.

[0132] Based on the same inventive concept, the present invention also provides a remote sensing image reconstruction device based on multi-scale window spatial channel attention, which is used to implement the remote sensing image reconstruction method based on multi-scale window spatial channel attention provided by the above embodiment of the present invention. The embodiment of the method is referred to above and will not be repeated here. The device includes a processor, a communication interface, a memory and a communication bus. The processor, the communication interface and the memory communicate with each other through the communication bus;

[0133] Memory, used to store computer programs;

[0134] The processor is used to implement the method provided in the above embodiment when executing the program stored in the memory.

[0135] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the term "include", "comprise" or any other variant is intended to cover non-exclusive inclusion, so that the article or device including a series of elements includes not only those elements, but also other elements that are not explicitly listed. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the article or device including the elements. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The orientation or position relationship indicated by "up", "down", "left", "right", etc. is based on the orientation or position relationship shown in the accompanying drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.

[0136] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.

[0137] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.

Claims

1. A remote sensing image reconstruction method based on multi-scale window spatial channel attention, characterized in that: include: Acquire low-resolution remote sensing images to be reconstructed; Using the trained reconstruction model to process the low-resolution remote sensing image to be reconstructed to obtain a reconstructed high-resolution remote sensing image; The trained reconstruction model uses data of preset categories as training data sets, and trains the initial reconstruction model for the purpose of extracting and fusing multi-scale feature information.

2. The remote sensing image reconstruction method based on multi-scale window spatial channel attention according to claim 1, characterized in that: The trained reconstruction model includes a trained first convolutional layer, a cascade of multiple trained multi-scale window spatial channel attention group modules, a trained feature fusion layer and a trained upsampling module; The method of using the trained reconstruction model to process the low-resolution remote sensing image to be reconstructed to obtain a high-resolution remote sensing image includes: Using the trained first convolutional layer to extract features from the low-resolution remote sensing image to be reconstructed, to obtain shallow features; Using a cascade of the trained multi-scale window spatial channel attention group modules to process the shallow features to obtain multiple features; Using the trained feature fusion layer to fuse the multiple features and the shallow features to obtain deep features; The trained upsampling module is used to process the deep features to obtain the reconstructed high-resolution remote sensing image.

3. The remote sensing image reconstruction method based on multi-scale window spatial channel attention according to claim 2, characterized in that: The trained multi-scale window space channel attention group module includes a trained multi-scale feature extraction module, a trained second convolutional layer, a trained normalization layer, a cascade of multiple trained window space channel attention modules and a trained third convolutional layer; The shallow features are processed by using a cascade of multiple trained multi-scale window space channel attention group modules to obtain multiple features, including: For the first-level trained multi-scale window spatial channel attention group module, the trained multi-scale feature extraction module is used to process the shallow features to obtain multi-scale features, and the multi-scale features are added to the shallow features to obtain a first fusion feature; Using the trained second convolutional layer to perform a convolution operation on the first fusion feature to obtain a first convolutional feature; Using the trained normalization layer to perform a normalization operation on the first convolutional feature to obtain a normalized feature; Using a cascade of the plurality of trained window space channel attention modules to process the normalized features to obtain aggregated features; Using the trained third convolution layer to process the aggregated feature to obtain a second convolution feature, and adding the second convolution feature to the first fusion feature to obtain a first feature; For the i-th level trained multi-scale window spatial channel attention group module, the output of the i-1-th level trained multi-scale window spatial channel attention group module is processed as input to obtain the i-th feature; for all cascaded trained multi-scale window spatial channel attention group modules, multiple features are obtained.

4. The remote sensing image reconstruction method based on multi-scale window spatial channel attention according to claim 2, characterized in that: The trained multi-scale window spatial channel attention group modules are provided with 6 groups.

5. The remote sensing image reconstruction method based on multi-scale window spatial channel attention according to claim 3, characterized in that: The trained multi-scale feature extraction module includes a plurality of feature extraction paths, a first splicing layer and a convolution layer 1 arranged in parallel; each of the feature extraction paths includes a first module and a second module, the first module includes a convolution layer 2 and an activation function 1, and the second module includes a second splicing layer, a convolution layer 3 and an activation function 2, wherein in the same feature extraction path, the convolution layer 1 and the convolution layer 2 have the same size; the shallow features are processed using the trained multi-scale feature extraction module to obtain multi-scale features, including: Using different first modules in the feature extraction paths to perform convolution and nonlinear activation function operations on the shallow features in sequence, respectively, to obtain first features of different scales; Using different second modules in the feature extraction paths to sequentially perform concatenation, convolution, and nonlinear activation function operations on all first features of different scales, respectively, to obtain second features of different scales; Using the first splicing layer to splice second features of different scales to obtain spliced ​​multi-scale features; The convolution layer 1 is used to perform a convolution operation on the concatenated multi-scale features to obtain multi-scale features.

6. The remote sensing image reconstruction method based on multi-scale window spatial channel attention according to claim 3, characterized in that: The trained window space channel attention module includes a third module and a fourth module, the third module includes a window space attention module, a first depth separation convolution module, a convolution layer four, an activation function three, a global average pooling layer one, a convolution layer five and an activation function four, and the fourth module includes a channel attention module, a second depth separation convolution module, a global average pooling layer two, a convolution layer six, an activation function five, a convolution layer seven and an activation function six; the plurality of the trained window space channel attention modules are cascaded to process the normalized features to obtain aggregated features, including: For the first-level trained window space channel attention module, the normalized features are subjected to a dimension change operation, and the dimension-changed normalized features are processed by the window space attention module to obtain window space attention features; The convolution layer 4 is used to perform a convolution operation on the window spatial attention feature, and the activation function 3 is used to perform a nonlinear operation on the result processed by the convolution layer 4 to obtain a first spatial weight matrix; Using the first depth separation convolution module to process the normalized features to obtain first depth separation features; The first depth separation feature is pooled by the global average pooling layer 1, the convolution layer 5 is convolved by the result processed by the global average pooling layer 1, and the activation function 4 is used to perform a nonlinear operation on the result processed by the convolution layer 5 to obtain a first channel weight matrix; Multiplying the first channel weight matrix by the window spatial attention feature to obtain feature one; multiplying the first spatial weight matrix by the first depth separation feature to obtain feature two; adding the feature one to the feature two to obtain a window spatial attention enhanced feature; The feature dimension of the window space attention enhancement is changed, and the channel attention module is used to process the feature dimension of the window space attention enhancement after the dimension change to obtain the channel attention feature; The global average pooling layer 2 is used to perform a pooling operation on the channel attention feature, the convolution layer 6 is used to perform a convolution operation on the result processed by the global average pooling layer 2, and the activation function 5 is used to perform a nonlinear operation on the result processed by the convolution layer 6 to obtain a second channel weight matrix; Using the second depth separation convolution module to process the window spatial attention enhanced feature to obtain a second depth separation feature; The convolution layer 7 is used to perform a convolution operation on the second depth separation feature, and the activation function 6 is used to perform a nonlinear operation on the result processed by the convolution layer 7 to obtain a second spatial weight matrix; Multiplying the channel attention feature by the second spatial weight matrix to obtain feature three; multiplying the second depth separation feature by the second channel weight matrix to obtain feature four; adding the feature three to the feature four to obtain a first aggregate feature; For the j-th level trained window space channel attention module, the output of the j-1-th level trained window space channel attention module is processed as input to obtain the j-th aggregate feature; for the output of the last level window space channel attention module, the aggregate feature is obtained.

7. The remote sensing image reconstruction method based on multi-scale window spatial channel attention according to claim 6, characterized in that: The window space attention module includes activation function seven and a third splicing layer; the window space attention module is used to process the normalized features after the dimension change to obtain the window space attention features, including: Obtain a query matrix, a key matrix, and a value matrix through linear mapping, and change the dimensions of the query matrix, the key matrix, and the value matrix; divide the query matrix, the key matrix, and the value matrix after the dimension change into multiple equal parts on the channel dimension; For each query matrix, key matrix, and value matrix after the dimension change, a first weight matrix is ​​calculated based on the query matrix and the key matrix; A preset relative position encoding method is used to obtain a position encoding, and the position encoding is added to the first weight matrix to obtain a second weight matrix; if the current window spatial attention module is in an even position, a preset mask is added to the second weight matrix as an updated second weight matrix; Normalizing the second weight matrix using the activation function seven to obtain a third weight matrix; Multiplying the third weight matrix by the value matrix to obtain a feature calculated by window spatial attention, and performing dimension transformation on the feature to obtain a window spatial attention feature; For the query matrix, key matrix and value matrix after all dimensions have changed, multiple window spatial attention features are obtained, and the third splicing layer is used to splice the multiple window spatial attention features to obtain the window spatial attention features.

8. The remote sensing image reconstruction method based on multi-scale window spatial channel attention according to claim 6, characterized in that: The channel attention module includes an activation function eight and a fourth splicing layer; the channel attention module is used to process the features of the window space attention enhancement after the dimension change to obtain the channel attention features, including: Obtain a query matrix, a key matrix, and a value matrix through linear mapping, and change the dimensions of the query matrix, the key matrix, and the value matrix; divide the query matrix, the key matrix, and the value matrix after the dimension change into multiple equal parts on the channel dimension; For each query matrix, key matrix, and value matrix after the dimension change, a first weight matrix is ​​calculated based on the query matrix and the key matrix; Using activation function eight to normalize the first weight matrix to obtain a second weight matrix; Multiplying the second weight matrix by the value matrix to obtain a feature calculated by channel attention, and performing dimension transformation on the feature to obtain a channel attention feature; For the query matrix, key matrix and value matrix after all dimensions have changed, multiple channel attention features are obtained, and the fourth splicing layer is used to splice the multiple channel attention features to obtain the channel attention features.

9. The remote sensing image reconstruction method based on multi-scale window spatial channel attention according to claim 1, characterized in that: The process of acquiring the trained reconstruction model includes: Acquire data of multiple preset categories, pre-process the data of the preset categories as samples in a training data set, and acquire labels of the samples in the training data set; The samples in the training data set are input into the reconstruction model in batches for training until the number of training times or the degree of convergence meets the preset conditions, thereby obtaining a trained reconstruction model.

10. A remote sensing image reconstruction device based on multi-scale window spatial channel attention, comprising a processor, a communication interface, a memory and a communication bus, characterized in that: The processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 9 when executing the program stored in the memory.

Citation Information

Cited By

  • Rotating perception enhanced variable window attention super-resolution method and system

    CN120563322A

  • A rotation-aware enhanced variable window attention super-resolution method and system

    CN120563322B