Remote sensing image detection method based on scale perception and spatial selection hierarchical interaction

By improving the U-net network's scale perception module SAM, the differential multi-scale attention module DMSAM and the spatial selection hierarchical interaction module SSIHM, the problem of insufficient feature interaction in remote sensing image change detection is solved, and the detection accuracy and recall rate are improved, especially in building detection.

CN120339824APending Publication Date: 2025-07-18CHINA THREE GORGES UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510342275.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing deep learning remote sensing image change detection method is not ideal when processing images with significant scale changes, and lacks effective interactions between features at different levels, resulting in loss of detailed information and insufficient semantic information during feature recovery.

Method used

Adopting an improved U-net network based on the interaction of scale perception and spatial selection hierarchical selection, multi-scale features are extracted through the scale perception module SAM, differentiated multi-scale attention module DMSAM enhances feature interaction, and the complementarity between shallow and deep features is achieved through the spatial selection hierarchical interaction module SSIHM, to improve feature resolution and semantic reconstruction effect.

Benefits of technology

It effectively realizes the selective interaction between the detailed information in shallow features and the semantic information in deep features, and improves the accuracy and recall of remote sensing image change detection, especially in building inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339824A_ABST
    Figure CN120339824A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image detection method based on scale perception and spatial selection hierarchical interaction, which comprises the following steps of: firstly, extracting features by partitioning and paralleling different sizes of depth separable convolution, introducing channel attention, and designing a scale perception module so as to effectively extract changing objects with different shapes and scales; secondly, shallow features and deep features are enhanced through spatial attention crossing, a spatial selection hierarchy interaction module is provided, and the characterization capacity of the features is refined; and finally, based on the difference chart of the two remote sensing images, giving a difference multi-scale attention module to highlight change information and suppress unchanged information. Compared with the existing six comparison change detection networks, the change detection results of the method provided by the invention on four public data sets of WHU, Google, LEVIR and GVLM are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing technology, and in particular to a remote sensing image detection method based on scale perception and spatial selection hierarchical interaction. Background Art

[0002] Change detection of remote sensing images has important application values in the fields of urban planning, disaster relief, etc. The method based on deep learning can automatically and efficiently perform change detection from remote sensing images.

[0003] Currently, the deep learning remote sensing image change detection method still has unsatisfactory effects when dealing with images with significant scale changes, and most methods lack effective interaction between different hierarchical features in the decoding stage.

[0004] DifUnet++: A Satellite Images Change Detection Network Based on Unet++ And Differential Pyramid improves the input and upsampling processes in Unet++, and at the same time adopts a multi-edge output fusion strategy to enhance the detection ability of different scale change regions and protect the detailed information of change edges. In the paper RemoteSensing Building Change Detection Based on Improved U-Net, in order to solve the problems of feature misalignment and loss of detailed information caused by deconvolution upsampling in the feature recovery process, a twin network is improved based on U-net, but the consideration of global information is lacking.

[0005] Although the above-mentioned deep learning change detection methods based on U-net (or U-net++) have their own advantages and can perform change detection more effectively. However, they have the following deficiencies: (1) In the encoding stage, the characteristics of variable scales of change targets in the image are not considered enough, and the multi-scale problem is mostly considered after the encoder extracts features; (2) In the decoding process, the shallow features and deep features are mostly fused by simple concatenation or addition, lacking the interaction and guidance between different hierarchical features; In view of the above problems, based on the classic U-net network, a high-resolution remote sensing image change detection method based on scale perception and spatial selection hierarchical interaction is proposed.

[0006] In view of the above problems, based on the classic U-net network, a high-resolution remote sensing image change detection method based on scale perception and spatial selection hierarchical interaction is proposed. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a remote sensing image detection method based on scale-aware and spatial selection hierarchical interaction, which realizes effective selective interaction between the detailed information in the shallow features and the semantic information in the deep features, and refines the spatial resolution restoration of the hierarchical features and the semantic reconstruction effect of the low-level features.

[0008] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A remote sensing image detection method based on scale-aware and spatial selection hierarchical interaction includes the following steps: Step1: Select remote sensing images of buildings in different time periods and different regions to obtain the dataset D1 as the detection dataset; preprocess the images in the dataset D1, and cut them into 256×256; divide the dataset into a training set and a test verification set according to a fixed ratio; Step2: Input the training set of the dataset D1 processed in Step1 into the improved U-net network encoder for feature extraction, and construct each layer of the difference auxiliary branch to extract spatio-temporal difference features: X d1, X d2, X d3, X d4, X d5 , each level generates an attention map with the help of the difference branch. Each encoding layer is composed of a convolution and a Scale-Aware Module (SAM), and the performance of the cascaded main branch is enhanced through the Difference Multi-Scale Attention Module (DMSAM) to obtain the main branch features of each layer: X c1, X c2, X c3, X c4, X c5 ; Step3: Input the enhanced main branch feature maps of each layer in Step2 into the decoding layer part. To comprehensively consider the shallow and deep features, define the resolutions of the shallow and deep features, and use the Spatial Selection Hierarchical Interaction Module (SSIHM) to introduce the detailed information of the shallow features into the deep features and the semantic information of the deep features into the shallow features, and the two guide each other to obtain the features after interactive guidance; Step4: Process the features after interactive guidance obtained in Step3 to generate a prediction probability map for building extraction for loss calculation; Step 5. Optimize the model using the evaluation metrics of the test validation set; re-optimize the process of Step 1 - Step 3 through the evaluation metrics of the test validation set, including precision, recall, and IoU, to improve the detection effect.

[0009] In the above Step 2, each encoding layer in the improved U-net network encoder consists of a 3×3 convolution and a Scale-Aware Module (SAM). A residual connection is introduced in the middle to alleviate overfitting. The 3×3 convolution is used to expand the number of channels, and the Scale-Aware Module (SAM) extracts multi-scale information and refines scale features.

[0010] The establishment formula of the Scale-Aware Module (SAM) in the above improved U-net network encoder is: ; ; ; (Conv1(CAM(F1)) F; LP ( F2 ) i {1, 2, 3, 4, 5}; In the formula, F represents the input feature map, represents linear processing of the feature map, represents extracting multi-scale features from the feature map. Conv1 represents a 1×1 convolution, CAM represents a feature enhancement module, and MLP represents a multi-layer perceptron operation to strengthen the fusion between different scale features. represents the feature map of each layer of the auxiliary branch, represents the feature map of each layer of the main branch. w represents the width of the feature map, c represents the number of channels of the feature map, and h represents the width of the feature map.

[0011] The above extraction of multi-scale features from the feature map , with the help of depthwise separable convolutions of four different sizes, captures spatial features at multiple scales while reducing the number of model parameters and computational costs. MSMC can be expressed as: ; In the formula, = , , , represents evenly dividing the input feature F′ into four parts in the channel dimension. Concat represents concatenation. represents the size of ki Depthwise separable convolution, where k i {5, 7, 9, 11}

[0012] The above enhanced feature module CAM selectively adjusts features according to the importance between different scales to optimize feature representation; the enhanced feature module CAM adopts the channel attention part of the classic lightweight attention module CBAM

[23] , and the specific formula is expressed as: ) FC(MaxPool( )))) In the formula, F′ represents the input feature, AvgPool represents global average pooling in the spatial dimension, MaxPool represents global max pooling in the spatial dimension, and FC represents the fully connected layer.

[0013] In the above Step2, after the improved U-net network encoder encodes one layer of features, it performs max pooling downsampling with a stride of 2 to generate five feature maps of different sizes. The five feature maps of different sizes are utilized by the difference auxiliary branch, and the attention feature maps are generated through the difference multi-scale attention module DMSAM to enhance the five feature maps of different sizes of the main branch. The number of channels of the cascaded branch feature maps is 32, 64, 128, 256, and 512 respectively, the number of channels in the first stage of the difference branch is 32, and the number of channels in each subsequent stage is 64.

[0014] The formula for the above difference auxiliary branch to use the difference multi-scale attention module DMSAM to generate attention feature maps to enhance the main branch is: = DMSAM( ; M i = GMP(GAP( )))) {1, 3, 5, 7} DMSAM( = Sigmoid( 7(CAM( , )))) In the formula, represents the feature maps of each layer of the auxiliary branch, represents the feature maps of each layer of the main branch, GMP and GAP respectively represent max pooling and average pooling in the channel dimension, uses 3×3 convolutions with four dilation rates of 1, 3, 5, and 7 to extract attention at different scales, and Sigmoid represents the activation function, 7 represents a 7×7 convolution, represents a concatenation operation.

[0015] In the above Step3, the shallow features and the deep features are feature-complemented through the Spatial Selection Hierarchical Interaction Module (SSIHM). First, the shallow features and the deep features go through convolution and transposed convolution to adjust the number of channels and the size of the feature maps to make them consistent. Then, the two are concatenated, and max-pooling and average-pooling are respectively applied to the concatenated features in the channel dimension. After 7×7 convolution and the Sigmoid activation function, a two-channel spatial attention map is obtained. Then, the attention map is divided into two parts in the channel dimension and multiplied by the adjusted shallow and deep features respectively. At this time, the feature maps go through the final weight assignment, embedding semantic information in the shallow features and detailed information in the deep features. Finally, the two are added and fused, and the number of channels is adjusted back to the original number of channels of the shallow features through 1×1 convolution to obtain the interacted features.

[0016] The calculation formula of the above Spatial Selection Hierarchical Interaction Module (SSIHM) is as follows: ; ; ) ; = ; In the formula, f i represents the shallow features, f j represents the deep features, and represent 3×3 convolution and transposed convolution with a stride of 2. GMP and GAP respectively represent max-pooling and average-pooling in the channel dimension. represents a concatenation operation, represents 7×7 convolution, Sigmoid represents the activation function, W represents the two-channel spatial attention map, represents 1×1 convolution operation to adjust the channels to obtain the interacted features, and finally the predicted probability map is obtained.

[0017] A remote sensing image detection method based on scale perception and spatial selection hierarchical interaction provided by the present invention has the following beneficial effects: 1) In the feature extraction stage, in order to capture multi-scale features and enhance the detection ability for objects with different shape and scale changes, a scale perception module (SAM) is proposed.

[0018] 2) In the feature interaction and fusion stage, in order to achieve effective selective interaction between the detailed information in the shallow features and the semantic information in the deep features, and to refine the spatial resolution recovery of the hierarchical features and the semantic reconstruction effect of the low-level features, a spatial selective hierarchical interaction module SSIHM is proposed.

[0019] 3) For object detection, considering the particularity of the difference map in two-phase remote sensing images, a difference multi-scale attention module DMSAM is proposed by extracting the spatio-temporal difference features of the difference map. Brief Description of the Drawings

[0020] The present invention will be further described below with reference to the drawings and embodiments: Figure 1 It is a structural diagram of a high-resolution remote sensing image change detection network with scale perception and spatial selective hierarchical interaction according to an embodiment of the present invention; Figure 2 It is a structural diagram of the SAM scale perception module according to an embodiment of the present invention; Figure 3 It is a structural diagram of the DMSAM difference multi-scale attention module according to an embodiment of the present invention; Figure 4 It is a structural diagram of the SSIHM spatial selective hierarchical interaction module according to an embodiment of the present invention. Detailed Description of the Invention

[0021] The technical solution of the present invention will be described in detail below with reference to the drawings and embodiments.

[0022] Embodiment 1: A remote sensing image detection method based on scale perception and spatial selective hierarchical interaction includes the following steps: Step1. Select remote sensing images of buildings in different time periods and different regions to obtain a dataset D1 as the detection dataset; the data source can include time series remote sensing images, that is, images of the same area obtained at different times, or remote sensing images of different geographical regions, to ensure the diversity and robustness of the data; preprocess the images in the dataset D1, and crop the images into small pieces of 256×256 to improve the training efficiency of the network model; then, divide the dataset into a training set and a test and validation set according to a fixed ratio to ensure that the model can perform generalization learning on different data; Step2. Input the training set of the dataset D1 processed in Step1 into the improved U-net network encoder for feature extraction, and construct each layer of the difference auxiliary branch to extract spatio-temporal difference features: X d1, X d2, X d3, X d4, X d5, at each level, an attention map is generated through differential branches. Each encoding layer consists of a convolution and a Scale-Aware Module (SAM), and the performance of the cascaded main branch is enhanced through a Difference Multi-Scale Attention Module (DMSAM) to obtain the main branch features of each layer: X c1, X c2, X c3, X c4, X c5 ; Step3: Input the enhanced main branch feature maps of each layer in Step2 into the decoding layer part. To comprehensively consider shallow and deep features, define the resolutions of shallow and deep features, and use a Spatial Selection Hierarchical Interaction Module (SSIHM) to introduce the detailed information of shallow features into deep features and the semantic information of deep features into shallow features, and the two guide each other to obtain the features after interactive guidance; Step4: Process the features after interactive guidance obtained in Step3 to generate a prediction probability map for building extraction for loss calculation, and the change detection result is significantly better than the original U-net network; Step5: Optimize the model through the evaluation metrics of the test validation set; re-optimize the process of Step1 - Step3 through the evaluation metrics of the test validation set, including precision, recall, and IoU, to improve the detection effect.

[0023] Shallow features refer to the features extracted by the early layers in a neural network, and the resolution of shallow features is larger than that of deep features. They capture the basic information of the image, usually low-level, local, and structural image information.

[0024] Processing steps of the test validation set: The test set is usually separated from the dataset to ensure that the model training process is not affected by the test data. The steps usually include: First, prepare and preprocess the training set and the test set; then, train the object detection model using the training set; next, apply the trained model to the test validation set to generate prediction results; finally, evaluate the performance of the model on the test set by calculating common evaluation metrics such as precision (Pre), recall (Rec), and intersection over union (IoU, the evaluation metrics used in this paper). This process can help developers understand the performance of the model in actual applications and make adjustments and optimizations as needed.

[0025] Optimize the training through the evaluation metrics of the test validation set: Through the evaluation metrics of the test validation set, including precision, recall, and IoU, the deficiencies of the model can be identified and the training steps Step1 - Step3 can be optimized. For example, if the precision is low, it may be due to a large number of false positives, and false positives can be reduced by adjusting the threshold of the model, increasing positive samples, or using a more effective loss function; if the recall is low, it may be because the model misses some buildings, and the recall can be improved by increasing minority class samples, improving data augmentation, or introducing more context information; if the IoU is low, it indicates that there are errors in the building localization of the model, and the IoU can be improved by increasing the network resolution, using a stronger feature extraction network, or optimizing the localization loss. According to these evaluation results, the model architecture, loss function, and dataset are adjusted accordingly to optimize the training process and improve the detection effect.

[0026] In the above Step2, each encoding layer in the improved U-net network encoder consists of a 3×3 convolution and a Scale-Aware Module (SAM). A residual connection is introduced in the middle to alleviate overfitting. The 3×3 convolution is used to expand the number of channels, and the Scale-Aware Module SAM extracts multi-scale information and refines scale features.

[0027] The establishment formula of the Scale-Aware Module SAM in the above improved U-net network encoder is as follows: ; ; ; (Conv1(CAM(F1)) F ; LP ( F2 ) i {1,2,3,4,5}; In the formula, F represents the input feature map, represents linear processing of the feature map, represents extracting multi-scale features from the feature map. Conv1 represents a 1×1 convolution, CAM represents a feature enhancement module, and MLP represents a multi-layer perceptron operation to strengthen the fusion between different scale features. represents the feature maps of each layer of the auxiliary branch, represents the feature maps of each layer of the main branch. w represents the width of the feature map, c represents the number of channels of the feature map, and h represents the width of the feature map.

[0028] The above extraction of multi-scale features from the feature map , by means of depthwise separable convolutions of four different sizes, spatial features at multiple scales are captured while reducing the number of model parameters and computational costs. Large convolutional kernels focus on extracting global and high-level semantic features, while small convolutional kernels are more focused on capturing local and detailed features. MSMC can be expressed as: ; In the formula, = , , , means evenly dividing the input feature F′ into four parts in the channel dimension, and Concat means concatenation. represents a depthwise separable convolution of size k i , where k i {5, 7, 9, 11}.

[0029] The above-mentioned enhanced feature module CAM selectively adjusts features according to the importance between different scales to optimize feature representation; the enhanced feature module CAM adopts the channel attention part of the classical lightweight attention module CBAM

[23] , and the specific formula is expressed as: ) FC(MaxPool( )))) ; In the formula, F′ represents the input feature, AvgPool represents global average pooling in the spatial dimension, MaxPool represents global maximum pooling in the spatial dimension, and FC represents the fully connected layer.

[0030] In the above-mentioned Step2, after the improved U-net network encoder encodes one layer of features, it performs max-pooling downsampling with a stride of 2, generating 5 feature maps of different sizes. The 5 feature maps of different sizes are utilized by the difference auxiliary branch, and attention feature maps are generated through the difference multi-scale attention module DMSAM to enhance the 5 feature maps of different sizes of the main branch. The number of channels of the cascaded branch feature maps are 32, 64, 128, 256, and 512 respectively. Considering the limited GPU, the number of channels in the first stage of the difference branch is 32, and the number of channels in each subsequent stage is 64.

[0031] The formula for the above-mentioned difference auxiliary branch to use the difference multi-scale attention module DMSAM to generate attention feature maps to enhance the main branch is: = DMSAM( ; M i = GMP(GAP( )))) , {1, 3, 5, 7}; DMSAM( = Sigmoid( 7(CAM( , )))); In the formula, represents the feature maps of each layer of the auxiliary branch, represents the feature maps of each layer of the main branch. GMP and GAP respectively represent max-pooling and average-pooling in the channel dimension, 3×3 convolutions with dilation rates of 1, 3, 5, and 7 are used to extract attention at different scales. Sigmoid represents the activation function, 7 represents a 7×7 convolution, represents the concatenation operation.

[0032] In Step 3 above, the shallow features are feature-complemented with the deep features through the Spatial Selection Hierarchical Interaction Module (SSIHM). First, the shallow features and the deep features go through convolution and transposed convolution to adjust the number of channels and the size of the feature maps to make them consistent. Then, the two are concatenated, and max-pooling and average-pooling are respectively performed on the concatenated features in the channel dimension. After passing through a 7×7 convolution and the Sigmoid activation function, a two-channel spatial attention map is obtained. Then, the attention map is divided into two parts along the channel dimension and multiplied with the adjusted shallow and deep features respectively. At this time, the feature maps go through the final weight assignment, embedding semantic information in the shallow features and detailed information in the deep features. Finally, the two are added and fused, and the number of channels is adjusted back to the original number of channels of the shallow features through a 1×1 convolution to obtain the interacted features.

[0033] The calculation formula of the above Spatial Selection Hierarchical Interaction Module (SSIHM) is: ; ; ); = ; In the formula, f i represents the shallow features, f j represents the deep features, and represent a 3×3 convolution and a transposed convolution with a stride of 2. GMP and GAP respectively represent max-pooling and average-pooling in the channel dimension, represents the concatenation operation, represents a 7×7 convolution. Sigmoid represents the activation function, W represents the two-channel spatial attention map, The 1×1 convolution operation is used to adjust the channels to obtain the interacted features, and finally the predicted probability map is obtained.

[0034] Example 2: For the remote sensing image detection method based on scale-aware and spatial selection hierarchical interaction, in the U-net network, in the feature extraction stage, in order to capture multi-scale features and enhance the detection ability for objects with different shape scale changes, a scale-aware module SAM is proposed. In the feature interaction and fusion stage, in order to achieve effective selective interaction between the detailed information in the shallow features and the semantic information in the deep features, and to refine the spatial resolution recovery of the hierarchical features and the semantic reconstruction effect of the low-level features, a spatial selection hierarchical interaction module SSIHM is proposed. And for object detection, considering the particularity of the difference map in two-phase remote sensing images, by extracting the spatio-temporal difference features of the difference map, a difference multi-scale attention module DMSAM is proposed.

[0035] It includes the following steps: S1: Select a large number of building pictures to obtain the dataset D1 as the dataset for this experiment; preprocess the pictures in the D1 dataset, cut them into a size of 256×256 to improve the training speed of the network model; divide them into a training set, a validation set and a test set according to a fixed ratio; S2: Input the training set of the processed D1 dataset into the improved U-net network encoder for feature extraction, and additionally construct a difference auxiliary branch to extract spatio-temporal difference features at each layer: X d1, X d2, X d3, X d4, X d5 , and at each level, an attention map is generated by means of the difference branch. Each encoding layer consists of a convolution and a scale-aware module (Scale-Aware Module, SAM), and the performance of the cascaded main branch is enhanced by a difference multi-scale attention module (Difference Multi-Scale Attention Module, DMSAM) to obtain the main branch features of each layer: X c1, X c2, X c3, X c4, X c5 ; As Figure 2 shown, in step S2, it includes the following sub-steps: 1) Each encoding layer of the improved U-Net encoder part consists of a 3×3 convolution and a Scale-Aware Module (SAM). A residual connection is introduced in the middle to alleviate overfitting. The 3×3 convolution is used to expand the number of channels, and the scale-aware module extracts multi-scale information and refines scale features. The scale-aware module (Scale-Aware Module, SAM) of the encoding layer The calculation formula is: ; ; ; (Conv1(CAM(F1)) F; LP ( F2 ) i {1, 2, 3, 4, 5}; In the formula, F represents the input feature map, represents linear processing of the feature map, represents extracting multi-scale features from the feature map. Conv1 represents a 1×1 convolution, CAM represents a feature enhancement module, and MLP represents a multi-layer perceptron operation to strengthen the fusion between different scale features, represents the feature map of each layer of the auxiliary branch, represents the feature map of each layer of the main branch, w represents the width of the feature map, c represents the number of channels of the feature map, and h represents the width of the feature map; 2) The following is a detailed introduction to the MSMC functional module: By means of depthwise separable convolutions of four different sizes, it aims to capture spatial features at multiple scales while reducing the number of model parameters and computational costs. Large convolutional kernels focus on extracting global and high-level semantic features, while small convolutional kernels are more focused on capturing local and detailed features. The MSMC module can be expressed as: ; In the formula, = , , , means evenly dividing the input feature F′ into four parts in the channel dimension. Concat represents concatenation, represents a depthwise separable convolution of size k i where k i {5, 7, 9, 11}.

[0036] 3) The CAM module is introduced after the MSMC module. This module selectively adjusts features according to the importance between different scales to optimize feature representation. Here, the channel attention part of the classic lightweight attention module CBAM

[23] is adopted for CAM. The specific formula is expressed as: ) FC(MaxPool( )))) ; In the formula, F′ represents the input feature, AvgPool represents global average pooling in the spatial dimension, MaxPool represents global maximum pooling in the spatial dimension, and FC represents the fully connected layer.

[0037] As Figure 3 shown, in step S2, the following steps are included: 1) After encoding one layer of features, max-pooling downsampling with a stride of 2 is adopted, which can generate five different sizes of feature maps. Among them, using the difference auxiliary branch, five different sizes of feature maps generate attention feature maps through the Difference Multi-Scale Attention Module (DMSAM) to enhance the five different sizes of feature maps of the main branch. The number of channels of the cascaded branch feature maps is 32, 64, 128, 256, and 512 respectively. Considering the limited GPU, the number of channels in the first stage of the difference branch is 32, and the number of channels in each subsequent stage is 64. The difference auxiliary branch uses the Difference Multi-Scale Attention Module (DMSAM) to generate attention feature maps to enhance the main branch. The calculation formula is: =DMSAM( ; M i = GMP(GAP( )))) , {1,3,5,7} ; DMSAM( =Sigmoid( 7(CAM( , )))) ; In the formula, represents the feature maps of each layer of the auxiliary branch, represents the feature maps of each layer of the main branch, GMP and GAP respectively represent max-pooling and average pooling in the channel dimension respectively, uses four 3×3 convolutions with dilation rates of 1, 3, 5, and 7 respectively to extract attention at different scales, and Sigmoid represents the activation function. 7 represents a 7×7 convolution, represents a concatenation operation; S3: Input the enhanced main branch feature maps of each layer in step S2 into the decoding layer part. To comprehensively consider shallow and deep features, a Spatial Selection Hierarchical Interaction Module (SSIHM) is designed. By introducing the detailed information of shallow features into deep features and the semantic information of deep features into shallow features, the two guide each other to obtain rich-information features; As Figure 4 shown, step S3 includes the following steps: 1) The shallow features are feature-complemented with the deep features through the Spatial Selection Hierarchical Interaction Module (SSIHM). First, the middle and shallow features and the deep features go through convolution and transposed convolution to adjust the number of channels and the size of the feature maps to be the same. Then, the two are concatenated. Max-pooling and average-pooling are respectively adopted for the concatenated features in the channel dimension. After 7×7 convolution and the Sigmoid activation function, a two-channel spatial attention map is obtained. Then, the attention map is divided into two parts in the channel dimension and multiplied by the adjusted shallow and deep features respectively. At this time, the feature maps go through the final weight assignment, embedding semantic information in the shallow features and detailed information in the deep features. Finally, the two are added and fused, and the number of channels is adjusted back to the original number of channels of the shallow features through 1×1 convolution to obtain the interacted features. The following is a detailed analysis of the SSIHM spatial selection hierarchical interaction module, and the calculation formula is: ; ; ); = ; In the formula, f i represents the shallow features, f j represents the deep features, and represent 3×3 convolution and transposed convolution with a stride of 2. GMP and GAP respectively represent max-pooling and average-pooling in the channel dimension, represents the concatenation operation, represents 7×7 convolution, Sigmoid represents the activation function, W represents the two-channel spatial attention map, represents the 1×1 convolution operation to adjust the channels to obtain the interacted features, and finally the predicted probability map is obtained.

[0038] Example 3: To verify the specific effect of the remote sensing image change detection network designed based on scale perception and spatial selection hierarchical interaction of the present invention, the trained model was experimented on the WHU dataset. The WHU building dataset comes from the GPCV team of Wuhan University and consists of satellite and aerial datasets. The experimental data uses the aerial dataset, with an image size of 512 × 512 pixels and a spatial resolution of 0.3m. In this study, according to the original resolution, each image was further cropped into smaller segments of 256 × 256 pixels without overlap. The final data was obtained.

[0039] To evaluate the proposed method, precision (Pre), recall (Rec), and intersection over union (IoU) were used as evaluation metrics. The higher the precision value, the better the model performance. The higher the recall, the lower the omission rate of the model. The higher the IoU, the better the localization accuracy of the model. The experimental environment was: Windows 10 operating system, 16GB of memory, Ryzen 5 5600G CPU, and NVIDIA GeForce GTX 1660 graphics card. The programming language was Python 3.8, and the deep learning framework was Pytorch 1.12. The test results were compared with the current mainstream building extraction models, and the results on the WHU dataset are shown in Table 1.

[0040] Table 1 Comparison of Results on WHU Dataset

[0041] The present invention reaches the maximum values in terms of precision, recall, and intersection, which are 93.46%, 90.04%, and 84.70% respectively. It shows more excellent performance compared with the advanced algorithms of many classic remote sensing image change detection networks.

[0042] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A remote sensing image detection method based on hierarchical interaction of scale perception and spatial selection, characterized in that, It includes the following steps: Step 1: Select remote sensing images of buildings in different time periods and different regions to obtain dataset D1 as the detection dataset; preprocess the images in dataset D1, cut them into a size of 256×256, and divide the dataset into a training set and a test validation set according to a fixed ratio; Step 2. Input the training set of the dataset D1 processed in Step 1 into the improved U-net network encoder for feature extraction, and construct each layer of the difference auxiliary branch to extract spatio-temporal difference features: X d1, X d2, X d3, X d4, X d5 , and each level generates an attention map through the difference branch. Each encoding layer is composed of a convolution and a Scale-Aware Module (SAM), and the performance of the cascaded main branch is enhanced through a Difference Multi-Scale Attention Module (DMSAM) to obtain the main branch features of each layer: X c1, X c2, X c3, X c4, X c5 ; Step 3: Input the enhanced main branch feature maps of each layer in Step 2 into the decoding layer part. To comprehensively consider shallow and deep features, define the resolutions of shallow and deep features, and use the Spatial Selection Hierarchical Interaction Module, that is, SSIHM, to introduce the detailed information of shallow features into deep features and the semantic information of deep features into shallow features, and the two guide each other to obtain the features after interactive guidance; Step 4: Process the features after interactive guidance obtained in Step 3 to generate a prediction probability map for building extraction for loss calculation; Step 5: Optimize the model through the evaluation metrics of the test validation set; through the evaluation metrics of the test validation set, including accuracy, recall rate, and IoU, re-optimize the process of Step 1 - Step 3 to improve the detection effect.

2. The remote sensing image detection method based on hierarchical interaction of scale perception and spatial selection according to claim 1, wherein In the aforementioned Step 2, each encoding layer in the improved U-net network encoder consists of a 3×3 convolution and a Scale-Aware Module, that is, SAM. A residual connection is introduced in the middle to alleviate overfitting. Among them, the 3×3 convolution is used to expand the number of channels, and the Scale-Aware Module SAM extracts multi-scale information and refines scale features.

3. The remote sensing image detection method based on hierarchical interaction of scale perception and spatial selection according to claim 2, wherein The establishment formula of the Scale-Aware Module SAM of the improved U-net network encoder is: ; ; ; (Conv1(CAM(F1)) F; LP ( F2 ) i {1,2,3,4,5}; Wherein, F represents the input feature map, represents linear processing of the feature map, represents extracting multi-scale features from the feature map, Conv1 represents a 1×1 convolution, CAM represents a feature enhancement module, MLP represents a multi-layer perceptron operation, and enhances the fusion between features of different scales. represents the feature maps of each layer of the auxiliary branch, represents the feature maps of each layer of the main branch, w represents the width of the feature map, c represents the number of channels of the feature map, and h represents the width of the feature map.

4. The remote sensing image detection method based on scale-aware and spatial selection hierarchical interaction according to claim 3, wherein Extracting multi-scale features from the feature map , by means of depthwise separable convolutions of four different sizes, captures spatial features at multiple scales while reducing the number of model parameters and computational costs. MSMC can be expressed as: ; In the formula, = , , , means that the input feature F′ is evenly divided into four parts in the channel dimension, and Concat means concatenation. represents a depthwise separable convolution of size k i , where k i ∈ {5, 7, 9, 11}.

5. The remote sensing image detection method based on scale-aware and spatial selection hierarchical interaction according to claim 4, characterized in that The enhanced feature module CAM selectively adjusts the features according to the importance between different scales to optimize the feature representation; The enhanced feature module CAM adopts the channel attention part of the classic lightweight attention module CBAM[23], and the specific formula is expressed as: ) FC(MaxPool( )))) ; In the formula, F′ represents the input feature, AvgPool represents global average pooling in the spatial dimension, MaxPool represents global maximum pooling in the spatial dimension, and FC represents the fully connected layer.

6. The remote sensing image detection method based on hierarchical interaction of scale perception and spatial selection according to claim 5, characterized in that, In the aforementioned Step 2, after encoding one layer of features in the improved U-net network encoder, max pooling with a stride of 2 is used for downsampling to generate 5 different sizes of feature maps. Use the differential auxiliary branch for 5 different sizes of feature maps and generate attention feature maps through the differential multi-scale attention module DMSAM to enhance the 5 different sizes of feature maps of the main branch. The number of channels of the cascaded branch feature maps is 32, 64, 128, 256, and 512 respectively. The number of channels in the first stage of the differential branch is 32, and the number of channels in each subsequent stage is 64.

7. The remote sensing image detection method based on scale-aware and spatial selection hierarchical interaction according to claim 6, characterized in that The calculation formula for the differential auxiliary branch to use the differential multi-scale attention module DMSAM to generate attention feature maps to enhance the main branch is: =DMSAM( ; M i = GMP(GAP( )))) , {1,3,5,7} ; DMSAM( = Sigmoid( 7(CAM( , )))); In the formula, represents the feature maps of each layer of the auxiliary branch, represents the feature maps of each layer of the main branch. GMP and GAP respectively represent maximum pooling and average pooling in the channel dimension, uses 3×3 convolutions with dilation rates of 1, 3, 5, and 7 respectively to extract attention at different scales. Sigmoid represents the activation function, 7 represents a 7×7 convolution, represents the concatenation operation.

8. The remote sensing image detection method based on scale-aware and spatial selection hierarchical interaction according to claim 7, wherein In the aforementioned Step 3, the shallow features are complementary to the deep features through the Spatial Selection Hierarchical Interaction Module (SSIHM). First, the shallow features and the deep features undergo convolution and deconvolution to adjust the number of channels and the size of the feature maps to make them consistent. Then, the two are concatenated, and max-pooling and average-pooling are respectively applied to the concatenated features in the channel dimension. After 7×7 convolution and the Sigmoid activation function, a two-channel spatial attention map is obtained. Next, the attention map is divided into two parts along the channel dimension and multiplied by the adjusted shallow and deep features respectively. At this time, the feature maps undergo the final weight allocation, embedding semantic information in the shallow features and detailed information in the deep features. Finally, the two are added and fused, and the number of channels is adjusted back to the original number of channels of the shallow features through 1×1 convolution to obtain the interacted features.

9. The remote sensing image detection method based on hierarchical interaction of scale perception and spatial selection according to claim 8, characterized in that The calculation formula of the aforementioned Spatial Selection Hierarchical Interaction Module (SSIHM) is as follows: ; ; ) ; = ; where, f i represents shallow features, and f j represents deep features, and represent 3×3 convolution and deconvolution with a stride of 2. GMP and GAP respectively represent max pooling and average pooling in the channel dimension. represents a concatenation operation, represents 7×7 convolution. Sigmoid represents an activation function. W represents a two-channel spatial attention map. represents a 1×1 convolution operation to adjust the channels to obtain the interacted features, and finally obtain the predicted probability map.

Citation Information

Cited By

  • Floating enteromorpha green tide remote sensing recognition model and method fused with multi-attention mechanism neural network

    CN121305389A

  • Remote Sensing Recognition Model and Method for Floating Ulva prolifera Green Tide Integrating Multi-Attention Mechanism Neural Networks

    CN121305389B