Remote sensing change detection method based on difference enhancement and mixed multi-scale downsampling

Through the DEHMD-Net network, two-stage differential enhancement and mixed multi-scale downsampling technology are used to solve the problem of pseudo-change information and information loss in remote sensing image change detection, improving detection accuracy and robustness.

CN120339790APending Publication Date: 2025-07-18CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285339.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing remote sensing image change detection methods, there are problems such as a lot of pseudo-change information and downsampling steps leading to information loss, feature redundancy and detection accuracy reduction.

Method used

The remote sensing change detection network DEHMD-Net based on differential reinforcement and mixed multi-scale downsampling is adopted. Through the two-stage differential reinforcement module, the hybrid multi-scale downsampling module and the hierarchical interactive fusion module, the learning of the change area is enhanced, pseudo-change information is suppressed, detailed information is retained, and feature redundancy is reduced.

Benefits of technology

It improves the accuracy and robustness of remote sensing change detection, can effectively identify changing areas, reduce pseudo-change information, enhance feature expression ability, and adapt to detection of targets at different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339790A_ABST
    Figure CN120339790A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing change detection method based on difference enhancement and mixed multi-scale downsampling. A model in the invention is composed of a double-branch encoder and a single decoder. The coding part is composed of a feature enhancement module (FEM), a two-stage difference module and a mixed multi-scale down-sampling module. A pair of dual-time images firstly enter the FEM for feature extraction operation. AttnConv is introduced into an FEM module, high-frequency local detail information and original context information are effectively extracted, defects of CNN can be made up, and features are fully extracted. And after the convolution operation, a dual-stage difference enhancement module is entered for enhancing the features of the changed object. Then, down-sampling operation is carried out, and target change information and detail information of different scales are effectively reserved by adopting a mixed multi-scale down-sampling module; and finally, in the decoding stage, a hierarchical interactive fusion module and a convolutional feature extraction (CFEM) module are used in the decoding stage, and the decoding stage is composed of the hierarchical interactive fusion module and the convolutional feature extraction (CFEM) module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, in particular to remote sensing image change detection technology, and specifically to a remote sensing change detection method based on difference enhancement and hybrid multi-scale downsampling. Background Art

[0002] Remote sensing change detection refers to the technology of obtaining changes in surface information by analyzing remote sensing images of the same area at different times, and is widely used in many fields such as land use, disaster assessment, forest monitoring, and urban scene research. In the past few decades, many change detection methods have been proposed. With the rapid development of deep learning technology, a large number of deep learning change detection methods have been proposed. Change detection networks can be constructed based on different deep learning models, such as Convolutional Neural Network (CNN), Stacked Autoencoder Network, and Deep Belief Network, etc., among which CNN is the most widely used.

[0003] As a classic CNN network, U-net combines the high-level features of the encoder with the low-level features of the decoder through skip connections, retains more detailed information, and improves the change detection accuracy. Based on these advantages, U-net has also been widely applied in other computer vision fields. Many existing deep learning change detection networks are proposed by improving the U-net (or U-net++) network. First, the literature "Remote Sensing Image Building Change Detection Based on FlowS-Unet" uses the refinement structure of FlowNet to improve U-net and proposes the FlowS-Unet network model to enhance the detailed localization of building change targets. The literature "A cross swin transformer based siamese u-shape network for change detection in remote sensing images" designs a new network architecture by combining cross Swin Transformer and Siamese U-Net to effectively extract the features of remote sensing images and perform change detection, and more accurately extract the feature differences between dual-temporal images.

[0004] To build the dependency of the global context, many scholars have introduced Transformer into U-net for change detection. The literature "Context and difference enhancement network for change detection" adds Transformer in the encoding to construct context information and difference information, highlighting the changed areas. The literature "Remote sensing image change detection with transformers" introduces the Transformer architecture, which effectively captures global context information using the self-attention mechanism, improving the accuracy of change detection. However, the above-mentioned models based on Transformer have a large number of model parameters, slow inference speed, and serious interference from "false changes".

[0005] Although these network models show their respective advantages, they still have the following problems:

[0006] (1) Some current networks tend to ignore the difference information between bi-temporal images or perform some simple subtraction operations. The complex background and high similarity make the effects of changed information and unchanged information inaccurate, including a lot of false change information;

[0007] (2) The network will lose the detailed information in the original image during downsampling and lacks consideration of different scale sizes of changed targets;

[0008] (3) Feature redundancy is likely to occur during the decoding process, reducing the accuracy of change detection.

[0009] To solve the above problems, the present invention proposes a change detection network DEHMD-Net based on differential-enhancement and hybrid multi-scale downsampling (DEHMD). First, a dual-stage differential enhancement (DE) module is proposed, which uses channel and spatial attention mechanisms in two stages to perform feature enhancement and differential feature enhancement respectively, strengthening the learning of the changed regions and suppressing pseudo-change information. Secondly, a hybrid multi-scale strategy and four different scale sizes are used for downsampling operations to construct a hybrid multi-scale downsampling (HMSD) module, thereby suppressing the loss of detailed information. Finally, a simple and effective hierarchical interaction fusion (HIF) module is designed to introduce shallow detailed information into deep features and embed the semantic information of deep features into shallow features, aiming to enhance the feature expression ability. Summary of the Invention

[0010] The object of the present invention is to solve the technical problems existing in the existing remote sensing image change detection methods, including a lot of pseudo-change information and information loss caused by the downsampling step, which leads to a reduction in the accuracy of remote sensing change detection, and thus the present invention is proposed.

[0011] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0012] A remote sensing change detection method based on differential-enhancement and hybrid multi-scale downsampling, comprising the following steps:

[0013] Step 1: Input the dual-time images, including image T0 and image T1, into the dual-branch encoder respectively, and output three output feature maps of five levels of the encoder; the three outputs are: two enhanced feature map outputs o i 1 and o i 2 (i = 1, 2, 3, 4, 5), and one differential enhanced feature map o i 3 (i = 1, 2, 3, 4, 5);

[0014] Step 2: Input the upsampled feature maps of the third output feature map o4 3 of the fourth level and the third output o5 3 of the fifth level in Step 1 into the first level of the decoder to obtain a fused change feature prediction map;

[0015] Step 3: Input the feature map after upsampling the fused feature map in Step 2 and the third output feature map of the third level in Step 1 into the second level of the decoder to obtain a fused change feature prediction map;

[0016] Step 4: Input the feature map after upsampling the fused feature map in Step 3 and the third output feature map of the second level in Step 1 into the third level of the decoder to obtain a fused change feature prediction map;

[0017] Step 5: Input the feature map after upsampling the fused feature map in Step 4 and the third output feature map of the first level in Step 1 into the fourth level of the decoder to obtain the final fused change detection feature map;

[0018] Finally, the above 5 steps can detect remote sensing changes well.

[0019] Among them, the output end of the dual-branch encoder is connected to the input end of the single decoder; the dual-time images are input to the dual-branch encoder, and the output end of the single decoder is used to output the final image.

[0020] The structure of the upper-branch encoder is as follows: The input of the remote sensing image T0 in the upper branch enters the FEM module of the first level of the encoder. Through the FEM module, enhanced feature extraction operations are performed on the input remote sensing image, which helps to extract high-frequency local detail information and original context information. The output of the FEM of the first-level encoder is connected to the input of the DE module of the first-level encoder, and the DE module operation is performed on the first-level encoder to obtain the first output feature map o1 of the first level of the encoder 1 、the second output feature map o1 2 and the third output feature map o1 3 ;

[0021] The first-level encoder outputs the first output feature map o1 1 as the input of the HMSD module in the second-level encoder. Through the HMSD module, downsampling of the feature map is realized. The output of the HMSD module in the second-level encoder is connected to the input of the FEM module in the second-level encoder, and the FEM module operation is performed on the second-level encoder to obtain an output enhanced feature map. The output of the FEM module of the second-level encoder is connected to the input of the DE module in the second-level encoder, and the DE module operation is performed on the second-level encoder to obtain the first output feature map o2 of the second level of the encoder 1 、the second output feature map o2 2 and the third output feature map o2 3 ;

[0022] The second-level encoder outputs the first output feature map o2 1As the input of the HMSD module in the third-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the third-level encoder is connected to the input of the FEM module in the third-level encoder. The FEM module in the third-level encoder is operated on to obtain an output enhanced feature map. The output of the FEM module of the third-level encoder is connected to the input of the DE module in the third-level encoder. The DE module of the third-level encoder is operated on to obtain the first output feature map o3 of the third level of the encoder 1 , the second output feature map o3 2 and the third output feature map o3 3 ;

[0023] The third level of the encoder outputs the first output feature map o3 1 As the input of the HMSD module in the fourth-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the fourth-level encoder is connected to the input of the FEM module in the fourth-level encoder. The FEM module in the fourth-level encoder is operated on to obtain an output enhanced feature map. The output of the FEM module of the fourth-level encoder is connected to the input of the DE module in the fourth-level encoder. The DE module of the fourth-level encoder is operated on to obtain the first output feature map o4 of the fourth level of the encoder 1 , the second output feature map o4 2 and the third output feature map o4 3 ;

[0024] The fourth level of the encoder outputs the first output feature map o4 1 As the input of the HMSD module in the fifth-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the fifth-level encoder is connected to the input of the FEM module in the fifth-level encoder. The FEM module in the fifth-level encoder is operated on to obtain an output enhanced feature map. The output of the FEM module of the fifth-level encoder is connected to the input of the DE module in the fifth-level encoder. The DE module of the fifth-level encoder is operated on to obtain the first output feature map o5 of the fifth level of the encoder 1 , the second output feature map o5 2 and the third output feature map o5 3 .

[0025] The structure of the lower-branch encoder is as follows: The input of the remote sensing image T1 of the lower branch enters the FEM module of the first layer of the encoder. Through the FEM module, enhanced feature extraction operations are performed on the input remote sensing image, which helps to extract high-frequency local detail information and original context information. The output of the FEM of the first-layer encoder is connected to the input of the DE module of the first-layer encoder, and the DE module operation is performed on the first-layer encoder to obtain the first output feature map o1 of the first layer of the encoder 1 the second output feature map o1 2 and the third output feature map o1 3 ;

[0026] The first layer of the encoder outputs the second output feature map o2 1 as the input of the HMSD module in the second layer of the encoder. Through the HMSD module, downsampling of the feature map is achieved. The output of the HMSD module in the second layer of the encoder is connected to the input of the FEM module in the second layer of the encoder, and the FEM module operation is performed on the second layer of the encoder to obtain the output enhanced feature map. The output of the FEM module of the second layer of the encoder is connected to the input of the DE module in the second layer of the encoder, and the DE module operation is performed on the second layer of the encoder to obtain the first output feature map o2 of the second layer of the encoder 1 the second output feature map o2 2 and the third output feature map o2 3 ;

[0027] The second layer of the encoder outputs the second output feature map o2 2 as the input of the HMSD module in the third layer of the encoder. Through the HMSD module, downsampling of the feature map is achieved. The output of the HMSD module in the third layer of the encoder is connected to the input of the FEM module in the third layer of the encoder, and the FEM module operation is performed on the third layer of the encoder to obtain the output enhanced feature map. The output of the FEM module of the third layer of the encoder is connected to the input of the DE module in the third layer of the encoder, and the DE module operation is performed on the third layer of the encoder to obtain the first output feature map o3 of the third layer of the encoder 1 the second output feature map o3 2 and the third output feature map o3 3 ;

[0028] The third layer of the encoder outputs the second output feature map o3 2As the input of the HMSD module in the fourth-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the fourth-level encoder is connected to the input of the FEM module in the fourth-level encoder. The FEM module in the fourth-level encoder is operated to obtain an output enhanced feature map. The output of the FEM module of the fourth-level encoder is connected to the input of the DE module in the fourth-level encoder. The DE module of the fourth-level encoder is operated to obtain the first output feature map o4 of the fourth level of the encoder 1 and the second output feature map o4 2 and the third output feature map o4 3 ;

[0029] The fourth level of the encoder outputs the second output feature map o4 2 As the input of the HMSD module in the fifth-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the fifth-level encoder is connected to the input of the FEM module in the fifth-level encoder. The FEM module in the fifth-level encoder is operated to obtain an output enhanced feature map. The output of the FEM module of the fifth-level encoder is connected to the input of the DE module in the fifth-level encoder. The DE module of the fifth-level encoder is operated to obtain the first output feature map o5 of the fifth level of the encoder 1 and the second output feature map o5 2 and the third output feature map o5 3 .

[0030] The specific single decoder is as follows: The third output feature map o4 of the DE module of the fourth level of the encoder 3 and the third output feature map o5 of the fifth-level DE 3 After the upsampled feature maps are used as the two inputs of the HIF module in the first level of the decoder, the HIF module performs feature fusion operations between different levels on the two input features; the output of the HIF module of the first-level decoder is connected to the input of the CFEM module of the first-level decoder. The CFEM module of the first-level decoder is operated to obtain the fused change feature prediction map of the first level of the decoder;

[0031] The feature map obtained by upsampling the output feature fusion map of the first level of the decoder and the third output feature map o3 of the DE module of the third level of the encoder 3 are used as the inputs of the HIF module in the second level of the decoder. The HIF module performs feature fusion operations between different levels on the two input features; the output of the HIF module of the second-level decoder is connected to the input of the CFEM module of the second-level decoder. The CFEM module of the second-level decoder is operated to obtain the fused change feature prediction map of the second level of the decoder;

[0032] The feature map after upsampling the output feature fusion map of the second layer of the decoder and the third output feature map o2 of the DE module in the second layer of the encoder 3 As the input of the HIF module in the third layer of the decoder, the HIF module performs feature fusion operations between different levels on the two input features; the output of the HIF module in the third layer of the decoder is connected to the input of the CFEM module in the third layer of the decoder, and the CFEM module operation is performed on the third layer of the decoder to obtain the fusion change feature prediction map of the third layer of the decoder;

[0033] The feature map after upsampling the output feature fusion map of the third layer of the decoder and the third output feature map o1 of the DE module in the first layer of the encoder 3 As the input of the HIF module in the fourth layer of the decoder, the HIF module performs feature fusion operations between different levels on the two input features; the output of the HIF module in the fourth layer of the decoder is connected to the input of the CFEM module in the fourth layer of the decoder, and the CFEM module operation is performed on the fourth layer of the decoder to obtain the final fusion change detection feature map of the fourth layer of the decoder.

[0034] When the DE module works, the following steps are adopted:

[0035] Step S1: Use the channel attention mechanism to capture the difference relationship between channels;

[0036] Step S2: Focus on the key area through the spatial attention mechanism, accurately identify local changes, and highlight the change area.

[0037] In step S1, the following sub-steps are specifically included:

[0038] Step 1-1) Subtract the feature maps of the two-phase images to obtain the preliminary difference features; for the two inputs f i 1 and f i 2 (i = 1, 2, 3, 4, 5); first subtract to obtain the preliminary difference feature f Di ;

[0039] f Di = f i 1 - f i 2 , i = 1, 2, 3, 4, 5(1)

[0040] Step 1-2) Perform global max pooling and global average pooling operations on the preliminary difference features respectively, and splice the two generated feature maps;

[0041] S i= σ(MLP(Concat(GMP(f Di ), GAP(f Di )))) i = 1, 2, 3, 4, 5 (2)

[0042] Step 1 - 3) Pass the result to a multi - layer perceptron for further processing;

[0043] Step 1 - 4) Input the result processed in the previous step into the Sigmoid activation function to generate the attention map S i , and multiply each element of the input feature map with the split attention weights respectively to generate two enhanced feature maps and The enhanced features will continue to be passed to the subsequent encoding module.

[0044]

[0045] Where: σ represents the Sigmoid function; MLP represents the multi - layer perceptron; Concat represents concatenation; GMP and GAP represent global max pooling and global average pooling respectively; Split represents the split operation, which divides the 1×1×2C weight map into two 1×1×C weights.

[0046] In step S2, it specifically includes the following sub - steps:

[0047] Step 2 - 1) Subtract the two enhanced features to obtain a further enhanced difference feature map d i ;

[0048]

[0049] Step 2 - 2) Perform average pooling, max pooling, and 1×1 convolution operations on the enhanced difference feature map;

[0050] S′ i = σ(Conv5(Concat(Avg(d i ), Conv1(d i ), Max(d i )))), i = 1, 2, 3, 4, 5 (6)

[0051] Step 2 - 3) Pass the result of the previous step through a 5×5 convolution and Sigmoid activation to extract the spatial information weights of the difference features, multiply this weight with the newly generated difference feature map to obtain an enhanced difference feature map, which is an important input for the decoding stage.

[0052]

[0053] Where: Conv1 represents a 1×1 convolution; Conv5 represents a 5×5 convolution; Avg represents average pooling; Max represents max pooling.

[0054] A method for using a two-stage difference enhancement module, comprising the following steps:

[0055] Step 1: Use the channel attention mechanism to capture the difference relationship between channels;

[0056] Step 1 specifically includes the following sub-steps:

[0057] Step 1-1) Perform a subtraction operation on the feature maps of two periods of images to obtain preliminary difference features; for the two inputs f of the dual-temporal image features i 1 and f i 2 (i = 1, 2, 3, 4, 5); First subtract to obtain the preliminary difference feature f Di ;

[0058] f Di = f i 1 - f i 2 , i = 1, 2, 3, 4, 5 (8)

[0059] Step 1-2) Perform global max pooling and global average pooling operations on the preliminary difference features respectively, and splice the two generated feature maps;

[0060] S i = σ(MLP(Concat(GMP(f Di ), GAP(f Di )))) i = 1, 2, 3, 4, 5 (9)

[0061] Step 1-3) Pass the result to a multi-layer perceptron for further processing;

[0062] Step 1-4) Input the result processed in the previous step into the Sigmoid activation function to generate the attention map S i , and multiply each element of the input feature map with the split attention weights respectively to generate two enhanced feature maps and The enhanced features will continue to be passed to the subsequent encoding module.

[0063]

[0064] Where: σ represents the Sigmoid function; MLP represents the multi-layer perceptron; Concat represents concatenation; GMP and GAP represent global max pooling and global average pooling respectively; Split represents the splitting operation that divides the weight map of 1×1×2C into two weight maps of 1×1×C.

[0065] Step 2: Focus on the key areas through the spatial attention mechanism, accurately identify local changes, and highlight the changed areas;

[0066] Step 2 specifically includes the following sub-steps:

[0067] Step 2-1) Subtract the two enhanced features to obtain a further enhanced difference feature map d i ;

[0068]

[0069] Step 2-2) Perform average pooling, max pooling, and 1×1 convolution operations on the enhanced difference feature map;

[0070] S i ′ = σ(Conv5(Concat(Avg(d i ), Conv1(d i ), Max(d i ))), i = 1, 2, 3, 4, 5 (13)

[0071] Step 2-3) Pass the result of the previous step through a 5×5 convolution and Sigmoid activation to extract the spatial information weight of the difference feature, multiply this weight by the newly generated difference feature map to obtain an enhanced difference feature map, which serves as an important input for the decoding stage.

[0072]

[0073] Where: Conv1 represents 1×1 convolution; Conv5 represents 5×5 convolution; Avg represents average pooling; Max represents max pooling.

[0074] When the HMSD module works, the following steps are adopted:

[0075] Step 1: Extract feature information using a hybrid multi-scale strategy;

[0076] Step 2: During the feature fusion process, the feature maps generated by each branch are concatenated after global pooling and average pooling in the channel dimension to form a fused feature map.

[0077] Step 3: The multi-scale spatial attention mechanism interaction stage.

[0078] In Step 1, specifically:

[0079] Step 1-1) The first branch performs a downsampling operation using max pooling with a kernel size of 2 and a stride of 2;

[0080] Step 1-2) The second branch performs a downsampling operation using a convolution with a stride of 2;

[0081] Step 1-3) The third branch's downsampling operation combines a convolution with a stride of 2, as well as 5×1 and 1×5 bar-shaped convolutions with a stride of 1;

[0082] Step 1-4) The fourth branch's downsampling operation uses a convolution with a stride of 2, supplemented by 3×1 and 1×3 bar-shaped convolutions with a stride of 1;

[0083] In step 3, specifically:

[0084] Step 3-1) Input the feature map output from step 2 into a convolutional layer and a Sigmoid activation function to generate four different spatial weights;

[0085] Step 3-2) Multiply each of the four spatial weights from the previous step element-wise with the corresponding feature map from step 1, and concatenate them into a comprehensive result feature map.

[0086] A method for using a hybrid multi-scale downsampling module, including the following steps:

[0087] Step 1: Extract feature information using a hybrid multi-scale strategy;

[0088] Step 1-1) The first branch performs a downsampling operation using 2×2 max pooling with a stride of 2;

[0089] Step 1-2) The second branch performs a downsampling operation using a convolution with a stride of 2;

[0090] Step 1-3) The third branch's downsampling operation combines a convolution with a stride of 2, as well as 5×1 and 1×5 bar-shaped convolutions with a stride of 1;

[0091] Step 1-4) The fourth branch's downsampling operation uses a convolution with a stride of 2, supplemented by 3×1 and 1×3 bar-shaped convolutions with a stride of 1.

[0092] Step 2: During the feature fusion process, the feature maps generated by each branch are globally pooled and average pooled along the channel dimension and then concatenated to form a fused feature map.

[0093] Step 3: Multi-scale spatial attention mechanism interaction stage.

[0094] Step 3-1) Input the feature map output from step 2 into a convolutional layer and a Sigmoid activation function to generate four different spatial weights;

[0095] Step 3-2) Multiply the four spatial weights from the previous step element-wise with the feature maps corresponding to Step 1, and concatenate them into a comprehensive resultant feature map.

[0096] When the Hierarchical Interaction Fusion Module HIF works, it is specifically as follows:

[0097] Step 1: Perform 1×1 convolution operations on both inputs to obtain two output features;

[0098] Step 2: Add the output features from the previous step;

[0099] Step 3: Input the added result into a Multi-Layer Perceptron MLP, and generate weights through the Sigmoid activation function;

[0100] Step 4: Multiply the generated weights element-wise with the two feature maps that have undergone 1×1 convolution;

[0101] Step 5: Add the two feature maps from the previous step again to generate a fused feature output.

[0102] Compared with the prior art, the present invention has the following technical effects:

[0103] 1) The present invention proposes a change detection network based on differential reinforcement and hybrid multi-scale downsampling to solve the problems that the coding stage of remote sensing image change detection usually contains a lot of unchanged information and the downsampling step will cause loss of detailed information;

[0104] 2) A Dual-stage differential enhancement (DE) module proposed by the present invention uses channel and spatial attention mechanisms in two stages respectively for feature enhancement and differential feature enhancement, which can strengthen the learning of the changed areas and suppress pseudo-change information;

[0105] 3) The present invention adopts a hybrid multi-scale strategy and four different scale sizes for downsampling operations to construct a Hybrid multi-scale downsampling (HMSD) module, thereby suppressing the loss of detailed information;

[0106] 4) The present invention designs a simple and effective Hierarchical interaction fusion (HIF) module, which introduces shallow detailed information into deep features and embeds deep feature semantic information into shallow features, aiming to enhance the feature expression ability. Description of the Drawings

[0107] The following further illustrates the present invention in conjunction with the drawings and embodiments:

[0108] Figure 1 Schematic diagram of the network framework of DEHMD-Net proposed in the present invention;

[0109] Figure 2 Schematic diagram of the structure of the dual-stage difference enhancement module DE in the present invention;

[0110] Figure 3 Schematic diagram of the structure of the hybrid multi-scale downsampling module HMSD in the present invention;

[0111] Figure 4 Schematic diagram of the structure of the hierarchical interaction fusion module HIF in the present invention;

[0112] Figure 5 Change detection diagrams of different methods on the WHU dataset in the embodiments of the present invention;

[0113] Figure 6 Change detection diagrams of different methods on the Google dataset in the embodiments of the present invention;

[0114] Figure 7 Change detection diagrams of different methods on the LEVIR dataset in the embodiments of the present invention. Detailed implementation manners

[0115] A remote sensing change detection method based on difference enhancement and hybrid multi-scale downsampling, comprising the following steps:

[0116] Step 1: Input the dual-time images, including image T0 and image T1, into the dual-branch encoder respectively, and output three output feature maps of five levels of the encoder; two of the three outputs are enhanced feature maps and one is a difference-enhanced feature map;

[0117] Step 2: Input the third output feature map of the fourth level in Step 1 and the upsampled feature map of the third output of the fifth level into the first level of the decoder to obtain a fused feature map;

[0118] Step 3: Input the upsampled feature map of the fused feature map in Step 2 and the third output feature map of the third level in Step 1 into the second level of the decoder to obtain a fused feature map;

[0119] Step 4: Input the upsampled feature map of the fused feature map in Step 3 and the third output feature map of the second level in Step 1 into the third level of the decoder to obtain a fused feature map;

[0120] Step 5: Input the upsampled feature map of the fused feature map in Step 4 and the third output feature map of the first level in Step 1 into the fourth level of the decoder to obtain the final fused change detection feature map;

[0121] Finally, the above five steps can detect remote sensing changes well.

[0122] Perform change detection on remote sensing images through the above steps;

[0123] In step 1, the encoder is composed of five levels based on the FEM module, DE module, and HMSD module. The first level is composed of the FEM module and the DE module, and the following four levels are all composed of the FEM module, the DE module, and the HMSD module.

[0124] The network framework of the present invention is as Figure 1 shown, and consists of a dual-branch encoder-single decoder. The encoding part consists of a Feature enhancement module (FEM), a two-stage difference module, and a hybrid multi-scale downsampling module. A pair of dual-temporal images first enter the FEM for feature extraction operations. AttnConv is introduced in the FEM module to effectively extract high-frequency local detail information and original context information, which helps to make up for the defects of CNN and fully extract features. After the convolution operation, it enters the two-stage difference enhancement module to enhance the features of the changed objects. Then, a downsampling operation is performed, and the hybrid multi-scale downsampling module is used to effectively retain the change information and detail information of targets of different scales. Finally, the decoding stage consists of a hierarchical interaction fusion module and a Convolution feature extraction module (CFEM). HIF uses multi-layer perceptron operations and enhanced difference features at different levels for skip connections, aiming to solve the disadvantages of feature redundancy and boundary blur easily caused by insufficient ordinary cascade fusion. The specific details of the two-stage difference enhancement module, the hybrid multi-scale downsampling module, and the hierarchical interaction fusion module are introduced in detail below.

[0125] As Figure 1 shown, the network framework of the present invention includes a dual-branch encoder and a single decoder. The output end of the dual-branch encoder is connected to the input end of the single decoder; the dual-temporal image is input to the dual-branch encoder, and the output end of the single decoder is used to output the final image;

[0126] The dual-branch encoder is divided into an upper-branch encoder and a lower-branch encoder;

[0127] The structure of the upper-branch encoder is as follows: The remote sensing image T0 of the upper branch enters the input of the FEM module 1 of the first layer of the encoder. The FEM module 1 performs enhanced feature extraction operations on the input remote sensing image, which helps to extract high-frequency local detail information and original context information. The output of the FEM1 of the first-layer encoder is connected to the input of the DE module 2 of the first-layer encoder, and the DE module 2 operation is performed on the first-layer encoder to obtain the first output feature map o1 of the first layer of the encoder. 1 and the second output feature map o1 2 and the third output feature map o1 3 ;

[0128] The first-layer encoder outputs the first output feature map o1 1 as the input of the HMSD module 4-1 in the second-layer encoder. The HMSD module 4-1 is used to perform downsampling of the feature map. The output of the HMSD module 4-1 in the second-layer encoder is connected to the input of the FEM module 4-2 in the second-layer encoder, and the FEM module 4-2 operation is performed on the second-layer encoder to obtain the output enhanced feature map. The output of the FEM module 4-2 of the second-layer encoder is connected to the input of the DE module 5 in the second-layer encoder, and the DE module 5 operation is performed on the second-layer encoder to obtain the first output feature map o2 of the second layer of the encoder. 1 and the second output feature map o2 2 and the third output feature map o2 3 ;

[0129] The second-layer encoder outputs the first output feature map o2 1 as the input of the HMSD module 7-1 in the third-layer encoder. The HMSD module 7-1 is used to perform downsampling of the feature map. The output of the HMSD module 7-1 in the third-layer encoder is connected to the input of the FEM module 7-2 in the third-layer encoder, and the FEM module 7-2 operation is performed on the third-layer encoder to obtain the output enhanced feature map. The output of the FEM module 7-2 of the third-layer encoder is connected to the input of the DE module 8 in the third-layer encoder, and the DE module 8 operation is performed on the third-layer encoder to obtain the first output feature map o3 of the third layer of the encoder. 1 and the second output feature map o3 2 and the third output feature map o3 3 ;

[0130] The third-layer encoder outputs the first output feature map o3 1As the input of the HMSD module 10-1 in the fourth-level encoder, the downsampling of the feature map is achieved through the HMSD module 10-1. The output of the HMSD module 10-1 in the fourth-level encoder is connected to the input of the FEM module 10-2 in the fourth-level encoder. The FEM module 10-2 in the fourth-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module 10-2 in the fourth-level encoder is connected to the input of the DE module 11 in the fourth-level encoder. The fourth-level encoder is operated on the DE module 11 to obtain the first output feature map o4 of the fourth level of the encoder 1 , the second output feature map o4 2 and the third output feature map o4 3 ;

[0131] The fourth level of the encoder outputs the first output feature map o4 1 As the input of the HMSD module 13-1 in the fifth-level encoder, the downsampling of the feature map is achieved through the HMSD module 13-1. The output of the HMSD module 13-1 in the fifth-level encoder is connected to the input of the FEM module 13-2 in the fifth-level encoder. The FEM module 13-2 in the fifth-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module 13-2 in the fifth-level encoder is connected to the input of the DE module 14 in the fifth-level encoder. The fifth-level encoder is operated on the DE module 14 to obtain the first output feature map o5 of the fifth level of the encoder 1 , the second output feature map o5 2 and the third output feature map o5 3 ;

[0132] The structure of the lower-branch encoder is as follows: In the lower-branch remote sensing image T1, it enters the input of the FEM module 3 in the first level of the encoder. Through the FEM module 3, the input remote sensing image is subjected to enhanced feature extraction operations, which helps to extract high-frequency local detail information and original context information. The output of the FEM3 in the first level of the encoder is connected to the input of the DE module 2 in the first level of the encoder. The first level of the encoder is operated on the DE module 2 to obtain the first output feature map o1 of the first level of the encoder 1 , the second output feature map o1 2 and the third output feature map o1 3 ;

[0133] The first level of the encoder outputs the second output feature map o2 1As the input of the HMSD module 6-1 in the second-level encoder, the downsampling of the feature map is achieved through the HMSD module 6-1. The output of the HMSD module 6-1 in the second-level encoder is connected to the input of the FEM module 6-2 in the second-level encoder. The FEM module 6-2 in the second-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module 6-2 in the second-level encoder is connected to the input of the DE module 5 in the second-level encoder. The second-level encoder is operated on by the DE module 5 to obtain the first output feature map o2 of the second level of the encoder 1 and the second output feature map o2 2 and the third output feature map o2 3 ;

[0134] The second-level encoder outputs the second output feature map o2 2 As the input of the HMSD module 9-1 in the third-level encoder, the downsampling of the feature map is achieved through the HMSD module 9-1. The output of the HMSD module 9-1 in the third-level encoder is connected to the input of the FEM module 9-2 in the third-level encoder. The FEM module 9-2 in the third-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module 9-2 in the third-level encoder is connected to the input of the DE module 8 in the third-level encoder. The third-level encoder is operated on by the DE module 8 to obtain the first output feature map o3 of the third level of the encoder 1 and the second output feature map o3 2 and the third output feature map o3 3 ;

[0135] The third-level encoder outputs the second output feature map o3 2 As the input of the HMSD module 12-1 in the fourth-level encoder, the downsampling of the feature map is achieved through the HMSD module 12-1. The output of the HMSD module 12-1 in the fourth-level encoder is connected to the input of the FEM module 12-2 in the fourth-level encoder. The FEM module 12-2 in the fourth-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module 12-2 in the fourth-level encoder is connected to the input of the DE module 11 in the fourth-level encoder. The fourth-level encoder is operated on by the DE module 11 to obtain the first output feature map o4 of the fourth level of the encoder 1 and the second output feature map o4 2 and the third output feature map o4 3 ;

[0136] The fourth-level encoder outputs the second output feature map o4 2As the input to the HMSD module 15-1 in the fifth-level encoder, the downsampling of the feature map is achieved through the HMSD module 15-1. The output of the HMSD module 15-1 in the fifth-level encoder is connected to the input of the FEM module 15-2 in the fifth-level encoder. By operating on the FEM module 15-2 in the fifth-level encoder, an output enhanced feature map is obtained. The output of the FEM module 15-2 in the fifth-level encoder is connected to the input of the DE module 14 in the fifth-level encoder. By operating on the DE module 14 in the fifth-level encoder, the first output feature map o5 of the fifth level of the encoder is obtained. 1 The second output feature map o5 2 and the third output feature map o5 3 ;

[0137] The structure of the single decoder is as follows: the third output feature map o4 of the DE11 module in the fourth level of the encoder 3 and the third output feature map o5 of the DE14 in the fifth level 3 After being upsampled, the feature maps are used as the two inputs of the HIF module 16-1 in the first level of the decoder. Through the HIF module 16-1, the two input features are subjected to feature fusion operations between different levels; the output of the HIF module 16-1 in the first level of the decoder is connected to the input of the CFEM module 16-2 in the first level of the decoder. By operating on the CFEM module 16-2 in the first level of the decoder, a fused change feature prediction map of the first level of the decoder is obtained.

[0138] The feature map after upsampling the output feature fusion map of the first level of the decoder and the third output feature map o3 of the DE8 module in the third level of the encoder 3 Are used as the inputs of the HIF module 17-1 in the second level of the decoder. Through the HIF module 17-1, the two input features are subjected to feature fusion operations between different levels; the output of the HIF module 17-1 in the second level of the decoder is connected to the input of the CFEM module 17-2 in the second level of the decoder. By operating on the CFEM module 17-2 in the second level of the decoder, a fused change feature prediction map of the second level of the decoder is obtained.

[0139] The feature map after upsampling the output feature fusion map of the second level of the decoder and the third output feature map o2 of the DE5 module in the second level of the encoder 3 Are used as the inputs of the HIF module 18-1 in the third level of the decoder. Through the HIF module 18-1, the two input features are subjected to feature fusion operations between different levels; the output of the HIF module 18-1 in the third level of the decoder is connected to the input of the CFEM module 18-2 in the third level of the decoder. By operating on the CFEM module 18-2 in the third level of the decoder, a fused change feature prediction map of the third level of the decoder is obtained.

[0140] The feature map after upsampling the output feature fusion map of the third layer of the decoder and the third output feature map o1 of the DE2 module in the first layer of the encoder 3 As the input to the HIF module 19-1 in the fourth-layer decoder, the two input features are subjected to feature fusion operations at different levels through the HIF module 19-1; the output of the HIF module 19-1 in the fourth-layer decoder is connected to the input of the CFEM module 19-2 in the fourth-layer decoder, and the CFEM module 19-2 operation is performed on the fourth-layer decoder to obtain the fusion change feature prediction map of the fourth layer of the decoder;

[0141] Among them, when the DE module works, the following steps are adopted:

[0142] Step S1: Use the channel attention mechanism to capture the differential relationship between channels;

[0143] Step S2: Focus on the key area through the spatial attention mechanism, accurately identify local changes, and highlight the changed area.

[0144] In step S1, it specifically includes the following sub-steps:

[0145] Step 1-1) Subtract the feature maps of the two-phase images to obtain the preliminary differential features; for the two inputs f i 1 and f i 2 (i = 1, 2, 3, 4, 5); first subtract to obtain the preliminary differential feature f Di ;

[0146] f Di = f i 1 - f i 2 , i = 1, 2, 3, 4, 5 (15)

[0147] Step 1-2) Perform global max pooling and global average pooling operations on the preliminary differential features respectively, and splice the two generated feature maps;

[0148] S i = σ(MLP(Concat(GMP(f Di ), GAP(f Di )))) i = 1, 2, 3, 4, 5 (16)

[0149] Step 1-3) Pass the result to the multi-layer perceptron for further processing;

[0150] Step 1-4) Input the result processed in the previous step into the Sigmoid activation function to generate the attention map S i, and the split attention weights are respectively multiplied element-wise with the input feature maps to generate two enhanced feature maps o i 1 and o i 2 , and the enhanced features will continue to be passed to the subsequent encoding modules.

[0151]

[0152] In the formula: σ represents the Sigmoid function; MLP represents the multi-layer perceptron; Concat represents concatenation; GMP and GAP respectively represent global max pooling and global average pooling; Split represents the splitting operation, which divides the 1×1×2C weight map into two 1×1×C weights.

[0153] In step S2, it specifically includes the following sub-steps:

[0154] Step 2-1): Subtract the two enhanced features generated in step 1-4) to obtain a further enhanced difference feature map d i ;

[0155]

[0156] Step 2-2): Perform average pooling, max pooling, and 1×1 convolution operations on the enhanced difference feature map;

[0157] S′ i =σ(Conv5(Concat(Avg(d i ),Conv1(d i ),Max(d i ))), i = 1, 2, 3, 4, 5 (20)

[0158] Step 2-3): Pass the result of the previous step through a 5×5 convolution and Sigmoid activation to extract the spatial information weight of the difference feature, multiply this weight with the newly generated difference feature map to obtain an enhanced difference feature map, and this feature is an important input for the decoding stage.

[0159]

[0160] In the formula: Conv1 represents 1×1 convolution; Conv5 represents 5×5 convolution; Avg represents average pooling; Max represents max pooling.

[0161] When the HMSD module works, it adopts the following steps:

[0162] Step 1: Extract feature information using a hybrid multi-scale strategy;

[0163] Step 2: During the feature fusion process, the feature maps generated by each branch are subjected to global pooling and average pooling in the channel dimension and then concatenated to form a fused feature map.

[0164] Step 3: The multi-scale spatial attention mechanism interaction stage.

[0165] In Step 1, specifically:

[0166] Step 1-1) The first branch performs a 2×2 max pooling downsampling operation with a stride of 2;

[0167] Step 1-2) The second branch uses a 3×3 convolution downsampling operation with a stride of 2;

[0168] Step 1-3) The third branch downsampling operation combines a 3×3 convolution with a stride of 2, and 5×1 and 1×5 bar-shaped convolutions with a stride of 1;

[0169] Step 1-4) The fourth branch uses a 3×3 convolution with a stride of 2 for the downsampling operation, supplemented by 3×1 and 1

[0170] ×3 bar-shaped convolutions with a stride of 1;

[0171] In Step 3, specifically:

[0172] Step 3-1) Input the feature map output from Step 2 into a 3×3 convolutional layer and a Sigmoid activation function to generate four different spatial weights;

[0173] Step 3-2) Multiply each of the four spatial weights from the previous step element-wise with the corresponding feature map from Step 1 and concatenate them into a comprehensive result feature map.

[0174] When the hierarchical interaction fusion module HIF works, it is as follows:

[0175] Step 1: Perform 1×1 convolution operations on both inputs to obtain two output features;

[0176] Step 2: Add the output features from the previous step;

[0177] Step 3: Input the added result into a multi-layer perceptron MLP and generate weights through a Sigmoid activation function;

[0178] Step 4: Multiply the generated weights element-wise with the two feature maps that have undergone 1×1 convolution;

[0179] Step 5: Add the two feature maps from the previous step again to generate the fused feature output.

[0180] Such as Figure 2As shown below, the two-stage differential enhancement module (DE) is as follows:

[0181] Traditional differential calculation methods usually detect changes by directly subtracting the later image from the earlier image. However, this method is easily affected by noise, which in turn affects the accurate identification of change targets and generates a lot of false change information. To solve this problem, the present invention proposes a two-stage differential enhancement module (as Figure 2 shown). The module aims to suppress redundant information, reduce noise interference, and enhance the expression ability of differential features, thereby effectively improving the detection accuracy of change targets. Through a two-stage design that takes into account both channels and space, the module can effectively enhance change information while suppressing the interference of false change information.

[0182] To solve the above technical problems, the following technical solutions are provided:

[0183] Step 1: Use the channel attention mechanism to capture the differential relationship between channels.

[0184] Step 1 specifically includes the following sub-steps:

[0185] Step 1-1) Perform a subtraction operation on the feature maps of two periods of images to obtain preliminary differential features; for the two inputs f i 1 and f i 2 (i = 1, 2, 3, 4, 5); first subtract to obtain the preliminary differential feature f Di ;

[0186] f Di = f i 1 - f i 2 , i = 1, 2, 3, 4, 5 (22)

[0187] Step 1-2) Perform global max pooling and global average pooling operations on the preliminary differential features respectively, and splice the two generated feature maps;

[0188] S i = σ(MLP(Concat(GMP(f Di ), GAP(f Di )))) i = 1, 2, 3, 4, 5 (23)

[0189] Step 1-3) Pass the result to a multi-layer perceptron for further processing;

[0190] Step 1-4) Input the result processed in the previous step into the Sigmoid activation function to generate the attention map S i, and use the split attention weights to multiply with the input feature map element by element to generate two enhanced feature maps and The enhanced features will continue to be passed to the subsequent encoding modules.

[0191]

[0192] In the formula: σ represents the Sigmoid function; MLP represents the multi-layer perceptron; Concat represents concatenation; GMP and GAP respectively represent global max pooling and global average pooling; Split represents the split operation, which divides the 1×1×2C weight map into two 1×1×C weights.

[0193] Step 2: Focus on the key areas through the spatial attention mechanism, accurately identify local changes, and highlight the changed areas.

[0194] Step 2 specifically includes the following sub-steps:

[0195] Step 2-1) Subtract the two enhanced features to obtain a further enhanced difference feature map d i ;

[0196]

[0197] Step 2-2) Perform average pooling, max pooling, and 1×1 convolution operations on the enhanced difference feature map;

[0198] S i ′ = σConv5(Concat(Avg(d i ), Conv1(d i ), Max(d i ))), i = 1, 2, 3, 4, 5 (27)

[0199] Step 2-3) Pass the result of the previous step through a 5×5 convolution and Sigmoid activation to extract the spatial information weight of the difference feature, multiply this weight with the newly generated difference feature map to obtain an enhanced difference feature map, and this feature is an important input for the decoding stage.

[0200]

[0201] In the formula: Conv1 represents 1×1 convolution; Conv5 represents 5×5 convolution; Avg represents average pooling; Max represents max pooling.

[0202] As Figure 3 shown, regarding the Hybrid Multi-Scale Downsampling Module (HMSD), it is as follows:

[0203] Traditional downsampling, due to its fixed receptive field and information loss, cannot flexibly capture the spatial and channel features of multi-scale objects, limiting the model's recognition ability and generalization performance. Although some studies have proposed multi-scale strategies to improve the feature extraction effect, these methods often fail to effectively integrate information at different scales, resulting in insufficient performance when dealing with complex scenarios. To solve this problem, the present invention proposes a hybrid multi-scale downsampling module. The HMSD module uses a hybrid multi-scale approach to enhance the diversity and richness of feature representation, covering receptive fields of different sizes. At the same time, the multi-scale processing method ensures the consideration of objects of different sizes, improves the network's adaptability to various spatial resolutions, and reduces information loss. Meanwhile, the HMSD module introduces a spatial attention mechanism for spatial interaction, not only achieving effective fusion of different scales but also improving the accuracy of edge detection and the integrity of feature extraction.

[0204] To solve the above technical problems, the following technical solutions are provided:

[0205] Step 1: Extract feature information using a hybrid multi-scale strategy.

[0206] Step 1-1) The first branch uses a 2×2 max pooling operation with a stride of 2 for downsampling;

[0207] Step 1-2) The second branch uses a 3×3 convolution operation with a stride of 2 for downsampling;

[0208] Step 1-3) The third branch's downsampling operation combines a 3×3 convolution with a stride of 2, and 5×1 and 1×5 strip convolutions with a stride of 1;

[0209] Step 1-4) The fourth branch's downsampling operation uses a 3×3 convolution with a stride of 2, supplemented by 3×1 and 1

[0210] ×3 strip convolutions with a stride of 1.

[0211] Step 2: During the feature fusion process, the feature maps generated by each branch are concatenated after global pooling and average pooling in the channel dimension to form a fused feature map.

[0212] Step 3: Multi-scale spatial attention mechanism interaction stage.

[0213] Step 3-1) Input the feature map output from Step 2 into a 3×3 convolutional layer and a Sigmoid activation function to generate four different spatial weights;

[0214] Step 3-2) Multiply each of the four spatial weights from the previous step element-wise with the corresponding feature map from Step 1, and concatenate them into a comprehensive result feature map.

[0215] Such as Figure 4As shown below, the Hierarchical Interaction Fusion Module (HIF) is as follows:

[0216] Traditional decoding generally performs simple splicing operations, which can lead to a large semantic gap between low-level details and high-level semantics and insufficient boundary information. Therefore, the present invention designs a simple and effective hierarchical interaction fusion module, and proposes a new feature fusion operation based only on a multi-layer perceptron and differential information. The multi-layer perceptron can reduce the semantic gap between features at different levels, avoid the accumulation of redundant information, and reduce the risk of overfitting. Specifically, the multi-layer perceptron can learn more complex and rich feature relationships through multi-layer non-linear transformations, which helps the effective fusion and transformation of low-level features and high-level features, and improves the consistency of feature representation.

[0217] To solve the above technical problems, the following technical solutions are provided:

[0218] Step 1: Perform 1×1 convolution operations on both inputs to obtain two output features;

[0219] Step 2: Add the output features from the previous step;

[0220] Step 3: Input the result of the addition into a multi-layer perceptron (MLP) and generate weights through the Sigmoid activation function;

[0221] Step 4: Multiply the generated weights element-wise with the two feature maps after 1×1 convolution;

[0222] Step 5: Add the two feature maps from the previous step again to generate a fused feature output.

[0223] Example:

[0224] 1) Experimental data and experimental design

[0225] Use three publicly available change detection datasets to test the performance of the proposed change detection network DEHMD-Net, namely WHU data, Google data, and LEVIR data for experiments. ① WHU data contains a pair of remote sensing images of 32507×15354 pixels, with a resolution of 0.2m. ② Google data contains 19 pairs of remote sensing images with sizes ranging from 1006×1168 to 4936×5224 pixels, with a resolution of 0.55m. The image resolution is high and the details are rich. ③ LEVIR data contains 637 pairs of remote sensing images of 1024×1024 pixels, with a resolution of 0.5m, mainly used for change detection and target detection tasks. All three of their datasets focus on building changes.

[0226] Due to the limitation of the CPU, the images in the experimental data are cropped into non-overlapping image patches of size 256×256 pixels. For the LEVIR dataset, based on the given training set, validation set, and test set partitions, the model is trained according to these partitions. For the other two datasets, since the provider did not provide a partitioning scheme, the present invention refers to the ratio in the classic change detection network BIT and randomly divides the image patches into a training set, a validation set, and a test set in a ratio of 8:1:1. To augment the training samples, data augmentation operations such as rotation, flipping, and scaling are performed on the training set.

[0227] To verify the effectiveness and characteristics of the method DEHMD-Net network of the present invention, the following comparative experiments are organized: (1) Compare with two classic CNN change detection networks, namely FC-EF and FC-Conc; (2) Compare with two groups of attention-improved CNN change detection networks, IFN and SNUNet; (3) Compare with the change detection networks BIT and MSCANet that integrate Transformer and CNN; (4) Compare with two newly proposed change detection networks, namely LightCDNet and STADE-CDNet.

[0228] Four commonly used accuracy metrics often used in change detection literature

[18] are used to quantitatively evaluate the change detection results, namely accuracy (precision, Pre), recall (recall, Rec), F1 value (F1), and intersection over union (IoU). For change detection tasks with unbalanced classes (changed and unchanged), F1 and IoU are more reliable than other metrics. The experiments of the present invention use the PyTorch deep learning framework and run in an environment of Geforce GTX1660 GPU (equipped with 6GB video memory). Due to the limitation of computing resources, the batch size is set to 4. To optimize the model, the stochastic gradient descent (SGD) algorithm with momentum is used to optimize the model, where the momentum parameter is 0.99 and the weight decay coefficient is 0.0005, and binary cross-entropy is selected as the loss function. The initial learning rate is 0.01, and the update strategy refers to the linear decay method in the classic paper BIT. The entire training process is carried out for 200 rounds. After each round of training, validation is performed, and the model with the best performance on the validation set is used to evaluate the test set.

[0229] 2) Experimental results and analysis

[0230] Figures 5 to 7Typical change detection maps of three different regions of WHU, Google, and LEVIR data are given respectively: (a) - (c) represent the images of the first period, the images of the second period, and the change reference map of the two periods respectively; (d) - (l) represent the change detection maps of FC-EF, FC-Conc, IFN, SNUNet, BIT, MSCANet, LightCDNet, STADE-CDNet, and the network DEHMD-Net of the present invention respectively. White and black represent the correctly detected changed regions and unchanged regions respectively; red represents false detection errors (i.e., unchanged regions are detected as changed regions); green represents missed detection errors (i.e., changed regions are detected as unchanged regions).

[0231] From Figure 5 it can be seen that for WHU data, the change detection map generated by the method DEHMD-Net of the present invention has the smallest areas of red false detection and green missed detection regions and is closest to the change reference map (i.e., the ground truth label map); while the performance of the 8 comparison methods is not ideal. For example, in Figure 6 the second row, the change detection maps of IFN, BIT, and MSCANet contain obvious green missed detections, and the other methods are not accurate enough in detecting the boundary regions of the changed objects.

[0232] From Figure 6 it can be seen that for Google data, the method DEHMD-Net of the present invention can also obtain the most accurate change detection map, and the area of its colored regions is the smallest: for example Figure 7 in the first row, FC-EF and FC-Conc contain a large number of missed detection errors, while the change detection maps of IFN, SNUNet, BIT, MSCANet, LightCDNet, and STADE-CDNet contain obvious false detection errors.

[0233] From Figure 7 the change detection results of the LEVIR data given, although all methods can effectively detect the general outline of the building changes, compared with the eight groups of comparative experiments, the building change objects detected by the method of the present invention are more complete, with the least false detections and missed detections, and the boundaries are more accurate. For example Figure 7 in the first row, both FC-EF and LightCDNet contain a large number of missed detection errors, and the other comparative experiments all have different degrees of missed detections and false detections, while the method of the present invention has few missed detection and false detection errors.

[0234] To more objectively evaluate the performance of different change detection network models, Table 1 presents the quantitative indicators of the change detection results of different methods on 3 datasets. It can be seen from Table 1 that on all datasets, the method of the present invention, DEHMD-Net, has achieved the best quantitative indicators: on the 3 datasets, the accuracy Pre and recall Rec of the method of the present invention, DEHMD-Net, either rank first or second, and the comprehensive evaluation indicators F1 and IoU values are the highest on the 3 datasets. For example, on the WHU dataset, the F1 value of the method of the present invention is 92.38%, which is at least 4.94% higher than other methods. On the Google dataset, the IoU value of the method of the present invention is 76.88%, which is at least 3.34% higher than other methods. Generally speaking, the DEHMD-Net proposed by the present invention has the best performance on the dataset, indicating that the network has strong anti-interference or noise ability, is more stable than other methods, has good robustness and generalization ability, can adapt to change detection in different complex scenarios, and demonstrates good detection ability.

[0235] Table 1 Precision Indicators of Change Detection Results for Different Network Models

[0236] Table 1Accuracy Indicators of Change Detection Results for DifferentNetwork Models

[0237]

[0238] 3) Ablation Experiments

[0239] To verify the effectiveness of DEHMD-Net, the present invention takes the WHU dataset as an example to conduct ablation experiments on the two-stage difference enhancement module, the hybrid multi-scale downsampling module, the hierarchical interaction fusion module, and AttnConv. The designs of the four groups of ablation experiments are as follows: ① Remove the two-stage difference enhancement module, delete the attention mechanisms of the two stages, and directly perform a subtraction operation on the two-phase feature maps to obtain the difference feature output; ② Remove the hybrid multi-scale downsampling module and directly use the traditional downsampling operation; ③ Remove the hierarchical interaction fusion module and directly perform a splicing operation during the skip connection; ④ Remove AttnConv and use ordinary convolution instead. Table 2 presents the change detection results of the above four groups of ablation experiments (× indicates removal, and √ indicates retention).

[0240] Table 2 Quantitative EvaluationResults of Ablation Experiments on the WHU Dataset

[0241]

[0242] Comparing the results of each row in Table 2, it can be seen that removing any module will lead to a decrease in the detection accuracy of the method of the present invention, indicating that these modules all play a role in improving the change detection accuracy of the method of the present invention. In the ablation experiment, the accuracy decreased significantly after removing the hybrid multi-scale downsampling module, hierarchical interaction fusion module, and difference enhancement module. The F1 values decreased by 1.54%, 1.49%, and 1.14% respectively, and the IoU values decreased by 2.61%, 2.52%, and 1.94% respectively. This shows that these three modules are key modules: the DE module can well highlight the change information and suppress the pseudo-change information; the HMSD module realizes the effective fusion of different-scale information and the integrity of feature extraction; the HIF module effectively enhances the edge detection ability and further highlights the change region information. After removing AttnConv, the model performance also decreased, proving that this module can also improve the model performance to a certain extent.

[0243] In summary, aiming at the characteristics of high-resolution remote sensing images, the present invention proposes a change detection network DEHMD-Net based on difference enhancement and hybrid multi-scale downsampling. The designed two-stage difference enhancement module effectively highlights the change region and suppresses the irrelevant pseudo-change region; the hybrid multi-scale downsampling module uses the hybrid multi-scale strategy operation to enhance feature diversity and reduce feature homogeneity, thereby suppressing information loss. At the same time, combined with the multi-scale attention mechanism, it can effectively retain the features of change targets of different scale sizes; the hierarchical interaction fusion module effectively avoids the accumulation of redundant information and effectively suppresses the interference of background noise;

[0244] Corresponding verification was carried out on three datasets and eight networks were selected for comparison. The final results show that the DEHMD-Net proposed by the present invention has good effects on the datasets, and the comprehensive indicators all reach the highest, effectively detecting the change region. However, DEHMD-Net still has certain limitations. The double-branch coding increases the computational cost, and later it will be considered whether the computational amount can be reduced while improving the accuracy.

Claims

1. A remote sensing change detection method based on differential reinforcement and hybrid multi-scale downsampling, characterized in that It includes the following steps: Step 1: Input the dual-time images, including image T0 and image T1, into the dual-branch encoder respectively, and output three output feature maps of the five levels of the encoder; the three outputs are: two enhanced feature map outputs o i 1 and o i 2 , and one differential enhanced feature map o i 3 ; Step 2: Input the feature map o4, which is the third output of the fourth level in Step 1 3 and the output o5, which is the third output of the fifth level 3 after upsampling into the first level of the decoder to obtain a fused variation feature prediction map; Step 3: Input the feature map after upsampling the fused feature map in Step 2 and the third output feature map of the third level in Step 1 into the second level of the decoder to obtain a fused change feature prediction map; Step 4: Input the feature map after upsampling the fused feature map in Step 3 and the third output feature map of the second level in Step 1 into the third level of the decoder to obtain a fused change feature prediction map; Step 5: Input the feature map after upsampling the fused feature map in Step 4 and the third output feature map of the first level in Step 1 into the fourth level of the decoder to obtain the final fused change detection feature map; Finally, the above 5 steps can detect remote sensing changes well.

2. The method according to claim 1, wherein Among them, The output end of the dual-branch encoder is connected to the input end of the single decoder; the dual-time images are input to the dual-branch encoder, and the output end of the single decoder is used to output the final image.

3. The method according to claim 1, wherein The structure of the upper-branch encoder is as follows: The input of the upper-branch remote sensing image T0 into the FEM module of the first layer of the encoder, through the FEM module, performs enhanced feature extraction operations on the input remote sensing image, which helps to extract high-frequency local detail information and original context information. The output of the FEM of the first-layer encoder is connected to the input of the DE module of the first-layer encoder, and the DE module operation is performed on the first-layer encoder to obtain the first output feature map o1 of the first layer of the encoder 1 and the second output feature map o1 2 and the third output feature map o1 3 ; The first - level encoder outputs the first output feature map o1 1 As the input of the HMSD module in the second - level encoder, the down - sampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the second - level encoder is connected to the input of the FEM module in the second - level encoder. By operating on the FEM module in the second - level encoder, the output enhanced feature map is obtained. The output of the FEM module in the second - level encoder is connected to the input of the DE module in the second - level encoder. By operating on the DE module in the second - level encoder, the first output feature map o2 of the second level of the encoder is obtained 1 The second output feature map o2 2 And the third output feature map o2 3 ; The second-level output of the encoder outputs the first output feature map o2 1 As the input of the HMSD module in the third-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the third-level encoder is connected to the input of the FEM module in the third-level encoder. The FEM module in the third-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module of the third-level encoder is connected to the input of the DE module in the third-level encoder. The third-level encoder is operated on the DE module to obtain the first output feature map o3 of the third level of the encoder 1 The second output feature map o3 2 And the third output feature map o3 3 ; The encoder's third-level output is the first output feature map o3 1 As the input of the HMSD module in the fourth-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the fourth-level encoder is connected to the input of the FEM module in the fourth-level encoder. The FEM module in the fourth-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module in the fourth-level encoder is connected to the input of the DE module in the fourth-level encoder. The DE module in the fourth-level encoder is operated on to obtain the first output feature map o4 of the fourth level of the encoder 1 The second output feature map o4 2 And the third output feature map o4 3 ; The encoder's fourth-level output is the first output feature map o4 1 As the input of the HMSD module in the fifth-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the fifth-level encoder is connected to the input of the FEM module in the fifth-level encoder. The FEM module in the fifth-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module of the fifth-level encoder is connected to the input of the DE module in the fifth-level encoder. The fifth-level encoder is operated on the DE module to obtain the first output feature map o5 of the fifth level of the encoder 1 and the second output feature map o5 2 and the third output feature map o5 3 .

4. The method according to claim 1, wherein The structure of the lower branch encoder is as follows: The input of the remote sensing image T1 of the lower branch enters the FEM module of the first layer of the encoder. The FEM module performs enhanced feature extraction operations on the input remote sensing image, which helps to extract high-frequency local detail information and original context information. The output of the FEM of the first layer encoder is connected to the input of the DE module of the first layer encoder, and the DE module operation is performed on the first layer encoder to obtain the first output feature map o1 of the first layer of the encoder 1 , the second output feature map o1 2 and the third output feature map o1 3 ; The second output feature map o2 of the first layer of the encoder 1 As the input of the HMSD module in the second layer encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the second layer encoder is connected to the input of the FEM module in the second layer encoder. The FEM module in the second layer encoder is operated on to obtain the output enhanced feature map. The output of the FEM module of the second layer encoder is connected to the input of the DE module in the second layer encoder. The DE module of the second layer encoder is operated on to obtain the first output feature map o2 of the second layer of the encoder 1 and the second output feature map o2 2 and the third output feature map o2 3 ; The second-level output of the encoder outputs the second output feature map o2 2 As the input of the HMSD module in the third-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the third-level encoder is connected to the input of the FEM module in the third-level encoder. The FEM module in the third-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module of the third-level encoder is connected to the input of the DE module in the third-level encoder. The third-level encoder is operated on the DE module to obtain the first output feature map o3 of the third level of the encoder 1 and the second output feature map o3 2 and the third output feature map o3 3 ; The third - level output of the encoder is the second output feature map o3 2 As the input of the HMSD module in the fourth - level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the fourth - level encoder is connected to the input of the FEM module in the fourth - level encoder. After operating on the FEM module in the fourth - level encoder, the output enhanced feature map is obtained. The output of the FEM module in the fourth - level encoder is connected to the input of the DE module in the fourth - level encoder. After operating on the DE module in the fourth - level encoder, the first output feature map o4 of the fourth - level encoder is obtained 1 and the second output feature map o4 2 and the third output feature map o4 3 ; The encoder's fourth-level output is the second output feature map o4 2 As the input of the HMSD module in the fifth-level encoder, the downsampling of the feature map is achieved through the HMSD module. The output of the HMSD module in the fifth-level encoder is connected to the input of the FEM module in the fifth-level encoder. The FEM module in the fifth-level encoder is operated on to obtain the output enhanced feature map. The output of the FEM module in the fifth-level encoder is connected to the input of the DE module in the fifth-level encoder. The fifth-level encoder is operated on the DE module to obtain the first output feature map o5 of the fifth level of the encoder 1 and the second output feature map o5 2 and the third output feature map o5 3 .

5. The method according to claim 2, wherein The single decoder specifically is: the third output feature map o4 of the DE module in the fourth layer of the encoder 3 and the third output feature map o5 of the DE in the fifth layer 3 The feature maps after upsampling are used as the two inputs of the HIF module in the first layer of the decoder. The HIF module performs feature fusion operations between different levels on the two input features; The output of the HIF module of the first-level decoder is connected to the input of the CFEM module of the first-level decoder, and the CFEM module operation is performed on the first-level decoder to obtain the fused change feature prediction map of the first level of the decoder; The feature map after upsampling the output feature fusion map of the first layer of the decoder and the third output feature map o3 of the DE module in the third layer of the encoder 3 As the input of the HIF module in the second layer decoder, the HIF module performs feature fusion operations between different levels on the two input features; The output of the HIF module of the second-level decoder is connected to the input of the CFEM module of the second-level decoder, and the CFEM module operation is performed on the second-level decoder to obtain the fused change feature prediction map of the second level of the decoder; The feature map obtained by upsampling the output feature fusion map of the second layer of the decoder and the third output feature map o2 of the DE module of the second layer of the encoder 3 As the input of the HIF module in the third-layer decoder, the HIF module performs feature fusion operations between different levels on the two input features; The output of the HIF module of the third-level decoder is connected to the input of the CFEM module of the third-level decoder, and the CFEM module operation is performed on the third-level decoder to obtain the fused change feature prediction map of the third level of the decoder; The feature map after upsampling the output feature fusion map of the third layer of the decoder and the third output feature map o1 of the DE module in the first layer of the encoder 3 As the input of the HIF module in the fourth-layer decoder, the HIF module performs feature fusion operations between different levels on the two input features; The output of the HIF module of the fourth-level decoder is connected to the input of the CFEM module of the fourth-level decoder, and the CFEM module operation is performed on the fourth-level decoder to obtain the final fused change detection feature map of the fourth level of the decoder.

6. The method according to any one of claims 1 to 5, characterized in that, When the DE module works, it adopts the following steps: Step S1: Use the channel attention mechanism to capture the difference relationship between channels; Step S2: Focus on the key areas through the spatial attention mechanism, accurately identify local changes, and highlight the changed areas.

7. The method according to claim 6, wherein In Step S1, it specifically includes the following sub-steps: Step 1-1) Subtract the feature maps of the two-phase images to obtain preliminary differential features; for the two inputs f of the dual-temporal image features i 1 and f i 2 (i = 1, 2, 3, 4, 5); First, subtract to obtain the preliminary differential feature f Di ; f Di = f i 1 -f i 2 , i = 1, 2, 3, 4, 5(1) Step 1-2) Perform global max pooling and global average pooling operations on the preliminary difference features respectively, and concatenate the two generated feature maps; S i = σ(MLP(Concat(GMP(f Di ), GAP(f Di )))) i = 1, 2, 3, 4, 5 (2) Step 1-3) Pass the result to the multi-layer perceptron for further processing; Step 1-4) Input the result processed in the previous step into the Sigmoid activation function to generate the attention map S i , and multiply the split attention weights with the input feature map element by element to generate two enhanced feature maps and The enhanced features will continue to be passed to the subsequent encoding module; In the formula: σ represents the Sigmoid function; MLP represents the multi-layer perceptron; Concat represents concatenation; GMP and GAP respectively represent global max pooling and global average pooling; Split represents the splitting operation, which divides the 1×1×2C weight map into two 1×1×C weights; In Step S2, it specifically includes the following sub-steps: Step 2-1) Subtract the two enhanced features to obtain a further enhanced difference feature map d i ; Step 2-2) Perform average pooling, max pooling, and 1×1 convolution operations on the enhanced difference feature map; S i ′ = σ(Conv5(Concat(Avg(d i ), Conv1(d i ), Max(d i ))), i = 1, 2, 3, 4, 5 (6) Step 2-3) Pass the result of the previous step through a 5×5 convolution and Sigmoid activation to extract the spatial information weight of the difference feature, multiply this weight by the newly generated difference feature map to obtain an enhanced difference feature map, and this feature is an important input in the decoding stage; Where: Conv1 represents a 1×1 convolution; Conv5 represents a 5×5 convolution; Avg represents average pooling; Max represents max pooling.

8. The method according to any one of claims 1 to 5, characterized in that, When the HMSD module works, the following steps are adopted: Step 1: Extract feature information using a hybrid multi-scale strategy; Step 2: During the feature fusion process, the feature maps generated by each branch are concatenated after global pooling and average pooling in the channel dimension to form a fused feature map; Step 3: The multi-scale spatial attention mechanism interaction stage.

9. The method according to claim 8, characterized in that In Step 1, specifically: Step 1-1) The first branch uses max pooling with a stride of 2 for downsampling; Step 1-2) The second branch uses convolution with a stride of 2 for downsampling; Step 1-3) The third branch's downsampling operation combines convolution with a stride of 2, and 5×1 and 1×5 strip convolutions with a stride of 1; Step 1-4) The fourth branch's downsampling operation uses convolution with a stride of 2, supplemented by 3×1 and 1×3 strip convolutions with a stride of 1; In Step 3, specifically: Step 3-1) Input the feature map output from Step 2 into a convolutional layer and a Sigmoid activation function to generate four different spatial weights; Step 3-2) Multiply each of the four spatial weights from the previous step element-wise with the corresponding feature map from Step 1, and concatenate them into a comprehensive result feature map.

10. The method according to any one of claims 1 to 5, characterized in that, When the hierarchical interaction fusion module HIF works, it is as follows specifically: Step 1: Perform 1×1 convolution operations on both inputs to obtain two output features; Step 2: Add the output features from the previous step; Step 3: Input the added result into a multi-layer perceptron MLP and generate weights through a Sigmoid activation function; Step 4: Multiply the generated weights element-wise with the two feature maps that have undergone 1×1 convolution; Step 5: Add the two feature maps from the previous step again to generate a fused feature output.