Deep learning-based remote sensing image change detection method
By introducing a multi-branch feature difference extraction module and an asymmetric integrated decision-making module in the remote sensing image change detection, the problems of information loss and scale deviation in the prior art are solved, and a higher accuracy and robust change detection effect is achieved.
Patent Information
- Application Number
- CN202510096071.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
The existing remote sensing image change detection methods have information loss and noise interference during data compression and change comparison, and the deep network in the decoding stage is prone to information distortion, making it difficult to effectively identify changes at different scales.
Multi-branch feature differential extraction (MFDE) module and asymmetric integrated decision-making (AED) module are used. The MFDE module extracts high-quality differential features from the dual-phase image through three strategies: differentiation, saving and fusion, and the AED module corrects the decoder's deviation on each scale by assigning different weight factors to different scales.
It significantly improves the accuracy and robustness of remote sensing image change detection, can more accurately identify complex scenes and subtle changes, and improves the overall performance of change detection.
Smart Images

Figure CN120032248A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer image processing technology, and in particular to a method for detecting remote sensing image changes based on deep learning. Background Art
[0002] Change detection technology plays an important role in many fields such as urban and coastal wetland change analysis, land management and environmental monitoring. This technology can observe and analyze the impact of human activities, climate change and other factors on the environment through various methods. Change detection mainly observes the same target at different time periods to identify its state changes. In the field of remote sensing, researchers usually use two or more remote sensing images acquired at different times from the same geographical location to identify changes in ground objects.
[0003] Due to the complexity of comparing image changes in different periods, traditional methods such as Euclidean distance and correlation coefficient often have difficulty in effectively modeling these differences. To this end, researchers have developed a number of change detection methods based on encoder and decoder structures. In this structure, the encoder first compresses the input sample data into a feature vector for calculation and processing, and then the decoder generates the final result based on the context vector.
[0004] The encoder and decoder network mainly includes two common structures: a network with multiple encoders and a single decoder, and a dual encoder and decoder network. The single decoder structure is easily disturbed by irrelevant differences in dual-phase data. For example, seasonal changes in vegetation can affect network performance, and low-level features generated by fusing encoder features may introduce noise interference; the dual encoder and decoder solves this problem by using two decoders, but there are still the following problems in the data compression and change comparison process: in the encoding stage, data compression will cause information loss, and the two sets of neural networks lack the necessary communication mechanism during the processing process. In the decoding stage, deep networks in long sequence networks are prone to information distortion. At the same time, there are obvious deviations when identifying changes at different scales. These problems will directly affect the final comparison results.
[0005] In response to these technical difficulties, researchers mainly explored two directions: feature difference extraction and receptive field research. In feature difference research, researchers usually extract feature differences at the encoder stage and directly generate the final prediction results through analysis, but the deepening of the network layer may cause information distortion. In receptive field research, researchers extract features of different dimensions by adjusting the shape or size of the receptive field kernel. These studies provide new ideas and directions for improving change detection performance. Summary of the invention
[0006] The purpose of the present invention is to provide a method for detecting remote sensing image changes based on deep learning in order to address the deficiencies of the prior art. This method can improve the accuracy and robustness of remote sensing image change detection and provide important support for the further development of remote sensing technology.
[0007] The technical solution for achieving the purpose of the present invention is:
[0008] A method for detecting changes in remote sensing images based on deep learning, comprising the following steps:
[0009] 1) Constructing datasets: Two different datasets, CLCD and NJDS, are used to verify the effectiveness of the method. The NJDS dataset collects remote sensing images of Nanjing City, including 563 pairs of 512×512 pixel image pairs, covering the period from 2016 to 2019, and mainly focusing on land use changes, including urban expansion, vegetation changes and other types of changes. The CLCD dataset is a public farmland dataset containing 2,400 pairs of farmland change samples taken by Gaofen, with an image size of 256×256 pixels. These dual-phase images were taken in 2017 and 2019 in Guangdong Province, China, and the spatial resolution of these images ranges from 0.5m to 2m. Each group of samples contains two images and a binary label representing the farmland changes. All samples are divided into training set, validation set and test set in a ratio of 6:2:2. The number of samples in the training set, validation set and test set are 1,440, 480 and 480 respectively. These two datasets have different geographical characteristics and change types, which can comprehensively evaluate the performance of the change detection algorithm.
[0010] 2) Construct a multi-branch feature difference extraction MFDE (Multibranch Feature Difference Extraction, MFDE for short) module: The MFDE module adopts a three-parallel branch network structure: a difference branch focuses on modeling the difference between two sets of features, a preservation branch retains the information lost during the encoding process, and a fusion branch fuses feature information of different scales. This design ensures the independent calculation of difference features and retains the original image information to the greatest extent. In the traditional encoder and decoder network architecture, the encoder cannot effectively capture image differences and there is information loss. Therefore, the MFDE of this technical solution fully extracts the difference features of the image from the dual-phase image. The MFDE module adopts three strategies of difference, preservation and fusion. The MFDE module intercepts features from each layer of the encoder to more accurately model the difference features and adopts multi-scale fusion to obtain high-quality difference features, which provides strong guidance for the final change detection, including:
[0011] 2-1) Difference branch: This branch takes the encoder's encoded features as input, calculates the difference matrix of the input features to capture the initial change information, and then calculates the corresponding threshold matrix. Based on the threshold matrix, the difference part of the feature expression is highlighted to obtain the feature difference. The mask module (MASK) is used to filter out non-change information from the features. The process formula is shown in (1):
[0012]
[0013] Among them, T 1 and T 2 is the feature of the dual-phase image, F d T 1 and T 2 The absolute value of the difference between two features indicates the difference between them. Avgpool indicates average pooling. MASK(x, 0) sets all elements in the matrix that are less than 0 to 0. MASK sets pixel values that are less than the difference to 0, eliminating their influence in the network training process and increasing the weight of pixels that actually contain difference information.
[0014] 2-2) Preservation branch: The preservation branch models the information lost in the difference branch. The subtraction operation in the difference branch will lead to the loss of pixel information. Therefore, a preservation branch is added to model the lost information and integrate the modeled features into the difference information to generate high-quality difference features. The process formula is shown in (2):
[0015] F P =Conv(Concat(Concat(T 1 +T 2 ),(T 1 +T 2 ))), (2),
[0016] Among them, Concat is used to connect the input vector; Conv is a 3×3 convolution; F P Represents the feature T 1 and T 2 Compared with direct addition, the input vector T is additionally linked in the channel dimension 1 and T 2 , which retains more data information and reduces the information loss in the network coding process;
[0017] 2-3) Fusion branch: Fusion of feature information of different scales to obtain difference features with complete information. Convolution and upsampling are used to match the size of deep features and shallow features. Then, convolution is used to model the fused difference features. The formula of the process is shown in (3):
[0018]
[0019] Among them, F up Record the information obtained from the first two branches of each layer of features in the encoding stage; Conv is a 3×3 convolution; F last is the difference feature obtained from the previous layer, F diff It is the difference feature obtained from the current layer. The fusion adopts a layer-by-layer progressive method to perform feature fusion: first, the first layer feature is fused with the second layer feature, and then the fusion result is fused with the third layer feature, and so on, finally realizing the hierarchical integration of the difference features of different scales in the encoding stage;
[0020] 3) Constructing an asymmetric ensemble decision (AED): The AED module constructs an independent learner based on the receptive field characteristics of different decoding stages, and corrects the deviation of the decoder at each scale by assigning different weight factors to different scales. The AED module establishes a parallel inference network for the vertical encoder group to further extract the feature information of each network layer, thereby achieving more accurate change detection: In the decoding stage, the decoder reconstructs the extracted feature vector, and uses jump connections to retain the original information of the bidirectional image in the encoding stage to make up for the detail information lost in the feature extraction process. However, the general decoding network structure ignores the difference in receptive field scale. Features of different scales tend to express features at different levels. Shallow features usually focus on color and texture, while deep features contain more pixel information and represent semantics. There is an urgent need for a method to deeply utilize the change information contained in the receptive fields of different scales. The AED module sets the corresponding confidence factor according to the difference in the receptive fields of features of different scales to measure the reliability of the prediction of the final change result and to correct the weights of features of different scales on the change. The AED module converts the original vertical linear network inference process into a parallel ensemble decision process to obtain the guiding features for change detection, including:
[0021] 3-1) Asymmetric learning: In the AED module, the change prediction map obtained from the last layer of the decoder stage has the same feature map size of 256×256 as the original input. Each pixel at this scale corresponds to each pixel of the change map, and the receptive field ratio is 1. The calculation method of the receptive field ratio at different scales is shown in formula (4):
[0022]
[0023] The calculation method of the confidence factor at different scales through the receptive field ratio is shown in formula (5):
[0024]
[0025] Among them, r i is the receptive field ratio of the i-th layer in the decoding stage; W t and H t is the size of the original input image; W ti and H ti Respectively represent the width and height of the i-th layer feature; I i represents the confidence factor of the i-th layer; d represents the number of network layers in the decoding stage. Through the Channel Avgpool Module (CAM) and the MASK module, the region expressing non-changing information in the feature is suppressed. Then, the output feature F of each layer is obtained according to different feature factors. dec ;
[0026] 3-2) Integrated decision: In the decoding stage, the confidence factor is multiplied with the feature maps at different scales to construct the feature vectors at the corresponding scales. After fusion, the final change detection prediction is obtained. The formal content of this process is shown in formula (6):
[0027]
[0028] Among them, d represents the number of layers in the network decoding stage; Conv is a 7×7 convolution; Extension repeats a single element of the feature at different scales in both width and height dimensions to achieve the same scale as the prediction result; Change refers to the change prediction result of the feature difference and receptive field network FDRFNet (Feature Difference and Receptive Field Network, referred to as FDRFNet) of the overall network architecture of this technical solution.
[0029] In order to solve the problem of insufficient encoder feature difference capture and information loss in this technical solution, a multi-branch feature difference extraction MFDE module is adopted. The MFDE module extracts difference information from the dual-phase image at each encoder layer and constructs high-quality difference features through cross-scale information fusion;
[0030] In order to solve the problem of feature expression deviation of the decoder at different scales, this technical solution adopts an asymmetric integrated decision AED module. The AED module constructs an independent learner according to the receptive field characteristics of different decoding stages, and corrects the deviation of the decoder at each scale by assigning different weight factors to different scales. The AED module establishes a parallel inference network for the vertical encoder group to further extract the feature information of each network layer, thereby achieving more accurate change detection.
[0031] The collaborative optimization of the MFDE and AED modules in this technical solution is an important innovation of this technical solution. MFDE provides high-quality difference feature expression, while AED further enhances the detection capability of edge details on this basis. The combination of the two significantly improves the detection capability of complex scenes and subtle changes.
[0032] This method can improve the accuracy and robustness of change detection in remote sensing images and provide important support for the further development of remote sensing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a schematic diagram of a data set sample in the embodiment;
[0034] Figure 2 is an example diagram of the MFDE module in the embodiment;
[0035] Figure 3 Schematic diagram of an AED module in an embodiment;
[0036] Figure 4 is a visualization diagram of an ablation experiment in an embodiment;
[0037] Figure 5 It is a visualization diagram of the comparative experiment in the embodiment. DETAILED DESCRIPTION
[0038] The content of the present invention is further described below in conjunction with the drawings and embodiments, but the present invention is not limited thereto.
[0039] Example:
[0040] A method for detecting changes in remote sensing images based on deep learning, comprising the following steps:
[0041] 1) Construct a data set: Figure 1As shown in the figure, two different datasets, CLCD and NJDS, are used to verify the effectiveness of the method. The NJDS dataset collects remote sensing images of Nanjing City, including 563 pairs of 512×512 pixel image pairs, covering the period from 2016 to 2019, and mainly focusing on land use changes, including urban expansion, vegetation changes and other types of changes. The CLCD dataset is a public farmland dataset containing 2400 pairs of farmland change samples taken by Gaofen, with an image size of 256×256 pixels. These dual-phase images were taken in 2017 and 2019 in Guangdong Province, China, and the spatial resolution of these images ranges from 0.5m to 2m. Each group of samples contains two images and a binary label representing the farmland changes. All samples are divided into training set, validation set and test set in a ratio of 6:2:2. The number of samples in the training set, validation set and test set are 1440, 480 and 480 respectively. The two datasets have different geographical characteristics and change types, which can comprehensively evaluate the performance of the change detection algorithm.
[0042] 2) Construct a multi-branch feature difference extraction MFDE module: The MFDE module adopts a three-parallel branch network structure: a difference branch focuses on modeling the difference between two sets of features, a preservation branch retains the information lost during the encoding process, and a fusion branch fuses feature information of different scales. This design ensures the independent calculation of difference features and retains the original image information to the greatest extent. In the traditional encoder and decoder network architecture, the encoder cannot effectively capture image differences and there is information loss. Therefore, the MFDE of this technical solution fully extracts the difference features of the image from the dual-phase image. The MFDE module adopts three strategies: difference, preservation and fusion, such as Figure 2 As shown in Figure 2, the MFDE module intercepts features from each layer of the encoder to more accurately model the difference features and adopts multi-scale fusion to obtain high-quality difference features, which provides strong guidance for the final change detection, including:
[0043] 2-1) Difference branch: This branch takes the encoder's encoded features as input, calculates the difference matrix of the input features to capture the initial change information, and then calculates the corresponding threshold matrix. Based on the threshold matrix, the difference part of the feature expression is highlighted to obtain the feature difference. The mask module (MASK) is used to filter out non-change information from the features. The process formula is shown in (1):
[0044]
[0045] Among them, T 1 and T 2 is the feature of the dual-phase image, F d T 1 and T 2The absolute value of the difference between two features indicates the difference between them. Avgpool indicates average pooling. MASK(x, 0) sets all elements in the matrix that are less than 0 to 0. MASK sets pixel values that are less than the difference to 0, eliminating their influence in the network training process and increasing the weight of pixels that actually contain difference information.
[0046] 2-2) Preservation branch: The preservation branch models the information lost in the difference branch. The subtraction operation in the difference branch will lead to the loss of pixel information. Therefore, a preservation branch is added to model the lost information and integrate the modeled features into the difference information to generate high-quality difference features. The process formula is shown in (2):
[0047] F P =Conv(Concat(Concat(T 1 +T 2 ),(T 1 +T 2 ))), (2),
[0048] Among them, Concat is used to connect the input vector; Conv is a 3×3 convolution; F P Represents the feature T 1 and T 2 Compared with direct addition, the input vector T is additionally linked in the channel dimension 1 and T 2 , which retains more data information and reduces the information loss in the network coding process;
[0049] 2-3) Fusion branch: Fusion of feature information of different scales to obtain difference features with complete information. Convolution and upsampling are used to match the size of deep features and shallow features. Then, convolution is used to model the fused difference features. The formula of the process is shown in (3):
[0050]
[0051] Among them, F up Record the information obtained from the first two branches of each layer of features in the encoding stage; Conv is a 3×3 convolution; F last is the difference feature obtained from the previous layer, F diff It is the difference feature obtained from the current layer. The fusion adopts a layer-by-layer progressive method to perform feature fusion: first, the first layer feature is fused with the second layer feature, and then the fusion result is fused with the third layer feature, and so on, finally realizing the hierarchical integration of the difference features of different scales in the encoding stage;
[0052] 3) Construct an asymmetric integrated decision AED module: The AED module constructs an independent learner according to the receptive field characteristics of different decoding stages, and corrects the deviation of the decoder at each scale by assigning different weight factors to different scales. The AED module establishes a parallel inference network for the vertical encoder group to further extract the feature information of each network layer, thereby achieving more accurate change detection: In the decoding stage, the decoder reconstructs the extracted feature vector, and uses jump connections to retain the original information of the bidirectional image in the encoding stage to make up for the detail information lost in the feature extraction process. However, the general decoding network structure ignores the difference in receptive field scale. Features of different scales tend to express features at different levels. Shallow features usually focus on color and texture, while deep features contain more pixel information and represent semantics. There is an urgent need for a method to deeply utilize the change information contained in the receptive fields of different scales, such as Figure 3 As shown in the figure, the AED module sets the corresponding confidence factor according to the difference in the receptive fields of features of different scales to measure the reliability of the prediction of the final change result and to correct the weight of features of different scales on the change. The AED module converts the original vertical linear network reasoning process into a parallel integrated decision process to obtain the guiding features for change detection, including:
[0053] 3-1) Asymmetric learning: In the AED module, the change prediction map obtained from the last layer of the decoder stage has the same feature map size of 256×256 as the original input. Each pixel at this scale corresponds to each pixel of the change map, and the receptive field ratio is 1. The calculation method of the receptive field ratio at different scales is shown in formula (4):
[0054]
[0055] The calculation method of the confidence factor at different scales through the receptive field ratio is shown in formula (5):
[0056]
[0057] Among them, r i is the receptive field ratio of the i-th layer in the decoding stage; W t and H t is the size of the original input image; W ti and H ti Respectively represent the width and height of the i-th layer feature; I i represents the confidence factor of the i-th layer; d represents the number of network layers in the decoding stage. Through the Channel Avgpool Module (CAM) and the MASK module, the region expressing non-changing information in the feature is suppressed. Then, the output feature F of each layer is obtained according to different feature factors. dec ;
[0058] 3-2) Integrated decision: In the decoding stage, the confidence factor is multiplied with the feature maps at different scales to construct the feature vectors at the corresponding scales. After fusion, the final change detection prediction is obtained. The formal content of this process is shown in formula (6):
[0059]
[0060] Among them, d represents the number of layers in the network decoding stage; Conv is a 7×7 convolution; Extension repeats a single element of the feature at different scales in both width and height dimensions to achieve the same scale as the prediction result; Change refers to the change prediction result of the feature difference and receptive field network FDRFNet (Feature Difference and Receptive Field Network, referred to as FDRFNet) of the overall network architecture of this technical solution.
[0061] In this example, PyTorch was used on an NVIDIA GTX 3090 GPU (24GB memory) to perform a verification experiment on FDRFNet. The programming language used was Python 3.10, and the deep learning framework PyTorch 2.0.1+cu117 was used. During the training process, the AdamW optimizer was used, the initial learning rate was 0.001, and the weight decay was 0.001. In addition, if the F1 verified every 12 epochs does not increase, this chapter will reduce the learning rate by 0.1, set the batch size to 32, and train FDRFNet in total. The method is tested over 300 epochs and compared with two different datasets, CLCD and NJDS, as well as ablation experiments to fully verify the effectiveness of the method. Four standard indicators are used to comprehensively evaluate the change detection performance: Precision (P) measures the accuracy of the detection results; Recall (R) measures the completeness of the detection results; F1-score is the harmonic mean of precision and recall; Intersection over Union (IoU) evaluates the detection accuracy by calculating the overlap between the predicted change area and the actual change area:
[0062] As shown in Table 1, the ablation experiment results show that the introduction of MFDE and AED modules significantly improves the model performance: First, after adding the MFDE module, the NJDS dataset increased from 64.40% to 66.96%, and the CLCD dataset increased from 71.14% to 72.66%, indicating that the MFDE module improves the overall performance of the model through effective feature difference extraction and information retention; after further adding the AED module, the F1 scores of the two datasets increased to 68.19% and 73.63% respectively, indicating that the AED module further optimizes the model performance by processing receptive field deviations of different scales. It is particularly noteworthy that in complex scenes, the performance improvement brought by the two modules is more significant: the accuracy of the NJDS dataset increased from 80.84% to 84.07%, and all indicators of the CLCD dataset were evenly improved, showing the stability of the model. These results confirm the effectiveness of the MFDE and AED modules in improving change detection performance, especially when dealing with complex scenes.
[0063] Table 1:
[0064] ;
[0065] like Figure 4 As shown in the figure, from the visualization results of the ablation experiment, it can be found that in the NJDS dataset, with the gradual addition of modules, the edges of buildings become clearer, the false positive areas are gradually reduced, and the integrity of the changing areas is gradually improved. In the CLCD dataset, the detection accuracy of water body boundaries is improved, the detail retention ability is enhanced, and the contours of the changing areas are more accurate. The addition of the MFDE module significantly improves the edge detection capability, while the AED module further optimizes the integrity and accuracy of the changing areas. The results from left to right show a gradual approach to GT. This ablation experiment clearly demonstrates the contribution of each module and proves the effectiveness and necessity of the method. Each newly added module brings significant performance improvements, especially in edge detection and integrity.
[0066] As shown in Table 2, the comparative experiments on the CLCD dataset show that FDRFNet has comprehensive performance advantages over other methods: the precision (P) reaches 73.78%, exceeding the second-best BiT (73.27%); the recall (R) reaches 73.48%, which is significantly higher than other methods and 1.58 percentage points higher than the second-best SGSLN (71.90%); the comprehensive performance indicators F1 and IoU reach 73.63% and 58.27% respectively, exceeding all the comparison methods, and compared with the second-best method SGSLN (F1 is 71.14%, IoU is 55.20%), it is 2.49 and 3.07 percentage points higher, respectively. In particular, compared with the traditional FC series methods, FDRFNet significantly improves the recall rate while maintaining a high precision, reflecting a more balanced detection capability; These results show that FDRFNet has stronger feature extraction and change detection capabilities when dealing with complex scenes such as the CLCD dataset:
[0067] Table 2:
[0068] Methods P(%) R(%) F1(%) IoU(%) FC-EF 70.82 62.37 66.32 49.62 FC-Sima-diff 71.7 47.6 57.22 40.07 FC-Siam-conc 61.42 62.75 62.08 45.01 SNUNet 64.26 52.33 57.69 40.54 BiT 73.27 52.91 61.45 44.35 SGSLN 70.39 71.90 71.14 55.20 FDRFNet 73.78 73.48 73.63 58.27 ;
[0069] As shown in Table 3, in the comparative experiment of NJDS dataset, FDRFNet shows significant performance advantages: in terms of precision (P), it reaches 84.07%, exceeding the second-best SGSLN (80.84%); in terms of recall (R), it reaches 57.35%, which is lower than 62.78% of DTCDSCN, but the comprehensive F1 score and IoU index reach 68.19% and 51.73% respectively, which is significantly better than other methods and higher than the second-best SGSLN (F1 is 64.40%, IoU is 47.50%), which are 3.79 and 4.23 percentage points higher than those of traditional methods such as U-Net (F1 is 49.35%) and PSPNet (F1 is 54.12%). FDRFNet has a more significant performance improvement than traditional methods such as U-Net (F1 is 49.35%) and PSPNet (F1 is 54.12%), and it also has a significant advantage over the improved attention model AttU-Net (F1 is 49.48%). These results show that FDRFNet has stronger feature extraction and change recognition capabilities when dealing with the complex land use change detection task of the NJDS dataset:
[0070] Table 3:
[0071] Methods P(%) R(%) F1(%) IoU(%) U-Net 46.45 52.64 49.35 32.76 AttU-Net 55.57 44.6 49.48 32.88 PSPNet 50.57 58.21 54.12 37.1 DTCDSCN 51.92 62.78 56.84 39.7 IFN 49.44 14.35 22.24 12.51 MTU-Net 65.29 62.82 64.03 47.09 SGSLN 80.84 53.52 64.4 47.5 FDRFNet 84.07 57.35 68.19 51.73 ;
[0072] like Figure 5As shown in the figure, from the visualization results of the comparative experiment, the performance comparison of three different methods (SGSLN, AMTNet, FDRF) on the two datasets of CLCD and NJDS is shown. In the CLCD dataset, the three methods perform differently when dealing with changes in water boundaries: SGSLN has a certain loss of edge details, and the results of AMTNet have some noise points and discontinuous areas, while the FDRF method is closest to Ground Truth in maintaining boundary integrity and detail performance. For the NJDS dataset, the change detection of building areas is mainly demonstrated. The detection results of SGSLN are relatively conservative and there are missed detections. AMTNet is overly sensitive in the detection of changed areas. The FDRF method maintains the integrity of the changed area while also avoiding false detections. The overall effect is closest to the true value label. From the comparison results of the two datasets, it can be seen that the FDRF method has shown strong robustness and accuracy in different scenarios, especially in edge preservation and detail characterization.
[0073] The experimental results of this method on two benchmark datasets, NJDS and CLCD, show that this method achieves 68.19% and 73.63% F1 scores respectively, which is significantly superior to existing methods, especially when dealing with complex urban changes and building changes, showing stronger feature extraction and boundary perception capabilities. This method has broad application prospects and can be applied to: urban construction planning and monitoring, dynamic assessment of vegetation cover, disaster assessment and emergency response, ecological environment protection and land use change analysis. With the continuous growth of remote sensing data and the deepening of its application in various fields, this method shows great potential in improving the efficiency and accuracy of change detection, which will provide important support for the further development of remote sensing technology.
Claims
1. A method for detecting changes in remote sensing images based on deep learning, comprising the following steps: 1) Constructing datasets: Two different datasets, CLCD and NJDS, are used to verify the effectiveness of the method. The NJDS dataset collects remote sensing images of Nanjing City, including 563 pairs of 512×512 pixel image pairs, covering a period from 2016 to 2019, including urban expansion, vegetation changes and other types of changes. The CLCD dataset is a public farmland dataset containing 2,400 pairs of farmland change samples taken by Gaofen. The image size is 256×256 pixels. These dual-phase images were taken in 2017 and 2019 in Guangdong Province, China. The spatial resolution of these images ranges from 0.5m to 2m. Each group of samples contains two images and a binary label representing farmland changes. All samples are divided into training set, validation set and test set in a ratio of 6:2:
2. The number of samples in the training set, validation set and test set are 1,440, 480 and 480 respectively. 2) Construct a multi-branch feature difference extraction MFDE module: The MFDE module adopts a three-parallel branch network structure: a difference branch focuses on modeling the difference between two sets of features, a preservation branch retains the information lost during the encoding process, and a fusion branch fuses feature information of different scales. The MFDE module adopts three strategies: difference, preservation, and fusion. The MFDE module intercepts features from each layer of the encoder and uses multi-scale fusion to obtain high-quality difference features, including: 2-1) Difference branch: This branch takes the encoder’s encoded features as input, calculates the difference matrix of the input features, captures the initial change information, and then calculates the corresponding threshold matrix. Based on the threshold matrix, the difference part of the feature expression is highlighted to obtain the feature difference. The mask module MASK is used to filter out non-change information from the features. The process formula is shown in (1): Among them, T1 and T2 are the characteristics of the dual-phase image, and F d is the absolute value of the difference between the two features T1 and T2, indicating the difference between them. Avgpool represents average pooling. MASK(x, 0) sets all elements in the matrix that are less than 0 to 0. MASK sets pixel values that are less than the difference to 0, eliminating their influence in the network training process and increasing the weight of pixels that truly contain difference information. 2-2) Preservation branch: The preservation branch models the information lost in the difference branch. The subtraction operation in the difference branch will lead to the loss of pixel information. Therefore, a preservation branch is added to model the lost information and integrate the modeled features into the difference information. The formula of the process is shown in (2): F P =Conv(Concat(Concat(T1+T2),(T1+T2))), (2), Among them, Concat is used to connect the input vector; Conv is a 3×3 convolution; F P Represents partial information of features T1 and T2; 2-3) Fusion branch: Fusion of feature information of different scales to obtain difference features with complete information. Convolution and upsampling are used to match the size of deep features and shallow features. Then, convolution is used to model the fused difference features. The formula of the process is shown in (3): Among them, F up Record the information obtained from the first two branches of each layer of features in the encoding stage; Conv is a 3×3 convolution; F last is the difference feature obtained from the previous layer, F diff It is the difference feature obtained from the current layer. The fusion adopts a layer-by-layer progressive method to perform feature fusion: first, the first layer feature is fused with the second layer feature, and then the fusion result is fused with the third layer feature, and so on, finally realizing the hierarchical integration of the difference features of different scales in the encoding stage; 3) Constructing an asymmetric integrated decision AED: The AED module constructs an independent learner based on the receptive field characteristics of different decoding stages, and corrects the deviation of the decoder at each scale by assigning different weight factors to different scales. The AED module establishes a parallel reasoning network for the vertical encoder group: In the decoding stage, the decoder reconstructs the extracted feature vector, and uses jump connections to retain the original information of the bidirectional image in the encoding stage. The AED module sets the corresponding confidence factor according to the difference in the receptive field of features at different scales. The AED module converts the original vertical linear network reasoning process into a parallel integrated decision process, including: 3-1) Asymmetric learning: In the AED module, the change prediction map obtained from the last layer of the decoder stage has the same feature map size of 256×256 as the original input. Each pixel at this scale corresponds to each pixel of the change map, and the receptive field ratio is 1. The calculation method of the receptive field ratio at different scales is shown in formula (4): The calculation method of the confidence factor at different scales through the receptive field ratio is shown in formula (5): Among them, r i is the receptive field ratio of the i-th layer in the decoding stage; W t and H t is the size of the original input image; W ti and H ti Respectively represent the width and height of the i-th layer feature; I i represents the confidence factor of the i-th layer; d represents the number of network layers in the decoding stage. Through the channel average pooling module CAM and the MASK module, the area expressing non-changing information in the feature is suppressed. Then, the output feature F of each layer is obtained according to different feature factors. dec ; 3-2) Integrated decision: In the decoding stage, the confidence factor is multiplied with the feature maps at different scales to construct the feature vectors at the corresponding scales. After fusion, the final change detection prediction is obtained, as shown in formula (6): Among them, d represents the number of layers in the network decoding stage; Conv is a 7×7 convolution; Extension repeats a single element of the feature at different scales in both width and height dimensions, and Change refers to the change prediction results of the overall network architecture difference features and the receptive field network FDRFNet.