Remote sensing image building change detection method based on depth features

The method addresses the issue of pseudo changes in deep learning-based building detection by using a twin ResNet18 encoder and DCEM/MDAM to enhance feature extraction and aggregation, achieving precise and detailed building change detection.

CN120318676APending Publication Date: 2025-07-15NORTHWEST A & F UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373657.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing deep learning methods are difficult to distinguish pseudo-changes caused by imaging time and sensor differences in remote sensing image building change detection, which affects detection accuracy.

Method used

The remote sensing image building change detection method based on depth features is adopted, and the depth features of the two-time phase image are extracted using the twin ResNet18 encoder with shared weights, the difference features are obtained through DCEM and MDAM modules, and the binary change detection result diagram is generated through the progressive multi-directional aggregation of differential features.

Benefits of technology

Accurately locate the building changes area, suppress interference such as light and seasonal changes, and generate a remote sensing image building changes detection result map with complete internal structure and clear edges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318676A_ABST
    Figure CN120318676A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image building change detection method based on depth features. The method specifically comprises the following steps: step 1, constructing a remote sensing change detection data set; step 2, constructing a remote sensing change detection model; step 3, extracting depth features of different levels of a dual-temporal remote sensing image; step 4, obtaining a difference characteristic # imgabs0 #; step 5, carrying out progressive multidirectional aggregation on the difference characteristics of each level to obtain a result graph containing rich context information; training the model through the training set; and step 6, obtaining a binary change detection result graph of the optimal model on the test set, and performing comparative analysis with other algorithms. According to the remote sensing image building change detection method based on the depth features, the problem of false change caused by imaging time and sensor difference in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning methods, and particularly relates to a method for detecting building changes in remote sensing images based on deep features. Background Art

[0002] At present, methods for detecting building changes in remote sensing images can be roughly classified into traditional methods based on image processing and methods based on deep learning. Among them, traditional methods include direct comparison method, image transformation method, and post-classification comparison method. The direct comparison method does not go through classification, and directly processes and compares the spectral information of remote sensing images of the same area at different times by subtracting (or dividing) the corresponding bands of two-phase remote sensing images, so as to determine the location and scope of changes, and then detects building changes through manual visual interpretation or classification methods. The image transformation method detects changes in surface buildings by performing coordinate transformation on multi-phase remote sensing images, converting the original image space into a manually designed feature space and then comparing them. The post-classification comparison method is a relatively simple and clear change detection method. First, each phase of remote sensing image is classified separately using a unified classification system, and then the type, quantity, and location of changes are directly obtained by comparing the classification results. Although the above three methods can simply and efficiently detect the change areas of buildings, in the direct comparison method, since its calculation process is through point-to-point operations, there are many noises in the difference image, and the anti-interference ability is weak, and it cannot provide the attributes of changes; while the image transformation method depends on the manually designed feature space, and the termination condition of the algorithm requires the selection of a manual threshold, so there are deficiencies in terms of accuracy and reliability, which limits the scope of related applications. The post-classification comparison method cannot detect the subtle changes within a certain land cover type when comparing the land cover classification data of different periods; moreover, the accuracy of the change map generated after comparing the classification data of two phases is only roughly equivalent to the product of the classification accuracy values of each phase, because the errors existing in each individual classification will be further amplified during the spatial comparison process.

[0003] In recent years, deep neural networks have gradually emerged in the field of computer vision. It can autonomously learn deep features from labeled data to detect building change areas, greatly reducing manual intervention and realizing end-to-end building change detection, showing great potential in the field of detecting building changes in remote sensing images. However, existing deep learning methods are still difficult to distinguish the pseudo-changes caused by imaging time and sensor differences, misdetecting the pseudo-change areas as change areas, which greatly affects the accuracy of the change detection algorithm and thus affects the detection of building changes. Summary of the Invention

[0004] The object of the present invention is to provide a remote sensing image building change detection method based on deep features, which solves the problem of pseudo-changes caused by imaging time and sensor differences in the prior art, thereby affecting the detection of building changes.

[0005] The technical solution adopted by the present invention is a remote sensing image building change detection method based on deep features, which specifically includes the following steps: Step 1, construct a remote sensing change detection dataset, obtain the LEVIR-CD and WHU-CD datasets and crop them into non-overlapping 256×256-sized image patches, including a training set, a validation set, and a test set; Step 2, construct a remote sensing change detection model, including a Siamese ResNet18 encoder with shared weights, a DCEM, and an MDAM; Step 3, use the Siamese ResNet18 encoder with shared weights in Step 2 to extract deep features of different levels of the two-temporal images, and obtain , ,..., , A total of 5 layers of two-temporal features with gradually decreasing spatial resolution; Step 4, obtain the difference feature by passing the extracted two-temporal features through the DCEM ; Step 5, pass the difference feature through the MDAM for progressive multi-directional aggregation of the difference feature to obtain a binary change detection result map; train the model with the training set and evaluate the model performance during the training process using the validation set to obtain the optimal model; Step 6, send the test set into the trained optimal model to classify the changed area and the unchanged area to obtain the final binary change detection result map.

[0006] The present invention is also characterized in that The specific extraction step of the Siamese ResNet18 encoder in Step 3 is as follows: the encoder is a Siamese architecture with shared weights, and each branch uses a pre-trained ResNet18 as a feature extractor, extracting a total of 5 layers of features. Among them, the first layer is taken from the feature map with a resolution of 128×128 before the maxpool max pooling layer after the 7×7 convolution operation, and the remaining layers are taken from after the Resblock, with feature maps of sizes 64×64, 32×32, 16×16, and 8×8 with gradually decreasing resolutions. The five layers of two-temporal features extracted are respectively represented as, , ,..., , where the superscript on the upper right represents the temporal phase, and the subscript on the lower right represents the level where the feature is extracted.

[0007] Step 4 is specifically as follows: Step 4.1: Obtain spatial, channel, and time difference features , , ; Step 4.2: Obtain the final difference features .

[0008] Step 4.1: Obtain spatial difference features and channel difference features Specifically as follows: Using the dual-temporal features of the corresponding layer , as the input, DCEM first calculates the absolute value of the difference between the dual-temporal features pixel by pixel to obtain spatial difference features ; Subsequently, use a 3×3 convolution according to to generate the mask weights of the building change area and the dual-temporal features , perform the Hadamard product, and enhance its discrimination ability for the building change area through a residual structure to suppress the interference of light and shadow, obtaining the refined dual-temporal features of the spatial difference information , ; Then, splice the channels of and features together and perform 3×3 convolution to obtain channel difference features , screen the channel dimensions that highlight building changes, capture spectral difference information, and the specific calculation formula is as follows: (1) (2) (3) (4) (5) In the formula, and respectively represent the dual-temporal features of the i-th layer at T1 and T2 time phases, represents the absolute value, represents the convolution operation with a convolution kernel size of k, represents the Hadamard product, represents splicing the dual-temporal features along the channel dimension.

[0009] The specific steps to obtain time difference features are as follows: To capture the temporal difference information, DCEM will Feed the input into the Temporal Difference Awareness (TDA) module with two ConvGRU units as the input of the first unit, and set the hidden state of the first unit to 0 to clear past information. The output of the first ConvGRU unit will be used as the hidden state of the second unit, and will be used as the input of the second unit. In the second ConvGRU, according to its hidden state, it will forget information similar to and retain the changing part as the output, so as to explore the temporal difference information and obtain the time difference features .

[0010] The Temporal Difference Awareness module TDA is specifically expressed as: (6) (7) Among them, the specific expression of ConvGRU is: (8) (9) (10) (11) In the formula, is the input feature of the TDA module, represents the hidden state of the previous moment. In the first ConvGRU unit, is initialized to 0, is . In the second ConvGRU unit, is initialized to the output of the first unit, is . represents the update gate, represents the reset gate, represents the candidate gate, represents the sigmoid activation function.

[0011] Step 4.2 is specifically as follows: After obtaining the spatial, channel, and time difference features , , , in order to construct the dependency relationship between the cross-domain difference features, a gated unit is used to dynamically fuse the above difference features, and the corresponding gated mask weights , , , then the differential features of each domain perform Hadamard product with the mask for itself to achieve feature screening and remove redundant differential information. After that, first add to the features screened by itself, that is and the screened features of itself, i.e., , then add the sum of the screened features and , to the Hadamard product sum of . Finally, obtain the final differential feature through 1×1 convolution. The specific formula is as follows: (12) (13) (14) (15); In the formula, , , respectively represent the gating masks corresponding to the spatial, channel, and temporal differential features of the i-th layer.

[0012] Step 5 is specifically as follows: This step includes three stages: generating the coarse-grained change detection result map, refining the differential features, and fusing the differential features. Using the different-level differential features obtained by the DCEM module, where i = 1,..., 5 as the input. In the stage of generating the coarse-grained change detection result map, the differential features of each level are aggregated together through the FPN method to generate the coarse-grained change detection result map, and the feature with the shallowest resolution of 128×128 in the FPN is used as the reference point for calculating the similarity between the differential features and the coarse-grained result map in the subsequent stages. In the stages of refining the differential features and fusing the differential features, the differential features of each level are obtained by multi-directional fusion of the differential features of different levels in the previous stage. The proportion of the differential feature at the i-th layer and the j-th stage in the fusion process is its corresponding adaptive weight, which is obtained by calculating the similarity between itself and . The specific formula is as follows: (18) In the formula, represents the adaptive weight of the differential feature . The mean value of the differential feature along the channel direction is regarded as the query vector Q, the differential feature is regarded as the key-value vector K, σ is the sigmoid activation function, and s represents the standard deviation of the differential feature along the channel dimension.

[0013] Differential features in the differential feature fusion stage After the upsampling operation, it is converted into a two-channel form of 256×256 size through convolution, and then the final prediction result is obtained through softmax; the above network is trained with the training set, and the validation set is used to evaluate the model performance during the training process to obtain the optimal model. During this process, the BCE and Dice loss functions are used to jointly supervise the coarse-grained change detection result map and the final prediction result of the network. The specific calculation formulas are as follows: (26) In the formula, represents the prediction result map, represents the ground truth label map, and the loss function constrains the prediction result maps of both coarse and fine granularities.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: (1)The remote sensing image building change detection method based on deep features provided by the present invention, which is included in the cross-domain difference enhancement module (DCEM) and the multi-directional difference feature aggregation module (MDAM) based on the duplex gating mechanism, can accurately locate the building change area and obtain a remote sensing image building change detection result map with a complete internal structure and clear edge details.

[0015] (2)The remote sensing image building change detection method based on deep features provided by the present invention, by fully modeling the dependence relationships of channel, spatial, and temporal difference features, progressively generates more representative cross-domain difference features, while accurately locating the building change area, suppressing the interference of irrelevant factors such as illumination and seasonal changes.

[0016] (3)The remote sensing image building change detection method based on deep features provided by the present invention includes a multi-directional difference feature aggregation module (MDAM) with three stages: generating a coarse-grained change detection result map, refining difference features, and fusing difference features. In the stage of generating a coarse-grained change detection result map, a coarse-grained change detection result map is generated through a feature pyramid network (FPN) to supervise the subsequent difference feature aggregation process; in the stages of refining and fusing difference features, the similarity between different-level and different-stage difference features and the coarse-grained result map is calculated through a non-parametric attention mechanism, and the weights of each difference feature in the aggregation process are adaptively generated without adding additional calculation parameters, achieving the purpose of refining difference features. In addition, in the stage of multi-directional aggregation of difference features, by aggregating difference features of different granularity sizes from horizontal and vertical directions, more abundant context information can be captured to obtain a detection result with clear edges while adapting to the scale change of buildings. Description of the Drawings

[0017] Figure 1It is a schematic flowchart of the remote sensing image building change detection method based on deep features of the present invention; Figure 2 It is a schematic structural diagram of the cross-domain difference enhancement module based on the duplex gating mechanism of the present invention; Figure 3 It is a schematic structural diagram of the temporal difference perception module of the present invention; Figure 4 It is a schematic structural diagram of the multi-directional difference feature aggregation module of the present invention; Figure 5 It is a schematic structural diagram of the adaptive weight generation module of the present invention; Figure 6 It is a comparison chart of the qualitative analysis results of the present invention on the LEVIR-CD dataset; Figure 7 It is a comparison chart of the qualitative analysis results of the present invention on the WHU-CD dataset. Specific embodiments

[0018] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments.

[0019] Embodiment 1 The present invention provides a remote sensing image building change detection method based on deep features, as Figure 1 shown, and is specifically implemented according to the following steps: Step 1, construct a remote sensing change detection dataset, obtain the LEVIR-CD and WHU-CD datasets and crop them into non-overlapping 256×256 size image pairs, including a training set, a validation set and a test set; Step 2, construct a remote sensing change detection model, including a Siamese ResNet18 encoder, a DCEM and an MDAM; Step 3, use the Siamese ResNet18 encoder in Step 2 to extract deep features of different levels of the two-temporal remote sensing images, and obtain , ,..., , a total of 5 layers of two-temporal features with gradually decreasing spatial resolution; Step 4, obtain the difference feature by passing the extracted two-temporal features through the DCEM; Step 5, pass the difference feature through the MDAM for progressive multi-directional aggregation of the difference features to obtain a binary change detection result map; train the model with the training set and evaluate the model performance during the training process using the validation set to obtain the optimal model; Step 6: Feed the test set into the trained optimal model to classify the changed areas and unchanged areas, and obtain the final binary change detection result map.

[0020] Example 2 On the basis of Example 1, the specific steps for the Siamese ResNet18 encoder to extract features in Step 3 are as follows: The encoder is a Siamese architecture with shared weights. Each branch uses the pre-trained ResNet18 as a feature extractor, and a total of 5 layers of features are extracted. Among them, the first layer is taken from the feature map with a resolution of 128×128 after the 7×7 convolution operation and before the maxpool max-pooling layer, and the remaining layers are taken from the 64×64, 32×32, 16×16, and 8×8 feature maps with decreasing resolutions after the Resblock. The five layers of dual-temporal features extracted are respectively represented as, , ,..., , , where the superscript on the right represents the temporal phase, and the subscript on the bottom represents the level where the feature is extracted.

[0021] Example 3 On the basis of Example 2, Step 4 is specifically as follows: Step 4.1: Obtain spatial, channel, and temporal difference features , , ; Step 4.2: Obtain the final difference features .

[0022] In Step 4.1, the method for obtaining the spatial difference feature and the channel difference feature is specifically as follows: Taking the dual-temporal features , at the corresponding level as the input, by modeling the dependencies between the channel, spatial, and temporal difference features, progressively generate cross-domain difference features; when constructing the dependencies of the channel and spatial difference features, DCEM first calculates the absolute value of the difference between the dual-temporal features pixel by pixel to obtain the spatial difference feature ; then, use a 3×3 convolution to generate the mask weight of the building change area according to and the dual-temporal features , perform the Hadamard product, and enhance its discrimination ability for the building change area through the residual structure to suppress the interference of light and shadow, and obtain the refined dual-temporal features , ; then, combine and The channels of the features are concatenated together and subjected to a 3×3 convolution to obtain channel difference features , filter the channel dimension that highlights the building changes, capture the spectral difference information, and the specific calculation formula is as follows: (1) (2) (3) (4) (5) In the formula, and represent the dual-temporal features at the i-th level for the T1 and T2 time phases respectively, represents the absolute value, represents the convolution operation with a convolution kernel size of k, represents the Hadamard product, represents the concatenation of the dual-temporal features along the channel dimension.

[0023] Obtain the time difference features The specific steps are as follows: To capture the temporal difference information, DCEM sends into the temporal difference awareness module (TDA) containing two ConvGRU units as the input of the first unit, and sets the hidden state of the first unit to 0 to clear the past information. The output of the first ConvGRU unit will be used as the hidden state of the second unit, and is used as the input of the second unit. In the second ConvGRU, according to its hidden state, it will forget the information similar to and retain the changing part as the output, so as to explore the temporal difference information to obtain the time difference features .

[0024] The temporal difference awareness module TDA is specifically expressed as: (6) (7) Among them, the specific expression of ConvGRU is: (8) (9) (10) (11) In the formula, is the input feature of the TDA module, Represents the hidden state at the previous moment. In the first ConvGRU cell, is initialized to 0, for ; in the second ConvGRU cell, is initialized to the output of the first cell, for . represents the update gate, represents the reset gate, represents the candidate gate, represents the sigmoid activation function.

[0025] Step 4.2 is specifically as follows: After obtaining the spatial, channel, and temporal difference features , , , to construct the dependencies between the cross-domain difference features, a gated unit is used to dynamically fuse the above difference features. By applying a 1×1 convolution and the sigmoid activation function respectively, the gated mask weights corresponding to each difference feature are obtained , , . Then, the difference features of each domain perform the Hadamard product of themselves and the corresponding masks to achieve feature screening and remove redundant difference information. After that, first add to its own screened feature, that is . Then, add the sum of the screened , and the Hadamard product of together. Finally, the final difference feature is obtained through a 1×1 convolution. The specific formula is as follows: (12) (13) (14) (15); In the formula, , , respectively represent the gated masks corresponding to the spatial, channel, and temporal difference features of the i-th layer.

[0026] Example 4 Based on Example 3, Step 5 is specifically as follows: This step includes three stages: generating the coarse-grained change detection result map, refining the difference features, and fusing the difference features; using the different-level difference features obtained by the DCEM module , with \(i = 1,\cdots,5\) as the input. In the stage of generating the coarse-grained change detection result map, the differential features at each level are aggregated together by the FPN method to generate the coarse-grained change detection result map, and the feature with the shallowest layer resolution of \(128\times128\) in the FPN is used as the reference point for calculating the similarity between the differential features and the coarse-grained result map in the subsequent stage; in the stage of refining the differential features and fusing the differential features, the differential features at each level are obtained by multi-directional fusion of the differential features at different levels in the previous stage. The proportion of the differential feature at the \(i\)-th layer and the \(j\)-th stage in the fusion process is its corresponding adaptive weight, which is obtained by calculating its similarity with and the specific formula is as follows: (18) In the formula, represents the adaptive weight of the differential feature . The mean value of the differential feature along the channel direction is regarded as the query vector \(Q\), the differential feature is regarded as the key-value vector \(K\), \(\sigma\) is the sigmoid activation function, and \(s\) represents the standard deviation of the differential feature along the channel dimension.

[0027] The differential feature in the differential feature fusion stage is transformed into a two-channel form of \(256\times256\) size through upsampling operation and then convolution, and the final prediction result is obtained through softmax; the above network is trained with the training set, and the performance of the model during the training process is evaluated using the validation set to obtain the optimal model. During this process, the BCE and Dice loss functions are used to jointly supervise the coarse-grained change detection result map and the final prediction result of the network. The specific calculation formulas are as follows: (26) In the formula, represents the prediction result map, represents the ground truth map, and the loss function constrains both the fine-grained and coarse-grained prediction result maps.

[0028] Example 5 This example provides a remote sensing image building change detection method based on deep features. The specific steps are as follows: S1. Construct a remote sensing change detection dataset, obtain the LEVIR-CD and WHU-CD datasets and crop them into image pairs of \(256\times256\) size, and obtain the corresponding training set, validation set, and test set according to the division standard of the official dataset; ​S2. Build a remote sensing change detection model based on deep features, which mainly includes three parts: a Siamese ResNet18 encoder, a cross-domain difference enhancement module (DCEM) based on a duplex gating mechanism, and a multi-directional difference feature aggregation module (MDAM). S3. The Siamese ResNet18 encoder in S2 is mainly used to extract deep features at different levels of the two-temporal remote sensing images, and obtain , ,..., , a total of 5 layers of two-temporal features; S4. The cross-domain difference enhancement module based on the duplex gating mechanism in S2 is mainly used to enhance the representation ability of different-level difference features for the changed areas. This module takes the two-temporal features at the corresponding levels in the encoder stage , as inputs, first calculates the spatial difference features corresponding to the two-temporal features at each level to recalibrate the spatial context information of their respective temporal features, and then performs channel concatenation convolution on the calibrated two-temporal features , to capture spectral difference information and generate channel difference features . Subsequently, the two-temporal features capture temporal difference information through a TDA module composed of cascaded ConvGRU units. The input of the first ConvGRU unit is , and the hidden state is 0. The input of the second ConvGRU unit is , so as to forget the part similar to and obtain the temporal difference feature . Finally, through a gating mechanism based on the duplex principle, using the gating masks corresponding to the spatial, channel, and temporal domain difference features , , dynamically models the dependencies between them, removes redundant difference information, and obtains difference features with stronger representation ability .

[0029] S5. The multi-directional difference feature aggregation module in S2 is mainly used to efficiently aggregate different-level difference information, capture richer context information to generate an accurate binary change detection result map, which mainly includes three stages: generating a coarse-grained change detection result map, refining difference features, and fusing difference features; this module takes the different-level difference features obtained by the DCEMs module Taking (i = 1, ..., 5) as the input, in the stage of generating the coarse-grained change detection result map, the difference features at each level are aggregated together in the way of FPN to generate the coarse-grained change detection result map, and the feature with the shallowest resolution of 128×128 in FPN is used as the reference point for calculating the similarity between the difference features and the coarse-grained result map in the subsequent stage; in the stage of refining the difference features and fusing the difference features, the difference features at each level are obtained by multi-directional fusion of the difference features at different levels in the previous stage. The proportion of the difference feature at the i-th layer and the j-th stage in the fusion process, that is, its corresponding adaptive weight, is obtained by calculating its similarity with . The formula is as follows: . The difference features in the stage of fusing the difference features After upsampling operation, it is transformed into a two-channel form of 256×256 size through convolution, and then the final prediction result is obtained through softmax.

[0030] S6. The above network is trained with the training set, and the validation set is used to evaluate the model performance during the training process to obtain the optimal model. In the above process, we use the BCE and Dice loss functions to jointly supervise the coarse-grained change detection result map and the final prediction result of the network. The specific calculation formulas are as follows: ; The network is implemented through the Pytorch framework and trained on an RTX 4090 graphics card. The SGD optimizer is used, and it is trained for 200 epochs. The initial learning rate is 0.01, and a linear decay learning rate adjustment strategy is adopted. The batch size is 8. To evaluate the performance of this method, the present invention uses Pre, Rec, OA, F1, IoU for quantitative evaluation, and the evaluation results are shown in Tables 1 and 2.

[0031] S7. The test set is sent into the trained optimal model to classify the changed area and the unchanged area, and the final binary change detection result map is obtained, where 0 represents the unchanged part, 1 represents the changed part, the unchanged part is black, the changed part is white, the missed detection part is green, and the misdetection part is red. The visualization comparison result with other methods is as Figure 6 , shown in Figure 7.

[0032] Furthermore, the encoder adopted by the present invention is a siamese architecture with shared weights. Each branch uses a pre-trained Resnet18 as the feature extractor, and a total of 5 layers of features are extracted. Among them, the first layer is taken from the feature map with a resolution of 128×128 after the 7×7 convolution operation and before the maxpool max-pooling layer, and the remaining layers are taken from after the Resblock, and the resolutions are 64×64, 32×32, 16×16, and 8×8 in turn. The five layers of double-temporal features extracted are respectively represented as , , ..., , , where the superscript represents the time phase and the subscript represents the level where the feature is extracted.

[0033] The process of the Dualplex Gating Mechanism-based Cross-Domain Enhancement Module (DCEM) is as follows: This module takes the dual-time-phase features , at the corresponding level as inputs, and progressively generates cross-domain difference features with stronger representational capabilities by modeling the dependencies between channel, spatial, and temporal difference features. When constructing the dependencies of channel and spatial difference features, DCEM first obtains the spatial difference feature by calculating the absolute value of the difference between dual-time-phase features pixel by pixel. Subsequently, it uses a 3×3 convolution to generate the mask weight of the building change area and the dual-time-phase features , performs a Hadamard product, and enhances its discrimination ability for the building change area through a residual structure to suppress the interference of light and shadow, obtaining the refined dual-time-phase features , . Then, it concatenates the channels of the and features together, performs a 3×3 convolution to obtain the channel difference feature , filters out the channel dimensions that highlight building changes, and captures spectral difference information. The specific calculation formula is as follows: (1) (2) (3) (4) (5) In the formula, and respectively represent the dual-time-phase features of the i-th level at time phases T1 and T2, represents the absolute value, represents the convolution operation with a convolution kernel size of k, represents the Hadamard product, represents concatenating the dual-time-phase features along the channel dimension.

[0034] To better capture the temporal difference information, DCEM will Feed the input into the Temporal Difference Awareness Module (TDA) containing two ConvGRU units as the input of the first unit, and set the hidden state of the first unit to 0 to clear past information. The output of the first ConvGRU unit will be used as the hidden state of the second unit, and will be used as the input of the second unit. In the second ConvGRU, it will forget information similar to according to its hidden state and retain the changing part as the output, thereby exploring temporal difference information to obtain temporal difference features , and the overall process is as shown in Figure 3 .

[0035] The entire process of TDA can be expressed as: (6) (7) Among them, the overall process of ConvGRU can be expressed as: (8) (9) (10) (11) In the formula, is the input feature of the TDA module, represents the hidden state of the previous moment. In the first ConvGRU unit, is initialized to 0, is . In the second ConvGRU unit, is initialized to the output of the first unit, is . represents the update gate, represents the reset gate, represents the candidate gate, represents the sigmoid activation function.

[0036] After obtaining the spatial, channel, and temporal difference features , , , in order to build the dependency relationship between cross-domain difference features, a gated unit is used to dynamically fuse the above difference features, and the corresponding gated mask weights , , , then the differential features of each domain perform Hadamard product with the corresponding masks to achieve feature screening and remove redundant differential information. After that, first add to its own screened features, that is and , then add the sum of , after feature screening to the Hadamard product sum of . Finally, the final differential feature is obtained through 1×1 convolution. The overall process of the above is as shown in Figure 2 Figure 2 .

[0037] (12) (13) (14) (15); In the formula, , , respectively represent the gating masks corresponding to the spatial, channel, and temporal differential features of the i-th layer.

[0038] As shown in Figure 4 , the multi-directional differential feature aggregation module (MDAM) serves as the decoder of the proposed network, which mainly includes the following stages: generation of the coarse-grained change detection result map, refinement of differential features, and fusion of differential features. First, the channel dimension of the multi-level differential features captured by DCEMs (cross-domain differential enhancement module based on duplex gating mechanism) is unified to 128 through convolution operations. Then, in the stage of generating the coarse-grained change detection result map, a bottom-up feature pyramid network (FPN) structure is adopted to gradually aggregate multi-scale differential features to generate the coarse-grained change detection result map. To solve the problem of spatial information loss caused by frequent upsampling operations in FPN and alleviate the semantic differences between different levels, in the stage of differential feature refinement and differential feature fusion, as shown in Figure 5 , the fusion ratio of differential features at different levels is adjusted through an adaptive weight strategy to better improve the detection ability for small-sized buildings. Finally, in the stage of differential feature fusion, by aggregating the differential features from the horizontal and vertical directions, richer context information is captured to generate a high-quality binary change detection result map.

[0039] Among them, after unifying the channel dimension of the differential features of each level to 128, the calculation formula in the stage of generating the coarse-grained change detection result map is as follows: (16) (17) It represents the differential features of the j-th part of the i-th layer in the decoder (coarse-grained change detection result map generation - 1, differential feature refinement - 2, differential feature fusion - 3). CB represents a series of convolutional operations with convolutional kernel sizes of 3, 1, and 3 respectively, and Up represents the upsampling operation. Through bottom-up hierarchical feature aggregation, differential features can be obtained. Subsequently, it is upsampled to the same size as the ground truth (GT), and a coarse-grained change detection result map is generated through a classifier.

[0040] In the differential feature refinement stage, MDAM adopts a difference-aware non-parametric attention mechanism (AW), which calculates and to assign adaptive weights to the differential feature through the spatial-spectral similarity between them. The calculation process is as follows: (18) In the formula, represents the adaptive weight of the differential feature . The mean value of the differential feature along the channel direction is regarded as the query vector Q, the differential feature is regarded as the key-value vector K, σ is the sigmoid activation function, and s represents the standard deviation of the differential feature along the channel dimension. The differential feature refinement process can be expressed as: (19) (20) On the one hand, the adaptive weight assignment strategy proposed by the present invention can autonomously focus the attention of different hierarchical differential features on the spatial-spectral characteristics similar to the differential feature , thereby accurately locating the real change area and reducing the probability of missed detection. On the other hand, this strategy effectively distinguishes the contribution degrees of different hierarchical differential features to the binary change detection result, rather than regarding them as equally important and fusing them in an equal proportion manner. In addition, the query-key value vectors (Q, K) in the non-parametric attention mechanism directly come from the differential features themselves, without learning additional parameters and without increasing the computational cost. Thanks to the guidance of the coarse-grained change detection result map, after differential feature refinement, the cross-scale semantic gap between different hierarchical differential features is effectively alleviated, providing a strong guarantee for the context interaction between them. It should be noted that during the differential feature refinement process, the differential features in all directions have been refined through different hierarchical differential information, enabling the scope of differential feature fusion to break through the limitation of the local neighborhood, thereby significantly enhancing the richness and representativeness of the captured context information. Therefore, as Figure 4As shown, in the multi-directional fusion stage of differential features, more abundant context information can be captured by aggregating differential features at different levels in multiple directions. This process can be expressed as: (21) (22) (23) Finally, the result of multi-directional aggregation of differential features is upsampled to the same size as the label map and fed into the classifier for classification to obtain the final binary change detection result map.

[0041] Furthermore, the loss function adopted in the present invention combines cross-entropy loss and dice loss function to alleviate the problem of unbalanced sample distribution caused by the fact that the unchanged regions are much more than the changed regions through the dice loss function: (24) (25) (26) Where represents the predicted result map, represents the true label map, and the loss function constrains both the fine-grained and coarse-grained predicted result maps.

[0042] Example 6 The examples adopted in the present invention are verified on the LEVIR-CD and WHU-CD datasets, and compared with the FC_EF, BIT, ChangeFormer, and DMINet algorithms. The qualitative analysis results are as Figure 6 , 7 shows, where white represents the changed region, black represents the unchanged region, green represents the missed detection region, and red represents the misdetection region. According to Figure 6 (b), (f), it can be observed that the method proposed in the present invention can effectively detect dense distribution building groups containing buildings of different scales, while the remaining methods such as BIT and ChangeFormer all have misdetection phenomena. In Figure 6 (a), Figure 7 (c), the present invention has obtained the most complete internal structure and clear and distinct edge details in the change detection of large-scale buildings, highlighting its advantages in dealing with the changes of large-scale buildings covering a wide area. In addition, this study uses F1, IoU, Pre, Rec, and OA as evaluation indicators for quantitative analysis. The corresponding index results are shown in Tables 1 and 2. The method proposed in the present invention has obtained the optimal results in terms of F1 and IoU indicators, further proving the effectiveness of the proposed method.

[0043] Table 1 Evaluation results of LEVIR-CD dataset

[0044] Table 2 Evaluation Results of WHU-CD Dataset

Claims

1. A method for detecting building changes in remote sensing images based on deep features, characterized in that, Specifically, it includes the following steps: Step 1: Construct a remote sensing change detection dataset, obtain the LEVIR-CD and WHU-CD datasets and crop them into non-overlapping 256×256 size image pairs, including training set, validation set and test set; Step 2: Construct a remote sensing change detection model, including Siamese ResNet18 encoder, DCEM and MDAM; Step 3: Use the twin ResNet18 encoder in Step 2 to extract the depth features of different levels of the dual-temporal remote sensing images, and obtain , , ..., , A total of 5 layers of dual-temporal features; Step 4, obtaining the differential features from the extracted dual-phase features through DCEM ; Step 5, for the differential features perform progressive multi-directional aggregation of the differential features through MDAM to obtain a binary change detection result map; train the model with the training set and evaluate the model performance during the training process using the validation set to obtain the optimal model; Step 6: Feed the test set into the trained optimal model to classify the changed areas and unchanged areas, and obtain the final binary change detection result map.

2. The method for detecting building changes in remote sensing images based on deep features according to claim 1, wherein The specific extraction steps of the twin ResNet18 encoder described in step 3 are as follows: The encoder is a twin architecture with shared weights. Each branch uses a pre-trained ResNet18 as a feature extractor and extracts 5 layers of features. Among them, the first layer is taken from the feature map with a resolution of 128×128 after the 7×7 convolution operation and before the maxpool max pooling layer. The remaining layers are taken from the feature maps with resolutions of 64×64, 32×32, 16×16, and 8×8 that decrease sequentially after the Resblock. The five layers of dual-temporal features extracted are respectively represented as, , ,..., , , where the superscript represents the temporal phase and the subscript represents the level where the feature is extracted.

3. The method for detecting building changes in remote sensing images based on deep features according to claim 1, characterized in that Specifically, Step 4 is as follows: Step 4.1 Obtain spatial, channel, and time difference features , , ; Step 4.2 Obtain the final differential features .

4. The method for detecting building changes in remote sensing images based on depth features according to claim 1, characterized in that Obtaining the spatial difference features described in step 4.1 and the channel difference features Specifically: With the dual-temporal features of the corresponding level , as the input, by modeling the dependencies among the channel, spatial, and temporal difference features, progressively generate cross-domain difference features; When constructing the spatial difference feature dependencies of the channels, DCEM first obtains the spatial difference features by calculating the absolute value of the difference between the two-temporal features pixel by pixel. Subsequently, it uses a 3×3 convolution to generate the mask weights for the building change regions and the two-temporal features , perform a Hadamard product, and enhance its discrimination ability for the building change regions through a residual structure to suppress the interference of light and shadow, obtaining the two-temporal features with refined spatial difference information , ; Then, it concatenates the channels of the and features together and performs a 3×3 convolution to obtain the channel difference feature , screening the channel dimensions that highlight the building changes to capture the spectral difference information. The specific calculation formula is as follows: (1) (2) (3) (4) (5) wherein, and represent the dual-temporal features at the i-th hierarchical levels T1 and T2 respectively, represents the absolute value, represents the convolution operation with a convolution kernel of size k, represents the Hadamard product, represents concatenating the dual-temporal features along the channel dimension.

5. The method for detecting building changes in remote sensing images based on deep features according to claim 3, wherein, The obtaining of the timing difference features is specifically carried out as follows: To capture the temporal difference information, DCEM will send it into the Temporal Difference Awareness Module (TDA) containing two ConvGRU units as the input of the first unit, and set the hidden state of the first unit to 0 to clear past information. The output of the first ConvGRU unit will be used as the hidden state of the second unit, and will be used as the input of the second unit. In the second ConvGRU, similar information to will be forgotten according to its hidden state, and the changing part will be retained as the output to explore the temporal difference information and obtain the time difference features .

6. The method for detecting building changes in remote sensing images based on depth features according to claim 5, wherein, The specific representation of the temporal difference awareness module TDA is as follows: (6) (7) Among them, the specific representation of ConvGRU is: (8) (9) (10) (11) Wherein, is the input feature of the TDA module, represents the hidden state at the previous moment. In the first ConvGRU cell, is initialized to 0, is . In the second ConvGRU cell, is initialized to the output of the first cell, is , represents the update gate, represents the reset gate, represents the candidate gate, represents the sigmoid activation function.

7. The method for detecting building changes in remote sensing images based on deep features according to claim 3, characterized in that Step 4.2 is specifically as follows: After obtaining the spatial, channel, and temporal difference features , , , in order to construct the dependency relationships between the cross-domain difference features, a gated unit is used to dynamically fuse the above-mentioned difference features, and the corresponding gated mask weights of each difference feature are obtained by applying a 1×1 convolution and a sigmoid activation function respectively , , . Then, the difference features of each domain perform the Hadamard product of themselves and the corresponding masks to achieve feature screening and the purpose of removing redundant difference information. After that, first add to its own screened feature, that is . Then, add the sum of the screened , to the Hadamard product of . Finally, the final difference feature is obtained through a 1×1 convolution. The specific formula is shown as follows: (12) (13) (14) (15); In the formula, , , respectively represent the gating masks corresponding to the spatial, channel, and temporal difference features of the i-th layer.

8. The method for detecting building changes in remote sensing images based on deep features according to claim 3, wherein, The specific steps of step 5 are as follows: This step includes three stages: generating a coarse-grained change detection result map, refining differential features, and fusing differential features; using the differential features at different levels obtained by the DCEM module , where i = 1,..., 5 as the input. In the stage of generating the coarse-grained change detection result map, the differential features at each level are aggregated together through the FPN method to generate a coarse-grained change detection result map, and the feature with the shallowest layer resolution of 128×128 in the FPN is used as the reference point for calculating the similarity between the differential features and the coarse-grained result map in the subsequent stage; in the stages of refining differential features and fusing differential features, the differential features at each level are obtained by multi-directional fusion of the differential features at different levels in the previous stage. The proportion of the differential feature at the i-th layer and the j-th stage in the fusion process, that is, its corresponding adaptive weight, is obtained by calculating its similarity with as follows: (18) In the formula, represents the differential feature of the adaptive weight. The mean value of the differential feature along the channel direction is regarded as the query vector Q, and the differential feature is regarded as the key-value vector K. σ is the sigmoid activation function, and s represents the differential feature standard deviation along the channel dimension; Differential features in the differential feature fusion stage After the upsampling operation, it is transformed into a two-channel form of 256×256 size through convolution, and then the final prediction result is obtained through softmax; the above network is trained with the training set, and the validation set is used to evaluate the model performance during the training process to obtain the optimal model. During this process, the BCE and Dice loss functions are used to jointly supervise the coarse-grained change detection result map and the final prediction result of the network. The specific calculation formulas are as follows: (26) In the formula, represents the predicted result graph, represents the true label graph, and the loss function constrains both the fine-grained and coarse-grained predicted result graphs.