Building remote sensing image change detection method based on dynamic context aggregation

Through dynamic context aggregation and twin network structure, combined with dynamic context attention module and deformable convolution operation, the shortcomings of existing remote sensing image change detection methods in multi-scale feature interaction and convolution operation are solved, and a more efficient change detection effect is achieved.

CN119992330AActive Publication Date: 2025-05-13SOUTHWEST FORESTRY UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510079654.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-18
Publication Date
2025-05-13
Estimated Expiration
2045-01-18

Smart Images

  • Figure CN119992330A_ABST
    Figure CN119992330A_ABST
Patent Text Reader

Abstract

The invention discloses a building remote sensing image change detection method based on dynamic context aggregation, and belongs to the field of image detection. The method comprises the following steps: constructing a data set and carrying out sample division; a DCA-NET network model is constructed; performing parameter adjustment on the constructed network model; inputting the divided data set into a model for detection to obtain a change result; and carrying out quantitative evaluation on the change result. According to the invention, by introducing a dynamic context attention mechanism, the model can better integrate local and global information, so that the detection capability of a target change area is improved; compared with the problems that the correlation of neighborhood features is not fully considered and the standard convolution operation is limited to a fixed local receptive field in the prior art, the shape and the size of the convolution kernel can be flexibly adjusted, and the capturing ability of the model in a large-scale global structure and small-scale detail features is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image processing, and in particular to a dynamic context aggregation building remote sensing image change detection method. Background Art

[0002] Change detection based on remote sensing technology aims to discover and identify the differences of ground objects using two or more images of the same geographical location at different times. Remote sensing change detection provides an important technical means for in-depth understanding of various surface changes caused by natural and human activities, and has been widely used in disaster monitoring, forest resource surveys, urban planning, agricultural production and other fields. In the past few decades, change detection technology has been widely studied and various change detection algorithms have been developed. Traditional change detection methods, such as image difference method, background subtraction method, threshold method, etc., are mostly based on manually designed spatiotemporal features, and class changes are identified by clustering or thresholding. However, with the emergence of massive and multimodal remote sensing data, the various limitations of such methods in complex scenes have also become prominent. In particular, manually designed features are difficult to describe the essential differences between land classes, which has become a bottleneck problem in improving the accuracy of change detection. Deep learning methods can adaptively learn from low-level visual cues to high-level semantic features, highlighting the essential differences between land classes. In recent years, they have been widely introduced into remote sensing image change detection research, greatly improving the accuracy of remote sensing change detection. Existing change detection methods are mainly based on deep learning networks, including early network models such as FC-EF, FC-Siam-conc and FC-Siam-diff based on the U-Net architecture, as well as improved models such as DSIFN that introduces the attention mechanism and BIT based on the Transformer architecture. In terms of feature extraction, existing technologies usually adopt multi-scale feature fusion. For example, SNUNet maintains high-resolution information by establishing dense skip connections between the encoder and the decoder, and UNet++ uses full-scale skip connections to learn low-level position information and high-level semantic information. In terms of feature interaction, TFIM proposes a temporal feature interaction module to enhance the perception of changing areas, and A2Net improves the ability to capture detailed changes through gradual feature aggregation.

[0003] However, the current technology has the following shortcomings: (1) Insufficient multi-scale feature interaction. Existing methods based on multi-scale contextual interaction fail to fully consider the relationship between neighborhood features and lack an effective dynamic adjustment mechanism in the feature learning process. (2) Convolution operations are limited by local receptive fields. Traditional convolution operations use fixed-size convolution kernels for feature extraction, which makes it difficult to capture both large-scale global structural features and small-scale local detail features at the same time. When performing feature extraction, the shape and size of the convolution kernel in existing methods are often fixed, and cannot be adaptively adjusted according to target features of different scales, which affects the feature expression ability of the model. Summary of the invention

[0004] To solve the above problems, the present invention provides a building remote sensing image change detection method based on dynamic context aggregation.

[0005] To implement the above technology, the steps are as follows:

[0006] S1, build the data set and divide the samples;

[0007] The dataset is obtained by augmenting the open source dataset;

[0008] The data augmentation used for open source datasets includes: geometric transformation and image enhancement;

[0009] Geometric changes include: random horizontal flipping, random vertical flipping, and random rotation;

[0010] Image enhancement includes: random adjustment of brightness, random adjustment of contrast, random adjustment of color and random addition of Gaussian noise;

[0011] The sample was divided according to the ratio of 8:1:1;

[0012] S2, build DCA-NET network model;

[0013] DCA-NET consists of a twin network structure, including: encoder, decoder, fusion module and dynamic context attention (DCA) module; the DCA module includes: upper GCP layer and lower DCK layer; the encoder includes: four patch embedding layers (Patch Embedding Layer) and encoder block; the encoder block includes: DCA module and convolutional multi-layer perceptron (MLP) unit;

[0014] The construction steps are as follows:

[0015] S2.1. Input the image into the encoder for encoding operation. The steps are as follows:

[0016] S2.1.1, input one image (H×W×3) from a pair of images in the dataset into the patch embedding layer in the encoder to split it into smaller patches;

[0017] S2.1.2, the image after the patch embedding layer enters the GCP layer in the DCA module for context fusion operation;

[0018] Specifically, the input image passes through the upper GCP layer in the DCA module, and the key, query, and value in the input two-dimensional feature map are defined as K = XW K ,Q=XW Q ,V=XW v; Different from the traditional self-attention mechanism that encodes each key independently through 1×1 convolution, the DCA module performs convolution on all neighboring keys in the grid through group convolution operation to achieve contextual key representation; the learned contextual keys reflect the static contextual information between local neighboring keys and are regarded as static contextual representation of the input;

[0019] The present invention concatenates the context key with the query and performs two consecutive 1×1 convolutions (W θ Using ReLU activation, W δ Without activation function), the attention matrix is ​​calculated as follows:

[0020] A=[K 1 ,Q]W θ W δ

[0021] Where W θ is the weight matrix of the first 1×1 convolution, W δ is the weight matrix of the second 1×1 convolution; K 1 Represents a context key;

[0022] For each attention head, the local attention at each spatial position in the attention matrix A is learned jointly based on the query features and the key features of the context, rather than relying solely on isolated query and key pairs; in this way, the additional guidance of static contextual information enhances the learning ability of self-attention; then, using the contextual attention matrix A, all values ​​are summarized to calculate the weighted feature map K 2 :

[0023] K 2 =V*A

[0024] In the formula, * represents the weighted sum operation;

[0025] S2.1.3, the image after the patch embedding layer enters the GCK layer in the DCA module for variable convolution operation;

[0026] The sampling grid of regular convolution is centered at (0,0), while irregular convolution kernels do not have a fixed center at various sizes; therefore, in order to adapt to the sizes of different convolution kernels, the present invention defines the upper left corner (0,0) as the sampling origin in the algorithm. On this basis, the initial sampling coordinates of irregular convolution are determined, and then the corresponding convolution operation is defined at position F0 as:

[0027] Conv(F0)=∑w×(F0+F n )

[0028] Where w represents the convolution parameter; F n Indicates offset at position F0;

[0029] The sampling position of standard convolution is fixed, which can only extract information within the local window and cannot capture contextual features in a larger range. Deformable convolution compensates for this limitation to a certain extent by learning offsets and dynamically adjusting the sampling grid. Specifically, the sampling grids of standard convolution and deformable convolution are still regular and cannot adapt to convolution kernels of arbitrary shapes or numbers of parameters. To solve this problem, DCA introduces a more flexible variable convolution operation DCK layer. The deformable convolution of DCA first generates an offset through the convolution layer, and its dimension is (B, 2N, H, W), where N represents the size of the convolution kernel.

[0030] S2.1.4, the DCP output result and the DCK output result are spliced ​​in the channel dimension as the output of the DCA module, which reduces the number of feature channels and promotes the further integration of information from different modules;

[0031] Specifically, by fusing static and dynamic context representations, the DCA module can form a more discriminative feature expression, thereby improving the accuracy of change detection. The outputs of these two parts are then merged in the channel dimension to achieve complementary advantages. Subsequently, the merged features are integrated through a 1x1 convolutional layer, which not only reduces the number of feature channels, but also promotes the further fusion of information from different modules.

[0032] S2.1.5, down-sampling the output of the DCA module;

[0033] Specifically, the output of the DCA module is downsampled by 1 / 4 to obtain a Image The output of the DCA module is downsampled by 1 / 8 to obtain a Image The output of the DCA module is downsampled by 1 / 16 to obtain a Image The output of the DCA module is downsampled by 1 / 32 to obtain a Image

[0034] At the same time, another image is subjected to downsampling operation after the operations of S2.1.1 to S2.1.4, including: performing 1 / 4 downsampling operation to obtain a Image The output of the DCA module is downsampled by 1 / 8 to obtain a Image The output of the DCA module is downsampled by 1 / 16 to obtain a Image The output of the DCA module is downsampled by 1 / 32 to obtain a Image

[0035] S2.2, performing a fusion operation on the encoded images;

[0036] Specifically, images and images image and images image and images and images and images Perform image feature fusion to obtain multi-scale images and

[0037] S2.3, inputting the fused image into a decoder for decoding operation;

[0038] Specifically, the multi-scale image and After connecting along the channel dimension, the 1×1 convolution layer is input for dimensionality reduction. After the transposed convolution operation, the spatial resolution of the feature map is gradually improved. The result of the transposed convolution operation passes through two 3×3 convolution layers and then performs the first element-by-element addition operation with the result of the transposed convolution operation. After the first element-by-element addition operation, the result passes through the transposed convolution operation, passes through two 3×3 convolution layers, and then performs the second element-by-element addition operation with the result of the transposed convolution operation. The result of the second element-by-element addition operation passes through the 3×3 convolution layer to obtain the final output result.

[0039] S3, adjusting parameters of the constructed network model;

[0040] S4, input the divided data set into the model for detection to obtain the change result;

[0041] The training set is input into the model for model training. After the trained model is input into the test set for testing, the validation set is input into the model for model verification. The verified model is used for change detection.

[0042] S5. Quantitatively evaluate the change results;

[0043] Quantitative evaluation methods include: F1-score (F1), Intersection over Union (IoU) and Overall Accuracy (OA);

[0044] Beneficial effects of the present invention:

[0045] In view of the problem that the existing technology does not fully consider the relationship between neighborhood features and the standard convolution operation is limited to a fixed local receptive field, the present invention uses dual-time images (captured by the same type of sensor and with similar features) as input and adopts a twin network structure with shared weights to effectively extract image features. The model uses rich contextual information to optimize the learning of the dynamic attention matrix and further enhances the ability to recognize changing areas. At the same time, through a novel coordinate generation algorithm, the initial positions of convolution kernels of different sizes are defined, so that the shape and size of the convolution kernel can be flexibly adjusted, thereby improving the model's ability to capture large-scale global structures and small-scale detail features.

[0046] By introducing a dynamic contextual attention mechanism, the model can better integrate local and global information, thereby improving the ability to detect target change areas; secondly, the proposed coordinate generation algorithm allows the initial position of the convolution kernel to be flexibly adjusted, enhancing the adaptability of the model to targets of different scales and shapes. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flow chart of the steps of the present invention;

[0048] Figure 2 It is the overall technical roadmap in the embodiment of the present invention;

[0049] Figure 3 is a change detection network of the present invention;

[0050] Figure 4 It is a schematic diagram of the DCA structure of the present invention;

[0051] Figure 5 is a schematic diagram of an encoder block of the present invention;

[0052] Figure 6 It is the dynamic context attention module proposed by the present invention;

[0053] Figure 7 is a decoder used in the present invention;

[0054] Figure 8 Comparison of change detection results of different methods on CDD-CD dataset;

[0055] Fig. 9 Comparison of change detection results of different methods on the LEVIR-CD dataset. DETAILED DESCRIPTION

[0056] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme is further described in detail in combination with the embodiments of the present invention and the accompanying drawings. It should be pointed out that the embodiments are only examples of part of the present invention, not all. All other embodiments that can be implemented by ordinary technicians in this field based on the embodiments of the present invention without the need for creative work should be deemed to belong to the protection scope of the present invention.

[0057] like Figure 1 and Figure 2 As shown, a dynamic context aggregation building remote sensing image change detection method includes the following steps:

[0058] S1, build the data set and divide the samples;

[0059] The dataset is obtained by augmenting the open source dataset;

[0060] The open source dataset LEVIR-CD dataset contains 637 pairs of dual-phase remote sensing images obtained from Google Earth, with a spatial resolution of 0.5 m / pixel; the open source dataset CDD-CD dataset contains 11 pairs of remote sensing images with seasonal changes, of which 7 pairs have a resolution of 4725×2700 pixels and 4 pairs have a resolution of 1900×1000 pixels; to facilitate deep learning model training, the images in the open source dataset are cut into image blocks of 256×256 pixels;

[0061] The data augmentation used for open source datasets includes: geometric transformation and image enhancement;

[0062] Geometric changes include: random horizontal flipping, random vertical flipping, and random rotation;

[0063] Image enhancement includes: random adjustment of brightness, random adjustment of contrast, random adjustment of color and random addition of Gaussian noise;

[0064] Through data augmentation methods, the training samples can be effectively expanded and the model's adaptability to various scenarios can be improved. During the training process, the above transformations are randomly applied to the input images, so that the model can learn more robust feature representations;

[0065] The sample was divided according to the ratio of 8:1:1;

[0066] The final results are shown in Table 1:

[0067] Table 1: Dataset division

[0068]

[0069] S2, such as Figure 3 As shown, the DCA-NET network model is constructed;

[0070] Change detection is a multi-scale, multi-level feature learning problem. Especially in remote sensing images, different types of changing targets (such as buildings, vegetation, water bodies, etc.) often have diverse appearances and forms. In order to address the limitations of existing methods in capturing multi-scale and complex background change features, this paper proposes DCA-NET (Dynamic Contextual Attention Network), which improves the change detection capability of the model through a dynamic context aggregation strategy.

[0071] like Figure 3 As shown in Figure 1, DCA-NET consists of a twin network structure, including: encoder, decoder, fusion module and dynamic context attention (DCA) module; the DCA module includes: the upper GCP layer and the lower DCK layer, as shown in Figure 1. Figure 4 As shown; the encoder includes: four patch embedding layers (Patch Embedding Layer) and an encoder block; the encoder block includes: a DCA module and a convolutional multi-layer perceptron (MLP) unit, as shown Figure 5 As shown;

[0072] This architecture design enhances the model's feature expression and detection accuracy in complex scenarios by combining static and dynamic context information and multimodal feature fusion;

[0073] The construction steps are as follows:

[0074] S2.1. Input the image into the encoder for encoding operation. The steps are as follows:

[0075] S2.1.1. Input one image (H×W×3) from a pair of images in the dataset into the patch embedding layer of the encoder and split it into smaller patches for multi-scale processing.

[0076] S2.1.2, if Figure 6 As shown, the image after the patch embedding layer enters the GCP layer in the DCA module for context fusion operation;

[0077] Specifically, the input image passes through the upper GCP layer in the DCA module, and the key, query, and value in the input two-dimensional feature map are defined as K = XW K ,Q=XW Q ,V=XW v; Different from the traditional self-attention mechanism that encodes each key independently through 1×1 convolution, the DCA module performs convolution processing on all neighboring keys in the grid through group convolution operation, thereby realizing contextual key representation; the learned context keys reflect the static context information between local neighboring keys and are regarded as the static context representation of the input; the present invention concatenates the context keys with the query and performs two consecutive 1×1 convolutions (W θ Using ReLU activation, W δ Without activation function), the attention matrix is ​​calculated as follows:

[0078] A=[K 1 ,Q]W θ W δ

[0079] Where W θ is the weight matrix of the first 1×1 convolution, W δ is the weight matrix of the second 1×1 convolution; K 1 Represents a context key;

[0080] For each attention head, the local attention at each spatial position in the attention matrix A is learned jointly based on the query features and the key features of the context, rather than relying solely on isolated query and key pairs; in this way, the additional guidance of static contextual information enhances the learning ability of self-attention; then, using the contextual attention matrix A, all values ​​are summarized to calculate the weighted feature map K 2 :

[0081] K 2 =V*A

[0082] In the formula, * represents the weighted sum operation;

[0083] The weighted feature graph captures the dynamic feature interaction between the input data, which is referred to as the dynamic context representation of the input data in the present invention. Therefore, the output of the GCP layer is the result of fusing the static context and the dynamic context through the attention mechanism, thereby enhancing the model's ability to express complex feature relationships. This mechanism enables the local attention matrix to be jointly learned based on the context key and the query feature, rather than relying solely on isolated query and key pairs, thereby significantly enhancing the ability of the self-attention mechanism in feature interaction.

[0084] S2.1.3, if Figure 6 As shown, the image after the patch embedding layer enters the GCK layer in the DCA module for variable convolution operation;

[0085] Convolutional neural networks (CNNs) are based on convolution operations, which map features to corresponding positions through a regular sampling grid. Usually, the sampling grid of convolution is regular, but in order to make irregularly shaped convolution kernels fit the sampling grid, DCA introduces an algorithm that supports convolution kernels of any size to generate the initial sampling coordinates of the convolution kernel; the algorithm first generates a regular sampling grid, then creates an irregular grid for the remaining sampling points, and finally splices the two to generate a complete sampling grid;

[0086] The sampling grid of regular convolution is centered at (0,0), while irregular convolution kernels do not have a fixed center at various sizes; therefore, in order to adapt to the sizes of different convolution kernels, the present invention defines the upper left corner (0,0) as the sampling origin in the algorithm. On this basis, the initial sampling coordinates of irregular convolution are determined, and then the corresponding convolution operation is defined at position F0 as:

[0087] Conv(F0)=∑w×(F0+F n )

[0088] Where w represents the convolution parameter; F n Indicates offset at position F0;

[0089] The sampling position of standard convolution is fixed, and it can only extract information within a local window, but cannot capture contextual features in a larger range. Deformable convolution compensates for this limitation to a certain extent by learning offsets and dynamically adjusting the sampling grid. Specifically, the sampling grids of standard convolution and deformable convolution are still regular and cannot adapt to convolution kernels of arbitrary shapes or numbers of parameters. To solve this problem, DCA introduces a more flexible variable convolution operation DCK layer. The deformable convolution of DCA first generates an offset through the convolution layer, and its dimension is (B, 2N, H, W), where N represents the size of the convolution kernel.

[0090] In this embodiment, N=3, the generated offset is added to the original sampling point coordinates to obtain a new sampling position; finally, the feature information at the position is obtained by interpolation and resampling;

[0091] This method effectively enhances the expressiveness of the convolution kernel by flexibly adjusting the sampling position. It not only effectively enhances the flexibility of the convolution kernel, but also improves the feature extraction efficiency of the model under different scales and complex backgrounds.

[0092] S2.1.4, the DCP output result and the DCK output result are spliced ​​in the channel dimension as the output of the DCA module, which reduces the number of feature channels and promotes the further integration of information from different modules;

[0093] Specifically, by fusing static and dynamic context representations, the DCA module can form a more discriminative feature expression, thereby improving the accuracy of change detection. The outputs of these two parts are then merged in the channel dimension to achieve complementary advantages. Subsequently, the merged features are integrated through a 1x1 convolutional layer, which not only reduces the number of feature channels, but also promotes the further fusion of information from different modules.

[0094] S2.1.5, down-sampling the output of the DCA module;

[0095] Specifically, the output of the DCA module is downsampled by 1 / 4 to obtain a Image The output of the DCA module is downsampled by 1 / 8 to obtain a Image The output of the DCA module is downsampled by 1 / 16 to obtain a Image The output of the DCA module is downsampled by 1 / 32 to obtain a Image

[0096] At the same time, another image is subjected to downsampling operation after the operations of S2.1.1 to S2.1.4, including: performing 1 / 4 downsampling operation to obtain a Image The output of the DCA module is downsampled by 1 / 8 to obtain a Image The output of the DCA module is downsampled by 1 / 16 to obtain a Image The output of the DCA module is downsampled by 1 / 32 to obtain a Image

[0097] S2.2, performing a fusion operation on the encoded images;

[0098] Specifically, images and images image and images image and images and images and images Perform image feature fusion to obtain multi-scale images and

[0099] S2.3, inputting the fused image into a decoder for decoding operation;

[0100] like Figure 7 As shown in the figure, the decoder of DCA-NET consists of multiple layers of convolution and transposed convolution operations, which is used to gradually restore high-resolution change detection maps;

[0101] The input of the decoder is the multi-scale features output by the encoder. These features are first concatenated along the channel dimension and reduced in dimension through a 1×1 convolutional layer to reduce computational complexity and prevent overfitting. The decoder gradually increases the spatial resolution of the feature map using a transposed convolution operation, and further optimizes the feature representation through a series of residual blocks to maintain the consistency of local details and global structure. The residual blocks enhance the expressiveness of the feature map through skip connections without increasing the computational burden. Finally, a convolutional layer is applied to generate a prediction score consisting of two channels, where one channel predicts the unchanged area and the other predicts the changed area, and an Argmax operation is used along the channel dimension to obtain a binary change map.

[0102] Specifically, the multi-scale image and After connecting along the channel dimension, the 1×1 convolution layer is input for dimensionality reduction. After the transposed convolution operation, the spatial resolution of the feature map is gradually improved. The result of the transposed convolution operation passes through two 3×3 convolution layers and then performs the first element-by-element addition operation with the result of the transposed convolution operation. After the first element-by-element addition operation, the result passes through the transposed convolution operation, passes through two 3×3 convolution layers, and then performs the second element-by-element addition operation with the result of the transposed convolution operation. The result of the second element-by-element addition operation passes through the 3×3 convolution layer to obtain the final output result.

[0103] S3, adjusting parameters of the constructed network model;

[0104] In this embodiment, all experiments are implemented based on the PyTorch deep learning framework, and the experimental hardware platform adopts a computing environment equipped with a single NVIDIA Tesla T4 GPU; in order to improve the robustness and generalization ability of the model, the present invention adopts a variety of data enhancement strategies, including random flipping, proportional cropping, color jittering and Gaussian blur;

[0105] During the training process, the batch size is set to 6 and the initial learning rate is set to 3.5e -4 , and the linear decay strategy is used to dynamically adjust the learning rate until the end of training; the model optimization uses Pixel-wise Cross Entropy as the loss function, and the parameters are updated through the back propagation algorithm. The number of training iterations is 300;

[0106] S4, input the divided data set into the model for detection to obtain the change result;

[0107] The training set is input into the model for model training. After the trained model is input into the test set for testing, the validation set is input into the model for model verification. The verified model is used for change detection.

[0108] Figure 8 and Fig. 9 The change detection results of different methods on CDD-CD and LEVIR-CD datasets are compared; from the visualization results, it can be observed that although each method can detect the approximate change area of ​​the building, there are significant differences in the details; among the compared methods, the FC-Siam-Conc and FC-Siam-Diff methods are greatly affected by pseudo-changes, resulting in high missed detection rates; DTCDSCN is difficult to effectively distinguish important features in a complex noise environment, and the detection area is incomplete; the BIT model has the problem of loss of mid- and bottom-level detail information during upsampling, and the detection area is incomplete. ChangeFormer still needs to be improved in terms of target integrity and boundary clarity, especially for large-scale buildings. In contrast, the DCA-Net proposed in the present invention can effectively extract small-scale and large-scale spatial features through the dynamic contextual attention mechanism, making the boundary contour of the building more complete; at the same time, the multi-scale feature fusion strategy is used to significantly improve the false detection and missed detection phenomena;

[0109] S5. Quantitatively evaluate the change results;

[0110] The quantitative evaluation methods include: F1-score (F1), Intersection over Union (IoU) and Overall Accuracy (OA); the expressions are as follows:

[0111]

[0112] In the formula, TP represents the number of samples correctly predicted as positive; TN represents the number of samples correctly predicted as negative; FP represents the number of samples incorrectly predicted as positive; FN represents the number of samples incorrectly predicted as negative; positive refers to pixels of the changed class; negative refers to pixels of the non-changed class;

[0113] The comparison of the evaluation results with other mainstream methods is shown in Table 2

[0114] Table 2: Performance comparison of different models on the LEVIR-CD dataset and CDD-CD dataset

[0115]

[0116] The present invention is compared and analyzed with other methods on the LEVIR-CD dataset and the CDD-CD dataset; FC-Siam-Di, FC-EF, FC-Siam-Conc, DTCDSCN, BIT, ChangeFormer and ScratchFormer are all methods that are currently widely used. Table 2 shows the comparison results of each method on the dataset. On the LEVIR-CD dataset, the F1-score of the method of the present invention reaches 91.69%, which is 1.29 percentage points higher than the currently better ChangeFormer; IoU reaches 85.89%, an increase of 3.41 percentage points; OA reaches 99.11%, an increase of 0.07 percentage points. On the CDD-CD dataset, the performance advantage is more significant, with F1-score reaching 95.5%, an increase of 5.67 percentage points over ChangeFormer; IoU reaches 91.77%, an increase of 10.24 percentage points; OA reaches 99.12%, an increase of 1.44 percentage points. The experimental results show that the DCA model proposed in the present invention can effectively improve the accuracy of change detection.

[0117] In summary, the DCANet proposed in this patent can effectively extract small-scale and large-scale spatial features through the dynamic context attention mechanism, making the building boundary contour more complete; at the same time, with the help of multi-scale feature fusion strategy, it significantly improves the false detection and missed detection phenomenon, and shows stable detection performance for buildings of different sizes and densities. In the qualitative comparison of the two data sets, the detection results of DCANet are closest to the true value labels, which effectively solves the problems of fragmented and incomplete detection results, and fully verifies the superiority of the proposed method.

Claims

1. A method for detecting changes in building remote sensing images based on dynamic context aggregation, characterized in that: The following steps are involved: S1, build the data set and divide the samples; The dataset is obtained by augmenting the open source dataset; The data augmentation used for open source datasets includes: geometric transformation and image enhancement; Geometric changes include: random horizontal flipping, random vertical flipping, and random rotation; Image enhancement includes: random adjustment of brightness, random adjustment of contrast, random adjustment of color and random addition of Gaussian noise; The sample was divided according to the ratio of 8:1:1; S2, build DCA-NET network model; The DCA-NET comprises: an encoder, a decoder, a fusion module and a dynamic context attention module; wherein the dynamic context attention module comprises: an upper GCP layer and a lower DCK layer; the encoder comprises: four patch embedding layers and an encoder block; the encoder block comprises: a DCA module and a convolutional multilayer perceptron unit; S3, adjusting parameters of the constructed network model; S4, input the divided data set into the model for detection to obtain the change result; S5. Quantitatively evaluate the change results; Quantitative evaluation methods include: F1-score, intersection over union, and overall accuracy.

2. The method for detecting changes in building remote sensing images based on dynamic context aggregation according to claim 1, characterized in that: The steps for constructing the DCA-NET network model are as follows: S2.1, input the image into the encoder for encoding operation; S2.2, performing a fusion operation on the encoded images; S2.

3. Input the fused image into the decoder for decoding operation.

3. The method for detecting changes in building remote sensing images based on dynamic context aggregation according to claim 2, characterized in that: The steps of the encoding operation are as follows: S2.1.1, one of the images in a pair in the dataset is input into the patch embedding layer of the encoder to split it into smaller patches; S2.1.2, the image after the patch embedding layer enters the GCP layer in the DCA module for context fusion operation; S2.1.3, the image after the patch embedding layer enters the GCK layer in the DCA module for variable convolution operation; S2.1.4, concatenate the DCP output result and the DCK output result in channel dimension as the output of the DCA module to reduce the number of feature channels; S2.1.5, down-sampling the output of the DCA module; The down sampling operation includes: 1 / 4 down sampling, 1 / 8 down sampling, 1 / 16 down sampling and 1 / 32 down sampling.

Citation Information

Patent Citations

  • Remote sensing image building change detection method based on multi-scale attention

    CN118298305A