Change detection method based on semantic priors and fusion spatial localization

By pre-training the BPNet network on the remote sensing extraction task and introducing a shared-weight Siamese network and a spatial consistency attention module, the data scarcity problem in building change detection is solved, and the performance and adaptability of the model in change detection tasks are improved.

CN119478542BActive Publication Date: 2025-12-26HOHAI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411683114.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-12-26
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

Existing change detection methods face the problem of data scarcity in building change detection tasks, and traditional pre-training datasets lack semantic features from a remote sensing perspective, making it difficult to effectively utilize data from similar tasks, thus limiting model performance.

Method used

A change detection method based on semantic prior and fusion spatial localization is adopted. By pre-training a BPNet network on a supervised remote sensing extraction task, a Siamese network structure with shared weights and a spatial consistency attention module are introduced. Combined with self-attention mechanism and large convolutional kernel decomposition technique, feature information is extracted and fused from dual temporal images.

Benefits of technology

It significantly improves the model's generalization ability and performance in change detection tasks, reduces the high cost of training from scratch and the risk of overfitting, and is able to better capture the details of bi-temporal changes and long-distance dependencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478542B_ABST
    Figure CN119478542B_ABST
Patent Text Reader

Abstract

The application discloses a change detection method based on semantic prior and fusion spatial positioning, relates to the technical field of image processing, and comprises the following steps: firstly, pre-training a BPNet network on a supervised remote sensing extraction task, and fully utilizing prior knowledge in the remote sensing extraction task; then, fine-tuning the pre-trained BPNet network, adopting a twin network structure with shared weights to expand the encoder and decoder structure, and introducing a spatial consistency attention module; and finally, combining a self-attention mechanism and a large convolution kernel decomposition technology to extract and fuse feature information of double-time images, so as to obtain remote sensing image target change information. Therefore, the change detection method based on semantic prior and fusion spatial positioning, through the spatial consistency attention module, helps to expand the effective receptive field of the deep encoder, enhances the modeling ability of long-distance dependency, and can effectively capture double-time change detail information and obtain more accurate feature maps.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly to a change detection method based on semantic prior and fusion spatial positioning. BACKGROUND

[0002] Change detection (CD) aims to identify ground changes by analyzing multi-temporal remote sensing images taken at different times of the same geographical area. Binary change detection (BCD) refers to the process of identifying changed objects of interest under the premise of known binary labels. Among them, building change detection is a binary change detection task focusing on buildings, and is widely used in various fields such as urban development monitoring, disaster assessment, and land use planning.

[0003] The development of CD technology is closely related to the progress of earth observation technology, the evolution of information technology (IT), and the rise of artificial intelligence. Initially, traditional CD mainly has two methods: pixel-based method and object-based method. The pixel-based change detection (PBCD) method mainly obtains a change map by performing pixel-by-pixel operations, and then determines the change area by setting a threshold. The object-based change detection (OBCD) method uses spectral, texture, and spatial context information to capture object-level changes. However, these methods still face some challenges. The PBCD method pays attention to the semantic information of each pixel, but ignores the spatial relationship between pixels. The OBCD method is very sensitive to the selection of object segmentation algorithms, and lacks effective modeling of pixel semantics, making it difficult to distinguish between false changes. In addition, these methods show limited adaptability when facing scene changes, and need to rely on a large amount of manual adjustment, and their accuracy is also easily questioned.

[0004] With the emergence of large-scale remote sensing data and the rapid development of deep learning technology, deep learning has shown great potential in change detection tasks in the field of remote sensing. Fully convolutional Siamese networks are considered the first deep learning method for change detection, opening up the extensive application of deep learning in this field. Current change detection methods include CNN-based methods (such as Changer, SGSLN), Transformer-based methods (such as ChangeFormer, BiT), and state space model-based methods (such as RSM-CD). Although these deep learning methods have achieved good results, their performance improvement has been limited due to the lack of high-quality labeled data. In addition, in the task of building change detection, obtaining and accurately registering remote sensing images of the same area often requires a large amount of human resources. Due to differences in sensors or changes in environmental conditions, the registration process can produce errors, which in turn limits the number of available image pairs. This data scarcity easily leads to a bottleneck in the performance of deep learning algorithms. In contrast, the scale of the data set for single-phase tasks is much larger, such as the WHU building data set, the Massachusetts building data set, the Inria building data set, and natural image data sets (such as ImageNet and MSCOCO). This imbalance in data size directly affects the training effect of the model in the field of change detection, so how to effectively use the data of similar tasks becomes a key problem in current research.

[0005] Pre-training on large-scale data sets can provide the model with rich prior knowledge of remote sensing images, thereby improving its performance when fine-tuning on the target data set. This is an effective strategy that can alleviate the problem of data scarcity in the target domain. For example, some studies have supervised pre-training on large-scale remote sensing scene classification data sets such as MillionAID to optimize model performance. Other studies have performed multi-task pre-training on the SAMRS data set, covering scene classification, semantic segmentation, instance segmentation, and rotating object detection. There are also studies that use unsupervised pre-training methods, such as using the MAE model on the MillionAID or fMoW Sentinel multispectral data set. In addition, the domain gap between remote sensing images and natural images also limits the performance of change detection, and traditional pre-training data sets lack semantic features (such as buildings, roads, and trees) from the perspective of remote sensing, making it difficult to transfer to remote sensing tasks. Therefore, existing research has not fully utilized semantic segmentation image data in the same visual domain as the building change detection task during the pre-training phase, and how to extend these models to support dual-input functions is still a problem worth exploring.

[0006] Therefore, there is an urgent need for a method integrating the two stages of pre-training and fine-tuning to adapt to multi-temporal image change recognition, improve the sensitivity of capturing the temporal relationship of double temporal images, and solve the problem of data scarcity, providing an effective framework for processing complex remote sensing tasks. SUMMARY

[0007] The purpose of the present application is to provide a change detection method based on semantic priori and fusion spatial positioning, by pre-training on an extraction dataset with highly similar features to the target task change detection, enabling the network to extract priori knowledge of corresponding features, making it more effective in dealing with the challenge of data sparsity in the change detection task.

[0008] To achieve the above purpose, the present application provides a change detection method based on semantic priori and fusion spatial positioning, comprising:

[0009] Firstly, the BPNet network is pre-trained on a supervised remote sensing extraction task, making full use of the priori knowledge in the remote sensing extraction task;

[0010] Then, the pre-trained BPNet network is fine-tuned, the encoder and decoder structure is expanded using a shared weight twin network structure, and a spatial consistency attention module is introduced, the feature information of double temporal images is extracted and fused by combining the self-attention mechanism and the large convolution kernel decomposition technology, and then the target change information of remote sensing images is obtained.

[0011] Preferably, the encoder and decoder of the BPNet network use a feature enhancement module for feature extraction, and the expression of the feature enhancement module is:

[0012]

[0013] In the formula, represents the input feature of the feature enhancement module, represents the feature expanded by the first 1x1 convolution, represents the output feature of the feature enhancement module, n represents the temporal state of the feature map, represents element-level addition, SE represents the self-attention mechanism, DWConv 3×3 respectively represent depth convolution with a convolution kernel of 3x3, Conv 1×1 represents 1x1 convolution.

[0014] Preferably, the spatial consistency attention module includes an ADLKA module and a TFU module: the ADLKA module is used to capture more extensive context information and fuse change information; the TFU module is used to retain the temporal information in the double temporal features.

[0015] Preferably, the expression of the output of the ADLKA module is:

[0016]

[0017] In the formula, represents the double temporal information features mixed after the TFU module, represents the attention map output by the ADLKA module, represents a depth dilated convolution with a convolution kernel of k2x k2, represents a depth convolution with a convolution kernel of k1x k1, 1×1 represents a point-wise convolution.

[0018] Preferably, the TFU module processes each sample independently by using layer normalization, and the expression is as follows:

[0019]

[0020] T l = GELU(LN(X s ));

[0021] In the formula, respectively represent input features of the first temporal state and the second temporal state, represents features reduced by 1x1 convolution, represents the double temporal features fused in time, concat represents channel connection, LN represents layer normalization, and GELU represents an activation function.

[0022] Preferably, a multi-scale fusion segmentation head is introduced after the encoder, and the feature maps extracted by each layer are convolved to generate a prediction map of the corresponding scale. The prediction maps obtained by each layer are adjusted to the size of the original image by bilinear interpolation, and a multi-scale prediction map is generated by fusion.

[0023] Preferably, in the fine-tuned network, deep supervision is introduced, and a multi-scale mask is generated by calculating a loss function at each layer of the network. The expression of the loss function is as follows:

[0024]

[0025] In the formula, l i represents the loss function of the i-th layer, Bce and Dice respectively represent a binary cross loss function and a similarity coefficient loss function, λ i represents the loss function coefficient of different layers, y、 respectively represent a real label and a predicted label.

[0026] Therefore, the change detection method based on the semantic prior and the fusion spatial positioning has the following technical effects:

[0027] (1) Using the "pre-training + fine-tuning" technology, first, the generalization ability and performance of the model are significantly improved through pre-training, and then based on the pre-trained model, through fine-tuning, the model is additionally trained on a specific dataset, so as to adapt to a specific task or field, while reducing the high cost of training the model from scratch and the risk of overfitting.

[0028] (2) The twin network structure with shared weights is used to expand the encoder and decoder structure, which can solve the adaptation problem of double image input; At the same time, the spatial consistency attention module is introduced, which enhances the modeling ability of long-distance dependence through large convolution kernel, thereby improving the performance of the model, and by fusing the features in the encoder branch, a feature map containing change information is generated, thereby better capturing the details of the double time change.

[0029] The technical solutions of the present application will be described in further detail below with the help of the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is the BPNet network structure diagram in the embodiment of the change detection method based on semantic priori and fusion spatial positioning;

[0031] Figure 2 is the fine-tuning network structure diagram in the embodiment of the change detection method based on semantic priori and fusion spatial positioning;

[0032] Figure 3 is the schematic diagram of each module in the fine-tuning network in the embodiment of the change detection method based on semantic priori and fusion spatial positioning, wherein (a) is the SCAM module, (b) is the ADLKA module, (c) is the TFU module, and (d) is the LFEM module;

[0033] Figure 4 is the Grad-CAM heat map visualization result corresponding to the LKA module in the embodiment of the change detection method based on semantic priori and fusion spatial positioning, wherein (a) is the first time image, (b) is the first time feature map of the first layer of the encoder, (c) is the first time feature map of the second layer of the encoder, (d) is the first time feature map of the third layer of the encoder, (e) is the first time feature map of the fourth layer of the encoder, (f) is the first time feature map of the fifth layer of the encoder, (g) is the second time image, (h) is the second time feature map of the first layer of the encoder, (k) is the second time feature map of the second layer of the encoder, (l) is the second time feature map of the third layer of the encoder, (m) is the second time feature map of the fourth layer of the encoder, and (n) is the second time feature map of the fifth layer of the encoder;

[0034] Figure 5Grad-CAM heat map visualization results of the SCAM module corresponding to the embodiment of the change detection method based on semantic priori and fusion spatial positioning, wherein (a) is the first time image, (b) is the first time feature map of the first layer of the encoder, (c) is the first time feature map of the second layer of the encoder, (d) is the first time feature map of the third layer of the encoder, (e) is the first time feature map of the fourth layer of the encoder, (f) is the first time feature map of the fifth layer of the encoder, (g) is the second time image, (h) is the second time feature map of the first layer of the encoder, (k) is the second time feature map of the second layer of the encoder, (l) is the second time feature map of the third layer of the encoder, (m) is the second time feature map of the fourth layer of the encoder, and (n) is the second time feature map of the fifth layer of the encoder;

[0035] Figure 6 The change detection result of scene one in the embodiment of the change detection method based on semantic priori and fusion spatial positioning, (a) is the first time image, (b) is the second time image, (c) is the real change map, (d) is the change detection result of the benchmark network, (e) is the change detection result after superimposing the LKA module, and (f) is the change detection result after superimposing the SCAM module;

[0036] Figure 7 The change detection result of scene two in the embodiment of the change detection method based on semantic priori and fusion spatial positioning, (a) is the first time image, (b) is the second time image, (c) is the real change map, (d) is the change detection result of the benchmark network, (e) is the change detection result after superimposing the LKA module, and (f) is the change detection result after superimposing the SCAM module;

[0037] Figure 8 The change detection result of scene three in the embodiment of the change detection method based on semantic priori and fusion spatial positioning, (a) is the first time image, (b) is the second time image, (c) is the real change map, (d) is the change detection result of the benchmark network, (e) is the change detection result after superimposing the LKA module, and (f) is the change detection result after superimposing the SCAM module. DETAILED DESCRIPTION

[0038] The present application can be explained in more detail by the following examples, and the purpose of the disclosure is to protect all changes and improvements within the scope of the present application, and the present application is not limited to the following examples.

[0039] In this embodiment, building changes are taken as the detection target to illustrate the framework and effect of the change detection method based on semantic priori and fusion spatial positioning, and the specific process is as follows:

[0040] Building change detection can be seen as a specific extension of building extraction task, whose core goal is to selectively identify and extract those buildings that have changed by analyzing the dual-time images, while ignoring the unchanged parts. In order to design an effective building change detection model, it is crucial to let the network learn the prior knowledge of building extraction first. This strategy enables the model to transform the change detection into a binary classification problem, i.e., judging whether the extracted building has changed or not, after mastering the basic ability of building extraction. To this end, in the present embodiment, the BPNet network is pre-trained on the supervised building extraction task to make full use of the prior knowledge in the building extraction task.

[0041] wherein the BPNet network is a classical U-shaped structure, as shown in Figure 1 including an encoder of five stages, a decoder of five stages, a Van module, and a multi-scale fusion module attached to the decoder stage, and the specific design is as follows:

[0042] (1) Encoder and decoder structure: the encoder and the decoder are respectively composed of two LFEMs. The first LFEM is used to adjust the dimension of the input feature, and the second LFEM focuses on extracting deep features. Before the encoder, a preliminary feature extraction module (stem) composed of 3x3 convolution, batch normalization and SiLU activation function is first used to expand the feature dimension. Between each encoder, a two-fold down-sampling is performed by max pooling. While in the decoder stage, up-sampling is performed by bilinear interpolation, which can effectively reduce the parameter quantity and computational quantity of the model.

[0043] (2) Integration of attention module: in the U-shaped structure, a Van module (Van Block) is configured between the symmetrical encoder and the decoder, which fuses the feature maps of the encoder stage and the up-sampled feature maps of the decoder stage through a skip connection. The input of each decoder stage is the sum of the up-sampled feature map of the previous stage and the symmetrical encoder feature map processed by the Van module.

[0044] (3) Processing and output of feature maps: after each decoder stage, the output feature map is saved, and finally five saliency feature maps are obtained. These feature maps are first processed by the first convolutional network of the multi-scale module and applied with sigmoid activation function; then, these processed feature maps are spliced by channel, and further processed by the second convolutional network of the multi-scale module, and finally the final output result is generated by the sigmoid function.

[0045] To adapt the pre-trained model to specific downstream tasks, fine-tuning transfer learning technology is needed. However, due to the special data structure of the change detection task involving dual image input, directly fine-tuning the semantic segmentation model trained based on single image is not applicable, and specific adjustments need to be made according to the characteristics of dual image input to ensure the effectiveness and accuracy of the model when processing the change detection task. Based on this, in the embodiment, a change enhancement network (CAN) is proposed, which assumes that the feature module F o has a fixed capacity, consisting of K layers L k ,k=1…K, and the hidden features of each layer are c k ∈R nk , where n k is the number of units in the kth layer. Let W k represent the weight matrix between the kth layer and the k-1th layer, and include the bias term, then c k =f(W k c k-1 ), where f(·) is a nonlinear function such as SiLU activation function. The fine-tuning network of the embodiment forms a change-enhanced representation module F k by constructing a new layer L c after the original network layer L c . L c can be regarded as a change detection adaptation layer, allowing new combinations of existing operations to avoid drastic modifications to the pre-trained layers to adapt to new tasks. The representation of layer L c is c c =f(W c c k ), where W c represents the weight matrix between layer L c and L k , and c c is the newly generated feature.

[0046] In summary, based on the specific needs of dual-time image input, two important improvements are made to the BPNet network structure based on traditional fine-tuning:

[0047] First, expand the encoder and decoder structure. To adapt to the input of dual-time images in the change detection task, this embodiment uses a twin network structure with shared weights, allowing the same encoder and decoder to process two different time point image inputs simultaneously and ensuring that both use the same pre-trained weights. The assumption underlying this approach is that t1 and t2 images come from the same distribution, i.e. By sharing the weights, the twin network is able to extract consistent feature representations between the two time points. Specifically, when the encoder and decoder networks receive images at t1 and t2 as input, respectively, the generated feature representations are c 1 k = F o (t 1 ), c 2 k = F o (t 2 ). In this way, the consistency of the diachronic images in the feature space is ensured, thus improving the performance of the model in the change detection task.

[0048] Secondly, diachronic information processing is introduced. In the change detection task, it is not enough to rely solely on the pre-trained network to effectively capture the diachronic relationship. The spatially consistent inductive bias is introduced into the model to increase the capacity of the model, and the Van module of the pre-trained network BPNet is replaced with the SCAM module to realize the fusion of diachronic information. In addition, in order to deal with the channel mismatch problem caused by the expansion of the twin network, a fusion block is introduced after the decoder to further reduce the dimension of the features. These two items are change detection adaptation layers L c Specifically, the fusion block combines the features at time points t1 and t2 together to generate richer feature representations, and reduces the feature dimension to half after channel connection to match the channel number of the multi-scale fusion segmentation head.

[0049] Through the above two processing operations, the stability of the model structure can be guaranteed, and the diachronic information processing capability can be effectively introduced to achieve more efficient change detection. Although the model is expanded and some new units are added, the overall parameter quantity and computational complexity are still low, and the complete fine-tuning cost is within an acceptable range. Through the building extraction model expanded by CAN, the designed model has rich multi-scale features and low computational and parameter quantity. The pre-trained weights of the building extraction network can be imported into the change detection architecture, only a small amount of parameters need to be introduced, so that the change detection network can learn more robust building features and improve performance. In addition, the building extraction model expanded by CAN needs to process multi-scale information, so deep supervision is introduced to improve the prediction accuracy of the model. Deep supervision generates multi-scale mask results by calculating the loss function at each layer of the network, which helps to improve the performance of the model at different scales. Specifically, the loss function takes into account the loss at each level and can be represented as:

[0050]

[0051] where l iThe loss function of the i-th layer is represented as Bce, Dice respectively represents the binary cross-entropy loss function and the similarity coefficient loss function, and λ i The loss function coefficients of different layers are represented as y, The real label and the predicted label are represented as y

[0052] In the change detection task, in order to effectively capture the dual-time information, the model needs to introduce appropriate inductive bias. However, the traditional twin network structure can only process the features of each time point independently, which means that each encoder-decoder branch can only focus on the data of a single time point, ignoring the change correlation between the dual time. Based on this, the spatial consistency attention module (SCAM module) is introduced, which generates a feature map containing change information by fusing the features in the encoder branch, so as to better capture the details of the dual-time change. In addition, the SCAM module also combines the advantages of self-attention mechanism and large convolution kernel, and adopts a large kernel convolution decomposition technology to generate more accurate attention feature maps.

[0053] As shown in Figure 3 , the SCAM module includes an ADLKA module and a TFU module: the ADLKA module can capture more extensive context information without destroying the integrity of the dual-time features, and fuse the change information; the TFU module aims to preserve the timing information in the dual-time features.

[0054] Firstly, the ADLKA module outputs an attention map, the expression is as follows:

[0055]

[0056] In the formula, represents the mixed dual-time information features after the TFU module; represents the attention map output by the ADLKA module; respectively represent a deep dilated convolution and a deep convolution, which are used to extract context features, in the embodiment, k1=5, k2=5, d=3 are set to balance performance and computational complexity; Conv 1×1 represents a point-wise convolution, which is used to fuse local and context features to generate an attention map.

[0057] Furthermore, since it is observed that the changed regions in the images of the two time points have spatial consistency, it is assumed that the change positions in the dual-time images have spatial consistency, and the inductive bias for change detection is integrated into the model. At this time, the same attention map can act on the features of the two time points at the same time, the expression is as follows:

[0058]

[0059] In the formula, input features representing two time states, element-level product.

[0060] Then, the features are extracted by batch normalization (BN), 1x1 convolution, GELU activation function, and FFN module in sequence, and the model training and inference process is completed.

[0061] In addition, since the dual-time-state feature fuses the time sequence information, directly using BN may cause confusion of the time sequence information and cause large fluctuations in the training process. Therefore, the TFU module uses LN to normalize each sample independently, which better preserves the time sequence information and reduces fluctuations in training. The specific expression is as follows:

[0062]

[0063] X s =Conv 1×1 (X m );

[0064] T l =GELU(LN(X s ));

[0065] In the formula, f represents the feature after channel connection, f represents the feature after dimension reduction by 1x1 convolution, f represents the time-fused dual-time-state feature, LN represents layer normalization, GELU represents the activation function, and concat represents channel connection.

[0066] As the model size increases, its deployment difficulty also increases. Based on this, the feature enhancement module (LFEM) is adopted, which can maintain the calculation and parameter efficiency while improving the feature extraction capability of the model. LFEM is similar to the mobile back-bottleneck convolution (MBConv) module, which includes a 1x1 convolution for dimension expansion, followed by a 3x3 deep convolution for cross-channel interaction, and an SE module is introduced for feature map weighting processing; finally, the feature map is mapped to the target output channel through a 1x1 convolution. The specific expression is as follows:

[0067]

[0068]

[0069] In the formula, X represents the input feature, X represents the feature after expansion by the first 1x1 convolution, The output features are represented by n=1, 2 represent the first and second time states, and ⊕ represents element-wise addition. Each convolutional layer is followed by batch normalization and the SiLU activation function. When the input channels equal the output channels, i.e., C... in =C out By using residual connections to shorten the gradient backpropagation path and using Droppath to randomly skip network paths, the network depth is made random, thereby reducing network complexity and preventing overfitting.

[0070] In addition, to achieve multi-scale prediction output and solve the problem that backpropagation is difficult to reach shallow layers, a multi-scale fusion segmentation head (MSFSH) is introduced. 3×3 convolution is used in each decoder layer of the network to generate prediction maps of the corresponding scale, and bilinear interpolation is used to adjust the prediction maps to the original image size. Finally, 1×1 convolution is used to fuse the multi-scale prediction maps to generate the final feature map output.

[0071] Verification Example 1

[0072] The building model extracted via CAN expansion is denoted as the Uchanger model. The Uchanger design features rich multi-scale characteristics and low computational and parameter requirements. In this embodiment, two instances of Uchanger are provided using different filter count configurations: the standard version Uchanger-base (2.782M) and the relatively smaller version Uchanger-small (0.593M). Detailed configuration files are shown in Table 1.

[0073] Table 1. Detailed configuration files for two instances of the Uchanger model.

[0074]

[0075] This verification example compares and analyzes the change detection method of this embodiment with existing change detection methods such as FC-EF, FC-Siam-Conc, FC-Siam-Diff, DTCDSCN, ChangeFormer, IFN, BiT, ChangerEx, SGSLN / 512, CGNet, C2Fnet, and SRCNet, as shown in Table 2. It can be seen that the method of this embodiment achieves high F1 scores for both Uchanger_small and Uchanger_base, at 92.45% and 92.87% respectively. This indicates that the method of this embodiment has advantages in accuracy and stability in identifying changing regions, and also demonstrates strong performance when handling complex scenes.

[0076] Table 2 Results of different change detection methods on the LEVIR-CD dataset

[0077] Method Precision Recall F1 Score FC-EF 86.91 80.17 83.40 FC-Siam-Conc 91.99 76.77 83.69 FC-Siam-Diff 89.53 83.31 86.31 DTCDSCN 88.53 86.83 87.67 ChangeFormer 92.59 89.68 91.11 IFN 94.02 82.93 88.13 BiT 91.97 88.62 90.26 ChangerEx 92.97 90.61 91.77 SGSLN / 512 93.07 91.61 92.33 CGNet 93.15 90.90 92.01 C2FNet 93.69 90.04 91.83 SRCNet 92.63 91.72 92.24 Uchanger_small 92.93 91.96 92.45 Uchanger_base 93.18 92.61 92.87

[0078] In another verification example, the method of the embodiment has good performance on LEVIR-CD+, S2looking, and CDD datasets. In particular, for the case of large style difference between the dual-time images in the rural area in the S2Looking dataset, the method of the embodiment improves the change detection induction bias, so that the change is captured more accurately and efficiently, and the method can better adapt to the complex and changeable scenes in the dataset, thereby achieving more significant performance improvement on the S2Looking dataset.

[0079] In another verification example, the method of the embodiment is compared and analyzed with existing models in terms of the number of parameters, computational complexity, and accuracy. Compared with the second-ranked model, the method of the embodiment realizes an F1 score of 92.45 using only 0.607M parameters (a decrease of 90%) and a computational complexity of 6.053 MACs (a decrease of 47%), and thus realizes performance superior to other change detection models with lower theoretical parameter amount and computational complexity.

[0080] Ablation experiment

[0081] In the embodiment, the MSFSH module aims to maximize the use of feature information of different scales to realize more accurate building change detection, which not only better captures key subtle changes, but also effectively enhances the learning of shallow features by shortening the path of gradient back propagation, reduces the influence of gradient vanishing or explosion problem, and thus improves the overall performance of the model. The SCAM module is a key component for enhancing the receptive field of the model and the interaction of dual-time information. The core purpose of the CAN strategy training model is to introduce prior knowledge of building extraction, thereby improving the performance of the model in the change detection task.

[0082] Table 3 Ablation experiment results of each module on the LEVIR-CD dataset

[0083]

[0084] In order to verify the effectiveness of the proposed modules, an ablation experiment was conducted on the LEVIR-CD dataset, and it can be seen that compared with the baseline model, the model using MSFSH improves the F1 score by 0.38%. This result significantly indicates that the multi-scale fusion segmentation head plays an important role and is necessary in improving the accuracy and efficiency of the model, as shown in Table 3.

[0085] Compared with the second group of experiments, it is also found that only introducing a large convolution kernel (the third group of experiments) significantly improves the recall rate (increases 1.41%) but causes the accuracy rate to decrease at the same time (decreases 1.37%). Compared with the third group of experiments, when the double temporal interaction function is added on the basis of the large convolution kernel (the fourth group of experiments), only 0.066M of the parameter amount is increased, but the MACs is reduced by 0.11G, and the recall rate and the accuracy rate are obviously improved, and finally the F1 score is increased by 0.57% compared with the baseline model. This shows that the model design not only needs to enhance the ability of a single branch, but also must pay attention to the interaction between different branches, so as to comprehensively improve the performance.

[0086] In addition, through the fourth and fifth groups of experiments, it can be seen that after introducing the prior knowledge of building extraction, the model performance is significantly improved after fine-tuning on the change detection dataset, and the recall rate and F1 score are increased by 0.64% and 0.33% respectively. These results fully prove the effectiveness of the building extraction prior knowledge in the change detection task. By introducing these prior knowledge, the model not only enhances the ability to capture target features, but also improves the accuracy and stability in complex change detection scenarios, effectively improving the model performance.

[0087] On the other hand, Grad-CAM is also used to analyze the decoder layer output of the network stacked with LKA module and SCAM module respectively, as shown in Figure 4 and Figure 5 It can be seen that the LKA module pays more attention to buildings and their surroundings more evenly, but cannot distinguish between changed and unchanged buildings. The SCAM module not only accurately marks the area to be built, but also pays special attention to the changed buildings while ignoring the unchanged parts. These results show that the SCAM module has a significant advantage in capturing change features, thereby greatly improving the overall performance of the model. In addition, in the contribution analysis of the SCAM module to the receptive field expansion, it is found that: (1) the decoder layer tends to focus on local features; (2) as the resolution of the decoder feature map decreases, the effective receptive field (ERF) becomes more extensive; (3) the SCAM module helps to expand the ERF of the deep decoder. This means that the SCAM module can enhance the modeling ability of long-distance dependencies through large convolution kernels, thereby improving the performance of the model.

[0088] In another verification example, the SCAM module is verified for target change detection in different scenarios using the LEVIR-CD dataset. In processing large-scale targets, the SCAM model performs well in detecting building change areas, especially in detail and edge processing, as shown in Figure 6 For tree-shaded scenes, SCAM shows higher robustness and can effectively prevent false detection of changed buildings, as shown in Figure 7are shown. In addition, SCAM shows good advantage in locating the changing buildings in complex scenes, such as Figure 8 are shown.

[0089] Therefore, the application adopts the above change detection method based on semantic prior and fusion space positioning, which can reduce the complexity of the network, prevent overfitting, realize effective fusion of double temporal images, better capture time series information, and thus improve the accuracy of target change detection.

[0090] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, but not to limit them, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand: its still can be modified or equivalent to replace the technical solutions of the present application, and these modifications or equivalent replacements also cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.

Claims

1. A change detection method based on semantic priors and fusion space localization, characterized in that, Comprise: First, pre-train the BPNet network on supervised remote sensing extraction tasks, making full use of prior knowledge in remote sensing extraction tasks; Then, fine-tune the pre-trained BPNet network, expand the encoder and decoder structure with a shared weight twin network structure, and introduce a spatial consistency attention module to replace the Van module of the pre-trained network BPNet. By combining the self-attention mechanism and large kernel decomposition technology, the feature information of the double-time image is extracted and fused to obtain the change information of the remote sensing image target; In the fine-tuned network, introduce deep supervision, calculate the loss function at each layer of the network to generate a multi-scale mask, and the expression of the loss function is: ; ; In the formula, denotes the first loss function of the layer, , denote binary cross-entropy loss function and similarity coefficient loss function, respectively, denotes the loss function coefficient of different layers, , are real labels and predicted labels, respectively; The spatial consistency attention module includes an ADLKA module and a TFU module: the ADLKA module is used to capture more extensive context information and fuse change information, and the expression of the ADLKA module output is: ; In the formula, bilateral temporal information features representing temporal fusion, attention maps output by the ADLKA module, a depth convolution with a convolution kernel a depth dilated convolution, a depth convolution with a convolution kernel a depth convolution with a convolution kernel a point-wise convolution; The TFU module is used to retain the timing information in the double-time feature, and the TFU module uses layer normalization to process each sample independently, and the expression is as follows: ; ; wherein, input features of the first and second time states, respectively, denotes a temporal fusion, convolutional reduced features, denotes a dual-time feature with temporal fusion, denotes a channel connection, denotes a layer normalization, denotes an activation function.

2. The method for change detection based on semantic priors and fused spatial localization of claim 1, wherein, The encoder and decoder of the BPNet network use feature enhancement modules for feature extraction, and the expression of the feature enhancement module is: ; ; ; In the formula, Input features of the feature enhancement module, The first Convolutionally expanded features, Output features of the feature enhancement module, Temporal of the feature map, Element-level addition, Self-attention mechanism, Deep separable convolution with Convolution with Convolution with Convolution with 3. The method for change detection based on semantic priors and fused spatial localization of claim 1, wherein, The multi-scale fusion segmentation head is introduced after the encoder to generate the prediction map of the corresponding scale by convolving the feature map extracted by each layer, and the prediction map obtained by each layer is adjusted to the original image size by bilinear interpolation, and a multi-scale prediction map is generated by fusion.

Citation Information

Patent Citations

  • Method for extracting building change area in double-time-phase remote sensing image based on twinborn mixed attention mechanism and multi-scale feature fusion

    CN118212532A