Remote Sensing Image Change Detection Method, Device and Electronic Device
Through the multi-scale cross-attention network, the interference caused by environmental and sensor differences in remote sensing images is solved, the detection ability of small change areas is enhanced, and the change detection with higher accuracy and integrity is achieved.
Patent Information
- Application Number
- CN202510386799.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-31
AI Technical Summary
When processing high-resolution remote sensing images, it is difficult for the prior art to accurately detect all changing areas, especially because the image differences caused by environmental and sensor differences increase detection difficulty, and the characteristics of the slight changing areas are not obvious enough, resulting in the omission of details.
Multi-scale cross-attention network (MSCA-Net) is adopted to integrate multi-scale information and cross-attention mechanisms to extract low-level features of dual-time phase remote sensing images, perform differential enhancement processing and context information linking, enhance the recognition ability of changing areas and solve the domain gap problem.
It improves the accuracy and completeness of remote sensing image change detection, and can perceive tiny changes more sensitively, ensuring that changes between images in different periods are comprehensive and accurate.
Smart Images

Figure CN119904760B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image change detection, and specifically relates to a method, device, and electronic device for remote sensing image change detection. Background Art
[0002] Remote sensing image change detection is a key technology for quickly identifying and evaluating surface feature changes, and is widely used in multiple fields such as resource exploration, environmental monitoring, urban planning, and management. By comparing remote sensing images at different time points, this technology can reveal the change patterns of the surface and provide important basis for decision-making in various industries.
[0003] However, when dealing with high-resolution remote sensing images, accurately detecting all change regions still faces challenges. On the one hand, due to differences in acquisition conditions (such as weather conditions, lighting angles) and sensor characteristics, even unchanged ground objects may appear significantly different in images taken at different time points. For example, the texture performance of the same building may show obvious differences under different sensors, which increases the difficulty of detecting change regions. On the other hand, for regions with small area and subtle changes, their features are often not obvious enough to be captured by existing change detection algorithms, resulting in the omission of change details in the final detection results.
[0004] To address the above problems, this application proposes a multi-scale cross-attention network (hereinafter referred to as MSCA-Net). MSCA-Net aims to improve the detection accuracy and integrity of change regions by integrating multi-scale information and using the cross-attention mechanism. This network can not only effectively reduce the interference caused by environmental and sensor differences, but also more sensitively perceive the existence of subtle changes, thus ensuring that the changes between images in different periods are comprehensively and accurately identified. This improvement is expected to significantly improve the effect of change detection and provide more reliable data support for users. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide a method, device, and electronic device for remote sensing image change detection, which can solve the problems that the differences in environment and sensors cause interference to image difference detection and the detail changes are difficult to be detected and are often ignored.
[0006] To solve the above technical problems, this application is implemented as follows:
[0007] In a first aspect, the embodiments of this application provide a method for remote sensing image change detection, which includes:
[0008] Input the dual-temporal remote sensing images into a multi-scale cross-attention network; wherein, the multi-scale cross-attention network is composed of multi-scale feature encoding, differential feature enhancement algorithm, and context information linking;
[0009] Extract low-level features of dual-temporal remote sensing images based on different feature scales, fuse the low-level features, and obtain low-level fused features; among them, the low-level fused features correspond to different time phases respectively.
[0010] Based on the cross-attention mechanism, perform differential enhancement processing on the low-level fused features to obtain an enhanced fused feature map; among them, the differential enhancement processing is used to reduce the domain gap of the low-level fused features and enhance the key change information between the low-level fused features.
[0011] Based on the context information linking algorithm, perform a reconstruction operation on the enhanced fused feature map to obtain a reconstructed feature map; among them, the reconstruction operation is used to enhance the key change information in the enhanced fused feature map based on spatial weights.
[0012] Perform a subtraction operation on the reconstructed feature map according to different time phases to obtain the final recognition image.
[0013] Compared with the prior art, the above technical solution provided by this application has at least the following beneficial effects:
[0014] In this application, first, input the dual-temporal remote sensing images into a multi-scale cross-attention network; second, extract low-level features of the dual-temporal remote sensing images based on different feature scales, fuse the low-level features, and obtain low-level fused features to enhance the feature characterization ability of the multi-scale cross-attention network for small change regions; then, based on the cross-attention mechanism, perform differential enhancement processing on the low-level fused features to obtain an enhanced fused feature map, so as to guide the network to focus on the change region through differential features, enhance the network's recognition ability for the change region, and at the same time solve the domain gap problem through information sharing of the unchanged region; then, based on the context information linking algorithm, perform a reconstruction operation on the enhanced fused feature map to obtain a reconstructed feature map, which can integrate the change information of small regions into the final feature expression and strengthen the network's recognition ability for small change regions; finally, perform a subtraction operation on the reconstructed feature map according to different time phases to obtain the final recognition image. By emphasizing the features of small change regions in remote sensing images and using the unchanged region to eliminate the domain gap at the same time, this application jointly expresses the final image detection result and improves the integrity and accuracy of identifying the change regions of remote sensing images in different periods.
[0015] Preferably, the specific steps of extracting low-level features of dual-temporal remote sensing images based on different feature scales, fusing the low-level features, and obtaining low-level fused features include:
[0016] Perform upsampling and downsampling operations on the dual-temporal remote sensing images to obtain multi-scale images.
[0017] Apply low-level feature extraction to the multi-scale images to obtain low-level features.
[0018] Fuse low-level features based on concatenation operation to obtain low-level fused features.
[0019] In this solution, by learning image features across different scales and integrating multi-scale feature information using concatenation operation, the feature expression ability of the multi-scale cross-attention network can be enhanced. In addition, by adding larger-scale feature information, it is easy for the network to focus on tiny change regions, thereby improving the accuracy and integrity of the change regions.
[0020] Preferably, the operation formula for fusing low-level features based on concatenation operation to obtain low-level fused features is:
[0021]
[0022] where represents the low-level fused features; represents the dual-temporal remote sensing image, and its subscript represents the scaling ratio of the upsampling and downsampling operations on the dual-temporal remote sensing image; i represents the classification identifier of the images or features in different periods; and represent the upsampling and downsampling operations respectively; represents the concatenation operation; represents the structure of the first 7 layers adopted by the VGG network.
[0023] Preferably, based on the cross-attention mechanism, the specific steps for performing differential enhancement processing on the low-level fused features to obtain the enhanced fused feature map include:
[0024] Adaptively fuse the low-level fused features through convolution to obtain convolution feature information;
[0025] Based on the cross-attention mechanism, calculate the cross-attention output according to the convolution feature information;
[0026] Add the cross-attention output and the convolution feature information according to the phase to obtain the fused feature combination; among them, the fused feature combination corresponds to different phases respectively;
[0027] Concatenate the fused feature combinations to obtain the concatenated fused features;
[0028] Process the concatenated fused features using a multi-layer perceptron to obtain the deep fused features;
[0029] Perform a deformation operation on the deep fused features to obtain the enhanced fused feature map.
[0030] In this solution, through the cross-attention mechanism, the feature differences in different phases are interacted with the initial features, enhancing the sensitivity of the multi-scale cross-attention network to change information and improving the ability to identify change regions;
[0031] Preferably, based on the cross-attention mechanism, according to the convolutional feature information, the specific steps for calculating the cross-attention output include:
[0032] Calculate the difference feature based on the convolutional feature information;
[0033] The operation formula for calculating the difference feature is:
[0034]
[0035] where D is the difference feature, and respectively represent the convolutional feature information at different times;
[0036] Construct the query feature according to the difference feature; the formula for constructing the query feature is:
[0037]
[0038] where Q represents the query feature, W q represents the query transformation matrix automatically learned by the multi-scale cross-attention network;
[0039] Construct the key feature and the value feature based on the convolutional feature information, and the formula for constructing the key feature and the value feature is:
[0040]
[0041] where K represents the key feature, V represents the value feature; W k and W v respectively represent the key transformation matrix and the value transformation matrix automatically learned by the multi-scale cross-attention network; i represents the classification identifier of the image or feature at different times;
[0042] Adopt the cross-attention mechanism to calculate the cross-attention output according to the query feature, the key feature and the value feature;
[0043] The operation formula for calculating the cross-attention output is:
[0044]
[0045] where d is the channel dimension.
[0046] In this solution, through cross-attention operations, the differential features are incorporated into the feature maps of each phase, enabling full interaction and fusion of features from different sources. This allows the multi-scale cross-attention network to accurately identify the changed regions while ignoring the unchanged regions. At the same time, for the unchanged regions in heterogeneous images, they are forced to be unified through differential operations and supervision information, thereby solving the domain gap problem in the feature space.
[0047] Preferably, based on the context information linking algorithm, the specific steps for reconstructing the enhanced fusion feature map to obtain the reconstructed feature map include:
[0048] Assign corresponding weights to the low-level features;
[0049] Based on the weights, splice the low-level features to obtain the spliced feature map;
[0050] Apply a convolution operation to the spliced feature map to obtain the spatial weight feature map;
[0051] Multiply the spatial weight feature map by the enhanced fusion feature map to obtain the reconstructed feature map.
[0052] In this solution, the ability of the network to identify changed regions and distinguish unchanged regions is enhanced respectively. In addition, the spatial weight feature map fully retains the change information, providing rich semantic information for obtaining a more comprehensive feature expression of the changed regions. At the same time, combined with the enhanced fusion feature map, it provides spatial distribution feature information for the unchanged regions.
[0053] Preferably, based on the context information linking algorithm, the operation formula for reconstructing the enhanced fusion feature map to obtain the reconstructed feature map is:
[0054]
[0055] Among them, represents the reconstructed feature map; represents the enhanced fusion feature map; 、 and respectively represent low-level features of different scales; i represents the classification identifier of images or features at different times; a 、 b and c respectively represent the weight values corresponding to the low-level features; RELU represents the activation function, BN represents batch normalization, and Conv represents the convolution operation.
[0056] Preferably, the specific steps of the remote sensing image change detection method further include:
[0057] Construct a loss function based on the image parameters of the final recognized image; wherein, the image parameters include the height, width, true pixel value and predicted pixel value of the final recognized image.
[0058] Train a multi-scale cross-attention network based on the loss function to obtain the target multi-scale cross-attention network.
[0059] In this solution, constructing the loss function plays a core role in guiding the optimization direction of the multi-scale cross-attention network. It clearly defines the difference metric between the predicted detection image and the actual detection image, so as to minimize such differences through an optimization algorithm, adjust the model parameters and improve the performance, and obtain the target multi-scale cross-attention network.
[0060] In a second aspect, an embodiment of the present application provides a remote sensing image change detection device, including:
[0061] A network construction module, configured to input dual-temporal remote sensing images into a multi-scale cross-attention network; wherein, the multi-scale cross-attention network is composed of multi-scale feature encoding, differential feature enhancement algorithm and context information linking.
[0062] A multi-scale feature fusion module, configured to extract low-level features of dual-temporal remote sensing images based on different feature scales, fuse the low-level features, and obtain low-level fusion features; wherein, the low-level fusion features respectively correspond to different time phases.
[0063] A differential feature encoding module, configured to perform differential enhancement processing on the low-level fusion features based on the cross-attention mechanism to obtain an enhanced fusion feature map; wherein, the differential enhancement processing is used to reduce the domain gap of the low-level fusion features and enhance the key change information between the low-level fusion features.
[0064] A context information linking module, configured to perform a reconstruction operation on the enhanced fusion feature map based on the context information linking algorithm to obtain a reconstructed feature map; wherein, the reconstruction operation is used to enhance the key change information in the enhanced fusion feature map based on spatial weights.
[0065] An identified image output module, configured to perform a subtraction operation on the reconstructed feature map according to different time phases to obtain the final recognized image.
[0066] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0067] It can be understood that the beneficial effects of the technical solutions provided in the above second aspect and third aspect can refer to the relevant descriptions in the above first aspect, and will not be repeated here.
[0068] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0070] Figure 1 is a schematic flowchart of a remote sensing image change detection method shown in some embodiments of the present application;
[0071] Figure 2 is a block diagram of a remote sensing image change detection device shown in some embodiments of the present application;
[0072] Figure 3 is a block diagram of an electronic device shown in some embodiments of the present application;
[0073] The following specific embodiments will further illustrate the present application in conjunction with the above drawings. SPECIFIC EMBODIMENTS
[0074] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0075] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.
[0076] A remote sensing image change detection method provided by the embodiments of the present application will be described in detail below with reference to the drawings, through specific embodiments and their application scenarios.
[0077] Figure 1 is a schematic flowchart of a remote sensing image change detection method shown in the first embodiment of the present application. Please refer to Figure 1 , the method includes:
[0078] Step S101: Input the dual-temporal remote sensing image into the multi-scale cross-attention network; where the multi-scale cross-attention network is composed of multi-scale feature encoding, differential feature enhancement algorithm, and contextual information linkage;
[0079] Specifically, the multi-scale cross-attention network consists of a multi-scale feature encoding module (Multi-Scale Feature Enhancement, MSFE), a differential feature enhancement module (Differential Feature Enhancement and Alignment module, DFEA), and a contextual information linkage module (Contextual Information Linkage, CIL);
[0080] The MSFE module is used to construct images of different scales, extract features respectively, and enhance the network's feature description ability for tiny change regions through multi-scale feature fusion;
[0081] The DFEA module is used to guide the network to focus on the change region through differential features, enhance the network's recognition ability for the change region, and at the same time solve the domain gap problem through the information sharing of the unchanged region;
[0082] The CIL module is used to perform weighted integration and reconstruction on the original multi-scale information, incorporate the change information of tiny regions into the final feature expression, and strengthen the network's recognition ability for tiny change regions.
[0083] Step S102: Extract the low-level features of the dual-temporal remote sensing image based on different feature scales, fuse the low-level features, and obtain low-level fused features; where the low-level fused features correspond to different time phases respectively;
[0084] Specifically, perform upsampling and downsampling operations on the dual-temporal remote sensing image to obtain multi-scale images; apply low-level feature extraction to the multi-scale images to obtain low-level features; fuse the low-level features based on concatenation operation to obtain low-level fused features.
[0085] In a possible implementation manner, to ensure the accurate recognition of tiny change regions, this embodiment designs the MSFE module. First, given the dual-temporal remote sensing image , where H , W and 3 represent the height, width, and number of channels of the image respectively. Based on the upsampling and downsampling operations, multi-scale images , , , , , where the superscripti = 1 or 2 represents images in different periods, represents the image being reduced by half, represents the original image size, represents the image being doubled.
[0086] Furthermore, use VGG-7 the network to extract low-level features from multi-scale images to learn the local features of different scales of the images, obtain low-level features of multiple scales, and fuse the low-level features of three different scales through a concatenation operation to obtain low-level fused features. The operation formula is as follows:
[0087]
[0088] Among them, represents the low-level fused feature, and respectively represent the upsampling and downsampling operations, represents the concatenation operation, represents VGG the network adopts the structure of the first 7 layers; represents the dual-temporal remote sensing image, and its subscript represents the scaling ratio of the upsampling and downsampling operations on the dual-temporal remote sensing image, i ∈ {1, 2}, representing the classification identifier of the images or features in different periods.
[0089] The MSFE module can enhance the feature expression ability of the multi-scale cross-attention network by learning image features across different scales and integrating multi-scale feature information using the concatenation operation. In addition, by adding larger-scale feature information, the multi-scale cross-attention network is more likely to focus on small change regions, thereby improving the accuracy and integrity of the change regions.
[0090] Step S103: Based on the cross-attention mechanism, perform differential enhancement processing on the low-level fused feature to obtain an enhanced fused feature map; among them, the differential enhancement processing is used to reduce the domain gap of the low-level fused feature and enhance the key change information between the low-level fused features;
[0091] Specifically, adaptively fuse the low-level fused feature through convolution to obtain convolution feature information; based on the cross-attention mechanism, calculate the cross-attention output according to the convolution feature information; add the cross-attention output and the convolution feature information according to the phase to obtain a fused feature combination; among them, the fused feature combination corresponds to different phases respectively; concatenate the fused feature combinations to obtain a concatenated fused feature; use a multi-layer perceptron to process the concatenated fused feature to obtain a deep fused feature; perform a deformation operation on the deep fused feature to obtain an enhanced fused feature map.
[0092] Further, based on the cross-attention mechanism, the specific steps for calculating the cross-attention output according to the convolutional feature information include: obtaining the difference feature through operations on the convolutional feature information; constructing the query feature based on the difference feature; constructing the key feature and value feature based on the convolutional feature information; and using the cross-attention mechanism to calculate the cross-attention output according to the query feature, key feature, and value feature.
[0093] In a possible implementation, due to environmental factors and the influence of different sensors, there are significant differences in images at different time periods, that is, the domain gap problem, which leads to large differences in the features of the same target and causes misidentifications in the changed and unchanged regions. To solve this problem, the DFEA module is proposed.
[0094] First, respectively apply 3×3 convolutional operations to the concatenated low-level fusion features and ∈ to adaptively fuse feature information at different scales to obtain convolutional feature information and , enabling the multi-scale cross-attention network to automatically learn features that are more beneficial to the change detection task.
[0095] First, respectively use 3×3 convolutional operations on the concatenated and ∈ to adaptively fuse feature information at different scales and , allowing the multi-scale cross-attention network to automatically learn features that are more conducive to the change detection task.
[0096] Secondly, map the convolutional feature information of different time phases to the feature space of the same dimension, and obtain the difference feature D of different time phases through subtraction operation. The operation formula is as follows:
[0097]
[0098] where D is the difference feature, and respectively represent the convolutional feature information in different periods;
[0099] Next, in order to enable the multi-scale cross-attention network to focus on the feature information of the changed region, the cross-attention mechanism is used for feature interaction and fusion. The calculation formula for the query feature is:
[0100]
[0101] And in and , the key features of the corresponding time phases are respectively constructed Sum feature , and its calculation formula is:
[0102]
[0103] Among them, , Sum , are the query transformation matrix, key transformation matrix, and value transformation matrix automatically learned by the multi-scale cross-attention network respectively, C is the product of the height and width of the input feature image, d is the channel dimension of the query feature, key feature, and value feature, i represents the classification identifier of the images or features in different periods. By adopting the cross-attention operation and using the difference feature as the key guiding information, the multi-scale cross-attention network can automatically identify and focus on the feature information of the changed area. The detailed calculation formula is as follows:
[0104]
[0105] Among them, d is the channel dimension.
[0106] It should be noted that the beneficial effects of using the cross-attention mechanism for feature processing are reflected in the following aspects:
[0107] First, through the cross-attention mechanism, the feature differences in different phases are interacted with the initial features, enhancing the sensitivity of the multi-scale cross-attention network to the changed information and improving the ability to identify the changed area;
[0108] Second, through the cross-attention operation, the difference features are incorporated into the feature maps of each phase, realizing the full interaction and fusion of features from different sources. This enables the multi-scale cross-attention network to accurately identify the changed area while ignoring the unchanged area. At the same time, for the unchanged area in the heterogeneous images, through the differential operation and supervision information, it is forced to be unified, thus solving the domain gap problem in the feature space;
[0109] Third, the subtraction operation of the feature maps in different phases will be accompanied by the loss of effective information. To reduce the impact of information loss on change detection, the cross-attention output is added to the convolutional feature information and respectively to obtain the fusion feature combination. At the same time, this method also retains the information of the unchanged area, enabling the multi-scale cross-attention network to establish the correlation between the feature expressions of data in different phases through the supervision signal of the unchanged area, thereby realizing feature alignment and further reducing the impact of the domain gap problem.
[0110] Finally, and Concatenate them to obtain concatenated fusion features, integrate the information of the changing features and the unchanging features, and perform deep feature fusion on the concatenated fusion features using a multi-layer perceptron (MLP). Subsequently, apply a deformation operation to obtain the output strong fusion feature map respectively. and .
[0111] Step S104: Based on the context information linking algorithm, perform a reconstruction operation on the enhanced fusion feature map to obtain a reconstructed feature map; wherein, the reconstruction operation is used to enhance the key changing information in the enhanced fusion feature map based on spatial weights.
[0112] Specifically, assign corresponding weights to the low-level features; concatenate the low-level features based on the weights to obtain a concatenated feature map; apply a convolution operation to the concatenated feature map to obtain a spatial weight feature map; multiply the spatial weight feature map by the enhanced fusion feature map to obtain a reconstructed feature map.
[0113] In a possible implementation manner, the embodiment of the present application adopts a multi-layer DEFA module to guide the network to pay more attention to the feature representation of the changing region. In order to improve the network's attention to the unchanging information and the details of the small changing region, thereby improving the integrity of the change detection result, the embodiment of the present application also proposes a CIL module. This module emphasizes the small-region change information, enhancing the integrity of the detection result while ensuring the detection accuracy.
[0114] First, concatenate the low-level features based on different weights to obtain a concatenated feature map, where the large-scale feature map is assigned a higher weight. Subsequently, perform reconstruction on the concatenated feature map through a 3×3 convolution operation to obtain a spatial weight feature map that pays more attention to the small target region . Multiply the spatial weight feature map by the enhanced fusion feature map output by the multi-level DFEA module to obtain a reconstructed feature map , thereby associating the reconstructed spatial weight feature with the fused high-level feature information.
[0115] This process enhances the network's ability to identify the changing region and distinguish the unchanging region respectively. In addition, the spatial weight feature map fully retains the changing information, providing rich semantic information for obtaining a more comprehensive feature expression of the changing region. At the same time, combined with the enhanced fusion feature map output by the DFEA module, it provides spatial distribution feature information for the unchanging region. The formula is as follows:
[0116]
[0117] wherein, represents the reconstructed feature map, represents the enhanced fusion feature map,a , b and c respectively represent the weight values corresponding to the low-level features. , and respectively represent the low-level features, RELU represents the activation function, BN represents batch normalization, and Conv represents the convolution operation.
[0118] In this solution, the multi-scale cross-attention network can more accurately identify the changed areas while retaining the important information of the unchanged areas, thereby achieving more accurate and complete change detection.
[0119] Step S105: Subtract the reconstructed feature maps according to different periods to obtain the final recognition image.
[0120] The specific steps of the remote sensing image change detection method further include:
[0121] Construct a loss function according to the image parameters of the final recognition image; the image parameters include the height, width, true pixel value and predicted pixel value of the final recognition image; train the multi-scale cross-attention network based on the loss function to obtain the target multi-scale cross-attention network.
[0122] In a possible implementation manner, this embodiment uses the cross-entropy loss function to minimize the parameters of the multi-scale cross-attention network. Change detection can be regarded as a binary classification problem. Therefore, the cross-entropy function is used as the loss function to quantify the difference between the predicted change of the model and the actual situation. The specific operation formula of the loss function is as follows:
[0123]
[0124] Among them, represent the height and width of the change detection result image, represents the position at the true pixel value of the pixel, and represents the predicted value of the pixel at this position .
[0125] Next, the accuracy evaluation of the multi-scale cross-attention network in the embodiments of the present application will be introduced.
[0126] The embodiments of the present application mainly use three evaluation indicators: F1 score, intersection over union (IoU), and overall accuracy (OA) for evaluation. The definitions of each evaluation indicator are as follows:
[0127] F1-Score: The F1-score is the harmonic mean of precision and recall, which is used to measure the performance of a model in a classification task. It comprehensively considers the prediction accuracy and completeness of the model for positive classes. The higher the value of the F1-score, the better the performance of the model.
[0128]
[0129]
[0130]
[0131] Among them, TP, TN, FP, and FN represent the numbers of true positives (1 Positives), true negatives (1 Negatives), false positives (0 Positives), and false negatives (0 Negatives), respectively.
[0132] TP (1 Positives, true positives): It represents the number of positive samples correctly predicted by the model.
[0133] TN (1 Negatives, true negatives): It represents the number of negative samples correctly predicted by the model.
[0134] FP (0 Positives, false positives): It represents the number of negative samples wrongly predicted as positive samples by the model.
[0135] FN (0 Negatives, false negatives): It represents the number of positive samples wrongly predicted as negative samples by the model.
[0136] Intersection over Union (IoU): The Intersection over Union is the ratio of the intersection area to the union area between the predicted region of the model and the ground truth region. It is an important metric for measuring the prediction accuracy of the model in change detection tasks. The higher the value of IoU, the closer the predicted region of the model is to the ground truth region.
[0137]
[0138] Overall Accuracy (OA): The Overall Accuracy is the ratio of the number of pixels correctly predicted by the model to the total number of pixels. It measures the overall performance of the model on the entire dataset. The higher the value of OA, the higher the accuracy of the model.
[0139]
[0140] These three evaluation metrics together constitute the main criteria for evaluating the model performance in this study, which helps to comprehensively and objectively measure the performance of the model in change detection tasks.
[0141] The main classification methods for comparison are: FC-EF (2018), FC-Siam-Di (2018), FC-Siam-Co (2018), STANet (2020), BIT (2022), SNUNet (2022), HANet (2023), VcT (2023), and C2FNet (2024). It should be noted that:
[0142] The FC-EF, FC-Siam-Di, and FC-Siam-Co methods are a hybrid convolutional neural network for hyperspectral image classification, integrating the advantages of three-dimensional convolution and two-dimensional convolution. The three-dimensional convolution part is responsible for extracting joint spatial-spectral feature representations from consecutive spectral bands, while the two-dimensional convolution part further learns more abstract high-level spatial representations based on the three-dimensional convolution. The STANet method uses a deep residual three-dimensional convolutional neural network to extract spectral-spatial features of the image, and at the same time learns a metric space to make samples of the same class more similar in the metric space and enhance the difference between samples of different classes. Finally, a nearest neighbor classifier is used to classify test samples in the learned metric space. The BIT method is a deep classification model based on the relational network and meta-learning idea. The feature learning module designed by the network extracts deep features from hyperspectral image samples, while the relational learning module conducts relational learning by comparing the similarities between different samples, that is, the relational score between samples of the same class is high, and the relational score between samples of different classes is low. The SNUNet method is a deep cross-domain few-shot learning method aiming to solve the domain shift problem in hyperspectral image classification. It uses the conditional adversarial domain adaptation strategy to overcome the domain shift problem and achieve domain distribution alignment; in addition, FSL is performed simultaneously in the source class and the target class, which can not only discover transferable knowledge in the source class but also learn a discriminative encoding model for the target class. The HANet method learns transferable knowledge from the mini-ImageNet dataset and then fine-tunes it on hyperspectral data to extract more discriminative spectral-spatial features and domain knowledge, improving the accuracy of HSI classification. The VcT method designs a self-supervised learning module with geometric transformation based on HFSL. By setting rotation labels as supervision, it extracts low-level features that can better represent different directions, and then conducts self-supervised learning and few-shot learning on the base class to learn transferable spatial meta-knowledge. The C2FNe method is a hyperspectral image classification method based on feature decoupling. Starting from the perspective of decoupled representation learning, it aims to solve the domain shift problem in hyperspectral image classification and reduce the bias of the source data on the representation, enabling the model to implicitly focus on the inherent knowledge of the target domain.
[0143] Table 1 Comparison results on three standard datasets, where the bold values are the highest scores
[0144]
[0145] For the remote sensing image change detection method provided by the above embodiments, first, input the dual-temporal remote sensing images into the multi-scale cross-attention network; secondly, extract the low-level features of the dual-temporal remote sensing images based on different feature scales, fuse the low-level features, and obtain the low-level fused features to enhance the ability of the multi-scale cross-attention network to depict the features of the tiny change regions; then, based on the cross-attention mechanism, perform differential enhancement processing on the low-level fused features to obtain the enhanced fused feature map, so as to guide the network to focus on the change regions through the differential features, enhance the network's recognition ability for the change regions, and at the same time solve the domain gap problem through the information sharing of the unchanged regions; next, based on the context information linking algorithm, perform a reconstruction operation on the enhanced fused feature map to obtain the reconstructed feature map, which can integrate the change information of the tiny regions into the final feature expression and strengthen the network's recognition ability for the tiny change regions; finally, perform a subtraction operation on the reconstructed feature map according to different periods to obtain the final recognition image. By emphasizing the features of the tiny change regions in the remote sensing images and using the unchanged regions to jointly express the final image detection results, the present application improves the integrity and accuracy of identifying the change regions in remote sensing images at different times.
[0146] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a remote sensing image change detection device shown according to the second embodiment of the present application. The remote sensing image change detection device 200 includes:
[0147] Network construction module 201: used to input the dual-temporal remote sensing images into the multi-scale cross-attention network; wherein, the multi-scale cross-attention network is composed of multi-scale feature encoding, differential feature enhancement algorithm, and context information linking;
[0148] Multi-scale feature fusion module 202: extract the low-level features of the dual-temporal remote sensing images based on different feature scales, fuse the low-level features, and obtain the low-level fused features; wherein, the low-level fused features correspond to different time phases respectively; specifically include:
[0149] Perform upsampling and downsampling operations on the dual-temporal remote sensing images to obtain multi-scale images; apply low-level feature extraction to the multi-scale images to obtain low-level features; fuse the low-level features based on concatenation operations to obtain low-level fused features.
[0150] Differential feature encoding module 203: based on the cross-attention mechanism, perform differential enhancement processing on the low-level fused features to obtain the enhanced fused feature map; wherein, the differential enhancement processing is used to reduce the domain gap of the low-level fused features and enhance the key change information between the low-level fused features; specifically include:
[0151] Adaptive fusion of low-level fusion features through convolution to obtain convolutional feature information; based on the cross-attention mechanism, calculate the cross-attention output according to the convolutional feature information; add the cross-attention output and the convolutional feature information according to the time phase to obtain a fusion feature combination; where the fusion feature combinations correspond to different time phases respectively; concatenate the fusion feature combinations to obtain a concatenated fusion feature; use a multi-layer perceptron to process the concatenated fusion feature to obtain a deep fusion feature; perform a deformation operation on the deep fusion feature to obtain an enhanced fusion feature map.
[0152] Further, based on the cross-attention mechanism, calculating the cross-attention output according to the convolutional feature information specifically includes: obtaining a difference feature by operating on the convolutional feature information; constructing a query feature according to the difference feature; constructing a key feature and a value feature based on the convolutional feature information; using the cross-attention mechanism to calculate the cross-attention output according to the query feature, the key feature and the value feature.
[0153] Context information linking module 204: Based on the context information linking algorithm, perform a reconstruction operation on the enhanced fusion feature map to obtain a reconstructed feature map; where the reconstruction operation is used to enhance the key change information in the enhanced fusion feature map based on spatial weights; specifically includes:
[0154] Assign corresponding weights to the low-level features; concatenate the low-level features based on the weights to obtain a concatenated feature map; apply a convolutional operation to the concatenated feature map to obtain a spatial weight feature map; multiply the spatial weight feature map by the enhanced fusion feature map to obtain a reconstructed feature map.
[0155] Image recognition output module 205: Used to perform a subtraction operation on the reconstructed feature map according to different periods to obtain a final recognized image.
[0156] Construct a loss function according to the image parameters of the final recognized image; the image parameters include the height, width, true pixel value and predicted pixel value of the final recognized image; train the multi-scale cross-attention network based on the loss function to obtain a target multi-scale cross-attention network.
[0157] It should be noted that for a remote sensing image change detection method provided in an embodiment of the present application, the execution subject can be a remote sensing image change detection device, or a control module in the remote sensing image change detection device for executing the loading of a remote sensing image change detection method. In an embodiment of the present application, taking a remote sensing image change detection device executing the loading of a remote sensing image change detection method as an example, a remote sensing image change detection method provided in an embodiment of the present application is described.
[0158] A remote sensing image change detection device in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0159] A remote sensing image change detection device in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0160] A remote sensing image change detection device provided in an embodiment of the present application can implement Figures 1 to 3 each process implemented by a remote sensing image change detection device in the method embodiment. To avoid repetition, it will not be elaborated here.
[0161] Optionally, please refer to Figure 3 , an embodiment of the present application further provides an electronic device 300, including a processor 301, a memory 302, and a computer program 303 stored on the memory 302 and executable on the processor 301. When the computer program 303 is executed by the processor 301, it implements each process of the above-mentioned remote sensing image change detection method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0162] An embodiment of the present application further provides a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, it implements each process of the above-mentioned remote sensing image change detection method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0163] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc.
[0164] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiment of a remote sensing image change detection method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0165] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.
[0166] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0168] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A remote sensing image change detection method, characterized in that, Including: Input the dual-temporal remote sensing image into a multi-scale cross-attention network; wherein, the multi-scale cross-attention network is composed of multi-scale feature encoding, differential feature enhancement algorithm, and context information linking; Extract low-level features of the dual-temporal remote sensing image based on different feature scales, fuse the low-level features, and obtain low-level fused features; wherein, the low-level fused features correspond to different time phases respectively; Convolutionally and adaptively fuse the low-level fused features to obtain convolutional feature information; Based on the cross-attention mechanism, calculate the cross-attention output according to the convolutional feature information; Add the cross-attention output and the convolutional feature information according to the time phase to obtain a fused feature combination; wherein, the fused feature combination corresponds to different time phases respectively; Concatenate the fused feature combination to obtain a concatenated fused feature; Process the concatenated fused feature using a multi-layer perceptron to obtain a deep fused feature; Perform a deformation operation on the deep fused feature to obtain an enhanced fused feature map; Based on the context information linking algorithm, perform a reconstruction operation on the enhanced fused feature map to obtain a reconstructed feature map; wherein, the reconstruction operation is used to enhance the key change information in the enhanced fused feature map based on spatial weights; Perform a subtraction operation on the reconstructed feature map according to different time phases to obtain a final recognition image.
2. The remote sensing image change detection method according to claim 1, characterized in that The specific steps of extracting low-level features of the dual-temporal remote sensing image based on different feature scales, fusing the low-level features, and obtaining low-level fused features include: Perform upsampling and downsampling operations on the dual-temporal remote sensing image to obtain a multi-scale image; Apply low-level feature extraction to the multi-scale image to obtain low-level features; Fuse the low-level features based on concatenation operation to obtain low-level fused features.
3. The remote sensing image change detection method according to claim 2, characterized in that, The operation formula for fusing the low-level features based on concatenation operation to obtain low-level fused features is: Among them, represents the low-level fusion feature; represents the bi-temporal remote sensing image, and its subscript represents the scaling ratio of the upsampling and downsampling operations on the bi-temporal remote sensing image; i represents the classification identifier of the images or features in different periods; and represent the upsampling and downsampling operations respectively; represents the concatenation operation; represents the structure of the first 7 layers adopted by the VGG network.
4. The remote sensing image change detection method according to claim 1, characterized in that The specific steps of calculating the cross-attention output based on the cross-attention mechanism according to the convolutional feature information include: Calculate a difference feature based on the convolutional feature information; The operation formula for calculating the difference feature is: Among them, D is the said differential feature, and respectively represent the said convolution feature information in different periods; Construct a query feature according to the difference feature; the formula for constructing the query feature is: Among them, Q represents the query feature, W q represents the query transformation matrix automatically learned by the multi-scale cross-attention network; Construct a key feature and a value feature based on the convolutional feature information, and the formulas for constructing the key feature and the value feature are: Among them, K represents the key feature, V represents the value feature; W k and W v respectively represent the key transformation matrix and the value transformation matrix automatically learned by the multi-scale cross-attention network; i represents the classification identifier of images or features in different periods; Adopt the cross-attention mechanism to calculate the cross-attention output according to the query feature, key feature, and value feature; The operation formula for calculating the cross-attention output is: Among them, d is the channel dimension.
5. The remote sensing image change detection method according to claim 1, characterized in that The specific steps of performing a reconstruction operation on the enhanced fused feature map based on the context information linking algorithm to obtain a reconstructed feature map include: Assign corresponding weights to the low-level features; Concatenate the low-level features based on the weights to obtain a concatenated feature map; Apply a convolution operation to the concatenated feature map to obtain a spatial weight feature map; Multiply the spatial weight feature map by the enhanced fused feature map to obtain a reconstructed feature map.
6. The remote sensing image change detection method according to claim 1, wherein The operation formula for reconstructing the enhanced fusion feature map based on the context information linking algorithm to obtain a reconstructed feature map is as follows: Among them, represents the reconstructed feature map; represents the enhanced fusion feature map; , and respectively represent the low-level features of different scales; i represents the classification identifier of images or features at different times; a , b and c respectively represent the weight values corresponding to the low-level features; RELU represents the activation function, BN represents batch normalization, and Conv represents the convolution operation.
7. The remote sensing image change detection method according to claim 1, wherein The specific steps further include: Constructing a loss function according to the image parameters of the final recognition image; wherein, the image parameters include the height, width, true pixel value and predicted pixel value of the final recognition image; Training the multi-scale cross-attention network based on the loss function to obtain a target multi-scale cross-attention network.
8. A remote sensing image change detection device, characterized in that, It includes: A network construction module for inputting dual-temporal remote sensing images into a multi-scale cross-attention network; wherein, the multi-scale cross-attention network is composed of multi-scale feature encoding, differential feature enhancement algorithm and context information linking; A multi-scale feature fusion module for extracting low-level features of the dual-temporal remote sensing images based on different feature scales, fusing the low-level features, and obtaining low-level fusion features; wherein, the low-level fusion features correspond to different time phases respectively; A differential feature encoding module for adaptively fusing the low-level fusion features through convolution to obtain convolution feature information; calculating a cross-attention output based on the cross-attention mechanism according to the convolution feature information; adding the cross-attention output and the convolution feature information according to the time phase to obtain a fusion feature combination; wherein, the fusion feature combinations correspond to different time phases respectively; concatenating the fusion feature combinations to obtain a concatenated fusion feature; processing the concatenated fusion feature with a multi-layer perceptron to obtain a deep fusion feature; performing a deformation operation on the deep fusion feature to obtain an enhanced fusion feature map; A context information linking module for performing a reconstruction operation on the enhanced fusion feature map based on the context information linking algorithm to obtain a reconstructed feature map; wherein, the reconstruction operation is used to enhance the key change information in the enhanced fusion feature map based on spatial weights; An identification image output module for performing a subtraction operation on the reconstructed feature map according to different time phases to obtain a final recognition image.
9. An electronic device, characterized in that, It includes: A memory, a processor, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, it performs the steps of the remote sensing image change detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
CNN-Transform-based remote sensing image water body extraction method and device, electronic equipment and medium
CN117036805A
Remote sensing image change detection method based on hierarchical cross-scale global feature fusion deep network
CN117853897A