High-resolution remote sensing image change detection method, system and equipment of cross attention network based on semantic guidance and medium
Multi-level feature fusion is carried out through semantic-guided cross attention network, which solves the problem of insufficient accuracy and reliability of high-resolution remote sensing image change detection in the prior art, and achieves more efficient change detection.
Patent Information
- Application Number
- CN202510297013.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
The existing change detection networks are difficult to make full use of the multi-level features extracted by the backbone network, resulting in low accuracy and reliability of change detection of high-resolution remote sensing images.
A semantic-guided cross-attention network is adopted, including an initial hierarchical feature extraction module, a cross-attention feature fusion module, a hierarchical semantic-guided fusion module and a classifier, and improve feature utilization through multi-level feature fusion and semantic guidance.
It improves the accuracy and reliability of high-resolution remote sensing image change detection, effectively identify key change information, and reduces false detection and missed detection.
Smart Images

Figure CN120236207A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular, to a high-resolution remote sensing image change detection method, system, device and medium based on a semantic-guided cross-attention network. Background Art
[0002] Change detection is the process of identifying the state differences of the same object or phenomenon by observing it at different times. Satellite remote sensing technology enables multi-temporal observations, provides an important technical means for understanding the continuous changes on the Earth's surface, and has been widely applied in many fields such as urban expansion research, land use and cover change, and ecological environment management. Especially in the field of disaster management, remote sensing change detection is crucial for monitoring natural disasters such as floods, earthquakes, and landslides.
[0003] With the continuous progress of earth observation technology, more and more remote sensing data with hyperspectral, high spatial resolution, and high temporal resolution have emerged. Different from medium- and low-resolution remote sensing images, high-resolution remote sensing images can provide rich surface details and spatial distribution information. How to effectively extract real change features from high-resolution remote sensing images while reducing the impact of pseudo-changes such as illumination, shadows, and noise remains an important challenge in remote sensing change detection.
[0004] Existing change detection networks often have difficulty in fully utilizing the multi-level features extracted by the backbone network. These features (from low-level spatial details to high-level semantic representations) are either not fully utilized or not correctly aligned, resulting in a decline in performance in complex changes and lower accuracy and reliability of change detection. Summary of the Invention
[0005] The purpose of the present application is to provide a high-resolution remote sensing image change detection method, system, device and medium based on a semantic-guided cross-attention network, which can improve the accuracy and reliability of high-resolution remote sensing image change detection.
[0006] To achieve the above purpose, the present application provides the following solutions:
[0007] In a first aspect, the present application provides a high-resolution remote sensing image change detection method based on a semantic-guided cross-attention network, including:
[0008] Obtain two remote sensing images to be detected;
[0009] Construct and train a semantic-guided cross-attention network; the semantic-guided cross-attention network includes an initial hierarchical feature extraction module, a cross-attention feature fusion module, a hierarchical semantic guidance fusion module, and a classifier;
[0010] Input the two remote sensing images to be detected into the trained semantic-guided cross-attention network to generate a change map; the change map is used to characterize the changes in the remote sensing images.
[0011] Among them, inputting the two remote sensing images to be detected into the trained semantic-guided cross-attention network to generate a change map specifically includes:
[0012] Input the two remote sensing images to be detected into the initial hierarchical feature extraction module for feature extraction to obtain multiple groups of feature maps.
[0013] Input the multiple groups of feature maps into the cross-attention feature fusion module respectively for feature fusion to obtain multiple fused feature maps.
[0014] Input the multiple fused feature maps into the hierarchical semantic-guided fusion module for multi-level feature fusion to obtain the final feature map.
[0015] Input the final feature map into the classifier to generate a change map.
[0016] In a second aspect, the present application provides a high-resolution remote sensing image change detection system based on a semantic-guided cross-attention network, including:
[0017] An image acquisition unit for acquiring two remote sensing images to be detected.
[0018] A network construction and training unit for constructing and training a semantic-guided cross-attention network; the semantic-guided cross-attention network includes an initial hierarchical feature extraction module, a cross-attention feature fusion module, a hierarchical semantic-guided fusion module, and a classifier.
[0019] A change detection unit for inputting the two remote sensing images to be detected into the trained semantic-guided cross-attention network to generate a change map; the change map is used to characterize the changes in the remote sensing images.
[0020] Among them, the initial hierarchical feature extraction module is used to perform feature extraction on the two remote sensing images to be detected to obtain multiple groups of feature maps.
[0021] The cross-attention feature fusion module is used to perform feature fusion on multiple groups of feature maps to obtain multiple fused feature maps.
[0022] The hierarchical semantic-guided fusion module is used to perform multi-level feature fusion on multiple fused feature maps to obtain the final feature map.
[0023] The classifier is used to classify the final feature map to generate a change map.
[0024] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned high-resolution remote sensing image change detection method based on a semantic-guided cross-attention network.
[0025] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned high-resolution remote sensing image change detection method based on a semantic-guided cross-attention network.
[0026] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0027] The present application provides a high-resolution remote sensing image change detection method, system, device, and medium based on a semantic-guided cross-attention network. A hierarchical semantic guidance fusion module is introduced into the semantic-guided cross-attention network, and this module uses high-level semantic information to guide low-level spatial details through an attention mechanism. In addition, in order to enhance the information interaction between bi-temporal images, the semantic-guided cross-attention network model integrates a cross-attention feature fusion module to establish the global correlation between bi-temporal images. By effectively mining complementary features, this module can better identify key change information, thereby improving the accuracy and reliability of change detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and for those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] Figure 1 It is a schematic flowchart of a high-resolution remote sensing image change detection method based on a semantic-guided cross-attention network provided by an embodiment of the present application;
[0030] Figure 2 It is an overall architecture diagram of a semantic-guided cross-attention network;
[0031] Figure 3 It is an architecture diagram of an initial hierarchical feature extraction module;
[0032] Figure 4 It is an architecture diagram of a cross-attention feature fusion module;
[0033] Figure 5 It is an architecture diagram of a multi-scale parallel convolution module;
[0034] Figure 6 Schematic diagram of a computer device provided in an embodiment of the present application. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0036] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the drawings and specific implementation manners.
[0037] In an exemplary embodiment, as Figure 1 shown, a high-resolution remote sensing image change detection method based on a semantic-guided cross-attention network is provided. This method is executed by a computer device, specifically, it can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking the application of this method to a server as an example for illustration, it includes the following steps S1 to S3. Among them:
[0038] S1: Obtain two remote sensing images to be detected.
[0039] S2: Construct and train a semantic-guided cross-attention network; the semantic-guided cross-attention network includes an initial hierarchical feature extraction module, a cross-attention feature fusion module, a hierarchical semantic guidance fusion module, and a classifier. The overall framework of the semantic-guided cross-attention network is as Figure 2 shown.
[0040] S3: Input the two remote sensing images to be detected into the trained semantic-guided cross-attention network to generate a change map; the change map is used to characterize the changes in the remote sensing images.
[0041] In a specific embodiment, step S3 specifically includes:
[0042] S31: Input the two remote sensing images to be detected into the initial hierarchical feature extraction module for feature extraction to obtain multiple groups of feature maps.
[0043] In this embodiment, the initial hierarchical feature extraction module is used to extract multi-level image features. First, two registered remote sensing images (i.e., the two remote sensing images to be detected) are input into the pre-trained VGG16_BN model, and features are gradually extracted through the five feature extraction blocks of the VGG16_BN model to generate five groups of feature maps and Here, i corresponds to the feature extraction stage from shallow to deep.
[0044] The core architecture of the VGG16 - BN model includes convolutional layers, pooling layers, activation layers, and fully - connected layers. To better meet the requirements of change detection in dual - temporal remote sensing images, the fully - connected layer is removed in this embodiment.
[0045] The two registered remote sensing images T1 and T2 as input have the same spatial and channel dimensions, denoted as H×W×C, where H and W represent the height and width of the image respectively, and C represents the number of channels.
[0046] Specifically, as Figure 3 shown, the five feature extraction blocks respectively correspond to the 0 - 5th layer, 5 - 12th layer, 12 - 22nd layer, 22 - 32nd layer, and 32 - 42nd layer of the VGG16_BN model. T1 and T2 sequentially pass through these five feature extraction blocks, gradually extract features, and generate five groups of feature maps and The spatial and channel dimensions of the feature maps output by each feature extraction block are as Figure 3 shown.
[0047] S32: Input the multiple groups of feature maps into the cross - attention feature fusion module respectively for feature fusion to obtain multiple fused feature maps. The cross - attention feature fusion module includes a hybrid pooling layer and a dual - path cross - attention mechanism module, and the dual - path cross - attention mechanism module includes two cross - attention temporal modules.
[0048] Step S32 specifically includes: Input the multiple groups of feature maps into the hybrid pooling layer respectively for average pooling and max - pooling processing to obtain multiple groups of pooled feature maps. Flatten the multiple groups of pooled feature maps and add positional encoding to obtain multiple groups of feature maps with positional encoding; Input the multiple groups of feature maps with positional encoding into the dual - path cross - attention mechanism module respectively for feature fusion to obtain multiple fused feature maps; The fused feature maps include multiple high - level feature maps and multiple low - level feature maps.
[0049] This embodiment uses a cross - attention feature fusion module to establish the global correlation between the dual - temporal feature maps. The feature maps from the same feature extraction block in the initial hierarchical feature extraction module and It is input into the cross-attention feature fusion module. First, the feature map is input into the hybrid pooling layer to compress the spatial dimension of the feature map, reducing the computational complexity. Then, the pooled feature map is passed to the dual-path cross-attention mechanism module to enhance the feature interaction between feature maps and extract key change information. In the dual-path cross-attention mechanism module, first, each pooled feature map is flattened into a set of tokens, learnable positional encoding is added, and the cross-attention mechanism is applied to enhance the information interaction between tokens. Finally, the outputs of the dual-path cross-attention mechanism are concatenated along the channel dimension and downsampled through a convolutional layer to obtain Fusion i , where i corresponds to the serial number of the feature extraction block in the initial hierarchical feature extraction module.
[0050] (1) As Figure 4 shown, the specific operation process of the hybrid pooling layer is as follows:
[0051] Given the input feature map F T1 , the hybrid pooling layer processes this feature map using average pooling F avg (·) and max pooling F max (·) respectively. After pooling, the spatial dimension of the feature map is compressed by the pooling factor S, generating and These pooling results will be combined in the following way:
[0052]
[0053] Here, λ1 and λ2 are learnable coefficients, with an initial value of 0.5 and satisfying λ1 + λ2 = 1.
[0054] The hybrid pooling layer combines the advantages of average pooling F avg (·) and max pooling F max (·), effectively compressing the spatial dimension of the feature map while retaining key information.
[0055] (2) As Figure 4 shown, the specific operation process of the dual-path cross-attention mechanism module is as follows:
[0056] The dual-path cross-attention module consists of two independent cross-attention phase modules: CAM T1 (Phase I) and CAM T2 (Phase II), and the weights are not shared between the two CAM modules. This design allows each module to adaptively optimize its input, ensuring effective feature interaction. Each CAM module regards the feature map of one phase as Query, and the feature maps of the other phase as Key and Value for cross-attention operations, thereby establishing the global association between dual-time features and effectively extracting key information.
[0057] Specifically, after being processed by the initial hierarchical feature extraction module and the hybrid pooling layer, a dual-temporal feature map (i.e., the pooled feature map) Feature is obtained. T1 and To effectively capture the location information (which is crucial for the change detection task), first flatten Feature T1 and Feature T2 to generate Tokens 1 and and add learnable position encoding (see Equation (2)) to enhance the model's spatial perception ability. The generated and with position encoding, that is, the feature map with position encoding, is then input into the dual-path cross-attention module for feature fusion.
[0058]
[0059] For clarity, the calculation mode of CAM T1 will be described in detail. CAM T2 follows a symmetric calculation mode.
[0060]
[0061] The calculation details of CAM T1 are as follows: First, the input dual-temporal features and are layer-normalized to enhance the stability of the network.
[0062] Next, is projected into matrices K and through different linear mappings T1 and V T1 . At the same time, is also projected into matrix Q through linear mapping T2 .
[0063]
[0064] In addition, the model adopts a multi-head cross-attention mechanism with eight parallel heads. By using three different weight matrices and to divide K T1 , V T1 , Q T2 into multiple attention heads enables the model to jointly understand the relationship between dual-temporal features from different perspectives.
[0065]
[0066] Here, h represents the serial number of different attention heads, and the weight matrix corresponding to each head.
[0067] Subsequently, for each attention head h, the result of the dot product of Q and K and After softmax regularization, the normalized attention matrix is obtained ensuring that the sum of all elements is 1.0.
[0068]
[0069] Here, d k represents the dimension.
[0070] After that, the normalized attention matrix of each head is used as the weight to multiply with the matrix to generate eight attention-weighted feature matrices Subsequently, these matrices are concatenated along the channel dimension to form a joint matrix, and the learnable weight matrix W O is applied to re-project the joint matrix back to the original feature space to generate Z T1 .
[0071]
[0072] Here, Concat(·) represents the concatenation operation.
[0073] After that, the input is added to Z T1 through a residual connection to retain the original input information. The mathematical expression can be represented as:
[0074]
[0075] At the end of CAM T1 the feature map is input into a feed-forward network (FFN) for further processing. The feed-forward network includes a linear layer and a non-linear activation function GELU. In addition, a residual connection is introduced to enhance the robustness of the model and prevent overfitting. The processing of the feed-forward network can be represented in the following form:
[0076]
[0077] where, W 1 and W 2represent the weight matrices of two linear layers in the feed-forward network, with channel dimensions of C×4C and 4C×C respectively. and represent the bias terms.
[0078] It should be noted that learnable coefficients are introduced in each branch of the residual connection. In equations (8) and (10), α, β, γ, and δ are trainable parameters with an initial value of 1, which allows the model to adaptively balance the contributions of different branches during the optimization process.
[0079] As mentioned before, CAM T2 adopts a computational mode symmetric to that of CAM T1 CAM T2 projects onto the Q T1 matrix, and projects onto the K T2 and V T2 matrices. This process can be represented in the following way:
[0080]
[0081] After completing the cross-attention calculation, the results output by CAM T1 and CAM T2 are concatenated along the channel dimension, and a 1×1 convolutional layer is applied to the concatenated result for downsampling to generate the fused feature map Fusion.
[0082]
[0083] Here, conv 1×1 (·) represents the convolutional operation with a kernel size of 1×1, which can reduce the channel dimension from 2C to C.
[0084] S33: Input multiple fused feature maps into the hierarchical semantic-guided fusion module for multi-level feature fusion to obtain the final feature map. The hierarchical semantic-guided fusion module includes: a multi-scale parallel convolution module, a progressive feature aggregation module, and a semantic-guided refinement module; the progressive feature aggregation module includes a first progressive feature aggregation module and a second progressive feature aggregation module.
[0085] Step S33 specifically includes: respectively inputting multiple high-level feature maps into the multi-scale parallel convolution module to extract multi-scale semantic information, obtaining corresponding high-level semantic information feature maps; inputting the high-level semantic information feature maps into the first progressive feature aggregation module to generate a semantic guidance map rich in high-level semantic information; inputting the semantic guidance map rich in high-level semantic information and multiple low-level feature maps into the semantic guidance refinement module to obtain multiple optimized low-level feature maps; inputting the multiple optimized low-level feature maps into the multi-scale parallel convolution module and the second progressive feature aggregation module to obtain the final feature map.
[0086] Through the hierarchical semantic guidance fusion module, high-level semantic information is used to guide low-level spatial features. First, the high-level feature maps Fusion 4 、Fusion 3 and Fusion 2 are input into the multi-scale parallel convolution module to extract multi-scale semantic information, generating and Then, these features are aggregated through the first progressive feature aggregation module to generate a semantic guidance map rich in high-level semantic information. This map is used to optimize Fusion 1 and Fusion 0 through the attention mechanism. Similar to the high-level feature fusion process, these optimized low-level feature maps are further processed through the multi-scale parallel convolution module and the second progressive feature aggregation module to generate the final feature map
[0087] (1) Multi-scale parallel convolution module
[0088] The multi-scale parallel convolution module aims to extract multi-scale features by balancing fine-grained local details and coarse-grained spatial dependencies. As Figure 5 shown, this module consists of 5 parallel branches, each branch is marked by the branch index N. The branch index determines the receptive field size of the corresponding branch, enabling the model to focus on specific spatial scales and feature dependencies.
[0089] Each branch is constructed using a basic convolution block, and the basic convolution block includes a two-dimensional convolutional layer, a batch normalization layer, and an activation function. All branches start with a basic convolution block with a kernel size of 1×1 to extract fine-grained local features. For N = 1, 2, 3, each branch further adds two asymmetric basic convolution blocks (with kernel sizes of 1×(2 N +1) and (2 N +1)×1) to enhance the directional feature representation. In addition, a 3×3 convolution block with a dilation rate of 2 N +1 is used to capture broader spatial dependencies.
[0090] Specifically, branch 0 (N = 0) only contains a 1×1 basic convolutional block, which focuses on local feature extraction. Branch 1 (N = 1) contains 1×1, 1×3, 3×1 basic convolutional blocks, as well as a 3×3 basic convolutional block with a dilation rate of 3, used to capture medium and short-range feature dependencies. Branches 2 (N = 2) and 3 (N = 3) use larger convolutional kernels. Branch 2 uses 1×5, 5×1 basic convolutional blocks; Branch 3 uses 1×7, 7×1 basic convolutional blocks. The dilation rates of branches 2 and 3 are 5 and 7 respectively, to extract medium and long-term dependencies and broader spatial features.
[0091] The outputs of branches 0 to 3 are concatenated along the channel dimension and the multi-scale features are fused through a 3×3 basic convolutional block. In addition, branch 4 (N = 4) is used to combine the original features with the multi-scale fused features through a residual connection, ensuring that the original information is retained. Finally, the ReLU function is used to process the generated feature maps to enhance the feature representation for downstream tasks.
[0092] (2) Progressive Feature Aggregation Module
[0093] To more effectively aggregate multi-level image features, this embodiment designs two progressive feature aggregation modules. The first progressive feature aggregation module aggregates the feature maps rich in high-level semantic information to generate a semantic guidance map; the second progressive feature aggregation module aggregates the low-level feature maps optimized under the guidance of high-level semantic information. Using this module, the model can gradually aggregate the feature information from different levels of the network in an iterative manner.
[0094] For ease of description, this embodiment only describes the first progressive feature aggregation module. The operation of the second progressive feature aggregation module is the same as that of the first progressive feature aggregation module. In the first progressive feature aggregation module, the feature maps from three input layers and are gradually aggregated. This module uses upsampling, convolution, and feature concatenation operations. Specifically, and The result of the upsampling operation is obtained by element-wise multiplication to get Here, the upsampling operation refers to magnifying the input feature map and then inputting it into a 3×3 basic convolutional block for further processing. Similarly, we combine and and respectively, after the upsampling operation, to obtain
[0095]
[0096] In the above formula, Upsample(·) is the upsampling with a scale factor of 2, used to magnify the input feature map.
[0097] Subsequently, and After the upsampling operation, the results are concatenated along the channel dimension, and the concatenated feature map is processed by a 3×3 basic convolution block. Finally, an aggregated output feature map X is generated through a 1×1 convolutional layer agg .
[0098]
[0099] (3) Semantic-guided Refinement Module
[0100] The semantic-guided refinement module uses the semantic guidance map generated by the first progressive feature aggregation module to optimize the low-level feature map. The semantic guidance map has rich global semantic information and optimizes the low-level feature map by integrating local and global contexts
[0101] First, the semantic guidance map is upsampled to match the spatial dimension of the target feature map. Specifically, for upsampling is performed with a scale factor of 4, and for Fusion 1 upsampling is performed with a scale factor of 2, and for Fusion 0 , the semantic feature map has the same spatial dimension as it and does not require upsampling. This enables the semantic guidance map to maintain the same resolution as feature maps at different levels. Then, the upsampled semantic guidance map is applied to the low-level feature map using matrix multiplication. This operation enhances the target feature map through global semantic guidance. Finally, the enhanced feature map is combined with the original feature map through a residual connection. This operation ensures that while retaining the original feature information, rich global context information is integrated. Therefore, the final feature map combines fine-grained local information with global semantic information, thereby improving the model's understanding of details and global information
[0102] S34: Input the final feature map into the classifier to generate a change map
[0103] Finally, the feature map that fully fuses multi-level features is input into the classifier to generate a change map. Here, the classifier is a binary convolutional layer
[0104] In this embodiment, to optimize the semantic-guided cross-attention network, a cross-entropy loss function is used to measure the difference between the predicted value and the ground truth. The loss function is defined as:
[0105]
[0106] where H0 and W0 respectively represent the height and width of the input image. y h,w represents the binary ground truth of the pixel (h,w), where 1 indicates that the pixel belongs to the changed area and 0 indicates no change Indicates the probability that the predicted pixel (h, w) belongs to the changed region.
[0107] The semantic-guided cross-attention network provided by this application effectively solves the limitation of making full use of multi-level features extracted by the backbone network by combining cross-attention feature fusion and hierarchical semantic-guided fusion modules; the cross-attention mechanism establishes global correlations between bi-temporal images, enabling the model to better extract and integrate complementary information.
[0108] Next, a large number of experiments will be conducted to verify the effectiveness of the model proposed in this application, including the introduction of experimental data, evaluation metrics, training details, and analysis of experimental results. The experimental results show that the model proposed in this application is superior to other models and achieves excellent performance.
[0109] (1) Experimental data
[0110] 1) IWHR-data dataset
[0111] The IWHE-data dataset is a publicly available change detection dataset collected from GF and ZY series satellite platforms, providing a spatial resolution of 2 meters per pixel. This dataset covers 26 key districts and counties in Shandong Province and Anhui Province of China. The project aims to support soil and water conservation and flood prevention work, focusing on land use changes related to large-scale alterations caused by buildings, river channels, pipelines, newly built roads, and human activities (such as excavation, occupation, dumping, and surface damage). The dataset consists of 2960 pairs of images, each image containing 512×512 pixels. By default, 2072 pairs are used for training, 592 pairs for validation, and 296 pairs for testing. To address GPU memory limitations and improve training efficiency, the images are cropped to 256×256 pixels before being input into the model.
[0112] 2) LEVIR-CD dataset
[0113] The LEVIR-CD dataset is a high-resolution building change detection dataset collected from Google Earth, with a spatial resolution of 0.5 meters per pixel. It contains significant land use changes, including various types of buildings such as villa residences, high-rise apartments, small garages, and large warehouses. This dataset focuses on three types of change scenarios: building addition, building demolition, and no change. The LEVIR-CD dataset consists of 637 pairs of 1024×1024 image patches, and by default, it is divided into a training set (445 pairs), a validation set (64 pairs), and a test set (128 pairs). Due to GPU memory limitations, the original image pairs are cropped into non-overlapping 256×256 patches, obtaining 7120 pairs, 1024 pairs, and 2048 pairs of patches for the training set, validation set, and test set respectively.
[0114] (2) Evaluation metrics
[0115] To comprehensively evaluate the performance of the proposed model, the following metrics are adopted in this embodiment: Precision, Recall, F1-score, Intersection over Union (IoU), and Overall Accuracy (OA). It is worth noting that higher F1-score and IoU values imply better model performance in the change detection task.
[0116]
[0117] Here TP TP is the number of changed pixels correctly identified as changed, and TN is the number of unchanged pixels correctly identified as unchanged. FP FP is the number of unchanged pixels misidentified as changed, while FN is the number of changed pixels misidentified as unchanged.
[0118] (3) Training details
[0119] The semantic-guided cross-attention network model proposed in this embodiment is implemented using PyTorch and trained on an NVIDIA TESLA-V100 GPU. To enhance the generalization ability of the model, data augmentation strategies are adopted during training, including horizontal and vertical flipping, scale cropping, and Gaussian blur.
[0120] To optimize the model parameters, Stochastic Gradient Descent (SGD) with a momentum of 0.9 and a weight decay of 0.0005 is used. Due to the limitation of GPU memory, the batch size is set to 8. The initial learning rate is set to 0.01, and the learning rate is linearly decayed to zero in 200 epochs. During training, the model parameters that achieve the best performance on the validation set are saved, and the saved parameters are applied to the test set for the final evaluation.
[0121] (3) Model comparison
[0122] To evaluate the performance of the semantic-guided cross-attention network model, three categories of state-of-the-art change detection models are selected for comparison: CNN-based models (FC-EF, FC-Siam-conc, FC-Siam-diff), attention mechanism-based models (STANet-BAM, STANet-PAM, SNUNet), and transformer-based models (ResNet18+Transformer, BIT, and ChangeFormer).
[0123] It should be noted that all the models compared with the semantic-guided cross-attention network model are implemented using the open-source code publicly available on GitHub with default hyperparameter settings. To ensure fairness, all models are trained, validated, and tested on the same dataset.
[0124] Tables 1 and 2 respectively present the quantitative comparison results of the semantic-guided cross-attention network model and other models on the IWHE-data dataset and the LEVIR-CD dataset. As shown in Tables 1 and 2, the semantic-guided cross-attention network model outperforms the other three types of models on both datasets.
[0125] Table 1
[0126]
[0127] Table 2
[0128]
[0129]
[0130] As can be seen from Table 2, the semantic-guided cross-attention network model has the best performance on the IWHE-data dataset, with an F1-score of 0.90796, an IoU score of 0.83143, and an OA score of 0.99153. In contrast, the semantic-guided cross-attention network model effectively reduces false detections while minimizing missed detections.
[0131] Table 3 presents the comparison results on the LEVIR-CD dataset. The semantic-guided cross-attention network model again achieves the best results in the three key metrics of F1-score, IoU, and OA, with scores of 0.91146, 0.83733, and 0.99116 respectively, further demonstrating the effectiveness of the semantic-guided cross-attention network model.
[0132] The research results show that compared with other models, the semantic-guided cross-attention network model makes full use of the multi-level features extracted by the initial hierarchical feature extraction module. By combining the cross-attention mechanism, it can more effectively capture the change information in the bi-temporal images, thus significantly improving the evaluation metrics.
[0133] Based on the same inventive concept, an embodiment of the present application further provides a high-resolution remote sensing image change detection system based on a semantic-guided cross-attention network. The implementation solution provided by this system to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the high-resolution remote sensing image change detection system based on the semantic-guided cross-attention network provided below can refer to the limitations on the high-resolution remote sensing image change detection method based on the semantic-guided cross-attention network in the above text, and will not be repeated here.
[0134] In an exemplary embodiment, a high-resolution remote sensing image change detection system based on a semantic-guided cross-attention network is provided, including:
[0135] An image acquisition unit for acquiring two remote sensing images to be detected;
[0136] A network construction and training unit for constructing and training a semantic-guided cross-attention network; the semantic-guided cross-attention network includes an initial hierarchical feature extraction module, a cross-attention feature fusion module, a hierarchical semantic guidance fusion module, and a classifier;
[0137] A change detection unit for inputting the two remote sensing images to be detected into the trained semantic-guided cross-attention network to generate a change map; the change map is used to represent the changes in the remote sensing image;
[0138] Wherein, the initial hierarchical feature extraction module is used to extract features from the two remote sensing images to be detected to obtain multiple groups of feature maps;
[0139] The cross-attention feature fusion module is used to perform feature fusion on multiple groups of feature maps to obtain multiple fused feature maps;
[0140] The hierarchical semantic guidance fusion module is used to perform multi-level feature fusion on multiple fused feature maps to obtain a final feature map;
[0141] The classifier is used to classify the final feature map to generate a change map.
[0142] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented. This computer device can be a server or a terminal, and its internal structure diagram can be as Figure 6As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data to be processed. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a high-resolution remote sensing image change detection method based on a semantically guided cross-attention network is implemented.
[0143] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0144] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0145] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0146] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0147] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0148] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0149] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A high-resolution remote sensing image change detection method based on semantically guided cross-attention network, characterized in that: include: Acquire two remote sensing images to be detected; Constructing and training a semantically guided cross-attention network; the semantically guided cross-attention network includes an initial hierarchical feature extraction module, a cross-attention feature fusion module, a hierarchical semantic guided fusion module and a classifier; Inputting the two remote sensing images to be detected into a trained semantically guided cross attention network to generate a change map; the change map is used to characterize the change of the remote sensing image; The two remote sensing images to be detected are input into the trained semantically guided cross attention network to generate a change map, which specifically includes: Inputting the two remote sensing images to be detected into an initial hierarchical feature extraction module for feature extraction to obtain multiple sets of feature maps; Inputting multiple groups of feature maps into the cross-attention feature fusion module for feature fusion to obtain multiple fused feature maps; Inputting multiple fused feature maps into the hierarchical semantic guided fusion module to perform multi-level feature fusion to obtain a final feature map; The final feature map is input into the classifier to generate a change map.
2. The high-resolution remote sensing image change detection method based on semantically guided cross-attention network according to claim 1 is characterized in that: The initial level feature extraction module adopts the VGG16-BN model, and the VGG16-BN model includes multiple feature extraction blocks.
3. The high-resolution remote sensing image change detection method based on semantically guided cross-attention network according to claim 1 is characterized in that: The cross-attention feature fusion module includes a mixed pooling layer and a dual-path cross-attention mechanism module.
4. The high-resolution remote sensing image change detection method based on semantically guided cross-attention network according to claim 3 is characterized in that: Inputting multiple groups of feature maps into the cross-attention feature fusion module for feature fusion, respectively, to obtain multiple fused feature maps, specifically including: Inputting the multiple groups of feature maps into the mixed pooling layer for average pooling and maximum pooling processing respectively, to obtain multiple groups of pooling feature maps; Flatten multiple groups of pooled feature maps and add position encoding to obtain multiple groups of feature maps with position encoding; Multiple groups of feature maps with position encoding are respectively input into the dual-path cross-attention mechanism module for feature fusion to obtain multiple fused feature maps; the fused feature maps include multiple high-level feature maps and multiple low-level feature maps.
5. The high-resolution remote sensing image change detection method based on semantically guided cross-attention network according to claim 3 is characterized in that: The dual-path cross-attention mechanism module includes two cross-attention phase modules.
6. The high-resolution remote sensing image change detection method based on semantically guided cross-attention network according to claim 4 is characterized in that: The hierarchical semantics-guided fusion module includes: a multi-scale parallel convolution module, a progressive feature aggregation module and a semantics-guided refinement module; the progressive feature aggregation module includes a first progressive feature aggregation module and a second progressive feature aggregation module.
7. The high-resolution remote sensing image change detection method based on semantically guided cross-attention network according to claim 6 is characterized in that: Inputting multiple fused feature maps into the hierarchical semantic guided fusion module to perform multi-level feature fusion to obtain a final feature map, specifically including: Inputting multiple high-level feature maps into the multi-scale parallel convolution module to extract multi-scale semantic information and obtain corresponding high-level semantic information feature maps; Inputting the high-level semantic information feature map into the first progressive feature aggregation module to generate a semantic guidance map rich in high-level semantic information; Inputting a semantic guidance map rich in high-level semantic information and multiple low-level feature maps into a semantic guidance refinement module to obtain multiple optimized low-level feature maps; The multiple optimized low-level feature maps are input into the multi-scale parallel convolution module and the second progressive feature aggregation module to obtain the final feature map.
8. A high-resolution remote sensing image change detection system based on semantically guided cross-attention network, characterized in that: include: An image acquisition unit, used for acquiring two remote sensing images to be detected; A network construction and training unit, used to construct and train a semantically guided cross-attention network; the semantically guided cross-attention network includes an initial hierarchical feature extraction module, a cross-attention feature fusion module, a hierarchical semantic guided fusion module and a classifier; A change detection unit, used for inputting the two remote sensing images to be detected into a trained semantically guided cross attention network to generate a change map; the change map is used to characterize the change of the remote sensing image; The initial level feature extraction module is used to extract features from the two remote sensing images to be detected to obtain multiple sets of feature maps; The cross-attention feature fusion module is used to perform feature fusion on multiple groups of feature maps to obtain multiple fused feature maps; The hierarchical semantic guided fusion module is used to perform multi-level feature fusion on multiple fusion feature maps to obtain a final feature map; The classifier is used to classify the final feature map and generate a change map.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the high-resolution remote sensing image change detection method based on a semantically guided cross-attention network as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the high-resolution remote sensing image change detection method based on a semantically guided cross-attention network as described in any one of claims 1 to 7.