Semi-supervised remote sensing image change detection method based on deep attention and class imbalance false label enhancement
By employing deep attention and class-imbalanced pseudo-label enhancement methods, this study addresses the issues of insufficient modeling of pixel relationships and low pseudo-label quality in remote sensing image change detection. This improves detection accuracy and model generalization ability, and significantly enhances the credibility of pseudo-labels, especially in the case of class imbalance.
Patent Information
- Application Number
- CN202510839154.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-28
AI Technical Summary
Existing semi-supervised remote sensing image change detection methods struggle to effectively model the spatial correlation and semantic similarity between pixels, leading to inaccurate predictions. Furthermore, they neglect the direct guidance of encoded features on pseudo-labels, resulting in poor pseudo-label quality, especially in cases of class imbalance, which negatively impacts model performance and generalization ability.
A deep attention module is used to combine channel attention, spatial attention and self-attention mechanisms to construct a similarity matrix to capture global dependencies between features. The pseudo-label quality is optimized by a class imbalance pseudo-label enhancement module. Combined with the prediction of high-confidence regions and change categories, the pseudo-label generation process is improved.
It improves the accuracy of remote sensing image change detection and the model's generalization ability, enhances the ability to detect subtle changes, optimizes the credibility of false labels, alleviates the class imbalance problem, and improves the performance of change detection tasks.
Smart Images

Figure CN120852901A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of remote sensing image processing technology, specifically relating to a semi-supervised remote sensing image change detection method based on depth attention and class-imbalanced pseudo-label enhancement. Background Technology
[0002] In recent years, the rapid development of remote sensing technology has made acquiring high-resolution remote sensing imagery more feasible, providing rich data support for surface monitoring, urban planning, and environmental assessment. In remote sensing data analysis, change detection (CD) is a key task, aiming to identify areas of change at the same location at different times, such as changes in land cover, structural morphology, and spatial location. CD technology is widely used in urban expansion, land use management, ecological environment monitoring, and disaster assessment, and its development is of great significance for improving the intelligent analysis capabilities of remote sensing imagery.
[0003] Traditional change detection (CD) methods mainly include analysis methods based on difference calculations and traditional machine learning. For example, image difference-based methods utilize pixel value changes in multi-temporal images to detect target regions, offering advantages such as computational simplicity and ease of implementation. However, they are susceptible to noise and cannot effectively utilize spatial structure information. Methods based on Change Vector Analysis (CVA) and Principal Component Analysis (PCA) extract major change patterns through feature transformations, providing more robust change information compared to simple difference methods, but are still limited by the ability to design features manually. Furthermore, CD methods based on traditional machine learning models such as Support Vector Machines (SVM) and Random Forests (RF) can learn the feature distribution of data, improving detection accuracy. However, these methods rely on feature engineering, have limited generalization ability, and are difficult to handle change detection tasks in complex scenes. Although traditional CD methods can identify changed regions to a certain extent, they are often limited by noise interference, illumination changes, and imaging conditions, making it difficult to guarantee detection accuracy.
[0004] In recent years, the introduction of deep learning has greatly promoted the development of change detection (CD) technology. End-to-end learning frameworks can automatically extract high-level features and improve the detection accuracy of changed regions. Deep learning-based CD methods mainly fall into two categories: single-stream networks and two-stream networks. Single-stream network methods use bi-temporal remote sensing images as a single input for feature extraction and change detection, achieving high computational efficiency, but failing to fully utilize the complementary information of the bi-temporal images. In contrast, two-stream networks use architectures similar to CNNs or Transformers to process bi-temporal images separately, and then use differential features for change detection, effectively capturing detailed features of changed regions and improving detection accuracy. However, these methods heavily rely on large amounts of labeled data, while in practical applications, obtaining high-quality labeled data is costly, limiting the practical application of fully supervised methods. Therefore, how to train using limited labeled data and a large amount of unlabeled data has become an important direction for CD research.
[0005] To alleviate the problem of insufficient labeled data, semi-supervised change detection (SSCD) methods have gradually become a research focus. Current SSCD methods are mainly divided into adversarial learning-based and consistency regularization-based methods. Adversarial learning-based methods improve the reliability of pseudo-labels through a game-like process of training the generator and discriminator, but are prone to introducing noise, leading to instability. Consistency regularization-based methods, on the other hand, construct a contrast mechanism that enhances strong and weak data, constraining the model to maintain prediction consistency under different perturbation conditions, and utilize confidence filtering strategies to optimize pseudo-label quality.
[0006] While existing SSCD methods can improve change detection accuracy with limited labeled data, they still face the following challenges: First, most methods select reliable pixels based on thresholds, neglecting the modeling of relationships between pixels and failing to fully utilize spatial structure and similarity information, leading to inaccurate predictions and limiting model performance and generalization ability. Second, previous methods mostly used confidence scores to determine the pseudo-label for the current pixel, assuming that high-confidence predictions are more accurate. However, these methods overemphasize the reliability of high-confidence predictions while ignoring the crucial guiding role of encoded features in pseudo-label generation. Especially when facing high uncertainty or imbalanced class distributions, this can easily lead to poor pseudo-label quality, thus affecting model performance. To address these issues, there is an urgent need for an SSCD method that can more effectively model pixel relationships, optimize pseudo-label quality, and improve class imbalance handling capabilities to enhance the accuracy and generalization ability of change detection tasks. Summary of the Invention
[0007] To overcome the shortcomings of the prior art, the present invention aims to provide a semi-supervised remote sensing image change detection method based on deep attention and class imbalance pseudo-label enhancement. This method solves the problems of existing methods, which have difficulty in effectively modeling the spatial correlation and semantic similarity between pixels, leading to inaccurate predictions, ignoring the direct guidance of encoded features on pseudo-labels, and being affected by class imbalance, making it difficult for the model to generate high-quality pseudo-labels. This method improves the accuracy of remote sensing image change detection and the generalization ability of the model.
[0008] This invention is achieved through the following technical solution:
[0009] A semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement includes the following steps:
[0010] S1. The remote sensing image dataset is preprocessed and divided into a training set, a validation set, and a test set. The training set includes labeled images and unlabeled images.
[0011] S2, Construct a semi-supervised remote sensing image change detection model. The semi-supervised remote sensing image change detection model includes a deep attention module and a class-imbalanced pseudo-label enhancement module. The deep attention module enhances the model's learning of the contextual relationships between image pixels by combining channel attention and spatial attention, and introduces a self-attention mechanism to construct a similarity matrix to capture the global dependencies between features. The class-imbalanced pseudo-label enhancement module optimizes pseudo-labels by fusing the similarity matrix with high-confidence regions.
[0012] S3, Given network training parameters, train the semi-supervised remote sensing image change detection model using the training set and validation set until the network converges;
[0013] S4. Input the test set into the semi-supervised remote sensing image change detection model trained in step S3, and output the remote sensing image change detection results.
[0014] Furthermore, the preprocessing of the remote sensing image dataset in step S1 specifically involves: randomly cropping the dual-temporal remote sensing images into image pairs of the same size, and performing data augmentation operations such as random angle rotation, random flipping, and feature perturbation.
[0015] Furthermore, in step S2, the semi-supervised remote sensing image change detection model uses DeepLabV3+ as the backbone network.
[0016] Furthermore, the training process in step S3 specifically includes:
[0017] Supervised training of labeled images: Labeled dual-temporal remote sensing images are input in parallel into a dual-encoder network. The dual encoders extract image features and perform difference calculations to capture the original difference features of the dual-temporal remote sensing images. Random dropout is applied to the extracted original difference features to obtain labeled feature enhancement features. The original difference features and labeled feature enhancement features are respectively used by the decoder to generate standard predictions. Feature correction is performed by a deep attention module to obtain supervised deep attention predictions.
[0018] Unsupervised training is performed on unlabeled images: Unlabeled dual-temporal remote sensing images are subjected to strong and weak enhancement processing to obtain strongly enhanced 1, strongly enhanced 2, and weakly enhanced unlabeled dual-temporal remote sensing images. These strongly enhanced 1, strongly enhanced 2, and weakly enhanced unlabeled dual-temporal remote sensing images are input in parallel into a dual encoder network. The extracted differential features are specifically strongly enhanced feature 1, strongly enhanced feature 2, and weakly enhanced feature. Then, the weakly enhanced feature is randomly dropped out to obtain the differential features of the unlabeled feature enhancement features. The strongly enhanced feature 1, strongly enhanced feature 2, weakly enhanced feature, and unlabeled feature enhancement features are then processed by the decoder to generate standard unsupervised predictions. The strongly enhanced feature 1, strongly enhanced feature 2, weakly enhanced feature, and unlabeled feature enhancement features are then processed by a deep attention module for feature correction to generate unsupervised deep attention predictions. The predictions obtained from the weakly enhanced features are then processed by argmax operation to generate original pseudo-labels. The original pseudo-labels are then optimized by a class-imbalanced pseudo-label enhancement module and further optimized by combining the prediction of the change category.
[0019] Furthermore, the feature correction via the deep attention module specifically involves:
[0020] 1.1) Channel attention calculation: The differential features are linearly mapped to obtain the mapped differential features. Where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map, respectively, for the mapped differential features. Perform channel attention calculation:
[0021]
[0022] Where CA(·) represents channel attention, which is calculated through global average pooling and a fully connected layer, and then... Perform a weighted operation. This represents the matrix multiplication operation;
[0023] 1.2) Based on the channel attention calculation results Perform spatial attention calculations:
[0024]
[0025] Where SA(·) represents spatial attention, This represents the matrix multiplication operation;
[0026] 1.3) Results of spatial attention calculation Perform a 1×1 convolution transformation:
[0027]
[0028] Among them, the similarity matrix After Softmax(·) normalization, the similarity between each location in the feature map and other locations is represented. Conv2d(·) represents the convolution operation. This indicates a transpose operation, where C is the number of channels;
[0029] 1.4) The similarity matrix The deep attention prediction is obtained by applying it to the logits output of the model decoder.
[0030]
[0031] in, This indicates an interpolation operation. This represents the logits output of the model decoder.
[0032] Furthermore, the pseudo-label optimization achieved through the class-imbalanced pseudo-label enhancement module specifically involves:
[0033] 2.1) For similarity matrices The row vectors of each row are normalized, and a binary mask is generated by thresholding and binarization.
[0034]
[0035] in, Indicates an indicator function, The expression represents the interpolation operation; z represents the similarity matrix. Each row vector in the array;
[0036] 2.2) By comparing the binarization masks and high confidence areas The overlap between the two predictions is considered if it exceeds a threshold τ, indicating a high degree of similarity. This is based on the original pseudo-labels generated through the argmax operation. Calculate high confidence regions The percentage of each category in the data, specifically the percentage of each category c∈[0,1], is calculated as follows:
[0037]
[0038] in, Indicates a pseudo tag. express Binary mappings of positions with high confidence. Specifically:
[0039]
[0040] 2.3) Select category c with the largest proportion. max Specifically:
[0041]
[0042] 2.4) Optimize pseudo-tags based on category proportions. If c max If a category's proportion in the entire image exceeds the threshold τ, it indicates that the category has high reliability and is suitable for expansion into pseudo-labels. The high-confidence regions are then updated, specifically:
[0043]
[0044] in, Indicates a pseudo tag. express A binary mapping of positions with high confidence.
[0045] 2.4) If c max If it is an unchanging class, it will be executed according to step 2.4); if c max If it is a variable class, then no longer
[0046] The modifications are as follows:
[0047]
[0048] Furthermore, the loss function of the semi-supervised remote sensing image change detection model for:
[0049]
[0050] in, To monitor losses, Loss due to lack of supervision;
[0051] The monitoring loss The definition is as follows:
[0052]
[0053] in, This indicates that the loss is optimized for supervised deep attention prediction. β1 and β2 represent the standard prediction loss, respectively, and their weighting coefficients.
[0054] Unsupervised loss The definition is as follows:
[0055]
[0056] in, Let these represent the unsupervised hard loss, the feature perturbation loss, and the KL divergence, respectively. This represents the loss for unsupervised deep attention prediction optimization. Let λ1, λ1, and λ1 represent the class-imbalanced pseudo-label optimization loss, where λ1, λ1, and λ1 correspond to their respective weight coefficients.
[0057] Furthermore, during the training process in step S3, a stochastic gradient descent optimizer is used to update the network parameters, with an initial learning rate of 0.005 and a weight decay of 1e-4.
[0058] Furthermore, the threshold τ is a fixed threshold of 0.85.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] The deep attention DA module designed in this invention uniquely combines pixel-level spatial structure information and similarity information. It can not only integrate global spatial context information, but also adaptively capture fine-grained similarity relationships between pixels in local regions, improving the detection capability of subtle changes, thereby enriching the model's prediction results and making the predictions more accurate and reliable. The class imbalance pseudo-label enhancement CIE module designed in this invention breaks through the limitations of traditional methods that rely solely on high-confidence regions. By directly guiding the pseudo-label generation process using encoded features and combining the prediction of high-confidence regions and change categories, it improves the quality of pseudo-labels, significantly improving the reliability of pseudo-labels while effectively alleviating the class imbalance problem.
[0061] This invention innovatively embeds spatial structure constraints into the consistency regularization process, effectively improving change detection performance. It can be widely applied to remote sensing change detection tasks, solving problems such as erroneous predictions and low quality of pseudo-labels in existing semi-supervised methods. It lays a good foundation for subsequent remote sensing change detection tasks and provides new research ideas and technical means for semi-supervised remote sensing change detection methods. Attached Figure Description
[0062] Figure 1 This is a flowchart of the method of the present invention.
[0063] Figure 2 This is a framework diagram of the DA-CIE proposed in this invention.
[0064] Figure 3 This is a structural diagram of the deep attention (DA) module designed in this invention.
[0065] Figure 4 This is a structural diagram of the class-unbalanced pseudo-label enhancement (CIE) module designed in this invention.
[0066] Figure 5 This is a visualization of the change detection results of the present invention on four datasets with 10% labeled data. Detailed Implementation
[0067] The present invention will be further described in detail below with reference to specific embodiments. These descriptions are for explanation purposes only and are not intended to limit the scope of the invention.
[0068] This invention proposes a semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement (DA-CIE). First, a deep attention (DA) approach is designed to extract key feature information from the spatial structure. Based on these key features, a correlation mapping between feature vectors is constructed, fully leveraging the similarity information between pixels to enrich the prediction. This overcomes the deficiency of most methods in neglecting relation modeling, significantly improving the accuracy of the prediction results. Furthermore, this invention designs a class-imbalanced pseudo-label enhancement strategy (CIE) that guides the pseudo-label generation process by incorporating encoded features, enhancing the semantic consistency of the prediction results and optimizing the quality of the pseudo-labels. This further improves the effective utilization of unlabeled data and the prediction accuracy of changed regions. The specific scheme is as follows:
[0069] like Figure 1 As shown, the image data is first preprocessed, then input into the DA-CIE model for training, and finally the test image is input into the trained DA-CIE to output the change detection results. The specific implementation steps are as follows:
[0070] (1) Preprocessing of bi-temporal remote sensing image pairs: First, the dataset was divided into training, validation, and test sets according to the number of images in the large public dataset, and the original bi-temporal remote sensing images were randomly cropped into 256×256 image pairs. Second, in order to ensure data diversity and prevent model overfitting, necessary data augmentation operations were performed to enhance the model's generalization ability, including random angle rotation, random flipping, and feature perturbation.
[0071] (2) Proposing and training the DA-CIE model: First, configure the parameters used in the network model training process; then, train the proposed DA-CIE model. During training, the dual-temporal images are input in parallel into the dual-encoder network for processing. The dual encoders extract features and then perform difference calculations to capture the changing features of the dual-temporal images. In the supervised training phase, the extracted changing features are then divided into two processing paths: one path generates prediction results through the standard decoder, and the other path performs feature correction through the DA designed in this invention to obtain more accurate prediction results. In the unsupervised training phase, the extracted intermediate features are divided into four branches: two strong enhancement branches, one weak enhancement branch, and one feature perturbation branch. These four branches are also divided into two processing paths: one path generates predictions through the standard decoder, and the other path corrects predictions through the DA. The prediction results obtained from the weak enhancement branch generate the original pseudo-labels through the argmax operation. Considering that the initial generation of pseudo-labels may be inaccurate due to factors such as label noise, this invention further introduces the CIE strategy to improve the pseudo-labels and combines it with the prediction of the change category to further optimize the quality of the pseudo-labels.
[0072] (3) Training: Given the network training parameters, the semi-supervised remote sensing image change detection model is trained using the training set and validation set until the network converges.
[0073] (4) Test reasoning: Load the optimal training weights for testing to obtain the final change detection results.
[0074] A semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement is described in detail below:
[0075] (1) Remote sensing dataset preprocessing and data augmentation: This invention uses four common remote sensing change detection datasets: LEVIR-CD, WHU-CD, SYSU-CD and CDD.
[0076] LEVIR-CD possesses a large ultra-high-resolution building change detection dataset, containing over 31,000 uniquely labeled objects showing changes. It comprises 637 pairs of high-resolution images taken between 2002 and 2018, each image being 1024×1024 pixels in size. To prevent overfitting, this invention employs data augmentation operations such as random cropping. In experiments, these image pairs were cropped into non-overlapping 256×256 pixel patches, and then divided into a training set of 12,000 pairs, a validation set of 4,000 pairs, and a test set of 4,000 pairs. The training set was further divided into 5%, 10%, 20%, and 40% labeled data for training.
[0077] WHU-CD is a publicly available building change detection dataset provided by Wuhan University. It contains a pair of high-resolution images (32507×15345 pixels) taken in the same area of Christchurch, New Zealand, between 2012 and 2016. Through data augmentation operations such as random cropping, the images were cropped into non-overlapping 256×256 pixel patches, and then divided into a training set of 5947 pairs, a validation set of 743 pairs, and a test set of 744 pairs. The training set was further divided into 5%, 10%, 20%, and 40% labeled data for training.
[0078] SYSU-CD is a multi-class binary change detection dataset collected in Hong Kong. It contains 20,000 image pairs, each pair being 256×256 pixels in size. In experiments, these image pairs were divided into a training set of 12,000 pairs, a validation set of 4,000 pairs, and a test set of 4,000 pairs. The training set was further divided into 5%, 10%, 20%, and 40% labeled data for training.
[0079] CDD is a seasonal change detection dataset collected by Google Earth. Proposed in 2018, it contains seven pairs of images, each 4725×2700 pixels in size. In experiments, these images were cropped into non-overlapping 256×256 pixel patches and then divided into a training set of 10,000 pairs, a validation set of 3,000 pairs, and a test set of 3,000 pairs. The training set was further divided into 5%, 10%, 20%, and 40% labeled data for training.
[0080] (2) Overall architecture of DA-CIE proposed in this invention: The DA-CIE framework proposed in this invention mainly includes two modules, such as... Figure 2 As shown, supervised training is performed on labeled images: labeled dual-temporal remote sensing images are input in parallel into a dual-encoder network. The dual encoders extract image features and perform difference calculations to capture the original difference features of the dual-temporal remote sensing images. Random dropout is performed on the extracted original difference features to obtain labeled feature enhancement features. The original difference features and labeled feature enhancement features are respectively used by the decoder to generate standard predictions. Feature correction is performed by a deep attention module to obtain supervised deep attention predictions.
[0081] Unsupervised training is performed on unlabeled images: Unlabeled dual-temporal remote sensing images are subjected to strong and weak enhancement processing to obtain strongly enhanced 1, strongly enhanced 2, and weakly enhanced unlabeled dual-temporal remote sensing images. These strongly enhanced 1, strongly enhanced 2, and weakly enhanced unlabeled dual-temporal remote sensing images are input in parallel into a dual encoder network. The extracted differential features are specifically strongly enhanced feature 1, strongly enhanced feature 2, and weakly enhanced feature. Then, the weakly enhanced feature is randomly dropped out to obtain the differential features of the unlabeled feature enhancement features. The strongly enhanced feature 1, strongly enhanced feature 2, weakly enhanced feature, and unlabeled feature enhancement features are then processed by the decoder to generate standard unsupervised predictions. The strongly enhanced feature 1, strongly enhanced feature 2, weakly enhanced feature, and unlabeled feature enhancement features are then processed by a deep attention module for feature correction to generate unsupervised deep attention predictions. The predictions obtained from the weakly enhanced features are then processed by argmax operation to generate original pseudo-labels. The original pseudo-labels are then optimized by a class-imbalanced pseudo-label enhancement module and further optimized by combining the prediction of the change category.
[0082] (3) Constructing a Deep Attention (DA) Module: Previous methods mostly relied on thresholding mechanisms to select reliable pixels. This approach not only ignores the spatial and semantic relationships between pixels, leading to a failure to fully utilize the contextual information of the image, but also severely limits the use of unlabeled data. Furthermore, in current SSCD research, few methods apply attention to the SSCD task, and even fewer combine channel, spatial, and self-attention. Therefore, this invention proposes a deep attention (DA) mechanism, which combines channel and spatial attention to focus on which channels and spatial locations are more important to the task, enhancing the model's understanding of the contextual relationships between image pixels. Further, a self-attention mechanism is introduced to construct a similarity matrix to capture global dependencies between features. For example... Figure 3 As shown in (a), firstly, by combining channel attention and spatial attention, we focus on which channels and spatial locations are more important to the task, and then introduce a self-attention mechanism to construct a similarity matrix; Figure 3 (b) is to put Figure 3 (a) The obtained similarity matrix is applied to the standard output. DA can enhance the model's learning of the contextual relationships between image pixels, thereby improving the utilization of unlabeled data.
[0083] (a) Channel Attention Calculation: The differential features obtained by the dual encoder contain contextual similarity relationships and can express the degree of similarity. Therefore, a simple linear mapping is performed to obtain... Where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map, respectively. For features... Perform channel attention (CA) calculations to enhance the model's focus on which channels are more important for change detection tasks:
[0084]
[0085] Where CA(·) represents channel attention, which is calculated through global average pooling and a fully connected layer, and then... Perform a weighted operation. This represents the matrix multiplication operation.
[0086] (b) Spatial attention computation: In Building upon this foundation, spatial attention (SA) is further calculated to focus on which spatial locations are more important to the task.
[0087]
[0088] Here, SA(·) represents spatial attention, which enhances the feature representation capability of important pixel locations. This represents the matrix multiplication operation.
[0089] (c) Self-attention calculation: For Perform a 1×1 convolution transformation to compute the spatial correlation matrix:
[0090]
[0091] The correlation matrix After Softmax(·) normalization, the similarity between each location in the feature map and other locations is represented. Conv2d(·) represents the convolution operation. This indicates a transpose operation, where C is the number of channels.
[0092] (d) Applying the correlation matrix to prediction: using the similarity matrix The corrected prediction is obtained by applying it to the logits output of the model decoder.
[0093]
[0094] in, This indicates an interpolation operation. This represents the logits output of the model decoder. This represents the similarity matrix.
[0095] (4) Construct the class-unbalanced pseudo-label enhancement strategy (CIE) module: Quantifying the similarity between different pixels or regions helps the model more accurately determine which pixels belong to the same category. This is especially important in CorelDRAW tasks, where pixel distribution in changing regions is typically sparse and significantly different from the background class, exhibiting a clear class imbalance problem. Therefore, this invention designs a Class Imbalanced Pseudo-Label Enhancement (CIE) strategy, such as... Figure 4 As shown, pseudo-label optimization is achieved by fusing the similarity matrix with high-confidence regions, which further alleviates the class imbalance problem.
[0096] (a) To enhance the difference between the changed area and the background, it is first necessary to extract change information from the similarity matrix. Each row is normalized, and a binary mask is generated through threshold binarization.
[0097]
[0098] in, Indicates an indicator function, This represents the interpolation operation, where z represents the similarity matrix. Each row vector in the array.
[0099] (b) In remote sensing change detection tasks, due to the relatively blurry edge detection and limited number of changed areas, models often tend to overlook changed areas, especially when they perform well in detecting invariant areas. To alleviate this problem, the CIE in this invention introduces a dynamic pseudo-label adjustment mechanism. This is mainly achieved by comparing binary masks. and high confidence areas If the overlap between the two predictions is greater than a threshold τ, it indicates that the two predictions are highly similar and the results are relatively reliable. Therefore, a binarized mask can be used. and high confidence areas Combined, the pseudo-tags can be further optimized. Specifically, the current pseudo-tags are known. Calculate high confidence regions The percentage of each category is calculated. The formula for calculating the percentage of category c∈[0,1] is as follows:
[0100]
[0101] in, Indicates a pseudo tag. express Binary mappings of positions with high confidence. The formula is as follows:
[0102]
[0103] (c) Select the category with the largest proportion. maxThe formula is as follows:
[0104]
[0105] (d) Optimize pseudo-tags based on category proportion. If c max If a category has a proportion greater than the threshold τ in the entire image, and the threshold τ is a fixed threshold of 0.85, it indicates that the category has high reliability and is suitable for expansion into pseudo-labels, and the high-confidence region is updated.
[0106]
[0107] in, Indicates a pseudo tag. express A binary mapping of positions with high confidence.
[0108] (e) In change detection tasks, most change maps have a small proportion of changed regions and a large proportion of unchanged regions. Therefore, in most cases, the class with the largest proportion is usually the unchanged class (i.e., class 0). To alleviate the class imbalance problem, if c max If it is an invariant class, this invention will incorporate variable classes; if c max If it's already a variable type, then no further modifications are needed. The formula is as follows:
[0109]
[0110] (5) Setting the loss function in this invention: The final loss of this invention The formula is as follows:
[0111]
[0112] in, To monitor losses, This is an unsupervised loss.
[0113] Monitoring losses The definition is as follows:
[0114]
[0115] in, This indicates that the loss is optimized for supervised deep attention prediction. This represents the standard prediction loss, with β1 and β2 set to 0.5 and 0.5 by default.
[0116] Unsupervised loss The definition is as follows:
[0117]
[0118] in, These represent the standard unsupervised hard loss, the characteristic perturbation loss, and the KL divergence, respectively. This represents the loss for unsupervised deep attention prediction optimization. This represents the class imbalance pseudo-label optimization loss, with λ1, λ1, and λ1 set to 0.5, 0.25, and 0.25 by default.
[0119] (6) Analysis of change detection results:
[0120] (a) Experimental setup and evaluation metrics:
[0121] To ensure a fair comparison, the DA-CIE proposed in this invention is compared with the SSCD performance of state-of-the-art (SOTA) methods. Hardware experimental platform: CPU: Intel Core i9-9900X 3.5GHz, GPU: NVIDIA GeForce RTX 3090Ti, VRAM: 24GB; Software experimental platform: PyTorch, Python, OpenCV, NumPy, and other open-source software and frameworks. DeepLabV3+ was used as the backbone of this invention, and a stochastic gradient descent (SGD) optimizer was used to update the network parameters with an initial learning rate of 0.005 and a weight decay of 1e-4. The model was trained for 80 epochs, with a batch size of 4 for both labeled and unlabeled data. In the semi-supervised training, the same data split was performed on each dataset, with 5%, 10%, 20%, and 40% of the RSI pairs in the training set randomly selected as labeled data, and the remainder as unlabeled data.
[0122] This invention calculates two widely used CD (Corrective Discrimination) evaluation metrics: Intersection over the Union (IoU) and Overall Accuracy (OA). The formulas for these two metrics are:
[0123]
[0124] In this IoU, TP represents the number of true positives, TN represents the number of true negatives, FP represents the number of false positives, and FN represents the number of false negatives. It's important to note that a higher IoU value indicates better detection performance.
[0125] (b) Analysis of experimental results:
[0126] To verify the superiority of the present invention, the DA-CIE of the present invention was compared with ten state-of-the-art methods on four datasets, including AdvEnt, SemiCDNet, S4GAN, SemiCD, RCL, Unimatch, FPA, ECPS, CutMix-CD, and AdaSemiCD.
[0127] The quantitative evaluation results of the LEVIR-CD dataset are shown in Table 1. The best values in all tables below are shown in bold red, and the second-best values are shown in bold blue. With 5%, 10%, 20%, and 40% labeled samples, the DA-CIE proposed in this invention outperforms the second-best method in IoU by 2.98%, 3.13%, 3.07%, and 2.33%, respectively. As the labeled proportion increases, the detection performance of each comparison method improves, but none surpasses the performance of this invention, which consistently ranks first. For ease of comparison, this invention is visualized against ten other comparison methods on the LEVIR-CD dataset (e.g., ...). Figure 5 (a) and (b), at a 10% marking ratio, where the red part represents the changed area and the black part represents the unchanged area. By observing the visualization results, it is demonstrated that the change prediction map obtained by the DA-CIE proposed in this invention is closer to the ground truth.
[0128] Table 1 shows the performance comparison results on the LEVIR-CD test set. Red indicates the best results, and blue indicates the second-best results.
[0129]
[0130]
[0131] The quantitative evaluation results of the WHU-CD dataset are shown in Table 2. Figure 5 Tables (c) and (d) visualize the experimental results on 10% labeled data, where red represents areas of change and black represents areas of no change. Observing the visualization, the change prediction map obtained by the proposed DA-CIE is more similar to the ground truth. Furthermore, it can be observed that DA-CIE outperforms other comparative methods in boundary detection. Table 2 summarizes the experimental results of all comparative methods and the proposed method on the WHU dataset. The proposed DA-CIE achieves best performance across most of the labeled data proportions.
[0132] Table 2 shows the performance comparison results on the WHU-CD test set. Red indicates the best results, and blue indicates the second-best results.
[0133]
[0134] The quantitative evaluation results of the SYSU-CD dataset are shown in Table 3. Figure 5Tables (e) and (f) visualize the experimental results with 10% labeled data, where red represents areas of change and black represents areas of no change. By observing the visualization, the change detection map obtained by the DA-CIE method of this invention is closer to the ground truth than other methods. Table 3 summarizes the experimental results of all comparative methods and DA-CIE on the SYSU dataset. With 5%, 10%, 20%, and 40% labeled samples, the DA-CIE proposed in this invention outperforms the second-best method in IoU by 2.64%, 2.42%, 2.74%, and 2.65%, respectively. Overall, DA-CIE achieves the best performance on the SYSU dataset.
[0135] Table 3 shows the performance comparison results on the SYSU-CD test set. Red indicates the best results, and blue indicates the second-best results.
[0136]
[0137]
[0138] The quantitative evaluation results of the CDD dataset are shown in Table 4. Figure 5 (g) and (h) illustrate the visualization of experimental results with 10% labeled data. Red areas represent areas of change, and black areas represent areas of no change. Observing the visualization, the change detection map obtained by the proposed DA-CIE preserves edge information more completely compared to other methods. Table 4 summarizes the experimental results of all comparative methods and the proposed DA-CIE on the CDD dataset. With 5%, 10%, and 20% labeled samples, the proposed DA-CIE outperforms the second-best method in IoU by 4.23%, 1.75%, and 0.67%, respectively. Overall, DA-CIE achieves relatively good results on the CDD dataset.
[0139] Table 4 shows the performance comparison results on the CDD test set. Red indicates the best results, and blue indicates the second-best results.
[0140]
[0141] To better validate the effects of the proposed DA mechanism and CIE strategy on the CD task, ablation experiments (as shown in Tables 5 and 6) were designed to demonstrate their synergistic effect and impact on model performance. As shown in Table 5, when the label ratio is 5%, the DA mechanism, compared to the base, improves performance by 1.9%, 1.98%, and 2.87% on the SYSU, LEVIR-CD, and CDD datasets, respectively; the CIE strategy, compared to the base, improves performance by 2.81% and 3.69% on the LEVIR-CD and CDD datasets, respectively. Compared to the baseline, the complete DA-CIE improves performance by 1.43%, 2.64%, 5.43%, and 7.9% on 5% labeled data across the four datasets; and by 3.04%, 4.24%, 4.52%, and 7.14% on 40% labeled data across the four datasets.
[0142] Table 5 shows the ablation experiments of the DA and CIE modules (5% labeled data from four datasets). The best results are indicated in bold black.
[0143]
[0144] Table 6 shows the ablation experiments of the DA and CIE modules (40% labeled data from four datasets). The best results are indicated in bold black.
[0145]
[0146] Tables 7 and 8 show the ablation experiments for each attention in the DA module. It can be observed that when the labeling ratio is 5%, Base is combined with three different attention types and validated on four datasets (as shown in Table 7). Finally, when combined with the CBAM hybrid attention, the metrics on the four datasets are improved by 0.37%, 0.58%, 0.87%, and 2.02%, respectively. At a labeling ratio of 40%, the best performance is also obtained by combining the CBAM hybrid attention (as shown in Table 8). Since the generated correlation matrix is usually large (e.g., ...), ... Figure 3As shown in (a), subsequent calculations become quite complex. More importantly, adjacent regions often have similar semantic information; therefore, using every row of the similarity matrix for pseudo-label optimization leads to redundant computation and unnecessary resource waste. Furthermore, uniform row selection may cause some regions to be oversampled while others are ignored, leading to overfitting in certain areas and neglecting more important ones. In contrast, random row selection introduces greater sample diversity, allowing the model to encounter more varied classes during training, thus avoiding bias towards learning the larger proportion of background classes. As shown in Table 9, experiments were conducted on 40% labeled data in the CDD dataset. The results show that randomly selecting 128 rows yields the highest performance; therefore, this invention chooses to randomly select 128 rows to balance computation and performance.
[0147] Table 7 shows the ablation experiments in the DA module (5% labeled data from four datasets). The best results are indicated in bold black.
[0148]
[0149]
[0150] Table 8 shows the ablation experiments in the DA module (40% labeled data from four datasets). The best results are indicated in bold black.
[0151]
[0152] Table 9 shows the ablation experiments of four randomly selected rows (CDD 40% labeled data). The best results are indicated in bold black.
[0153] 32 64 128 256 IoU 82.17 84.85 86.41 85.45 OA 97.79 98.13 98.69 98.61
Claims
1. A semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement, characterized in that, Includes the following steps: S1. The remote sensing image dataset is preprocessed and divided into a training set, a validation set, and a test set. The training set includes labeled images and unlabeled images. S2, Construct a semi-supervised remote sensing image change detection model. The semi-supervised remote sensing image change detection model includes a deep attention module and a class-imbalanced pseudo-label enhancement module. The deep attention module enhances the model's learning of the contextual relationship between image pixels by combining channel attention and spatial attention. At the same time, it introduces a self-attention mechanism to construct a similarity matrix to capture the global dependency relationship between features. The class-imbalanced pseudo-label enhancement module optimizes pseudo-labels by fusing the similarity matrix with high-confidence regions. S3, Given network training parameters, train the semi-supervised remote sensing image change detection model using the training set and validation set until the network converges; S4. Input the test set into the semi-supervised remote sensing image change detection model trained in step S3, and output the remote sensing image change detection results.
2. The semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement according to claim 1, characterized in that, The preprocessing of the remote sensing image dataset in step S1 specifically involves: randomly cropping the dual-temporal remote sensing images into pairs of images of the same size, and performing data augmentation operations such as random angle rotation, random flipping, and feature perturbation.
3. The semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement according to claim 1, characterized in that, In step S2, the semi-supervised remote sensing image change detection model uses DeepLabV3+ as the backbone network.
4. The semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement according to claim 1, characterized in that, The training process in step S3 is specifically as follows: Supervised training of labeled images: Labeled dual-temporal remote sensing images are input in parallel into a dual encoder network. The dual encoders extract image features separately and then perform difference calculation to capture the original difference features of the dual-temporal remote sensing images. Random dropout is performed on the extracted original difference features to obtain labeled feature enhancement features. The original difference features and labeled feature enhancement features are respectively used by the decoder to generate standard predictions. The features are then corrected by the deep attention module to obtain supervised deep attention predictions. Unsupervised training on unlabeled images: Unlabeled dual-temporal remote sensing images are subjected to strong and weak enhancement processing to obtain strongly enhanced 1, strongly enhanced 2, and weakly enhanced unlabeled dual-temporal remote sensing images. These strongly enhanced 1, strongly enhanced 2, and weakly enhanced unlabeled dual-temporal remote sensing images are then input in parallel into a dual encoder network. The extracted differential features are specifically strongly enhanced feature 1, strongly enhanced feature 2, and weakly enhanced feature. Then, the weakly enhanced feature is randomly dropped out to obtain the differential features of the unlabeled feature enhancement features. The strongly enhanced feature 1, strongly enhanced feature 2, weakly enhanced feature, and unlabeled feature enhancement features are then processed by a decoder to generate standard unsupervised predictions. Finally, the strongly enhanced feature 1, strongly enhanced feature 2, weakly enhanced feature, and unlabeled feature enhancement features are processed by a deep attention module for feature correction to generate unsupervised deep attention predictions. The predictions obtained from the weakly enhanced features are used to generate the original pseudo-labels through the argmax operation. The original pseudo-labels are then optimized by the class imbalance pseudo-label enhancement module, and further optimized by combining the predictions of the changing categories.
5. The semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement according to claim 4, characterized in that, The feature correction via the deep attention module specifically involves: 1.1) Channel attention calculation: The differential features are linearly mapped to obtain the mapped differential features. Where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map, respectively, for the mapped differential features. Perform channel attention calculation: Where CA(·) represents channel attention, which is calculated through global average pooling and a fully connected layer, and then... Perform a weighted operation. This represents the matrix multiplication operation; 1.2) Based on the channel attention calculation results Perform spatial attention calculations: Where SA(·) represents spatial attention, This represents the matrix multiplication operation; 1.3) Results of spatial attention calculation Perform a 1×1 convolution transformation: Among them, the similarity matrix After Softmax(·) normalization, the similarity between each location in the feature map and other locations is represented. Conv2d(·) represents the convolution operation. This indicates a transpose operation, where C is the number of channels; 1.4) The similarity matrix The deep attention prediction is obtained by applying it to the logits output of the model decoder. in, This indicates an interpolation operation. This represents the logits output of the model decoder.
6. The semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement according to claim 5, characterized in that, The pseudo-label optimization achieved through the class-unbalanced pseudo-label enhancement module specifically involves: 2.1) For similarity matrices The row vectors of each row are normalized, and a binary mask is generated by thresholding and binarization. in, Indicates an indicator function, The expression represents the interpolation operation; z represents the similarity matrix. Each row vector in the array; 2.2) By comparing the binarization masks and high confidence areas The overlap between the two predictions is considered if it exceeds a threshold τ, indicating a high degree of similarity. This is based on the original pseudo-labels generated through the argmax operation. Calculate high confidence regions The percentage of each category in the data, specifically the percentage of each category c∈[0,1], is calculated as follows: in, Indicates a pseudo tag. express Binary mappings of positions with high confidence. Specifically: 2.3) Select category c with the largest proportion. max Specifically: 2.4) Optimize pseudo-tags based on category proportions. If c max If a category's proportion in the entire image exceeds the threshold τ, it indicates that the category has high reliability and is suitable for expansion into pseudo-labels. The high-confidence regions are then updated, specifically: in, Indicates a pseudo tag. express A binary mapping of positions with high confidence. 2.4) If c max If it is an unchanging class, it will be executed according to step 2.4); if c max If it is a variable class, then no longer The modifications are as follows:
7. The semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement according to claim 1, characterized in that, The loss function of the semi-supervised remote sensing image change detection model for: in, To monitor losses, Loss due to lack of supervision; The monitoring loss The definition is as follows: in, This indicates that the loss is optimized for supervised deep attention prediction. β1 and β2 represent the standard prediction loss, respectively, and their weighting coefficients. Unsupervised loss The definition is as follows: in, These represent the standard unsupervised hard loss, the characteristic perturbation loss, and the KL divergence, respectively. This represents the loss for unsupervised deep attention prediction optimization. Let λ1, λ1, and λ1 represent the class-imbalanced pseudo-label optimization loss, where λ1, λ1, and λ1 correspond to their respective weight coefficients.
8. The semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement according to claim 1, characterized in that, During the training process in step S3, a stochastic gradient descent optimizer is used to update the network parameters, with an initial learning rate of 0.005 and a weight decay of 1e-4.
9. A semi-supervised remote sensing image change detection method based on deep attention and class-imbalanced pseudo-label enhancement according to claim 6, characterized in that, The threshold τ is a fixed threshold of 0.85.