Semi-supervised Remote Sensing Image Change Detection Method and Device Based on Federated Learning

Through a semi-supervised method based on joint learning, using a small number of single-time phase remote sensing images to generate pseudo-labels and combined with consistency regularization training, the problem of lack of label data in remote sensing image change detection is solved, the detection accuracy and robustness are improved, and the cost of manual labeling is reduced.

CN117173561BActive Publication Date: 2025-07-18NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311053412.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-21
Publication Date
2025-07-18
Estimated Expiration
2043-08-21

AI Technical Summary

Technical Problem

The existing remote sensing image change detection technology requires a large amount of label data for training, resulting in high cost of manual labeling and low detection accuracy of unsupervised learning methods, which cannot meet the actual application needs.

Method used

A semi-supervised method based on joint learning is adopted to generate pseudo-dual-time phase remote sensing images and pseudo-change detection tags using a small number of single-time phase remote sensing images and their building extraction labels. The model is trained through consistency regularization, combining building extraction tasks and change detection tasks to improve the robustness and generalization of the model.

Benefits of technology

Under the condition of a small amount of label data, the accuracy and robustness of remote sensing image change detection are improved, the cost of manual labeling is reduced, and efficient change detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173561B_ABST
    Figure CN117173561B_ABST
Patent Text Reader

Abstract

This application relates to a semi-supervised remote sensing image change detection method and device based on collaborative learning. The method includes: generating a large number of pseudo double-temporal remote sensing images and pseudo change detection labels from a small amount of single-temporal remote sensing image data and their building extraction labels to provide label data support for model training; training the building extraction task and the change detection task together to further enhance the change detection model's understanding of remote sensing image information and improve the change detection performance. At the same time, a semi-supervised method based on the principle of consistency regularization is used to force the model to produce consistent change detection results for unlabeled remote sensing images before and after data perturbation, improving the robustness and generalization of the model, effectively improving the change detection accuracy of remote sensing images, and having a wide range of application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of remote sensing image change detection, and particularly to a semi-supervised remote sensing image change detection method and device based on joint learning. Background Art

[0002] The task of remote sensing image change detection (CD) aims to identify the changes that occur in remote sensing images of the same area acquired at different times. Here, the changes generally refer to semantic changes, such as the changes in buildings. For many years, remote sensing image change detection has always been one of the research hotspots in the remote sensing field. With the continuous growth of the number of remote sensing images and the development of learning technology, the use of remote sensing image change detection technology can quickly obtain the change information of the areas we are concerned about, including the changes in natural features and artificial buildings, providing strong support for the decision-making of governments, companies, and organizations. So far, remote sensing image change detection technology has been widely applied in the fields of ecosystem monitoring, land resource and land use mapping, damage assessment, urban expansion monitoring, etc.

[0003] Currently, the mainstream CD algorithms are based on fully supervised deep learning methods, mostly based on convolutional neural networks (CNNs). However, using remote sensing image change algorithms based on fully supervised deep learning requires a large number of labeled dual-temporal remote sensing image pairs for neural network training, and annotating these labels requires a large amount of manual and time costs. If an unsupervised learning method is used to train the model, due to the lack of guidance from labeled data, the change detection accuracy of the model is often low and insufficient to support actual change detection applications.

[0004] To solve this problem, in recent years, change detection (CD) methods based on semi-supervised learning (SSL) have received increasing attention. Semi-supervised CD methods aim to train a model using a small amount of labeled data and a large amount of unlabeled data to achieve an effect close to that of a fully supervised algorithm trained with a large amount of labeled data. Training the model with a small amount of labeled data can guide the model to understand the task, while the large amount of unlabeled data can effectively prevent the model from overfitting on the small amount of labeled data, thereby improving the robustness and generalization of the model. In recent years, semi-supervised learning algorithms based on consistency regularization have developed rapidly. This method relies on making the network make consistent change detection predictions for remotely sensed image pairs before and after significant distortion, thereby using a large amount of unlabeled remotely sensed image pairs to perform self-supervised training on the model. Coupled with the guidance of a small amount of remotely sensed images with labels, the purpose of achieving high-precision change detection with a small number of labeled tags is achieved. There are mainly two key problems in this type of method. The first is how to make good use of the large amount of unlabeled data and introduce it into network training. The second is how to make good use of the information contained in the small amount of labeled data so that it can play the greatest role. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a semi-supervised remote sensing image change detection method and device based on joint learning that can improve the CD accuracy by using a large amount of unlabeled data when only a small number of single-temporal remote sensing images and their building extraction labels (such as ten) are available.

[0006] A semi-supervised remote sensing image change detection method based on joint learning, the method comprising:

[0007] Obtain single-temporal remote sensing images and corresponding building extraction labels, as well as unlabeled double-temporal remote sensing image pairs;

[0008] Preprocess the single-temporal remote sensing images and corresponding building extraction labels to obtain pseudo double-temporal remote sensing image pairs and corresponding building extraction labels, perform an exclusive OR operation on the building extraction labels of the pseudo double-temporal remote sensing image pairs to obtain pseudo change detection labels;

[0009] Input the pseudo double-temporal remote sensing image pairs and corresponding building extraction labels, as well as the pseudo change detection labels into a pre-constructed change detection model, process the pseudo double-temporal remote sensing image pairs according to the feature extraction network, building extraction decoder, and change detection decoder in the change detection model, output a supervised building extraction result and a supervised change detection prediction result, calculate a segmentation loss based on the supervised building extraction result and the building extraction labels of the pseudo double-temporal remote sensing image pairs, and calculate a change detection loss based on the supervised change detection prediction result and the pseudo change detection labels;

[0010] The unlabeled dual-temporal remote sensing image pairs are successively subjected to flipping and translation and data perturbation to obtain the flipped and translated image pairs and the distorted image pairs. The flipped and translated image pairs and the distorted image pairs are input into the change detection model for processing, and the building extraction results of the flipped and translated image pairs, the building extraction results of the distorted image pairs, the change detection prediction results of the flipped and translated image pairs, and the change detection prediction results of the distorted image pairs are output. According to the building extraction results of the flipped and translated image pairs and the building extraction results of the distorted image pairs, a consistency segmentation loss is calculated. According to the change detection prediction results of the flipped and translated image pairs and the change detection prediction results of the distorted image pairs, a consistency change detection loss is calculated;

[0011] The segmentation loss, the change detection loss, the consistency segmentation loss, and the consistency change detection loss are added together to obtain the total loss. Taking the minimum total loss as the objective function, the parameters of the change detection model are trained and optimized until a trained change detection model is obtained;

[0012] The remote sensing image pairs to be detected are input into the trained change detection model for change detection, and the change detection results are output.

[0013] In one embodiment, the single-temporal remote sensing image and the corresponding building extraction label are preprocessed to obtain a pseudo-dual-temporal remote sensing image pair and the corresponding building extraction label. The exclusive OR operation is performed on the building extraction labels of the pseudo-dual-temporal remote sensing image pair, and a pseudo-change detection label is obtained, including:

[0014] The single-temporal remote sensing image X l and the corresponding building extraction label Y be are subjected to the first random sampling and flipping and translation to obtain the single-temporal remote sensing image X l1 after the first preprocessing and the corresponding building extraction label Y be_1 ;

[0015] X l and Y be are subjected to the second random sampling and flipping and translation to obtain the single-temporal remote sensing image X l2 after the second preprocessing and the corresponding building extraction label Y be_2 ;

[0016] X l1 and X l2 are paired to obtain a pseudo-dual-temporal remote sensing image pair, Y be_1 and Y be_2 are paired to obtain the building extraction label of the pseudo-dual-temporal remote sensing image pair, and the paired Y be_1 and Y be_2 are subjected to the exclusive OR operation to obtain the pseudo-change detection label Ycd 。

[0017] In one embodiment, the pseudo-bitemporal remote sensing image pair, the corresponding building extraction labels, and the pseudo-change detection labels are input into a pre-constructed change detection model. The pseudo-bitemporal remote sensing image pair is processed according to the feature extraction network, the building extraction decoder, and the change detection decoder in the change detection model, and a supervised building extraction result and a supervised change detection prediction result are output, including:

[0018] Input X in the pseudo-bitemporal remote sensing image pair l1 and X l2 , Y in the building extraction labels of the pseudo-bitemporal remote sensing image pair be_1 and Y be_2 and the pseudo-change detection label Y cd into the change detection model;

[0019] According to the first feature extraction network and the second feature extraction network with shared weights in the change detection model, feature extraction is respectively performed on X l1 and X l2 to obtain a first type of supervised feature extraction result and a second type of supervised feature extraction result;

[0020] The first type of supervised feature extraction result and the second type of supervised feature extraction result are input into the building extraction decoder for processing to obtain a first type of supervised building extraction result and a second type of supervised building extraction result

[0021] The first type of supervised feature extraction result and the second type of supervised feature extraction result are concatenated, and the concatenated supervised feature extraction result is input into the change detection decoder for processing to obtain a supervised change detection prediction result

[0022] In one embodiment, according to the supervised building extraction result and the building extraction labels of the pseudo-bitemporal remote sensing image pair, a segmentation loss is calculated, and according to the supervised change detection prediction result and the pseudo-change detection labels, a change detection loss is calculated, including:

[0023] According to the first type of supervised building extraction result and Y in the building extraction labels of the pseudo-bitemporal remote sensing image pair be_1 a first type of segmentation loss L seg1 is calculated, and according to the second type of supervised building extraction result and Y in the building extraction labels of the pseudo-bitemporal remote sensing image pair be_2 a second type of segmentation loss L seg2;

[0024] According to the prediction result of supervised change detection and the pseudo-change detection label Y cd perform calculations to obtain the change detection loss L cd .

[0025] In one embodiment, the unlabeled dual-temporal remote sensing image pair is sequentially subjected to flipping and translation and data perturbation to obtain a flipped and translated image pair and a distorted image pair, including:

[0026] For the unlabeled dual-temporal remote sensing image pair X u in the first unlabeled dual-temporal remote sensing image X u1 and the second unlabeled dual-temporal remote sensing image X u2 perform the same random flipping and translation to obtain the first flipped and translated image X u1_aug and the second flipped and translated image X u2_aug , X u1_aug and X u2_aug constitute a flipped and translated image pair;

[0027] Perform the same random data perturbation on X u1_aug and X u2_aug to obtain the first distorted image X u1_per and the second distorted image X u2_per , X u1_per and X u2_per constitute a distorted image pair.

[0028] In one embodiment, input the flipped and translated image pair and the distorted image pair into the change detection model for processing, and output the building extraction results of the flipped and translated image pair, the building extraction results of the distorted image pair, the change detection prediction results of the flipped and translated image pair, and the change detection prediction results of the distorted image pair, including:

[0029] Input X u1_aug and X u2_aug in the flipped and translated image pair into the first feature extraction network and the second feature extraction network with shared weights in the change detection model for feature extraction respectively, to obtain the first flipped and translated image feature extraction result and the second flipped and translated image feature extraction result. Input X u1_per and X u2_per in the distorted image pair into the first feature extraction network and the second feature extraction network for feature extraction respectively, to obtain the first distorted image feature extraction result and the second distorted image feature extraction result;

[0030] Input the first flipped and translated image feature extraction result, the second flipped and translated image feature extraction result, the first distorted image feature extraction result, and the second distorted image feature extraction result into the building extraction decoder in the change detection model for processing to obtain the building extraction result of the first flipped and translated image The building extraction result of the second flipped and translated image The building extraction result of the first distorted image And the building extraction result of the second distorted image

[0031] Concatenate the first flipped and translated image feature extraction result and the second flipped and translated image feature extraction result to obtain the flipped and translated image pair feature extraction result. Concatenate the first distorted image feature extraction result and the second distorted image feature extraction result to obtain the distorted image pair feature extraction result;

[0032] Input the flipped and translated image pair feature extraction result and the distorted image pair feature extraction result into the change detection decoder in the change detection model for processing to obtain the change detection prediction result of the flipped and translated image pair And the change detection prediction result of the distorted image pair

[0033] In one embodiment, calculate based on the building extraction result of the flipped and translated image pair and the building extraction result of the distorted image pair to obtain the consistency segmentation loss. Calculate based on the change detection prediction result of the flipped and translated image pair and the change detection prediction result of the distorted image pair to obtain the consistency change detection loss, including:

[0034] For the building extraction result of the first flipped and translated image The building extraction result of the second flipped and translated image And the change detection prediction result of the flipped and translated image pair Perform data perturbation and threshold filtering in sequence to obtain the first building extraction pseudo-label The second building extraction pseudo-label And the change detection pseudo-label

[0035] According to And the building extraction result of the first distorted image Perform calculation to obtain the first type of consistency segmentation loss L cov_seg1 According to And the building extraction result of the second distorted image Perform calculation to obtain the second type of consistency segmentation loss L cov_seg2 According to And the change detection prediction result of the distorted image pair Perform calculations to obtain the consistency change detection loss L cov_cd .

[0036] In one embodiment, add the segmentation loss, change detection loss, consistency segmentation loss, and consistency change detection loss to obtain the total loss, including:

[0037] Add the first type of segmentation loss L seg1 , the second type of segmentation loss L seg2 , the change detection loss L cd , the first type of consistency segmentation loss L cov_seg1 , the second type of consistency segmentation loss L cov_seg2 , and the consistency change detection loss L cov_cd to obtain the total loss, expressed as

[0038] L = L seg1 + L seg2 + L cd + L cov_seg1 + L cov_seg2 + L cov_cd .

[0039] In one embodiment, the data perturbation includes color perturbation and / or shape perturbation with a preset intensity and preset combination for the image, building extraction result, and change detection prediction result.

[0040] A semi-supervised remote sensing image change detection device based on joint learning, the device includes:

[0041] A data acquisition module for acquiring a single-temporal remote sensing image and the corresponding building extraction label, as well as an unlabeled dual-temporal remote sensing image pair;

[0042] A data preprocessing module for preprocessing the single-temporal remote sensing image and the corresponding building extraction label to obtain a pseudo-dual-temporal remote sensing image pair and the corresponding building extraction label, performing an exclusive OR operation on the building extraction labels of the pseudo-dual-temporal remote sensing image pair to obtain a pseudo-change detection label;

[0043] A supervised learning module for inputting the pseudo-dual-temporal remote sensing image pair and the corresponding building extraction label, as well as the pseudo-change detection label into a pre-constructed change detection model, processing the pseudo-dual-temporal remote sensing image pair according to the feature extraction network, building extraction decoder, and change detection decoder in the change detection model, outputting a supervised building extraction result and a supervised change detection prediction result, calculating according to the supervised building extraction result and the building extraction label of the pseudo-dual-temporal remote sensing image pair to obtain a segmentation loss, and calculating according to the supervised change detection prediction result and the pseudo-change detection label to obtain a change detection loss;

[0044] An unsupervised learning module for sequentially performing flipping and translation and data perturbation on unlabeled dual-temporal remote sensing image pairs to obtain flipped and translated image pairs and distorted image pairs, inputting the flipped and translated image pairs and the distorted image pairs into a change detection model for processing, and outputting the building extraction results of the flipped and translated image pairs, the building extraction results of the distorted image pairs, the change detection prediction results of the flipped and translated image pairs, and the change detection prediction results of the distorted image pairs. Calculate based on the building extraction results of the flipped and translated image pairs and the building extraction results of the distorted image pairs to obtain a consistency segmentation loss, and calculate based on the change detection prediction results of the flipped and translated image pairs and the change detection prediction results of the distorted image pairs to obtain a consistency change detection loss;

[0045] A model training module for adding the segmentation loss, the change detection loss, the consistency segmentation loss, and the consistency change detection loss to obtain a total loss, using the minimum total loss as the objective function to train and optimize the parameters of the change detection model until a trained change detection model is obtained;

[0046] A model testing module for inputting the remote sensing image pairs to be detected into the trained change detection model for change detection and outputting the change detection results.

[0047] The above semi-supervised remote sensing image change detection method and device based on joint learning specifically solve the problems of lack of labels for remote sensing image change detection and the need to consume a large amount of labor and time costs for labeling. The present invention introduces a single-temporal joint learning technique in the problem of remote sensing image change detection. On the one hand, through a small amount of single-temporal remote sensing image data and their building extraction labels, a large number of pseudo-dual-temporal remote sensing images and pseudo-change detection labels are generated to provide label data support for model training. On the other hand, by training the building extraction task and the change detection task together, the change detection model's understanding of remote sensing image information is further strengthened, and the change detection performance is improved. At the same time, a semi-supervised method based on the principle of consistency regularization is used to force the model to make consistent change detection results for unlabeled remote sensing images before and after data perturbation, improving the robustness and generalization of the model, thereby improving the change detection accuracy and having good application prospects. Description of the Drawings

[0048] Figure 1 It is a schematic flowchart of a semi-supervised remote sensing image change detection method based on joint learning in an embodiment;

[0049] Figure 2 It is a schematic flowchart of the generation process of pseudo-dual-temporal image pairs and pseudo-change detection labels in an embodiment;

[0050] Figure 3Schematic diagram of the network structure of the building extraction decoder in an embodiment;

[0051] Figure 4 Schematic diagram of the network structure of the change detection decoder in an embodiment;

[0052] Figure 5 Visual display diagram of the data perturbation method in an embodiment;

[0053] Figure 6 Visual display diagram of the data perturbation method randomly combined in an embodiment;

[0054] Figure 7 Visual display diagram of threshold filtering in an embodiment;

[0055] Figure 8 Schematic diagram of the test results of the change detection model in an embodiment. Detailed implementation manners

[0056] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0057] In one embodiment, as Figure 1 shown, a semi-supervised remote sensing image change detection method based on joint learning is provided, including a change detection model training stage and a test prediction stage. The change detection model is composed of two feature extraction networks with shared parameters, a building extraction decoder and a change detection decoder. The specific structures of the building extraction decoder and the change detection decoder are as Figure 3 and Figure 4 shown. The training of the change detection model can be divided into a single-temporal supervised learning part and an unsupervised learning part.

[0058] (1) Change detection model training process

[0059] Input: Single-temporal remote sensing image X with labels l ∈R 256×256×3 , and its building extraction label Y be ∈{0,1} 256×256 , and unlabeled double-temporal remote sensing image pair X u (including X u1 , X u2 ∈R 256×256×3 ). Herein, R represents the dimension.

[0060] Output: Trained change detection model M.

[0061] Its specific training steps include the following:

[0062] 1. First is data preprocessing, that is, generating pseudo dual-temporal image pairs and pseudo change detection labels. As Figure 2 shown, specifically, the single-temporal image X l and its building extraction label Y be are subjected to the first random sampling, flipping, and translation to obtain the preprocessed single-temporal remote sensing image X l1 and the corresponding building extraction label Y be_1 ; then, the first different random sampling, flipping, and translation are performed to obtain the second preprocessed single-temporal remote sensing image X l2 and the corresponding building extraction label Y be_2 ; afterwards, X l1 and X l2 are paired one by one to obtain pseudo dual-temporal remote sensing image pairs for training the change detection network. Y be_1 and Y be_2 are paired one by one to obtain the building extraction labels of the pseudo dual-temporal remote sensing image pairs. The paired Y be_1 and Y be_2 are subjected to an exclusive OR operation (i.e., "Xor" in Figure 2 ) to obtain the pseudo change detection label Y cd .

[0063] The above data preprocessing steps are executed once each time the single-temporal data is input into the model for training. It is a key means to obtain a large number of pseudo dual-temporal remote sensing images from a small number of single-temporal remote sensing images. Through random and repeatable sampling, flipping, and translation, different positions between the same or different images can be superimposed to compare semantic differences, which is equivalent to being able to utilize the semantic differences between any positions of any image, thereby providing pseudo dual-temporal remote sensing images with a quantity far more than the input single-temporal remote sensing images for the training of the change detection model.

[0064] 2. Single-temporal supervised learning part. In this part, first, X l1 , X l2 , Y be_1 , Y be_2 , and Y cd are input into the change detection model. According to the first feature extraction network and the second feature extraction network with shared weights in the change detection model, feature extraction is performed on X l1 and X l2 respectively to obtain the first type of supervised feature extraction result and the second type of supervised feature extraction result. Then it is divided into two parts, namely the building extraction part and the change detection part.

[0065] In the building extraction part, the first type of supervised feature extraction result and the second type of supervised feature extraction result are input into the building extraction decoder for processing, and the first type of supervised building extraction result is obtained and the second type of supervised building extraction result According to the first type of supervised building extraction result and the Y in the building extraction label of the pseudo double-temporal remote sensing image pair be_1 a calculation is performed to obtain the first type of segmentation loss L seg1 , and according to the second type of supervised building extraction result and the Y in the building extraction label of the pseudo double-temporal remote sensing image pair be_2 a calculation is performed to obtain the second type of segmentation loss L seg2 ;

[0066] In the change detection part, the first type of supervised feature extraction result and the second type of supervised feature extraction result are concatenated (i.e., "Cat" in Figure 1 ), and the concatenated supervised feature extraction result is input into the change detection decoder for processing to obtain the supervised change detection prediction result According to the supervised change detection prediction result and the pseudo change detection label Y cd a calculation is performed to obtain the change detection loss L cd .

[0067] 3. Unsupervised learning part. In this part, first, the first unlabeled double-temporal remote sensing image pair X u in the unlabeled double-temporal remote sensing image pair X u1 and the second unlabeled double-temporal remote sensing image X u2 are subjected to the same random flipping and translation to obtain the first flipped and translated image X u1_aug and the second flipped and translated image X u2_aug , X u1_aug and X u2_aug constitute a flipped and translated image pair; then, X u1_aug and X u2_aug are subjected to the same random data perturbation to obtain the first distorted image X u1_per and the second distorted image X u2_per , X u1_per and X u2_per constitute a distorted image pair.

[0068] Among them, the role of random flipping and translation is to enhance the diversity of data and prevent model overfitting, while the role of random data perturbation is to cause huge distortions in the images for consistency regularization training. Specifically, the data perturbation includes color perturbation and / or shape perturbation with a preset intensity and a preset combination for the images, the building extraction results, and the change detection prediction results. The use of data perturbation mainly has two purposes. One is to increase or decrease the color difference between the features of the bi-temporal images through color perturbation. The second is to deform the ground objects through shape perturbation. Consistency regularization requires the model to ignore these changes and focus on the changes of the objects, thereby improving the semantic understanding ability and robustness of the model.

[0069] Figure 5 Lists the visual display diagrams of the data perturbation methods used in the change detection model in this application. Each time it is used, this application randomly selects three methods (random intensity) except for cropping, and then combines them with the cropping method to form a strong enhancement method. Figure 6 Lists the visual display diagrams of some randomly combined data perturbation methods. Table 1 lists the random intensity ranges of the data perturbation methods used in this application for reference.

[0070] Table 1 Data perturbation methods and their intensity ranges and descriptions

[0071] Enhancement method Strength range Description Brightness [0.05,0.95] Change the brightness of the image Color [0.05,0.95] Change the color balance of the image Contrast [0.05,0.95] Change the contrast of the image Equalization / Perform histogram equalization on the image Posterization [4,8] Reduce the number of bits per color channel Rotation [-30,30] Rotate the image Sharpening [0.05,0.95] Adjust the sharpness of the image Horizontal shear [-0.3,0.3] Shear the image along the horizontal axis Vertical shear [-0.3,0.3] Shear the image along the vertical axis Exposure [0,256] Invert all pixel values above the threshold Horizontal translation [-0.3,0.3] Translate in the horizontal direction Vertical translation [-0.3,0.3] Translate in the vertical direction Cropping [0.25,0.35] Crop a region from the image

[0072] Next, the X in the flipped and translated image pair u1_aug and X u2_aug are respectively input into the first feature extraction network and the second feature extraction network with shared weights in the change detection model for feature extraction, obtaining the first flipped and translated image feature extraction result and the second flipped and translated image feature extraction result. The X in the distorted image pair u1_per and X u2_per are respectively input into the first feature extraction network and the second feature extraction network for feature extraction, obtaining the first distorted image feature extraction result and the second distorted image feature extraction result.

[0073] The unsupervised learning part is also divided into a building extraction part and a change detection part. In the building extraction part, the first flipped and translated image feature extraction result, the second flipped and translated image feature extraction result, the first distorted image feature extraction result, and the second distorted image feature extraction result are input into the building extraction decoder in the change detection model for processing, obtaining the building extraction result of the first flipped and translated image The building extraction result of the second flipped and translated image The building extraction result of the first distorted image And the building extraction result of the second distorted image Put and The loss function is calculated for these two to obtain the consistency segmentation loss L cov_seg1 . Before calculating the loss function, in order to make and correspond pixel by pixel, a data perturbation needs to be performed first. And since X u1_aug the predicted result of building extraction obtained there will be many mispredictions in the early stage of training. If all predicted pixels are directly used to calculate the loss, it will have a negative impact on model training. Therefore, before calculating the loss function, a filtering needs to be performed first using the threshold τ to leave the predicted pixels with high confidence and filter out the predicted pixels with low confidence. After performing data perturbation and threshold filtering on in sequence, the first building extraction pseudo-label is obtained Then and are used to calculate the loss function to obtain the first type of consistency segmentation loss L cov_seg1 . Similarly, after is subjected to data perturbation and threshold filtering, the second building extraction pseudo-label is obtained According to and the loss function is calculated to obtain the second type of consistency segmentation loss L cov_seg2 .

[0074] In the change detection part, the first flipped and translated image feature extraction result and the second flipped and translated image feature extraction result are spliced to obtain the flipped and translated image pair feature extraction result. The first distorted image feature extraction result and the second distorted image feature extraction result are spliced to obtain the distorted image pair feature extraction result. The flipped and translated image pair feature extraction result and the distorted image pair feature extraction result are input into the change detection decoder in the change detection model for processing to obtain the change detection prediction result of the flipped and translated image pair and the change detection prediction result of the distorted image pair After is subjected to data perturbation and threshold filtering, the change detection pseudo-label is obtained According to and the loss function is calculated to obtain the consistency change detection loss L cov_cd .

[0075] Specifically, the role of threshold filtering is to screen out the pixel prediction values (between 0 and 1) that exceed the confidence threshold τ. The pixel prediction values less than the confidence threshold τ will not participate in the calculation of the loss function. As Figure 7 shown, this avoids the pixel prediction values with low confidence being used as pseudo-labels and misleading the model.

[0076] 4. Loss Function Calculation

[0077] Considering the characteristic that the number of unchanged pixels in the task is much larger than the number of changed pixels, this application adopts a combination of cross - entropy loss (L ce ) and dice loss (L dice ) as the loss function, which is defined as follows:

[0078] L hybrid = L ce + L dice

[0079] Let the label be Y, and the changed image where H is the height of the image, W is the width of the image, is the pixel value of the k - th pixel of the image. Let c be 0 or 1, representing whether the k - th pixel in the label has changed. The calculation of the cross - entropy loss is as follows:

[0080]

[0081] In addition, the calculation of the dice loss is as follows:

[0082]

[0083] Therefore, the calculation of each part of the loss function is as follows:

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090] Adding the above - mentioned loss functions of each part, the total loss is obtained as:

[0091] L = L seg1 + L seg2 + L cd + L cov_seg1 + L cov_deg2 + L cov_cd

[0092] Taking the minimum total loss as the objective function, the parameters of the change detection model are trained and optimized using the AdamW optimization algorithm until the trained change detection model M is obtained.

[0093] (2) Main processes of testing and prediction

[0094] Input: The remote sensing image pair Img to be detected that the network has not learned A , Img B ∈R 256×256×3 Input the trained change detection model M.

[0095] Output: Change detection image D

[0096] Directly input the input image into the model M, and the output of the model is the change detection image D.

[0097] In the above semi-supervised remote sensing image change detection method based on joint learning, on the one hand, aiming at the change detection problem when the labels of the remote sensing image change detection task are insufficient, a semi-supervised change detection algorithm is proposed. This algorithm can achieve high-precision change detection by using a small amount of single-temporal remote sensing images and their building extraction labels and a large number of unlabeled double-temporal remote sensing image pairs. By forcing the model to make consistent change detection results for the unlabeled remote sensing images before and after data perturbation, the robustness and generalization of the model are improved, thereby improving the change detection accuracy. On the other hand, aiming at the problem that the production of remote sensing image change detection labels requires a large amount of time and labor costs, a large number of pseudo double-temporal remote sensing image data and pseudo change detection labels are generated from a small amount of single-temporal remote sensing image data with building extraction labels. The expansion of a large amount of data provides data support for the performance of the change detection model. And, aiming at the problem that there is too little remote sensing image information and label information to be fully utilized, a method of joint learning of the building extraction task and the change detection task is proposed. Since the training method of single-temporal change detection is used, the present invention can only use the building extraction labels, combine the remote sensing image change detection task and the building extraction task to improve the training effect of the network, without the need to use both building extraction labels and change detection labels at the same time. Joint learning makes the model use remote sensing image information more fully, and uses the building extraction task to enhance the performance of the feature extraction network, improving the change detection effect.

[0098] Furthermore, to verify the effect of the semi-supervised remote sensing image change detection method based on joint learning proposed by the present invention, 10 single-temporal remote sensing images and corresponding building extraction labels and 2000 pairs of unlabeled double-temporal remote sensing images are used for training to obtain a trained change detection model. As Figure 8As shown in the figure, the first two columns are the input images, the third column is the building extraction labels corresponding to the single-temporal remote sensing images, the fourth column is the change detection effect of the model trained with 10 single-temporal remote sensing images and the corresponding building extraction labels in a fully supervised manner, and the fifth column is the change detection effect of the model trained with 10 single-temporal remote sensing images, the corresponding building extraction labels, and 2000 pairs of unlabeled double-temporal remote sensing images in a semi-supervised manner. Through comparison, it can be found that the remote sensing image change detection results obtained by the semi-supervised remote sensing image change detection method based on joint learning proposed in the present invention are closer to the true labels, and the accuracy of change detection is higher.

[0099] It should be understood that although Figures 1 - 2 the steps in the flowchart of Figures 1 - 2 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,

[0100] In one embodiment, a semi-supervised remote sensing image change detection device based on joint learning is provided, including:

[0101] A data acquisition module, configured to acquire single-temporal remote sensing images and corresponding building extraction labels, as well as unlabeled double-temporal remote sensing image pairs;

[0102] A data preprocessing module, configured to preprocess the single-temporal remote sensing images and corresponding building extraction labels to obtain pseudo double-temporal remote sensing image pairs and corresponding building extraction labels, perform an exclusive OR operation on the building extraction labels of the pseudo double-temporal remote sensing image pairs to obtain pseudo change detection labels;

[0103] A supervised learning module, configured to input the pseudo double-temporal remote sensing image pairs and corresponding building extraction labels, as well as the pseudo change detection labels into a pre-constructed change detection model, process the pseudo double-temporal remote sensing image pairs according to the feature extraction network, building extraction decoder, and change detection decoder in the change detection model, output the supervised building extraction results and supervised change detection prediction results, calculate the segmentation loss according to the supervised building extraction results and the building extraction labels of the pseudo double-temporal remote sensing image pairs, and calculate the change detection loss according to the supervised change detection prediction results and the pseudo change detection labels;

[0104] An unsupervised learning module is used to sequentially perform flipping and translation and data perturbation on unlabeled dual-temporal remote sensing image pairs to obtain flipped and translated image pairs and distorted image pairs. The flipped and translated image pairs and the distorted image pairs are input into a change detection model for processing, and the building extraction results of the flipped and translated image pairs, the building extraction results of the distorted image pairs, the change detection prediction results of the flipped and translated image pairs, and the change detection prediction results of the distorted image pairs are output. According to the building extraction results of the flipped and translated image pairs and the building extraction results of the distorted image pairs, a consistency segmentation loss is calculated. According to the change detection prediction results of the flipped and translated image pairs and the change detection prediction results of the distorted image pairs, a consistency change detection loss is calculated;

[0105] A model training module is used to add the segmentation loss, the change detection loss, the consistency segmentation loss, and the consistency change detection loss to obtain a total loss. Taking the minimum total loss as the objective function, the parameters of the change detection model are trained and optimized until a trained change detection model is obtained;

[0106] A model testing module is used to input the remote sensing image pairs to be detected into the trained change detection model for change detection and output the change detection results.

[0107] For the specific limitations of the semi-supervised remote sensing image change detection device based on joint learning, reference can be made to the limitations of the semi-supervised remote sensing image change detection method in the above text, which will not be elaborated here. Each module in the above semi-supervised remote sensing image change detection device based on joint learning can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0108] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0109] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A semi-supervised remote sensing image change detection method based on joint learning, characterized in that, The method includes: Obtaining a single-temporal remote sensing image and its corresponding building extraction label, as well as an unlabeled dual-temporal remote sensing image pair; Preprocessing the single-temporal remote sensing image and its corresponding building extraction label to obtain a pseudo-dual-temporal remote sensing image pair and its corresponding building extraction label, performing an exclusive OR operation on the building extraction labels of the pseudo-dual-temporal remote sensing image pair to obtain a pseudo-change detection label; Inputting the pseudo-dual-temporal remote sensing image pair and its corresponding building extraction label, as well as the pseudo-change detection label, into a pre-constructed change detection model, processing the pseudo-dual-temporal remote sensing image pair according to the feature extraction network, building extraction decoder, and change detection decoder in the change detection model, outputting a supervised building extraction result and a supervised change detection prediction result, calculating a segmentation loss based on the supervised building extraction result and the building extraction label of the pseudo-dual-temporal remote sensing image pair, and calculating a change detection loss based on the supervised change detection prediction result and the pseudo-change detection label; Performing flipping, translation, and data perturbation on the unlabeled dual-temporal remote sensing image pair in sequence to obtain a flipped and translated image pair and a distorted image pair, inputting the flipped and translated image pair and the distorted image pair into the change detection model for processing, outputting the building extraction result of the flipped and translated image pair, the building extraction result of the distorted image pair, the change detection prediction result of the flipped and translated image pair, and the change detection prediction result of the distorted image pair, calculating a consistency segmentation loss based on the building extraction result of the flipped and translated image pair and the building extraction result of the distorted image pair, and calculating a consistency change detection loss based on the change detection prediction result of the flipped and translated image pair and the change detection prediction result of the distorted image pair; Adding the segmentation loss, change detection loss, consistency segmentation loss, and consistency change detection loss to obtain a total loss, taking the minimum total loss as the objective function, training and optimizing the parameters of the change detection model until a trained change detection model is obtained; Inputting the remote sensing image pair to be detected into the trained change detection model for change detection, and outputting a change detection result.

2. The method according to claim 1, wherein Preprocessing the single-temporal remote sensing image and its corresponding building extraction label to obtain a pseudo-dual-temporal remote sensing image pair and its corresponding building extraction label, performing an exclusive OR operation on the building extraction labels of the pseudo-dual-temporal remote sensing image pair to obtain a pseudo-change detection label, including: For single-temporal remote sensing image X l and the corresponding building extraction label Y be Perform the first random sampling, flipping, and translation to obtain the preprocessed single-temporal remote sensing image X l1 and the corresponding building extraction label Y be_1 ; Perform a second random sampling and flipping translation on the said X l and Y be to obtain the single-temporal remote sensing image X after the second preprocessing l2 and the corresponding building extraction label Y be_2 ; Pair the X l1 and X l2 to obtain a pseudo-bitemporal remote sensing image pair. Pair the Y be_1 and Y be_2 to obtain the building extraction label of the pseudo-bitemporal remote sensing image pair. Perform an exclusive OR operation on the paired Y be_1 and Y be_2 to obtain the pseudo-change detection label Y cd .

3. The method according to claim 2, wherein Inputting the pseudo-dual-temporal remote sensing image pair and its corresponding building extraction label, as well as the pseudo-change detection label, into a pre-constructed change detection model, processing the pseudo-dual-temporal remote sensing image pair according to the feature extraction network, building extraction decoder, and change detection decoder in the change detection model, outputting a supervised building extraction result and a supervised change detection prediction result, including: The X in the pseudo-bitemporal remote sensing image pair l1 and X l2 , the Y in the building extraction label of the pseudo-bitemporal remote sensing image pair be_1 and Y be_2 and the pseudo-change detection label Y cd are input into the change detection model; The first feature extraction network and the second feature extraction network with shared weights in the change detection model respectively perform feature extraction on the X l1 and X l2 to obtain a first type of supervised feature extraction result and a second type of supervised feature extraction result; Input the first type of supervised feature extraction result and the second type of supervised feature extraction result into the building extraction decoder for processing to obtain the first type of supervised building extraction result and the second type of supervised building extraction result Concatenate the first type of supervised feature extraction result and the second type of supervised feature extraction result, and input the concatenated supervised feature extraction result into the change detection decoder for processing to obtain a supervised change detection prediction result 4. The method according to claim 3, characterized in that, Calculate the segmentation loss based on the supervised building extraction result and the building extraction label of the pseudo-bispectral remote sensing image pair, and calculate the change detection loss based on the supervised change detection prediction result and the pseudo-change detection label, including: According to the supervised building extraction results of the first category and the Y in the building extraction labels of the pseudo bi-temporal remote sensing image pairs be_1 perform calculations to obtain the first type of segmentation loss L seg1 , according to the supervised building extraction results of the second category and the Y in the building extraction labels of the pseudo bi-temporal remote sensing image pairs be_2 perform calculations to obtain the second type of segmentation loss L seg2 ; According to the prediction results of supervised change detection and the pseudo-change detection label Y cd calculate to obtain the change detection loss L cd .

5. The method according to claim 1, wherein Perform flipping and translation and data perturbation on the unlabeled bispectral remote sensing image pair in sequence to obtain a flipped and translated image pair and a distorted image pair, including: For the unlabeled dual-temporal remote sensing image pair X u The first unlabeled dual-temporal remote sensing image X in u1 and the second unlabeled dual-temporal remote sensing image X u2 are subjected to the same random flipping and translation to obtain the first flipped and translated image X u1_aug and the second flipped and translated image X u2_aug , X u1_aug and X u2_aug constitute a flipped and translated image pair; Perform the same random data perturbation on the said X u1_aug and X u2_aug to obtain the first distorted image X u1_per and the second distorted image X u2_per , X u1_per and X u2_per constitute a pair of distorted images.

6. The method according to claim 5, wherein Input the flipped and translated image pair and the distorted image pair into the change detection model for processing, and output the building extraction result of the flipped and translated image pair, the building extraction result of the distorted image pair, the change detection prediction result of the flipped and translated image pair, and the change detection prediction result of the distorted image pair, including: Input the X in the flipped and translated image pair u1_aug and X u2_aug into the first feature extraction network and the second feature extraction network with shared weights in the change detection model respectively for feature extraction, obtaining the first flipped and translated image feature extraction result and the second flipped and translated image feature extraction result. Input the X in the distorted image pair u1_per and X u2_per into the first feature extraction network and the second feature extraction network respectively for feature extraction, obtaining the first distorted image feature extraction result and the second distorted image feature extraction result; Input the first flipped and translated image feature extraction result, the second flipped and translated image feature extraction result, the first distorted image feature extraction result, and the second distorted image feature extraction result into the building extraction decoder in the change detection model for processing to obtain the building extraction result of the first flipped and translated image The building extraction result of the second flipped and translated image The building extraction result of the first distorted image And the building extraction result of the second distorted image Stitch the first flipped and translated image feature extraction result and the second flipped and translated image feature extraction result to obtain a flipped and translated image pair feature extraction result, and stitch the first distorted image feature extraction result and the second distorted image feature extraction result to obtain a distorted image pair feature extraction result; Input the feature extraction results of the flipped and translated image pair and the feature extraction results of the distorted image pair into the change detection decoder in the change detection model to obtain the change detection prediction results of the flipped and translated image pair and the change detection prediction results of the distorted image pair 7. The method according to claim 6, characterized in that, Calculate the consistency segmentation loss based on the building extraction result of the flipped and translated image pair and the building extraction result of the distorted image pair, and calculate the consistency change detection loss based on the change detection prediction result of the flipped and translated image pair and the change detection prediction result of the distorted image pair, including: Building extraction results of the first flipped and translated image Building extraction results of the second flipped and translated image And change detection prediction results of the flipped and translated image pair Perform data perturbation and threshold filtering in sequence to obtain the first building extraction pseudo-label The second building extraction pseudo-label And the change detection pseudo-label According to the building extraction result of the first distorted image perform calculations to obtain the first type of consistency segmentation loss L cov_seg1 According to the building extraction result of the second distorted image perform calculations to obtain the second type of consistency segmentation loss L cov_seg2 According to the change detection prediction result of the distorted image pair perform calculations to obtain the consistency change detection loss L cov_cd .

8. The method according to claim 1, wherein Add the segmentation loss, the change detection loss, the consistency segmentation loss, and the consistency change detection loss to obtain the total loss, including: Add the first type of segmentation loss \(L\) seg1 , the second type of segmentation loss \(L\) seg2 , the change detection loss \(L\) cd , the first type of consistent segmentation loss \(L\) cov_seg1 , the second type of consistent segmentation loss \(L\) cov_seg2 and the consistent change detection loss \(L\) cov_cd to obtain the total loss, denoted as L = L seg1 + L seg2 + L cd + L cov_seg1 + L cov_seg2 + L cov_cd 。 9. The method according to claim 1 or 5 or 7, characterized in that, The data perturbation includes performing color perturbation and / or shape perturbation with a preset intensity and preset combination on the image, the building extraction result, and the change detection prediction result.

10. A semi-supervised remote sensing image change detection device based on joint learning, characterized in that The device includes: A data acquisition module for acquiring a single-temporal remote sensing image and the corresponding building extraction label, and an unlabeled bispectral remote sensing image pair; A data preprocessing module for preprocessing the single-temporal remote sensing image and the corresponding building extraction label to obtain a pseudo-bispectral remote sensing image pair and the corresponding building extraction label, and performing an exclusive OR operation on the building extraction labels of the pseudo-bispectral remote sensing image pair to obtain a pseudo-change detection label; A supervised learning module for inputting the pseudo-bispectral remote sensing image pair and the corresponding building extraction label, and the pseudo-change detection label into a pre-constructed change detection model, processing the pseudo-bispectral remote sensing image pair according to the feature extraction network, the building extraction decoder, and the change detection decoder in the change detection model, and outputting a supervised building extraction result and a supervised change detection prediction result, calculating the segmentation loss based on the supervised building extraction result and the building extraction label of the pseudo-bispectral remote sensing image pair, and calculating the change detection loss based on the supervised change detection prediction result and the pseudo-change detection label; An unsupervised learning module is used to perform flipping and translation and data perturbation on the unlabeled dual-temporal remote sensing image pairs in sequence to obtain the flipped and translated image pairs and the distorted image pairs. The flipped and translated image pairs and the distorted image pairs are input into the change detection model for processing, and the building extraction results of the flipped and translated image pairs, the building extraction results of the distorted image pairs, the change detection prediction results of the flipped and translated image pairs, and the change detection prediction results of the distorted image pairs are output. According to the building extraction results of the flipped and translated image pairs and the building extraction results of the distorted image pairs, a consistency segmentation loss is calculated. According to the change detection prediction results of the flipped and translated image pairs and the change detection prediction results of the distorted image pairs, a consistency change detection loss is calculated; A model training module is used to add the segmentation loss, the change detection loss, the consistency segmentation loss, and the consistency change detection loss to obtain a total loss. Taking the minimum total loss as the objective function, the parameters of the change detection model are trained and optimized until a trained change detection model is obtained; A model testing module is used to input the remote sensing image pairs to be detected into the trained change detection model for change detection and output the change detection results.

Citation Information

Patent Citations

  • Method for flood disaster monitoring and disaster analysis based on vision transformer

    US11521379B1

  • System and method for whole body landmark detection, segmentation and change quantification in digital images

    US20070081712A1