Semi-supervised multi-temporal satellite image time-varying information extraction method

Through semi-supervised learning and pseudo-label optimization strategies, combined with multi-scale feature fusion and space-time coupling mechanism, the problems of high labeling costs and low detection accuracy in remote sensing image change detection are solved, and the detection performance and robustness of the model in complex shape change areas are improved.

CN120564070AActive Publication Date: 2025-08-29HARBIN AEROSPACE STAR DATA SYST TECH CO LTD +1

Patent Information

Application Number
CN202510699214.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-29
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

In the detection of remote sensing image change, especially in the detection of irregular shape change areas, there are problems such as high labeling cost, low efficiency, insufficient model generalization ability and low detection accuracy, especially in complex shape change areas and small samples categories, which are difficult to accurately identify.

Method used

A semi-supervised learning strategy is adopted, combining pseudo-label optimization and consistency regularization, and a twin-structure encoder-decoder architecture is constructed, multi-scale feature fusion and space-time coupling mechanisms are used, confidence consistency correction and confidence threshold filtering are combined to generate high-quality pseudo-labels, and semi-supervised training is carried out to improve model performance.

Benefits of technology

The detection accuracy and semantic change detection performance of complex shape change areas are improved, the annotation cost is reduced, the robustness and generalization ability of the model are enhanced, and efficient semantic change detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564070A_ABST
    Figure CN120564070A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised multi-temporal satellite image time-varying information extraction method, and belongs to the technical field of remote sensing image processing. The semantic change detection performance of the model and the detection precision of complex shape change ground objects are improved. The method comprises the following steps: constructing a semantic change detection model, carrying out full-supervised training on the semantic change detection model by using a binary change detection supervised loss function and a semantic segmentation supervised loss function by using a small amount of labeled dual-temporal remote sensing images to obtain an initial model, and obtaining a semantic change detection prediction result of each pair of images; using a pseudo label optimization strategy to optimize the semantic change detection prediction result of each pair of images; combining the obtained pseudo label data with high confidence and a small amount of labeled dual-temporal remote sensing images into a new training set, using the new training set to perform semi-supervised training on the initial model in a semantic change detection model, and using a consistency regularization combination loss function to perform supervised training to obtain a new model; and until a preset number of iterations is reached.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of remote sensing image processing, and in particular relates to a semi-supervised method for extracting time-varying information from multi-temporal satellite images. Background Art

[0002] The changes in land features in multi-temporal remote sensing satellite imagery are often random and complex, resulting in a variety of irregular shapes in the areas of change. In urban scenes, the areas of change resulting from processes such as building expansion and demolition usually do not conform to rectangular or regular polygonal forms; in natural scenes, changes in river morphology due to erosion and sedimentation are manifested as irregular, tortuous, and bifurcated boundaries. These irregularly shaped areas of change not only require change detection models to simultaneously complete the two tasks of regional detection and internal semantic classification, but also place higher demands on the model's structural design. It is necessary to extract global semantics while accurately depicting detailed boundaries. Otherwise, pseudo-changes such as discontinuities, holes, or noise are likely to occur. The complex landscape and spatial heterogeneity in remote sensing images are one of the main factors that make change detection difficult.

[0003] At the same time, semantic change detection methods based on deep learning are highly dependent on the quality and quantity of annotated data. This is especially true for accurately labeling irregularly shaped regions of change, which often requires manual pixel-level segmentation or object-oriented multi-category labeling. Labeling high-resolution, multi-temporal remote sensing data not only requires specialized domain knowledge but also consumes significant manpower and time, significantly increasing project costs. The low efficiency of manual labeling of large-scale remote sensing data has become a bottleneck restricting the research and application of change detection.

[0004] At the same time, for irregular water system change categories such as rivers and wetlands, the number of real change samples is limited, making it difficult to form a sufficient training set, resulting in generally low detection accuracy for these small-sample categories. On the one hand, the lack of diverse training samples can reduce the model's generalization ability. On the other hand, because the spatiotemporal changes of natural landscapes such as rivers are influenced by multiple factors such as climate and hydrological conditions, their morphological changes are highly heterogeneous, making it more difficult for a small number of samples to capture typical change patterns, further exacerbating missed detections and false detections. Furthermore, research on large-scale building damage detection has also shown that it is difficult to learn stable local and neighborhood relationships in small-sample scenarios, which affects the effectiveness of change discrimination.

[0005] Semi-supervised strategies have shown potential in reducing labeling requirements and accelerating model training by jointly training with a small amount of precisely labeled data and a large amount of unlabeled data. Based on the consistency regularization method, pseudo labels are generated from unlabeled dual-temporal images to improve model stability and robustness. Further combined with graph convolution, self-attention, and multi-scale network structures, feature enhancement and iterative updates of pseudo labels under weak supervision are achieved, significantly alleviating the problem of label scarcity. However, the performance of semi-supervised strategies in complex irregular shape change areas is still unsatisfactory: pseudo labels cannot accurately depict curved boundaries, and reducing the amount of annotations exacerbates the loss of shape details. Therefore, how to enhance the feature expression and background discrimination of irregular change areas under the semi-supervised framework remains a key direction for future research. Summary of the Invention

[0006] The problem to be solved by the present invention is to improve the semantic change detection performance of the model and the detection accuracy of complex shape-changing objects, and propose a semi-supervised method for extracting time-varying information from multi-temporal satellite images.

[0007] To achieve the above object, the present invention is implemented through the following technical solutions:

[0008] A semi-supervised method for extracting time-varying information from multi-temporal satellite images includes the following steps:

[0009] S1. Collect a small amount of labeled bi-temporal remote sensing images and their corresponding semantic change detection label data, collect a large amount of unlabeled bi-temporal remote sensing image data, and preprocess the collected bi-temporal remote sensing images to obtain a preprocessed bi-temporal remote sensing image set.

[0010] S2. Build a semantic change detection model, including two encoders and decoders, a multi-scale feature fusion module, a bi-temporal feature interaction module, a change detection module, and two semantic segmentation modules.

[0011] S3. Using the small amount of preprocessed, annotated bi-temporal remote sensing images obtained in step S1, perform fully supervised training on the semantic change detection model constructed in step S2 using a binary change detection supervised loss function and a semantic segmentation supervised loss function to obtain the initial model model_1.

[0012] S4. Inputting the large amount of unlabeled dual-temporal remote sensing image data obtained in step S1 into the initial model model_1 obtained in step S3 for prediction, obtaining semantic change detection prediction results for each pair of images;

[0013] S5. Optimize the semantic change detection prediction results for each pair of images obtained in step S4 using a pseudo-label optimization strategy, wherein the pseudo-label optimization strategy includes confidence consistency correction and confidence threshold filtering to generate high-confidence pseudo-label data;

[0014] S6. Combine the high-confidence pseudo-label data obtained in step S5 with the small amount of annotated bi-temporal remote sensing images obtained in step S1 to form a new training set. Use this new training set to perform semi-supervised training on the initial model model_1 in the semantic change detection model, and supervise the training using a consistency regularized combined loss function to obtain a new model model_2.

[0015] S7. Iterate steps S4 to S6 until a predetermined number of iterations is reached, and finally output model model_N. Model model_N is used to extract time-varying information from multi-temporal satellite images.

[0016] Furthermore, the preprocessing method in step S1 is to match the size of the dual-phase remote sensing image, and detect whether the dual-phase remote sensing image is a three-channel image, set the dual-phase remote sensing images of the previous time and the next time with matching size and three channels as an image pair, convert the image pair into a numpy format file, and obtain the preprocessed dual-phase remote sensing image.

[0017] Furthermore, the specific implementation method of step S2 includes the following steps:

[0018] S2.1. The semantic change detection model uses a twin-structure encoder-decoder architecture. The multi-scale feature fusion module is embedded at the end of the encoder-decoder architecture. The output of the multi-scale feature fusion module is connected to a bi-temporal feature interaction module, which is connected to the change detection module and two semantic segmentation modules.

[0019] S2.2. The multi-scale feature fusion module is composed of two complementary submodules: the global semantic enhancement submodule and the local edge optimization submodule. The output expression of the global semantic enhancement submodule is:

[0020]

[0021] in, is the output of the global semantic enhancer module, represents the feature concatenation operation of the channel dimension, represents the sigmoid activation function, The expansion rate is The dilated convolution operation, represents pixel-by-pixel multiplication, Represents the multi-scale features of the 4th or 3rd layer of the encoder, Represents the multi-scale features of the first or second layer of the encoder, i is a cubic value from 1 to 3;

[0022] The output expression of the local edge optimization submodule is:

[0023]

[0024] in, is the output of the local edge optimization submodule;

[0025] S2.3. A dual-temporal feature interaction module is established to achieve collaborative modeling of cross-temporal information through a spatiotemporal coupling mechanism. The core process consists of two stages: obtaining a change weight map and obtaining spatial semantic features.

[0026] The change weight map of the dual-temporal feature interaction module and the generation process of spatial semantic features with spatiotemporal information are shown in the following formula:

[0027]

[0028]

[0029] in, is the change weight graph, Optimize the features for the local edge at time 1, Optimize features for local edges at time 2, To carry the spatial semantic features of rich spatiotemporal information, represents the convolutional layer operation, Represents a concatenation operation in the channel dimension.

[0030] Furthermore, during the training process of step S3, the loss function is composed of multivariate cross entropy loss and binary change cross entropy loss The composition is expressed as:

[0031]

[0032] in, is the loss function for fully supervised training, is the multivariate cross entropy loss at time 1, is the multivariate cross entropy loss at time 2, is the binary change cross entropy loss.

[0033] Furthermore, step S4 uses the initial model model_1 to make predictions, obtain the semantic change detection prediction results of each pair of images, and use the results as the initial pseudo labels .

[0034] Furthermore, the confidence consistency correction method in step S5 is:

[0035] For the first round of model_1 prediction of a large amount of unlabeled data, confidence consistency correction is not used, but confidence threshold filtering is performed directly to the initial pseudo label. Pixels with low confidence are shielded;

[0036] For the update of the i-1th round model model_i, retain the pseudo labels used to train the i-1th round model model_i , while retaining the predicted pseudo labels of model_i for unlabeled data Through fusion and , correct some erroneous pixels in the pseudo-label of the same unlabeled remote sensing image pair and generate the optimized pseudo-label of the i-th round , used for training model model_i+1, fusion and The expression is:

[0037]

[0038]

[0039]

[0040] in, is the confidence at pixel position (x, y) in the mask image of round i-1 and round i, is the probability distribution of the nth feature generated by the pseudo-label in the i-1th round, is the probability distribution of the nth feature generated by the pseudo-label in the i-th round, is any one of N, indicating the nth type of land feature, is the probability of the nth type of feature at (x, y) in the i-th round of training, is the total number of land feature categories.

[0041] Furthermore, the confidence threshold filtering method in step S5 is to set the confidence level below the preset threshold. The pixel points are shielded, and the shielded pixel samples do not participate in the calculation of the current round loss function and the optimization of model parameters, and are retained. and Medium confidence is higher than The pixel samples are used as optimized pseudo labels for the next round of model training;

[0042] The expression of confidence is:

[0043]

[0044] in, yes The confidence of the sample at the pixel, represents the entropy of the probability distribution of the pixel value at (x, y), represents the normalization factor.

[0045] Furthermore, the consistency regularization combined loss function of step S6 is The expression is:

[0046]

[0047] in, is the consistency loss function, is a hyperparameter;

[0048] Use KL divergence loss to calculate the consistency loss function, the calculation formula is:

[0049]

[0050]

[0051] Where M represents the number of unlabeled remote sensing image pairs, represents a pseudo label, Represents the prediction result of the unlabeled image, Num represents the number of changing object categories in the dataset, represents the probability of category j in the pseudo-label probability distribution, Represents the probability of category j in the prediction result probability distribution; Represents the difference between the pseudo label with random perturbations and the predicted results of remote sensing images.

[0052] Beneficial effects of the present invention:

[0053] The semi-supervised method for extracting time-varying information from multi-temporal satellite images described in the present invention addresses the limitations of pixel-level semantic change detection, which is expensive, time-consuming, and requires a high level of expertise, as well as the limited number of samples of complex shape-varying objects. A semi-supervised training strategy for semantic change detection models is proposed, aiming to train a model with better performance using a small amount of labeled data and a large amount of unlabeled data. First, a confidence consistency correction strategy and a confidence threshold filtering strategy are proposed for the auxiliary model to learn pseudo-labels of object distribution information in unlabeled data to optimize the quality of the pseudo-labels. Furthermore, the KL divergence calculation is introduced into the consistency regularization loss function to calculate the loss between unlabeled data and optimized pseudo-labels. During this process, image and label enhancement operations are performed to impose consistency regularization constraints on the model, enhancing its resistance to interference factors such as noise, thereby improving the model's semantic change detection performance and the detection accuracy of complex shape-varying objects in the absence of insufficient labels. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a flow chart of a semi-supervised method for extracting time-varying information from multi-temporal satellite images according to the present invention;

[0055] Figure 2 This is a flowchart of the extraction of time-varying information from temporal satellite images according to the present invention;

[0056] Figure 3 Schematic diagram of the process of training a semantic change detection network under the semi-supervised strategy of the present invention;

[0057] Figure 4 Schematic diagram of the semantic change detection network of the present invention. DETAILED DESCRIPTION

[0058] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present invention and are not intended to limit the present invention. That is, the specific embodiments described herein are only some embodiments of the present invention, not all embodiments. Generally, the components of the specific embodiments of the present invention described and illustrated in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0059] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but is merely representative of selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0060] In order to further understand the content, features and effects of the present invention, the following specific embodiments are given as examples, and the attached Figure 1 -Attached Figure 4 The detailed instructions are as follows:

[0061] Example 1:

[0062] A semi-supervised method for extracting time-varying information from multi-temporal satellite images includes the following steps:

[0063] S1. Collect a small amount of labeled bi-temporal remote sensing images and their corresponding semantic change detection label data, collect a large amount of unlabeled bi-temporal remote sensing image data, and preprocess the collected bi-temporal remote sensing images to obtain a preprocessed bi-temporal remote sensing image set.

[0064] Furthermore, the preprocessing method in step S1 is to match the size of the dual-temporal remote sensing image, detect whether the dual-temporal remote sensing image is a three-channel image, set the dual-temporal remote sensing image of the previous time and the next time with matching size and three channels as an image pair, convert the image pair into a numpy format file, and obtain the preprocessed dual-temporal remote sensing image;

[0065] S2. Build a semantic change detection model, including two encoders and decoders, a multi-scale feature fusion module, a bi-temporal feature interaction module, a change detection module, and two semantic segmentation modules.

[0066] Furthermore, the specific implementation method of step S2 includes the following steps:

[0067] S2.1. The semantic change detection model uses a twin-structure encoder-decoder architecture. The multi-scale feature fusion module is embedded at the end of the encoder-decoder architecture. The output of the multi-scale feature fusion module is connected to a bi-temporal feature interaction module, which is connected to the change detection module and two semantic segmentation modules.

[0068] S2.2. The multi-scale feature fusion module is composed of two complementary submodules: the global semantic enhancement submodule and the local edge optimization submodule. The output expression of the global semantic enhancement submodule is:

[0069]

[0070] in, is the output of the global semantic enhancer module, represents the feature concatenation operation of the channel dimension, represents the sigmoid activation function, The expansion rate is The dilated convolution operation, represents pixel-by-pixel multiplication, Represents the multi-scale features of the 4th or 3rd layer of the encoder, Represents the multi-scale features of the first or second layer of the encoder, i is a cubic value from 1 to 3;

[0071] The output expression of the local edge optimization submodule is:

[0072]

[0073] in, is the output of the local edge optimization submodule;

[0074] Furthermore, a multi-scale feature fusion module is embedded at the end of the encoder-decoder architecture, responsible for integrating the spatial-semantic information of cross-level features. The dimensions of the four-layer feature tensors output by the encoder are 64×W / 2×H / 2, 128×W / 4×H / 4, 256×W / 8×H / 8, and 512×W / 8×H / 8 (the input size is 3×W×H), respectively. The deep features maintain a W / 8×H / 8 resolution to preserve high-dimensional semantic information. After upsampling and reconstruction by the decoder, a feature sequence with increasing resolution is obtained, which is divided into two groups of heterogeneous features: a high-level semantic feature group (F3, F4) and a low-level detail feature group (F1, F2), which are input into a two-way fusion structure respectively.

[0075] Furthermore, the global semantic enhancement module extracts multi-scale contextual features through three-way parallel dilated convolution (dilation rates r = 1, 2, and 4), and generates a spatial attention weight map through sigmoid activation. This weight map is then fused with the underlying high-resolution features in a channel-wise weighted manner to highlight the shape characteristics and spatial positioning information of irregular objects.

[0076] S2.3. A dual-temporal feature interaction module is established to achieve collaborative modeling of cross-temporal information through a spatiotemporal coupling mechanism. The core process consists of two stages: obtaining a change weight map and obtaining spatial semantic features.

[0077] Furthermore, the initial change feature is obtained by concatenating the outputs of the multi-scale feature fusion module on both branches, a common approach for obtaining change features in most SCD methods. Next, the initial change feature is passed through a series of convolutional layers to fully extract spatiotemporal information, which is crucial for guiding the generation of the final mask. The change feature, enhanced in both spatiotemporal dimensions, is then normalized using a sigmoid function to highlight the change target in the feature map. Subsequently, inverse attention is applied to further amplify the distinction between the change target and the background, extracting detailed edge information. This series of processes generates a change weight map that not only highlights the overall location of the change target but also pays special attention to its detailed edges. The change weight map is then used to weight the output features of the previous module on both branches, resulting in a dual-branch spatial semantic feature that carries rich spatiotemporal information and models the correlation between temporal, spatial, and semantic information. Finally, these two enhanced spatial semantic features are concatenated in the channel dimension to produce the enhanced change feature.

[0078] The change weight map of the dual-temporal feature interaction module and the generation process of spatial semantic features with spatiotemporal information are shown in the following formula:

[0079]

[0080]

[0081] in, is the change weight graph, Optimize the features for the local edge at time 1, Optimize features for local edges at time 2, To carry the spatial semantic features of rich spatiotemporal information, represents the convolutional layer operation, Represents a concatenation operation in the channel dimension.

[0082] Furthermore, the enhanced spatial semantic features are fed into the dual-temporal feature interaction module, and the change features obtained through concatenation, which carry spatiotemporal information, are used as auxiliary information. First, the multi-scale fusion features extracted by the two branches of the Siamese network are concatenated in the channel dimension to form an initial spatiotemporal joint representation. This is then processed through a series of convolutional layers to extract spatiotemporal contextual features and discriminative change features. This feature is normalized with a sigmoid function to generate a spatial attention weight map, which highlights change-sensitive areas and suppresses background interference. Subsequently, reverse attention is applied to further amplify the distinction between the changing target and the background, extracting detailed edge information and achieving secondary optimization of the feature map. The change weight map is used to weight the output features of the previous module on both branches, resulting in a dual-branch spatial semantic feature with rich spatiotemporal information. This allows for modeling the correlation between temporal, spatial, and semantic information. Finally, these two enhanced spatial semantic features are concatenated in the channel dimension to produce an enhanced change feature.

[0083] Furthermore, the function of the dual-temporal feature interaction module is to integrate multi-layer features carrying semantic information and spatial information on different time branches to obtain change features carrying spatiotemporal information, and use the change features to provide guidance for the identification of land feature types in the change area. Before performing the above operations, it is first necessary to connect the output of the previous module on a single time branch. The 4-1 fusion features and the 3-2 fusion features on a single branch are cascaded in the channel dimension to obtain a single-branch spatial semantic enhancement feature. The single-branch spatial semantic enhancement feature and the 4th layer feature of the encoder-decoder are cascaded again to ensure that important semantic information in the image is not lost. After the operation is completed on both branches, the final dual-branch spatial semantic feature can be input into the dual-branch feature interaction module to perform the process of semantic change mask generation guided by change information described above;

[0084] Furthermore, a twin-structured encoder-decoder architecture is used to process dual-temporal remote sensing imagery. During the feature extraction phase, a ResNet34-based backbone network extracts multi-level spatial features from the two temporal images. Through a designed cross-scale feature fusion mechanism, the network integrates low-level fine-grained features with high-level abstract features, effectively coupling local texture details with global semantic information. The fused enhanced features are then input into the temporal interaction module, which constructs spatiotemporal associations through feature splicing and information interaction mechanisms, providing a joint spatial-spectral representation for the identification of changed areas. Finally, a classifier consisting of deconvolutional and convolutional layers outputs a binary changed area detection map and a multi-category land cover change semantic map in parallel, enabling a joint analysis from pixel-level change detection to semantic-level change interpretation.

[0085] Furthermore, the temporal satellite image semantic change detection dataset was divided into a training set and a test set. A training group was extracted from the training set, and each training group was used as a training batch. The training set was used to train the semantic change detection network to obtain a trained semantic change detection model. The trained semantic change detection network was then evaluated using the test set to obtain an evaluation of the semantic change detection model.

[0086] S3. Using the small amount of preprocessed, annotated bi-temporal remote sensing images obtained in step S1, perform fully supervised training on the semantic change detection model constructed in step S2 using a binary change detection supervised loss function and a semantic segmentation supervised loss function to obtain the initial model model_1.

[0087] Furthermore, during the training process of step S3, the loss function is composed of multivariate cross entropy loss and binary change cross entropy loss The composition is expressed as:

[0088]

[0089] in, is the loss function for fully supervised training, is the multivariate cross entropy loss at time 1, is the multivariate cross entropy loss at time 2, is the binary change cross entropy loss.

[0090] S4. Inputting the large amount of unlabeled dual-temporal remote sensing image data obtained in step S1 into the initial model model_1 obtained in step S3 for prediction, obtaining semantic change detection prediction results for each pair of images;

[0091] Furthermore, step S4 uses the initial model model_1 to make predictions, obtain the semantic change detection prediction results of each pair of images, and use the results as the initial pseudo labels .

[0092] S5. Optimize the semantic change detection prediction results for each pair of images obtained in step S4 using a pseudo-label optimization strategy, wherein the pseudo-label optimization strategy includes confidence consistency correction and confidence threshold filtering to generate high-confidence pseudo-label data;

[0093] Furthermore, confidence consistency correction modifies some erroneous pixels in the initial pseudo-label according to the changes in ground objects in real remote sensing images, so as to improve the accuracy of identifying semantic changes; confidence threshold filtering improves the quality of the new training set by selecting samples with confidence higher than a preset threshold in the initial pseudo-label as pseudo-label data to expand the new training set.

[0094] Furthermore, the confidence consistency correction method in step S5 is:

[0095] For the first round of model_1 prediction of a large amount of unlabeled data, confidence consistency correction is not used, but confidence threshold filtering is performed directly to the initial pseudo label. Pixels with low confidence are shielded;

[0096] For the update of the i-1th round model model_i, retain the pseudo labels used to train the i-1th round model model_i , while retaining the predicted pseudo labels of model_i for unlabeled data Through fusion and , correct some erroneous pixels in the pseudo-label of the same unlabeled remote sensing image pair and generate the optimized pseudo-label of the i-th round , used for training model model_i+1, fusion and The expression is:

[0097]

[0098]

[0099]

[0100] in, is the confidence at pixel position (x, y) in the mask image of round i-1 and round i, is the probability distribution of the nth feature generated by the pseudo-label in the i-1th round, is the probability distribution of the nth feature generated by the pseudo-label in the i-th round, is any one of N, indicating the nth type of land feature, is the probability of the nth type of feature at (x, y) in the i-th round of training, is the total number of land feature categories.

[0101] Further, and The fusion strategy is as follows: First, generate corresponding probability distribution maps for the pseudo labels of round i-1 and round i respectively. and , both represent the probability set of the pixel at position (x, y) on the mask map belonging to each feature category, where Num represents the number of feature categories in the dataset. The two rounds of probability distribution maps are then aligned in the spatial dimension, and the probability values ​​of each pixel belonging to Num feature categories on the two rounds of probability distribution maps are compared. The maximum probability value is taken as the category probability distribution of the pixel. Finally, the pseudo label is corrected according to this category. The pixel value at (x, y);

[0102] Furthermore, the confidence threshold filtering method in step S5 is to set the confidence level below the preset threshold. The pixel points are shielded, and the shielded pixel samples do not participate in the calculation of the current round loss function and the optimization of model parameters, and are retained. and Medium confidence is higher than The pixel samples are used as optimized pseudo labels for the next round of model training;

[0103] The expression of confidence is:

[0104]

[0105] in, yes The confidence of the sample at the pixel, represents the entropy of the probability distribution of the pixel value at (x, y), represents the normalization factor.

[0106] Furthermore, the pseudo-labels corrected for confidence consistency After correcting some erroneous pixels, the essence of this is to select the category with the higher predicted probability on the two probability maps as the category of the pixel. However, the confidence of the selected category may still be low. This is because some ground objects in the remote sensing image pair are difficult to detect and identify due to factors such as occlusion and shadows, and the model's probability values ​​for the ground object category in this area are similar. There are still some noise samples that are misleading for model training. In order to reduce the impact of these noise pixels, a confidence threshold filtering strategy is used to filter Noise pixels in . Updated pseudo labels and its probability distribution to calculate the confidence of each pixel sample;

[0107] S6. Combine the high-confidence pseudo-label data obtained in step S5 with the small amount of annotated bi-temporal remote sensing images obtained in step S1 to form a new training set. Use this new training set to perform semi-supervised training on the initial model model_1 in the semantic change detection model, and supervise the training using a consistency regularized combined loss function to obtain a new model model_2.

[0108] Furthermore, the consistency regularization combined loss function of step S6 is The expression is:

[0109]

[0110] in, is the consistency loss function, is a hyperparameter;

[0111] Use KL divergence loss to calculate the consistency loss function, the calculation formula is:

[0112]

[0113]

[0114] Where M represents the number of unlabeled remote sensing image pairs, represents a pseudo label, Represents the prediction result of the unlabeled image, Num represents the number of changing object categories in the dataset, represents the probability of category j in the pseudo-label probability distribution, Represents the probability of category j in the prediction result probability distribution; Represents the difference between the pseudo-labels with random perturbations and the predicted results of remote sensing images. During model training and parameter optimization, minimizing this difference constrains the model's response to random perturbations and improves model robustness and generalization.

[0115] Furthermore, in the iterative semi-supervised phase, in each iteration, a new training set consisting of labeled data and pseudo-labeled data is generated. Input model_1 / i, where Generate prediction results , and with The loss is calculated using the formula. First, perform random combination enhancement operations, then input the model to generate prediction results . Through the same random combination enhancement operation, and then with Calculate the consistency regularization loss. and tags Apply the same noise and minimize the prediction result and tags The gap in the probability distribution of pixels between the two images is used to impose consistency constraints on the model and enhance the model's ability to resist different disturbances and noises.

[0116] The loss in the semi-supervised stage is determined by the prediction results and pseudo labels Calculated, where Unlabeled image after applying enhanced perturbation Generate pseudo labels through model_1 / i The prediction results generated by the model model_0 / i-1 after confidence consistency correction and confidence threshold filtering, as well as the above combined enhanced perturbations.

[0117] S7. Iterate steps S4 to S6 until a predetermined number of iterations is reached, and finally output model model_N. Model model_N is used to extract time-varying information from multi-temporal satellite images.

[0118] In this example, the initial learning rate is set to 1×10 -3 , until a total of 200 training cycles are completed. Furthermore, the SGD optimizer is used for training to accelerate the model convergence process. This embodiment effectively achieves high-precision recognition of time-varying information under semi-supervision, with its mIoU stably maintained above 70%.

[0119] Example 2:

[0120] Based on Example 1, this example further illustrates the semantic change detection network training process under the semi-supervised strategy as follows:

[0121] A large number of dual-temporal remote sensing images are obtained and divided into training sets and test sets. The semi-supervised training strategy is divided into a supervised part using a small amount of labeled data and an iterative semi-supervised part using a large amount of unlabeled data. The dataset is also divided into two parts, namely (1) labeled dataset ,in represents an image pair, represents the label pair, N represents the dataset size; (2) unlabeled dataset ,in Denotes an image pair, and M denotes the dataset size. Furthermore, M >> N. A training group is extracted from the training set, and each training group is used as a training batch. The semantic change detection network is trained using the training set to obtain a trained semantic change detection model. The trained semantic change detection network is then evaluated using the test set to obtain an evaluation of the semantic change detection model.

[0122] In the supervised part, 20% of the labeled data is used to train the initial model, and the prediction results are calculated and narrowed. and tags The loss between is used to optimize the parameters of the initial model to obtain model_1. The loss function used here is a combined loss function that includes semantic class multivariate cross entropy loss and binary change cross entropy loss.

[0123] In the iterative semi-supervised part, there are four steps. The following is a detailed description of each step: (1) Input 80% of the unlabeled data pairs into model_1 / model_i to obtain the prediction results . (2) Through the pseudo-label optimization strategy, the confidence consistency correction is used to correct the low-confidence prediction results, and the confidence threshold is used to filter out the prediction results with high confidence, and finally the pseudo-label is generated. (3) Combine 20% of the labeled data and their labels and 80% of the unlabeled data and their pseudo labels into a new training set, which can be expressed as , where N + M is the number of all remote sensing image pairs, and then use the new training set Retrain the initial model. By shrinking the prediction results 、 and labels (pseudo labels) 、 The loss between updates the model parameters so that the model fits the semantic change detection task and obtains the model model_i, where i is the number of rounds of model updates. (4) In order to continuously optimize the quality of pseudo labels and the degree of fit of model parameters to the task, repeat steps 1, 2, and 3 until the model completes the specified round of updates. In step 1, the latest updated model model_i replaces model_i-1, and the new model is used to generate prediction results for unlabeled data; in step 2, the pseudo labels of round i-1 are retained to correct the pseudo labels of round i; in step 3, the pseudo labels of round i-1 are supervised. and The loss function is a combination of binary cross entropy loss and semantic multivariate cross entropy loss. and The loss function is the consistency regularization loss.

[0124] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0125] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and components may be substituted with equivalents without departing from the scope of the present application. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of these combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions within the scope of the claims.

Claims

1. A semi-supervised method for extracting time-varying information from multi-temporal satellite images, characterized in that: The steps include: S1. Collect a small amount of labeled bi-temporal remote sensing images and their corresponding semantic change detection label data, collect a large amount of unlabeled bi-temporal remote sensing image data, and preprocess the collected bi-temporal remote sensing images to obtain a preprocessed bi-temporal remote sensing image set. S2. Build a semantic change detection model, including two encoders and decoders, a multi-scale feature fusion module, a bi-temporal feature interaction module, a change detection module, and two semantic segmentation modules. S3. Using the small amount of preprocessed, annotated bi-temporal remote sensing images obtained in step S1, perform fully supervised training on the semantic change detection model constructed in step S2 using a binary change cross entropy loss function and a semantic multivariate cross entropy loss function to obtain the initial model model_1. S4. Inputting the large amount of unlabeled dual-temporal remote sensing image data obtained in step S1 into the initial model model_1 obtained in step S3 for prediction, obtaining semantic change detection prediction results for each pair of images; S5. Optimize the semantic change detection prediction results for each pair of images obtained in step S4 using a pseudo-label optimization strategy, wherein the pseudo-label optimization strategy includes confidence consistency correction and confidence threshold filtering to generate high-confidence pseudo-label data; S6. Combine the high-confidence pseudo-label data obtained in step S5 with the small amount of annotated bi-temporal remote sensing images obtained in step S1 to form a new training set. Use this new training set to perform semi-supervised training on the initial model model_1 in the semantic change detection model, and supervise the training using a consistency regularized combined loss function to obtain a new model model_2. S7. Iterate steps S4 to S6 until a predetermined number of iterations is reached, and finally output model model_N. Model model_N is used to extract time-varying information from multi-temporal satellite images.

2. The semi-supervised method for extracting time-varying information from multi-temporal satellite images according to claim 1, characterized in that: The preprocessing method in step S1 is to match the size of the dual-phase remote sensing image, detect whether the dual-phase remote sensing image is a three-channel image, set the dual-phase remote sensing images of the previous time and the next time with matching size and three channels as an image pair, convert the image pair into a numpy format file, and obtain the preprocessed dual-phase remote sensing image.

3. A semi-supervised method for extracting time-varying information from multi-temporal satellite images according to claim 1 or 2, characterized in that: The specific implementation method of step S2 includes the following steps: S2.

1. The semantic change detection model uses a twin-structure encoder-decoder architecture. The multi-scale feature fusion module is embedded at the end of the encoder-decoder architecture. The output of the multi-scale feature fusion module is connected to a bi-temporal feature interaction module, which is connected to the change detection module and two semantic segmentation modules. S2.

2. The multi-scale feature fusion module is composed of two complementary submodules: the global semantic enhancement submodule and the local edge optimization submodule. The output expression of the global semantic enhancement submodule is: in, is the output of the global semantic enhancer module, represents the feature concatenation operation of the channel dimension, represents the sigmoid activation function, The expansion rate is The dilated convolution operation, represents pixel-by-pixel multiplication, Represents the multi-scale features of the 4th or 3rd layer of the encoder, Represents the multi-scale features of the first or second layer of the encoder, i is a cubic value from 1 to 3; The output expression of the local edge optimization submodule is: in, is the output of the local edge optimization submodule; S2.

3. A dual-temporal feature interaction module is established to achieve collaborative modeling of cross-temporal information through a spatiotemporal coupling mechanism. The core process consists of two stages: obtaining a change weight map and obtaining spatial semantic features. The change weight map of the dual-temporal feature interaction module and the generation process of spatial semantic features with spatiotemporal information are shown in the following formula: in, is the change weight graph, Optimize the features for the local edge at time 1, Optimize features for local edges at time 2, To carry the spatial semantic features of rich spatiotemporal information, represents the convolutional layer operation, Represents a concatenation operation in the channel dimension.

4. The semi-supervised method for extracting time-varying information from multi-temporal satellite images according to claim 3, characterized in that: During the training process of step S3, the loss function is composed of multivariate cross entropy loss and binary change cross entropy loss The composition is expressed as: in, is the loss function for fully supervised training, is the multivariate cross entropy loss at time 1, is the multivariate cross entropy loss at time 2, is the binary change cross entropy loss.

5. The semi-supervised method for extracting time-varying information from multi-temporal satellite images according to claim 4, characterized in that: Step S4 uses the initial model model_1 to make predictions, obtain the semantic change detection prediction results for each pair of images, and use the results as the initial pseudo labels .

6. The semi-supervised method for extracting time-varying information from multi-temporal satellite images according to claim 5, characterized in that: The confidence consistency correction method in step S5 is: For the first round of model_1 prediction of a large amount of unlabeled data, confidence consistency correction is not used, but confidence threshold filtering is performed directly to the initial pseudo label. Pixels with low confidence are shielded; For the update of the i-1th round model model_i, retain the pseudo labels used to train the i-1th round model model_i , while retaining the predicted pseudo labels of model_i for unlabeled data Through fusion and , correct some erroneous pixels in the pseudo-label of the same unlabeled remote sensing image pair and generate the optimized pseudo-label of the i-th round , used for training model model_i+1, fusion and The expression is: in, is the confidence at pixel position (x, y) in the mask image of round i-1 and round i, is the probability distribution of the nth feature generated by the pseudo-label in the i-1th round, is the probability distribution of the nth feature generated by the pseudo-label in the i-th round, is any one of N, indicating the nth type of land feature, is the probability of the nth type of feature at (x, y) in the i-th round of training, is the total number of land feature categories.

7. The semi-supervised method for extracting time-varying information from multi-temporal satellite images according to claim 6, characterized in that: The confidence threshold filtering method in step S5 is to set the confidence level below the preset threshold. The pixel points are shielded, and the shielded pixel samples do not participate in the calculation of the current round loss function and the optimization of model parameters, and are retained. and Medium confidence is higher than The pixel samples are used as optimized pseudo labels for the next round of model training; The expression of confidence is: in, yes The confidence of the sample at the pixel, represents the entropy of the probability distribution of the pixel value at (x, y), represents the normalization factor.

8. The semi-supervised method for extracting time-varying information from multi-temporal satellite images according to claim 7, characterized in that: Consistency regularization combined loss function in step S6 The expression is: in, is the consistency loss function, is a hyperparameter; Use KL divergence loss to calculate the consistency loss function, the calculation formula is: Where M represents the number of unlabeled remote sensing image pairs, represents a pseudo label, Represents the prediction result of the unlabeled image, Num represents the number of changing object categories in the dataset, represents the probability of category j in the pseudo-label probability distribution, Represents the probability of category j in the prediction result probability distribution; Represents the difference between the pseudo label with random perturbations and the predicted results of remote sensing images.

Citation Information

Patent Citations

  • Semi-supervised remote sensing image change detection method and device based on joint learning

    CN117173561A

  • Semi-supervised 3D image segmentation method based on quality-driven cross learning

    CN118397275A

  • Semi-supervised change detection method and device based on visual language model

    CN118447344A

Cited By

  • Cross-scale space-time fusion ground feature classification method based on double-branch architecture

    CN121121529A

  • Satellite image in-orbit change detection method and device, storage medium and electronic equipment

    CN121438011A

  • Multi-source satellite data driven irrigation district ecological hydrological basic model construction method

    CN121542656A

  • Semi-supervised scene change detection method based on edge-guided double alignment

    CN121686383A

  • Satellite same frequency interference detection method based on semi-supervised learning

    CN121984572A