Remote sensing image unsupervised change detection method and system for joint learning of time domain difference and consistency

Through the joint learning method of time domain difference and consistency, the coordinated use of contrast learning and the visual basic model VFM, the problems of data dependence and noise interference in unsupervised change detection of high-resolution remote sensing images are solved, and efficient and accurate change detection effect is achieved.

CN120107747APending Publication Date: 2025-06-06Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510163996.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art faces the problems of data dependence and noise interference in unsupervised change detection of high-resolution remote sensing images, especially in environments with high spatial and temporal complexity, and it is difficult to effectively distinguish semantic changes from time domain noise.

Method used

Using the joint learning time domain difference and consistency method, through the collaborative utilization of the basic visual model VFM, training loss function is constructed to drive the model to learn the cross-time domain consistency and correlation of multi-time phase remote sensing images. Specific steps include data augmentation, LoRA fine-tuning, time domain and spatial comparison learning, and grid regularization constraints.

Benefits of technology

The change detection without explicit supervision of high-resolution remote sensing images is realized, which significantly improves the accuracy and efficiency of change detection, and can accurately capture semantic differences in the image under the conditions of simulating time-domain noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107747A_ABST
    Figure CN120107747A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image change detection, in particular to a remote sensing image unsupervised change detection method and system for joint learning of time domain difference and consistency, and the method comprises the steps: constructing a multi-temporal remote sensing image sample through data enhancement transformation; constructing a training loss function based on time domain contrast learning and space contrast learning of consistency constraints, and training the image change detection model based on the training loss function and by using the multi-temporal remote sensing image sample; and inputting the multi-temporal remote sensing image of the target area into the trained image change detection model, and obtaining change information in the multi-temporal remote sensing image of the target area by using the image change detection model. According to the method, difference and consistency among multi-temporal observation are captured in combination with comparative learning, pixel-level semantic representation is obtained in combination with a visual basic model, sparse and compact change mapping learned by a neural network is promoted by utilizing grid sparseness loss, and unsupervised change detection of a high-resolution remote sensing image can be accurately and efficiently realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image change detection, and in particular to a remote sensing image unsupervised change detection method and system for jointly learning temporal difference and consistency. Background Art

[0002] Change detection (CD) in remote sensing is the process of identifying and segmenting changed areas using multi-temporal observations of the same geographic area. In the past decade, great progress has been made in change detection using deep learning (DL) techniques. Cutting-edge methods have achieved F scores of over 90% on various benchmark datasets for change detection. 1 Accuracy. However, most of these DL-based change detection methods require a large amount of high-quality labeled data, while change samples require multi-temporal and differentiated regional observations, which are usually scarce and difficult to collect. Therefore, the deployment of change detection algorithms in practical applications still faces great challenges. Due to the inherent spatiotemporal complexity of the data, unsupervised change detection (UCD) of remote sensing images, especially high-resolution remote sensing images, remains a difficult problem.

[0003] In order to reduce the dependence on training data, more and more research works have focused on unsupervised change detection. However, most of these studies focus on unsupervised change detection of low- and medium-resolution remote sensing images. Unsupervised change detection algorithms for high-resolution remote sensing images are more challenging due to the increased spatial and temporal complexity. Several cases of noise are encountered in change detection of high-resolution remote sensing images, including: i) Spectral changes. This can be attributed to seasonal changes in vegetation or differences in imaging sensors and lighting conditions. ii) Non-significant changes. In change detection applications, only specific changes are of interest, such as building changes in urban management and farmland changes in agricultural monitoring. Certain temporary changes, such as parked vehicles, are usually considered irrelevant noise. iii) Spatial misalignment. This may be caused by changes in imaging angle, optical distortion, or image registration errors. Therefore, deep neural networks (DNNs) face significant challenges in learning to distinguish semantic changes from temporal noise in an unsupervised manner.

[0004] In current research, contrastive learning (CL) and visual foundation models (VFMs) have been identified as two effective methods to alleviate data dependency in change detection. The former mines and exploits the intrinsic semantic consistency of the data, while the latter learns a universal semantic representation in combination with knowledge pre-trained from external data. Using contrastive learning techniques, semi-supervised change detection for high-resolution remote sensing images has made significant progress. However, due to the inherent spatiotemporal complexity of the task, unsupervised change detection for high-resolution remote sensing images using contrastive learning or visual foundation models (VFMs) remains challenging. Among the contrastive learning methods based on change detection, most studies adopt a training framework based on consistency regularization constraints. Although this method significantly enhances the generalization and robustness of feature representation, it still requires a certain proportion of training data. In addition, existing contrastive learning methods mainly focus on temporal similarity embedding and have significant deficiencies in explicit difference modeling. VFM-based change detection usually uses VFMs, such as SegmentAnything Model, to decode change maps using semantic features. Nevertheless, there is a significant domain difference between the training domains of VFMs and remote sensing images, which has an adverse impact on their recognition ability. In addition, accurately interpreting semantic features into change detection results still requires a certain amount of supervised training. Although this approach significantly enhances the generalization and robustness of feature representation, it still requires a certain proportion of training data. Summary of the invention

[0005] To this end, the present invention provides a remote sensing image unsupervised change detection method and system for jointly learning temporal differences and consistency, which solves the unsatisfactory problems in the existing unsupervised change detection of remote sensing images, and realizes change detection of high-resolution remote sensing images without explicit supervision by synergistically utilizing contrastive learning and visual basis model VFM.

[0006] According to the design scheme provided by the present invention, on the one hand, a remote sensing image unsupervised change detection method for jointly learning temporal differences and consistency is provided, comprising:

[0007] Acquire multi-temporal remote sensing image data of the same area, and construct multi-temporal remote sensing image samples through data enhancement transformation, wherein the data enhancement transformation includes spatial transformation and / or spectral transformation for simulating temporal noise between imaging of multi-temporal remote sensing images;

[0008] The LoRA fine-tuning technology is introduced to achieve remote sensing domain adaptation of the visual basic model. While freezing the original model parameters, only the low-rank decomposition matrix is ​​trained to achieve efficient parameter update.

[0009] Based on the consistency constraint, temporal contrast learning and spatial contrast learning are used to construct a training loss function for driving the model to learn the consistency and correlation of multi-temporal remote sensing images across time domains during the training process. Based on the training loss function and using multi-temporal remote sensing image samples, the image change detection model is trained. The image change detection model includes: an encoder for extracting remote sensing image features and a decoder for capturing change information in the image based on the remote sensing image features. Both the encoder and the decoder are constructed based on the visual basis model VFM.

[0010] The multi-temporal remote sensing images of the target area are input into the trained image change detection model, and the image change detection model is used to obtain the change information in the multi-temporal remote sensing images of the target area.

[0011] As an unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency of the present invention, further, multi-temporal remote sensing image samples are constructed through data enhancement transformation, including:

[0012] Constructing a transformation function set for enhancing transformation of remote sensing image data, the transformation function set includes a spatial transformation function for simulating spatial dislocation, imaging degradation and distortion, and a spectral variation function for simulating imaging spectrum and seasonal variation, the spatial transformation function includes a random shift function and a random downsampling function, and the spectral variation function includes an RGB value random drift function and an image cross-domain principal component adaptation transformation function;

[0013] A spatial transformation function and a spectral variation function are randomly selected from a transformation function set to augment the multi-temporal remote sensing image data, and an augmented copy image is generated, so as to construct a multi-temporal remote sensing image sample by fusing the augmented copy image and the multi-temporal remote sensing image data.

[0014] As an unsupervised change detection method for remote sensing images that jointly learns temporal differences and consistency of the present invention, further, a training loss function is constructed to drive the model to learn the consistency and correlation of multi-temporal remote sensing images across time domains during the training process, including:

[0015] Using multi-temporal remote sensing image data and its augmented images, a time domain contrast learning objective function with consistency constraints is constructed based on cosine distance.

[0016] Using multi-temporal remote sensing image data and its augmented images, a spatial contrast learning objective function with consistency constraints is constructed based on the similarity of image spatial positions.

[0017] A remote sensing image sparsity loss function is constructed based on the variation of high and low frequency features of remote sensing images with spatial variation rate and using a grid regularization constraint, wherein the grid regularization constraint is used to evaluate the sparsity of the changing objects in the remote sensing image at each local grid level of the remote sensing image;

[0018] Based on the time domain contrast learning objective function, the space contrast learning objective function and the remote sensing image sparsity loss function, a training loss function for model training is established. The training loss function is expressed as: L = L tri +αL info +βL spa , where L tri is the objective function of time domain contrast learning, L info is the spatial contrast learning objective function, L spa is the sparseness loss function of remote sensing images, and α and β are weight parameters.

[0019] As the unsupervised change detection method of remote sensing images by jointly learning temporal difference and consistency of the present invention, further, two remote sensing images y and The similarity calculation formula of the image space position P is expressed as: Among them, w and h are the spatial dimensions of the remote sensing image.

[0020] As the unsupervised change detection method of remote sensing images by jointly learning temporal difference and consistency of the present invention, further, the grid regularization constraint process includes:

[0021] Set the sparsity threshold and grid size;

[0022] Calculate the average intensity of each grid in the remote sensing image change prediction map based on the sparsity threshold and the grid size, wherein the average intensity is represented by a change probability value;

[0023] The grids are sorted according to their average intensity to select a specified number of grids with the lowest average intensity to calculate the remote sensing image sparsity loss.

[0024] As an unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency of the present invention, further, using an image change detection model to obtain change information in multi-temporal remote sensing images of a target area, the method comprises:

[0025] The encoder in the trained image change detection model is used to extract the semantic features of the multi-temporal remote sensing images of the target area, and the semantic features are mapped to a rough change probability map;

[0026] A decoder in a trained image change detection model is used to generate spatial cues by specifying a response area in a coarse change probability map, and the map is segmented into two sets of multi-temporal masks, so as to obtain a mask map representing a false alarm area by matching and merging the two sets of multi-temporal masks;

[0027] An intersection-and-union analysis is performed between the coarse change probability map and the mask map representing the false alarm area, and the objects in the mask map that match the coarse change probability map are used as the mask of the changed object area. The mask of the changed object area replaces the corresponding area in the coarse change probability map to obtain and output the fine change information in the multi-temporal remote sensing image of the target area.

[0028] As the unsupervised change detection method of remote sensing images by jointly learning temporal difference and consistency of the present invention, further, the rough change probability map is expressed as: y c =σ[-cos(y 1 ,y 2 )*η], where y 1 ,y 2 are the bi-temporal remote sensing images of the target area, σ is the Sigmoid function, and η is the cosine inverse distance scaling factor of the projected bi-temporal semantic features.

[0029] On the other hand, the present invention also provides a remote sensing image unsupervised change detection system for joint learning of temporal difference and consistency, comprising: a sample acquisition module, a model training module and a target detection module, wherein:

[0030] A sample acquisition module is used to acquire multi-temporal remote sensing image data of the same area and construct multi-temporal remote sensing image samples through data enhancement transformation, wherein the data enhancement transformation includes a spatial transformation and / or a spectral transformation for simulating temporal noise between imaging of multi-temporal remote sensing images;

[0031] A model training module is used for constructing a training loss function for driving the model to learn the consistency and correlation of multi-temporal remote sensing images across time domains based on consistency constraint temporal contrast learning and spatial contrast learning during training, and training an image change detection model based on the training loss function and using multi-temporal remote sensing image samples, wherein the image change detection model comprises: an encoder for extracting remote sensing image features and a decoder for capturing change information in the image based on the remote sensing image features, wherein both the encoder and the decoder are constructed based on a visual basis model VFM;

[0032] The target detection module is used to input the multi-temporal remote sensing images of the target area into the trained image change detection model, and use the image change detection model to obtain the change information in the multi-temporal remote sensing images of the target area.

[0033] Beneficial effects of the present invention:

[0034] This paper combines consistency-regularized spatiotemporal contrastive learning (CTC / CSC) with the visual base model for the first time, realizes domain adaptation through LoRA fine-tuning, and innovatively introduces two change detection contrastive learning paradigms, namely consistency-regularized temporal contrast (CTC) and consistency-regularized spatial contrast (CSC), to capture the differences and consistencies between multi-temporal observations respectively, and accurately and efficiently realize unsupervised change detection of high-resolution remote sensing images. Contrastive learning and VFM are complementary technologies. Contrastive learning provides a self-supervised training optimization goal, which is crucial for adapting VFMs to the remote sensing domain and mapping changes. At the same time, VFM can obtain pixel-level semantic representation; under the condition of simulated temporal noise interference, contrastive learning is used to mine the consistency and difference of multi-temporal images in time and space, and the mapping of semantic differences is achieved under the mutual constraint of the two; grid sparsity loss is used to promote neural networks to learn sparse and compact change mappings, and sparsity calculations are performed on the grid rather than pixel scale to avoid training collapse and ensure computational efficiency. It has good application prospects in the field of unsupervised change detection in remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of the unsupervised change detection process of remote sensing images by jointly learning temporal differences and consistency in the embodiment;

[0036] Figure 2 Schematic diagram of noise types for change detection of high-resolution remote sensing images in an embodiment;

[0037] Figure 3 The S2C algorithm framework for unsupervised change detection in the embodiment is shown;

[0038] Figure 4 It is a comparative illustration of the contrastive learning paradigm in change detection in the embodiments;

[0039] Figure 5 Schematic diagram of the quantitative results obtained by the unsupervised change detection method in the embodiment on different benchmark data sets. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below in conjunction with the accompanying drawings and technical solutions.

[0041] To address the challenges of unsupervised change detection in the absence of explicit supervision, three main strategies have been proposed, including feature difference mapping, generative representation, and knowledge transfer. Due to the lack of explicit supervision, feature difference mapping methods usually use pre-trained deep neural networks (DNNs) for feature extraction. However, pre-trained networks are usually not adaptable to inference data and have weak semantic representation. Methods based on generative representation usually use image generation methods to reduce the style differences between multi-temporal observations. This strategy is called generative transcoding. This strategy can achieve change detection between heterogeneous images, such as optical and synthetic aperture radar (SAR) images. Commonly used image generation methods include autoencoders and generative adversarial networks (GANs). Different from generative transcoding, some scholars have introduced a generative framework to iteratively optimize change detection results. The defect of the generative method is the instability of the generated results, which is easy to destroy the semantic distribution of the image, and the transcoded results still do not directly correspond to the semantic changes.

[0042] Contrastive Learning is a self-supervised learning method that extracts semantic consistency from unlabeled data by constructing and comparing positive and negative pairs. A common paradigm of contrastive learning in visual recognition is to introduce noise perturbations of varying strengths during training and learn robust semantic representations by regularizing the constrained DNN. These perturbations can be applied directly to the input image or to the features of the network map.

[0043] In change detection, dual-temporal images of the same and different regions are usually used to construct contrast sample pairs. Some scholars have proposed a simple contrastive learning paradigm for change detection, in which change pairs are constructed by remote sensing clipped at the same and different locations. Other scholars introduced perturbations on the dual-temporal differential features and performed consistency regularized contrastive learning. Other scholars have proposed a contrastive learning framework for change detection under long-term time observation.

[0044] In general, contrastive learning is mainly used for semi-supervised change detection and weakly supervised change detection to enhance the generalization ability of the network and make better use of the scarce supervision signals. Unsupervised change detection based on contrastive learning mainly models and learns temporal similarity embedding, but there is a significant gap in exploring temporal difference embedding.

[0045] In recent years, there has been a research trend to train and build visual foundation models (VFMs), such as CLIP and SegmentAnything Model (SAM), to obtain comprehensive recognition capabilities. VFMs are trained on large datasets to capture general features applicable to a variety of tasks. However, since most of these VFMs are trained in common natural scenes, they exhibit problems such as semantic drift and insufficient recognition capabilities when applied to the recognition of remote sensing scenes.

[0046] Taking into account the spectral and temporal characteristics of multi-temporal remote sensing data, some researchers have developed some basic models in the field of remote sensing, including SpectralGPT and SkySense. However, using these models for change detection still requires adding a neural network module specific to change detection and performing fully supervised fine-tuning (retraining).

[0047] Considering that visual basis models contain implicit knowledge of image content, some recent methods have explored using visual basis to achieve sample-efficient change detection. For example, SAM-CD, the first work to adapt VFMs to the remote sensing domain using semantic latent alignment techniques, achieved better accuracy than strongly supervised change detection methods and demonstrated comparable sample efficiency to several semi-supervised change detection methods.

[0048] Some scholars pointed out that SAM can be used to generate pseudo labels as prompts from blurred change maps. Chen et al. used SAM to implement unsupervised change detection between optical images and map data. Some scholars pointed out that zero-sample change detection is achieved by measuring the similarity of SAM encoded features. Dong et al. used the CLIP model to learn visual language representation to improve the accuracy of change detection.

[0049] Although previous works have utilized VFM for change detection, unsupervised change detection on high-resolution RSI remains a challenging task and the accuracy of state-of-the-art methods is still low.

[0050] To this end, the embodiments of the present invention, see Figure 1 As shown, a remote sensing image unsupervised change detection method for jointly learning temporal differences and consistency is provided, including:

[0051] S101, acquiring multi-temporal remote sensing image data of the same area, and constructing multi-temporal remote sensing image samples through data enhancement transformation, wherein the data enhancement transformation includes spatial transformation and / or spectral transformation for simulating temporal noise between imaging of multi-temporal remote sensing images.

[0052] Specifically, multi-temporal remote sensing image samples are constructed through data enhancement transformation, which can be designed to include:

[0053] Constructing a transformation function set for enhancing transformation of remote sensing image data, wherein the transformation function set includes a spatial transformation function for simulating spatial dislocation, imaging degradation and distortion, and a spectral variation function for simulating imaging spectrum and seasonal variation, wherein the spatial transformation function includes a random shift function and a random downsampling function, and the spectral variation function includes an RGB value random drift function and an image cross-domain PCA adaptive transformation function;

[0054] A spatial transformation function and a spectral variation function are randomly selected from a transformation function set to augment the multi-temporal remote sensing image data, and an augmented copy image is generated, so as to construct a multi-temporal remote sensing image sample by fusing the augmented copy image and the multi-temporal remote sensing image data.

[0055] like Figure 2 As shown in FIG. 1 , several situations in which noise is encountered in change detection of high-resolution remote sensing images, including (a) spectral change, (b) non-significant change, and (c) spatial misalignment. In this embodiment, data enhancement transformation is used to enhance the robustness to temporal noise.

[0056] S102. Based on consistency constraints, temporal contrast learning and spatial contrast learning are used to construct a training loss function for driving the model to learn the consistency and correlation of multi-temporal remote sensing images across time domains during the training process. Based on the training loss function and using multi-temporal remote sensing image samples, an image change detection model is trained. The image transformation detection model includes: an encoder for extracting remote sensing image features and a decoder for capturing change information in the image based on the remote sensing image features. Both the encoder and the decoder are constructed based on the visual fundamental model VFM.

[0057] Among them, the training loss function can be designed to include:

[0058] Using multi-temporal remote sensing image data and its augmented images, a time domain contrast learning objective function with consistency constraints is constructed based on cosine distance.

[0059] Using multi-temporal remote sensing image data and its augmented images, a spatial contrast learning objective function with consistency constraints is constructed based on the similarity of image spatial positions.

[0060] A remote sensing image sparsity loss function is constructed based on the variation of high and low frequency features of remote sensing images with spatial variation rate and using grid regularization constraints, wherein the grid regularization constraints are used to evaluate the sparsity of the changing objects in the remote sensing image at each local grid level of the remote sensing image;

[0061] A training loss function for model training is established based on the time domain contrast learning objective function, the spatial contrast learning objective function and the remote sensing image sparsity loss function.

[0062] In the embodiment of this case, contrastive learning technology is used to jointly utilize temporal consistency and difference features, extract candidate changes through difference learning, and reduce the detection of pseudo changes through consistency learning, thereby achieving change detection without explicit supervision.

[0063] Change detection based on deep learning is essentially learning to project multi-temporal remote sensing images I1, I2 into a binary change map y c Assume fθ is an encoding function with parameters and g is a projection function. This process can be expressed as:

[0064] y 1 =f θ (I 1 ),y 2 =f θ (I 2 ),y c =g(y 1 ,y 2 )

[0065] in is the learned latent semantics, s, h, and w are the semantic and spatial dimensions respectively.

[0066] The S2C algorithm architecture in this embodiment is as follows: Figure 3 As shown in Figure 1, it is an unsupervised change detection framework that first performs self-supervised contrastive learning to extract task-specific semantic features. Subsequently, these semantic representations are further transformed into change detection results through mapping and purification algorithms. This method can also be trained from scratch using ordinary neural networks, but using VFM as a feature encoder will bring better accuracy. Therefore, a VFM with an additional parameter w is used to adapt the VFM in the RS domain. The LoRA optimization method can be used to learn and adjust the VFM parameters, which correspond to the encoding function f θ+w , where θ is the pre-trained VFM weight and w is the weight learned by LoRA technology. VFM can use any existing pre-trained model because its internal parameters are not modified in the architecture of this method. For example, SAM, efficient-SAM, Dino and other models can be used as θ. During the training phase, θ is frozen to retain the pre-trained visual knowledge, while w is the LoRA weight trained using the contrastive learning paradigm to fully extract the semantic features of the time domain.

[0067] The training process is carried out in two contrastive learning paradigms to learn semantic representations related to change detection. In this embodiment, two contrastive learning paradigms for learning difference and consistency representations are introduced. The relevant loss functions are L tri and L info , and grid regularization constraints to learn sparse and compact variation representation, denoted as L spaThe joint training objective function can be expressed as:

[0068] L=L tri +αL info +βL spa

[0069] Where α and β are two weight parameters.

[0070] In the reasoning stage, the semantic latent y1 and y2 are first mapped into a coarse change map, and then the VFM decoder and a matching algorithm based on intersection region detection are used to purify the change regions.

[0071] Consistency Regularization, such as Figure 4 As shown in (a), the neural network f θ For samples under noise interference, we learn to improve the feature representation with robustness and generalization. First, we use weak transformation and strong transformation to augment the image I to obtain two copies. and The distance loss between the two copies is then calculated to ensure a consistent representation that is resistant to noise perturbations. Since this learning paradigm does not explicitly model the difference region, it is often adopted in weakly supervised learning to extend the generalization of change detection learned from limited samples. However, this method cannot be directly applied to unsupervised change detection learning.

[0072] Spatial Contrast (SC), such as Figure 4 As shown in (b), f θ Learn to distinguish bi-temporal image pairs of the same region i and dual-phase image pairs of different regions i and j This drives f θ Learn consistent representations that are independent of temporal changes. Regions with high similarity are considered unchanged, and opposite regions are considered changing. However, this paradigm has some limitations: 1) This method identifies changes through negative distance embedding of similarity rather than through explicit modeling. This often leads to sensitivity to noise. 2) The f θ Focusing on discriminative elements within the region, such as certain edges or corners, instead of effectively utilizing local semantic context.

[0073] In the embodiment of this case, the consistency-regularized temporal contrast learning (CTC) first transforms the remote sensing image I 1 Augment and create a copy Then, I1 As anchor points and positive samples And negative sample I 2 φ(·) simulates the spectral and spatial noise between multi-temporal observations, such as Figure 2 Therefore, f θ Learning to exploit difference representations that are invariant to noise, i.e., semantic variations.

[0074] Figure 3 The left side shows the CTC paradigm in detail. and Bidirectional comparison within . The transformation function φ(·) contains a series of spatial and spectral enhancements that are randomly performed in each training iteration, including random shift, random downsampling, random drift of RGB values, and image cross-domain PCA adaptation to simulate temporal noise between multi-phase imaging. Among them, the first two are spatial changes, which are used to simulate spatial dislocation, imaging degradation and distortion, while the latter two are spectral changes, which are used to simulate imaging spectrum and seasonal changes. These transformations can enhance the robustness of the algorithm to temporal noise.

[0075] In this embodiment, the cosine distance-based triplet training optimization target L can be used. tri , compare the elements within the triple. This is to maintain the inverse embedding with the cosine distance during the inference phase. The calculation can be expressed as follows:

[0076]

[0077] where y 1 ,y 2 f θ From I 1 ,I 2 The semantic features of the code, Yes θ from The encoded semantic features, m is the marginal parameter (usually set to 1) that improves the separation between anchor points and positive points.

[0078] The consistency-regularized spatial contrast learning (CSC) in the present embodiment combines consistency constraints into the typical SC learning paradigm to obtain regional consistency semantic features that are resistant to noise interference. CSC alleviates the inherent sensitivity to noise of the SC paradigm by introducing the transformation φ(·). This transformation, especially the spatial transformation, reduces the dependence on high-frequency spatial details, and thus can force the neural network to learn to utilize local semantic context (such as color and texture patterns, etc.). An additional change is introduced in the calculation of the loss function, that is, the consistency of each spatial position (pixel level) is calculated. Given N pairs of remote sensing images First, φ(·) is applied to each temporal image, resulting in two sets of augmented images. These images are further augmented using f θ+w Encoding, we get 4 sets of features: and Then calculate the co-occurrence relationship between them and obtain two N×N dimensional matrices, such as Figure 4 As shown, the infoNCE loss function is used to effectively train f θ+w To distinguish the original image pairs corresponding to the same region. This loss function is calculated for two phases separately and is expressed as:

[0079]

[0080] Where ⊙ represents the similarity metric function. Instead of pooling spatial features into a single vector for similarity calculation (a common method), it calculates the similarity of each spatial position p, which can be expressed as:

[0081]

[0082] L tri and All are calculated based on cosine similarity. tri Promotes learning of temporal difference features with appearance invariance, This forces the network to learn temporal consistency that is resistant to noise. Therefore, when certain temporal consistency representations are captured in CSC, the different representations of the same region in CTC are suppressed.

[0083] The changing objects are usually sparsely distributed in remote sensing images, and each change term is represented as a compact small target. In contrast, the edges and points in the change prediction graph are often noise. Although there are studies that can promote the training optimization objectives of sparse representation, they usually directly calculate y c Therefore, this method cannot guarantee sparsity because there is a c A suboptimal solution with an additional bias value added.

[0084] In this embodiment, a grid sparsity loss is used, where the sparsity is evaluated at the level of each local grid rather than at the level of each pixel. Considering that the high and low frequency features of remote sensing images vary with spatial resolution, the grid regularization constraint process may include:

[0085] Set the sparsity threshold and grid size;

[0086] Calculate the average intensity of each grid in the remote sensing image change prediction map based on the sparsity threshold and the grid size, wherein the average intensity is represented by a change probability value;

[0087] The grids are sorted according to their average intensity to select a specified number of grids with the lowest average intensity to calculate the remote sensing image sparsity loss.

[0088] First, a sparsity threshold T and a grid size d are defined. Then, the average strength (probability of change value) of each grid g is calculated and sorted, and 1-T grids with the lowest strength (probability of change value) are selected for loss calculation.

[0089] n=wh*(1-T) / d 2 ,

[0090]

[0091] For high-resolution remote sensing images, d = 16 can be set, and for data with sparse changes, T = 0.2 can be set. This objective optimization ensures that grids with a ratio less than 1-T in the change feature map can present higher values, while unimportant changes in other grids are penalized, thereby minimizing their values.

[0092] S103, inputting the multi-temporal remote sensing images of the target area into the trained image change detection model, and using the image change detection model to obtain change information in the multi-temporal remote sensing images of the target area.

[0093] Specifically, using the image change detection model to obtain change information in multi-temporal remote sensing images of the target area may include:

[0094] The encoder in the trained image change detection model is used to extract the semantic features of the multi-temporal remote sensing images of the target area, and the semantic features are mapped to a rough change probability map;

[0095] A decoder in a trained image change detection model is used to generate spatial cues by specifying a response area in a coarse change probability map, and the map is segmented into two sets of multi-temporal masks, so as to obtain a mask map representing a false alarm area by matching and merging the two sets of multi-temporal masks;

[0096] An intersection-and-union analysis is performed between the coarse change probability map and the mask map representing the false alarm area, and the objects in the mask map that match the coarse change probability map are used as the mask of the changed object area. The mask of the changed object area replaces the corresponding area in the coarse change probability map to obtain and output the fine change information in the multi-temporal remote sensing image of the target area.

[0097] Through the model training step, a rough representation of changes in a region can be learned from unlabeled training data. In the inference stage, the main challenge is to accurately map fine-grained changes. In this embodiment, a coarse-to-fine refinement strategy is adopted. First, a rough change probability map y is obtained by projecting the cosine inverse distance of the bi-temporal semantic features.c :

[0098] y c =σ[-cos(y 1 ,y 2 )*η]

[0099] Among them, σ is the Sigmoid function, and η=ln(1 / 0.07) is the scaling factor determined according to the literature.

[0100] Then, the pre-trained VFM decoder g is used γ , using y c Generate spatial cues on medium and high response areas and segment two sets of bi-temporal masks and Considering the logical meaning of "change", M 1 and M 2 The highly overlapping objects in can be inferred as false alarm areas. Therefore, a matching algorithm similar to XOR can be used (denoted as ) to merge M 1 , M 2 , while removing objects with large overlap:

[0101] M 12 =M 1 ⊕M 2

[0102] Further information can be found at c and M 12 Intersection-over-Union (IoU) analysis is performed between the mask generated by VFM and y c The high confidence region in M ​​is used for matching. 12 Zhongneng and y c The matched object can be regarded as the changed feature area, and its precision is higher, so replace y c The pseudo code of the IoU analysis and matching algorithm is provided in detail in Algorithm 1. Using this change mapping algorithm, the coarse predictions obtained from the DNN are refined into detailed change detection results that are consistent with the spatial details present in the high-resolution image.

[0103]

[0104] Furthermore, based on the above method, an embodiment of the present invention also provides an unsupervised change detection system for remote sensing images by jointly learning temporal differences and consistency, comprising: a sample acquisition module, a model training module and a target detection module, wherein:

[0105] A sample acquisition module is used to acquire multi-temporal remote sensing image data of the same area and construct multi-temporal remote sensing image samples through data enhancement transformation, wherein the data enhancement transformation includes a spatial transformation and / or a spectral transformation for simulating temporal noise between imaging of multi-temporal remote sensing images;

[0106] A model training module is used for constructing a training loss function for driving the model to learn the consistency and correlation of multi-temporal remote sensing images across time domains based on consistency constraint temporal contrast learning and spatial contrast learning during training, and training an image change detection model based on the training loss function and using multi-temporal remote sensing image samples, wherein the image change detection model comprises: an encoder for extracting remote sensing image features and a decoder for capturing change information in the image based on the remote sensing image features, wherein both the encoder and the decoder are constructed based on a visual basis model VFM;

[0107] The target detection module is used to input the multi-temporal remote sensing images of the target area into the trained image change detection model, and use the image change detection model to obtain the change information in the multi-temporal remote sensing images of the target area.

[0108] In order to verify the effectiveness of this solution, the following is a further explanation based on experimental data:

[0109] The following experimental data can prove that the accuracy of this solution is greatly improved compared with the current cutting-edge methods.

[0110] Overall, the improvements in F1 are approximately 22%, 9.2%, and 15% on the three benchmark datasets for change detection. Various SOTA methods for unsupervised change detection in RS are compared through experiments. Background technology methods close to the scheme of this case include non-parametric methods based on differential analysis, including CVA-based methods, ISFA, DCVA, DSFA, KPCA-MNet, and SiROC. In addition, a generative method CDRL, an enhancement-based method I3PE, and a recent VFM-based method AnyChange are also compared in the experiment. In order to facilitate a comprehensive evaluation of the accuracy, the fully supervised change detection method SAM-CD is also included in the comparison, representing the SOTA accuracy of supervised change detection.

[0111] like Figure 5 The experimental quantitative results are shown in Figure 2. Compared with the SOTA method, the S2C method in this case has achieved significant improvement in accuracy. The rough prediction results of S2C have an F-value of more than 10% on both CLCD and Levir datasets. 1Improved accuracy. The subsequent precise post-processing using the IoU matching algorithm further enhances the balance between precision Pre and recall Rec. The solution in this case has also been greatly improved on the Levir dataset, with an F1 value increase of more than 16%. This dataset is recognized as a challenging benchmark for unsupervised change detection algorithms because it has changes that focus on buildings rather than generality, and the change samples are very sparse.

[0112] The above experimental data show that this scheme can, under the condition of simulating temporal noise interference, mine the consistency and difference of multi-temporal images in time and space through contrastive learning, realize the mapping of semantic differences under the mutual constraint of the two, and realize unsupervised change detection of high-resolution remote sensing images. It has good application prospects in the field of remote sensing image change detection.

[0113] Unless otherwise specifically stated, the relative steps, numerical expressions and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0114] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0115] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.

[0116] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.

[0117] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency, characterized in that: Include: Acquire multi-temporal remote sensing image data of the same area, and construct multi-temporal remote sensing image samples through data enhancement transformation, wherein the data enhancement transformation includes spatial transformation and / or spectral transformation for simulating temporal noise between imaging of multi-temporal remote sensing images; Based on the consistency constraint, temporal contrast learning and spatial contrast learning are used to construct a training loss function for driving the model to learn the consistency and correlation of multi-temporal remote sensing images across time domains during the training process. Based on the training loss function and using multi-temporal remote sensing image samples, the image change detection model is trained. The image change detection model includes: an encoder for extracting remote sensing image features and a decoder for capturing change information in the image based on the remote sensing image features. Both the encoder and the decoder are constructed based on the visual basis model VFM. The multi-temporal remote sensing images of the target area are input into the trained image change detection model, and the image change detection model is used to obtain the change information in the multi-temporal remote sensing images of the target area.

2. The unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency according to claim 1, characterized in that: Construct multi-temporal remote sensing image samples through data enhancement transformation, including: Constructing a transformation function set for enhancing transformation of remote sensing image data, wherein the transformation function set includes a spatial transformation function for simulating spatial dislocation, imaging degradation and distortion, and a spectral variation function for simulating imaging spectrum and seasonal variation, wherein the spatial transformation function includes a random shift function and a random downsampling function, and the spectral variation function includes an RGB value random drift function and an image cross-domain PCA adaptive transformation function; A spatial transformation function and a spectral variation function are randomly selected from a transformation function set to augment the multi-temporal remote sensing image data, and an augmented copy image is generated, so as to construct a multi-temporal remote sensing image sample by fusing the augmented copy image and the multi-temporal remote sensing image data.

3. The unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency according to claim 1, characterized in that: Construct a training loss function to drive the model to learn the consistency and correlation of multi-temporal remote sensing images across time domains during the training process, including: Using multi-temporal remote sensing image data and its augmented images, a time domain contrast learning objective function with consistency constraints is constructed based on cosine distance. Using multi-temporal remote sensing image data and its augmented images, a spatial contrast learning objective function with consistency constraints is constructed based on the similarity of image spatial positions. A remote sensing image sparsity loss function is constructed based on the variation of high and low frequency features of remote sensing images with spatial variation rate and using a grid regularization constraint, wherein the grid regularization constraint is used to evaluate the sparsity of the changing objects in the remote sensing image at each local grid level of the remote sensing image; Based on the time domain contrast learning objective function, the space contrast learning objective function and the remote sensing image sparsity loss function, a training loss function for model training is established. The training loss function is expressed as: L = L tri +αL info +βL spa , where L tri is the objective function of time domain contrast learning, L info is the spatial contrast learning objective function, L spa is the sparseness loss function of remote sensing images, and α and β are weight parameters.

4. The unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency according to claim 3 is characterized in that: Two remote sensing images y and The similarity calculation formula of the image space position P is expressed as: Among them, w and h are the spatial dimensions of the remote sensing image.

5. The unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency according to claim 3, characterized in that: The grid regularization constraint process includes: Set the sparsity threshold and grid size; Calculate the average intensity of each grid in the remote sensing image change prediction map based on the sparsity threshold and the grid size, wherein the average intensity is represented by a change probability value; The grids are sorted according to their average intensity to select a specified number of grids with the lowest average intensity to calculate the remote sensing image sparsity loss.

6. The unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency according to claim 1, characterized in that: The image change detection model is used to obtain change information in multi-temporal remote sensing images of the target area, including: The encoder in the trained image change detection model is used to extract the semantic features of the multi-temporal remote sensing images of the target area, and the semantic features are mapped to a rough change probability map; A decoder in a trained image change detection model is used to generate spatial cues by specifying a response area in a coarse change probability map, and the map is segmented into two sets of multi-temporal masks, so as to obtain a mask map representing a false alarm area by matching and merging the two sets of multi-temporal masks; An intersection-and-union analysis is performed between the coarse change probability map and the mask map representing the false alarm area, and the objects in the mask map that match the coarse change probability map are used as the mask of the changed object area. The mask of the changed object area replaces the corresponding area in the coarse change probability map to obtain and output the change information in the multi-temporal remote sensing image of the target area.

7. The unsupervised change detection method for remote sensing images by jointly learning temporal differences and consistency according to claim 6, characterized in that: The rough change probability map is represented as: c =σ[-cos(y1,y2)*η], where y1 and y2 are the dual-temporal remote sensing images of the target area, σ is the Sigmoid function, and η is the cosine inverse distance scaling factor of the projected dual-temporal semantic features.

8. An unsupervised change detection system for remote sensing images by jointly learning temporal differences and consistency, characterized in that: It includes: sample acquisition module, model training module and target detection module, among which, A sample acquisition module is used to acquire multi-temporal remote sensing image data of the same area and construct multi-temporal remote sensing image samples through data enhancement transformation, wherein the data enhancement transformation includes a spatial transformation and / or a spectral transformation for simulating temporal noise between imaging of multi-temporal remote sensing images; A model training module is used for constructing a training loss function for driving the model to learn the consistency and correlation of multi-temporal remote sensing images across time domains based on consistency constraint temporal contrast learning and spatial contrast learning during training, and training an image change detection model based on the training loss function and using multi-temporal remote sensing image samples, wherein the image change detection model comprises: an encoder for extracting remote sensing image features and a decoder for capturing change information in the image based on the remote sensing image features, wherein both the encoder and the decoder are constructed based on a visual basis model VFM; The target detection module is used to input the multi-temporal remote sensing images of the target area into the trained image change detection model, and use the image change detection model to obtain the change information in the multi-temporal remote sensing images of the target area.

9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.

Citation Information

Cited By

  • Dual-time-phase remote sensing image change detection method, device, equipment and medium

    CN121033670A

  • Low-altitude remote sensing image AI large model identification training method and system

    CN121191001A

  • General model pre-training and adaptive optimization system and method for remote sensing image

    CN121415186A