Remote sensing image change detection method based on fusion state space model
By using a method based on a fusion state space model in remote sensing image change detection, multi-level features are extracted and fused, the problem of low accuracy caused by pseudo-change interference in the prior art is solved, and more accurate change detection results are achieved.
Patent Information
- Application Number
- CN202510348741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The prior art is difficult to effectively eliminate pseudo-change interference in remote sensing image change detection, resulting in low accuracy of change detection results.
Using a method based on the fusion state space model, hierarchical features including low-level features, intermediate features, deep features and high-level features are extracted through the hierarchical feature extraction unit, and these features are hierarchically spliced and compressed by using the stitching and fusion unit to aggregate remote context information to obtain fusion features with rich information and scale.
This method can better characterize the important information of the input image group, improve the accuracy of change detection, effectively eliminate false change interference, and obtain accurate change detection results.
Smart Images

Figure CN120182827A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and more specifically, to a remote sensing image change detection method based on a fused state space model. Background Art
[0002] Change detection refers to the technique of performing two or more repeated observations of the same location and analyzing the ground object changes in the observed area by comparing the ground object types and ground object ranges of different observations. Remote sensing images, with their wide coverage, efficient data acquisition, and rich information content, have become a key technical means for monitoring surface changes. In particular, high-resolution remote sensing images are widely used in fields such as natural resource monitoring, urban expansion analysis, and agricultural security.
[0003] With the continuous development of remote sensing sensors and image processing technologies, remote sensing change detection methods represented by deep learning are driving the development of remote sensing change detection towards a more refined direction. However, high-resolution remote sensing images exhibit significant spatio-temporal heterogeneity, resulting in a large number of false changes in the detection results.
[0004] Existing methods attempt to enhance the robustness of the model in dealing with confusing backgrounds through data augmentation or attention mechanisms. However, existing methods can only simulate the illumination changes between two-temporal images, and the model still has difficulties in coping with the temporal perturbations brought about by seasonal changes. Recently, some studies have started to focus on the learning of false change features in multi-temporal remote sensing images. However, most of these methods rely on the learning of global features, lack sensitivity to local detail changes, and are difficult to capture seasonal sensitivity when dealing with high-resolution images. Summary of the Invention
[0005] In order to overcome the defect that the existing technology cannot exclude the interference of false changes, resulting in a low accuracy of the change detection result, the present invention provides a remote sensing image change detection method based on a fused state space model that can exclude the interference of false changes.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] A fused state space model, comprising: a hierarchical feature extraction unit, a splicing and fusion unit, and an output unit;
[0008] The hierarchical feature extraction unit is used to receive the image group input to the hierarchical feature extraction unit. The image group includes at least two or more observation images corresponding to different time points at the same location, and at least one observation image in the image group is regarded as the base image, and other observation images are regarded as comparison images. The hierarchical feature extraction unit is also used to extract the hierarchical feature representations of each observation image in the image group. The hierarchical feature representations include low-level features, intermediate features, deep features, and high-level features.
[0009] The splicing and fusion unit is used to splice and compress the hierarchical feature representations of each observation image hierarchically in the channel dimension, and is used to aggregate the long-range context information of the compressed hierarchical feature representations hierarchically to obtain the fusion feature corresponding to each observation image.
[0010] The output module is used to output the predicted probability that the comparison image has changed compared with the base image based on the fusion feature of the base image in the image group and the fusion feature of the comparison image in the image group.
[0011] The present invention also provides a remote sensing image change detection method based on a fusion state space model, including the following steps:
[0012] Obtain the image group to be detected. At least one observation image in the image group to be detected is regarded as the base image, and other observation images are regarded as comparison images.
[0013] Input the image group to be detected into the fusion state space model.
[0014] The fusion state space model outputs the predicted probability that the comparison image in the image group to be detected has changed compared with the base image in the image group to be detected.
[0015] Regard the area where the predicted probability in the comparison image is greater than or equal to the preset threshold as the area where the comparison image has changed compared with the base image.
[0016] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0017] The present invention uses the hierarchical feature extraction unit to extract hierarchical features including low-level features, intermediate features, deep features, and high-level features, and uses the splicing and fusion unit to splice and compress the hierarchical features hierarchically in the channel dimension, and hierarchically aggregate the long-range context information of the compressed hierarchical feature representations to obtain rich-information and rich-scale fusion features, so as to better represent the important information including changes in the input image group, enable the output module to have better change detection ability, and combined with the remote sensing image change detection method, it can further exclude the interference of false changes and obtain accurate change detection results. Description of the Drawings
[0018] Figure 1 The first structural schematic diagram of the fusion state space model proposed in Embodiment 1;
[0019] Figure 2 The second structural schematic diagram of the fusion state space model proposed in Embodiment 1;
[0020] Figure 3 The process schematic diagram of the remote sensing image change detection method based on the fusion state space model proposed in Embodiment 2;
[0021] Figure 4 The schematic diagram of the contrast learning training strategy based on the time-series image extension proposed in Embodiment 3;
[0022] Figure 5 The schematic diagram of the land cover change detection results on the dataset BCDD proposed in Embodiment 3;
[0023] Figure 6 The schematic diagram of the land cover change detection results on the dataset CLCD proposed in Embodiment 3. Detailed implementation manners
[0024] The drawings are only for illustrative purposes and should not be construed as limitations of this patent;
[0025] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;
[0026] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0027] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0028] Embodiment 1
[0029] This embodiment proposes a fusion state space model, Figure 1 which is the first structural schematic diagram of the fusion state space model proposed in this embodiment; Figure 2 which is the second structural schematic diagram of the fusion state space model proposed in this embodiment, Figure 2 showing the structures of the hierarchical feature extraction unit and the splicing and fusion unit, where DownSampling in the hierarchical feature extraction unit represents the downsampling operation.
[0030] As Figure 1 and Figure 2 shown, the fusion state space model of this embodiment includes: a hierarchical feature extraction unit, a splicing and fusion unit, and an output unit;
[0031] The hierarchical feature extraction unit is used to receive the image group input to the hierarchical feature extraction unit. The image group includes at least two or more observation images corresponding to different time points at the same location, and at least one observation image in the image group is regarded as the base image, and the other observation images are regarded as reference images. The hierarchical feature extraction unit is also used to extract the hierarchical feature representation of each observation image in the image group, and the hierarchical feature representation includes low-level features, intermediate features, deep features, and high-level features.
[0032] The splicing and fusion unit is used to splice and compress the hierarchical feature representations of each observation image hierarchically in the channel dimension, and is used to hierarchically aggregate the long-range context information of the compressed hierarchical feature representations to obtain the fusion feature corresponding to each observation image.
[0033] The output module is used to output the prediction probability of the change of the reference image compared with the base image based on the fusion feature of the base image in the image group and the fusion feature of the reference image in the image group.
[0034] In the specific implementation process, the present invention uses the hierarchical feature extraction unit to extract hierarchical features including low-level features, intermediate features, deep features, and high-level features, and uses the splicing and fusion unit to splice and compress the hierarchical features hierarchically in the channel dimension, and hierarchically aggregate the long-range context information of the compressed hierarchical feature representations to obtain rich and multi-scale fusion features, so as to better represent the important information including changes in the input image group, enable the output module to have better change detection ability, and further exclude the interference of false changes in combination with the remote sensing image change detection method to obtain accurate change detection results.
[0035] In an optional embodiment, a VMamba backbone network is arranged in the hierarchical feature extraction unit. The VMamba backbone network includes at least a first stacking module, a second stacking module, a third stacking module, and a fourth stacking module connected in sequence. The first stacking module stacks 2 VSS blocks to initially extract low-level features, the second stacking module stacks 2 VSS blocks to extract intermediate features, the third stacking module stacks 5 VSS blocks to extract deep features, and the fourth stacking module stacks 2 VSS blocks to extract high-level features.
[0036] In an optional embodiment, the splicing and fusion unit includes a first convolutional layer, a second convolutional layer, a first SS1D module, a third convolutional layer, a second SS1D module, a fourth convolutional layer, and a third SS1D module connected in sequence;
[0037] Among them, each convolutional layer is used to reduce the number of channels of the features input to the convolutional layer;
[0038] Each SS1D module includes a two-dimensional image stretching layer, an SSM layer, a feature dimension reduction layer, and a normalization layer connected in sequence; the two-dimensional image stretching layer is used to stretch the received features into one-dimensional feature vectors in two directions, forward and backward in the horizontal direction, or downward and upward in the vertical direction; the SSM layer is used to selectively scan the one-dimensional feature vectors input by the two-dimensional image stretching layer, screen and retain valid information, and optimize the feature expression, the feature dimension reduction layer is used to restore the output of the SSM layer to two-dimensional features by adding and combining, and the normalization layer is used to standardize the output of the feature dimension reduction layer to unify the feature distributions of different scales;
[0039] The high-level features are input into the first convolutional layer, and the first convolutional layer outputs after reducing the dimension of the high-level features. After the high-level features with reduced dimensions and the depth features are concatenated by channels, they are output after passing through the second convolutional layer and the information aggregation of the first SS1D module in sequence; after the output of the first SS1D module and the intermediate features are concatenated by channels, they are output after passing through the third convolutional layer and the information aggregation of the second SS1D module in sequence; after the output of the second SS1D module and the low-level features are concatenated by channels, they are output after passing through the fourth convolutional layer and the information aggregation of the third SS1D module in sequence, and the output of the third SS1D module is the fusion feature.
[0040] As an exemplary illustration, the normalization layer can improve the stability of the training process.
[0041] In an optional embodiment, a classifier and a Softmax layer are provided in the output module. The classifier includes a fifth convolutional layer, a first batch normalization layer, a first activation function layer, and a sixth convolutional layer connected in sequence. The output of the sixth convolutional layer is input into the Softmax layer, and the Softmax layer outputs the prediction probability that the comparison image has changed compared with the base image.
[0042] Embodiment 2:
[0043] This embodiment proposes a remote sensing image change detection method based on the fusion state space model described in Embodiment 1.
[0044] The remote sensing image change detection method based on the fusion state space model includes the following steps:
[0045] S1: Obtain the image group to be detected, and at least one observation image in the image group to be detected is regarded as the base image, and other observation images are regarded as comparison images;
[0046] S2: Input the image group to be detected into the fusion state space model;
[0047] S3: The fused state space model outputs the prediction probability that the control image in the image group to be detected has changed compared with the base image in the image group to be detected.
[0048] S4: Regions in the control image where the prediction probability is greater than or equal to a preset threshold are regarded as regions where the control image has changed compared with the base image.
[0049] In an optional embodiment, before inputting the image group to be detected into the fused state space model, the fused state space model is trained to obtain a trained fused state space model. When inputting the image group to be detected into the fused state space model, the image group to be detected is input into the trained fused state space model;
[0050] The steps of training the fused state space model include:
[0051] Obtain an image change detection data set, where the image change detection data set includes several groups of image groups. Among them, a group of image groups includes several observation images taken at several time points at the same location, and one observation image in each group of image groups is selected as the base image, and the other observation images in the image group are regarded as control images. Each control image is labeled with a change label of the changed region that has changed compared with the base image;
[0052] Perform a random color enhancement operation on the base image in each group of image groups in the image change detection data set to generate a corresponding temporal extension image for each group of image groups;
[0053] Add a global projection head and a dense projection head to the fused state space model;
[0054] Input each group of image groups in the image change detection data set and its corresponding temporal extension image into the fused state space model, and the splicing and fusion unit of the fused state space model outputs the fusion features corresponding to the base image, the control image, and the temporal extension image;
[0055] The output module outputs the prediction probability that the control image has changed compared with the base image based on the fusion features corresponding to the base image and the control image, and constructs a cross-entropy loss function based on the prediction probability and the change label of the control image;
[0056] The global projection head respectively converts the fusion features corresponding to the base image, the control image, and the temporal extension image into global feature vectors corresponding to the base image, the control image, and the temporal extension image;
[0057] The dense projection head respectively converts the fusion features corresponding to the base image, the reference image, and the temporal extension image into dense feature vector groups corresponding to the base image, the reference image, and the temporal extension image respectively;
[0058] Construct a global contrast learning loss function based on the global feature vectors corresponding to the base image, the reference image, and the temporal extension image respectively;
[0059] Construct a dense contrast learning loss function based on the dense feature vector groups corresponding to the base image, the reference image, and the temporal extension image respectively;
[0060] Construct the target loss function of the model based on the cross-entropy loss function, the global contrast learning loss function, and the dense contrast learning loss function;
[0061] During the training process, iteratively solve the target loss function. When the number of iterations reaches the preset value or the target loss function reaches the minimum value, end the training to obtain the trained fusion state space model.
[0062] In an optional embodiment, the global projection head includes a first adaptive average pooling layer, a Dropout layer, a first fully connected layer, a second activation function layer, a second fully connected layer, a second batch normalization layer, and a third activation function layer connected in sequence.
[0063] In an optional embodiment, the dense projection head includes a seventh convolutional layer and a second adaptive average pooling layer connected in sequence.
[0064] As an exemplary illustration, Figure 3 is a schematic flowchart of the remote sensing image change detection method based on the fusion state space model proposed in this embodiment; Figure 3 shows the process of contrast learning training using the fusion state space model, where F0 represents the fusion feature corresponding to the temporal extension image, F1 represents the fusion feature corresponding to the base image, and F2 represents the fusion feature corresponding to the reference image.
[0065] In an optional embodiment, the expression of the target loss function includes:
[0066]
[0067]
[0068] a = 0 and b = 1, or a = 1 and b = 0
[0069]
[0070] In the formula, represents the target loss function, Represents the change detection loss function, Represents the global contrast loss function, Represents the dense contrast loss function, where λ1 represents The control coefficient of, and λ2 represents The control coefficient of; y i Represents the label corresponding to the i-th group of image groups, Represents the change prediction result corresponding to the i-th group of image groups; N represents the total number of image pairs in the image change detection dataset; And Respectively represent the global feature vectors of the temporal extended image, the base image, and the reference image of the i-th group of image groups, And Respectively represent the j-th local feature vector located in the change area extracted from the change label of the temporal extended image, the base image, and the reference image of the i-th group of image groups in the image change detection dataset; n i Represents the total number of local feature vectors of the change area corresponding to a temporal extended image, a base image, or a reference image in the i-th group of images; sim(i, j) represents the cosine similarity between i and j; τ is the temperature parameter; where, for the base images corresponding to N groups of image groups, N enhanced views are generated through data augmentation. The positive sample of any base image is the enhanced image corresponding to this base image, and the negative samples of any base image include the reference image, the remaining (N - 1) base images except this base image and their corresponding (N - 1) enhanced views and (N - 1) reference images, Represents the global feature vector of the k-th negative sample among the aforementioned 3(N - 1) negative samples except the reference image, Represents the m-th local feature vector among the top M negative samples with the largest cosine similarity to the local positive sample among all local negative samples, where M represents the number of selected negative samples.
[0071] This embodiment also proposes a computer device, including a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor executes the steps of the remote sensing image change detection method based on the fusion state space model as described in this embodiment.
[0072] Embodiment 3:
[0073] Based on the fusion state space model proposed in Embodiment 1 and the remote sensing image change detection method based on the fusion state space model proposed in Embodiment 2, this embodiment proposes a specific implementation example:
[0074] Step 1: Obtain a large, publicly available change detection dataset based on high-resolution remote sensing images and preprocess it. Consult existing data sources, organize the sources of existing large, publicly available high-resolution remote sensing change detection datasets, and download relevant data resources. Since there are samples without change labels in the dataset, and change labels are required in the dense contrast learning strategy, the dataset samples are screened to remove samples without change labels and re-divided into a training set, a validation set, and a test set.
[0075] As an exemplary illustration, the specific steps for obtaining a large, publicly available change detection dataset based on high-resolution remote sensing images and preprocessing it are as follows: On websites such as IEEE Xplore and Github, consult relevant papers and literature to organize the sources of existing publicly available high-resolution remote sensing change detection datasets. According to the organized data sources, select a dataset with a relatively large amount of data, better data quality, and an image type closer to the images of the detection task for download and use for pre-training in subsequent experiments. The change label refers to image data that annotates the changed areas (such as buildings, newly added farmland, etc.). Check the obtained high-resolution remote sensing change detection dataset to verify whether the image data and labels in the dataset meet the expected format requirements (such as TIFF, JPEG, PNG, etc.) to ensure that the data can be imported into the training environment without errors. For samples without change labels, manually remove them to ensure that all screened samples have change labels. To ensure data balance, the qualified samples are randomly divided into a training set, a validation set, and a test set in a ratio of 6:2:2 for subsequent model training and prediction.
[0076] Step 2: Design a fusion state space model. Among them, the hierarchical feature extraction unit and the splicing and fusion unit are mainly used to extract multi-scale features from remote sensing images and fuse these features. First, the hierarchical feature extraction unit uses VMamba as the backbone network and extracts low-level to high-level features of the image through stacked VSS modules (Visual State Space), gradually constructing a hierarchical feature representation. Secondly, the splicing and fusion unit replaces the traditional simple splicing operation, splices and compresses feature maps at multiple different stages in the channel dimension, and then aggregates remote context information through the SS1D module (1-Direction Selective Scan) designed in this application, which is beneficial to improving the fusion effect of multi-scale features. Finally, feature maps rich in information and scales are generated, and these feature maps can be used as inputs for subsequent classifiers or contrast learning branches.
[0077] As an illustrative example, the core component of the VMamba backbone network is SS2D (2D-Selective-Scan), which scans from the four corners of the input feature map to achieve efficient computational efficiency while maintaining the global receptive field. The overall network is a multi-stage stacked structure that extracts feature representations at different levels in each stage through multiple levels of VSS blocks. Among them, 2 VSS blocks are stacked in stage 1 to initially extract low-level features; 2 VSS blocks are stacked in stage 2 to extract intermediate-level features; 5 VSS blocks are stacked in stage 3 to extract deep features; 2 VSS blocks are stacked in stage 4 to extract high-level features and form the output. Each stage needs to downsample the input feature map of the previous stage, and the generated multi-scale feature maps provide the basis for subsequent fusion.
[0078] As an illustrative example, the main function of the splicing and fusion unit is to fuse the multi-scale feature maps generated by the VMamba backbone network, thereby generating a feature representation that is rich in information and consistent in scale. The core component, the SS1D module, uses a selective scanning method: first, the image features are stretched into 1D vectors in both the forward and backward directions; then, they are processed by the SSM; finally, the 1D vectors are restored to 2D feature maps through the Merge module to complete further feature fusion and information extraction.
[0079] As an illustrative example, the feature map of stage 4 (high-level features) is concatenated with the feature map of stage 3 (deep features) in the channel dimension. Before concatenation, the feature maps are compressed to 32 channels through 1×1 convolution to reduce computational complexity. The concatenated features are input into the SS1D module for aggregating remote context information. The output feature map of SS1D is then concatenated with the feature map of stage 2 (intermediate-level features), and the above operations are repeated until the fusion with the feature map of stage 1 (low-level features) is completed.
[0080] Step 3: Design a contrastive learning training strategy based on temporal image extension. First, randomly enhance the color of the pre-change remote sensing image to generate temporal extended images, simulating the pseudo-changes caused by changes in lighting conditions and seasonal changes. Input the extended temporal images, pre-change images, and post-change images into the fusion state space model to extract features. Input the extracted feature maps into the Global Projection Head (GPH) and Dense Projection Head (DPH) in the contrastive learning branch to extract global feature vectors and dense feature vector groups respectively. Among them, GPH guides the model to learn the global similarity of time series images to reduce the impact of pseudo-changes on change detection, and DPH focuses on the dense features in remote sensing images, guiding the model to learn the local differences of time series images to enhance the model's sensitivity to local changes. Based on the global feature vectors and dense feature vector groups of temporal images, pre-change images, and post-change images, utilize the characteristics that the semantic information of pre- and post-change images changes while the semantic information of temporal extended images remains unchanged to construct global contrastive learning loss and dense contrastive learning loss (dense contrastive learning loss) respectively. The two cooperate with the change detection loss to maximize the global and local similarities between positive samples and maximize the global and local differences between negative samples, enabling the model to simultaneously capture global changes and local details in the change detection task, improve the ability to identify pseudo-changes, and accurately identify the change area at the pixel level.
[0081] As an exemplary illustration, GPH and DPH in the contrastive learning branch share weights.
[0082] As an exemplary illustration, generate temporal extended images that visually change but the ground objects do not change by randomly transforming the brightness, color saturation, contrast, hue, etc. of the pre-change image, thereby simulating the pseudo-change phenomenon caused by changes in lighting conditions and seasonal changes. The temporal extended images generated by the random color enhancement method participate in the loss calculation as positive samples in the contrastive learning loss.
[0083] As an exemplary illustration, GPH guides the model to learn the global similarity of time series images and improves the model's robustness to pseudo-changes such as lighting and seasonal changes. The feature map output by the fusion state space model is used as the input of the global contrastive learning branch. First, it is processed by an adaptive average pooling layer to generate a global feature vector, and then through two linear projection layers, the high-dimensional feature vector is mapped to a low-dimensional space for subsequent calculation of the contrastive learning loss.
[0084] As an exemplary illustration, DPH, as a supplement to GPH, captures pixel-level change features at the local scale, especially for identifying minor local changes in images. The feature maps output from the fusion state space model are used as the input to the dense contrast learning branch. First, a 1×1 convolution is applied to compress the channel dimension, and then it is processed by an adaptive pooling layer to generate a set of dense feature vectors to represent the dense features.
[0085] As an exemplary illustration, the change detection loss is calculated using the cross-entropy loss function. For each pixel, the predicted change probability of the model is compared with the true change label to calculate the loss value. Figure 4 This is a schematic diagram of the contrast learning training strategy based on the extension of temporal images proposed in this embodiment; Figure 4 shows a schematic diagram of the process of contrast learning training using the global contrast learning loss and the dense contrast learning loss. As Figure 4 shown, the global contrast learning loss is calculated by maximizing the similarity between positive samples and minimizing the similarity between negative samples. Specifically, the cosine similarity of each image pair is calculated and optimized using the formula based on the normalized temperature cross-entropy loss. The dense contrast learning loss is optimized by calculating the cosine similarity between dense feature vectors, and the most relevant negative samples are selected during the calculation process. This method not only greatly reduces the computational amount but also improves the sensitivity of the model to dense feature differences. The fusion objective function is the sum of the change detection loss and the contrast learning loss, and the contrast learning loss is the sum of the global contrast learning loss and the dense contrast learning loss multiplied by the weight coefficient λ.
[0086] As an exemplary illustration, after the model extracts the feature maps of the two-temporal images, they are concatenated in the channel dimension and input into a simple classifier to output the change results. Among them, the change classifier consists of a convolutional layer, a BN layer, a ReLU activation layer, and a convolutional layer. For the change detection results, the standard binary cross-entropy loss function is used for calculation and optimization. During the training process, the results output by the classifier are processed by Softmax to obtain the change probability, and then the loss is calculated according to the cross-entropy loss formula and the change label. This process can be expressed as:
[0087]
[0088] In the formula, represents the change detection loss function, y i represents the label of the i-th pixel of the image, represents the predicted probability of the i-th pixel of the image.
[0089] As an illustrative example, the core of the contrastive learning loss lies in maximizing the similarity between positive samples and minimizing the similarity between negative samples. Given a batch size of N, N augmented views are generated through data augmentation, resulting in a total of 2N samples. Each sample uses the augmented view as the positive sample, and the remaining 2(N - 1) images as negative samples.
[0090] As an illustrative example, in the global contrastive learning loss of the fusion state space model, the positive sample pair is the original image I1 at time T1 and I0 obtained by temporal extension. The negative samples are I1 and the remote sensing images from other imaging locations in the same batch and their temporally extended views I2. Different from contrastive learning methods such as SimCLR, there are ground object changes between the dual-temporal remote sensing images I1 and I2. Therefore, in Global CL, negative sample pairs (I0, I2) and (I1, I2) can also be constructed. So for image I0 or I1, there is 1 positive sample and 3(N - 1)+1 negative samples. The features F0, F1, and F2 output by the model, after passing through the global projection head, obtain global feature vectors where d is the dimension of the feature vector, which is set to 128 in the experiment. According to the defined positive and negative samples, the loss is calculated according to the following formula:
[0091]
[0092] a = 0 and b = 1, or a = 1 and b = 0
[0093]
[0094] In the formula, represents the global contrast loss function, i represents the three-phase remote sensing images of the i-th group in the batch, i ∈ N. represents the global feature vector of the k-th negative sample among the aforementioned 3(N - 1) negative samples excluding the control image. τ is the temperature, which is set to 0.7 in the experiment.
[0095] As an illustrative example, the dense contrastive learning loss is based on the global contrastive learning, extracting multiple dense feature vectors to form a dense feature vector group. The loss function of the dense contrastive learning loss (Dense CL) is constructed based on the dense feature vector (local feature vector) group. In remote sensing images, changes and non-changes usually occur only in local areas, resulting in the possible coexistence of positive and negative samples among the dense feature vector groups, making it difficult to construct the loss with a unified formula. Therefore, Dense CL uses change labels to screen the dense feature vector groups and flexibly selects positive and negative samples according to the occurrence of changes.
[0096] Specifically, the dense projection head outputs a dense feature vector group Where S is the size of the dense feature vector group, which is set to 9 in CLCDNet. d is the dimension of the feature, which is consistent with the global feature vector and is set to 128. The change labels are downsampled to obtain a mask of size S×S, and this mask is used to extract the dense feature vectors for calculating the dense contrast learning loss. According to the mask, n i dense feature vectors are extracted, with the corresponding dense feature vectors in I0 as the positive samples, and the dense feature vectors not extracted by the mask as the negative samples. i represents the i-th group of images in the batch, and j represents the j-th local feature vector extracted by the mask from this group of images. At the same time, the negative samples also include 3*(S*S - n i ) dense feature vectors not extracted by the mask, and 3*(N - 1)*S*S dense feature vectors of other images in the same batch. Compared with global contrast learning, the number of negative samples in dense contrast learning is larger, which not only places higher requirements on the video memory but also reduces the training efficiency. Therefore, in CLCDNet, among the 3*(S*S - n)+3*(N - 1)*S*S negative samples, the top M dense feature vectors with the largest cosine similarity are selected as the negative samples, so as to reduce the video memory, improve the training efficiency, and enhance the model's discriminative ability for positive and negative samples at the same time. The calculation formula of the dense contrast loss can be expressed as:
[0097]
[0098] where, represents the m-th local feature vector among the top M negative samples with the largest cosine similarity to the local positive sample among all local negative samples.
[0099] Finally, the fusion state space model will be optimized through a hybrid loss function. Specifically, it consists of three parts, including the change detection loss the global contrast learning loss and the dense contrast learning loss Therefore, the objective loss function of the fusion state space model can be expressed as:
[0100]
[0101] As an exemplary illustration, λ1 and λ2 are both set to 0.05 in specific experiments.
[0102] Step 4: Based on the publicly available high-resolution remote sensing change detection dataset, train the proposed change detection model.
[0103] First, pre-train VMamba using the publicly available ImageNet-1K dataset; then, load the pre-trained VMamba to initialize the parameters of the model; after that, select appropriate parameters and train the entire model on the high-resolution remote sensing change detection dataset; finally, obtain a change detection model that can distinguish pseudo-changes and real changes.
[0104] As an illustrative example, the proposed model and all related experiments were implemented in PyTorch and executed on a GeForce RTX 3090. VMamba was initially pre-trained on the ImageNet-1K dataset. During the training process based on the contrastive learning strategy, the number of channels of the feature map output by the network encoder was set to 32, and the dimensionality of the feature vectors of the global contrastive learning branch and the dense contrastive learning branch was 128. The temperature scaling factor in the global contrastive learning loss and the dense contrastive learning loss was set to 0.7. The size of the feature vector group output by the dense contrastive learning was 9×9, and the number M of negative samples selected in the dense contrastive learning loss was 50. The weight coefficient λ of the contrastive learning loss was 0.05. All models were trained using the Adam optimizer with an initial learning rate of 10 -4 , and the training period was 100. According to the test experiments on different datasets, the test set in the dataset was input into the trained deep learning model to obtain the corresponding image change map.
[0105] As an illustrative example, Figure 5 FIG. shows the schematic diagram of the land cover change detection results on the dataset BCDD proposed in this embodiment. The full English name of the dataset BCDD is: Building Change Detection Dataset, Figure 6 FIG. shows the schematic diagram of the land cover change detection results on the dataset CLCD proposed in this embodiment. The full English name of the dataset CLCD is: CropLand Change Dection(CLCD)Dataset, Figure 5 and Figure 6 in which (a) is the image at time I1; (b) is the image at time I2; (c) is the real change label; (d) is the change detection result; Figure 5 and Figure 6 show the accurate remote sensing image change detection results output by the fusion state space model.
[0106] The beneficial effects of the present invention mainly include: the proposed model is applicable to various scenarios such as natural resource monitoring, farmland management, and urban expansion analysis. Its lightweight design and efficient training strategy enable the model to have good practicability, providing technical support for future diverse remote sensing change detection tasks. The fusion state space model can extract multi-level feature representations, and through multi-scale feature fusion and aggregation of long-range context information, provides a more distinguishable feature map for change detection, thereby enhancing the adaptability and robustness of the model in complex scenarios. The contrastive learning training strategy based on the extension of temporal images effectively reduces the influence of pseudo-changes caused by illumination and seasonal changes, significantly improves the accuracy of high-resolution remote sensing change detection, and ensures the accurate capture of real changes.
[0107] The same or similar reference numerals correspond to the same or similar components;
[0108] The terms describing the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0109] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A fusion state space model, characterized in that: include: Hierarchical feature extraction unit, splicing and fusion unit and output unit; The hierarchical feature extraction unit is used to receive an image group input into the hierarchical feature extraction unit, the image group at least includes observation images corresponding to two or more time points of the same location, and at least one observation image of the image group is regarded as a basic image, and the other observation images are regarded as reference images; and is used to extract a hierarchical feature representation of each observation image of the image group, the hierarchical feature representation including low-level features, mid-level features, deep features and high-level features; The splicing and fusion unit is used to splice and compress the hierarchical feature representation of each observation image hierarchically in the channel dimension, and to hierarchically aggregate the remote context information of the compressed hierarchical feature representation to obtain the fusion feature corresponding to each observation image; The output module is used to output the predicted probability that the control image has changed compared with the basic image based on the fusion features of the basic image in the image group and the fusion features of the control image in the image group.
2. The fusion state space model according to claim 1, characterized in that: A VMamba backbone network is provided in the hierarchical feature extraction unit, and the VMamba backbone network includes at least a first stacking module, a second stacking module, a third stacking module and a fourth stacking module connected in sequence, the first stacking module stacks 2 VSS blocks to preliminarily extract low-level features, the second stacking module stacks 2 VSS blocks to extract intermediate features, the third stacking module stacks 5 VSS blocks to extract deep features, and the fourth stacking module stacks 2 VSS blocks to extract high-level features.
3. The fusion state space model according to claim 1, characterized in that: The splicing and fusion unit includes a first convolution layer, a second convolution layer, a first SS1D module, a third convolution layer, a second SS1D module, a fourth convolution layer and a third SS1D module connected in sequence; Among them, each convolution layer is used to reduce the number of channels of the features of the input convolution layer; Each SS1D module includes a two-dimensional image stretching layer, an SSM layer, a feature dimension restoration layer and a normalization layer connected in sequence; the two-dimensional image stretching layer is used to stretch the received features into a one-dimensional feature vector in the forward and reverse directions in the horizontal direction, or in the downward and upward directions in the vertical direction; the SSM layer is used to selectively scan the one-dimensional feature vector input by the two-dimensional image stretching layer, filter and retain effective information, and optimize the feature expression; the feature dimension restoration layer is used to restore the output of the SSM layer to a two-dimensional feature by addition and merging; the normalization layer is used to standardize the output of the feature dimension restoration layer and unify the feature distribution of different scales; The high-level features are input into the first convolution layer, and the first convolution layer reduces the dimension of the high-level features and outputs them. The high-level features after dimension reduction are spliced with the deep features by channels, and then are aggregated through the second convolution layer and the first SS1D module in sequence and then output; the output of the first SS1D module is spliced with the intermediate features by channels, and then are aggregated through the third convolution layer and the second SS1D module in sequence and then output; the output of the second SS1D module is spliced with the low-level features by channels, and then are aggregated through the fourth convolution layer and the third SS1D module in sequence and then output, and the output of the third SS1D module is the fused feature.
4. The fusion state space model according to claim 1, characterized in that: The output module is provided with a classifier and a Softmax layer, wherein the classifier includes a fifth convolutional layer, a first batch normalization layer, a first activation function layer and a sixth convolutional layer connected in sequence, the output of the sixth convolutional layer is input into the Softmax layer, and the Softmax layer outputs the predicted probability that the control image has changed compared with the base image.
5. A remote sensing image change detection method based on a fusion state space model, characterized in that: The following steps are involved: Acquire an image group to be detected, wherein at least one observed image in the image group to be detected is regarded as a basic image, and other observed images are regarded as reference images; Inputting the image group to be detected into the fusion state space model; The fusion state space model outputs a predicted probability that the control image in the image group to be detected has changed compared with the basic image in the image group to be detected; The region in the control image where the prediction probability is greater than or equal to a preset threshold is regarded as the region where the control image has changed compared with the basic image.
6. The remote sensing image change detection method based on fusion state space model according to claim 5 is characterized in that: Before inputting the image group to be detected into the fusion state space model, the fusion state space model is trained to obtain a trained fusion state space model, and when inputting the image group to be detected into the fusion state space model, the image group to be detected is input into the trained fusion state space model; The steps of training the fusion state space model include: Acquire an image change detection dataset, wherein the image change detection dataset includes a plurality of image groups, wherein one image group includes a plurality of observation images taken at a plurality of time points at the same location, and one observation image is selected from each image group as a basic image, and other observation images in the image group are regarded as reference images, wherein each reference image is annotated with a change label of a change region that has changed compared with the basic image; Performing a random color enhancement operation on the basic images in each image group of the image change detection data set to generate a time-series extended image corresponding to each image group; Adding a global projection head and a dense projection head to the fused state space model; Input each image group of the image change detection data set and its corresponding time-series extended image into the fusion state space model, and the splicing fusion unit of the fusion state space model outputs the fusion features corresponding to the basic image, the control image and the time-series extended image; The output module outputs the predicted probability that the control image has changed compared with the base image based on the fusion features corresponding to the base image and the control image, and constructs a cross entropy loss function based on the predicted probability and the change label of the control image; The global projection head converts the fusion features corresponding to the basic image, the control image and the time-series extended image into global feature vectors corresponding to the basic image, the control image and the time-series extended image respectively; The dense projection head converts the fusion features corresponding to the basic image, the control image and the time-series extended image into dense feature vector groups corresponding to the basic image, the control image and the time-series extended image respectively; Constructing a global contrast learning loss function based on the global feature vectors corresponding to the basic image, the control image and the time-series extended image; Constructing a dense contrast learning loss function based on the dense feature vector groups corresponding to the basic image, the control image and the time-series extended image; Constructing a target loss function of the model based on the cross entropy loss function, the global contrastive learning loss function and the dense contrastive learning loss function; During the training process, the target loss function is solved iteratively. When the number of iterations reaches a preset value or the target loss function reaches a minimum value, the training is terminated to obtain a trained fusion state space model.
7. The remote sensing image change detection method based on fusion state space model according to claim 6 is characterized in that: The global projection head includes a first adaptive average pooling layer, a Dropout layer, a first fully connected layer, a second activation function layer, a second fully connected layer, a second batch normalization layer and a third activation function layer which are connected in sequence.
8. The remote sensing image change detection method based on fusion state space model according to claim 6 is characterized in that: The dense projection head includes a seventh convolutional layer and a second adaptive average pooling layer connected in sequence.
9. The remote sensing image change detection method based on fusion state space model according to any one of claims 6 to 8, characterized in that: The expression of the objective loss function includes: a=0 and b=1, or a=1 and b=0 In the formula, represents the target loss function, represents the change detection loss function, represents the global contrast loss function, represents the dense contrast loss function, λ1 represents The control coefficient of The control coefficient of i represents the label corresponding to the i-th group of images, represents the change prediction result corresponding to the i-th image group; N represents the total number of image pairs in the image change detection dataset; and Respectively represent the global feature vectors of the time-series extended image, basic image and control image of the i-th image group, and Respectively represent the jth local feature vector located in the change area extracted by the change label of the time-series extended image, the basic image and the control image of the i-th image group of the image change detection data set; n i represents the total number of local feature vectors of the change area corresponding to a time-series extended image, basic image or control image in the i-th group of images; sim(i,j) represents the cosine similarity between i and j; τ temperature parameter; wherein, for the basic images corresponding to the N groups of image groups, N enhanced views are generated by data enhancement, and the positive sample of any basic image is the enhanced image corresponding to the basic image, and the negative sample of any basic image includes the control image, the remaining (N-1) basic images except the basic image and their corresponding (N-1) enhanced views and (N-1) control images, represents the global feature vector of the kth negative sample among the aforementioned 3(N-1) negative samples excluding the control image, It represents the mth local feature vector among the first M negative samples with the largest cosine similarity with the local positive samples among all local negative samples, and M represents the number of selected negative samples.
10. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, characterized in that: When the computer-readable instructions are executed by the processor, the processor executes the steps of the remote sensing image change detection method based on the fusion state space model as claimed in any one of claims 5 to 9.
Citation Information
Patent Citations
Two-stage high-resolution remote sensing image change detection method in technical field of remote sensing
CN110263705A
Road network change detection method and device, model training method, equipment and medium
CN113807198A
Attention mechanism and Mama combined farmland non-agrochemical detection method
CN119169478A
Remote sensing image change detection method and device based on iterative Mama architecture
CN119205638A
Remote sensing image semantic change detection method and device based on Mamba model
CN119580258A