Remote sensing image change detection method based on visual state space model

The visual state space model with frequency domain and spatial-channel attention modules addresses the inefficiencies of existing methods by efficiently extracting global and local features in remote sensing images, improving change detection accuracy and reducing computational demands.

CN120318705APending Publication Date: 2025-07-15DALIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510403414.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing remote sensing image change detection methods are highly complex in computing, inefficient in training and inference when processing high-resolution images, and it is difficult to fully capture global information and local details of the image.

Method used

A twin encoder based on visual state space model is adopted, combined with the frequency domain fusion module and the space-channel attention fusion module, and the feature differences are processed through Fourier transform and the full connection layer to realize global spatiotemporal modeling and capture local details.

Benefits of technology

It realizes efficient and low-complex remote sensing image change detection, which can accurately identify changing areas, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318705A_ABST
    Figure CN120318705A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of deep learning and computer vision, and discloses a remote sensing image change detection method based on a visual state space model, which comprises the following steps of: constructing a twin feature encoder, and extracting multi-scale features of a dual-time-phase image through the visual state space model; features from different time phases are fused through a frequency domain fusion module, and noise filtering is further carried out in a frequency domain through Fourier transform; local change features are enhanced through a space-channel attention fusion module, and accurate decoding is achieved; and through deep supervision, the model is assisted to learn an accurate change mask. Compared with other methods, the method has the advantages that global semantics and local details can be considered under the condition of low calculation complexity, high-frequency noise is effectively filtered by the aid of a frequency domain fusion mechanism, false change caused by the noise is reduced, and change masks with higher quality are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning and computer vision, and relates to a remote sensing image change detection method based on a visual state space model. Background Art

[0002] Remote Sensing Image Change Detection (RSCD) refers to judging whether the ground objects have changed, and where they have changed, based on remote sensing images collected at different times at the same location, and generating corresponding change masks. Remote sensing image change detection is one of the basic tasks in remote sensing image processing, and is also the basis for many ground object monitoring related applications, such as urban development planning, mineral resource exploration, agricultural crop detection, etc. Therefore, remote sensing image change detection has extremely broad application prospects in many fields such as public management, military, agriculture, industry and mining.

[0003] With the development of deep learning technology, remote sensing image change detection algorithms have gradually transitioned from traditional methods based on algebra or pixel-level statistics to methods based on deep learning. These include methods based on convolutional neural networks and methods based on Transformer. The convolutional neural network-based method extracts original image features through convolutional layers, which has the advantages of low computational complexity, small number of parameters, and fast training and inference speed. However, due to the limitations of the convolution operation, this method can only model the relationship between objects in a certain area and it is difficult to fully capture the global information of the image. The Transformer-based method can use the attention mechanism to capture the global spatiotemporal relationship of the image content. However, the quadratic computational complexity of the self-attention mechanism causes this method to face challenges such as high computational complexity, low training and inference efficiency when processing high-resolution remote sensing images. In addition, this method sometimes pays too much attention to global information and ignores some key fine-grained visual clues, which has certain limitations in practical applications.

[0004] Therefore, designing a new remote sensing image change detection model with high computational efficiency, low complexity and efficient aggregation of global information is one of the important challenges that need to be urgently addressed in the current field of remote sensing image change detection. Summary of the invention

[0005] The present invention aims to provide a remote sensing image change detection method based on a visual state space model. By designing a frequency domain fusion module and a spatial-channel attention fusion module, and combining a siamese encoder based on the visual state space model, high-efficiency and high-performance remote sensing image change detection can be achieved. Compared with the method based on Transformer, the present invention can use fewer computing resources to achieve global spatio-temporal modeling and can capture local details of remote sensing images. While ensuring the model accuracy, the dependence on computing power is reduced. The technical solution of the present invention:

[0006] A remote sensing image change detection method based on a visual state space model is as follows:

[0007] Step 1: Collect registered dual-temporal optical remote sensing images as a data set; for each pair of dual-temporal optical remote sensing images, according to the definition of the concept of change in different downstream tasks, perform pixel-by-pixel manual accurate annotation to obtain a change mask;

[0008] Step 2: Randomly allocate the collected and annotated data set to the training set, validation set, and test set at a ratio of 8:1:1;

[0009] Step 3: Perform data processing, uniformly crop the images into slices of the same resolution, and perform normalization processing according to the statistical mean and variance;

[0010] Step 4: Train the change detection model, use the training set to train the remote sensing image change detection model, adopt the cross-entropy loss function for supervision, monitor the change of the loss function during the training process, and end the training after it converges;

[0011] By using the siamese encoder based on the visual state space model as a feature extractor, process the input dual-temporal optical remote sensing images, and respectively extract the basic feature maps of the dual-temporal optical remote sensing images at four different scales; fuse the basic feature maps in dual-temporal, and apply the frequency domain fusion module to map them to the frequency domain through Fourier transform for processing during the fusion process, and input them into the spatial-channel attention fusion module after feature fusion to construct a remote sensing image change detection model;

[0012] The siamese encoder based on the visual state space model includes a convolutional module and a state space module; the convolutional module includes a convolutional layer, a normalization layer, and an activation function, which can be expressed as:

[0013] CBR(x) = ReLU(BN(Conv(x)))

[0014] Among them, Conv(.) is a convolutional operation, BN(.) is a batch normalization layer, and the ReLU activation function can be expressed as:

[0015]

[0016] Among them, x is the input feature of the convolutional module;

[0017] The state space module can be expressed as:

[0018] VSS(x) = FFN(SS2D(x))

[0019] SS2D(x) = x + Linear(LN(SSM(SiLU(DWConv(Linear(x))))

[0020] FFN(x) = x + Linear(LN(x))

[0021] Among them, LN(.) is the normalization layer, Linear(.) is the fully connected layer, SSM(.) is the state space equation, DWConv(.) is the depthwise separable convolution, VSS(.) is the state space module, SS2D(.) is the 2D selective scan layer, FFN(.) is the feedforward neural network, and SiLU is the activation function, which can be expressed as:

[0022]

[0023] For each temporal image, the features extracted by the convolutional module and the features extracted by the state space module at the same scale are fused using the fully connected layer and the features extracted by the state space module to obtain four different-scale features and the preliminary fusion features of two different temporal phases Subsequently, for each scale, the difference between the preliminary fusion features of two different temporal phases is taken and the absolute value is taken. After being processed by the fully connected layer, the double-temporal difference feature F of each scale is obtained; finally, the obtained difference feature F is subjected to the fast Fourier transform, and is processed by the fully connected layer and the activation function in the frequency domain. After being transformed back to the spatial domain through the inverse Fourier transform, it is multiplied by the learnable scale adjustment factor α and added to the original difference feature F to obtain the final output difference feature F'; this process is expressed as:

[0024]

[0025] F' = F + α * IFFT(ReLU(Linear(FFT(F)))

[0026] Among them, [.] represents concatenation on the channels, |.| represents taking the absolute value, and FFT(.), IFFT(.) represent the fast Fourier transform and the inverse Fourier transform respectively;

[0027] Four different-scale differential features F′ are then input into four corresponding-scale spatial-channel attention fusion modules. The first spatial-channel attention fusion module only receives the differential features F′ of the same size, while the subsequent second, third, and fourth spatial-channel attention fusion modules receive both the differential features F′ of the corresponding size and the decoded features output from the previous spatial-channel attention fusion module. Subsequently, the two input features are accumulated and fused, which is expressed as:

[0028]

[0029] where i is the decoder layer number;

[0030] The extracted features are further screened using spatial attention and channel attention to filter out useless redundant information and increase the weight of the representation of key changes. Finally, the fused features are processed using a state space module that matches the same-level feature extraction part in terms of size and space, and this process is expressed as:

[0031]

[0032] where S(.) is spatial attention and CA(.) is channel attention;

[0033] The spatial attention module can be expressed as:

[0034] SA(f) = x ⊙ σ(Linear[AvgPool(f), MaxPool(f)])

[0035] where σ represents the Sigmoid activation function, AvgPool(.), MaxPool(.) represent channel average pooling and channel maximum pooling respectively, and f is the input feature;

[0036] The channel attention module can be expressed as:

[0037] CA(f) = x ⊙ σ(Linear(AvgPool(f)) + Linear(MaxPool(f)))

[0038] Subsequently, the features output from other spatial-channel attention fusion modules will be directly classified through a convolutional layer to obtain prediction change masks of different sizes for deep supervision. The decoded features output from the fourth spatial-channel attention fusion module will be upsampled and then classified through a convolutional layer to obtain the final prediction result, and this process is expressed as:

[0039]

[0040] Among them, Upsample(.) represents 4-fold upsampling;

[0041] During the training process, the cross-entropy loss function is applied to supervise the network model, which can be expressed as:

[0042]

[0043] Among them, n is the number of pixels, Y(i) represents the ground truth of the change mask, and y(i) is the probability that the pixel belongs to a certain class of the mask after Softmax normalization;

[0044] The complete loss function is:

[0045]

[0046] Among them, through simple classification of the intermediate features and calculation of the loss function for deep supervision, the model can more effectively mine deep semantic features during the training stage, thereby enhancing the understanding ability of complex semantic content; in the inference stage, the system only outputs a change mask with the same size as the input image as the final detection result;

[0047] Step 5: Evaluate the prediction accuracy, change detection sensitivity and specificity of the remote sensing image change detection model; if the performance of the remote sensing image change detection model meets the expectations, the training process ends; if not, the hyperparameters in the training of the remote sensing image change detection model, such as the learning rate, batch size, etc., need to be adjusted, and then retrained until the performance meets the requirements;

[0048] Among them, the specific engineering and technical indicators include: accuracy, recall rate, F1 score and intersection over union. The beneficial effects of the present invention:

[0049] (1) The present invention discloses a remote sensing image change detection method based on a visual state space model, which can accurately identify the change regions and objects in dual-temporal remote sensing images. It can assist the application work in multiple fields such as urban planning, disaster prevention and mitigation, and agricultural monitoring, and has good market application prospects.

[0050] (2) The remote sensing image change detection method based on the visual state space model provided by the present invention effectively integrates convolutional neural network, state space model, frequency domain analysis and attention mechanism, and constructs an end-to-end learning framework. It has the ability to extract global semantics and local details, and has a low computational complexity, which can provide a low-cost and high-accuracy basis and support for subsequent practical applications. Description of the Drawings

[0051] Figure 1 is a schematic diagram of the twin encoder structure of the visual state space model of the present invention.

[0052] Figure 2 It is a schematic diagram of the frequency-domain fusion module of the present invention.

[0053] Figure 3 It is a schematic diagram of the spatial-channel attention fusion module of the present invention.

[0054] Figure 4 It is a schematic diagram of the overall structure of the present invention. Detailed implementation manners

[0055] The following further describes the detailed implementation manners of the present invention in combination with the accompanying drawings and technical solutions.

[0056] Embodiment

[0057] A remote sensing image change detection method based on a visual state space model comprises the following steps:

[0058] (1) First, obtain the data required for training.

[0059] Collect registered dual-temporal optical remote sensing images as a data set; for each pair of dual-temporal optical remote sensing images, according to the definition of the concept of change in different downstream tasks, perform pixel-by-pixel manual precise annotation to obtain a change mask, and divide it into a training set, a validation set, and a test set. Uniformly crop the images into image slices with the same resolution and perform normalization processing.

[0060] (2) Construct and implement the proposed Siamese encoder.

[0061] The Siamese feature encoder first uses a convolutional layer to perform preliminary feature extraction on the input remote sensing image, reducing the image resolution from 256x256 to 64x64 while increasing the number of channels to 128, thereby achieving preliminary compression of image information and accelerating the subsequent calculation process. Immediately afterwards, the remote sensing images of each time phase respectively flow through four groups of state space modules and convolutional blocks arranged in parallel. Among them, the resolutions of the four groups of output feature maps are 64x64, 32x32, 16x16, and 8x8 respectively; the numbers of channels are 128, 256, 512, and 1024 respectively; the initial convolutional layer is composed of two convolutional layers with a convolutional kernel size of 2x2 and a stride of 2 stacked together.

[0062] (3) Embed the frequency-domain fusion module to achieve feature fusion with noise suppression.

[0063] The dual-temporal multi-scale features output by the twin encoders will pass through the frequency-domain fusion module. For each scale, the convolutional features and the state-space model features extracted from the corresponding phases will be concatenated along the channels, and the original number of channels will be restored through a fully connected layer. Subsequently, the difference between the features of two different phases will be calculated and the absolute value will be taken to extract the dual-temporal difference features. Then, the difference features will be mapped to the frequency domain through the fast Fourier transform, processed through fully connected layers and activation functions, and finally restored to the spatial domain through the inverse Fourier transform. After multiplying by a learnable scale factor, the difference features will be added to the original difference features to obtain the optimized dual-temporal difference features.

[0064] Specifically, for the four fully connected layers in the frequency-domain fusion module, the input channels are set to 256, 512, 1024, and 2048 in sequence, while the corresponding output channels are 128, 256, 512, and 1024 respectively.

[0065] (4) Utilize the spatial-channel attention fusion module and deep supervision to achieve efficient decoding of the change features.

[0066] The present invention sets four independent spatial-channel attention fusion modules to achieve hierarchical fusion of the features output at different scales. The fused features are sequentially input into the fusion modules at the corresponding levels in the order of increasing resolution. After upsampling, the features output by the first three fusion modules are concatenated with the fused features at the same level and fused through a fully connected layer for subsequent processing. In each level of the fusion module, the input features first filter the effective information through spatial attention and channel attention, and the enhanced features are input into the state-space block for further processing to obtain the final change features.

[0067] The change features corresponding to the scale output by the fusion module will be directly connected to a change detection head to directly perform pixel-wise classification based on the current change features to generate change masks at different scales; the change features output by the fourth fusion module will be upsampled by a factor of 4 to restore the resolution of the original image and generate the final change mask through the change detection head.

[0068] During the model training process, the label image will be downsampled to three sizes of 1 / 4, 1 / 8, and 1 / 16, and the losses of the change masks of different sizes output by the three intermediate layers will be calculated for supervision, and the original label will be used to calculate the loss with the final output change mask for supervision.

Claims

1. A remote sensing image change detection method based on a visual state space model, characterized in that, The steps are as follows: Step 1: Collect registered dual-temporal optical remote sensing images as the dataset; for each pair of dual-temporal optical remote sensing images, according to the definition of the concept of change in different downstream tasks, perform pixel-by-pixel manual precise annotation to obtain a change mask. Step 2: Randomly allocate the collected and annotated dataset into a training set, a validation set, and a test set at a ratio of 8:1:

1. Step 3: Perform data processing, uniformly crop the images into slices with the same resolution, and perform normalization processing according to the statistical mean and variance. Step 4: Train the change detection model, use the training set to train the remote sensing image change detection model, adopt the cross-entropy loss function for supervision, monitor the change of the loss function during the training process, and end the training after it converges. By using the Siamese encoder based on the visual state space model as the feature extractor, process the input dual-temporal optical remote sensing images, and extract the basic feature maps of the dual-temporal optical remote sensing images at four different scales respectively. Fuse the basic feature maps in a dual-temporal manner, and apply the frequency-domain fusion module to map them to the frequency domain through Fourier transform for processing during the fusion process, and input them into the spatial-channel attention fusion module after feature fusion to construct the remote sensing image change detection model. The Siamese encoder based on the visual state space model includes a convolutional module and a state space module; the convolutional module includes a convolutional layer, a normalization layer, and an activation function, which can be expressed as: CBR(x) = ReLU(BN(Conv(x))) where Conv(.) is the convolution operation, BN(.) is the batch normalization layer, and the ReLU activation function can be expressed as: where x is the input feature of the convolutional module. The state space module can be expressed as: VSS(x) = FFN(SS2D(x)) SS2D(x) = x + Linear(LN(SSM(SiLU(DWConv(Linear(x))))) FFN(x) = x + Linear(LN(x)) where LN(.) is the normalization layer, Linear(.) is the fully connected layer, SSM(.) is the state space equation, DWConv(.) is the depthwise separable convolution, VSS(.) is the state space module, SS2D(.) is the 2D selective scanning layer, FFN(.) is the feed-forward neural network, and SiLU is the activation function, which can be expressed as: For each phase image, a fully connected layer is used to extract the features extracted by the convolution module at the same scale. and features extracted by the state-space module Fusion is performed to obtain four different scale features and two preliminary fusion features at different phases. Then, for each scale, the initial fusion features of two different phases are subtracted and the absolute value is taken. After being processed by the fully connected layer, the dual-phase difference feature F of each scale is obtained; finally, the obtained difference feature F is subjected to fast Fourier transform, and processed by the fully connected layer and activation function in the frequency domain. After being transformed to the spatial domain by the inverse Fourier transform, it is multiplied by the learnable scale adjustment factor α and added to the original difference feature F to obtain the final output difference feature F d ; This process is expressed as: F′ = F + α * IFFT(ReLU(Linear(FFT(F))) where [.] represents concatenation on the channel, |.| represents taking the absolute value, and FFT(.), IFFT(.) represent the fast Fourier transform and the inverse Fourier transform respectively; Four different-scale differential features F′ are then input into four corresponding-scale spatial-channel attention fusion modules. The first spatial-channel attention fusion module only receives the differential features F′ of the same size, while the subsequent second, third, and fourth spatial-channel attention fusion modules receive both the differential features F′ of the corresponding size and the decoded features output from the previous spatial-channel attention fusion module. Subsequently, the two input features are accumulated and fused, denoted as: where i is the number of decoder layers; Use spatial attention and channel attention to further screen the extracted features, filter out useless redundant information and increase the weight of the key change representations. Finally, use the state space module that matches the same-level feature extraction part in size and space to process the fused features, and this process is expressed as: where SA(.) is the spatial attention and CA(.) is the channel attention; The spatial attention module can be expressed as: SA(f) = x ⊙ σ(Linear[AvgPool(f), MaxPool(f)]) where σ represents the Sigmoid activation function, AvgPool(.) and MaxPool(.) represent channel average pooling and channel maximum pooling respectively, and f is the input feature; The channel attention module can be expressed as: CA(f) = x ⊙ σ(Linear(AvgPool(f)) + Linear(MaxPool(f))) Subsequently, the features output by other spatial-channel attention fusion modules will be directly classified through a convolutional layer to obtain prediction change masks of different sizes for deep supervision. The decoded features output by the fourth spatial-channel attention fusion module will be upsampled and then classified through a convolutional layer to obtain the final prediction result. This process is expressed as: where Upsample(.) represents 4x upsampling; During the training process, the cross-entropy loss function is applied to supervise the network model, which can be expressed as: where n is the number of pixels, Y(i) represents the ground truth of the change mask, and y(i) is the probability that the pixel belongs to a certain class of the mask after Softmax normalization; The complete loss function is: where, by simply classifying the intermediate features and calculating the loss function for deep supervision, the model can more effectively mine deep semantic features during the training stage, thereby enhancing the understanding of complex semantic content; in the inference stage, the system only outputs a change mask consistent with the size of the input image as the final detection result; Step 5: Evaluate the prediction accuracy, change detection sensitivity and specificity of the remote sensing image change detection model; if the performance of the remote sensing image change detection model meets the expectations, the training process ends; if not, the hyperparameters in the training of the remote sensing image change detection model, such as the learning rate, batch size, etc., need to be adjusted, and then retrained until the performance meets the requirements; where the specific engineering and technical indicators include: accuracy, recall rate, F1 score and intersection over union.

Citation Information

Cited By

  • Image change detection method and device based on state space, equipment and medium

    CN120612496A