Remote sensing change detection method based on space-time perception and redundancy suppression

By introducing spatiotemporal perception and redundancy suppression technologies in remote sensing change detection, TAFIM and multi-stage decoder capture the spatiotemporal relationship of the dual-time phase features and suppress redundant information, the problems of information loss and time-dependent insensitivity in traditional methods when dealing with the interaction of dual-time phase features are solved, and higher detection accuracy and robustness are achieved.

CN119992333APending Publication Date: 2025-05-13XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108650.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional remote sensing image change detection methods tend to lose small target or edge information when processing bi-time phase feature interactions, and are insensitive to time dependence, resulting in a decrease in accuracy.

Method used

Using a remote sensing change detection method based on space-time perception and redundancy suppression, a time-sequence-aware feature interaction module (TAFIM) and a multi-stage decoder are constructed, combining an extended long and short-term memory network (LSTM) and a multi-channel feature optimization mechanism (MCFOM) to capture the spatiotemporal relationship of the two-time phase features and suppress redundant information.

Benefits of technology

Improve the accuracy and robustness of remote sensing change detection, and enhance the sensitivity to changing areas, especially in areas with smaller changes or complex details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992333A_ABST
    Figure CN119992333A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing change detection method based on space-time perception and redundancy suppression, and belongs to the field of remote sensing image change detection.The method comprises the specific implementation steps that S1, a remote sensing image data set is collected, and it is ensured that each image has a pixel-level label; s2, constructing a remote sensing change detection network architecture and features based on space-time perception and redundancy suppression; s3, adding a time sequence sensing feature interaction module behind the last layer of extracted features of the backbone network; s4, designing a multi-stage decoder; s5, obtaining a change detection result of the test image; according to the method, the characteristics of parallelism and time sequence modeling capability of the extended long and short-term memory network are utilized to effectively capture time attributes among change detection data, and modeling is carried out on spatio-temporal information. And meanwhile, redundant information is suppressed layer by layer by utilizing multi-level decoding, so that the accuracy and robustness of the detection method are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image change detection, and in particular to a remote sensing change detection method based on spatiotemporal perception and redundancy suppression. Background Art

[0002] In recent years, many researchers have been conducting research on change detection methods for remote sensing images. These methods can be roughly divided into two categories: change detection methods based on traditional methods and change detection methods based on deep learning. According to the different research objects, traditional change detection algorithms can be divided into two categories: pixel-based change detection algorithms and object-based change detection algorithms. Among them, pixel-based change detection methods ignore the relationship between adjacent pixels in remote sensing images, resulting in large detection noise and large errors, and cannot meet the accuracy requirements. In object-based change detection methods, the quality of image preprocessing greatly affects the accuracy of the change detection algorithm. At the same time, these traditional methods also have many limitations, such as sensitivity to noise, insufficient information utilization, the need for manual feature extraction, poor generalization, and difficulty in processing high-dimensional data. In addition, traditional change detection methods require a lot of manual intervention and are very susceptible to the influence of light intensity and sensors, resulting in poor robustness. When the changes become more complex, traditional change detection methods can no longer meet the requirements in terms of accuracy. Therefore, change detection methods based on deep learning have gradually become mainstream. At present, change detection methods based on deep learning have achieved many breakthrough results and achieved good performance, but most methods usually use splicing or subtraction to fuse change features during the interaction of dual-phase features, which easily leads to the loss of small targets or edge information during the interaction of dual-phase features. At the same time, the time dependency of dual-phase images is ignored, resulting in insensitivity to detection targets that change over time. In addition, after feature fusion, different channels may contain repeated or highly correlated redundant features. This will also affect the model's sensitivity to change areas, especially in areas with small changes or complex detail information. This leads to a decrease in accuracy. Summary of the invention

[0003] The present invention intends to provide a remote sensing change detection method based on spatiotemporal perception and redundancy suppression to solve the problems raised in the above background technology.

[0004] In order to achieve the above object, the present invention provides the following technical solutions:

[0005] A remote sensing change detection method based on spatiotemporal perception and redundancy suppression, the specific implementation steps include:

[0006] S1: Collect remote sensing image datasets and ensure that each image has pixel-level annotations; divide the dataset into training set, validation set, and test set for use in different stages, and cut the size of the dataset photos into 256×256;

[0007] S2: Build a remote sensing change detection network architecture and feature extractor based on spatiotemporal perception and redundancy suppression, the feature extractor is ResNet50; use the pre-trained ResNet50 to extract features from remote sensing dual-phase images, aggregate these extracted features along the channel dimension, and obtain new features, so that the network can learn more compact and effective feature representation;

[0008] S3: A temporal-aware feature interaction module, or TAFIM, is added after the last layer of features extracted by the backbone network. It aims to integrate important information in space and time to comprehensively and effectively capture and process the complex changing features in the dual-temporal features, and add the aggregated shallow features to the mixed spatiotemporal features to prevent the loss of shallow texture details.

[0009] S4: Design a multi-stage decoder with dual decoding branches. The decoder uses a multi-channel feature optimization mechanism, MCFOM, to process complementary information from different upsampling methods to obtain redundantly suppressed features, and then predicts the change map with fine details. The multi-layer prediction map and loss function output by the multi-stage decoder are used for deep supervision to learn richer information.

[0010] S5: Save the best model weights in the training process, use the test set for testing, and obtain the change detection result of the test image.

[0011] Preferably, the remote sensing change detection network architecture based on spatiotemporal perception and redundancy suppression proposed in step S2 can be divided into three parts, including feature extraction, feature interaction and multi-level decoding;

[0012] The specific steps include:

[0013] S21: The dual-phase image is input into the weight-sharing pre-trained twin network, i.e., ResNet50, to extract features

[0014] S22: Aggregate the extracted feature maps in the channel dimension to obtain new This process can be expressed as:

[0015]

[0016] In the above formula, Conv1×1(·) represents a 1×1 convolutional layer with batch normalization, so that the network can learn a more concise and effective feature representation;

[0017] S23: The aggregated features Input to TAFIM to capture the spatiotemporal relationship between the two-phase features and reduce the impact of environmental changes on feature matching, which can be expressed as:

[0018]

[0019] In the above formula, T(·) represents the function of TAFIM, and pi represents the mixed feature with spatiotemporal information;

[0020] S24: Integrate the fused initial features into the dual-phase mixed features with spatiotemporal information to prevent the loss of texture information and edge information and enhance the robustness of the model, namely:

[0021]

[0022] In the above formula, Cat(·) represents the cascade operation along the channel dimension, and Mi represents the spatiotemporal hybrid feature that integrates shallow detail information;

[0023] Preferably, in step S3, it is proposed to design TAFIM after the last layer of features extracted by the backbone network, integrate important information in space and time, and comprehensively and efficiently capture and process complex change features in dual-phase features;

[0024] The specific steps include:

[0025] S31: Set the bi-phase input features to X, Y∈R C×H×W , and input it into TAFIM to capture the spatiotemporal information between the two temporal phases.

[0026] S32: Flatten X and Y and perform layer normalization operation so that each pixel can be processed independently and stably, and the result is:

[0027] X n ,Y n ∈R S×C

[0028] In the above formula, S is the product of H and W.

[0029] S33: Use causal convolution to perceive changes and dynamic processes in remote sensing images, and use SiLU function to introduce nonlinearity to enhance the expressive power of the model.

[0030] S34: Map the input tensor to a high-dimensional space to obtain the query (Q), key (K) and value (V) corresponding to the bi-phase feature. This process can be expressed by the mathematical formula:

[0031]

[0032] In the above formula, NH is the number of heads, d represents the dimension of each head, C represents the input dimension, S represents the length of the pixel sequence, and Rea represents the rearrangement operation;

[0033] S35: For each head, perform linear transformation to obtain:

[0034]

[0035] at the same time,

[0036]

[0037] In the above formula, WV, WQ, WV represent the weight matrices corresponding to V, Q, and K respectively. represents the inverse operation of the rearrangement operation, b represents the bias vector, and Conv1d(·) represents the causal convolution operation;

[0038] S36: Combine the cross attention idea to realize the interactive operation of dual-phase features in mLSTM, match QX from the X branch with KY, VY from the Y branch, and obtain the pre-activation calculation of the input gate and the forget gate by simply concatenating QX, KY, VY. This process can be expressed as:

[0039] input gate =[Q X ,K Y ,V Y ]

[0040] g in =Gate_in(input_gate)

[0041] g forget =Gate_forget(input_gate)

[0042] In the above formula, [·] represents the concatenation operation along the last dimension; Gate_in(·) and Gate_forget(·) represent linear transformation operations; gin and gforget represent the pre-activation of the input gate and the forget gate, respectively.

[0043] S37: {QX, KY, VY, gin, gforget} are input into an mLSTM block for modeling; the gating mechanism and memory unit in mLSTM are used to selectively weight the features to reduce the negative impact of noise and occlusion on change detection; then, a learnable jump input is added after the output of the mLSTM block. This process can be expressed as:

[0044] Y′=μ(Q X ,K Y ,V Y , g in , g forget )+α·SiLU(Conv1d(Y_n)))∈R S×C

[0045] In the above formula, α represents a learnable parameter; μ(·) represents the operation of the mLSTM function; Y' represents the output result of a single branch after modeling; similarly, {QY, KX, VX} and its corresponding gin, gforget are input into another mLSTM block for modeling to obtain X', so as to realize the information interaction between the two-phase features.

[0046] S37: Restore the feature to its original shape to maintain the consistency between input and output.

[0047] Preferably, the multi-channel feature optimization mechanism proposed in step S4, in this mechanism, assumes that the input feature is D∈R C×H×W After using the convolutional block to perform basic enhancement on D, it is divided into D according to the channel dimension in the ratio of {2,2,3,1} 1 ,D 2 ,D 3 ,D 4 Four parts are used to achieve multi-level optimization of feature maps. This process can be expressed as:

[0048] D'=CBR 3×3 (CBR 3×3 (D)

[0049] D 1 ,D 2 ,D 3 ,D 4 =split(D',{2,2,3,1})

[0050] In the above formula, CBR3×3(·) represents a 3×3 convolutional layer with batch normalization and ReLU activation function; D' represents the enhanced features; split(·) represents the segmentation operation;

[0051] Will After convolution with kernel 3, Cascade and get To enhance the information expression ability of features, and then dynamically suppress redundant information through the gating mechanism;

[0052] The specific steps include:

[0053] S41: Flatten to Map it to a high-dimensional space through a linear layer to obtain D 12 ′∈R L×2C , to increase the information capacity of the feature; then D 12 ′ is evenly divided into D 1 ′∈R L×C ,D 2 ′∈R L ×C Two parts, for D 2 ′ performs depth-separable convolution and GeLU processing, and the processed D 2 ′ and D 1 'Element-by-element multiplication to complete the gating operation, this process can be expressed as:

[0054] D 12 =Cat(Conv2d 3×3 (D 1 ),D 2 )

[0055]

[0056] In the above formula, Conv2d3×3(·) represents a 3×3 convolution layer; Reshape(·) represents the reshaping operation; * represents element-by-element multiplication; φ(·) represents the linear layer operation; dw(·) represents the depthwise separable convolution operation; LN(·) represents the layer normalization operation; the final linear layer and layer normalization layer map the features back to the original dimension and enhance the stability of the model during training, obtaining a change feature after redundancy suppression.

[0057] S42: During feature recovery, The SSM algorithm in VMamba is introduced to obtain global information; at this time, the model needs to be simplified, and only the common flattening direction is used for the global scanning operation. This process can be expressed as:

[0058]

[0059] In the above formula, expand(·) represents the scanning expansion operation in one direction; Merge(·) represents the feature map recovery operation; represents a model of selective scanning spatial state sequences; With the addition of , the model can utilize the fine expression of local features and obtain the supplement of global context information.

[0060] S43: Keep the original input features and The process of splicing and fusion is performed in the channel dimension. This process can be expressed by mathematical formula:

[0061]

[0062] S44: In the change detection task, a hybrid loss is cited, that is, a combination of binary cross entropy loss and dice loss. The binary cross entropy loss, that is, BCE loss, can be expressed as

[0063]

[0064] In the above formula, N represents the total number of pixels, gt i ∈{0,1} represents the binary ground truth label of the ith pixel, Represents the predicted probability of the i-th pixel position changing. The mathematical expression of Dice loss is:

[0065]

[0066] In the above formula, Dice loss measures the spatial overlap between the predicted area and the true area;

[0067] S45: The mathematical formula of hybrid loss is:

[0068]

[0069] In the above formula, n represents the total number of predicted images involved in the loss calculation.

[0070] Compared with the prior art, the present invention has the following beneficial effects:

[0071] This paper proposes a remote sensing change detection method based on spatiotemporal perception and redundancy suppression. It uses the parallelism and time series modeling capabilities of the extended long short-term memory network to effectively capture the time attributes between change detection data and model the spatiotemporal information. At the same time, multi-level decoding is used to suppress redundant information layer by layer, thereby improving the accuracy and robustness of the detection method. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is a schematic diagram of the remote sensing change detection network architecture based on spatiotemporal perception and redundancy suppression;

[0073] Figure 2 It is a schematic diagram of the Temporal Aware Feature Interaction Module (TAFIM);

[0074] Figure 3 Schematic diagram of multi-channel feature optimization mechanism;

[0075] Figure 4 Flowchart of a remote sensing change detection method based on spatiotemporal perception and redundancy suppression. DETAILED DESCRIPTION

[0076] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments:

[0077] The specific implementation process is as follows:

[0078] like Figure 4 As shown in FIG. 1 , a remote sensing change detection method based on spatiotemporal perception and redundancy suppression is implemented in the following specific steps:

[0079] Step 1: Collect remote sensing image datasets and ensure that each image has pixel-level annotations. Divide the dataset into training, validation, and test sets for use in different stages, and cut the size of the dataset photos into 256×256;

[0080] Step 2: Build a remote sensing change detection network architecture and feature extractor (ResNet50) based on spatiotemporal perception and redundancy suppression; use the pre-trained ResNet50 to extract features from remote sensing dual-temporal images, aggregate these extracted features along the channel dimension, and obtain new features, so that the network can learn more compact and effective feature representation;

[0081] Step 3: Add the time-aware feature interaction module (TAFIM) after the last layer of features extracted by the backbone network, aiming to integrate important information in space and time to fully and effectively capture and process the complex change features in the dual-phase features. And add the aggregated shallow features to the mixed spatiotemporal features to prevent the loss of shallow texture details;

[0082] Step 4: Design a multi-stage decoder. The decoder uses dual decoding branches to process complementary information from different upsampling methods through a multi-channel feature optimization mechanism (MCFOM) to obtain redundant suppressed features, and then predicts the change map with fine details. The multi-layer prediction map and loss function (binary cross entropy loss and dice coefficient) output by the multi-stage decoder are used for deep supervision to learn richer information;

[0083] Step 5: Save the best model weights during the training process, test it using the test set, and obtain the change detection result of the test image.

[0084] The paper uses weight-shared ResNet50 to extract multi-scale features from dual-phase images; then, a temporal-aware feature interaction module (TAFIM) is used to capture the spatiotemporal relationship between dual-phase features, focusing on areas with significant changes, while integrating shallow features to prevent the loss of details; next, a multi-level decoder is constructed by combining different upsampling methods and a multi-channel feature optimization mechanism (MCFOM) to gradually suppress the redundant information in the dual-phase mixed features; and finally a change map is output.

[0085] The specific implementation steps are as follows:

[0086] Furthermore, the remote sensing change detection network architecture based on spatiotemporal perception and redundancy suppression proposed in step 2 is as follows Figure 1 As shown in the figure, the proposed network can be divided into three parts: feature extraction, feature interaction and multi-level decoding; by designing the time-aware feature interaction module (TAFIM) and the multi-channel feature optimization mechanism (MCFOM), a new approach is provided for inter-layer feature fusion and feature reconstruction. First, the dual-phase image is input into the weight-sharing pre-trained Siamese network (ResNet50) to extract the features. Then, the extracted feature maps are aggregated in the channel dimension to obtain a new This process can be expressed as:

[0087]

[0088] In the above formula, Conv1×1(·) represents a 1×1 convolutional layer with batch normalization, so that the network can learn a more concise and effective feature representation.

[0089] Furthermore, the aggregated features Input to TAFIM to capture the spatiotemporal relationship between the two-phase features and reduce the impact of environmental changes on feature matching, which can be expressed as:

[0090]

[0091] In the above formula, T(·) represents the function of TAFIM, and pi represents the mixed feature with spatiotemporal information.

[0092] Furthermore, the fused initial features are integrated into the dual-phase mixed features with spatiotemporal information to prevent the loss of texture information and edge information and enhance the robustness of the model, namely:

[0093]

[0094] In the above formula, Cat(·) represents the cascade operation along the channel dimension, and Mi represents the spatiotemporal hybrid feature that integrates shallow detail information.

[0095] Furthermore, the temporal-aware feature interaction module (TAFIM) module mentioned in step 3 ( Figure 2 As shown in Figure 2, the remote sensing change detection task has both spatial and temporal attributes. Spatial and temporal modeling between change detection data can often effectively identify change trends and improve the effectiveness of change detection methods. The lack of temporal modeling may make it difficult for the model to capture the normal changes of remote sensing images over time, and thus perform poorly in the face of environmental changes (such as seasonal changes). In recent studies, the introduction of visual LSTM has confirmed the potential of extended long short-term memory networks (xLSTM) in visual backbone networks. Considering the powerful spatiotemporal modeling capabilities of xLSTM, and the matrix LSTM (mLSTM) therein supports highly parallel processing. The present invention explores the strategy of using mLSTM to model the spatiotemporal information of dual-phase features in the feature interaction stage. This technology complements the ability of the CNN backbone in temporal feature analysis. Based on this, the present invention designs TAFIM behind the last layer of features extracted by the backbone network. It aims to integrate important information in space and time, and comprehensively and efficiently capture and process complex change features in dual-phase features.

[0096] Furthermore, the bi-phase input features are set as X,Y∈R C×H×W , and input it into TAFIM to capture the spatiotemporal information between the two time phases. First, the present invention flattens X and Y and then performs layer normalization operation so that each pixel can be independently and stably processed. n ,Y n ∈R S×C . Where S is the product of H and W. Then, causal convolution (kernel size is 4) is used to perceive the changes and dynamic processes in the remote sensing image. And the SiLU function is used to introduce the nonlinearity to enhance the expressive power of the model. Then, the present invention maps the input tensor to a high-dimensional space to obtain the query (Q), key (K) and value (V) corresponding to the dual-phase features. The process can be expressed by mathematical formula:

[0097]

[0098] In the above formula, NH is the number of heads, d represents the dimension of each head, C represents the input dimension, S represents the pixel sequence length, and Rea represents the rearrangement operation.

[0099] Furthermore, for each head, a linear transformation is performed to obtain:

[0100]

[0101] at the same time,

[0102]

[0103] In the above formula, WV, WQ, WV represent the weight matrices corresponding to V, Q, and K respectively. represents the inverse operation of the rearrangement operation, b represents the bias vector, and Conv1d(·) represents the causal convolution operation.

[0104] Furthermore, the idea of ​​cross attention is combined to realize the interactive operation of dual-phase features in mLSTM, so that the difference between the information of the other phase can be highlighted while retaining its own feature information, enhancing the model's ability to focus on the change area. First, match QX from the X branch with KY and VY from the Y branch. By simply concatenating QX, KY, and VY, the pre-activation calculation of the input gate and the forget gate is obtained. This process can be expressed as:

[0105] input gate =[Q X ,K Y ,V Y ]

[0106] g in =Gate_in(input_gate)

[0107] g forget =Gate_forget(input_gate)

[0108] In the above formula, [·] represents the concatenation operation along the last dimension; Gate_in(·) and Gate_forget(·) represent linear transformation operations; gin and gforget represent the pre-activation of the input gate and the forget gate, respectively.

[0109] Furthermore, {QX, KY, VY, gin, gforget} are jointly input into an mLSTM block for modeling; the gating mechanism and memory unit in mLSTM are used to selectively weight the features to reduce the negative impact of noise and occlusion on change detection; then, a learnable jump input is added after the output of the mLSTM block to improve information transfer and enhance the robustness of the model. This process can be expressed as:

[0110] Y′=μ(Q X ,K Y ,V Y , g in , g forget )+α·SiLU(Conv1d(Y_n)))∈R S×C

[0111] In the above formula, α represents a learnable parameter; μ(·) represents the operation of the mLSTM function; Y' represents the output result of a single branch after modeling; similarly, {QY, KX, VX} and its corresponding gin, gforget are input into another mLSTM block for modeling to obtain X', so as to realize the information interaction between the dual-phase features. Finally, the features are restored to their original shape to maintain the consistency of input and output. In short, this part combines the advantages of mLSTM and cross-attention ideas. Adaptively learn the spatiotemporal relationship of dual-phase images. Reduce the interference caused by environmental differences.

[0112] Furthermore, the multi-channel feature optimization mechanism proposed in step 4 is as follows Figure 3 As shown in Figure 2. Due to the influence of shooting time and environmental changes, remote sensing images often present redundant background information or low-contrast areas in different channels. These redundant information may mask the real change areas in the image. In addition, the fused features may contain repeated or highly correlated information, which further increases the redundancy and thus increases the computational burden. To solve the above problem, the present invention proposes MBFOM. In this mechanism, let the input feature be D∈R C×H×W After the convolution block performs basic enhancement processing on D, the present invention divides it into D according to the channel dimension in the ratio of {2,2,3,1}. 1 ,D 2 ,D 3 ,D 4 Four parts are needed to achieve multi-level optimization of feature maps. This process can be expressed as:

[0113] D'=CBR 3×3 (CBR 3×3 (D)

[0114] D 1 ,D 2 ,D 3 ,D 4 =split(D',{2,2,3,1})

[0115] In the above formula, CBR3×3(·) represents a 3×3 convolutional layer with batch normalization and ReLU activation function; D' represents the enhanced feature; split(·) represents the segmentation operation.

[0116] First, After convolution with kernel 3, Cascade and get In order to enhance the information expression ability of the features, the redundant information can be dynamically suppressed through the gating mechanism.

[0117] Specifically, Flatten to Map it to a high-dimensional space through a linear layer to obtain D 12 ′∈R L×2C , to increase the information capacity of the features, which is crucial for introducing locality into the architecture. 12 ′ is evenly divided into D 1 ′∈R L×C ,D 2 ′∈R L×C Two parts. 2 ′ performs depth-separable convolution and GeLU processing, and then the processed D 2 ′ and D 1 ′ multiply element by element to complete the gating operation. This process can be expressed as:

[0118] D 12 =Cat(Conv2d 3×3 (D 1 ),D 2 )

[0119]

[0120] In the above formula, Conv2d3×3(·) represents a 3×3 convolution layer; Reshape(·) represents the reshape operation; * represents element-by-element multiplication; φ(·) represents the linear layer operation; dw(·) represents the depthwise separable convolution operation; LN(·) represents the layer normalization operation. The final linear layer and layer normalization layer map the features back to the original dimension and enhance the stability of the model during training, obtaining a change feature after redundancy suppression.

[0121] Furthermore, in the feature recovery process, if only local features are relied upon, the overall shape and distribution of the object will be ignored, thus affecting the accuracy of change detection. The SSM algorithm in VMamba is introduced to obtain global information. Studies have shown that "remote sensing images are different from traditional images in terms of features. In the Mamba-based method, a simple flattening operation is sufficient." Therefore, in order to simplify the model, the present invention only uses the common flattening direction when performing the global scanning operation. This process can be expressed as:

[0122]

[0123] In the above formula, expand(·) represents the scanning expansion operation in one direction; Merge(·) represents the feature map recovery operation; Represents the Selective Scanning Space State Sequence Model, which is the core SSM operator of this part. With the addition of , the model can not only utilize the fine expression of local features, but also obtain the supplement of global context information, thereby further improving the hierarchical expression of feature maps.

[0124] Keep the original input features and The splicing and fusion are performed in the channel dimension. The purpose of this is to avoid the loss of key information caused by over-processing and provide a stable benchmark for the feature optimization process. This process can be expressed in mathematical formula as follows:

[0125]

[0126] In the change detection task, in most cases, the proportion of changed areas is much smaller than the proportion of unchanged areas, which will lead to class imbalance problem. In order to effectively alleviate this problem, the present invention refers to the hybrid loss, which is a combination of binary cross entropy (BCE) loss and dice loss (Dice). BCE loss can be expressed as

[0127]

[0128] Where N is the total number of pixels, gt i ∈{0,1} represents the binary ground truth label of the ith pixel, represents the predicted probability of the i-th pixel position changing. The mathematical expression of Dice loss is

[0129]

[0130] In the above formula, Dice loss effectively solves the category imbalance problem by measuring the spatial overlap between the predicted area and the true area, which makes it particularly suitable for change detection tasks where the change area is usually sparse.

[0131] Furthermore, the mathematical formula of the hybrid loss is expressed as:

[0132]

[0133] In the above formula, n represents the total number of predicted images involved in the loss calculation. By combining BCE and Dice loss, pixel-level accuracy and region-level consistency complement each other, thereby suppressing the sample imbalance problem.

[0134] Furthermore, the present invention designs a state space model (SSM) assisted multi-stage decoder in this paper. The decoder adopts dual decoding branches and predicts the change map with fine details through complementary information brought by different upsampling methods.

[0135] Specifically, the present invention upsamples the inter-layer features through the Patch Expanding layer and bilinear interpolation. The Patch Expanding method focuses on feature expansion based on local areas and can retain more complete spatial information. Subsequently, the upper branch hierarchical features are fused into the lower branch hierarchical features, and the designed MCFOM is used to optimize the fused hierarchical features. This process can be expressed as

[0136] B 1 =E 1 =Conv 1×1 (Cat(M 1 ,M 2 ))

[0137] E i+1 =Up(E i ), i∈1,2,3

[0138] B i+1 =Conv 1×1 (Cat(up(O(B i )),E i+1 )), i∈1,2,3

[0139] In the above formula, Ei and Bi represent the hierarchical features of the dual decoding branch, O(·) represents the function of the MCFOM designed by the present invention, Up(·) represents the patch expanding layer upsampling operation, and up(·) represents the bilinear interpolation layer upsampling operation. Through this process, redundant information is reconstructed step by step and suppressed, and finally a binary change map with clear edges is obtained. In addition, a deep supervision method is introduced to enhance the learning ability and convergence speed of the network.

[0140] The above is only an embodiment of the present invention, and the common knowledge such as the known specific technical solutions and / or characteristics in the solution is not described in detail here. It should be pointed out that for those skilled in the art, without departing from the technical solution of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A remote sensing change detection method based on spatiotemporal perception and redundancy suppression, characterized in that: The specific implementation steps include: S1: Collect remote sensing image datasets and ensure that each image has pixel-level annotations; divide the dataset into training set, validation set, and test set for use in different stages, and cut the size of the dataset photos into 256×256; S2: Build a remote sensing change detection network architecture and feature extractor based on spatiotemporal perception and redundancy suppression, the feature extractor is ResNet50; use the pre-trained ResNet50 to extract features from remote sensing dual-phase images, aggregate these extracted features along the channel dimension, and obtain new features, so that the network can learn more compact and effective feature representation; S3: A temporal-aware feature interaction module, or TAFIM, is added after the last layer of features extracted by the backbone network. It aims to integrate important information in space and time to comprehensively and effectively capture and process the complex changing features in the dual-temporal features, and add the aggregated shallow features to the mixed spatiotemporal features to prevent the loss of shallow texture details. S4: Design a multi-stage decoder with dual decoding branches. The decoder uses a multi-channel feature optimization mechanism, MCFOM, to process complementary information from different upsampling methods to obtain redundantly suppressed features, and then predicts the change map with fine details. The multi-layer prediction map and loss function output by the multi-stage decoder are used for deep supervision to learn richer information. S5: Save the best model weights in the training process, use the test set for testing, and obtain the change detection result of the test image.

2. The remote sensing change detection method based on spatiotemporal perception and redundancy suppression according to claim 1 is characterized in that: The remote sensing change detection network architecture based on spatiotemporal perception and redundancy suppression proposed in step S2 can be divided into three parts, including feature extraction, feature interaction and multi-level decoding; The specific steps include: S21: The dual-phase image is input into the weight-sharing pre-trained twin network, i.e., ResNet50, to extract features S22: Aggregate the extracted feature maps in the channel dimension to obtain new This process can be expressed as: In the above formula, Conv1×1(·) represents a 1×1 convolutional layer with batch normalization, so that the network can learn a more concise and effective feature representation; S23: The aggregated features Input to TAFIM to capture the spatiotemporal relationship between the two-phase features and reduce the impact of environmental changes on feature matching, which can be expressed as: In the above formula, T(·) represents the function of TAFIM, and pi represents the mixed feature with spatiotemporal information; S24: Integrate the fused initial features into the dual-phase mixed features with spatiotemporal information to prevent the loss of texture information and edge information and enhance the robustness of the model, namely: In the above formula, Cat(·) represents the cascade operation along the channel dimension, and Mi represents the spatiotemporal hybrid feature that integrates shallow detail information.

3. The remote sensing change detection method based on spatiotemporal perception and redundancy suppression according to claim 1 is characterized in that: In the step S3, it is proposed to design TAFIM after the last layer of features extracted by the backbone network, integrate important information in space and time, and comprehensively and efficiently capture and process the complex change features in the dual-phase features; The specific steps include: S31: Set the bi-phase input features to X, Y∈R C×H×W , and input it into TAFIM to capture the spatiotemporal information between the two temporal phases. S32: Flatten X and Y and perform layer normalization operation so that each pixel can be processed independently and stably, and the result is: X n ,AND n ∈T S×C In the above formula, S is the product of H and W. S33: Use causal convolution to perceive changes and dynamic processes in remote sensing images, and use SiLU function to introduce nonlinearity to enhance the expressive power of the model. S34: Map the input tensor to a high-dimensional space to obtain the query (Q), key (K) and value (V) corresponding to the bi-phase feature. This process can be expressed by the mathematical formula: In the above formula, NH is the number of heads, d represents the dimension of each head, C represents the input dimension, S represents the length of the pixel sequence, and Rea represents the rearrangement operation; S35: For each head, perform linear transformation to obtain: at the same time, In the above formula, WV, WQ, WV represent the weight matrices corresponding to V, Q, and K respectively. represents the inverse operation of the rearrangement operation, b represents the bias vector, and Conv1d(·) represents the causal convolution operation; S36: Combine the cross attention idea to realize the interactive operation of dual-phase features in mLSTM, match QX from the X branch with KY, VY from the Y branch, and obtain the pre-activation calculation of the input gate and the forget gate by simply concatenating QX, KY, VY. This process can be expressed as: input gate =[Q X ,K Y ,V Y ] g in =Gate_in(input_gate) g forget =Gate_forget(input_gate) In the above formula, [·] represents the concatenation operation along the last dimension; Gate_in(·) and Gate_forget(·) represent linear transformation operations; gin and gforget represent the pre-activation of the input gate and the forget gate, respectively. S37: {QX, KY, VY, gin, gforget} are input into an mLSTM block for modeling; the gating mechanism and memory unit in mLSTM are used to selectively weight the features to reduce the negative impact of noise and occlusion on change detection; then, a learnable jump input is added after the output of the mLSTM block. This process can be expressed as: Y′=μ(Q X ,K Y ,V Y ,g in ,g forget )+α·SiLU(Conv1d(Y_n)))∈R S×C In the above formula, α represents a learnable parameter; μ(·) represents the operation of the mLSTM function; Y' represents the output result of a single branch after modeling; similarly, {QY, KX, VX} and its corresponding gin, gforget are input into another mLSTM block for modeling to obtain X', so as to realize the information interaction between the two-phase features. S37: Restore the feature to its original shape to maintain the consistency between input and output.

4. The remote sensing change detection method based on spatiotemporal perception and redundancy suppression according to claim 1 is characterized in that: The multi-channel feature optimization mechanism proposed in step S4 is as follows: In this mechanism, let the input feature be D∈R C×H×W After using the convolution block to perform basic enhancement processing on D, it is divided into four parts D1, D2, D3, and D4 in the ratio of {2, 2, 3, 1} according to the channel dimension to achieve multi-level optimization of the feature map. This process can be expressed as: D‘=CBR 3×3 (CBR 3×3 (D)) D1,D2,D3,D4=split(D',{2,2,3,1}) In the above formula, CBR3×3(·) represents a 3×3 convolutional layer with batch normalization and ReLU activation function; D' represents the enhanced feature; split(·) represents the split operation; Will After convolution with kernel 3, Cascade and get To enhance the information expression ability of features, and then dynamically suppress redundant information through the gating mechanism; The specific steps include: S41: Flatten to (L = H × W), and map it to a high-dimensional space through a linear layer to obtain D 12 ′∈R L×2C , to increase the information capacity of the feature; then D 12 ′ is evenly divided into D1′∈R L×C ,D2′∈R L×C There are two parts. D2′ is processed by depth-wise separable convolution and GeLU. The processed D2′ is multiplied element-wise with D1′ to complete the gating operation. This process can be expressed as: D 12 =Cat(Conv2d 3×3 (D1),D2) In the above formula, Conv2d3×3(·) represents a 3×3 convolution layer; Reshape(·) represents the reshaping operation; * represents element-by-element multiplication; φ(·) represents the linear layer operation; dw(·) represents the depthwise separable convolution operation; LN(·) represents the layer normalization operation; the final linear layer and layer normalization layer map the features back to the original dimension and enhance the stability of the model during training, obtaining a change feature after redundancy suppression. S42: During feature recovery, The SSM algorithm in VMamba is introduced to obtain global information; at this time, the model needs to be simplified, and only the common flattening direction is used for the global scanning operation. This process can be expressed as: In the above formula, expand(·) represents the scanning expansion operation in one direction; Merge(·) represents the feature map recovery operation; represents a model of selective scanning spatial state sequences; With the addition of , the model can utilize the fine expression of local features and obtain the supplement of global context information. S43: Keep the original input features and The process of splicing and fusion is performed in the channel dimension. This process can be expressed by mathematical formula: S44: In the change detection task, a hybrid loss is cited, that is, a combination of binary cross entropy loss and dice loss. The binary cross entropy loss, that is, BCE loss, can be expressed as In the above formula, N represents the total number of pixels, gt i ∈{0,1} represents the binary ground truth label of the ith pixel, Represents the predicted probability of the i-th pixel position changing. The mathematical expression of Dice loss is: In the above formula, Dice loss measures the spatial overlap between the predicted area and the true area; S45: The mathematical formula of hybrid loss is: In the above formula, n represents the total number of predicted images involved in the loss calculation. By combining BCE and Dice loss, pixel-level accuracy and region-level consistency complement each other, thereby suppressing the sample imbalance problem.