A remote sensing image change detection method based on spatiotemporal interactive Transformer model

By introducing a space-time interaction Transformer model across time and across space interaction modules, the problem of time-space interaction defects in remote sensing image change detection is solved, and efficient and accurate remote sensing image change detection is achieved.

CN117095287BActive Publication Date: 2025-09-02ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310933742.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-09-02
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

The existing remote sensing image change detection methods lack consideration of the time and space dimensional characteristics, resulting in time-space interaction defects, high computational complexity and high redundancy, making it difficult to effectively detect small area changes.

Method used

The remote sensing image change detection method based on the spatial and temporal interaction Transformer model is adopted to extract features across time and across space interaction modules, and combine frequency domain information to design a lightweight network structure to achieve efficient fusion of time-space features.

Benefits of technology

The efficiency and accuracy of remote sensing image change detection are improved, the calculation complexity is reduced, and a highly efficient remote sensing image change detection scheme is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095287B_ABST
    Figure CN117095287B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote sensing image change detection method based on a spatiotemporal interactive Transformer model. In view of the problem that the common paradigm of existing remote sensing image change detection methods lacks consideration of temporal and spatial dimensional features, resulting in spatiotemporal interaction defects, the present invention designs a spatiotemporal interactive Transformer model for multi-temporal feature extraction, which is the first general backbone network specially designed for the remote sensing image change detection task. The present invention also proposes a parameter-free multi-frequency token mixer for integrating frequency domain features that provide spectral information. The present invention not only utilizes the spectral information in the remote sensing image to enrich the frequency domain features of the image, but also utilizes the spatiotemporal interactive Transformer model to enhance the spatiotemporal interaction, thereby achieving efficient remote sensing image change detection. The present invention combines temporal features and spatial features to provide a new solution for remote sensing image change detection, which can achieve a satisfactory balance between efficiency and accuracy in the field of remote sensing image change detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention applies technologies related to deep learning and computer vision, and specifically invents and applies a remote sensing image change detection method based on a spatiotemporal interactive Transformer model. Background Art

[0002] With the development of Earth observation technology, the amount of remote sensing imagery has increased dramatically, prompting the Earth science and remote sensing communities to adopt deep learning techniques to accomplish related tasks. Remote sensing image change detection focuses on comparing two or more images of the same area taken at different times to quantitatively and qualitatively assess changes in geographic entities and environmental factors, typically across multiple scales and temporal contexts. It serves a wide range of purposes, including environmental monitoring, urban planning, disaster assessment, and land use, and possesses both high scientific significance and practical value.

[0003] The task of change detection in remote sensing imagery can be viewed as a binary semantic segmentation problem, which assigns a binary label to each pixel, indicating whether the object of interest in the corresponding region has changed. In practical applications, frequent changes of no interest, caused by seasonal illumination variations, irrelevant motion, and even differences in sensor and imaging conditions, pose significant challenges to change detection in remote sensing imagery. Furthermore, within a certain time span, the size of the changed region can be much smaller than the target region, requiring rich spatial details for detection.

[0004] Traditional change detection methods for remote sensing images are mostly based on algebraic summation. While simple to implement, these methods rely on handcrafted features, resulting in high computational complexity and noise sensitivity. The recent rise of deep learning techniques, particularly convolutional neural networks, has significantly advanced the field of remote sensing image change detection due to their outstanding nonlinear fitting capabilities, enabling the extraction of high-quality discriminative features. Some methods incorporate Siamese neural networks into remote sensing image change detection, extracting bi-temporal features through concatenation or summation, and then applying them to a change detection head. This paradigm can be further improved by using a weight-sharing cascaded classification network as the backbone to enhance the performance of the change detection head. For example, methods based on spatial attention and channel attention to enhance feature representation optimize concatenation or phase subtraction to refine temporal feature interactions. However, the multi-level features obtained by cascading classification networks still lack significant semantic information and spatial detail, and the high redundancy in deep feature channels leads to significant computational costs. Furthermore, U-shaped structures can stack and fuse features from different levels, improving the method's ability to distinguish between changed and unchanged regions. However, the dense connectivity involved also leads to the aforementioned computational issues and high redundancy.

[0005] Recent research has employed Transformer models (which can be translated as converter models) for remote sensing image change detection to circumvent the limitations of convolutional neural networks in terms of fixed perceptual fields and weak capture of long-range dependencies. For example, a pure Transformer model remote sensing image change detection network was proposed using SwinTransformer; coarse-grained and fine-grained features were extracted from dual-temporal images by constructing a pair of twin neural networks with layered Transformer model encoders; and the Transformer model encoder was used to model context in a compact token-based space-time, where the learned context-rich tokens were fed into the pixel space, and the decoder refined the original features. However, these methods also follow the cascade design of the classification network, and the attention mechanism requires high computational complexity. Summary of the Invention

[0006] The technical problem to be solved by the present invention is how to simultaneously consider the characteristics of remote sensing images while following the paradigm of non-interactive twin neural networks and change detection heads, improve feature expression capabilities by fusing cross-temporal and cross-spatial interactions of features during feature extraction, and provide a remote sensing image change detection method based on a spatiotemporal interactive Transformer model. By introducing cross-temporal and cross-spatial interaction modules, the present invention extracts and integrates the spatial and temporal features of features at each stage, and by enriching feature representations with frequency domain information, achieves a linear complexity, lightweight model, and improves the accuracy and robustness of the model.

[0007] The specific technical solutions adopted in the present invention are as follows:

[0008] A remote sensing image change detection method based on a spatiotemporal interactive Transformer model is proposed. The specific approach is as follows: bi-temporal remote sensing images at two moments to be detected are input into a trained spatiotemporal interactive Transformer model network to obtain the final change detection results;

[0009] The spatiotemporal interaction Transformer model network uses the spatiotemporal interaction module as an encoder and a multi-layer perceptron as a decoder;

[0010] The spatiotemporal interaction module includes four cascaded stages, each of which has the same network structure and is composed of two patch embedding modules, two cross-temporal interaction modules and a cross-space interaction module. In the spatiotemporal interaction module, two remote sensing images in the input dual-temporal remote sensing image are first used as the input of the first stage to generate dual-temporal features, and the dual-temporal features output by the previous stage are used as the input of the next stage. In each stage, the initial input is first converted into embedded tokens by a patch embedding module, and then fed into the respective cross-space interaction modules to extract multi-scale features. The deepest features extracted by each cross-space interaction module are passed to the cross-temporal interaction module as the encoding stage features, and cross-temporally interact with the deepest features extracted by another cross-space interaction module to generate enhanced features after enhancing the time difference. The enhanced features corresponding to each deepest feature are returned to the cross-space interaction module that generated the deepest feature, and then after multi-level upsampling and jump connection, the spatial details are restored to form the output features, realizing cross-temporal and cross-space interaction of the dual-phase features at each stage. The output features of the two cross-space interaction modules are used as the final output dual-temporal features.

[0011] In the encoder, the bi-temporal features output by each of the four stages are input into a multi-layer perceptron decoder for decoding. The bi-temporal features output from each of the four stages are spliced ​​into a change representation along the channel dimension. Then, all four change representations are upsampled to the same resolution through bilinear interpolation and spliced ​​along the channel dimension. The spliced ​​change representation is convolved with 1*1 and then restored to the size of the original remote sensing image through upsampling to generate the final remote sensing image change detection result.

[0012] Preferably, the input of the cross-temporal interaction module is the deepest features extracted by each of the two cross-spatial interaction modules. Each deepest feature is element-wise subtracted from the other deepest feature to obtain a rough change representation. Each deepest feature is then concatenated with the rough change representation to form a concatenated feature. Each concatenated feature is then processed using depth-wise separable convolution and Sigmoid activation function to obtain an enhanced difference weight map. Finally, the deepest feature of each input is weightedly summed with the corresponding enhanced difference weight map to obtain an enhanced feature after enhancing the time difference corresponding to each deepest feature.

[0013] Preferably, the cross-spatial interaction module adopts a U-shaped network architecture consisting of a contraction path and an expansion path, and a total of four basic blocks are used in the two paths for feature extraction; the original input features of the cross-spatial interaction module are first input into the contraction path, passed through the first basic block for feature extraction, and then down-sampled and passed through the second basic block for feature extraction, and then down-sampled and passed to the cross-temporal interaction module as the deepest feature; the enhanced features returned by the cross-temporal interaction module are input into the expansion path, and after up-sampling, they are jump-connected with the features extracted from the second basic block, and then input into the third basic block for feature extraction, and then after up-sampling, they are jump-connected with the features extracted from the first basic block, and continue to be input into the fourth basic block for feature extraction, and finally one temporal feature of the dual-temporal feature is obtained;

[0014] The basic block adopts the Transformer model architecture. The original input features of the basic block are first subjected to a regularization function to increase nonlinear features, and then a multi-frequency mixer is used to enrich the frequency domain information of the feature representation to obtain the first intermediate feature. The first intermediate feature is then residually connected with the original input feature and input into the regularization processing and channel multi-layer perceptron module with residual connection to obtain the output features of the basic block.

[0015] Preferably, the input of the multi-frequency mixer is the regularized feature map in each basic block, and the feature map is encoded using a two-dimensional discrete cosine transform algorithm based on a basis corresponding to a plurality of pre-selected effective frequencies to obtain an encoded spectrum; the feature map is split into multiple sub-feature maps along the channel dimension, the spectrum is weighted to each sub-feature map and all the weighted sub-feature maps are re-spliced ​​to obtain the output of the multi-frequency mixer.

[0016] Preferably, when a two-dimensional discrete cosine transform algorithm is used for encoding, bases corresponding to multiple effective frequencies need to be pre-selected in advance using a frequency selection strategy.

[0017] Preferably, the frequency selection strategy includes a pre-training prior strategy, a random selection strategy and a dynamic programming strategy; wherein, the pre-training prior strategy is to conduct experiments on ImageNet, explore the importance of frequency by selecting only one frequency at a time, and thus select the most important multiple frequencies; the random selection strategy is to randomly select several frequency values ​​for token mixing based on the information that the signal energy tends to maintain low frequencies, while maintaining the lowest frequency value; the dynamic programming strategy is to incorporate frequency selection into model training, send the spectrogram into the convolution module, and use the Sigmoid activation function to obtain weight values, from which multiple frequencies with the highest weight values ​​are selected.

[0018] Preferably, the frequency selection strategy adopts a pre-training prior strategy.

[0019] Preferably, in the decoder, all four variation representations are upsampled to the same H / 2×W / 2 resolution by bilinear interpolation, where H and W are the height and width of the original remote sensing image, respectively.

[0020] Preferably, the loss function adopted by the spatiotemporal interactive Transformer model network is a weighted sum of a focal loss function and a Dice loss function.

[0021] Preferably, the remote sensing image is a high-resolution remote sensing image with a spatial resolution of less than 1 m.

[0022] Compared with the prior art, the present invention has the following benefits:

[0023] The present invention discloses a remote sensing image change detection method based on a spatiotemporal interactive Transformer model. In view of the problem that existing remote sensing image change detection methods generally follow a fixed paradigm and lack consideration of temporal and spatial dimensional features, resulting in temporal-spatial interaction defects, the present invention designs a temporal-spatial interactive Transformer model for multi-temporal feature extraction. It is the first general backbone network designed specifically for remote sensing image change detection tasks. At the same time, a parameter-free multi-frequency token mixer is proposed to integrate frequency domain features that provide spectral information. The present invention not only enriches the frequency domain features of the image using the spectral information in the remote sensing image, but also enhances the spatiotemporal interaction of the remote sensing image change detection method through the spatiotemporal interactive Transformer model, thereby achieving efficient remote sensing image change detection. The present invention combines temporal features and spatial features to provide a new solution for remote sensing image change detection tasks, achieving a satisfactory balance between efficiency and accuracy in the field of remote sensing image change detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Visualize the results of the remote sensing image change detection challenge;

[0025] Figure 2 This is the structure diagram of the encoder part in the STeInFormer model;

[0026] Figure 3 Schematic diagram of the multi-frequency mixer structure;

[0027] Figure 4 A training and testing flow chart of the STeInFormer model in an embodiment of the present invention;

[0028] Figure 5 This is the test visualization result in the embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to make the above-mentioned objects, features and advantages of the present invention more clearly understood, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined accordingly without conflicting with each other.

[0030] Remote sensing image change detection can be viewed as a binary semantic segmentation problem, which assigns a binary label to each pixel, indicating whether the object of interest in the corresponding area has changed. Figure 1 As shown in the figure, in practical applications, frequent non-interesting changes caused by seasonal illumination variations, unrelated motion, and even differences in sensor and imaging conditions pose significant challenges to remote sensing image change detection. Furthermore, within a certain time span, the size of the changed region may be much smaller than the target region, requiring rich spatial details for detection. Related methods mostly follow the paradigm of non-interactive Siamese neural networks and change detection heads, but rarely consider the characteristics of remote sensing image change detection. It is hypothesized that integrating cross-temporal and cross-spatial interactions of features in feature extraction can improve the performance of remote sensing image change detection. Following this hypothesis, this paper proposes a novel spatiotemporal interactive Transformer model (i.e., STeInFormer) for remote sensing image change detection. It should be noted that this network is the first architecture designed entirely for remote sensing image change detection. Its capabilities have been extensively verified experimentally, making it suitable as a general backbone for change detection tasks. Furthermore, due to the incorporation of frequency domain information, it has linear complexity, enabling lightweight network construction.

[0031] The present invention provides a remote sensing image change detection method based on a spatiotemporal interactive Transformer model. Specifically, the spatiotemporal interactive Transformer model network STeInFormer is trained and used as a change detection model. When performing a detection task, two remote sensing images to be detected at different times are input into the STeInFormer change detection model, which has a backbone network composed of a cross-spatiotemporal interaction module and a cross-spatial interaction module as an encoder and a multi-layer perceptron as a decoder, to obtain change detection results. The images in the present invention are preferably remote sensing images, and more preferably high-resolution remote sensing images with a spatial resolution of less than 1 meter.

[0032] The specific structure and principle of the above change detection model STeInFormer are described in detail below.

[0033] The backbone network, or encoder, in the aforementioned SteInFormer consists primarily of a cross-temporal interaction module and a cross-spatial interaction module. The cross-temporal interaction module employs a gating mechanism to emphasize interesting changes while suppressing non-interesting changes in feature extraction. The cross-spatial interaction module, as the encoding stage based on a U-shaped architecture, integrates semantic and detail information to achieve more robust feature representation. During operation, the backbone network first converts the input bi-temporal image into embedded tokens via a patch embedding module, which are then fed into the cross-spatial interaction module to extract multi-scale features. Simultaneously, the deepest features of each cross-spatial interaction module are passed as encoding-stage features to the corresponding cross-temporal interaction module for cross-temporal interaction, generating features that enhance temporal differences. These enhanced features are then returned to the cross-spatial interaction module, where they undergo multi-level upsampling and skip connections to restore spatial details. Bi-temporal features are then output at four scales, enabling cross-temporal and cross-spatial interaction of bi-temporal features at each stage.

[0034] The decoder of the STeInFormer model is a multi-layer perceptron. The bi-temporal features of different scales are input into the multi-layer perceptron decoder for decoding. The bi-temporal features output from each of the four stages are initially concatenated into a change representation. All four change representations are then upsampled to the same resolution through bilinear interpolation and concatenated. The concatenated change representation is then convolved and upsampled to generate the final remote sensing image change detection result.

[0035] The specific structure of the STeInFormer model of the present invention is described in detail below. Figure 2 The encoder architecture of the STeInFormer model consists of four cascaded stages, each with a common network structure: two patch embedding modules (PEs), two U-Blocks (U-Blocks in the four stages are denoted as U-Block-H4, U-Block-H3, U-Block-H2, and U-Block-H1, respectively), and a cross-spatial interaction module (CTI). The CTI module extracts multi-scale features, while the CTI enriches the temporal features of the image. The CTI also includes a multi-frequency mixer to enrich the frequency domain information of the image. In the CTI module, two remote sensing images from the input bi-temporal remote sensing image are used as input to the first stage to generate bi-temporal features (i.e., feature maps at two different moments in time). The bi-temporal features output by the previous stage serve as the input to the next stage, and the bi-temporal features output by all four stages serve as the input to the decoder.

[0036] Specifically, see Figure 2As shown in the figure, in each stage, the initial input of the stage is first converted into an embedded token through a patch embedding module, and then fed into the respective cross-spatial interaction modules to extract multi-scale features. The deepest features extracted by each cross-spatial interaction module are passed to the cross-temporal interaction module as the encoding stage features, and are cross-temporally interacted with the deepest features extracted by another cross-spatial interaction module to generate enhanced features after enhancing the time difference. The enhanced features corresponding to each deepest feature are returned to the cross-spatial interaction module that generates the deepest feature, and then after multi-level upsampling and jump connections, the spatial details are restored to form the output features, realizing cross-temporal and cross-spatial interaction of the dual-phase features at each stage; the output features of the two cross-spatial interaction modules are used as the final output dual-temporal features.

[0037] In an embodiment of the present invention, if the remote sensing image dimension in the original bi-temporal image is H×W×C, the bi-temporal image is first converted into embedded tokens through a patch embedding module, and then fed into two cross-spatial interaction modules to extract features. At this time, the feature size is H / 2×W / 2×C; the cross-spatial interaction module uses a U-shaped architecture to refine and downsample the features through basic blocks, repeating twice. The deepest features of the two cross-spatial interaction modules are passed to the cross-temporal interaction module of the corresponding stage for cross-temporal interaction to generate features with enhanced temporal differences. The enhanced features are returned to the cross-spatial interaction module, and after processing by multiple basic blocks, upsampling and skip connections, the spatial details are restored and passed to the patch embedding module, cross-spatial interaction module and cross-temporal interaction module of the next stage to further enrich the details. At the same time, the feature size is reduced by 2 times at each stage. Therefore, the feature sizes fed to the cross-spatial interaction module in the next three stages are H / 4×W / 4×2C, H / 8×W / 8×4C, and H / 16×W / 16×8C, respectively. Ultimately, bi-temporal features are output at four scales, enabling cross-temporal and cross-spatial interaction of bi-temporal features at each stage. The output of the backbone network is passed as input to the decoder of the multi-layer perceptron. The bi-temporal features output from each of the four stages are initially concatenated into a change representation. All four change representations are then upsampled to the same spatial resolution of H / 2 × W / 2 using bilinear interpolation and concatenated. The concatenated change representation is then convolved and upsampled to restore the dimensions to H × W × 1, generating the final remote sensing image change detection result. In this change detection result, the H × W dimensional image is a binary image representing whether each pixel in the remote sensing image has changed at two moments in time.

[0038] The cross-temporal interaction module in the present invention is inspired by the gating mechanism. It enhances feature differences by learning weights. The input of the cross-temporal interaction module at each stage is the deepest feature of the input dual-phase feature processed by the cross-space interaction module, and the output is the feature enhanced by the enhanced feature difference weight. The module first uses element-level subtraction on the input feature to obtain a rough change representation, then splices the dual-phase feature with the change representation feature, and then uses depth-wise separable convolution and Sigmoid activation function to process the spliced ​​features to obtain an enhanced difference weight map. Finally, the enhanced weight and the dual-phase feature are weighted summed to obtain a feature map containing rich temporal information.

[0039] In this embodiment, the main purpose of the cross-temporal interaction module is to enrich the temporal features of the image and realize cross-temporal information interaction. The input of the cross-temporal interaction module is the deepest features extracted by each of the two cross-spatial interaction modules. Each deepest feature is element-wise subtracted from the other deepest feature to obtain a rough change representation. Each deepest feature is then concatenated with the rough change representation to form a concatenated feature. Each concatenated feature is then processed using a depthwise separable convolution and a Sigmoid activation function to obtain an enhanced difference weight map. Finally, the deepest feature of each input is weighted summed with the corresponding enhanced difference weight map to obtain an enhanced feature after enhancing the temporal difference corresponding to each deepest feature.

[0040] Specifically, see Figure 2 As shown in the figure, the specific approach of the cross-temporal interaction module is as follows: first, the deepest bi-temporal features F1 and F2 of the input cross-spatial interaction module are subjected to element-wise subtraction to obtain the rough change feature R c Then, the bi-temporal features F1 and F2 are respectively combined with the change feature R c Cascade and use depth-separable convolution and Sigmoid activation function to process the two cascaded features respectively to obtain weight maps W1 and W2. Finally, we weight F1 and F2 with W1 and W2 respectively to adjust the enhanced feature representation after temporal difference enhancement.

[0041] The cross-spatial interaction module in the present invention is inspired by the traditional U-Net architecture, uses a U-shaped architecture, and relies on basic blocks to perform feature extraction stage by stage. The input of the cross-spatial interaction module at each stage is the feature representation processed by the previous level module. The input feature is passed to the basic block for feature extraction. Following the Transformer model, nonlinear features are added through the regularization function, and the frequency domain information of the feature representation is enriched using a multi-frequency mixer. Then, the feature representation is input into the regularization processing and channel multi-layer perceptron module with residual connection to obtain the feature representation processed by the basic block. Following the U-shaped architecture, the basic blocks are used to perform multi-level feature extraction on the input features in turn. At the same time, the deepest features are transmitted to the corresponding cross-time interaction module through the gating mechanism for feature separation, and the output dual-temporal features are returned to the cross-spatial interaction module. The feature representation enhanced over time differences is spliced ​​and upsampled step by step. Finally, the enhanced feature representation after the interactive fusion of spatiotemporal information is obtained.

[0042] Specifically, see Figure 2 As shown, the cross-spatial interaction module adopts a U-shaped network architecture consisting of a contraction path and an expansion path. The two paths use a total of four base blocks (Base Block, abbreviated as B) for feature extraction. The original input features of the cross-spatial interaction module are first input into the contraction path, passed through the first base block for feature extraction, then downsampled and passed through the second base block for feature extraction, and then downsampled again and passed to the cross-temporal interaction module as the deepest feature; the enhanced features returned by the cross-temporal interaction module are input into the expansion path, upsampled and skip-connected with the features extracted from the second base block, and then input into the third base block for feature extraction. After upsampling, they are skip-connected with the features extracted from the first base block and continue to be input into the fourth base block for feature extraction, ultimately obtaining one temporal feature of the dual-temporal feature.

[0043] Continue to see Figure 2 As shown in the figure, the above basic block adopts the Transformer model architecture. The original input features of the basic block are first normalized by the regularization function to increase the nonlinear features, and then the frequency domain information of the feature representation is enriched by the multi-frequency mixer to obtain the first intermediate feature. The first intermediate feature is then residually connected with the original input feature and input into the regularization processing and channel multi-layer perceptron module with residual connection to obtain the output features of the basic block.

[0044] The cross-space interaction module is mainly used to enrich the spatial features of the image and realize cross-space information interaction. The purpose of using a multi-frequency mixer in this embodiment is to add frequency domain information that can enrich the feature representation. The multi-frequency mixer introduces the effective frequency information in the spatial domain into the multi-head attention mixer. The input of the multi-frequency mixer is the feature representation after regularization in each basic block. The mixer uses two-dimensional discrete cosine transform encoding to obtain the pattern features of each frequency of the input feature, selects multiple effective frequency bases for calculation to improve the efficiency of batch processing, and uses projection and separation operations to perform weighted summation on the input features to obtain the pattern features of the corresponding frequencies. All pattern features are spliced ​​and projected to obtain the final output of the multi-frequency mixer.

[0045] like Figure 3 As shown, the input of the multi-frequency mixer is the regularized feature map in each basic block. The multi-frequency mixer can encode the feature map based on the basis corresponding to multiple pre-selected effective frequencies using a two-dimensional discrete cosine transform algorithm (2D DCT) to obtain an encoded spectrum; the feature map is split into multiple sub-feature maps along the channel dimension, the spectrum is weighted to each sub-feature map and all the weighted sub-feature maps are re-joined to obtain the output of the multi-frequency mixer. In an embodiment of the present invention, the specific method in the multi-frequency mixer is as follows: first, the feature map R input to the mixer is p , R p Split into M+1 sub-feature maps A0, A1, ..., A along the channel dimension M ; For each sub-feature map A m Corresponding to the preselection of a basis (DCTbase), these bases can be used to calculate the spectrum f after 2D DCT encoding h,w , the spectrum f h,w Weighted to each sub-feature map and all weighted sub-feature maps are re-joined in the split order to obtain the output feature map R f .

[0046] It should be noted that the two-dimensional discrete cosine coding algorithm belongs to the prior art, wherein the basis in the two-dimensional discrete cosine coding algorithm is represented as in, So we can get the two-dimensional discrete cosine transform formula Among them A x,y represents the input image, whose dimension is H×W, f h,w Represents the spectrum after two-dimensional discrete cosine coding. Before using this algorithm, it is necessary to use a frequency preselection strategy to select the basis of different frequencies. Screening is performed to obtain the basis corresponding to the most important several effective frequencies, which are used to calculate the spectrum required for weighting.

[0047] It should be noted that when the above-mentioned change detection model selects the basis corresponding to the effective frequency, the frequency selection strategy can adopt a pre-training prior strategy, a random selection strategy and a dynamic programming strategy. Among them, the pre-training prior strategy is to conduct experiments on ImageNet, explore the importance of frequency by selecting only one frequency at a time, and thus select the most important multiple frequencies; the random selection strategy is to randomly select several frequency values ​​for token mixing based on the information that the signal energy tends to maintain low frequencies, while maintaining the lowest frequency value; the dynamic programming strategy is to incorporate frequency selection into model training, send the spectrum graph into the convolution module, and use the Sigmoid activation function to obtain the weight value, from which multiple frequencies with the highest weight value are selected. The number of bases corresponding to the effective frequency is M+1, and the specific value can be optimized through experiments based on actual data. In an embodiment of the present invention, the frequency selection strategy adopts a pre-training prior strategy according to the experimental results.

[0048] In order to expand the training samples, the training data can be enhanced. The loss function used by the model is a hybrid loss function L=L that combines the focus loss function and the Dice loss function. focal +L dice The specific calculation process will not be described in detail. The point loss function and the Dice loss function each belong to the existing technology, among which the focal loss function Among them, α and γ are two hyperparameters that control the weight of positive and negative samples and the attention of the method to difficult samples, p is the probability, y is the binary label of the pixel (0 or 1) corresponding to the unchanged and changed values, and the Dice loss function E′={e′ k},k∈[1,H×W], where E represents the ground truth, E′ is of dimension H×W×2 and represents the change map, e′ k represents a two-dimensional pixel in E′.

[0049] The remote sensing image change detection method based on the spatiotemporal interactive Transformer model is applied to a specific embodiment to demonstrate the technical effects that can be achieved.

[0050] Example

[0051] The network structure of the change detection model used in this embodiment is as described above and will not be repeated here. Figure 4 As shown in Figure 3, the overall process of change detection for remote sensing images can be divided into three stages: data preprocessing, model training, and image prediction.

[0052] 1. Data preprocessing stage

[0053] For the original remote sensing images obtained (this embodiment takes the WHU-CD dataset as an example), image preprocessing is performed, and operations such as random rotation and flipping are performed to perform data enhancement.

[0054] 2. Model training

[0055] Step 1: Construct the training set data and divide the training data set into batches according to a fixed batch size, with a total of N.

[0056] Step 2: Sequentially select a batch of training samples with index i from the training dataset, where i∈{0,1,...,N}. Use each batch of training samples to train the change detection model (i.e., the aforementioned spatiotemporal interactive Transformer model STeInFormer). During the training process, the hybrid loss function L formed by adding the focal loss function and the Dice loss function for each training sample is calculated. b , and the loss L is calculated based on all training samples in the batch b The total loss L is calculated, and the network parameters of the entire STeInFormer model are adjusted based on the total loss until all batches of the training dataset participate in the model training. After reaching the specified number of iterations, the model converges and the training is completed. The final SteInFormer model is used as the change detection model for testing or application.

[0057] 3. Image Prediction

[0058] The images in the test set are directly used as input to the trained change detection model, and the probability vector of each changed pixel class is finally predicted. The class with the highest probability is selected as the change classification output through activation functions such as sigmoid, thereby realizing change detection.

[0059] In this embodiment, the test visualization results are as follows: Figure 5 The test data results are shown in Table 1:

[0060] Table 1 Test data results

[0061] Dataset F1 Pre. Rec. IoU OA WHU-CD 89.61 91.01 88.26 79.87 98.68

[0062] Depend on Figure 5 As can be seen from Table 1, the change detection model of the present invention can well process the change detection results for remote sensing images. It relies on the spatiotemporal interaction module, fully considers the characteristics of the time and space dimensions of remote sensing images, solves the problem of time-space interaction defects, and designs a parameter-free multi-frequency token mixture to enrich the frequency domain characteristics of the image, realizes efficient change detection, and provides an efficient general backbone network design scheme for remote sensing image change detection tasks.

[0063] The embodiment described above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.

Claims

1. A remote sensing image change detection method based on a spatiotemporal interactive Transformer model, characterized by: Input the two-time dual-temporal remote sensing images to be detected into the trained spatiotemporal interactive Transformer model network to obtain the final change detection results; The spatiotemporal interaction Transformer model network uses the spatiotemporal interaction module as an encoder and a multi-layer perceptron as a decoder; The spatiotemporal interaction module consists of four cascaded stages, each of which has the same network structure, consisting of two patch embedding modules, two cross-temporal interaction modules, and a cross-space interaction module. In the spatiotemporal interaction module, two remote sensing images in the input bi-temporal remote sensing image are first used as input to the first stage to generate bi-temporal features, and the bi-temporal features output by the previous stage are used as input to the next stage. In each stage, the initial input is first converted into an embedded token through a patch embedding module, and then fed into the respective cross-spatial interaction modules to extract multi-scale features. The deepest features extracted by each cross-spatial interaction module are passed to the cross-temporal interaction module as the encoding stage features, and cross-temporally interact with the deepest features extracted by another cross-spatial interaction module to generate enhanced features after enhancing the temporal difference. The enhanced features corresponding to each deepest feature are returned to the cross-spatial interaction module that generated the deepest feature. After multi-level upsampling and jump connections, the spatial details are restored to form the output features, realizing cross-temporal and cross-spatial interaction of bi-temporal features at each stage; the output features of the two cross-spatial interaction modules are used as the final output bi-temporal features; In the encoder, the bi-temporal features output by each of the four stages are input into a multi-layer perceptron decoder for decoding. The bi-temporal features output from each of the four stages are spliced ​​into a change representation along the channel dimension. Then, all four change representations are upsampled to the same resolution through bilinear interpolation and spliced ​​along the channel dimension. The spliced ​​change representation is convolved with 1*1 and then restored to the size of the original remote sensing image through upsampling to generate the final remote sensing image change detection result.

2. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 1, characterized in that: The input of the cross-temporal interaction module is the deepest features extracted by each of the two cross-spatial interaction modules. Each deepest feature is element-wise subtracted from the other deepest feature to obtain a rough change representation. Each deepest feature is then concatenated with the rough change representation to form a concatenated feature. Each concatenated feature is then processed using depthwise separable convolution and Sigmoid activation function to obtain an enhanced difference weight map. Finally, the deepest feature of each input is weightedly summed with the corresponding enhanced difference weight map to obtain an enhanced feature after enhancing the time difference corresponding to each deepest feature.

3. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 1, characterized in that: The cross-spatial interaction module adopts a U-shaped network architecture consisting of a contraction path and an expansion path, and a total of four basic blocks are used in the two paths for feature extraction; the original input features of the cross-spatial interaction module are first input into the contraction path, passed through the first basic block for feature extraction, then downsampled and passed through the second basic block for feature extraction, and then downsampled and passed to the cross-temporal interaction module as the deepest feature; the enhanced features returned by the cross-temporal interaction module are input into the expansion path, upsampled and skip-connected with the features extracted from the second basic block, and then input into the third basic block for feature extraction, and then upsampled and skip-connected with the features extracted from the first basic block, and then input into the fourth basic block for feature extraction, ultimately obtaining one temporal feature of the dual-temporal feature; The basic block adopts the Transformer model architecture. The original input features of the basic block are first subjected to a regularization function to increase nonlinear features, and then a multi-frequency mixer is used to enrich the frequency domain information of the feature representation to obtain the first intermediate feature. The first intermediate feature is then residually connected with the original input feature and input into the regularization processing and channel multi-layer perceptron module with residual connection to obtain the output features of the basic block.

4. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 3, characterized in that: The input of the multi-frequency mixer is the regularized feature map in each basic block. Based on the basis corresponding to multiple pre-selected effective frequencies, the feature map is encoded using a two-dimensional discrete cosine transform algorithm to obtain an encoded spectrum; the feature map is split into multiple sub-feature maps along the channel dimension, the spectrum is weighted to each sub-feature map and all the weighted sub-feature maps are re-spliced ​​to obtain the output of the multi-frequency mixer.

5. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 4, characterized in that: When using the two-dimensional discrete cosine transform algorithm for encoding, bases corresponding to multiple valid frequencies need to be pre-selected in advance using a frequency selection strategy.

6. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 5, characterized in that: The frequency selection strategies include pre-training prior strategy, random selection strategy and dynamic programming strategy; among them, the pre-training prior strategy is to conduct experiments on ImageNet, explore the importance of frequency by selecting only one frequency at a time, and thus select multiple most important frequencies; the random selection strategy is to randomly select several frequency values ​​for token mixing based on the information that signal energy tends to maintain low frequencies, while maintaining the lowest frequency value; the dynamic programming strategy is to incorporate frequency selection into model training, send the spectrogram into the convolution module, and use the Sigmoid activation function to obtain the weight value, from which multiple frequencies with the highest weight values ​​are selected.

7. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 6, characterized in that: The frequency selection strategy adopts a pre-training prior strategy.

8. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 1, characterized in that: In the decoder, all four variation representations are upsampled to the same H / 2×W / 2 resolution by bilinear interpolation, where H and W are the height and width of the original remote sensing image, respectively.

9. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 1, characterized in that: The loss function used by the spatiotemporal interactive Transformer model network is the weighted sum of the focal loss function and the Dice loss function.

10. The remote sensing image change detection method based on the spatiotemporal interactive Transformer model according to claim 1, characterized in that: The remote sensing image is a high-resolution remote sensing image with a spatial resolution of less than 1m.

Citation Information

Patent Citations

  • Remote sensing image change detection method based on interactive feature perception

    CN116363527A

  • Remote sensing image change detection method based on local-global Transform network

    CN116434069A