A Progressive Fusion Remote Sensing Change Detection Method for Spatiotemporal Pixel Features Based on ViT

By performing low-rank fine-tuning and gradual fusion of multi-level features on the visual basic model backbone, the problem of insufficient homogeneity and heterogeneity characteristics in remote sensing change detection is solved, and the accuracy and generalization ability of remote sensing change detection is improved, which is suitable for multiple application scenarios.

CN120088655BActive Publication Date: 2025-07-22GUIZHOU SECOND INST OF SURVEYING & MAPPING

Patent Information

Application Number
CN202510554150.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-22
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing remote sensing change detection method based on Transformer faces the problems of insufficient synergistic enhancement of homogeneous features and limited differentiated characterization capabilities of heterogeneous features, resulting in blurred boundaries and loss of details in changes targets, and insufficient generalization capabilities.

Method used

The visual basic model backbone is fine-tuned by low-rank matrix, combining global context and timing dependence branches, and enhancing the dual-time phase remote sensing image features through a multi-level feature gradual fusion, and using multi-head self-attention mechanism and channel slicing technology to capture the global context information and timing dependence of remote sensing images to achieve gradual fusion and retention of features.

Benefits of technology

It significantly improves the clarity and detection accuracy of target boundaries in complex scenarios, reduces information loss, and enhances the generalization ability of the model. It is suitable for surface change detection in fields such as arable land protection, urban management, environmental monitoring and disaster assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088655B_ABST
    Figure CN120088655B_ABST
Patent Text Reader

Abstract

The present invention discloses a spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT, including: using a pre-trained Vision Transformer (ViT) as the backbone network, fine-tuning the model parameters by introducing a low-rank matrix, and using the fine-tuned model to extract the features of dual-temporal remote sensing images; enhancing the global-local feature interaction of dual-temporal remote sensing images through the global context branch and the temporal dependence branch, and combining a multi-level progressive fusion mechanism to perform three-stage fusion (shallow, middle, and high levels) on the features of the front and back temporal remote sensing images respectively to obtain the enhanced features of dual-temporal remote sensing images; after the features of dual-temporal remote sensing images are connected by channels, convolutional fusion is performed, and a binary classification prediction change image is generated through upsampling, non-linear feature transformation, and convolution. The present invention can significantly improve the clarity of the target boundary and the detection accuracy in complex scenes, effectively enhance the generalization ability of the model, and reduce information loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image change detection, and specifically to a spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT. Background Art

[0002] Change detection is an important research field in remote sensing interpretation and is the key to observing and analyzing surface changes. Compared with traditional differential information extraction methods, deep learning-based methods have high robustness. They can not only stably process large amounts of data but also model complex change information, which has greatly improved the accuracy of change detection tasks. Deep learning methods represented by the Transformer architecture have injected new development impetus into this field. With its powerful global dependency capture ability and advantages in remote spatio-temporal relationship modeling, this architecture has broken through the bottleneck of the receptive field limitation of traditional convolutional neural networks and provided an innovative solution for change detection tasks that require high-level semantic understanding. However, the performance of the Transformer architecture highly depends on large-scale training data that matches the parameter scale. The scarcity of labeled data commonly existing in the remote sensing field severely restricts the generalization ability of the model in complex and changeable real scenarios.

[0003] Recently, vision transformers (ViTs), which are Transformer-based visual foundation models, have demonstrated excellent transfer learning capabilities through large-scale pre-training in the natural image field, providing new ideas for solving the remote sensing data annotation bottleneck. Although the global representation of a single-temporal remote sensing image constructed by such models through the dense attention mechanism has excellent feature expression capabilities, they still face great challenges when processing dual-temporal remote sensing images: First, the collaborative enhancement mechanism for homogeneous features between cross-temporal remote sensing images is insufficient, making it difficult to effectively maintain the feature consistency of unchanged areas; second, the differential representation ability of heterogeneous features is limited, resulting in prominent problems such as blurred boundaries and lost details of changed targets in complex backgrounds. Therefore, making full use of the powerful feature expression ability of ViT, effectively migrating it to the remote sensing scenario, and at the same time enhancing the homogeneous and heterogeneous feature representation capabilities between dual-temporal remote sensing images through method improvement to reduce the loss of change information during feature extraction has become a current research hotspot. Summary of the Invention

[0004] An object of the present invention is to provide a spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT that can significantly improve the target boundary clarity and detection accuracy in complex scenarios, effectively enhance the generalization ability of the model, and reduce information loss to solve the above problems.

[0005] To achieve the above object, the present invention adopts the following technical solutions.

[0006] A spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT of the present invention includes the following steps:

[0007] S1 Visual basic model backbone low-rank fine-tuning: Use a low-rank matrix to perform low-rank fine-tuning on the visual basic model backbone, and extract the features X1 and X2 of the pre-temporal and post-temporal remote sensing images;

[0008] S2 Dual-temporal remote sensing feature interaction enhancement: Extract the significant features of the pre-temporal and post-temporal remote sensing image features X1 and X2, which are respectively expressed as , Connect these two features by channel to obtain , where b is the batch size and d is the feature vector dimension; The dual-temporal remote sensing feature interaction enhancement includes two branches:

[0009] 2.1 Global context branch: Input into the Dense Multi-Head Self-Attention layer to capture the global context information of the remote sensing image, and obtain the global features of the dual-temporal remote sensing image ;

[0010] 2.2 Temporal dependence branch: Input into the temporal dependence branch to capture the temporal dependence relationship, and obtain the local features of the dual-temporal remote sensing image ;

[0011] 2.3 Dual-temporal remote sensing feature decoupling based on channel splitting: Concatenate and project the global features of the dual-temporal remote sensing image and the local features of the dual-temporal remote sensing image to obtain the significant features of the dual-temporal remote sensing image . The significant features of the dual-temporal remote sensing image are evenly divided into two parts along the channel dimension, and are respectively extracted as the significant features of the pre-temporal remote sensing image and the significant features of the post-temporal remote sensing image ;

[0012] S3 Multi-level feature progressive fusion enhancement:

[0013] 3.1 Multi-level feature progressive fusion enhancement: Starting from the 12th layer to the 24th layer of the significant features of the pre-temporal remote sensing image , extract one feature every four layers to obtain multi-level features , and perform progressive fusion enhancement on to obtain the enhanced features of the pre-temporal remote sensing image F 1 ,

[0014] 3.2 Multi - level Feature Progressive Fusion Enhancement: Starting from the 12th layer of the significant features of the later - phase remote - sensing image and ending at the 24th layer, extract one feature every four layers to obtain multi - level features , and perform progressive fusion enhancement to obtain the enhanced features of the later - phase remote - sensing image F 2 ;

[0015] S4 Dual - phase Remote - Sensing Image Fusion Feature Decoding:

[0016] 4.1 Feature Fusion and Upsampling: Concatenate the enhanced features of the earlier - phase remote - sensing image F 1 and the enhanced features of the later - phase remote - sensing image F 2 by channels, then perform convolutional fusion, and upsample back to the input image size through bilinear interpolation;

[0017] 4.2 Feature Mapping and Prediction: First, perform non - linear feature transformation on the fused feature map F through two fully - connected layers and the GELU activation function; then multiply each pixel value of the feature map by the 1×1 convolutional parameter value (the 1×1 convolutional parameters are adaptively updated during the training process) to obtain a prediction result with a size of 1024×1024 and pixel values between [0 - 1]. Each pixel value in the prediction result represents the predicted probability that the pixel has changed;

[0018] 4.3 Classification: Assign pixel values greater than or equal to 0.5 in the prediction result to 1, and pixel values less than 0.5 to 0 to obtain a binary - classification predicted change image; a pixel value of 1 in the binary - classification predicted change image represents that the predicted ground object has changed, and a pixel value of 0 represents that the predicted ground object has not changed.

[0019] The above - mentioned spatio - temporal pixel feature progressive fusion remote - sensing change detection method based on ViT, where the method of using a low - rank matrix for low - rank fine - tuning of the backbone of the visual basic model in step S1 includes the following steps:

[0020] 1.1 Data Pre - processing: Use 3 - band earlier - phase and later - phase remote - sensing images with a size of 1024×1024 as input data for pre - processing to form feature vector embeddings with a size of 768×64×64;

[0021] 1.2 Introducing low-rank matrix parameters: The backbone of the vision base model adopts the ViT (Vision Transformer) network architecture with 24 low-order fine-tuned Transformer layers. The Transformer encoder of ViT is stacked by a linear layer, a multi-head self-attention mechanism layer, and a feed-forward neural network layer; two low-rank matrices are introduced in the multi-head self-attention mechanism layer A and B, are denoted as: ΔWr = B ⋅ A , where d is the feature vector embedding dimension, is a hyperparameter set to 16; during the fine-tuning process, the parameters of the original vision base model backbone are frozen, and only the low-rank matrices A and B are updated;

[0022] 1.3 Initialization strategy: At the beginning of training, matrix A is initialized according to a random Gaussian distribution, while matrix B is initialized as a zero matrix;

[0023] 1.4 Parameter replacement and update: Replace the linear layer parameters in each Transformer encoder with the fine-tuned parameters: , , , , where W q , W k , W v , W d represent the original parameters of the query matrix Q, the key matrix K, the value matrix V, and the encoder output linear layer respectively; , , , represent the parameters after low-rank fine-tuning; the original weight matrix W is frozen and does not participate in gradient calculation, and only B and A are trained;

[0024] 1.5 Feature extraction:

[0025] 1.5.1 Layer normalization: Calculate the mean and variance of the feature vector embedding obtained in step (1) in the dimension, and perform layer normalization (LayerNormalization) on it using the mean and variance;

[0026] 1.5.2 Multi-head self-attention mechanism layer: Multiply the layer-normalized feature vector embedding with the fine-tuned model parameters , , to obtain the query matrix Q , the key matrix K and the value matrix V respectively; first, based on the query matrix Q and the key matrix KCalculate the attention scores, and then multiply the attention scores with the value matrix V to obtain the attention weights;

[0027] 1.5.3 Residual connection and layer normalization: Add the attention weights to the feature vector embedding to obtain new attention weights, which are used as the input of layer normalization. Further enhance the feature expression ability through layer normalization again to obtain the weighted features;

[0028] 1.5.4 Feed-forward neural network layer: The weighted features pass through a linear layer and a feed-forward neural network layer (FFN) for non-linear transformation. The feed-forward neural network layer contains two fully connected layers, two Dropout layers, and an activation function GELU for non-linear mapping. The output features are added to the attention weights to obtain the remote sensing image features X1 and X2 of the front and back time phases respectively.

[0029] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, wherein the step 1.1 data preprocessing includes the following steps:

[0030] 1.1.1 Segmentation operation: Use a convolutional layer with a kernel size of 16×16 and a stride of 16 to segment the 3-band front and back time phase remote sensing images with an input size of 1024×1024. The convolutional kernel sliding window covers an area of 16×16 each time and moves in a stride of 16. The two 1024×1024 3-band remote sensing images are segmented into 64×64 3-band remote sensing image blocks of size 16×16;

[0031] 1.1.2 Flattening operation: Flatten each 16×16 remote sensing image block into a one-dimensional vector in channel order. Each remote sensing image block contains 16×16×3 pixel values. After flattening, the length of the remote sensing image vector is 768;

[0032] 1.1.3 Mapping to a high-dimensional space: Use a fully connected layer to map the flattened remote sensing image block vector to an embedding space of 768 dimensions, and the generated high-dimensional embedding vector size is 768×64×64;

[0033] 1.1.4 Adding positional encoding: On the basis of the high-dimensional embedding vector, add positional encoding (PositionalEncoding). The dimension of the positional encoding is the same as that of the high-dimensional embedding vector, both are 768 dimensions. After adding the positional encoding, a feature vector embedding is formed, and the size is still 768×64×64.

[0034] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, in which the formula of layer normalization in the feature extraction step 1.5 is as follows:

[0035]

[0036] Among them, represents the input of layer normalization (here refers to the feature vector embedding), X LN represents the result after layer normalization, and represent the mean and standard deviation respectively, and are learnable scaling and translation parameters.

[0037] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, in which the query matrix Q , key matrix K and value matrix V in the multi-head self-attention mechanism layer of the feature extraction in step 1.5 are calculated as follows:

[0038]

[0039] Based on the query matrix Q and the key matrix K calculate the attention score 。 The formula for calculating the attention score is:

[0040]

[0041] Among them, d represents the dimension of the feature vector embedding.

[0042] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, in which the method of the temporal dependence branch in step 2.2 includes the following steps:

[0043] 2.2.1 Input feature projection: The input feature is multiplied by the linear projection matrix and mapped to a high-dimensional space to obtain the feature , , and the process is as follows:

[0044]

[0045] , is the linear projection matrix;

[0046] 2.2.2 Convolutional feature extraction: The projected feature Xproj After the convolution operation, local temporal features are extracted X conv , and the process is as follows:

[0047]

[0048] is the convolution operation;

[0049] 2.2.3 Parameter discretization: The linear projection matrix and the locally temporal features after layer normalization X conv are multiplied to obtain the time step , the input mapping matrix B , and the output mapping matrix C , and the process is as follows:

[0050]

[0051] GroupNorm is layer normalization, W 3 is the linear projection matrix, is the sampling interval of the state space, is the state transition matrix. A is constrained to be a negative definite matrix, and A = −exp(logA) is used for logarithmic parameterization. logA is a learnable parameter matrix, initialized with random values; is the mapping matrix from input to state; maps the state to the output, and 128 is the hidden state dimension;

[0052] The continuous parameter is converted into a discrete parameter , is the state transition matrix, and the method is as follows:

[0053]

[0054]

[0055] Among them, is the identity matrix;

[0056] 2.2.4 Hidden state update: Update the hidden state at the current moment through the discretized parameter t h t , and map the hidden state to temporal features through the output mapping matrix C t X t , and the method is as follows:

[0057]

[0058]

[0059] h t is the hidden state at the current moment, h t-1 is the hidden state at the previous moment;

[0060] 2.2.5 Result processing: Temporal features X t and the feature are multiplied by the result after passing through the activation function to obtain the feature X g ,X g After passing through layer normalization and multiplying by the linear projection matrix, the local features of the dual-temporal remote sensing image are obtained , and the method is as follows:

[0061]

[0062]

[0063] Sigmoid is the activation function, GroupNorm is the layer normalization, W 4 is the linear projection matrix.

[0064] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, wherein the step 3.1 The method for progressive fusion and enhancement of multi-level features includes the following steps:

[0065] 3.1.1 First stage, shallow feature fusion and enhancement: Connect the feature and the feature to obtain the feature through a 1×1 convolutional layer. The feature is added to the feature to obtain the feature ;

[0066] 3.1.2 Second stage, middle feature fusion and enhancement: Connect the feature and the feature to obtain the feature through a 1×1 convolutional layer and then add it to the feature . The feature is added to the feature to obtain the feature ;

[0067] 3.1.3 Third stage, high-level feature fusion and enhancement: Connect the feature Connect with the feature and then, through a 1×1 convolutional layer, connect with the feature and add them to obtain the feature . The feature is added to the feature to obtain the feature , that is, the enhanced feature of the previous temporal remote sensing image F 1 .

[0068] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, wherein step 3.1 The mathematical expression of multi-level feature progressive fusion enhancement is expressed as:

[0069]

[0070]

[0071] represents the significant feature of the previous temporal remote sensing image output by the backbone network, , represents a 1×1 convolutional layer, represents connection.

[0072] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, wherein the process of feature fusion and upsampling in step 4.1 is as follows:

[0073]

[0074] represents channel connection, is a 1x1 convolution, represents bilinear interpolation of the feature, and F represents the generated fused feature map.

[0075] Compared with the prior art, the present invention has obvious beneficial effects. From the above technical solutions, it can be seen that the present invention performs low-rank fine-tuning based on a vision foundation model as a backbone network to extract remote sensing image features of two temporal phases before and after; adopts a global-local collaborative modeling strategy, utilizes the multi-head self-attention mechanism to capture the global context information of the remote sensing image, and through the temporal dependence branch, uses a state space model to establish the context relationship of local features, effectively taking into account the scene adaptation ability in the spatial dimension and the evolution analysis ability in the temporal dimension, realizing the selective retention of double-temporal related change features and the continuous tracking of the temporal evolution pattern, and enhancing the model's ability to represent the homogeneous and heterogeneous features between double-temporal remote sensing images; at the same time, adopts a channel compression and residual learning mechanism to adaptively fuse the significant features of the pre-temporal remote sensing image and the significant features of the post-temporal remote sensing image at shallow, middle, and high levels, realizing three-stage cross-level feature interaction, reducing the loss of change information in the feature extraction process, and further improving the change detection accuracy. The present invention can significantly improve the recognition accuracy of change targets, effectively enhance the clarity of the boundaries of changed ground objects, detail retention, and generalization performance in complex scenes. It provides reliable surface change detection technical support for fields such as cultivated land protection, urban management, environmental monitoring, disaster assessment, and natural resource management. Description of the Drawings

[0076] Figure 1 is the flowchart of the present invention;

[0077] Figure 2 is the flowchart of feature extraction for low-rank fine-tuning of the vision foundation model backbone of the present invention.

[0078] Figure 3 is the flowchart of enhanced interaction of double-temporal remote sensing features of the present invention.

[0079] Figure 4 is the flowchart of enhanced progressive fusion of multi-level features of the present invention.

[0080] Figure 5 is the comparison diagram of the effects of the present invention. Detailed Embodiments

[0081] The following, in conjunction with the accompanying drawings and preferred embodiments, details the specific embodiments, structures, features, and effects of a remote sensing change detection method based on progressive fusion of spatio-temporal pixel features based on ViT proposed according to the present invention.

[0082] Embodiment 1:

[0083] Refer to Figure 1 , a remote sensing change detection method based on progressive fusion of spatio-temporal pixel features based on ViT of the present invention includes the following steps:

[0084] S1 Visual foundation model backbone low-rank fine-tuning: To adapt to the change detection task while maintaining the performance of the visual foundation model backbone, a low-rank matrix is used to perform low-rank fine-tuning on the visual foundation model backbone, and the features X1 and X2 of the pre-phase and post-phase remote sensing images are extracted.

[0085] The specific method is as follows:

[0086] 1.1 Data preprocessing: The 3-band pre-phase and post-phase remote sensing images with a size of 1024×1024 are used as input data for preprocessing. The steps are as follows:

[0087] 1.1.1 Segmentation operation: A convolutional layer with a kernel size of 16×16 and a stride of 16 is used to segment the input remote sensing image. The convolutional kernel sliding window covers an area of 16×16 each time and moves in a stride of 16. The two 1024×1024 3-band remote sensing images are segmented into 64×64 3-band remote sensing image patches of size 16×16;

[0088] 1.1.2 Flattening operation: Each 16×16 remote sensing image patch is flattened into a one-dimensional vector in channel order. Each remote sensing image patch contains 16×16×3 pixel values, and the length of the flattened remote sensing image vector is 768;

[0089] 1.1.3 Mapping to a high-dimensional space: A fully connected layer is used to map the flattened remote sensing image patch vector to a 768-dimensional embedding space, and the generated high-dimensional embedding vector has a size of 768×64×64;

[0090] 1.1.4 Adding positional encoding: On the basis of the high-dimensional embedding vector, positional encoding (PositionalEncoding) is added to provide spatial position information. The dimension of the positional encoding is the same as that of the high-dimensional embedding vector, both are 768 dimensions. After adding the positional encoding, a feature vector embedding is formed, and the size is still 768×64×64;

[0091] 1.2 Introducing low-rank matrix parameters: The visual foundation model backbone adopts a ViT (Vision Transformer) network architecture with 24 low-order fine-tuning Transformer layers. The Transformer encoder of ViT is stacked by a linear layer, a multi-head self-attention mechanism layer, and a feed-forward neural network layer; two low-rank matrices are introduced in the multi-head self-attention mechanism layer A and B, can be regarded as parameterized learning of this process, expressed as: ΔWr = B ⋅ A , where d is the dimension of the feature vector embedding, is a hyperparameter, set to 16, which is used to control the rank size of the low-rank matrix; during the fine-tuning process, the parameters of the original vision backbone model are frozen, and only the low-rank matrices A and B are updated;

[0092] 1.3 Initialization strategy: At the beginning of training, matrix A is initialized according to a random Gaussian distribution, while matrix B is initialized as a zero matrix; this initialization method ensures that in the initial stage of training, the low-rank update will not interfere with the original performance of the model.

[0093] 1.4 Parameter replacement and update: Replace the linear layer parameters in each Transformer encoder with the fine-tuned parameters: , , , , where W q , W k , W v , W d respectively represent the original parameters of the query matrix Q, key matrix K, value matrix V, and the encoder output linear layer; , , , represents the parameters after low-rank fine-tuning; the original weight matrix W is frozen and does not participate in gradient calculation, only B and A are trained; through the above low-rank fine-tuning, the vision backbone model can capture key information and adapt to the requirements of the remote sensing image change detection task.

[0094] 1.5 Feature extraction: Perform feature extraction on the feature vector embedding obtained in step (1); the method is as follows (as Figure 2 shown):

[0095] 1.5.1 Layer normalization: Calculate the mean and variance of the feature vector embedding in the dimension, and perform layer normalization (LayerNormalization) on it using the mean and variance. The layer normalization formula is:

[0096]

[0097] where, represents the input of layer normalization (here refers to the feature vector embedding), X LN represents the result after layer normalization, and respectively represent the mean and standard deviation, and are learnable scaling and translation parameters;

[0098] 1.5.2 Multi-head self-attention mechanism layer: Use the layer-normalized feature vector embedding and the fine-tuned model parameters , , Multiply them respectively to obtain the query matrix Q , the key matrix K and the value matrix V . The calculation formulas for the query matrix Q , the key matrix K and the value matrix V are as follows:

[0099]

[0100] Calculate the attention scores based on the query matrix Q and the key matrix K . The calculation formula for the attention scores is:

[0101]

[0102] where d represents the dimension of the feature vector embedding;

[0103] Multiply the attention scores by the value matrix V to obtain the attention weights;

[0104] 1.5.3 Residual Connection and Layer Normalization: Add the attention weights to the feature vector embedding to obtain new attention weights, which are used as the input of layer normalization. Further enhance the feature expression ability through layer normalization again to obtain the weighted features;

[0105] 1.5.4 Feed-Forward Neural Network Layer: The weighted features are non-linearly transformed through the linear layer and the feed-forward neural network layer (FFN). The feed-forward neural network layer contains two fully connected layers, two Dropout layers, and the activation function GELU, which are used to implement non-linear mapping. The output features are added to the attention weights respectively to obtain the remote sensing image features X1 and X2 of the front and back time phases;

[0106] S2 Dual-Temporal Remote Sensing Feature Interaction Enhancement: As Figure 3 shown, extract the significant features of the remote sensing image features X1 and X2 of the front and back time phases, which are respectively represented as . Connect these two features by channel to obtain , where b is the batch size and d is the feature vector dimension; The dual-temporal remote sensing feature interaction enhancement contains two branches:

[0107] 2.1 Global Context Branch: Input into the Dense Multi-Head Self-Attention layer to capture the global context information of the remote sensing image and obtain the global features of the dual-temporal remote sensing image;

[0108] 2.2 Temporal Dependence Branch: Input into the temporal dependence branch to capture the temporal dependence relationship and obtain the local features of the dual-temporal remote sensing image , while accelerating the training while maintaining a linear time complexity. The method is as follows:

[0109] 2.2.1 Input Feature Projection: Input feature is multiplied by the linear projection matrix and mapped to a high-dimensional space to obtain feature , , and the process is as follows:

[0110]

[0111] , is the linear projection matrix;

[0112] 2.2.2 Convolutional Feature Extraction: The projected feature X proj undergoes a convolutional operation to extract local temporal features X conv , and the process is as follows:

[0113]

[0114] is the convolutional operation;

[0115] 2.2.3 Parameter Discretization: The linear projection matrix is multiplied by the locally temporal feature X conv after layer normalization to obtain the time step , the input mapping matrix B , and the output mapping matrix C , and the process is as follows:

[0116]

[0117] GroupNorm is layer normalization, W 3 is the linear projection matrix, is the sampling interval of the state space, is the state transition matrix (describing the transition relationship between states, used to determine how many hidden states should be propagated from the previous token to the next token). To stabilize the training, A is constrained to be a negative definite matrix, and logarithmic parameterization A=− exp(log A ) , log Ais a learnable parameter matrix, initialized to random values; is the mapping matrix from input to state, used to determine how much input enters the hidden state; maps the state to the output, used to determine how the hidden state is transformed into the output, and 128 is the dimension of the hidden state;

[0118] Convert the continuous parameter into a discrete parameter , is the state transition matrix, describing the transition relationship between states, used to determine how many hidden states should be propagated from the previous feature to the next feature, as follows:

[0119]

[0120]

[0121] where is the identity matrix;

[0122] 2.2.4 Hidden state update: Update the hidden state at the current time through the discretized parameter h t , and map the hidden state to the temporal feature through the output mapping matrix C X t , as follows:

[0123]

[0124]

[0125] h t is the hidden state at the current time, h t-1 is the hidden state at the previous time;

[0126] 2.2.5 Result processing: The temporal feature X t is multiplied by the result after the activation function of the feature to obtain the feature X g ,X g is multiplied by the layer normalization and the linear projection matrix to obtain the local feature of the dual-temporal remote sensing image , as follows:

[0127]

[0128] ​​

[0129] The Sigmoid is the activation function, and GroupNorm is the layer normalization, W 4 is the linear projection matrix;

[0130] 2.3 Dual-temporal remote sensing feature decoupling based on channel splitting: The global features of the dual-temporal remote sensing images and the local features of the dual-temporal remote sensing images are stitched and projected to obtain the significant features of the dual-temporal remote sensing images . The significant features of the dual-temporal remote sensing images are evenly divided into two parts along the channel dimension and are respectively extracted as the significant features of the pre-temporal remote sensing images and the significant features of the post-temporal remote sensing images ;

[0131] The dual-temporal remote sensing feature interaction enhancement selectively retains the features related to the cross-temporal changes through the gating mechanism, suppresses the background interference information, and at the same time uses the long-term memory characteristics of the state space model to continuously track the temporal evolution pattern, significantly improving the discriminability of the key features in the change detection task.

[0132] S3 Multi-level feature progressive fusion enhancement:

[0133] 3.1 Multi-level feature progressive fusion enhancement: Starting from the 12th layer of the significant features of the pre-temporal remote sensing images and ending at the 24th layer, one feature is extracted every four layers to obtain the multi-level features . The is progressively fused and enhanced (as shown in Figure 4 ) to obtain the enhanced features of the pre-temporal remote sensing images F 1 . The method is as follows:

[0134] 3.1.1 The first stage, shallow feature fusion enhancement: The feature is concatenated with the feature and the feature is obtained through a 1×1 convolutional layer. The feature is added to the feature to obtain the feature ;

[0135] 3.1.2 The second stage, middle feature fusion enhancement: The feature is concatenated with the feature and then added to the feature through a 1×1 convolutional layer to obtain the feature . The feature is added to the feature Add to obtain features ;

[0136] 3.1.3 Third stage, high-level feature fusion enhancement: Connect the feature with the feature , and then add it to the feature through a 1×1 convolutional layer to obtain the feature . Add the feature to the feature to obtain the feature , that is, the enhanced feature of the previous-temporal remote sensing image F 1 ;

[0137] Through three-stage cross-level feature interaction, adopting the channel compression and residual learning mechanism, progressive fusion enhancement of multi-level feature information is achieved, and its mathematical expression can be expressed as::

[0138]

[0139]

[0140] represents the significant feature of the previous-temporal remote sensing image output by the backbone network, , represents the 1×1 convolutional layer, represents the connection;

[0141] 3.2 Progressive fusion enhancement of multi-level features: Starting from the 12th layer to the 24th layer of the significant feature of the later-temporal remote sensing image, extract one feature every four layers to obtain the multi-level feature , and perform progressive fusion enhancement on (as shown in Figure 4 ) to obtain the enhanced feature F 2 (the method is the same as progressive fusion enhancement of multi-level features);

[0142] S4 Decoding of the fused features of dual-temporal remote sensing images:

[0143] 4.1 Feature fusion and upsampling: Connect the enhanced feature F 1 of the previous-temporal remote sensing image and the enhanced feature F 2 of the later-temporal remote sensing image by channel, then perform convolutional fusion, and upsample and restore it to the input image size through bilinear interpolation. The process is as follows:

[0144]

[0145] Indicates a channel connection Is a 1x1 convolution Indicates bilinear interpolation of features, and F represents the generated fused feature map;

[0146] 4.2 Feature mapping and prediction: First, perform a non-linear feature transformation on the fused feature map F through two fully connected layers and the GELU activation function; then multiply each pixel value of the feature map by the 1x1 convolution parameter value through a 1×1 convolution (the 1×1 convolution parameters are adaptively updated during the training process) to obtain a prediction result with a size of 1024×1024 and pixel values between [0-1]. Each pixel value in the prediction result represents the predicted probability that the pixel has changed;

[0147] 4.3 Classification: Assign pixel values greater than or equal to 0.5 in the prediction result to 1, and pixel values less than 0.5 to 0 to obtain a binary classification predicted change image; a pixel value of 1 in the binary classification predicted change image represents that the predicted ground object has changed, and a pixel value of 0 represents that the predicted ground object has not changed.

[0148] Experimental example: Evaluation of the effectiveness of the present invention

[0149] Construct a loss function based on the true change label and the binary classification predicted change image, and screen the remote sensing change detection method based on ViT's spatio-temporal pixel feature progressive fusion of the present invention to obtain the remote sensing change detection method with the best performance in evaluation indicators; the determination method of the loss function is the binary cross-entropy function, as follows:

[0150]

[0151] In the formula, Represents the true label Represents the predicted label;

[0152] Experiments were conducted on the LEVIR-CD dataset to verify the effectiveness of the present invention. This dataset contains 637 pairs of remote sensing images, and the size of each image is 1024 × 1024. The present invention follows the official standard and divides the dataset into three subsets: training, validation, and test, containing 445, 64, and 128 pairs of remote sensing images respectively. Each pair of remote sensing images includes a pre-temporal remote sensing image, a post-temporal remote sensing image, and the corresponding true change label. Set the model input size to 1024×1024, and use rotation, flipping, and random cropping data augmentation techniques to increase the sample size. Use the AdamW optimizer with a learning rate of 0.0004 and cosine annealing adjustment to decay the learning rate. Set the batch size to 4 and the maximum epoch to 300. The change detection results obtained on the LEVIR-CD dataset (such as Figure 5As shown in [figure number], the predicted map is compared with the ground truth. In addition, in this test example, the effectiveness of the method is evaluated quantitatively. Table 1 shows the performance evaluation metrics of the test example on the LEVIR-CD dataset, including the F1 score (F1), intersection over union (IoU), and overall accuracy (OA), all of which range from [0,1]. As can be seen from Table 1, the evaluation metrics of the detection method of the present invention are all relatively high, indicating reliable detection accuracy.

[0153] Table 1 Precision Verification of Change Detection Results

[0154] F1(%) IoU (%) OA (%) This invention 92.46 86.1 99.2

[0155] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT, comprising the following steps: S1 Visual foundation model backbone low-rank fine-tuning: Use a low-rank matrix to perform low-rank fine-tuning on the visual foundation model backbone, and extract the features X1 and X2 of the pre-temporal and post-temporal remote sensing images; S2 Dual-temporal Remote Sensing Feature Interactive Enhancement: Extract the significant features of the remote sensing image features X1 and X2 in the pre- and post-temporal phases, and concatenate these two significant features by channel to obtain , where b is the batch size and d is the feature vector dimension; the dual-temporal remote sensing feature interactive enhancement includes two branches: 2.1 Global context branch: Input into the dense multi-head self-attention mechanism layer to capture the global context information of the remote sensing image and obtain the global features of the dual-temporal remote sensing image ; 2.2 Temporal Dependence Branch: Input into the temporal dependence branch to capture the temporal dependence relationship and obtain the local features of the dual-temporal remote sensing image ; 2.3 Dual-temporal remote sensing feature decoupling based on channel splitting: The global features of dual-temporal remote sensing images and the local features of dual-temporal remote sensing images are stitched and projected to obtain the significant features of dual-temporal remote sensing images . The significant features of dual-temporal remote sensing images are evenly divided into two parts along the channel dimension, and are respectively extracted as the significant features of the former-temporal remote sensing images and the significant features of the latter-temporal remote sensing images ; S3 Multi-level feature progressive fusion enhancement: 3.1 Multi-level feature progressive fusion enhancement: Starting from the 12th layer of the significant features of the previous-phase remote sensing image and ending at the 24th layer, one feature is extracted every four layers to obtain multi-level features , and perform progressive fusion enhancement to obtain the enhanced features of the previous-phase remote sensing image F 1 , 3.2 Multi - level feature progressive fusion enhancement: Starting from the 12th layer of the significant features of the later - phase remote - sensing image and ending at the 24th layer, extract one feature every four layers to obtain multi - level features , and perform progressive fusion enhancement on to obtain the enhanced features of the later - phase remote - sensing image F 2 ; S4 Dual-temporal remote sensing image fusion feature decoding: 4.1 Feature Fusion and Upsampling: Enhance the features of the previous-temporal remote sensing image F 1 and the enhanced features of the subsequent-temporal remote sensing image F 2 Connect them by channel, then perform convolutional fusion, and upsample them back to the size of the input image through bilinear interpolation; 4.2 Feature mapping and prediction: First, perform a non-linear feature transformation on the fusion feature map F through two fully connected layers and the GELU activation function; then multiply each pixel value of the feature map by the 1×1 convolution parameter value whose convolution parameters are adaptively updated during the training process, to obtain a prediction result with a size of 1024×1024 and pixel values between 0 and 1. Each pixel value in the prediction result represents the predicted probability that the pixel has changed; 4.3 Classification: Assign pixel values greater than or equal to 0.5 in the prediction result to 1, and pixel values less than 0.5 to 0, to obtain a binary classification predicted change image; a pixel value of 1 in the binary classification predicted change image represents that the predicted ground object has changed, and a pixel value of 0 represents that the predicted ground object has not changed.

2. A spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT according to claim 1, wherein the method of using a low-rank matrix to perform low-rank fine-tuning on the visual foundation model backbone in step S1 comprises the following steps: 1.1 Data preprocessing: Use the 3-band pre-temporal and post-temporal remote sensing images with a size of 1024×1024 as input data for preprocessing to form feature vector embeddings with a size of 768×64×64; 1.2 Introduce low-rank matrix parameters: The backbone of the vision base model adopts the ViT network architecture with 24 low-order fine-tuning Transformer layers. The Transformer encoder of ViT is stacked by a linear layer, a multi-head self-attention mechanism layer, and a feed-forward neural network layer; two low-rank matrices are introduced in the multi-head self-attention mechanism layer A and B, which are denoted as: , where d is the feature vector embedding dimension and r is a hyperparameter, set to 16; during the fine-tuning process, the parameters of the original vision base model backbone are frozen, and only the low-rank matrices A and B are updated; 1.3 Initialization strategy: At the beginning of training, matrix A is initialized according to a random Gaussian distribution, while matrix B is initialized as a zero matrix; 1.4 Parameter Replacement and Update: Replace the parameters of the linear layer in each Transformer encoder with the fine-tuned parameters: , , , , where W q , W k , W v , W d represent the original parameters of the query matrix Q, the key matrix K, the value matrix V, and the encoder output linear layer, respectively; represents the parameters after low-rank fine-tuning; the original weight matrix W is frozen and does not participate in gradient calculation, and only B and A are trained; 1.5 Feature extraction: 1.5.1 Layer normalization: Calculate the mean and variance of the feature vector embeddings obtained in step (1) in the dimension, and perform layer normalization on them using the mean and variance; 1.5.2 Multi-Head Self-Attention Mechanism Layer: Multiply the feature vector embedding after layer normalization with the fine-tuned model parameters to obtain the query matrix Q , the key matrix K and the value matrix V ; First, calculate the attention scores based on the query matrix Q and the key matrix K , and then multiply the attention scores by the value matrix V to obtain the attention weights; 1.5.3 Residual connection and layer normalization: Add the attention weights to the feature vector embeddings to obtain new attention weights, which are used as the input for layer normalization, and further enhance the feature expression ability through layer normalization again to obtain weighted features; 1.5.4 Feedforward neural network layer: The weight features are non-linearly transformed through a linear layer and a feedforward neural network layer. The feedforward neural network layer contains two fully connected layers, two Dropout layers, and an activation function GELU to achieve non-linear mapping. The output features are then added to the attention weights to obtain the remote sensing image features X1 and X2 of the front and back time phases respectively.

3. A spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT according to claim 2, wherein the step 1.1 data preprocessing comprises the following steps: 1.1.1 Segmentation operation: Use a convolutional layer with a kernel size of 16×16 and a stride of 16 to segment the 3-band pre-temporal and post-temporal remote sensing images with an input size of 1024×1024. The convolutional kernel sliding window covers an area of 16×16 each time and moves in a stride of 16, and segment the two 1024×1024 3-band remote sensing images into 64×64 3-band remote sensing image blocks with a size of 16×16; 1.1.2 Flattening operation: Flatten each 16×16 remote sensing image patch into a one-dimensional vector in channel order. Each remote sensing image patch contains 16×16×3 pixel values, and the length of the resulting remote sensing image vector after flattening is 768; 1.1.3 Mapping to a high-dimensional space: Use a fully connected layer to map the flattened remote sensing image patch vector to an embedding space of 768 dimensions, and the size of the generated high-dimensional embedding vector is 768×64×64; 1.1.4 Adding positional encoding: On the basis of the high-dimensional embedding vector, add positional encoding. The dimension of the positional encoding is the same as that of the high-dimensional embedding vector, both being 768 dimensions. After adding the positional encoding, a feature vector embedding is formed, and the size remains 768×64×64.

4. A spatio-temporal pixel feature progressive fusion remote sensing change detection method according to claim 2, wherein the layer normalization formula for feature extraction in step 1.5 is: , Among them, X represents the input of layer normalization, X LN represents the result after layer normalization, where μ and σ represent the mean and standard deviation respectively, and γ and β are learnable scale and translation parameters.

5. The spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT according to claim 2, wherein the query matrix Q , key matrix K and value matrix V in the multi-head self-attention mechanism layer of feature extraction in step 1.5 are calculated as follows: , Based on a query matrix Q and a key matrix K calculate attention scores ; The attention score calculation formula is: , Among them, d Denotes the dimension of the feature vector embedding.

6. A spatio-temporal pixel feature progressive fusion remote sensing change detection method according to any one of claims 1-5, wherein the method of the temporal dependence branch in step 2.2 comprises the following steps: 2.2.1 Input Feature Projection: Input Features are multiplied by a linear projection matrix and mapped to a high-dimensional space to obtain features The process is as follows: , W1 and W2 are linear projection matrices; 2.2.2 Convolution Feature Extraction: Projected Features X proj Through convolution operations, local temporal features are extracted X conv , and the process is as follows: , is a convolution operation; 2.2.3 Parameter Discretization: Local Temporal Features after Linear Projection Matrix and Layer Normalization X conv Multiply to obtain the time step , input mapping matrix B’ , output mapping matrix C , the process is as follows: , GroupNorm is layer normalization, W 3 is a linear projection matrix, is the sampling interval of the state space, is the state transition matrix, which constrains A’ to be a negative definite matrix. Use logarithmic parameterization A’ = −exp(logA’), where logA’ is a learnable parameter matrix, initialized to random values; is the mapping matrix from input to state; maps the state to the output, and 128 is the hidden state dimension; Convert the continuous parameter through the following formula into a discrete parameter , A’ is the state transition matrix, and the method is as follows: , , where I is the identity matrix; 2.2.4 Hidden state update: through discretized parameters Update the hidden state at the current moment h t , through the output mapping matrix C Map the hidden state to temporal features X t , the method is as follows: , , h t is the hidden state at the current moment, h t-1 is the hidden state at the previous moment; 2.2.5 Result Processing: Temporal Features X t and the feature are multiplied by the result after the activation function to obtain the feature X g ,X g After multiplying by the layer normalization and the linear projection matrix, the local features of the dual-temporal remote sensing image are obtained , and the method is as follows: , , The sigmoid function is used as the activation function, and GroupNorm is used for layer normalization. W 4 is the linear projection matrix.

7. A spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT according to claim 6, wherein the step 3.1 The method for enhancing progressive fusion of multi-level features includes the following steps: 3.1.1 First stage, shallow feature fusion enhancement: Connect feature with feature to obtain feature through a 1×1 convolutional layer. Add feature and feature to get feature ; 3.1.2 Second stage, middle-level feature fusion enhancement: Connect the feature with the feature , pass through a 1×1 convolutional layer and then add it to the feature to obtain the feature . Add the feature to the feature to obtain the feature . 3.1.3 Third stage, high-level feature fusion enhancement: Connect the feature with the feature , and then add it to the feature through a 1×1 convolutional layer to obtain the feature . Add the feature to the feature to obtain the feature , that is, the enhanced feature of the previous temporal remote sensing image F 1 .

8. A spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT as described in claim 7, wherein step 3.1 The mathematical expression of multi-level feature progressive fusion enhancement is expressed as: , , Represents the significant features of the early-phase remote sensing images output by the backbone network, , represents a 1×1 convolutional layer, represents a connection.

9. A spatio-temporal pixel feature progressive fusion remote sensing change detection method according to claim 8, wherein the feature fusion and upsampling process in step 4.1 is as follows: , Indicates a channel connection, which is a 1x1 convolution, indicates bilinear interpolation of features, and F represents the generated fused feature map.

Citation Information

Patent Citations

  • End-to-end infrared small target detection method based on Transform decoder network

    CN118505965A

  • Remote sensing image change detection method and system based on semantic fusion

    CN119068351A

Cited By

  • Remote sensing cultivated land change detection method

    CN122657735A