ViT-based space-time pixel feature progressive fusion remote sensing change detection method
By applying the ViT-based spatial and temporal pixel features a progressive fusion method in the field of remote sensing, the problem of insufficient feature consistency and differentiated representation capabilities in dual-time phase remote sensing image processing is solved, and higher detection accuracy and boundary clarity are achieved.
Patent Information
- Application Number
- CN202510554150.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-29
AI Technical Summary
In the field of remote sensing, when processing bi-time phase remote sensing images, it is difficult to effectively maintain the consistency of the characteristics of unchanged areas, and the differentiated characterization ability of heterogeneous features is limited, resulting in blurring of boundaries and loss of details of changing targets in complex backgrounds.
The ViT-based spatial-temporal pixel features progressive fusion remote sensing change detection method is used to extract the front and back time-phase remote sensing image features through low-rank fine-tuning, and combine the global context and timing-dependent branches to perform interaction enhancement of dual-time phase remote sensing features, and through multi-level feature progressive fusion enhancement, the loss of change information in the feature extraction process is reduced.
It significantly improves the clarity and detection accuracy of target boundaries in complex scenarios, enhances the generalization ability of the model, reduces information loss, and improves the overall performance of the change detection task.
Smart Images

Figure CN120088655A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image change detection, and particularly to a spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT. Background Art
[0002] Change detection is an important research field in remote sensing interpretation and is the key to observing and analyzing surface changes. Compared with traditional differential information extraction methods, deep learning-based methods have high robustness. They can not only stably process large amounts of data but also model complex change information, which has greatly improved the accuracy of change detection tasks. Deep learning methods represented by the Transformer architecture have injected new development impetus into this field. With its powerful global dependency capture ability and advantages in remote spatio-temporal relationship modeling, this architecture has broken through the bottleneck of the receptive field limitation of traditional convolutional neural networks and provided an innovative solution for change detection tasks that require high-level semantic understanding. However, the performance of the Transformer architecture highly depends on large-scale training data that matches the order of the number of parameters. The problem of scarce labeled data commonly existing in the remote sensing field seriously restricts the generalization ability of the model in complex and variable real scenarios.
[0003] Recently, vision transformers (ViTs), which are Transformer-based visual foundation models, have demonstrated excellent transfer learning capabilities through large-scale pre-training in the field of natural images, providing a new idea for solving the remote sensing data annotation bottleneck. Although the global representation of single-temporal remote sensing images constructed by such models through the dense attention mechanism has excellent feature expression capabilities, they still face great challenges when processing double-temporal remote sensing images: First, the collaborative enhancement mechanism for homogeneous features between cross-temporal remote sensing images is insufficient, making it difficult to effectively maintain the feature consistency of unchanged areas; second, the differential representation ability of heterogeneous features is limited, resulting in prominent problems such as blurred boundaries and lost details of changed targets in complex backgrounds. Therefore, making full use of the powerful feature expression ability of ViT, effectively migrating it to the remote sensing scenario, and at the same time enhancing the homogeneous and heterogeneous feature representation capabilities between double-temporal remote sensing images through method improvement to reduce the loss of change information during feature extraction has become a current research hotspot. Summary of the Invention
[0004] The purpose of the present invention is to provide a spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT that can significantly improve the clarity of target boundaries and detection accuracy in complex scenarios, effectively enhance the generalization ability of the model, and reduce information loss to solve the above problems.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions.
[0006] A spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT of the present invention comprises the following steps: S1 Visual basic model backbone low-rank fine-tuning: Use a low-rank matrix to perform low-rank fine-tuning on the visual basic model backbone, and extract the features X of the pre-temporal and post-temporal remote sensing images 1 , X 2 ; S2 Dual-temporal remote sensing feature interaction enhancement: Extract the significant features of the pre-temporal and post-temporal remote sensing image features X 1 , X 2 , which are respectively denoted as . Connect these two features by channel to obtain , where b is the batch size and d is the feature vector dimension; The dual-temporal remote sensing feature interaction enhancement includes two branches: 2.1 Global context branch: Input into the DenseMulti-Head Self-Attention layer to capture the global context information of the remote sensing image and obtain the global features of the dual-temporal remote sensing image ; 2.2 Temporal dependence branch: Input into the temporal dependence branch to capture the temporal dependence relationship and obtain the local features of the dual-temporal remote sensing image ; 2.3 Dual-temporal remote sensing feature decoupling based on channel splitting: Concatenate and project the global features of the dual-temporal remote sensing image and the local features of the dual-temporal remote sensing image to obtain the significant features of the dual-temporal remote sensing image . The significant features of the dual-temporal remote sensing image are evenly divided into two parts along the channel dimension, and are respectively extracted as the significant features of the pre-temporal remote sensing image and the significant features of the post-temporal remote sensing image ; S3 Multi-level feature progressive fusion enhancement: 3.1 Multi-level feature progressive fusion enhancement: Starting from the 12th layer to the 24th layer of the significant features of the pre-temporal remote sensing image , extract one feature every four layers to obtain multi-level features . Perform progressive fusion enhancement on to obtain the enhanced features of the pre-temporal remote sensing image F 1 , 3.2 Multi-level feature progressive fusion enhancement: Starting from the significant features of the post-temporal remote sensing image Starting from the 12th layer to the 24th layer, extract one feature every four layers to obtain multi-level features , and perform progressive fusion enhancement to obtain enhanced features of the post-temporal remote sensing image F 2 ; S4 Decoding of dual-temporal remote sensing image fusion features: 4.1 Feature fusion and upsampling: Connect the enhanced features of the pre-temporal remote sensing image F 1 and the enhanced features of the post-temporal remote sensing image F 2 by channels, then perform convolutional fusion, and upsample back to the input image size through bilinear interpolation; 4.2 Feature mapping and prediction: First, perform non-linear feature transformation on the fused feature map F through two fully connected layers and the GELU activation function; then multiply each pixel value of the feature map by the 1×1 convolutional parameter value through a 1×1 convolution (the 1×1 convolutional parameters are adaptively updated during the training process) to obtain a prediction result with a size of 1024×1024 and pixel values between [0-1]. Each pixel value in the prediction result represents the predicted probability that the pixel has changed; 4.3 Classification: Assign pixel values greater than or equal to 0.5 in the prediction result to 1, and pixel values less than 0.5 to 0 to obtain a binary classification prediction change image; a pixel value of 1 in the binary classification prediction change image represents that the predicted ground object has changed, and a pixel value of 0 represents that the predicted ground object has not changed.
[0007] The above-mentioned spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT, wherein the step S1 uses a low-rank matrix to perform low-rank fine-tuning on the backbone of the visual basic model, including the following steps: 1.1 Data preprocessing: Use 3-band pre-temporal and post-temporal remote sensing images with a size of 1024×1024 as input data for preprocessing to form feature vector embeddings with a size of 768×64×64; 1.2 Introduce low-rank matrix parameters: The backbone of the visual basic model uses a ViT (Vision Transformer) network architecture with 24 low-order fine-tuning Transformer layers. The Transformer encoder of ViT is stacked by a linear layer, a multi-head self-attention mechanism layer, and a feed-forward neural network layer; introduce two low-rank matrices A and B, in the multi-head self-attention mechanism layer, which are expressed as: ΔWr = B ⋅ A , where d is the dimension of the feature vector embedding, is a hyperparameter and is set to 16; during fine-tuning, the parameters of the original vision backbone model are frozen, and only the low-rank matrices A and B are updated; 1.3 Initialization strategy: At the beginning of training, matrix A is initialized according to a random Gaussian distribution, while matrix B is initialized as a zero matrix; 1.4 Parameter replacement and update: Replace the parameters of the linear layer in each Transformer encoder with the fine-tuned parameters: , , , , where W q , W k , W v , W d represent the original parameters of the query matrix Q, key matrix K, value matrix V, and the encoder output linear layer respectively; , , , represents the parameters after low-rank fine-tuning; the original weight matrix W is frozen and does not participate in gradient calculation, and only B and A are trained; 1.5 Feature extraction: 1.5.1 Layer normalization: Calculate the mean and variance of the feature vector embeddings obtained in step (1) on the dimension, and perform layer normalization (LayerNormalization) on them using the mean and variance; 1.5.2 Multi-head self-attention mechanism layer: Use the layer-normalized feature vector embeddings and the fine-tuned model parameters , , to multiply, and obtain the query matrix Q , key matrix K and value matrix V respectively; first calculate the attention scores based on the query matrix Q and the key matrix K , and then multiply the attention scores by the value matrix V to obtain the attention weights; 1.5.3 Residual connection and layer normalization: Add the attention weights to the feature vector embeddings to obtain new attention weights, which are used as the input of layer normalization, and further enhance the feature expression ability through layer normalization again to obtain weighted features; 1.5.4 Feed-forward neural network layer: The weighted features are non-linearly transformed through the linear layer and the feed-forward neural network layer (FFN). The feed-forward neural network layer contains two fully connected layers, two Dropout layers, and the activation function GELU for non-linear mapping. The output features are added to the attention weights to obtain the remote sensing image features X 1 , X2 。
[0008] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, wherein the step 1.1 data preprocessing includes the following steps: 1.1.1 Segmentation operation: Use a convolutional layer with a kernel size of 16×16 and a stride of 16 to segment the 3-band pre-temporal and post-temporal remote sensing images with an input size of 1024×1024. The convolutional kernel sliding window covers an area of 16×16 each time and moves in a stride of 16, dividing the two 1024×1024 3-band remote sensing images into 64×64 3-band remote sensing image patches of size 16×16; 1.1.2 Flattening operation: Flatten each 16×16 remote sensing image patch into a one-dimensional vector in channel order. Each remote sensing image patch contains 16×16×3 pixel values, and the length of the flattened remote sensing image vector is 768; 1.1.3 Mapping to a high-dimensional space: Use a fully connected layer to map the flattened remote sensing image patch vector to an embedding space of 768 dimensions, and the generated high-dimensional embedding vector size is 768×64×64; 1.1.4 Adding positional encoding: On the basis of the high-dimensional embedding vector, add positional encoding (PositionalEncoding). The dimension of the positional encoding is the same as that of the high-dimensional embedding vector, both are 768 dimensions. After adding the positional encoding, a feature vector embedding is formed, and the size is still 768×64×64.
[0009] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, wherein the formula for layer normalization in the feature extraction step 1.5 is: Among them, represents the input of layer normalization (here referring to the feature vector embedding), X LN represents the result after layer normalization, and represent the mean and standard deviation respectively, and are learnable scaling and translation parameters.
[0010] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, wherein the query matrix Q , key matrix K and value matrix V in the multi-head self-attention mechanism layer of the feature extraction step 1.5 are calculated as follows: Based on the query matrix Q and the key matrix K calculate the attention scores 。 The formula for calculating the attention scores is as follows: where d represents the dimension of the feature vector embedding.
[0011] The above-mentioned spatio-temporal pixel feature progressive fusion remote sensing change detection method based on ViT, wherein the method of the temporal dependence branch in step 2.2 includes the following steps: 2.2.1 Input feature projection: The input feature is multiplied by the linear projection matrix and mapped to a high-dimensional space to obtain the feature , , and the process is as follows: , is the linear projection matrix; 2.2.2 Convolutional feature extraction: The projected feature X proj undergoes a convolutional operation to extract local temporal features X conv , and the process is as follows: is the convolutional operation; 2.2.3 Parameter discretization: The linear projection matrix is multiplied by the locally temporally normalized feature X conv to obtain the time step , the input mapping matrix B , and the output mapping matrix C , and the process is as follows: GroupNorm is layer normalization, W 3 is the linear projection matrix, is the sampling interval of the state space, is the state transition matrix, and A is constrained to be a negative definite matrix. Use logarithmic parameterization A = −exp(logA), and logA is a learnable parameter matrix, initialized to random values; is the mapping matrix from the input to the state; maps the state to the output, and 128 is the hidden state dimension; The continuous parameter Convert to discrete parameters , is the state transfer matrix, the method is as follows: in, It is a unit array; 2.2.4 Hidden state update: by discretizing parameters Update the current hidden state h t , through the output mapping matrix C Mapping hidden states to temporal features X t , the method is as follows: h t is the hidden state at the current moment, h t-1 is the hidden state at the previous moment; 2.2.5 Result processing: time series characteristics X t With features The result after the activation function is multiplied to obtain the feature X g , X g After layer normalization and linear projection matrix multiplication, the local features of the dual-temporal remote sensing image are obtained. , the method is as follows: Sigmoid is the activation function, GroupNorm is the layer normalization, W 4 is the linear projection matrix.
[0012] The above-mentioned remote sensing change detection method based on ViT's spatiotemporal pixel feature progressive fusion, wherein the step 3.1 The method of progressive fusion enhancement of multi-level features includes the following steps: 3.1.1 The first stage, shallow feature fusion enhancement: the features With features Connect and get features through 1×1 convolution layer ,feature With features Add up the features ; 3.1.2 Second stage, intermediate feature fusion enhancement: Connect the feature with the feature , and then add it to the feature through a 1×1 convolutional layer to obtain the feature . Add the feature to the feature to obtain the feature ; 3.1.3 Third stage, high-level feature fusion enhancement: Connect the feature with the feature , and then add it to the feature through a 1×1 convolutional layer to obtain the feature . Add the feature to the feature to obtain the feature , that is, the enhanced feature F 1 of the previous temporal remote sensing image. .
[0013] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, where step 3.1 The mathematical expression of multi-level feature progressive fusion enhancement is: represents the significant features of the previous temporal remote sensing image output by the backbone network, , represents a 1×1 convolutional layer, represents connection.
[0014] The above-mentioned remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT, where the process of feature fusion and upsampling in step 4.1 is as follows: represents channel connection, is a 1x1 convolution, represents bilinear interpolation of the feature, and F represents the generated fused feature map.
[0015] Compared with the prior art, the present invention has obvious beneficial effects. As can be seen from the above technical solutions: The present invention performs low-rank fine-tuning based on a vision foundation model as a backbone network to extract remote sensing image features of two temporal phases before and after; adopts a global-local collaborative modeling strategy, utilizes the multi-head self-attention mechanism to capture the global context information of the remote sensing image, and through the temporal dependence branch, uses a state space model to establish the context relationship of local features, effectively taking into account the scene adaptation ability in the spatial dimension and the evolution analysis ability in the temporal dimension, realizing the selective retention of double-temporal related change features and the continuous tracking of the temporal evolution pattern, and enhancing the model's ability to represent homogeneous and heterogeneous features between double-temporal remote sensing images; at the same time, adopts a channel compression and residual learning mechanism to adaptively fuse the significant features of the pre-temporal remote sensing image and the significant features of the post-temporal remote sensing image at shallow, medium, and high levels, realizing three-stage cross-level feature interaction, reducing the loss of change information in the feature extraction process, and further improving the change detection accuracy. The present invention can significantly improve the recognition accuracy of change targets, effectively enhance the clarity of the boundaries of changed ground objects, detail retention, and generalization performance in complex scenes. It provides reliable surface change detection technical support for fields such as cultivated land protection, urban management, environmental monitoring, disaster assessment, and natural resource management. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of the present invention; Figure 2 is a flowchart of feature extraction for low-rank fine-tuning of the vision foundation model backbone of the present invention.
[0017] Figure 3 is a flowchart of enhanced interaction of double-temporal remote sensing features of the present invention.
[0018] Figure 4 is a flowchart of enhanced progressive fusion of multi-level features of the present invention.
[0019] Figure 5 is a comparison diagram of the effects of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following, in conjunction with the accompanying drawings and preferred embodiments, details the specific embodiments, structures, features, and effects of a remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT according to the present invention.
[0021] Embodiment 1: Refer to Figure 1 , a remote sensing change detection method based on progressive fusion of spatio-temporal pixel features of ViT of the present invention includes the following steps: S1 Visual foundation model backbone low-rank fine-tuning: To adapt to the change detection task while maintaining the performance of the visual foundation model backbone, a low-rank matrix is used to perform low-rank fine-tuning on the visual foundation model backbone, and the features X of the pre-temporal and post-temporal remote sensing images are extracted. 1 , X 2 .
[0022] The specific method is as follows: 1.1 Data preprocessing: The 3-band pre-temporal and post-temporal remote sensing images with a size of 1024×1024 are used as input data for preprocessing. The steps are as follows: 1.1.1 Segmentation operation: A convolutional layer with a kernel size of 16×16 and a stride of 16 is used to segment the input remote sensing image. The convolutional kernel sliding window covers an area of 16×16 each time and moves in a way with a stride of 16. The two 1024×1024 3-band remote sensing images are segmented into 64×64 blocks of 3-band remote sensing image blocks with a size of 16×16; 1.1.2 Flattening operation: Each 16×16 remote sensing image block is flattened into a one-dimensional vector in the channel order. Each remote sensing image block contains 16×16×3 pixel values. After flattening, the length of the remote sensing image vector obtained is 768; 1.1.3 Mapping to a high-dimensional space: A fully connected layer is used to map the flattened remote sensing image block vector to a 768-dimensional embedding space, and the size of the generated high-dimensional embedding vector is 768×64×64; 1.1.4 Adding positional encoding: On the basis of the high-dimensional embedding vector, positional encoding (PositionalEncoding) is added to provide spatial position information. The dimension of the positional encoding is the same as that of the high-dimensional embedding vector, both are 768 dimensions. After adding the positional encoding, a feature vector embedding is formed, and the size is still 768×64×64; 1.2 Introducing low-rank matrix parameters: The visual foundation model backbone adopts a ViT (Vision Transformer) network architecture with 24 low-order fine-tuning Transformer layers. The Transformer encoder of ViT is stacked by a linear layer, a multi-head self-attention mechanism layer, and a feed-forward neural network layer; Two low-rank matrices are introduced in the multi-head self-attention mechanism layer A and B, can be regarded as parametric learning of this process, expressed as: ΔWr = B ⋅ A , where d is the dimension of the feature vector embedding, is a hyperparameter, set to 16, which is used to control the rank size of the low-rank matrix; during the fine-tuning process, the parameters of the original visual foundation model backbone are frozen, and only the low-rank matrices A and B are updated; 1.3 Initialization Strategy: At the beginning of training, matrix A is initialized according to a random Gaussian distribution, while matrix B is initialized as a zero matrix; this initialization method ensures that in the initial stage of training, the low-rank update will not interfere with the original performance of the model.
[0023] 1.4 Parameter Replacement and Update: Replace the linear layer parameters in each Transformer encoder with the fine-tuned parameters: , , , , where W q , W k , W v , W d represent the original parameters of the query matrix Q, key matrix K, value matrix V, and the encoder output linear layer respectively; , , , represent the parameters after low-rank fine-tuning; the original weight matrix W is frozen and does not participate in gradient calculation, only B and A are trained; through the above low-rank fine-tuning, the vision base model can capture key information and adapt to the requirements of the remote sensing image change detection task.
[0024] 1.5 Feature Extraction: Perform feature extraction on the feature vector embedding obtained in step (1); the method is as follows (as Figure 2 shown): 1.5.1 Layer Normalization: Calculate the mean and variance of the feature vector embedding in the dimension, and perform layer normalization (LayerNormalization) on it using the mean and variance. The layer normalization formula is: where, represents the input of layer normalization (here refers to the feature vector embedding), X LN represents the result after layer normalization, and represent the mean and standard deviation respectively, and are learnable scaling and translation parameters; 1.5.2 Multi-Head Self-Attention Mechanism Layer: Multiply the feature vector embedding after layer normalization with the fine-tuned model parameters , , to obtain the query matrix Q , key matrix K and value matrix V respectively. The query matrix Q , key matrix K and value matrixV The calculation formula is: Based on the query matrix Q and the key matrix K calculate the attention scores. The attention score calculation formula is: where d represents the dimension of the feature vector embedding; Multiply the attention scores by the value matrix V to obtain the attention weights; 1.5.3 Residual connection and layer normalization: Add the attention weights to the feature vector embedding to obtain new attention weights, which are used as the input of layer normalization. Further enhance the feature expression ability through layer normalization again to obtain weighted features; 1.5.4 Feed-forward neural network layer: The weighted features are non-linearly transformed through the linear layer and the feed-forward neural network layer (FFN). The feed-forward neural network layer contains two fully connected layers, two Dropout layers, and the activation function GELU, which are used to implement non-linear mapping. The output features are added to the attention weights to obtain the remote sensing image features X 1 , X 2 ; S2 Dual-temporal Remote Sensing Feature Interaction Enhancement: As Figure 3 shown, extract the significant features of the front and back temporal remote sensing image features X 1 , X 2 , which are respectively represented as . Concatenate these two features by channel to obtain , where b is the batch size and d is the feature vector dimension; The dual-temporal remote sensing feature interaction enhancement contains two branches: 2.1 Global context branch: Input into the DenseMulti-Head Self-Attention layer to capture the global context information of the remote sensing image and obtain the global features of the dual-temporal remote sensing image ; 2.2 Temporal dependence branch: Input into the temporal dependence branch to capture the temporal dependence relationship and obtain the local features of the dual-temporal remote sensing image , while maintaining linear time complexity and accelerating training. The method is as follows: 2.2.1 Input feature projection: Multiply the input feature by the linear projection matrix to map it to a high-dimensional space to obtain the feature , , the process is as follows: , is a linear projection matrix; 2.2.2 Convolutional Feature Extraction: The projected features X proj Undergo a convolution operation to extract local temporal features X conv , the process is as follows: is the convolution operation; 2.2.3 Parameter Discretization: The linear projection matrix and the locally temporally normalized features X conv are multiplied to obtain the time step , the input mapping matrix B , the output mapping matrix C , the process is as follows: GroupNorm is layer normalization, W 3 is the linear projection matrix, is the sampling interval of the state space, is the state transition matrix (describing the transition relationship between states, used to determine how many hidden states should be propagated from the previous token to the next token). For stable training, A is constrained to be a negative definite matrix and logarithmic parameterization is used A = −exp(log A ) , log A is a learnable parameter matrix, initialized to random values; is the mapping matrix from input to state, used to determine how much input enters the hidden state; maps the state to the output, used to determine how the hidden state is transformed into the output, and 128 is the hidden state dimension; The continuous parameter is transformed into a discrete parameter , is the state transition matrix, describing the transition relationship between states, used to determine how many hidden states should be propagated from the previous feature to the next feature, the method is as follows: where is the identity matrix; 2.2.4 Hidden State Update: Through Discretized Parameters Update the hidden state at the current moment h t , through the output mapping matrix C Map the hidden state to temporal features X t , the method is as follows: h t is the hidden state at the current moment, h t-1 is the hidden state at the previous moment; 2.2.5 Result Processing: Temporal Features X t Multiply the result after the activation function of the temporal features with the features X g , X g After layer normalization and multiplication by the linear projection matrix, obtain the local features of the dual-temporal remote sensing image , the method is as follows: Sigmoid is the activation function, GroupNorm is the layer normalization, W 4 is the linear projection matrix; 2.3 Dual-Temporal Remote Sensing Feature Decoupling Based on Channel Splitting: Concatenate and project the global features of the dual-temporal remote sensing image with the local features of the dual-temporal remote sensing image to obtain the significant features of the dual-temporal remote sensing image , the significant features of the dual-temporal remote sensing image are evenly divided into two parts along the channel dimension, and are respectively extracted as the significant features of the former-temporal remote sensing image and the significant features of the latter-temporal remote sensing image ; Dual-temporal remote sensing feature interaction enhancement selectively retains features related to cross-temporal changes through a gating mechanism, suppresses background interference information, and simultaneously uses the long-term memory characteristics of the state space model to continuously track the temporal evolution pattern, significantly improving the discriminability of key features in the change detection task.
[0025] S3 Multi-Level Feature Progressive Fusion Enhancement: 3.1 Multi-level Feature Progressive Fusion Enhancement: Significant Features of Previous Temporal Remote Sensing Images Starting from the 12th layer until the 24th layer, extract one feature every four layers to obtain multi-level features , and perform progressive fusion enhancement (as shown in Figure 4 ) to obtain the enhanced features of the previous temporal remote sensing image F 1 . The method is as follows: 3.1.1 First stage, shallow feature fusion enhancement: Connect feature with feature and obtain feature through a 1×1 convolutional layer. Add feature to feature to obtain feature ; 3.1.2 Second stage, middle feature fusion enhancement: Connect feature with feature and add it to feature through a 1×1 convolutional layer to obtain feature . Add feature to feature to obtain feature ; 3.1.3 Third stage, high-level feature fusion enhancement: Connect feature with feature and add it to feature through a 1×1 convolutional layer to obtain feature . Add feature to feature to obtain feature , which is the enhanced feature of the previous temporal remote sensing image F 1 ; Through three-stage cross-level feature interaction and adopting the channel compression and residual learning mechanism, progressive fusion enhancement of multi-level feature information is achieved. Its mathematical expression can be represented as:: represents the significant features of the previous temporal remote sensing image output by the backbone network, , represents the 1×1 convolutional layer, represents connection; 3.2 Multi-level Feature Progressive Fusion Enhancement: From the significant features of the later temporal remote sensing image Starting from the 12th layer to the 24th layer, a feature is extracted every four layers to obtain multi-level features , and is subjected to progressive fusion enhancement (as shown in Figure 4 ), to obtain the enhanced features of the post-temporal remote sensing image F 2 (the method is the same as progressive fusion enhancement of multi-level features); S4 Decoding of the fused features of dual-temporal remote sensing images: 4.1 Feature fusion and upsampling: The enhanced features of the pre-temporal remote sensing image F 1 and the enhanced features of the post-temporal remote sensing image F 2 are concatenated by channel, then convolution fusion is performed, and upsampling is carried out by bilinear interpolation to restore to the size of the input image. The process is as follows: represents concatenation by channel, is a 1x1 convolution, represents bilinear interpolation of the features, and F represents the generated fused feature map; 4.2 Feature mapping and prediction: First, the fused feature map F is subjected to non-linear feature transformation through two fully connected layers and the GELU activation function; then, each pixel value of the feature map is multiplied by the 1x1 convolution parameter value through a 1×1 convolution (the 1x1 convolution parameters are adaptively updated during the training process) to obtain a prediction result with a size of 1024×1024 and pixel values between [0-1]. Each pixel value in the prediction result represents the prediction probability that the pixel has changed; 4.3 Classification: Pixel values greater than or equal to 0.5 in the prediction result are assigned a value of 1, and pixel values less than 0.5 are assigned a value of 0 to obtain a binary classification prediction change image; a pixel value of 1 in the binary classification prediction change image represents that the predicted ground object has changed, and a pixel value of 0 represents that the predicted ground object has not changed.
[0026] Experimental example: Evaluation of the effectiveness of the present invention A loss function is constructed based on the true change label and the binary classification prediction change image to screen the method for remote sensing change detection based on progressive fusion of spatio-temporal pixel features of the present invention, and the best-performing remote sensing change detection method in terms of evaluation indicators is obtained; the determination method of the loss function is the binary cross-entropy function, as follows: In the formula, represents the true label, represents the predicted label; Experiments were conducted on the LEVIR-CD dataset to verify the effectiveness of the present invention. This dataset contains 637 pairs of remote sensing images, and the size of each image is 1024 × 1024. The present invention follows the official standard and divides the dataset into three subsets: training, validation, and testing, which contain 445, 64, and 128 pairs of remote sensing images respectively. Each pair of remote sensing images includes a pre-temporal remote sensing image, a post-temporal remote sensing image, and the corresponding ground truth change label. The input size of the model is set to 1024×1024, and data augmentation techniques such as rotation, flipping, and random cropping are used to increase the sample size. The AdamW optimizer with a learning rate of 0.0004 and cosine annealing adjustment are used to decay the learning rate. The batch size is set to 4, and the maximum number of epochs is 300. The change detection results obtained on the LEVIR-CD dataset (as Figure 5 shown), the predicted map is compared with the ground truth. In addition, this experimental example evaluates the effectiveness of the method quantitatively. Table 1 shows the performance evaluation metrics of the experimental example on the LEVIR-CD dataset, which includes the F1 score (F1), intersection over union (IoU), and overall accuracy (OA), and their value ranges are all [0,1]. As can be seen from Table 1, the evaluation metrics of the detection method of the present invention are all relatively high, and it has reliable detection accuracy.
[0027] Table 1 Precision Verification of Change Detection Results F1(%) IoU (%) OA (%) This invention 92.46 86.1 99.2 The above description is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A remote sensing change detection method based on ViT based on progressive fusion of spatiotemporal pixel features, comprising the following steps: S1 Low-rank fine-tuning of the visual basic model backbone: Use a low-rank matrix to perform low-rank fine-tuning on the visual basic model backbone to extract the front-phase and back-phase remote sensing image features X1, X2; S2 Dual-phase remote sensing feature interactive enhancement: Extract the significant features of the remote sensing image features X1 and X2 before and after the time phase, expressed as , connect these two features by channel, and get , where b is the batch size and d is the feature vector dimension; the dual-temporal remote sensing feature interactive enhancement consists of two branches: 2.1 Global context branch: Input into the dense multi-head self-attention mechanism layer (Dense Multi-Head Self-Attention) to capture the global context information of the remote sensing image and obtain the global features of the dual-temporal remote sensing image ; 2.2 Timing-dependent branches: Input into the temporal dependency branch to capture the temporal dependency and obtain the local features of the dual-temporal remote sensing image ; 2.3 Decoupling of dual-temporal remote sensing features based on channel segmentation: Decoupling the global features of dual-temporal remote sensing images Local features of dual-temporal remote sensing images Stitching and projection to obtain the salient features of dual-temporal remote sensing images , significant features of dual-temporal remote sensing images The channel dimension is equally divided into two parts, and the salient features of the remote sensing image in the previous phase are extracted respectively. The significant features of remote sensing images in the later phases ; S3 multi-level feature progressive fusion enhancement: 3.1 Multi-level feature progressive fusion enhancement: salient features of remote sensing images from previous phases Starting from the 12th layer to the 24th layer, a feature is extracted every four layers to obtain multi-level features. ,Will Perform progressive fusion enhancement to obtain the enhanced features of the remote sensing image in the previous phase F 1 , 3.2 Multi-level feature progressive fusion enhancement: salient features from posterior-phase remote sensing images Starting from the 12th layer to the 24th layer, a feature is extracted every four layers to obtain multi-level features. ,Will Perform progressive fusion enhancement to obtain the enhanced features of post-phase remote sensing images F 2 ; S4 dual-temporal remote sensing image fusion feature decoding: 4.1 Feature fusion and upsampling: Enhance the features of the previous phase remote sensing image F 1 and post-phase remote sensing image enhancement features F 2 Connect by channel, then perform convolution fusion, and upsample to the input image size through bilinear interpolation; 4.2 Feature mapping and prediction: First, the fused feature map F is transformed nonlinearly through two fully connected layers and the GELU activation function. Then, each pixel value of the feature map is multiplied by the 1×1 convolution parameter value through a 1×1 convolution whose convolution parameter is adaptively updated during the training process to obtain a prediction result with a size of 1024×1024 and a pixel value between [0-1]. Each pixel value in the prediction result represents the predicted probability that the pixel has changed. 4.3 Classification: Assign pixel values greater than or equal to 0.5 in the prediction results to 1, and assign pixel values less than 0.5 to 0, to obtain a binary classification prediction change image; in the binary classification prediction change image, a pixel value of 1 represents that the predicted object has changed, and a pixel value of 0 represents that the predicted object has not changed.
2. A remote sensing change detection method based on ViT based on progressive fusion of spatiotemporal pixel features as claimed in claim 1, wherein the step S1 uses a low-rank matrix to perform a low-rank fine-tuning method on the backbone of the visual basic model, comprising the following steps: 1.1 Data preprocessing: The three-band pre-phase and post-phase remote sensing images with a size of 1024×1024 are preprocessed as input data to form feature vector embedding with a size of 768×64×64; 1.2 Introducing low-rank matrix parameters: The backbone of the visual base model adopts the ViT (Vision Transformer) network architecture with 24 low-rank fine-tuned Transformer layers. The Transformer encoder of ViT is composed of a stack of linear layers, multi-head self-attention mechanism layers, and feedforward neural network layers; two low-rank matrices are introduced in the multi-head self-attention mechanism layer A and B, Expressed as: ΔWr=B⋅A , where d is the feature vector embedding dimension and r is a hyperparameter set to 16; during fine-tuning, the parameters of the original visual base model backbone are frozen, and only the low-rank matrices A and B are updated; 1.3 Initialization strategy: At the beginning of training, matrix A is initialized according to a random Gaussian distribution, while matrix B is initialized to a zero matrix; 1.4 Parameter replacement and update: Replace the linear layer parameters in each Transformer encoder with the fine-tuned parameters: , , , , where W q , W k , W v , W d They represent the query matrix Q, key matrix K, value matrix V, and the original parameters of the encoder output linear layer respectively; Represents the parameters after low-rank fine-tuning; the original weight matrix W is frozen and does not participate in gradient calculation, only B and A are trained; 1.5 Feature extraction: 1.5.1 Layer Normalization: Calculate the mean and variance of the feature vector embedding obtained in step (1) in the dimension, and use the mean and variance to perform layer normalization (LayerNormalization); 1.5.2 Multi-head self-attention mechanism layer: using layer-normalized feature vector embedding and fine-tuning model parameters Multiply them to get the query matrix Q , key matrix K Sum Matrix V ; First, based on the query matrix Q and key matrix K Calculate the attention score and then add the attention score to the value matrix V Multiply them together to get the attention weight; 1.5.3 Residual connection and layer normalization: Add the attention weight to the feature vector embedding to obtain a new attention weight as the input of layer normalization. The feature expression ability is further enhanced by layer normalization again to obtain the weight feature. 1.5.4 Feedforward Neural Network Layer: Weight Features Pass Through Linear Layer The feedforward neural network layer (FFN) performs nonlinear transformation. The feedforward neural network layer contains two fully connected layers, two Dropout layers and an activation function GELU, which are used to realize nonlinear mapping. The output features are then added to the attention weights to obtain the remote sensing image features X1 and X2 of the previous and next phases, respectively.
3. A remote sensing change detection method based on ViT-based spatiotemporal pixel feature progressive fusion as claimed in claim 2, wherein the step 1.1 data preprocessing comprises the following steps: 1.1.1 Segmentation operation: Use a convolution layer with a convolution kernel size of 16×16 and a stride of 16 to segment the 3-band front-phase and post-phase remote sensing images with an input size of 1024×1024. The convolution kernel sliding window covers an area of 16×16 each time and moves with a stride of 16, segmenting the two 1024×1024 3-band remote sensing images into 64×64 blocks of 16×16 3-band remote sensing image blocks; 1.1.2 Flattening operation: Flatten each 16×16 remote sensing image block into a one-dimensional vector in channel order. Each remote sensing image block contains 16×16×3 pixel values. The length of the remote sensing image vector obtained after flattening is 768; 1.1.3 Mapping to high-dimensional space: Use a fully connected layer to map the flattened remote sensing image block vector to a 768-dimensional embedding space. The generated high-dimensional embedding vector size is 768×64×64. 1.1.4 Add positional encoding: Based on the high-dimensional embedding vector, positional encoding is added. The dimension of positional encoding is consistent with the high-dimensional embedding vector, both of which are 768 dimensions. After adding positional encoding, feature vector embedding is formed, and the size is still 768×64×64.
4. A remote sensing change detection method based on ViT-based spatiotemporal pixel feature progressive fusion as claimed in claim 2, wherein the layer normalization formula of the feature extraction in step 1.5 is: , in, X represents the normalized input of the layer, X LN represents the result after layer normalization, μ and σ represent the mean and standard deviation respectively, and γ and β are learnable scaling and translation parameters.
5. A remote sensing change detection method based on ViT-based spatiotemporal pixel feature progressive fusion as claimed in claim 2, wherein the query matrix in the multi-head self-attention mechanism layer of the feature extraction in step 1.5 Q , key matrix K Sum Matrix V The calculation formula is: , Based on the query matrix Q and key matrix K Calculating attention scores ; The attention score calculation formula is: , in, d Represents the dimension of feature vector embedding.
6. A remote sensing change detection method based on ViT based on progressive fusion of spatiotemporal pixel features as claimed in any one of claims 1 to 5, wherein the method of temporal dependent branching in step 2.2 is as follows: The following steps are involved: 2.2.1 Input feature projection: Input features Multiply with the linear projection matrix and map it to a high-dimensional space to get the features , the process is as follows: W1, W2 are linear projection matrices; 2.2.2 Convolutional feature extraction: features after projection X proj After convolution operation, local time series features are extracted X conv , the process is as follows: , is the convolution operation; 2.2.3 Parameter Discretization: Local Temporal Features after Linear Projection Matrix and Layer Normalization X conv Multiply to get the time step , input mapping matrix B , output mapping matrix C , the process is as follows: , GroupNorm is layer normalization, W 3 is the linear projection matrix, is the sampling interval of the state space, is the state transfer matrix, constrain A to be a negative definite matrix, use logarithmic parameterization A=−exp(logA), logA is a learnable parameter matrix initialized to random values; is the mapping matrix from input to state; Map the state to the output, 128 is the hidden state dimension; The continuous parameters are expressed as follows: Convert to discrete parameters , A is the state transfer matrix, the method is as follows: , , Among them, I is the unit matrix; 2.2.4 Hidden state update: by discretizing parameters Update the current hidden state h t , through the output mapping matrix C Mapping hidden states to temporal features X t , the method is as follows: , , h t is the hidden state at the current moment, h t-1 is the hidden state at the previous moment; 2.2.5 Result processing: time series characteristics X t With features The result after the activation function is multiplied to obtain the feature X g ,X g After layer normalization and linear projection matrix multiplication, the local features of the dual-temporal remote sensing image are obtained. , the method is as follows: , , Sigmoid is the activation function, GroupNorm is the layer normalization, W 4 is the linear projection matrix.
7. A remote sensing change detection method based on ViT-based spatiotemporal pixel feature progressive fusion as claimed in claim 6, wherein the step 3.1 The method of progressive fusion enhancement of multi-level features includes the following steps: 3.1.1 The first stage, shallow feature fusion enhancement: the features With features Connect and get features through 1×1 convolution layer ,feature With features Add up the features ; 3.1.2 The second stage, mid-level feature fusion enhancement: the features With features Connect, pass through 1×1 convolution layer and then with feature Add up the features ,feature With features Add up the features ; 3.1.3 The third stage, high-level feature fusion enhancement: With features Connect, pass through 1×1 convolution layer and then with feature Add up the features ,feature With features Add up the features , namely the enhancement feature of the front phase remote sensing image F 1 .
8. A remote sensing change detection method based on ViT-based spatiotemporal pixel feature progressive fusion as claimed in claim 7, wherein the step 3.1 The mathematical expression of multi-level feature progressive fusion enhancement is: , , represents the salient features of the previous phase remote sensing image output by the backbone network, , represents a 1×1 convolutional layer, Indicates a connection.
9. A remote sensing change detection method based on ViT-based spatiotemporal pixel feature progressive fusion as claimed in claim 8, wherein the feature fusion and upsampling process in step 4.1 is as follows: , Indicates channel connection, is a 1x1 convolution, It represents bilinear interpolation of features, and F represents the generated fusion feature map.
Citation Information
Patent Citations
End-to-end infrared small target detection method based on Transform decoder network
CN118505965A
Remote sensing image change detection method and system based on semantic fusion
CN119068351A
Remote sensing image change detection method based on adaptive Transform and deformable convolution
CN119418204A
Remote sensing image land classification method based on SAM multi-order fine tuning
CN119494988A
Hyperspectral remote sensing image classification method based on self-attention context network
WO2022073452A1
Cited By
Remote sensing image building change detection system and method based on parallel branch feature interaction
CN120411812A
Optical remote sensing image salient target detection method based on progressive attention enhancement
CN120894536A
Low-rank adaptation-based few-sample coronal mass ejection segmentation method and system
CN121616830A
Remote sensing image adaptive enhancement method based on segmented hybrid mapping
CN122115293A