A Hyperspectral Image Change Detection Method Based on Spectral Transformer Network
The depth characteristics of high-spectral images are extracted through the spectral Transformer network and explored the spectral and spatial correlation, which solves the problem of insufficient utilization of spectral information in the existing methods, and achieves higher accuracy change detection.
Patent Information
- Application Number
- CN202310053970.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-03
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-02-03
AI Technical Summary
The existing hyperspectral image change detection methods fail to effectively utilize the spectral information and spatial information of hyperspectral images, resulting in low detection accuracy and failure to fully express the correlation of the characteristics of the changing region.
Using a hyperspectral image change detection method based on the spectral Transformer network, depth features are extracted through Unnet and combined with the Transformer encoder to fuse spectral information, differential flow and point multiplication flow to explore the feature correlation, and the final change detection result diagram is generated.
Improve the accuracy and robustness of change detection, and improve the detection accuracy and performance.
Smart Images

Figure CN116958802B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing information processing, and specifically relates to change detection of hyperspectral images. It aims to detect the changes of ground objects at the same location in different times through dual-temporal hyperspectral images, and can be applied to fields such as urban management, natural disaster monitoring, and forest monitoring. Background Technique
[0002] Hyperspectral image change detection aims to detect the changes of ground objects at the same location through dual-temporal hyperspectral images, and then analyze the changed areas from quantitative and qualitative perspectives. Hyperspectral images have hundreds or even thousands of spectral bands, which contain rich spectral and detailed information and can be used to detect or identify targets. Therefore, hyperspectral image change detection has gradually become a popular research field in the field of change detection.
[0003] Hyperspectral image change detection methods can be roughly divided into traditional change detection methods and deep learning-based methods. Traditional change detection methods mainly include algebra-based methods such as image difference and image ratio; transformation-based methods such as change vector analysis (CVA) and principal component analysis (PCA); classification-based methods and other methods such as Markov models and fuzzy clustering. Most traditional change detection methods rely on artificial features or shallow features and cannot well express the features of ground objects, resulting in low accuracy of change detection.
[0004] Due to its ability to extract abstract, hierarchical, and high-level semantic features, deep learning technology has been widely applied in the field of computer vision. Therefore, deep learning technology has been applied to remote sensing change detection tasks. Many deep learning structures have been used to process hyperspectral images, such as feature pyramid networks, Siamese networks, Unet networks, etc. However, these change detection methods do not pay attention to the extraction of features of changed areas, and the effectiveness of features of changed areas is more important for improving detection performance. Change detection methods based on the attention mechanism can focus on the changed areas of interest and extract more representative features through the generated spatial or channel attention maps. Long-range dependencies have been proven to improve detection performance. Transformer-based change detection methods capture the spatial structure relationship of images by focusing on spatial long-range dependencies and extract discriminative features.
[0005] Although the above-mentioned Transformer-based change detection methods have achieved good detection results, most of these methods only introduce the Transformer mechanism to mine the positional dependence of spatial information from the spatial dimension, and do not use the Transformer mechanism to mine the spectral dependence of spectral information from the spectral dimension, ignoring the positive effect of spectral dependence on the change detection task. However, due to the large number of spectral bands in hyperspectral images, the spectral information of hyperspectral images is more important for detecting changes. In addition, most deep learning-based change detection methods only use differential features to express the correlation between the extracted deep features, and do not explore the correlation between dual-temporal hyperspectral images from other perspectives, resulting in the problem of insufficient comprehensive expression of feature correlation. Summary of the Invention
[0006] In view of the above deficiencies of the prior art, the present invention proposes a hyperspectral image change detection method based on a spectral Transformer network. The hyperspectral image change detection network based on the spectral Transformer network includes a feature extraction module, a Transformer-based correlation expression module, and a detection module. The feature extraction module uses Unet to extract the deep features of dual-temporal hyperspectral images, and fuses the Transformer encoder in the skip connection to fuse the spectral information. The Transformer-based correlation expression module explores the correlation of the extracted deep features from two perspectives: the difference flow and the dot product flow. Both the difference flow and the dot product flow use the Transformer encoder to fuse the spectral information of each band to obtain the correlation of spectral features. Finally, a weighted fusion operation is performed on the change detection result map based on the difference flow, the change detection result map based on the dot product flow, and the change detection result map based on the cascade flow to generate the final change detection result map, making full use of the learned features, thereby improving the accuracy and robustness of change detection.
[0007] The method includes the following steps: A hyperspectral image change detection method based on a spectral Transformer network, the method includes the following steps:
[0008] Step 1: Data segmentation;
[0009] Step 2: Divide the training set and the test set;
[0010] Step 3: Deep feature extraction. Input the image patches pair of the training set into the UnetTrans feature extraction module. The UnetTrans feature extraction module consists of two components: Unet and Transformer encoder. Among them, Unet is composed of an encoder, a decoder, and skip connections. It obtains more representative features through continuous upsampling and downsampling, and its characteristics can better represent abstract and global information. For the input dual-temporal hyperspectral image patches pair, the image patch at time T1 is denoted as The image patch at time T2 is denoted as C is the number of bands of the hyperspectral image. After passing through Unet, the deep feature can be obtained from T1. After passing through Unet, the deep feature can be obtained from T2. The Transformer encoder is composed of L identical sub-encoders connected in series. Each sub-encoder consists of a multi-head self-attention mechanism (MSA) and a multi-layer perceptron (MLP). For the input feature 1 ≤ k ≤ L, H is the height of the feature Z k-1 and W is the width of the feature Z k-1 . C is the number of bands of the hyperspectral image. After passing through MSA, Z' k can be obtained. After passing through MLP, the output can be obtained. Finally, the final output after passing through L sub-encoders is
[0011] Step 4: Differential feature and dot product feature extraction. Perform pixel-by-pixel subtraction and multiplication operations on the feature maps F T1 and F T2 to obtain the differential feature and the dot product feature. Then, send the difference feature into the Transformer encoder to obtain the feature Send the dot product feature into the Transformer encoder to obtain the feature to mine the effective information in the spectral dimension;
[0012] Step 5: Generate the difference flow, feature flow, and cascaded flow change detection result maps;
[0013] Step 6: Generate the final change detection result map;
[0014] Step 7: Optimization of the spectral Transformer network;
[0015] Step 8: Testing process. Obtain the optimized model and detect the test samples.
[0016] Further Step 1: Data segmentation includes dividing the hyperspectral image data pixel by pixel through a 7×7 sliding window to split the dataset, and the image at time T1 is split to generate a series of image patches P 1,i , where i = 1,..., N, and the image at time T2 is split to generate a series of image patches P 2,i , where i = 1,..., N.
[0017] Furthermore, for each image patch, the height is denoted by H and the width by W, where the value of H is 7 and the value of W is 7.
[0018] Further, Step 3: Each sub-encoder operation can be expressed as:
[0019] Z′ k = MSA(LN(z k-1 )) + z k-1 (1)
[0020] Z k = MLP(LN(z′ k )) + z′ k (2)
[0021] Among them, LN represents LayerNorm, MSA represents the multi-head self-attention mechanism, MLP represents the multi-layer perceptron, and the calculation formula of the self-attention mechanism (self-attention) is as follows:
[0022]
[0023] Among them, represents Query, represents Key, represents Value. Q, K, and V are all learnable parameters obtained from the input, d is the channel dimension. σ represents the activation function. Then, MSA can be expressed as:
[0024] MSA(Q, K, V) = Concat(head1,..., head h )W O (4)
[0025]
[0026] Among them, is the linear projection matrix, represents Query, represents Key, represents Value, Q, K, and V are all learnable parameters obtained from the input, Concat is the concatenation operation, Attention is the attention mechanism, head iDenote the $i$-th head in the multi-head self-attention mechanism, and $h$ is the number of heads set to 8.
[0027] Further step four: The process is as follows:
[0028]
[0029]
[0030] where TE represents the Transformer encoder, $\alpha_1$ and $\alpha_2$ represent the weights of the difference flow and the dot product flow respectively, and the output $F$ of this module DS and are obtained after the reshape operation, is the difference, is the dot product.
[0031] Further set $\alpha_1 = 10$ and $\alpha_2 = 1$.
[0032] Further step five: First, pass the difference flow feature $F$ DS and the multiplication flow feature $F$ MS through the softmax function respectively to generate the change map CM DS of the difference flow and the change map of the multiplication flow DS and
[0033] Then, concatenate the difference flow feature $F$ DS and the multiplication flow feature $F$ MS and input the concatenated features into the cascade flow. Input the concatenated features into a fully connected network and generate the change map of the concatenated features through the softmax function which can be expressed as:
[0034] CM con = σ(FC(concat(F DS , F MS ))) (8)
[0035] where concat(F DS , F MS ) represents the concatenation of the feature F DS and the feature F MS , FC represents the fully connected network, and σ represents the softmax function.
[0036] Further step six: Generate the final change detection result map, which includes: weighted sum and aggregation of the change maps detected from the difference flow, multiplication flow, and cascade flow to generate the final change map D total , which can be expressed as:
[0037]
[0038] Among them, w1, w2, and w3 represent the weights of the differential flow change diagram, the multiplication flow change diagram, and the cascade flow change diagram, respectively.
[0039] Furthermore, set w1 = 0.4, w2 = 0.1, and w3 = 0.5.
[0040] The further step seven includes that the loss function used is a composite loss function, mainly including two terms: binary cross-entropy loss and contrast loss. The total loss can be written as:
[0041] Loss = λ1L BCE + λ2L CE (10)
[0042] Among them, L BCE is the binary cross-entropy loss, and L CE is the contrast loss. λ1 and λ2 are the penalty parameters of L BCE and L CE respectively, where λ1 = 1 and λ2 = 0.25; the binary cross-entropy loss is the cross-entropy loss of a binary classification problem, and the formula is as follows:
[0043]
[0044] Among them, N is the batch size of the samples, y i is the label information, y i = 0 indicates dissimilarity, and y i = 1 indicates similarity. is the predicted value; the mathematical expression of the contrast loss is written as follows:
[0045]
[0046] Among them, y is the label, y = 0 indicates dissimilarity, and y = 1 indicates similarity; d is the Euclidean distance between the features of two samples; N is the batch size of the samples, and margin represents the threshold, whose value is 0.5. Table 1 shows the parameters of each neural network in the UnetTrans feature extraction network.
[0047] Table 1 Parameters of each neural network in the UnetTrans feature extraction network
[0048]
[0049] Description of the Drawings
[0050] Figure 1 Flowchart of the hyperspectral image change detection method based on the spectral Transformer network
[0051] Figure 2 Transformer Encoder Flowchart
[0052] Figure 3 The three datasets and their corresponding ground truths used in the present invention are specifically as follows: (1) SantaBarbara dataset, (2) BayArea dataset, and (3) Hermiston City dataset. Detailed Implementation Manner
[0053] To more clearly illustrate the technical solution implemented by the present invention, the following further describes in detail the steps of implementing the present invention with reference to the accompanying drawings.
[0054] Refer to Figure 1 , the steps of implementing the present invention are as follows:
[0055] Step 1: Slice the hyperspectral image data of the dataset pixel by pixel through a 7×7 sliding window to generate image patches pairs.
[0056] Step 2: Input the image pairs into the UnetTrans feature extraction module in Figure 1 , and we can respectively obtain the deep features F T1 and
[0057] Step 3: Input the feature vectors obtained in Step 2 into the Transformer Encoder. After performing a reshape operation on the obtained feature vectors, use formulas (6) and (7) to obtain F DS and
[0058] Step 4: Pass F DS and through the softmax function respectively to obtain the differential flow change map and the dot multiplication change map. Pass F DS and through formula (8) to obtain the cascaded features, and then obtain the change map of the cascaded features through the softmax function.
[0059] Step 5: Perform weighted sum aggregation on the change maps detected from the differential flow, multiplication flow, and cascaded flow to generate the final change map.
[0060] In the change detection results, if an unchanged sample pair is detected as a changed sample pair, it is considered a (False Positive, FP). Detecting a changed sample pair as a changed sample pair is considered a (True Positive, TP). Detecting an unchanged sample pair as an unchanged sample pair is considered a (True Negative, TN). Misdetecting a changed sample as an unchanged sample is denoted as (False Negative, FN). The overall accuracy (OA) is calculated as follows:
[0061]
[0062] The Kappa coefficient is calculated as follows:
[0063]
[0064]
[0065] The F1 score is the harmonic mean of precision and recall. Therefore, this criterion can balance the effects of precision and recall and reflect the average performance of the model. These formulas can be written as follows:
[0066]
[0067] Simulation experiment:
[0068] 1. Simulation conditions:
[0069] Training and testing were performed using a single NVIDIA GeForce GTX 3080Ti GPU with 12G of memory. The datasets used were the three hyperspectral datasets as shown Figure 2 below.
[0070] 2. Simulation content:
[0071] To prove the effectiveness of the present invention, six relatively popular and newly proposed methods in recent years were compared in the experiment. The quantitative detection accuracies OA, Kappa values, precision values, and F1 values of the algorithm proposed in the present invention and the other six comparison algorithms on the three databases. As shown in Tables 2, 3, and 4
[0072] Table 2 Four algorithm metrics for the Santa Barbara dataset
[0073] Method OA kappa precision F1 CVA 0.8677 0.4216 0.3345 0.4812 RCVA 0.8987 0.4926 0.4007 0.5417 DSAMNet 0.9602 0.6797 0.7593 0.7008 SSAN 0.9462 0.6851 0.5769 0.7129 PA-Former 0.9814 0.8746 0.7970 0.8845 BIT 0.9802 0.8657 0.7934 0.8762 ours 0.9825 0.8765 0.8319 0.8859
[0074] Table 3 Four algorithm metrics for the BayArea dataset
[0075]
[0076]
[0077] Table 4 Four algorithm metrics of Hermiston City dataset
[0078] Method OA kappa precision F1 CVA 0.9241 0.5959 0.4828 0.6336 RCVA 0.925 0.5893 0.4857 0.6271 DSAMNet 0.968 0.8697 0.8038 0.8881 SSAN 0.9701 0.8102 0.706 0.8259 PA-Former 0.9881 0.9487 0.9195 0.9555 BIT 0.9895 0.9541 0.9343 0.9601 ours 0.9910 0.9607 0.9471 0.9658
[0079] As can be seen from Tables 2 - 4, the effectiveness of the present invention has been verified on three real hyperspectral datasets. The experimental results show that the present invention achieves the best detection results in most cases compared with other existing methods.
Claims
1. A hyperspectral image change detection method based on a spectral Transformer network, the method comprising the following steps: Step 1: Data segmentation; Step 2: Divide the training set and the test set; Step 3: Deep feature extraction. Input the image patches pair of the training set into the UnetTrans feature extraction module. The UnetTrans feature extraction module consists of two components: Unet and Transformer encoder. Among them, Unet is composed of an encoder, a decoder, and skip connections. It obtains more representative features through continuous upsampling and downsampling, and its characteristics can better represent abstract and global information. For the input dual-temporal hyperspectral image patches pair, the image patch at time T1 is denoted as The image patch at time T2 is denoted as C is the number of bands of the hyperspectral image. After passing through Unet, the deep feature can be obtained from T1. After passing through Unet, the deep feature can be obtained from T2. The Transformer encoder is composed of L identical sub-encoders connected in series. Each sub-encoder consists of a multi-head self-attention mechanism (MSA) and a multi-layer perceptron (MLP). For the input feature 1 ≤ k ≤ L, H is the height of the feature Z k-1 and W is the width of the feature Z k-1 . C is the number of bands of the hyperspectral image. After passing through MSA, Z' k can be obtained. After passing through MLP, the output can be obtained. 1 ≤ k ≤ L. Finally, the final output after passing through L sub-encoders is Step 4: Differential feature and dot product feature extraction. Perform pixel-by-pixel subtraction and multiplication operations on the feature maps F T1 and F T2 to obtain differential features and dot product features. Then, send the difference features to the Transformer encoder to obtain features Send the dot product features to the Transformer encoder to obtain features Mine the effective information in the spectral dimension; Step 5: Generate difference flow, feature flow and cascade flow change detection result maps; Step 6: Generate the final change detection result map; Step 7: Optimization of the spectral Transformer network; Step 8: Testing process, obtain the optimized model, and detect the test samples.
2. The hyperspectral image change detection method based on the spectral Transformer network according to claim 1, characterized in that Step 1: Data segmentation includes dividing the hyperspectral image data pixel by pixel through a 7×7 sliding window to split the dataset, and splitting the image at time T1 to generate a series of image patches Pi, where i = 1,..., N. Splitting the image at time T2 generates a series of image patches Pj, where j = 1,..., N. 1,i , where i = 1,..., N, and splitting the image at time T2 generates a series of image patches P 2,i , where i = 1,..., N.
3. The hyperspectral image change detection method based on the spectral Transformer network according to claim 2, characterized in that The height of each image block is represented by H, and the width is represented by W, where the value of H is 7 and the value of W is 7.
4. A hyperspectral image change detection method based on a spectral Transformer network according to claim 1, characterized in that The step 3: Each sub-encoder operation can be expressed as: Z′ k = MSA(LN(z k-1 )) + z k-1 (1) Z k = MLP(LN(z′ k )) + z′ k (2) Where, LN represents LayerNorm, MSA represents multi-head self-attention mechanism, MLP represents multi-layer perceptron, and the calculation formula of the self-attention mechanism (self-attention) is as follows: Among them, denotes Query, denotes Key, denotes Value. Q, K, and V are all learnable parameters obtained from the input. d is the channel dimension, and σ represents the activation function. Then, MSA can be expressed as: MSA(Q, K, V) = Concat(head1,..., head h )W O (4) Among them, is a linear projection matrix, represents Query, represents Key, represents Value. Q, K, and V are all learnable parameters obtained from the input. Concat is a concatenation operation, and Attention is an attention mechanism. head i represents the i-th head in the multi-head self-attention mechanism, and h is the number of heads set to 8.
5. A hyperspectral image change detection method based on a spectral Transformer network according to claim 1, characterized in that The step 4: The process is as follows: Among them, TE represents the Transformer encoder, α1 and α2 represent the weights of the differential flow and the dot product flow respectively, and the output F of this module DS and is obtained after the reshape operation, is the difference, is the dot product.
6. A hyperspectral image change detection method based on a spectral Transformer network according to claim 5, characterized in that Set α1 = 10 and α2 = 1.
7. A hyperspectral image change detection method based on a spectral Transformer network according to claim 1, characterized in that Step Five: First, the differential flow feature F DS and the multiplication flow feature F MS are respectively passed through the softmax function to generate the change map CM DS of the differential flow and the change map of the multiplication flow, where CM DS and Then, the differential flow feature F DS and the multiplication flow feature F MS are subjected to a feature concatenation operation and then input into the cascaded flow. The concatenated feature is input into a fully connected network, and a change map of the concatenated feature is generated through the softmax function which can be expressed as: CM con = σ(FC(concaat(F DS ,F MS ))) (8) Among them, concat(F DS , F MS ) represents the concatenation of feature F DS and feature F MS . FC represents a fully connected network, and σ represents the softmax function.
8. A hyperspectral image change detection method based on a spectral Transformer network according to claim 1, characterized in that Step six: generating the final change detection result map includes: performing weighted summation on the change maps detected from the difference flow, multiplication flow, and cascade flow to generate the final change map D total , which can be expressed as: Where, w1, w2 and w3 respectively represent the weights of the difference flow change map, the multiplication flow change map and the cascade flow change map.
9. A hyperspectral image change detection method based on a spectral Transformer network according to claim 8, characterized in that Set w1 = 0.4, w2 = 0.1 and w3 = 0.
5.
10. A hyperspectral image change detection method based on a spectral Transformer network according to claim 1, characterized in that The step 7 includes that the loss function used is a composite loss function, including two terms: binary cross-entropy loss and contrast loss, and the total loss can be written as: Loss=λ1L BCE +λ2L CE (10) Among them, L BCE is the binary cross-entropy loss, and L CE is the contrastive loss. λ1 and λ2 are the penalty parameters of L BCE and L CE respectively, where λ1 = 1 and λ2 = 0.
25. The binary cross-entropy loss is the cross-entropy loss for binary classification problems, and the formula is as follows: where N is the batch number of samples, and y i is the label information, where y i = 0 indicates dissimilarity, and y i = 1 indicates similarity, is the predicted value; the mathematical expression of the contrast loss is written as follows: Where, y is the label, y = 0 indicates dissimilarity, y = 1 indicates similarity; d is the Euclidean distance between the features of two samples; N is the batch number of samples, and margin represents the threshold, and its value is 0.5.
Citation Information
Patent Citations
Hyperspectral image classification method based on Transform enhanced non-local U-shaped network
CN114445665A
Remote sensing image change detection method
CN114881916A