A hyperspectral image classification method based on TRANSFORMER feature fusion
Through the feature fusion method based on TRANSFORMER, combined with 3D-2D convolutional neural network and Transformer encoder and decoder, the problem of difficulty in effectively mining spectral and spatial information in hyperspectral images is solved in the existing technology, and the classification accuracy of hyperspectral images has been significantly improved.
Patent Information
- Application Number
- CN202210144122.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-17
AI Technical Summary
The existing hyperspectral image classification methods are difficult to effectively explore spectral and spatial information in the image, resulting in low classification accuracy.
A hyperspectral image classification method based on TRANSFORMER feature fusion is proposed. Through the space-spectral information mining module and the Transformer-based feature fusion module, combined with the 3D-2D convolutional neural network and the Transformer encoder and decoder, the spectral and spatial information are fused to generate higher-level features.
It significantly improves the classification accuracy of hyperspectral images, can more effectively explore spectral and spatial information in the images, and improves the accuracy of classification results.
Smart Images

Figure CN114627370B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a hyperspectral image classification method, in particular to a hyperspectral image classification method based on TRANSFORMER feature fusion, and belongs to the technical field of remote sensing information processing. Background Art
[0002] Hyperspectral images contain hundreds of continuous spectral bands, which carry rich spectral information and can be used in crop analysis, geological mapping, mineral exploration, national defense research, urban measurement, military monitoring and other fields. However, due to the complex imaging mechanism and large amount of data of hyperspectral images, the classification of hyperspectral images has brought great challenges.
[0003] Deep learning methods are currently the most popular hyperspectral classification methods. Compared with traditional hand-crafted features, deep features are more abstract, robust, and not affected by local changes. Therefore, many researchers are committed to optimizing the basic network framework. For example, Li et al. proposed an effective classification framework that uses a deep belief network (DBN) to extract deep spectral features and adopts an active learning algorithm to iteratively select high-quality labeled samples as training samples. Subsequently, Zhong et al. proposed an improved DBN model in which a multivariate DBN and a standardized fine-tuning procedure were used in the pre-training stage. However, DBN is mainly used to extract spectral information and loses a lot of spatial information. The same problem also exists in one-dimensional convolutional neural networks (1D CNN), which only act on a single spectral pixel and ignore the spatial correlation around the pixel. In addition, recurrent neural networks (RNNs) have also been used to learn the dependencies between channels in spectral sequences. For example, Mou et al. first applied RNNs to hyperspectral image classification, in which RNNs used a new activation function to analyze hyperspectral sequence data.
[0004] However, this discards the spatial information contained in the original HSI. To this end, Zhang et al. further improved the one-dimensional CNN and designed a one-dimensional CNN-ConvcapsNet. It enables it to extract spatial features. In addition, two-dimensional CNN (2D CNN) has gradually become the mainstream backbone for hyperspectral classification tasks to extract representative spatial features. Lee et al. proposed a contextual CNNs, which can optimize the exploration of local contextual interactions by jointly utilizing the local spatial-spectral relationship of adjacent single pixels. However, 2D CNN only performs convolution operations on (height and width) dimensions. This makes it unable to fully mine spectral and spatial information from different dimensions. Ying et al. proposed three-dimensional CNNs (3D CNNs), which do not rely on any pre-processing or post-processing. At the same time, it can effectively extract spatial features. However, 3D CNN is more complex and computationally expensive than 2D CNN. In order to reduce the complexity of 3D CNN and make better use of it. Roy et al. combined 2D CNN and 3D CNN to construct HybridSN, in which 3D CNN is used to extract joint spatial-spectral features, while 2D CNN further learns more abstract spatial representations.
[0005] After feature extraction, fusion is another important step for classification tasks. Traditional fusion is implemented at the feature level (early fusion) or at the decision score level (late fusion). For feature-level fusion, deep features are fused by directly stacking them. For late fusion. The most common strategy is unified weight fusion. It assigns a uniform weight value to each probability matrix. However, none of them takes into account the correlation between features. So we turned our attention to the field of natural language processing (NLP), whose models are more flexible. Transformer is a seq2seq model (consisting of an encoder and a decoder) that has been widely used. Transformer is a multi-head model that has been integrated with CNN for different tasks. For example, Zhang et al. proposed a novel dual-branch architecture (i.e., TransFuse). It consists of a CNN branch and a Transformer branch, which are then fused in the BiFusion module. Another approach is that the spatial-spectral transformer (SST) is constructed by fusing CNN and Transformer. Meanwhile, Qing et al. combined the attention mechanism with the encoder of Transformer to build an end-to-end spatial-spectral transformer model. The model uses spectral attention mechanism and self-attention mechanism to extract spectral and spatial features of HSI respectively. However, most works use encoder to extract features and then directly input them into the classifier, and the function of Transformer decoder has not been fully explored.
[0006] In this context, we proposed a hyperspectral image classification method based on TRANSFORMER feature fusion, which more effectively mines the spectral and spatial information in the hyperspectral image, and at the same time uses Transformer to extract features for fusion. In subsequent classification problems, the use of fused features can significantly improve the classification accuracy. Summary of the invention
[0007] The purpose of the present invention is to improve the classification accuracy of hyperspectral remote sensing images, and a hyperspectral image classification method based on TRANSFORMER feature fusion (expressed as TransCNN) is proposed.
[0008] This paper proposes a hyperspectral image classification method based on TRANSFORMER feature fusion, which consists of three modules: spatial-spectral information mining, Transformer-based feature fusion and prediction. The technical details of the three modules are as follows:
[0009] Spatial-spectral information mining module: This module consists of a three-channel 3D-2D convolutional neural network. Since hyperspectral data has a high degree of redundancy, principal component analysis of hyperspectral data is required before feature extraction. After principal component analysis, the image cube is transposed twice to obtain three image cubes of different sizes (including the original image cube). Through this operation, the correlation between channel and height, channel and width, and height and width can be fully mined. In other words, if the size of the original image is B×M×N, then after transposition, its size will become B×N×M, B×M×N, and M×B×N. Subsequently, the acquired transposed images are respectively input into the three-channel 3D-2D convolutional network to mine the spatial-spectral information of the image.
[0010] Transformer-based feature fusion module: First, the features obtained above are input into the semantic tagger to convert them into a one-dimensional sequence. Then, the attention mechanism is used to recalculate the spatial-spectral information features. The weighted average of the pixels in the image is used to obtain a compact L sequences (i.e., {1, 2, 3…, L}). At the same time, one-dimensional position embedding is used to embed each Embed position information. Then, the sequence containing the position embedding can be Input to the Transformer Encoder to extract higher-level deep correlation information and obtain feature sequences , will get After cascading, the feature map obtained from the spectral spatial information mining module is input into the Decoder module of the Transformer for fusion to obtain the fused spatial-spectral feature map. X fusion .
[0011] Prediction module: It consists of convolutional layers and softmax layers, which integrate the fusion features obtained from the decoder. X fusion Directly input into the prediction module for classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 , implementation block diagram of the hyperspectral image classification method based on TRANSFORMER feature fusion.
[0013] Figure 2 , spatial-spectral information mining module.
[0014] Figure 3 , Transformer-based feature fusion module.
[0015] Figure 4 , the effects of different training ratios and spatial window sizes. (a) Indian Pines overall classification accuracy; (b) Pavia University overall classification accuracy; (c) Salinas overall classification accuracy.
[0016] Figure 5 , Loss curves of the training phase on different datasets. a) Indian Pines; (b) Pavia University; (c) Salinas.
[0017] Figure 6 ,Visual comparison of classification maps of Indian Pines dataset obtained by different methods. (h)CTF; (i)TDF; (j) Proposed method. DETAILED DESCRIPTION
[0018] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only embodiments of a part of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present invention.
[0019] The method proposed in this invention includes three steps: spatial-spectral information mining, Transformer-based feature fusion, and prediction. The specific analysis steps are as follows:
[0020] Step S1: Spatial-spectral information mining
[0021] Step S1.1: Perform principal component analysis (PCA) on the original hyperspectral image;
[0022] Step S1.2: image transposition processing;
[0023] make represents the image cube obtained after PCA, where B, M, and N are the number of channels, height, and width, respectively. and is the image cube after transposing X. represents the original X, and the formula is as follows:
[0024]
[0025] Step S1.3: spatial-spectral information mining;
[0026] Will , , As the input of three convolutional neural networks, each convolutional neural network is as follows Figure 2 As shown in Figure 2, all three convolutional neural networks are composed of three-dimensional and two-dimensional convolutional neural networks, but the size of the convolution kernel is adjusted according to the input image.
[0027] In experiments, the height and width are usually set equal ( M=N ), which will result in the same size of the two image cubes. In other words, the two convolutional neural networks are composed of the same network. The specific network composition and parameter settings of the three convolutional neural network pipelines are summarized in Table 1.
[0028] Table 1
[0029]
[0030] Taking the first convolutional neural network as an example, the size of the original input is 30×25×25, and after transposition, the sizes become 30×25×25 and 25×30×25 respectively. If we set the size of the 3D convolution kernel in the first convolutional neural network to 7×3×3, 5×3×3, and 3×3×3 respectively, the image size will be reshaped to 32×18×9×9 after passing through the three-dimensional convolutional neural network. Then, it is further set to 576×9×9 to input the two-dimensional convolutional neural network of three convolutional layers (whose kernels are all set to 3×3). Therefore, the size of the final output feature map is 16×3×3. The other convolutional neural network pipeline setting is similar to the first one. In addition, we also added an attention mechanism to the 3D-2D convolutional network, namely one channel attention and one spatial attention, which can better extract spectral and spatial features.
[0031] let , , Represents the feature map obtained after passing through the three-channel convolutional neural network, where H, M, and C represent the height, width, and channel of the feature map, respectively.
[0032] (i=1, 2, 3)
[0033] in Represents the i-th convolutional neural network channel.
[0034] Step S2: Transformer-based feature fusion
[0035] Step S2.1: Convert features into sequences;
[0036] set up is passed through a semantic tagger , , The obtained sequences can be expressed as:
[0037]
[0038]
[0039] in Represents a learnable kernel Point-wise convolution, is used to standardize the feature map .
[0040] Step S2.2: Obtain higher-level association information;
[0041] In order to further obtain We input the high-level association information of the three Transformers into three encoders, which have the same structure consisting of a multi-head self-attention (MSA) and a multi-layer perception module (MLP).
[0042] However, unlike the iterative operations in RNN, Transformer performs parallel processing, so Position encoding can be done by To calculate, Represents position encoding. Unlike the original Transformer, our proposed structure is pre-regularized (PreNorm). That is, the normalization layer (NL) occurs before the multi-head self-attention mechanism. Some studies have shown that pre-regularization is more efficient and stable.
[0043] Suppose there are three input variables (query Q, key K, value V), which can be obtained by The specific formula is as follows:
[0044]
[0045]
[0046]
[0047]
[0048] in, , , are three linear matrices, d Indicates the number of channels contained in the image. Q, K, V ) is used when we use dot product for calculation. For the multi-head attention mechanism, Q, K, V First, a linear transformation is performed and then input into the dot product module. This process will execute h Second-rate, h Indicates the number of longs.
[0049] The multi-head attention mechanism enables the model to obtain input different subspaces of . Through the multi-head self-attention mechanism, and The residual part of is input into the next normalization layer and multi-layer perception. Considering that the multi-layer perception is mainly composed of a fully connected layer, a GELU activation function and a downsampling layer. We can get the following formula:
[0050]
[0051] Where Lin(.) is the linear transformation, GELU(.) is the activation function, and Dro(.) is used for downsampling.
[0052] Step S2.3: spatial-spectral feature fusion;
[0053] For the decoder block, we use it to perform feature fusion, such as Figure 3 As shown. We added three feature maps , that is, obtained from the three-channel CNN as Q value input, the encoder output As K Value and V Value input.
[0054] Obviously, the decoding part fuses the spatial information obtained from CNN and the deep high-order sequence obtained from the Transformer encoder, and uses the fused features as the input of the final classifier. It should be noted here that the difference between the multi-head attention mechanism (MA) and the multi-head self-attention mechanism (MSA) is that the MSA Q, K, V From the same sequence, while MA K and V From ,and Q Generated from CNNs , calculated as follows:
[0055]
[0056]
[0057] in is the linear projection matrix.
[0058] Step S3: Prediction
[0059] The fused features obtained from the decoder X fusion Directly input the prediction module for classification. It can be expressed as:
[0060] Prediction = Softmax(Conv(X fusion )) .
[0061] In order to illustrate the effectiveness of the present invention, the following experimental demonstration is conducted. The experimental data comes from the AVIRIS hyperspectral remote sensing image of Indian agriculture and forestry in northwest Indiana, USA, acquired in 1992, which contains There are 16 pixels, 220 bands, and 20 bands are removed due to noise and other factors.
[0062] Set the experimental environment parameters of the present invention:
[0063] Ten independent tests were performed using Pytorch on a computer with an Intel Core i5 processor equipped with an RTX3090 GPU.
[0064] The first set of experiments focused on the ratio of training samples ( r ) and the input image patch neighborhood width ( s ) for the performance of the present invention, the experimental results are as follows Figure 4 As shown. The proportion of training data in the total data (denoted by r) and the spatial window size of the input image (denoted by s) are two important factors affecting the final performance. In this test, the number of principal components is set to 30. We randomly select 1%, 3%, 5%, 7%, 10%, 13%, 15% and 20% of all samples as the training sample ratio of the dataset. Spatial window size s Variations in {21×21, 23×23, 25×25, 27×27, 29×29} datasets. Figure 4 It can be seen that when r changes from 1% to 20%, although the classification accuracy fluctuates slightly, the general trend is that the classification accuracy will increase with the increase of r. In addition, when r is set to 10%, the classification performance converges to a relatively stable value. Therefore, in subsequent experiments, we set r equal to 10% (ie, r=10). Figure 4 (a) It can be seen that when s=25×25, the overall accuracy (OA) of the Indian Pines dataset can reach the maximum value (98.36%). The experiment shows that the performance of the invention is related to the proportion of training samples ( r ) and the input image patch neighborhood width ( s ), so determining these two parameters is crucial to the present invention.
[0065] The second group of experiments focuses on studying the changes in the loss function during the training phase of the present invention, such as Figure 5 As shown. In this experiment, the initial learning rate is set to 0.001, the cross entropy function is selected to train the entire network, and the Adam optimization function is selected. The number of training times is initially set to 300. It can be seen that after 300 training cycles, the loss function is close to convergence. For this reason, we selected 300 training epochs for the Indian Pines dataset.
[0066] The third set of experiments focuses on comparing the performance of the present invention with other classification methods, and the experimental results are summarized in Table 2. Table 2 compares the classification accuracy provided by different methods using 10% training samples (Indian Pines).
[0067] Table 2
[0068]
[0069] The comparison methods include the basic hyperspectral image classification methods support vector machine (SVM), two-dimensional convolutional network (CNN), three-dimensional convolutional network (3DCNN), HybridSN, deep pyramid residual network (DPRN), and the recently emerged Transformer and Vision Transformer (ViT). The comparison method CNN-Transformer (CTF) is a cascade structure of CNN and Transformer. But it does not include the fusion framework between the encoder and the decoder, and CNN and Transformer are simple serial structures. In addition, in order to verify the effectiveness of transposition and fusion, we also constructed a comparison method called dual data fusion (TDF), which only fuses the original image cube (B×N×M) and the transposed image cube (B×M×N). In this test, we set the training ratio r to 10%. Analyzing Table 2, we can draw the following results: 1) Our proposed method OA is 19.76%, 23.52%, 10.51% and 13.42% higher than SVM, CNN, 3D CNN and ViT, respectively, and about 1% higher than Transformer, HybridSN and DPRN. 2) The best performance of the proposed method is OA=98.36%, AA=98.11% and Kappa=98.13%, indicating that our proposed architecture can respond well to data information. 3) The performance of CTF, TDF and the proposed method is relatively better than other reference methods. This is because they are based on the combination of CNN and Transformer, and the transposition operation can well mine spectral and spatial information. At the same time, Figure 6 The classification diagram of the Indian Pines dataset is shown in Figure 2. Obviously, there is a lot of noise in (h), and the effect of (i) is not ideal. It can be clearly seen that the classification diagram of this method is clearer.
[0070] The above is only a specific implementation method of the present application, and does not limit the present application in any form. Any simple modification, equivalent change or modification made to the above implementation method based on the technical essence of the present application still falls within the protection scope of the technical solution of the present application.
Claims
1. A hyperspectral image classification method based on TRANSFORMER feature fusion, characterized in that: The following steps are involved: S1: Principal component analysis (PCA) processing of the original hyperspectral image; S2: Image transposition processing, let X∈R B×M×N represents the image cube obtained after PCA, where B , M and N are the number of channels, height and width respectively, X 1 and X 3 is transposed X The sizes of the image cubes are B×N×M and M×B×N , X 2 Represents the original X ; S3: X 1 , X 2 , X 3 As the input of the three-channel convolutional neural network, the spatial-spectral information mining is carried out. The three convolutional neural networks are composed of three-dimensional and two-dimensional convolutional neural networks. The feature maps of the spatial-spectral information mined by the three-channel convolutional neural network are represented as X 1new , X 2new , X 3new ; S4: The feature map described in S3 X 1new , X 2new , X 3new , passed through a semantic tagger to convert into a sequence T 1 , T 2 and T 3 ; S5: T 1 , T 2 and T 3 The three encoders input into the three Transformers have the same structure, consisting of multi-head self-attention (MSA) and multi-layer perception module (MLP) to obtain higher-level correlation information of features of different dimensions. T 1new , T 2new and T 3new ; S6: Cascading spatial-spectral information feature maps X new =concat{ X 1new , X 2new , X 3new }, while concatenating deep correlation sequence information T new =concat{ T 1new , T 2new , T 3new }; S7: Use the Transformer decoder module to perform feature fusion. X new and T new Fusion, obtain fusion features X fusion , more effectively utilize the spectral and spatial characteristics of images; S8: Fusion features obtained from the decoder X fusion Directly input into the prediction module for classification.
2. The hyperspectral image classification method based on TRANSFORMER feature fusion according to claim 1, characterized in that: The formula for image transposition processing in step S2 is: X 1 , X 3 = Transpose ( X ), Transpose ( . ) indicates a transposition.
3. The hyperspectral image classification method based on TRANSFORMER feature fusion according to claim 1, characterized in that: The size of the convolution kernel of the three-channel convolutional neural network in step S3 needs to be adjusted according to the input image size.
4. The hyperspectral image classification method based on TRANSFORMER feature fusion according to claim 1, characterized in that: In step S4, the semantic tagger is used to pass X 1new , X 2new , X 3new Get sequence T 1 , T 2 and T 3 , the formula is: T i =(A i ) T X inew A i =(σ(φ(X inew ; W))) T Where φ(.) represents a learnable kernel W∈R C×L Point-wise convolution, σ(.) is used to normalize the feature map A i ∈R (HW)×L .
Citation Information
Patent Citations
Hyperspectral image classification method based on deep learning
CN106845418A
Hyperspectral image classification method
CN112348097A