A hyperspectral image classification method based on scale interaction transformer
By extracting multi-scale features from hyperspectral images through directional separable convolution and scale interaction Transformer modules, the problem of insufficient utilization of spectral features in existing methods is solved, and higher classification accuracy and ground feature recognition capabilities are achieved.
Patent Information
- Application Number
- CN202410804064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-06-20
AI Technical Summary
Existing Transformer-based hyperspectral image classification methods do not fully utilize multi-scale spectral features, resulting in insufficient classification accuracy and making it difficult to meet the needs of real-world applications.
A directional separation convolution module is used to extract local and non-local scale spectral information, and a scale interaction Transformer module is used to realize global correlation modeling from multiple scale perspectives. The classification is then performed in conjunction with a fully connected network.
It improves the extraction efficiency and classification accuracy of spectral features, achieving higher ground cover identification capabilities and more accurate land cover classification.
Smart Images

Figure CN118608861B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and hyperspectral remote sensing image processing technology, and in particular to a hyperspectral image classification method based on scale-interactive Transformer. Background Technology
[0002] This remote sensing technology, with its excellent performance, demonstrates great potential in ground feature identification and monitoring. Hyperspectral remote sensing imaging systems can acquire images containing two-dimensional spatial distribution information and one-dimensional continuous narrow-band spectral information of ground features. The acquired images not only reveal the spatial geometric distribution of ground features on the observed surface, but also allow for the extraction of the radiation intensity reflected by the ground features by analyzing the continuous narrow-band spectral information. Due to the differences in the chemical composition of different substances, they each possess unique spectral characteristics; the unique spectral information of different ground features in hyperspectral images can be used for the identification of these features.
[0003] The main goal of hyperspectral image classification is to classify each pixel in an image into different land cover categories using the spectral and spatial geometric information of the land cover being measured. Hyperspectral image classification is the most important step in realizing many hyperspectral remote sensing applications because the classification results directly affect the accuracy of hyperspectral data acquisition, thus impacting subsequent application effectiveness. Machine learning-based hyperspectral image classification methods typically first use feature extraction or feature selection techniques to obtain the spectral or spatial features of the image, and then use a conventional classifier to complete the classification. Because machine learning methods heavily rely on expert knowledge for manually designed features, and the models only have shallow structures with poor feature representation capabilities, they often struggle to achieve satisfactory classification accuracy. In recent years, deep learning technologies, especially convolutional neural networks (CNNs), have made significant breakthroughs in many computer vision tasks. Sellami et al. proposed a fusion of 3D-CNNs for spatial-spectral classification of hyperspectral images, aiming to integrate multiple 3D-CNNs applied to a set of similar spectral bands. Compared to machine learning feature extraction methods, CNNs can automatically extract discriminative deep features from images in an end-to-end manner and complete the image-to-label mapping through a classifier. Although CNN-based methods have achieved good results in hyperspectral image classification, they still have some limitations when applied to this field. First, due to the high spectral dimensionality of hyperspectral images, the limited receptive field of convolutional operations makes it difficult to capture long-range inter-spectral band correlation features.
[0004] Therefore, Transformer-based techniques are currently being applied to hyperspectral image classification tasks to extract long-range correlation features. In 2021, Hong et al. published a paper titled "SpectralFormer: Rethinking hyperspectral image classification with transformers" in IEEE Trans. Geosci. Remote Sens., disclosing a classification method using Transformer to extract spectral features, achieving advanced classification performance. However, this method lacks consideration of global weights across multiple receptive fields when calculating global correlation weights, and cannot fully extract the spectral features of hyperspectral images to obtain more accurate classification results, thus failing to meet the needs of real-world applications. Summary of the Invention
[0005] Purpose of the invention: The purpose of this invention is to address the shortcomings of the existing technology by proposing a hyperspectral image classification method based on scale-interactive Transformer. First, it proposes a directional separation convolution module to efficiently extract the multi-scale spectral features of hyperspectral images. Then, it uses the constructed scale-interactive Transformer to achieve global correlation extraction from multiple scale perspectives, thereby solving the problem that existing Transformer-based hyperspectral image classification methods do not fully utilize multi-scale spectral features.
[0006] To achieve the above objectives, the specific implementation steps of the present invention include the following:
[0007] (1) A hyperspectral image X with size M×N×B was selected from the hyperspectral image library, where M and N represent the width and height of the image, respectively, B represents the number of bands of the image, M and N are both positive integers greater than 0, and B is a positive integer greater than or equal to 100.
[0008] (2) Extract each pixel and its corresponding spectral information from the hyperspectral image X. Perform planarization on each pixel;
[0009] (3) Use the directional separation convolution module DSCB to extract local scale spectral information M spe_A Nonlocal scale spectral information M spe_NA ;
[0010] (4) For feature M spe_A and M spe_NA Serialize them separately to meet the input requirements of the scale interaction Transformer module;
[0011] (5) Feature M spe_A and M spe_NAInputting into the scale-interaction Transformer module yields a global feature correlation model from multiple scale perspectives;
[0012] (6) The output of the scale interaction Transformer module is passed through a fully connected network to obtain the classification result of each pixel;
[0013] In step (2), the pixel x i The planarization process is implemented as follows:
[0014] (2a) in x i The spectral dimension is obtained by padding the ends with p zero elements. To meet Where p satisfies the following formula:
[0015]
[0016] in, This represents the floor function;
[0017] (2b) the above Remodeling
[0018] In step (3), local scale spectral features M are extracted using the directional separation convolution module. spe_A Nonlocal scale spectral features M spe_NA Its implementation is as follows: the directional separation convolution module includes two branches, and the local scale spectral feature extraction branch uses a two-dimensional convolutional layer with a kernel size of 1×k from the description in claim 2. Extracting local-scale spectral features The nonlocal scale spectral feature extraction branch uses a two-dimensional convolutional layer with a kernel size of k×1 from the feature extraction described in claim 2. Extracting local-scale spectral features The two-dimensional convolutional layer consists of two-dimensional convolution operations, batch normalization operations, and the Mish activation function.
[0019] In step (4), feature M spe_A and M spe_NA Serialization is performed separately, as follows:
[0020] (4a) M spe_A and M spe_NA Convert to one-dimensional sequence and in
[0021] (4b) Using a linear layer and Mapping the feature dimension to D dimensions yields... and
[0022]
[0023]
[0024] Where ω is a learnable parameter,
[0025] (4c) will and Learnable class tags and Connect them, and then add position code PE. spe ,in The sequence used as input to subsequent Transformers is obtained:
[0026]
[0027]
[0028] Step (5) specifically includes the following steps:
[0029] (5a) z spe_A The sequence is linearly mapped to Q. spe_A K spe_A V spe_A Three matrices, z spe_NA The sequence is linearly mapped to Q. spe_NA K spe_NA V spe_NA Three matrices;
[0030] (5b) Construct a scale-interactive multi-head attention module SiMHA, which is composed of multiple scale-interactive attention layers SiA stacked together. Each SiA takes two feature sequences reflecting information at different scales as inputs and calculates the single-scale attention weight SAW of the two sequences. 1 / 2 The results of the scale-interactive multi-head attention module (SIMHA) are weighted stepwise and then output as two weighted feature sequences. The calculation formula for SIMHA is as follows:
[0031]
[0032]
[0033] MAW = Softmax((V spe_A +V spe_NA (V) spe_A +V spe_NA )),
[0034] SiA spe_A =SiAttention(Q spe_A K spe_A V spe_A V spe_NA ) = MAW(SAW spe_A (V spe_A )),
[0035] SiA spe_NA =SiAttention(Q spe_NA K spe_NA V spe_NA V spe_A ) = MAW(SAW spe_NA (V spe_NA )),
[0036]
[0037] Softmax() is the activation function used to obtain the weights, and Concat() is used to concatenate different headers. is the feature dimension of K1 or K2, h is the number of stacked SiA arrays, and w is a learnable parameter used to map the dimension of the merged SiA output to the dimension of the original input feature vector.
[0038] (5c) The weighted features are input into the layer normalization LN layer and the multilayer perceptron MLP layer. The LN layer is for stabilizing the feature distribution. The MLP layer consists of an input layer, a GELU activation function, and an output layer. Both the input and output layers are composed of fully connected networks. Skip connections are established between the input of the LN layer and the output of the MLP to alleviate gradient vanishing. Finally, the first dimension of the output of the residual connection and the one-dimensional feature vector are used as the output of the scale interaction Transformer.
[0039] Step (6) specifically includes the following steps:
[0040] (6a) will obtain and Perform splicing along the direction of the second-dimensional feature;
[0041] (6b) Input the concatenated features into an MLP layer to obtain the final classification result.
[0042] Compared with the prior art, the present invention has the following advantages:
[0043] First, the multi-scale spectral feature extraction strategy used in this invention reduces the phase redundancy of spectral features at different scales by using directional separation convolution after planarizing the spectral features.
[0044] Second, compared with existing methods, the scale-interactive Transformer proposed in this invention adds modeling of the correlation of multiple global spectral bands from multiple scale perspectives, thereby improving the efficiency of spectral feature extraction.
[0045] Simulation results show that the present invention has higher classification accuracy than existing hyperspectral image classification algorithms. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 A schematic diagram of the method flow provided for an example of the present invention;
[0048] Figure 2 This is a schematic diagram of the spectral planarization process proposed in this invention.
[0049] Figure 3 This is the intent of the directional separation convolution module proposed in the example of the present invention;
[0050] Figure 4 This is a structural diagram of the scale-interaction Transformer module proposed in the example of this invention;
[0051] Figure 5 This is a diagram of the model structure for scale-interactive Transformer-based hyperspectral image classification proposed in this invention.
[0052] Figure 6 This is a pseudo-color image of a real hyperspectral image used in the embodiments of the present invention and the corresponding truth label image;
[0053] Figure 7 This is a simulation of the land cover classification of a real hyperspectral image using the existing SpectralFormer hyperspectral image classification method that uses Transformer to extract spectral features.
[0054] Figure 8 This is a simulation image of the land cover classification of the same real hyperspectral image using the method of this invention. Detailed Implementation
[0055] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0056] like Figure 1 As shown, the present invention provides a hyperspectral image classification method based on scale-interactive Transformer, comprising the following steps.
[0057] Step 1: Obtain the hyperspectral image X to be classified. A hyperspectral image X with dimensions M×N×B is selected from the hyperspectral image library, where M and N represent the width and height of the image, respectively, and B represents the number of bands in the image. M and N are both positive integers greater than 0, and B is a positive integer greater than or equal to 100. In this example, the hyperspectral image X to be detected is a real hyperspectral image of Pavia University acquired by the ROSIS imaging spectrometer, which has 103 spectral bands, a size of 610×340, and 9 land cover of interest to be classified.
[0058] Step 2: Extract each pixel and its corresponding spectral information from the hyperspectral image X. Each pixel undergoes planarization, specifically implemented as follows:
[0059] (2.1) In x i The spectral dimension is obtained by padding the ends with p zero elements. To meet Where p satisfies the following formula:
[0060]
[0061] in, This represents the floor function;
[0062] (2b) the above Remodeling
[0063] Step 3: Use the Directional Separating Convolutional Module (DSCB) to extract local scale spectral information M. spe_A Nonlocal scale spectral information M spe_NA Its specific implementation is as follows: the directional separation convolution module includes two branches. The local scale spectral feature extraction branch uses a two-dimensional convolutional layer with a kernel size of 1×k from the description in claim 2. Extracting local-scale spectral features The nonlocal scale spectral feature extraction branch uses a two-dimensional convolutional layer with a kernel size of k×1 from the feature extraction described in claim 2. Extracting local-scale spectral features The two-dimensional convolutional layer consists of two-dimensional convolution operations, batch normalization operations, and the Mish activation function.
[0064] Step 4, for feature M spe_A and M spe_NA Sequencing is performed separately to meet the input requirements of the scale interaction Transformer module. The specific implementation method is as follows:
[0065] (4a) M spe_A and M spe_NA Convert to one-dimensional sequence and in
[0066] (4b) Using a linear layer and Mapping the feature dimension to D dimensions yields... and
[0067]
[0068]
[0069] Where ω is a learnable parameter,
[0070] (4c) will and Learnable class tags and Connect them, and then add position code PE. spe ,in The sequence used as input to subsequent Transformers is obtained:
[0071]
[0072]
[0073] Step 5, Feature M spe_A and M spe_NA The input is fed into the scale-interaction Transformer module to obtain a global feature correlation model from multiple scale perspectives, and the implementation method is as follows:
[0074] (5a) z spe_A The sequence is linearly mapped to Q. spe_A K spe_A V spe_A Three matrices, z spe_NA The sequence is linearly mapped to Q. spe_NA K spe_NAV spe_NA Three matrices;
[0075] (5b) Construct a scale-interactive multi-head attention module SiMHA, which is composed of multiple scale-interactive attention layers SiA stacked together. Each SiA takes two feature sequences reflecting information at different scales as inputs and calculates the single-scale attention weight SAW of the two sequences. 1 / 2 The results of the multi-scale attention weighted MAW are weighted stepwise, and two weighted feature sequences are output. The calculation formula of the scale-interactive multi-head attention module SiMHA is as follows:
[0076]
[0077]
[0078] MAW = Softmax((V spe_A +V spe_NA (V) spe_A +V spe_NA )),
[0079] SiA spe_A =SiAttention(Q spe_A K spe_A V spe_A V spe_NA ) = MAW(SAW spe_A (V spe_A )),
[0080] SiA spe_NA =SiAttention(Q spe_NA K spe_NA V spe_NA V spe_A ) = MAW(SAW spe_NA (V spe_NA )),
[0081]
[0082] Softmax() is the activation function used to obtain the weights, and Concat() is used to concatenate different headers. is the feature dimension of K1 or K2, h is the number of stacked SiA arrays, and w is a learnable parameter used to map the dimension of the merged SiA output to the dimension of the original input feature vector.
[0083] (5c) The weighted features are input into the layer normalization LN layer and the multilayer perceptron MLP layer. The LN layer is for stabilizing the feature distribution. The MLP layer consists of an input layer, a GELU activation function, and an output layer. Both the input and output layers are composed of fully connected networks. Skip connections are established between the input of the LN layer and the output of the MLP to alleviate gradient vanishing. Finally, the first dimension of the output of the residual connection and the one-dimensional feature vector are used as the output of the scale interaction Transformer.
[0084] Step 6: The output of the scale interaction Transformer module is passed through a fully connected network to obtain the classification result for each pixel. The specific implementation method is as follows:
[0085] (6a) will obtain and Perform splicing along the direction of the second-dimensional feature;
[0086] (6b) Input the concatenated features into an MLP layer to obtain the final classification result.
[0087] The effects of the present invention will be further explained below with reference to simulation experiments.
[0088] 1. Simulation conditions:
[0089] The simulation experiments of this invention were conducted in a hardware environment with a GTX1660 GPU and 32GB of memory, and a software environment including PyCharm and Torch 1.9.1.
[0090] The simulation experiment of this invention uses a real hyperspectral image of Pavia University acquired by the ROSIS imaging spectrometer with a reflective optical system as the target for detection. This image has 103 spectral bands, a size of 610 pixels × 340 pixels, and contains 9 ground features of interest to be classified. The pseudo-color image of this real hyperspectral image is shown below. Figure 6 As shown in (a), the ground truth label image attached to the real hyperspectral image is as follows: Figure 6 As shown in (b) Figure 6 (b) Different colors represent different land cover categories. Table 1 lists the names of all land cover categories and the sample size for each category.
[0091] Table 1. Land cover names and sample sizes at Pavia University
[0092] serial number Name of surface cover Sample size C1 Asphalt 6631 C2 Meadows 18649 C3 Gravel 2099 C4 Trees 3064 C5 Metal Sheets 1345 C6 Bair Soil 5029 C7 Bitumen 1330 C8 Briks 3682 C9 Shadows 947 total 42776
[0093] To effectively evaluate the performance of the proposed method, the simulation experiment of this invention uses the overall accuracy OA, average accuracy AA, and Kappa coefficient as quantitative evaluation indicators for the classification performance of the hyperspectral image classification task, and plots the classification results of land cover as qualitative evaluation indicators.
[0094] 2. Simulation content and result analysis:
[0095] Under the above simulation conditions, the existing SpectralFormer hyperspectral image classification method using Transformer to extract spectral features and the method of this invention were used to classify the land cover of interest in the real hyperspectral images used in the simulation experiment. The classification results are shown in Table 2. The table provides the classification accuracy, overall classification accuracy OA, average accuracy AA, and kappa coefficient for different land cover types. Figure 7 and Figure 8 The two methods are presented in a visual representation of the classification results of the hyperspectral image Pavia University.
[0096] Table 2. Comparison of Ground Cover Classification Accuracy of Existing Techniques and the Method of the Present Invention on Real Hyperspectral Images
[0097] method SpectralFormer The method proposed in this invention C1 90.7% 92.26% C2 98.07% 96.77% C3 71.18% 77.04% C4 94.61% 99.55% C5 100% 100% C6 73.63% 84.17% C7 80.3% 95.71% C8 82.65% 98.32% C9 99.89% 99.27% OA 90.71% 93.33% AA 87.89% 93.68% Kappa 0.8755 0.9124
[0098] As can be seen from the quantitative results in Table 2, the method of the present invention has better accuracy in classifying ground features in real hyperspectral images than existing Transformer-based methods. This is because the present invention improves the efficiency of spectral feature extraction and classification accuracy by exploring the global correlation of spectral features in different domains when mining the correlation of spectral bands.
[0099] like Figure 7 and Figure 8 As shown, the visualization classification results of the method of the present invention are more refined and have less intra-class noise compared with existing methods, indicating that the method of the present invention has a stronger ability to extract discriminative features, thus having a more accurate ability to identify ground features and achieving higher classification accuracy.
[0100] In summary, the method of this invention solves the problem of insufficient spectral feature extraction in existing Transformer-based spectral feature extraction algorithms by extracting multi-scale spectral features from a global perspective. The spectral feature extraction method proposed in this invention can more fully model and reconstruct the inter-band correlation features in hyperspectral images, thereby mining the class representative features contained in each pixel and achieving high-precision land cover classification.
Claims
1. A hyperspectral image classification method based on scale-interactive Transformer, characterized in that, Includes the following steps: (1) An image of size 1 was selected from the hyperspectral image library. hyperspectral images ,in and These represent the width and height of the image, respectively. Represents the number of bands in the image. and All are positive integers greater than 0. It is a positive integer greater than or equal to 100; (2) From hyperspectral images Extract each pixel and its corresponding spectral information. Regarding the above Perform planarization; (3) Use the Directional Separating Convolution Module (DSCB) to extract local scale spectral features. Nonlocal scale spectral features ; (4) Features and Serialize them separately to meet the input requirements of the scale interaction Transformer module; (5) Serialize the features and The input is fed into the scale interaction Transformer module to obtain global feature correlation modeling under multi-scale perspective. The scale interaction Transformer module includes: a scale interaction multi-head attention module SiMHA formed by stacking several scale interaction attention layers SiA. A layer normalization layer (LN) and a multilayer perceptron (MLP) are connected in series, and a residual connection is established between the input of the LN layer and the output of the MLP. (6) The output of the scale interaction Transformer module is passed through a fully connected network to obtain the classification result of each pixel.
2. The hyperspectral image classification method based on scale-interactive Transformer according to claim 1, characterized in that, In step (2), the pixel points and their corresponding spectral information are processed. The planarization process is implemented as follows: (2a) in The end of the spectral dimension uses Fill with zero elements to obtain to satisfy ,in Satisfy the following formula: , in, This represents the floor function; (2b) The above Remodeling .
3. The hyperspectral image classification method based on scale-interactive Transformer according to claim 2, characterized in that, In step (3), local scale spectral features are extracted using the directional separation convolution module. Nonlocal scale spectral features Its implementation is as follows: the directional separation convolution module contains two branches, and the local scale spectral feature extraction branch uses a convolution kernel with a size of [missing information]. The two-dimensional convolutional layer from the Extracting local-scale spectral features The nonlocal scale spectral feature extraction branch uses a convolution kernel with a size of [missing value]. The two-dimensional convolutional layer from the Extracting nonlocal scale spectral features The two-dimensional convolutional layer consists of two-dimensional convolution operations, batch normalization operations, and the Mish activation function.
4. The hyperspectral image classification method based on scale-interactive Transformer according to claim 1, characterized in that, In step (4), the features and Serialization is performed separately, as follows: (4a) will and Convert to one-dimensional sequence and ,in ; (4b) Using a linear layer and Mapping the feature dimension to D dimensions yields... and : , , in, These are learnable parameters. ; (4c) will and Each with learnable class tags and Connect them, and then add position coding. ,in , This yields the sequence used as input to subsequent Transformers: , 。 5. The hyperspectral image classification method based on scale-interactive Transformer according to claim 4, characterized in that, Step (5) specifically includes the following steps: (5a) will The sequence is linearly mapped to , , Three matrices, The sequence is linearly mapped to , , Three matrices; (5b) Construct a scale-interactive multi-head attention module SiMHA, which is composed of multiple scale-interactive attention layers SiA stacked together. Each SiA takes two feature sequences reflecting information at different scales as inputs and calculates the single-scale attention weights of the two sequences. , The results of the multi-scale attention weighted MAW are weighted stepwise, and two weighted feature sequences are output. The calculation formula of the scale-interactive multi-head attention module SiMHA is as follows: , , , , , , in, It is an activation function used to obtain weights. Used to connect different heads. and They are and Feature dimensions, It is the number of stacked SiA atoms. It is a learnable parameter used to map the dimension of the merged SiA output to the dimension of the original input feature vector; (5c) The weighted features are input into a Layer Normalization (LN) layer and a Multilayer Perceptron (MLP) layer. The LN layer is for stabilizing the feature distribution. The MLP layer consists of an input layer, a GELU activation function, and an output layer. Both the input and output layers are composed of fully connected networks. Skip connections are established between the input of the LN layer and the output of the MLP to alleviate gradient vanishing. Finally, the first dimension of the output of the residual connection is included. and One-dimensional feature vectors are used as the output of the scale-interaction Transformer.
6. The hyperspectral image classification method based on scale-interactive Transformer according to claim 1, characterized in that, Step (6) specifically includes the following steps: (6a) will be obtained and Perform splicing along the direction of the second-dimensional feature; (6b) Input the concatenated features into an MLP layer to obtain the final classification result.
Citation Information
Patent Citations
Hyperspectral image recognition method based on deep sequence convolutional network
CN115830461A
Hyperspectral image classification method based on hybrid transformation
CN117788875A