Hyperspectral Image Classification Method Based on Spectral Enhancement and Dense-Connected Transformer

By building a spectral enhancement module and a densely connected transformer module, the problem of insufficient utilization of local similarity and long-distance spatial information in hyperspectral image classification is solved, and the accuracy and consistency of hyperspectral image classification is improved.

CN115410085BActive Publication Date: 2025-07-04XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211033767.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-26
Publication Date
2025-07-04
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

The existing hyperspectral image classification methods fail to fully utilize the local similarity and long-distance spatial information of hyperspectral images, resulting in poor misjudgment of pixel and region consistency in the classification results.

Method used

The spectral enhancement module and densely connected transformer module are built to extract the spatial spectral information of high-spectral images through spectral enhancement, and the transmission of shallow features to the deep layer is promoted through densely connected transformer modules, improving classification performance.

Benefits of technology

It improves the accuracy and consistency of hyperspectral image classification, reduces large blocks of misjudgment of pixels, and enhances the regional consistency and robustness of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410085B_ABST
    Figure CN115410085B_ABST
Patent Text Reader

Abstract

The present invention discloses a hyperspectral image classification method based on spectral enhancement and densely connected transformers, which mainly solves the problems of poor classification performance and inconsistent classification regions of existing hyperspectral images. The implementation scheme is as follows: obtain a hyperspectral image dataset, and generate a training sample set and a test sample set; respectively construct a spectral enhancement module and a densely connected transformer module to generate a spectral enhancement and densely connected transformer model; train the spectral enhancement and densely connected transformer model; input the test set into the trained spectral enhancement and densely connected transformer model to output the classification results of the hyperspectral image. By using the constructed spectral enhancement and densely connected transformers, the present invention can extract and fuse the global and local features and long-distance spatial information of hyperspectral images, improve the accuracy and consistency of hyperspectral image classification, and can be used for land cover mapping, precision agriculture, urban planning, tree species classification, and mineral exploration of hyperspectral images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a hyperspectral image classification method, which can be used for land cover mapping, precision agriculture, urban planning, tree species classification and mineral exploration of ground objects in hyperspectral images. Background Art

[0002] The purpose of the task of hyperspectral image classification is to classify hyperspectral images pixel by pixel through various methods. In the early stage of hyperspectral image classification research, most methods focused on exploring the role of spectral features in classification. These methods mainly focused on pixel-level classification methods. However, the classification maps obtained by these pixel-level classifiers are not satisfactory because the spatial context is not considered. The spectral resolution and channel dimension of hyperspectral images are much higher than those of natural images, while the expression ability of the models of traditional algorithms is often limited and cannot maintain the same good processing performance for such high-dimensional problems without prior knowledge.

[0003] Recently, deep learning has become a growing trend in big data analysis and has made significant breakthroughs in many computer vision tasks and has also been widely applied in hyperspectral classification tasks. The hyperspectral classification task is essentially a pixel-by-pixel classification task based on hyperspectral images, and a large number of pixel classification algorithms have laid the foundation for this field. Some methods have migrated the algorithms of natural images to the task of remote sensing image classification and proposed several remote sensing image classification methods based on deep learning. However, since these methods apply convolutional neural networks to remote sensing images in the same or similar way as natural images, they regard the pixel-by-pixel classification of images as an image classification task. Therefore, it limits the input of larger local information and can only selectively use less pixel local information.

[0004] In 2021, Hong et al. proposed using transformers to solve the image classification problem in the paper "SpectralFormer: Rethinking Hyperspectral Image Classification With transformers" published on TGARS. Compared with the convolutional neural network (CNN) which fails to well mine and represent the sequential attributes of spectral features, this method regards the pixel classification problem as a classification problem, bundles the pixels before or after a pixel together as the input of a pixel, and then adds position encoding information and sends it into the encoder. Subsequently, a linear layer is used to predict this pixel block. Although this method is the pioneer work of the vision transformer model in the field of hyperspectral image classification, since the input of this network is only the pixels before and after a single pixel and does not make full use of the local similarity of hyperspectral images, there are many small misjudged pixels in the middle of large blocks in the final classification results. Moreover, since the vision transformer is originally a classification network while the hyperspectral classification network is pixel-by-pixel classification, the accuracy of the classification network is not very satisfactory.

[0005] In 2021, Zheng et al. published a paper "Spectral-Spatial transformer Network for Hyperspectral Image Classification: A Factorized Architecture Search Framework" on TGARS, proposing a new spectral-spatial transformer network, which consists of a spatial attention and a spectral correlation module to overcome the limitations of convolutional kernels. In addition, a factorized architecture search (FAS) is designed. Although this method takes both spatial and spectral information into account in the network, due to the use of attention modules weakening the extraction of spectral-spatial information and being unable to fully extract the local information of the spectrum, the classification performance is not good.

[0006] Zhejiang University of Technology disclosed a "Method for Hyperspectral Image Classification and Regression of Leaves Based on Multi-Scale Cascade Convolutional Neural Network" in the patent application document with the application number CN202210450076.9. It first embeds dilated convolutions in a 3D-CNN to construct spectral-spatial feature extraction structures of different scales to achieve multi-scale feature fusion. However, since this method uses a CNN network lacking the ability to extract long-distance spatial information, it cannot fully extract the spatial features of hyperspectral images, resulting in poor classification results. Summary of the Invention

[0007] The purpose of the present invention is to propose a hyperspectral image classification method based on spectral enhancement and dense connection transformers in view of the deficiencies of the existing methods above, so as to improve the classification accuracy and achieve the extraction of long-distance spatial information.

[0008] The technical idea for achieving the object of the present invention is as follows: By performing spectral enhancement on the input image blocks, the spatial spectral information of the hyperspectral image is fully extracted and then input into the transformer module to achieve the extraction of long-distance spatial information; through the dense connection of the transformer module, the transfer from shallow features to deep features is better promoted, improving the classification performance.

[0009] According to the above idea, the implementation steps of the present invention are as follows:

[0010] 1. A hyperspectral image classification method based on spectral enhancement and dense connection transformer, characterized by comprising the following steps:

[0011] (1) Construct a training sample set and a test sample set:

[0012] 1a) Download hyperspectral images with labels from a public website, generate a sample set according to the labeled pixels. For each pixel in the hyperspectral image, a spatial window of size 11×11 is delimited around it as a pixel block. Each data block is a data cube, and all the data cubes form the sample set of the hyperspectral image;

[0013] 1b) In the sample set of the hyperspectral image, randomly take 5% of each class as training samples to form the training sample set of the hyperspectral image, and the remaining 95% of the samples form the test sample set of the hyperspectral image;

[0014] (2) Build a spectral enhancement module composed of a max pooling layer, an average pooling layer, and a sigmoid activation function;

[0015] (3) Construct a dense connection transformer module:

[0016] 3a) Build an encoder module composed of a first normalization layer, a multi-head self-attention layer, a second normalization layer, and a multi-layer perceptron layer;

[0017] 3b) Stack 5 encoder modules and add dense connections between each encoder, that is, the outputs of all encoder modules before the current encoder module are successively added to the input of the current encoder module to form a dense connection transformer module;

[0018] (4) Cascade the spectral enhancement module, the dense connection transformer module, and the fully connected layer in sequence to generate a spectral enhancement and dense connection transformer model, and use the cross-entropy function as the loss function L of the model:

[0019] (5) Use the training samples and adopt the gradient descent method to train the spectral enhancement and dense connection transformer model to obtain a trained spectral enhancement and dense connection transformer model;

[0020] (6) Input the test sample set of the hyperspectral image into the trained spectral enhancement and densely connected transformer model one by one, and use the output of the fully connected layer as the predicted label of the test sample to obtain the classification result.

[0021] Compared with the prior art, the present invention has the following advantages:

[0022] First, since the present invention constructs an encoder module, it can fully extract the long-distance spatial information of the spectrum, overcoming the problem that traditional convolutions cannot extract global spatial information; at the same time, since the present invention adds dense connections between multiple encoders, the shallow features of the model can be better transmitted to the deep layer, overcoming the problem in the prior art that shallow information cannot be well transmitted to the deep layer, and improving the accuracy of hyperspectral image classification.

[0023] Second, since the present invention constructs a spectral enhancement module, it can make full use of the spectral-spatial features of the hyperspectral image, overcoming the problem in the prior art that the regional consistency features of the hyperspectral image are not fully utilized, so that there are few misjudged pixels in large blocks in the image classification result, alleviating the problem of poor regional consistency of the classification result, and improving the robustness of hyperspectral image classification. Description of the Drawings

[0024] Figure 1 is the implementation flowchart of the present invention;

[0025] Figure 2 is the simulation result diagram of the present invention. Detailed Embodiments

[0026] The following further describes the embodiments and effects of the present invention in detail with reference to the drawings.

[0027] Refer to Figure 1 , the implementation steps of this example are as follows:

[0028] Step 1. Obtain a hyperspectral image dataset.

[0029] Download the hyperspectral dataset marked with ground object categories from a public website.

[0030] In this example, the Pavia University dataset of the University of Pavia is downloaded from a public website. This dataset is a hyperspectral image of 610×340 pixels collected by the Reflective Optics System Imaging Spectrometer (ROSIS) sensor of the University of Pavia. The ROSIS sensor can collect information on 103 spectral bands with a spatial resolution of 1.3 m in the wavelength range of 0.43 to 0.86 μm. This dataset contains a total of 9 categories and 42,776 labeled pixels.

[0031] Step 2. Generate the training sample set and the test sample set.

[0032] 2.1) With each pixel in the hyperspectral image dataset as the center, delimit a spatial window of 11×11 pixels; form a data cube with all the pixels within each spatial window; form the sample set of the hyperspectral image with all the data cubes.

[0033] 2.2) Randomly select 5% of the samples in the sample set of the hyperspectral image to form the training sample set of the hyperspectral image, and form the test sample set of the hyperspectral image with the remaining samples.

[0034] Step 3. Construct a spectral enhancement module including a max pooling layer, an average pooling layer, and an activation function layer.

[0035] 3.1) Set the parameters of each layer:

[0036] Set both the input and output sizes of the activation function layer to 1×1.

[0037] Set the pooling kernel size of the max pooling layer to 11×11, the convolution stride to 1, the input size to 11×11, and the output size to 1×1.

[0038] Set the pooling kernel size of the average pooling layer to 11×11, the stride to 1, the input size to 11×11, and the output size to 1×1.

[0039] 3.2) Connect the max pooling layer and the average pooling layer in parallel, connect the output vector to the input of the activation function layer, and perform a dot product between the output of the activation function layer and the original data to obtain the final output of the spectral enhancement module.

[0040] Step 4. Linearize the output of the spectral enhancement module.

[0041] Perform a one-dimensional process on the output of the spectral enhancement module, that is, first stretch the 11×11 vector into a vector of size 121, and then add digital position encodings from 0 to n to the corresponding vector.

[0042] Step 5. Construct a dense connection transformer module.

[0043] 5.1) Build an encoder module including a first normalization layer, a multi-head self-attention layer, a second normalization layer, and a multi-layer perceptron layer. The structural relationship is:

[0044] The first normalization layer, the multi-head self-attention layer, the second normalization layer, and the multi-layer perceptron layer are connected in sequence. And one branch of the input of the first normalization layer makes a residual with the output of the multi-head attention layer, and one branch of the input of the second normalization layer makes a residual with the output of the multi-layer perceptron.

[0045] The parameters of each layer are as follows:

[0046] The input and output of the normalization layer are both vectors of size 200×121;

[0047] The input and output of the multi-layer perceptron layer are both vectors of size 200×121;

[0048] The number of heads in the multi-head attention layer is 10, and the generation formula for multiple heads is as follows:

[0049] MultiHead(Q,K,V)=Concat(head1,...head i ..,head h )W O

[0050] where head i =Attention(QW i Q ,KW i K ,VW i V ) represents the i-th head generated by the multi-head attention layer, and the calculation formula for the attention mechanism of the i-th head is: In the formula, Q = XW Q is the query matrix of the input X of the multi-head attention layer, K = XW K is the key matrix of the input X of the multi-head attention layer, V = XW V is the value matrix of the input X of the multi-head attention layer; W Q is the weight matrix corresponding to the query matrix, W K is the weight matrix corresponding to the key matrix, W V is the weight matrix corresponding to the value matrix, W O is the weight matrix for concatenating multiple heads in the multi-head attention layer; MultiHead is the concatenation of multiple Heads, QW i Q ,KW i K ,VW i V are the query matrix, key matrix, and value matrix of the i-th head generated by the multi-head attention layer, W i Q ,W i K ,W i V are the weight matrices of the query matrix, key matrix, and value matrix corresponding to the i-th head respectively, d k is a vector with the same dimension size as the Q matrix, used for normalization operation, and softmax is the activation function.

[0051] 5.2) Stack the 5-layer encoder modules and add dense connections between each encoder, that is, sum the outputs of all encoder modules before each encoder module and add them to the input of the current encoder module in turn to form a dense connection transformer module. The specific relationship is as follows:

[0052] The input of the second-layer encoder is the output of the first-layer encoder,

[0053] The input of the third-layer encoder is the sum of the outputs of the first and second-layer encoders,

[0054] The input of the fourth-layer encoder is the sum of the outputs of the first, second, and third-layer encoders,

[0055] The input of the fifth-layer encoder is the sum of the outputs of the first, second, third, and fourth-layer encoders;

[0056] The input size of each layer of encoder is a vector of size 200×121.

[0057] Step 6. Generate a spectral enhancement and dense connection transformer model.

[0058] Cascade the spectral enhancement module, the dense connection transformer module, and the fully connected layer in sequence to generate a spectral enhancement and dense connection transformer model, and use the cross-entropy function as the loss function L of this model.

[0059] The cross-entropy formula is as follows:

[0060]

[0061] Among them, L represents the cross-entropy between the predicted label vector and the true label vector, y i represents the i-th element in the predicted label vector, represents the m-th element in the predicted label vector.

[0062] Step 7. Train the spectral enhancement and dense connection transformer model.

[0063] 7.1) Initialize the model: Set epochs to 200, the number of heads in the multi-head attention layer to 10, the initial learning rate to 3e-4, the number of samples input to the model to 200, the size of the patch to 11×11, and the optimizer to Adam;

[0064] 7.2) Input 200 hyperspectral image patches of size 11×11×103 pixels in the training samples into the spectral enhancement and dense connection transformer model, and they respectively go through the following processes:

[0065] First, it passes through the max - pooling layer and the average - pooling layer, then they are added together, passed through the sigmoid activation function, and multiplied point - by - point to the input image patches. Image patches with a size of 11×11×103 and a quantity of 200 are output. These pixel patches are convolved into 200 one - dimensional image patches of 11×11;

[0066] Next, the 200 one - dimensional image patches of 11×11 are linearly processed to obtain 200 vectors of size 121. Then, a digital position encoding from 0 to n is added to each corresponding vector to obtain 200 vectors of size 121 with position encoding;

[0067] Subsequently, these 200 vectors of size 121 with position encoding are processed through 5 encoder layers to obtain 200 feature vectors of size 121. The feature vectors pass through a fully - connected layer to obtain the predicted labels. The input size of this fully - connected layer is a vector of 24200, and the output is a vector of size 9;

[0068] 7.3) Calculate the cross - entropy between the predicted label vector and the true label vector using the cross - entropy formula. Adopt the gradient - descent method to optimize the model parameters using the cross - entropy until the model parameters converge, and obtain the trained spectral enhancement and densely - connected transformer model.

[0069] Step 8. Classify the hyperspectral image.

[0070] Input the test sample set of the hyperspectral image into the trained spectral enhancement and densely - connected transformer model one by one. The output of its fully - connected layer is the predicted label of the test sample, and the classification result is obtained.

[0071] The following further illustrates the effect of the present invention in combination with simulation experiments:

[0072] 1. Simulation experiment conditions:

[0073] The hardware platform for the simulation experiment: The processor is an Intel i7 - 7820X CPU with a main frequency of 3.6 GHz and a memory of 64 GB. The used graphics card is a GeForce RTX2080 Ti.

[0074] The software platform for the simulation experiment is: Windows 10 operating system and python 3.7.

[0075] The input image used in the simulation experiment is the Pavia University dataset. The image size is 610×340×103 pixels. The image contains 103 bands and 9 types of ground objects, and the image format is mat.

[0076] 2. Simulation content

[0077] The hyperspectral images of the Pavia University dataset input are classified respectively using the present invention and three existing image classification methods, 3D-CNN, SSRN, and SSTN, to obtain classification results Figure 2 , where:

[0078] Figure 2 (a) is the classification result diagram of the true labels in the Pavia University dataset

[0079] Figure 2 (b) is the classification result diagram of the Pavia University dataset using the existing 3D-CNN method

[0080] Figure 2 (c) is the classification result diagram of the Pavia University dataset using the existing SSRN method

[0081] Figure 2 (d) is the classification result diagram of the Pavia University dataset using the existing SSTN method

[0082] Figure 2 (e) is the classification result diagram of the Pavia University dataset using the existing present invention method

[0083] The 3D-CNN method refers to the image classification algorithm proposed by Ji S et al. in "3D Convolutional Neural Networks for Human Action Recognition," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 1, pp. 221-231, Jan. 2013, doi: 10.1109 / TPAMI.2012.59., abbreviated as the 3D-CNN algorithm

[0084] The SSTN method refers to the image classification algorithm proposed by Zhong Z et al. in "Spectral–Spatial Transformer Network for Hyperspectral Image Classification: A Factorized Architecture Search Framework," in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-15, 2022, Art no. 5514715, doi: 10.1109 / TGRS.2021.3115699., abbreviated as the SSTN algorithm.

[0085] The SSRN method refers to the image classification algorithm proposed by Zhong Z et al. in "Spectral–Spatial Residual Network for Hyperspectral Image Classification: A 3-D Deep Learning Framework," in IEEE Transactions on Geoscience and Remote Sensing, vol. 56, no. 2, pp. 847-858, Feb. 2018, doi: 10.1109 / TGRS.2017.2755542., abbreviated as the SSRN algorithm.

[0086] From Figure 2 it can be seen that:

[0087] In Figure 2 (b), the 3D-CNN method misclassifies asphalt 2 as asphalt 1, and there are many misjudged pixels in some classification blocks;

[0088] In Figure 2 (c), the SSRN method has misjudged pixels in the classification of gravel and asphalt, and misjudges some small blocks on the meadow as other pixels, such as trees and bare soil;

[0089] In Figure 2 (d), the SSTN method has misjudged pixels in the classification of gravel and asphalt, and there are misjudged pixels in the judgment of asphalt and self-locking bricks;

[0090] In Figure 2 (e), compared with several other classification algorithms, the present invention has fewer misjudged pixels, and the classification effect is significantly better than several other methods, and it is closer to the classification result of the true label shown in Figure 2 (a).

[0091] 3. Result Analysis:

[0092] The classification results of the above four methods are evaluated respectively using three evaluation indicators: overall accuracy OA, average accuracy AA, and Kappa coefficient. The results are shown in Table 1:

[0093] Table 1 Quantitative analysis of the classification results of the present invention and each existing technology in the simulation experiment (%)

[0094]

[0095] In Table 1, OA is the overall classification accuracy index, AA is the average classification index, and Kappa is the index for consistency test. Their calculation formulas are as follows:

[0096] Overall classification accuracy

[0097] Average classification accuracy

[0098] Consistency index

[0099] In the formula p o is the sum of the number of correctly classified samples in each class divided by the total number of samples, that is, the overall classification accuracy; p e is the sum of the "product of actual and predicted quantities" corresponding to all classes divided by the "square of the total number of samples"; C is the total number of classes; T i is the number of samples correctly classified in each class. The number of true samples in each class is a1, a2,..., aC, the number of samples predicted for each class is b1, b2,..., bC, and the total number of samples is n.

[0100] As can be seen from Table 1, the overall classification accuracy OA of the present invention is 99.21%, the average classification accuracy AA is 99.35%, and the consistency index Kappa is 98.95. These three indicators are all higher than those of the 3 existing technology methods, proving that the present invention can obtain higher hyperspectral image classification accuracy.

[0101] The above simulation experiments show that: the method of the present invention can make good use of the spatial consistency principle of hyperspectral images by using the built spectral enhancement module, enhance the spatial spectral information of hyperspectral images, and can make good use of the long-distance spatial information of hyperspectral images by using the built dense connection transformer module, and can better transfer the shallow features to the deep layer, solving the problems existing in the existing technology methods, such as insufficient utilization of spatial spectral information, inability to transfer shallow features to the deep layer well, resulting in poor consistency of classification regions and low accuracy. It is a very practical hyperspectral image classification method.

Claims

1. A hyperspectral image classification method based on spectral enhancement and densely connected transformers, characterized in that, It includes the following steps: (1) Construct a training sample set and a test sample set: 1a) Download hyperspectral images with labels from a public website, generate a sample set based on labeled pixels. For each pixel in the hyperspectral image, a spatial window of size 11×11 is defined around it as a pixel block. Each data block is a data cube, and all data cubes form the sample set of the hyperspectral image; 1b) In the sample set of the hyperspectral image, randomly select 5% of each class as training samples to form the training sample set of the hyperspectral image, and the remaining 95% of the samples form the test sample set of the hyperspectral image; (2) Build a spectral enhancement module composed of a max pooling layer, an average pooling layer, and a sigmoid activation function; (3) Construct a densely connected transformer module: 3a) Build an encoder module composed of a first normalization layer, a multi-head self-attention layer, a second normalization layer, and a multi-layer perceptron layer; 3b) Stack 5 encoder modules and add dense connections between each encoder, that is, the outputs of all encoder modules before each encoder module are sequentially added to the input of the current encoder module to form a densely connected transformer module; (4) Cascade the spectral enhancement module, the densely connected transformer module, and the fully connected layer in sequence to generate a spectral enhancement and densely connected transformer model, and use the cross-entropy function as the loss function L of this model; (5) Use the training samples and adopt the gradient descent method to train the spectral enhancement and densely connected transformer model to obtain a trained spectral enhancement and densely connected transformer model; (6) Input the test sample set of the hyperspectral image into the trained spectral enhancement and densely connected transformer model one by one, and use the output of the fully connected layer as the predicted label of the test sample to obtain the classification result.

2. The method according to claim 1, wherein The structure and parameters of the spectral enhancement module built in step (2) are as follows: The max pooling layer and the average pooling layer are connected in parallel, the vector output by them is connected to the input of the activation function layer, and the output of the activation function layer is dot-multiplied with the original data to obtain the final output of the spectral enhancement module; The input and output sizes of the activation function layer are both 1×1; The pooling kernel size of the max pooling layer is 11×11, the convolution stride is 1, the input size is 11×11, and the output size is 1×1; The pooling kernel size of the average pooling layer is 11×11, the stride is 1, the input size is 11×11, and the output size is 1×1.

3. The method according to claim 1, wherein The structure and parameters of each layer of the encoder built in (3a) are as follows: The first normalization layer, the multi-head self-attention layer, the second normalization layer, and the multi-layer perceptron layer are connected in sequence. One branch of the input of its first normalization layer makes a residual with the output of the multi-head attention layer, and one branch of the input of the second normalization layer makes a residual with the output of the multi-layer perceptron; The input and output of the layer normalization layer are both vectors of size 200×121; The number of heads of the multi-head attention layer is 10; The input and output layers of the multi-layer perceptron layer are both vectors of size 200×121.

4. The method according to claim 1, wherein (3b) The structural relationship and parameters of building the densely connected transformer module are as follows: The input of the second-layer encoder is the output of the first-layer encoder. The input of the third-layer encoder is the sum of the outputs of the first and second-layer encoders. The input of the fourth-layer encoder is the sum of the outputs of the first, second, and third-layer encoders. The input of the fifth-layer encoder is the sum of the outputs of the first, second, third, and fourth-layer encoders. The input size of each encoder is a vector of size 200×121.

5. The method according to claim 1, characterized in that, The parameters of the spectral enhancement module, dense connection transformer module, and fully connected layer in step (4) are as follows: For the spectral enhancement module, both its input and output are data blocks of size 200 with a size of 11×11. For the dense connection transformer module, both its input and output are vectors of size 200 with a size of 121. For the fully connected layer, its input and output are 24200 and 9 respectively.

6. The method according to claim 1, characterized in that The cross-entropy formula in step (4) is expressed as follows: Among them, L represents the cross-entropy between the predicted label vector and the true label vector, y i represents the i-th element in the predicted label vector, and represents the m-th element in the predicted label vector.

7. The method according to claim 1, characterized in that (5) Using the training samples, the spectral enhancement and dense connection transformer model is trained by the gradient descent method as follows: 5a) Initialize the model: Set the number of epochs to 200, the number of heads in the multi-head attention layer to 10, the initial learning rate to 3e-4, the number of samples input to the network to 200, the size of the taken blocks to 11×11, and the optimizer to Adam. 5b) Input the training samples into the model, and pass through the spectral enhancement module, dense connection transformer module, and fully connected layer in sequence to output the predicted label vector of the training samples. 5c) Use the cross-entropy formula to calculate the cross-entropy between the predicted label vector and the true label vector, and use the gradient descent method to optimize the model parameters with the cross-entropy until the model parameters converge, obtaining the trained spectral enhancement and dense connection transformer model.

Citation Information

Patent Citations

  • A classification and regression method for leaf hyperspectral images based on multi-scale cascade convolutional neural networks

    CN114821321B

  • Hyperspectral image ground object classification method based on spectral segmentation and homogeneous region detection

    CN112308152A

  • Hyperspectral image classification method based on depth spectral space inverse residual network

    CN113935433A