Hyperspectral image classification method and device combining EMP features and TNT module

By combining EMP features and TNT module methods, the problem of insufficient remote dependency modeling and global context information acquisition in hyperspectral images is solved, and higher classification accuracy and robustness are achieved.

CN113850315BActive Publication Date: 2025-05-23Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111107476.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-05-23
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

Existing deep learning models based on convolutional neural networks (CNN) are not good at modeling remote dependencies and obtaining global context information in hyperspectral images, resulting in insufficient classification accuracy.

Method used

Combining the extended morphological profile (EMP) features and Transformer-iN-Transformer (TNT) module, EMP cubes of hyperspectral images are extracted and end-to-end classification is performed through the TNT module to make full use of spatial and spectral information.

Benefits of technology

The classification accuracy of hyperspectral images is improved, and the classification performance is higher than that of traditional CNN models, which can better handle the characteristics of high-dimensional, nonlinear and spatial-spectral information fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113850315B_ABST
    Figure CN113850315B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of hyperspectral image classification, and particularly relates to a hyperspectral image classification method and device combining EMP features and TNT modules. The method first uses an extended morphological profile to extract the EMP features of the entire hyperspectral image, and divides the generated EMP cube into a number of patches in turn; then each patch is expanded and linearly transformed to obtain a patch embedding and a number of pixel embeddings; finally, the patch embedding and the pixel embedding are respectively added to their corresponding position codes, and the obtained vectors are input together into a deep network model containing L TNT modules for classification. Compared with support vector machines and other CNN deep learning models, the present invention can obtain higher classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hyperspectral image classification, and in particular relates to a hyperspectral image classification method and device combining EMP features and TNT modules. Background Art

[0002] Hyperspectral image classification is one of the most important links in hyperspectral image processing and analysis, and its accurate classification results can provide strong data support for subsequent tasks. At present, hyperspectral image classification has been widely used in many fields such as precision agriculture, urban planning, and resource exploration. Hyperspectral images contain rich spectral information, and each pixel has an approximately continuous spectral curve, which makes it possible to accurately classify and identify objects. However, the high-dimensional complexity and correlation between bands of hyperspectral images have an impact on classification and identification. In order to make full use of the spectral features in hyperspectral images for classification, spectral feature extraction techniques such as principal component analysis (PCA), independent component analysis (ICA) and local linear embedding (LLE) have been widely used. At the same time, affected by factors such as the environment and equipment, hyperspectral images also have the phenomenon of "same object, different spectrum" and "same spectrum, different objects". In order to further improve the accuracy and robustness of hyperspectral image classification, hyperspectral image classification methods based on spatial feature extraction such as extended morphological profile (EMP) and local binary pattern (LBP) have received widespread attention. At the same time, spectral and spatial feature extraction techniques are often combined with machine learning classifiers such as support vector machines (SVM) for classification, which can improve classification accuracy to a certain extent. However, the traditional feature extraction plus classifier model cannot fully adapt to the high-dimensional, nonlinear, and spatial-spectral information fusion characteristics of hyperspectral images.

[0003] Compared with traditional machine learning methods, deep learning methods can automatically learn deep abstract features that are beneficial to the target task layer by layer. These features have large information content and strong robustness. At present, deep learning models such as stacked auto encoder (SAE), recurrent neural network (RNN), deep belief network (DBN) and convolutional neural network (CNN) have been widely used in hyperspectral image classification, and have achieved better classification performance than traditional classification methods with sufficient training samples. In the above network models, CNN has always been an indispensable and important module.

[0004] Despite this, CNN still has shortcomings in long-range dependency modeling and global context information acquisition. In contrast, the converter model treats the input image as a sequence of patches, which can better utilize global context information in a large range. It has achieved good results in computer vision tasks such as image segmentation and object detection. In addition, the converter model itself also contains a self-attention mechanism, which can more accurately capture features and information that are beneficial to the target task, thereby obtaining more accurate and robust classification results. Summary of the invention

[0005] In view of the problems that the deep learning model based on convolutional neural network (CNN) in the prior art is not good at modeling long-range dependencies and acquiring global context information, the present invention proposes a hyperspectral image classification method and device combining EMP features and TNT (Transformer-iN-Transformer) module. The method first extracts the EMP features of the hyperspectral image, and then directly inputs the obtained EMP cube into the constructed deep network model based on the TNT module for end-to-end classification, thereby improving the classification accuracy of the hyperspectral image.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention provides a hyperspectral image classification method combining EMP features and TNT modules, comprising the following steps:

[0008] The EMP features of the entire hyperspectral image are extracted using the extended morphological profile, and the generated EMP cube is divided into several patches in turn;

[0009] Each patch is expanded and linearly transformed to obtain a patch embedding and several pixel embeddings;

[0010] The patch embedding and pixel embedding are added to their respective corresponding position encodings, and the resulting vectors are input together into a deep network model containing L TNT modules for classification.

[0011] Furthermore, the EMP features are extracted to generate an EMP cube, including:

[0012] The hyperspectral image is processed by principal component analysis to reduce its dimensionality, and the first three principal components are retained;

[0013] By using the structural element to perform opening and closing operations on the three principal components, one principal component image generates 9 EMP feature maps including itself, and 3 principal components generate 27 EMP feature maps;

[0014] After EMP feature extraction, the hyperspectral image of size W×H×C is converted into an image of size W×H×27, where W, H, and C represent the width, height, and number of bands of the image, respectively;

[0015] All data around the center pixel are selected to generate the EMP cube.

[0016] Furthermore, cross-shaped structural elements with sizes of 3, 5, 7, and 9 were selected to perform four opening and closing operations on the input principal component image to extract spatial features at different scales.

[0017] Furthermore, the EMP feature extraction formula is as follows:

[0018]

[0019]

[0020] In the formula, I is the input image, m is the number of principal components, and n is the number of opening and closing operations. and γ R Respectively represent the opening operation and closing operation, MP (n) (I) is the morphological profile feature formed after opening and closing operations are performed on image I.

[0021] Further, the EMP cube is divided into n patches X = [[X 1 ,X 2 ,…,X n ]]∈R n×p×p×27 , where (p,p) is the size of each patch.

[0022] Furthermore, each patch is expanded and linearly transformed to obtain a patch embedding and several pixel embeddings, including:

[0023] Each patch is converted into multiple (p′, p′) pixel embeddings by expansion and linear transformation, then the patch tensor sequence is expressed as:

[0024] In the formula, each patch tensor Considered as a pixel embedding sequence, the total number of pixels after conversion is m = p′ 2 , A vector representing a pixel.

[0025] Further, the TNT module includes an external converter and an internal converter, the external converter is used to process patch-level features, and the internal converter is used to process pixel-level features;

[0026] For pixel embedding, the internal transformer extracts pixel-level features using the following formula:

[0027]

[0028]

[0029] In the formula, index layer l = 1, 2, ..., L, L is the total number of layers, Y l i represents the pixel embedding sequence of the lth layer, represents the l-1th layer pixel embedding sequence, LN represents the layer normalization of the visual converter, MLP represents the multi-layer perceptron of the visual converter, and MHA represents the multi-head attention mechanism of the visual converter;

[0030] For the patch layer, create a new patch embedding memory to store the patch-level feature sequence: Where Z class represents the class label, n represents the number of patches, and d represents The dimension of , j = 1, 2, ..., n; the patch tensor is transformed into the patch embedding domain by linear projection and added to the patch embedding, the formula is as follows:

[0031]

[0032] In the formula, Vec represents the flattening operation, W and b are weights and biases respectively. represents the patch embedding sequence of the l-1th layer;

[0033] The calculation formula for extracting patch-level features by the external converter is as follows:

[0034]

[0035]

[0036] In the formula, represents the embedding sequence of the first layer of patches, Represents the patch embedding sequence of the l-1th layer.

[0037] The present invention also provides a hyperspectral image classification device combining EMP features and TNT modules, comprising:

[0038] The EMP feature extraction module is used to extract the EMP features of the entire hyperspectral image using the extended morphological profile and divide the generated EMP cube into several patches in turn;

[0039] The patch conversion module is used to expand each patch and perform a linear transformation to obtain a patch embedding and several pixel embeddings;

[0040] The TNT deep network classification module is used to add the patch embedding and pixel embedding to their respective corresponding position codes, and input the obtained vectors together into a deep network model containing L TNT modules for classification.

[0041] Compared with the prior art, the present invention has the following advantages:

[0042] The hyperspectral image classification method combining EMP features and TNT modules of the present invention uses extended morphological profiles (EMP) to extract features from hyperspectral images, effectively utilizing spatial information and spectral information in hyperspectral images, while reducing the number of bands of hyperspectral images; the internal converter and external converter in the TNT module can respectively extract pixel-level features and patch-level features of hyperspectral images, making full use of the global and local information of the input EMP cube data, and further improving the classification performance of hyperspectral images. Compared with support vector machines and other CNN deep learning models, this method can achieve higher classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0044] Figure 1 is a workflow diagram of a hyperspectral image classification method combining EMP features and TNT modules according to an embodiment of the present invention, wherein MLP represents a multi-layer perceptron;

[0045] Figure 2 is a workflow diagram of EMP cube generation according to an embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of a network structure of a visual converter model according to an embodiment of the present invention;

[0047] Figure 4 is a schematic structural diagram of a TNT module according to an embodiment of the present invention;

[0048] Figure 5 Schematic diagram of the effect of network depth on classification accuracy in an embodiment of the present invention, where L represents the number of TNT modules;

[0049] Figure 6 is a classification diagram of different classification methods of the embodiment of the present invention on the UP data set;

[0050] Figure 7is a classification diagram of different classification methods of an embodiment of the present invention on an IP data set;

[0051] Figure 8 is a classification diagram of different classification methods of the embodiment of the present invention on the SA data set;

[0052] Fig. 9 1 is a classification accuracy curve of different methods according to the embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] like Figure 1 As shown, the hyperspectral image classification method combining EMP features and TNT modules in this embodiment takes the hyperspectral image as input and the classification result as output, and includes the following steps:

[0055] Step S1, using extended morphological profiles to extract the EMP features of the entire hyperspectral image, and dividing the generated EMP cube into a number of patches in turn; EMP feature extraction can reduce the dimension of the hyperspectral image and effectively utilize the spatial spectral information in the data.

[0056] Step S2, each patch is expanded and linearly transformed to obtain a patch embedding and several pixel embeddings.

[0057] In step S3, the patch embedding and pixel embedding are added to their respective corresponding position codes, and the obtained vectors are input together into a deep network model containing L TNT modules for classification.

[0058] Extended Morphological Profile (EMP) can simultaneously extract the spatial and spectral features of hyperspectral images and has been widely used in hyperspectral image classification. Figure 2 As shown, EMP features are extracted and an EMP cube is generated, including:

[0059] Step S11, using principal component analysis (PCA) to perform dimensionality reduction processing on the hyperspectral image, and retain the first three principal components.

[0060] Step S12, use the structure element (SE) to perform opening and closing operations on the three principal components. Preferably, cross-shaped structure elements of sizes 3, 5, 7, and 9 are selected to perform four opening and closing operations on the input principal component image to extract spatial features at different scales. Then, one principal component image generates 9 EMP feature maps including itself, and 3 principal components generate 27 EMP feature maps.

[0061] Step S13, after EMP feature extraction, the hyperspectral image of size W×H×C is converted into an image of size W×H×27, where W, H, and C represent the width, height, and number of bands of the image, respectively.

[0062] Step S14, select all data near the central pixel to generate an EMP cube. The EMP feature extraction formula is as follows:

[0063]

[0064]

[0065] In the formula, I is the input image, m is the number of principal components, and n is the number of opening and closing operations. and γ R Respectively represent the opening operation and closing operation, MP (n) (I) is the morphological profile feature formed after opening and closing operations are performed on image I.

[0066] Here we first briefly introduce the visual converter. The network structure of the visual converter model is as follows Figure 3 As shown in Figure 1, the model mainly consists of three parts: multi-head attention mechanism (MHA), multi-layer perceptron (MLP) and layer normalization (LN). MHA is the key part of feature learning in the converter model. MLP is introduced between different MHAs to achieve feature conversion and nonlinearity. LN, as a data normalization layer, can ensure the stability and fast convergence of model training. In addition, residual connections are introduced to make full use of abstract features at different levels. The self-attention mechanism is the basic unit of MHA. The input embedding of the self-attention mechanism is X∈R n ×d , first convert the input embedding into a query matrix Key Matrix Sum Matrix n and d represent the length and size of the input embedding, d k and d v Represents the dimensions of the query (key) matrix and the value matrix respectively.

[0067] The output of the self-attention mechanism is:

[0068]

[0069] From the above formula, we can see that the output of the self-attention mechanism is actually the weighted sum of the value vectors, and the weight assigned to each value vector is calculated from the query vector and the corresponding key vector.

[0070] M=MHA(X)=Concat(a 1 ,a 2 ,…,a p )W O (2)

[0071] Similar to the idea of ​​using different convolution kernels to extract different features in CNN, MHA is used to improve the representation ability of the converter model. Specifically, multiple query matrices, key matrices, and value matrices are generated simultaneously in a converter model, and multiple output vectors a are calculated; then, all output vectors are concatenated and linearly projected to obtain the final output vector of formula (2). In formula (2), a represents the output vector obtained according to formula (1), W O is a trainable parameter matrix for linear transformation. In short, MHA transforms the input vector X into a feature matrix M that contains the original input vector information and the relationship between vectors.

[0072] MLP can deepen the network to improve the nonlinear representation ability of the model. In this paper, two fully connected layers are used to construct MLP:

[0073] MLP(M)=σ(XW 1 +b 1 )W 2 +b 2 (3)

[0074] in, and Represent the weights of the two fully connected layers, and b 2 ∈R d is the bias term, σ represents the GELU activation function, d and d m Represents the size of the weight matrix.

[0075] Introducing layer normalization LN after MHA and MLP can effectively speed up the convergence of the model and improve the stability of the training process. LN acts on each sample x∈R d :

[0076]

[0077] Where μ and δ are the mean and standard deviation of the input samples, respectively, and γ and β are the affine transformation parameters.

[0078] The visual transformer model can be widely used in the field of image classification. It treats the input image as a series of image patches, but ignores the intrinsic structural information of each patch. In order to make full use of the global and local information of the EMP cube, the TNT module is introduced as the core of the deep network model. Here, the TNT module is an improvement on the visual transformer above. For the input EMP cube, first divide it into n patches X = [X 1 ,X 2 ,…,X n ]∈R n×p×p×27 , where (p, p) is the size of each patch, and then each patch is converted into multiple (p′, p′) pixel embeddings by expansion and linear transformation. Then the patch tensor sequence is expressed as:

[0079]

[0080] In the formula, each patch tensor Considered as a pixel embedding sequence, the total number of pixels after conversion is m = p′ 2 , A vector representing a pixel.

[0081] Specifically, the TNT module includes an external converter and an internal converter. The external converter is used to process patch-level features and can make full use of the global information of the EMP cube data. The internal converter is used to process pixel-level features and can make full use of the local information of the EMP cube data. Figure 4 shown.

[0082] For pixel embedding, an internal transformer is used to express the relationship between pixels. The calculation formula for extracting pixel-level features by the internal transformer is as follows:

[0083]

[0084]

[0085] In the formula, index layer l = 1, 2, ..., L, L is the total number of layers, Y l i represents the pixel embedding sequence of the lth layer, represents the l-1th layer pixel embedding sequence, LN represents the layer normalization of the visual transformer, MLP represents the multi-layer perceptron of the visual transformer, and MHA represents the multi-head attention mechanism of the visual transformer. This process establishes the relationship between pixels by calculating the interaction between two pixel embeddings, thereby effectively utilizing the local information within the EMP cube.

[0086] For the patch layer, create a new patch embedding memory to store the patch-level feature sequence: Where Z class represents the class label, n represents the number of patches, and d represents The dimension of is, j = 1, 2, …, n; in each TNT module, the patch tensor is transformed into the patch embedding domain by linear projection and added to the patch embedding, the formula is as follows:

[0087]

[0088] In the formula, Vec represents the flattening operation, W and b are the weight and bias respectively. represents the l-1th layer patch embedding sequence. Similarly, the calculation formula for extracting patch-level features by the external transformer is as follows:

[0089]

[0090]

[0091] In the formula, represents the embedding sequence of the first layer of patches, represents the l-1th layer patch embedding sequence. This process can effectively learn patch-level features by acquiring the inherent information of the patch sequence. In other words, the external transformer can make full use of the global information in the input data.

[0092] The TNT module in this example can process both pixel-level and patch-level data at the same time, which means that the deep network model built by stacking TNT modules can make full use of the global and local information in the EMP cube, learn richer and more robust features, and improve the classification accuracy of hyperspectral images.

[0093] The following three hyperspectral public datasets are used for classification experiments.

[0094] The experiment uses three hyperspectral image datasets. The computer hardware environment is Intel Core i7-9750H processor, 16G memory, NVIDIA GeoForce GTX 2070 graphics card, and the software environment is Python3.6, PyTorch and sklearn.

[0095] 1. Experimental data

[0096] In order to verify the effectiveness of the proposed method, experiments were carried out using three public hyperspectral image datasets from the University of Pavia (UP), Indian Pines (IP) and Salinas (SA) acquired with different sensors and with different spatial resolutions and spectral ranges.

[0097] The UP dataset is a hyperspectral image of the University of Pavia in Italy acquired by the ROSIS imaging spectrometer. The spectral coverage range is 430-860nm, the image size is 610×340 pixels, and the spatial resolution is 1.3m. After removing the bands that are greatly affected by noise, 103 bands remain for the experiment. The dataset contains 9 types of land features, including asphalt roads, grass, gravel, and trees. The training sample and test sample information are shown in Table 1.

[0098] Table 1 UP dataset sample information

[0099]

[0100] The IP dataset is a hyperspectral image of vegetation in northwest Indiana, USA, collected by the Airborne Visible Infrared Imaging Spectrometer (AVIRIS). The spectral imaging range is 400-2500nm, the image size is 145×145 pixels, and the spatial resolution is about 20m. The dataset contains 16 types of land features, including alfalfa, corn, grassland, and soybean. Since the number of samples of 7 types of land features, such as alfalfa, is very small, only 9 types of land features, such as no-till corn, with a sample size of more than 200 are used in the experiment. The training sample and test sample information are shown in Table 2.

[0101] Table 2 IP dataset sample information

[0102]

[0103]

[0104] The SA dataset is a hyperspectral image of the Salinas Valley in California, USA, collected by AVIRIS. The spectral imaging range is 430-860nm, the image size is 512×217 pixels, the spatial resolution is about 3.7m, and there are 204 bands in total. The dataset includes 16 types of land features such as fallow land and celery. The training sample and test sample information are shown in Table 3.

[0105] Table 3 SA dataset sample information

[0106]

[0107] 2. Hyperparameter settings

[0108] The selection of hyperparameters has a significant impact on the performance of deep learning models, and appropriate hyperparameters can effectively improve classification and recognition accuracy.

[0109] The network depth is directly related to the nonlinear representation ability of the model. Generally speaking, the more layers the network has, the stronger the model's abstract modeling ability is, and the deeper and more robust the features that can be learned are. The impact of network depth on classification accuracy is as follows: Figure 5As shown in the figure, it can be seen that for IP, SA and UP data sets, the classification accuracy of the model generally increases first and then decreases with the increase in the number of TNT modules. This shows that an appropriate network structure can obtain the best classification performance, while too many network layers may lead to overfitting, resulting in a decrease in classification accuracy.

[0110] In order to improve the feature learning ability of the model, this paper introduces a multi-head attention mechanism in the TNT module. Theoretically, appropriately increasing the number of attention heads can enable the model to learn richer and more robust features, thereby obtaining better classification results. Therefore, we analyzed the impact of the number of attention heads on the classification accuracy, and the results are shown in Table 4. From the results in the table, it can be seen that: In general, as H increases, the classification accuracy of the model first gradually increases and then slowly decreases. When H is equal to 6 (SA) or 8 (UP and IP), the classification accuracy reaches the maximum value.

[0111] Table 4 Relationship between the number of attention heads (H) and classification accuracy on UP, SA and IP datasets

[0112]

[0113] The training process mainly adopts a method combining large number of iterations and small learning rate. The relevant hyperparameter settings are directly adopted from the reference literature, such as: the number of training iterations is set to 500, the learning rate is set to 0.00001, and the batch size is set to 64. First, the input EMP cube is divided into 16 small blocks in spatial order, and each small block is further divided into small blocks with a width of 2. In the TNT module, the patch embedding size is set to 128 and the pixel embedding size is set to 64.

[0114] 3. Classification results and analysis

[0115] In order to verify the effectiveness of the proposed classification method, the experiment uses RBF-SVM classic machine learning method, CNN-PPF, CDCNN, RES-3D-CNN, DCCNN, S-CNN+SVM and other five advanced CNN classification models for comparative analysis. The overall classification accuracy (OA), average accuracy (AA) and kappa coefficient are used as evaluation indicators to quantitatively compare and analyze different classification methods. In addition, in order to reduce the fluctuation of classification results caused by the randomness of sample selection, all experimental results are the average of 10 results, which further enhances the persuasiveness of the experimental results. The classification results of different methods on UP, IP and SA datasets are shown in Tables 5-7.

[0116] Table 5 Classification results of different algorithms on UP dataset (%)

[0117]

[0118] Table 6 Classification results of different algorithms on IP dataset (%)

[0119]

[0120] Table 7 Classification results of different algorithms on SA dataset (%)

[0121]

[0122] From the statistical results in the above table, we can conclude that:

[0123] (1) The classification accuracy of support vector machine is significantly lower than that of the other six deep learning methods. As a traditional shallow classifier, SVM cannot extract the deep features contained in hyperspectral images, so it cannot obtain satisfactory classification results. In contrast, the deep learning model can extract deep abstract features with greater information content and stronger robustness, thereby achieving higher classification accuracy.

[0124] (2) By constructing a contextual deep network, CDCNN can make more full use of the spatial information in hyperspectral images, so its classification accuracy is higher than that of CNN-PPF based on one-dimensional convolution. Res-3D-CNN uses three-dimensional convolution to directly extract spatial-spectral features in hyperspectral images, DCCNN uses one-dimensional convolution and two-dimensional convolution to extract spectral and spatial features respectively and fuse them, and S-CNN based on the idea of ​​metric learning can extract highly discriminative deep features. All three methods can effectively utilize the spatial-spectral information in hyperspectral images, thus achieving better classification performance.

[0125] (3) The classification results of our method on the UP dataset are comparable to those of DCCNN. We achieve the highest classification accuracy on the SA and IP datasets, with the overall classification accuracy being 1.8% and 0.88% higher than the second place, respectively.

[0126] Compared with statistical results, the classification result diagram can more intuitively display the classification results of different methods. The classification diagrams of different classification methods for the three sets of data sets are shown in Figure 2. Figure 6-8 As shown in the figure, it can be seen that the SVM method does not use spatial feature information, so its classification map contains a lot of noise; CNN-PPF uses a voting strategy in the neighborhood of pixels to determine the category, which uses spatial information to a certain extent, thereby reducing the noise in the classification map; CDCNN, RES-3D-CNN, DCCNN and SCNN+SVM use two-dimensional or three-dimensional convolution to use the spatial and spectral information in hyperspectral images to obtain smoother classification maps; compared with other methods, the classification map obtained by the method proposed in this paper has the best visual effect and is closest to the ground truth data, which once again proves the effectiveness of the proposed method from a visual perspective.

[0127] In addition, in order to verify the impact of sample size on classification performance, we randomly selected 50, 100 and 150 labeled samples per category as training samples for experiments and analyzed the adaptability of different methods to the number of training samples. Fig. 9 As shown in the figure, as the number of training samples increases, the classification accuracy of all methods gradually improves, but the classification accuracy of the method in this paper is the highest, which shows that the method in this paper has the best adaptability to the change of the number of training samples.

[0128] In order to improve the classification accuracy of hyperspectral images, this embodiment proposes a hyperspectral image classification method that combines EMP features and TNT modules. The main advantages of this method are: first, the EMP feature extraction method effectively utilizes the spatial information and spectral information in the hyperspectral image, while reducing the number of bands of the hyperspectral image; second, the internal converter model and the external converter model in the TNT module can respectively extract the pixel-level features and patch-level features of the hyperspectral image, making full use of the global and local information of the input EMP cube data, further improving the classification performance of the hyperspectral image. Experimental results on three public hyperspectral image datasets show that the performance of this method is better than that of support vector machines and other CNN deep learning models.

[0129] Corresponding to the above-mentioned hyperspectral image classification method combining EMP features and TNT modules, this embodiment further proposes a hyperspectral image classification device combining EMP features and TNT modules, including:

[0130] The EMP feature extraction module is used to extract the EMP features of the entire hyperspectral image using the extended morphological profile and divide the generated EMP cube into several patches in turn;

[0131] The patch conversion module is used to expand each patch and perform a linear transformation to obtain a patch embedding and several pixel embeddings;

[0132] The TNT deep network classification module is used to add the patch embedding and pixel embedding to their respective corresponding position codes, and input the obtained vectors together into a deep network model containing L TNT modules for classification.

[0133] It should be noted that, in this article, the terms "comprises", "includes" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or apparatus.

[0134] Finally, it should be noted that the above is only a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A hyperspectral image classification method combining EMP features and TNT modules, It is characterized in that The following steps are involved: The EMP features of the entire hyperspectral image are extracted using the extended morphological profile, and the generated EMP cube is divided into several patches in turn; Each patch is expanded and linearly transformed to obtain a patch embedding and several pixel embeddings; The patch embedding and pixel embedding are respectively added to their respective corresponding position codes, and the obtained vectors are input together into a deep network model containing L TNT modules for classification; the TNT module includes an external converter and an internal converter, the external converter is used to process patch-level features, and the internal converter is used to process pixel-level features.

2. The hyperspectral image classification method combining EMP features and TNT modules according to claim 1, It is characterized in that Extract EMP features and generate EMP cube, including: The hyperspectral image is processed with principal component analysis to reduce its dimensionality, and the first three principal components are retained; By using the structural element to perform opening and closing operations on the three principal components, one principal component image generates 9 EMP feature maps including itself, and 3 principal components generate 27 EMP feature maps; After EMP feature extraction, the hyperspectral image of size W′H′C is converted into an image of size W′H′27, where W, H, and C represent the width, height, and number of bands of the image, respectively; All data around the center pixel are selected to generate the EMP cube.

3. The hyperspectral image classification method combining EMP features and TNT modules according to claim 2, It is characterized in that Cross-shaped structural elements of sizes 3, 5, 7, and 9 are selected to perform four opening and closing operations on the input principal component image to extract spatial features at different scales.

4. The hyperspectral image classification method combining EMP features and TNT modules according to claim 2, It is characterized in that The EMP feature extraction formula is as follows: In the formula, I is the input image, m is the number of principal components, and n is the number of opening and closing operations. and γ R Respectively represent the opening operation and closing operation, MP (n) (I) is the morphological profile feature formed after opening and closing operations are performed on image I.

5. The hyperspectral image classification method combining EMP features and TNT modules according to claim 2, It is characterized in that Divide the EMP cube into n patches X = [X 1 ,X 2 ,…,X n ]∈R n′p′p′27 , where (p,p) is the size of each patch.

6. The hyperspectral image classification method combining EMP features and TNT modules according to claim 5, It is characterized in that Each patch is expanded and linearly transformed to obtain a patch embedding and several pixel embeddings, including: Each patch is converted into multiple (p′, p′) pixel embeddings by expansion and linear transformation, then the patch tensor sequence is expressed as: In the formula, each patch tensor Considered as a pixel embedding sequence, the total number of pixels after conversion is m = p′ 2 , A vector representing a pixel.

7. The hyperspectral image classification method combining EMP features and TNT modules according to claim 6, It is characterized in that For pixel embedding, the internal transformer extracts pixel-level features using the following formula: AND l i =And l ′ i +LN(MLP(Y l ′ i )) In the formula, index layer l = 1, 2, ..., L, L is the total number of layers, Y l i represents the pixel embedding sequence of the lth layer, represents the l-1th layer pixel embedding sequence, LN represents the layer normalization of the visual converter, MLP represents the multi-layer perceptron of the visual converter, and MHA represents the multi-head attention mechanism of the visual converter; For the patch layer, create a new patch embedding memory to store the patch-level feature sequence: Where Z class represents the class label, n represents the number of patches, and d represents The dimension of , j = 1, 2, ..., n; the patch tensor is transformed into the patch embedding domain by linear projection and added to the patch embedding, the formula is as follows: In the formula, Vec represents the flattening operation, W and b are the weight and bias respectively. represents the patch embedding sequence of the l-1th layer; The calculation formula for extracting patch-level features by the external converter is as follows: In the formula, represents the embedding sequence of the first layer of patches, Represents the patch embedding sequence of the l-1th layer.

8. A hyperspectral image classification device combining EMP features and TNT modules, It is characterized in that include: The EMP feature extraction module is used to extract the EMP features of the entire hyperspectral image using the extended morphological profile and divide the generated EMP cube into several patches in turn; The patch conversion module is used to expand each patch and perform a linear transformation to obtain a patch embedding and several pixel embeddings; The TNT deep network classification module is used to add patch embedding and pixel embedding to their respective corresponding position codes, and input the obtained vectors together into a deep network model containing L TNT modules for classification; the TNT module includes an external converter and an internal converter, the external converter is used to process patch-level features, and the internal converter is used to process pixel-level features.