Hyperspectral image classification method based on LCTNet novel lightweight model

Through the new LCTNet lightweight model, the spectral extraction-dimensionality reduction, space-spectral double labeling and TKformer modules are used to solve the problems of feature extraction difficulties and high model complexity in hyperspectral image classification, and efficient and accurate hyperspectral image classification is achieved.

CN120339699APending Publication Date: 2025-07-18GUIZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510419772.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional methods are difficult to effectively extract the features of hyperspectral images. Convolutional neural networks (CNNs) have receptive field limitations and insufficient spectral features utilization, while Transformer has high computational complexity, resulting in low classification accuracy of hyperspectral images and high model complexity.

Method used

The new LCTNet lightweight model is adopted, and through the spectral extraction-dimensionality reduction module, the space-spectral double marking module and the TKformer module, combined with three-dimensional convolution, receptive field attention convolution, channel attention and sparse attention mechanism, the spectral dimension is reduced and feature extraction and classification capabilities are enhanced.

Benefits of technology

While improving classification accuracy, it effectively controls the complexity of the model and realizes efficient hyperspectral image classification, which significantly improves the lightweightness and classification efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339699A_ABST
    Figure CN120339699A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral image classification method based on an LCTNet novel lightweight model, and the method comprises the following steps: S1, processing an original hyperspectral image, and enabling the original hyperspectral image to accord with a classification task and a tensor shape inputted by the model; s2, key spectral features are extracted from the original hyperspectral data through a spectrum extraction-dimensionality reduction module, and the dimensionality of the data is reduced; s3, carrying out dual feature labeling on space and spectrum information while effectively reducing the spectrum dimension through a space-spectrum dual labeling module; s4, the TKSA is introduced into a Transform module, and a TKform module is constructed; and S5, carrying out global average pooling on the feature map obtained in the previous step, compressing spatial dimensions, extracting key channel features, flattening pooled feature vectors into one-dimensional vectors, inputting the one-dimensional vectors into a full connection layer, carrying out classification, and outputting ground covering types. According to the method, the problems of over-high spectral dimension, complex features in the hyperspectral image, limitation of CNN and Transform and the like in a hyperspectral image classification task can be relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a hyperspectral image classification method based on a new lightweight model of LCTNet, belonging to the technical field of deep learning image classification. Background Art

[0002] Hyperspectral images have been widely used in multiple fields due to their ability to provide rich spectral and spatial information. By recording spatial information and hundreds of continuous spectral bands, hyperspectral images can accurately capture the subtle features and complex structures of ground objects, which is of great significance for distinguishing ground object types and monitoring their changes. Hyperspectral images have demonstrated important value in fields such as environmental monitoring, geological exploration, urban planning, precision agriculture, and military reconnaissance, providing key support for detailed ground object analysis and accurate classification.

[0003] Traditional methods are difficult to effectively extract features. Convolutional neural networks (CNNs) have problems of limited receptive fields and insufficient utilization of spectral features when dealing with hyperspectral data, while Transformers, although having advantages, have high computational complexity. Therefore, a new classification method is needed to overcome these problems, improve classification accuracy, and control model complexity. In view of this, the present invention proposes a hyperspectral image classification method based on a new lightweight model of LCTNet. Summary of the Invention

[0004] The purpose of the present invention is to provide a hyperspectral image classification method based on a new lightweight model of LCTNet. This method can alleviate problems such as too high spectral dimensions, complexity of features in hyperspectral images, and limitations of CNNs and Transformers in hyperspectral image classification tasks.

[0005] The technical solution of the present invention: A hyperspectral image classification method based on a new lightweight model of LCTNet, comprising the following steps:

[0006] S1: Process the original hyperspectral image to make it conform to the tensor shape of the classification task and model input;

[0007] S2: Extract key spectral features from the original hyperspectral data through a spectral extraction - dimensionality reduction module and reduce the dimensionality of the data;

[0008] S3: While effectively reducing the spectral dimension through a spatial - spectral double - labeling module, perform double - feature labeling on spatial and spectral information;

[0009] S4: Introduce TKSA into the Transformer module to construct a TKformer module;

[0010] S5: Perform global average pooling on the feature map obtained in the previous step to compress the spatial dimension and extract key channel features. Subsequently, flatten the pooled feature vector into a one-dimensional vector, input it into the fully connected layer, and perform classification to output the type of ground cover.

[0011] In the aforementioned hyperspectral image classification method based on the novel lightweight LCTNet model, in S1, processing the original hyperspectral image specifically involves: dividing the dataset into several small cubes centered on pixel points that contain both spectral features and spatial features.

[0012] In the aforementioned hyperspectral image classification method based on the novel lightweight LCTNet model, in S2, the spectral extraction - dimensionality reduction module consists of three - dimensional convolutions with two different kernels and a one - dimensional convolution. The specific process is as follows:

[0013] X Conv = Conv(Reshape(Conv3D2(Conv3D1(X HSI ))))

[0014] Among them, X HSI is the processed hyperspectral image, X Conv is the output of the spectral extraction - dimensionality reduction module. Conv3D1(·) is a convolution with a kernel of (7×1×1), 8 output channels, and a stride of 3; Conv3D2(·) is a convolution with a kernel of (3×1×1), 8 output channels, and a stride of 2; Reshape(·) represents splicing the feature map along the spectral dimension.

[0015] In the aforementioned hyperspectral image classification method based on the novel lightweight LCTNet model, in step S3, the spatial - spectral feature marking module consists of a receptive field attention convolution RFAConv, an efficient channel attention ECA, and a 1×1 convolution. X Conv is respectively sent into RFAConv and ECA, and the obtained results are concatenated together along the channels. Finally, a convolution with a kernel of 1×1 is used for dimensionality reduction and channel interaction. The specific process can be expressed as:

[0016] X s = Conv(Cat(RFA(X Com ), ECA(X Com )))

[0017] = Conv(Cat(X RFA , X ECA ))

[0018] Among them, X Sis the output of the spatial - feature marking module, Conv(·) is a 2D convolution with a kernel of 1×1 and 64 output channels, RFA(·) is the receptive - field attention convolution, and ECA(·) is the efficient channel attention operation.

[0019] In the foregoing hyperspectral image classification method based on the novel lightweight LCTNet model, the implementation processes of the RFAConv and ECA are as follows:

[0020] 1) RFAConv implementation: RFAConv generates a receptive - field feature map by combining spatial features and attention weights, and its implementation steps are as follows:

[0021] Step 1: Use grouped convolution to obtain the spatial features of the original feature map. The output channels of the grouped convolution are set to the square of the convolution kernel. The specific process can be expressed as:

[0022] F rf =σ(g C (X Conv , C, k 2 C, k×k))

[0023] where F rf is the spatial feature, σ(·) is the activation function, g C (·) is the grouped convolution, C is the number of channels of the original feature map, k 2 C is the number of channels of the feature map after convolution, k is the convolution - kernel size, and gC(X Conv , C, k 2 C, k×k) represents performing convolution on X 2 using a grouped convolution with a convolution kernel of k×k, a group of C, an input channel of C, and an output channel of k Conv C;

[0024] Step 2: While obtaining the spatial features, obtain the attention weights. After aggregating the global information of each receptive - field feature through average pooling, then use 1×1 grouped - convolution operations to interact the information to obtain the attention weights. Finally, use Softmax to emphasize the importance of each feature in the receptive - field feature. The specific process can be expressed as:

[0025] A rf =Softmax(g(AvgPool(X Conv ), C, k 2 C, 1×1))

[0026] where, is the attention weight;

[0027] Step 3: After obtaining the spatial feature F rf and the attention weight A rfAfter that, perform dot multiplication on them, and through feature recombination, recombine the extra channel features into the space, so that the new feature map contains the features of all receptive fields and there is no duplication. The specific process can be expressed as:

[0028] F = Reshape(F rf × A rf , (C, kH, kW))

[0029] Among them, F is the receptive field feature map. Each k×k region in F represents the features of a receptive field slider, and H and W are the length and width of the feature map respectively;

[0030] Step 4: Finally, use convolution to extract features. The convolution process can be expressed as:

[0031] X RFA = Conv(F)

[0032] Among them, Conv(·) is a convolution with equal stride and convolution size, and X RFA is the output feature;

[0033] 2) ECA implementation: The ECA module uses global average pooling and one-dimensional convolution techniques to accurately extract channel features. This module generates weights by learning the feature relationships between channels, thereby strengthening key features and weakening redundant information. The steps are as follows:

[0034] Step 1: Perform global average pooling on the input feature map to obtain the global information of each channel, and compress the spatial dimension of the feature map to 1. It can be specifically expressed as:

[0035]

[0036] Among them, X pool is the value after global average pooling of X Conv , H, W are the spatial sizes of the feature map, is the output X Conv of the preliminary extraction module, and the element value at the h-th row and w-th column;

[0037] Step 2: Perform dimensional transformation on X pool after global average pooling, and then apply one-dimensional convolution. It can be specifically expressed as:

[0038] X’ pool = Conv1D k (X pool )

[0039] Among them, K is the kernel size of one-dimensional convolution, which is determined by an adaptive method. The calculation formula for the kernel size k is as follows:

[0040]

[0041] Among them, C is the number of channels, and γ and b are hyperparameters (set as γ = 2, b = 1);

[0042] Step 3: Finally, apply the Sigmoid activation function to the output of the one-dimensional convolution to obtain the attention weights for each channel, and then multiply the attention weight tensor element-wise with the feature map X Conv The multiplication is specifically expressed as:

[0043]

[0044] Among them, X ECA is the output feature, X Conv is the output of the preliminary extraction module, and Sigmoid(X' pool ) is the channel attention weight.

[0045] In the above-mentioned hyperspectral image classification method based on the novel lightweight LCTNet model, in step S4, the TKformer module consists of a positional encoding, TKSA, and a multi-layer perceptron layer MLP, and is specifically expressed as:

[0046] X T = MLP(Norm(TKSA(Norm(X S ))))

[0047] Among them, X T is the output of the Transformer, MLP(·) is the multi-layer perceptron layer, TKSA(·) performs Top-k Sparse Attention, and X S is the output of the double-label module.

[0048] In the above-mentioned hyperspectral image classification method based on the novel lightweight LCTNet model, the implementation steps of introducing TKSA are as follows:

[0049] Step 1: Implement channel information interaction with a 1×1 convolution kernel, and then use a 3×3 depth convolution kernel to obtain the Q, K, and V matrices, which is specifically expressed as:

[0050] QKV = Conv3(Conv1(X s ))

[0051] Step 2: Calculate the pixel pair similarity between all rearranged queries and keys, then mask out unnecessary elements in the transposed attention matrix M, and retain the top K scores in the attention matrix M that contribute the most to the result to retain the most significant components and remove the useless parts. K is an adjustable parameter, and the sparsity is dynamically controlled by weighted averaging several appropriate ratios. Also, only normalize the Top-k values within the range of each row of M and then calculate Softmax. For other elements with scores less than the Top-k scores, we use a scattering function to replace their probabilities at a given exponent with 0. This dynamic selection makes the attention change from dense to sparse, which is obtained in the following way:

[0052]

[0053] where λ is an optional temperature factor, T k (·) is a learnable top-k selection operator, T k (·) can be expressed as:

[0054]

[0055] where t i is the k-th maximum value in the j-th row;

[0056] Step 3: Use the multi-head strategy, concatenate all the outputs of the multi-head attention, and then obtain the final result through linear projection. This process can be expressed as:

[0057] TKSA(Q, K, V) = Concat(SA1, SA2, …, SA h )

[0058] where h is the number of attention heads and SA is the attention score of a single head.

[0059] In the aforementioned hyperspectral image classification method based on the novel lightweight LCTNet model, the specific step S5 is as follows: perform global average pooling on the feature map learned in the previous convolution operation, compress the spatial dimension to 1×1, flatten the feature vector after global average pooling into a one-dimensional vector, and finally, input the flattened feature vector into a fully connected layer to perform weighted summation and activation on the input features, thereby outputting the classification of each category and completing the classification task of ground cover.

[0060] Advantages of the present invention: Compared with the prior art, the method of the present invention consists of three core modules connected in sequence, aiming to achieve efficient classification of hyperspectral images while maintaining the lightweight of the model. First, the spectral extraction-dimension reduction module extracts key spectral features from the processed hyperspectral data and reduces the dimension of the data, thereby reducing the computational burden of subsequent processing. Secondly, the spatial-spectral feature marking module is used to enhance the key spatial and spectral features in the image, ensuring the efficient fusion and utilization of spatial and spectral information. Finally, the TKformer module is adopted, which introduces an adaptive sparse attention mechanism, aiming to improve the model's ability to capture global features while maintaining the lightweight of the model. Experiments show that the present invention can effectively control the complexity of the model while improving the classification accuracy.

[0061] The method of the present invention can alleviate problems such as too high spectral dimension, complexity of features in hyperspectral images, and limitations of CNN and Transformer in the hyperspectral image classification task. Experimental results show that compared with other advanced hyperspectral image classification methods in recent years, the present invention shows significant advantages in improving the classification accuracy. At the same time, through means such as optimizing the model structure, the effective control of the model complexity is achieved, ensuring that the model has good efficiency and applicability while maintaining high accuracy. Description of the Drawings

[0062] Figure 1 is the specific structure of the LCTNet model;

[0063] Figure 2 is the specific structure of the receptive field attention convolution;

[0064] Figure 3 is the specific structure of the efficient channel attention;

[0065] Figure 4 is the specific structure of the TKSA. Detailed Embodiment

[0066] The following will combine Figures 1 to 4 to further elaborate on the technical solution of the present invention, but it is not used as a basis for limiting the present invention.

[0067] The present invention provides a hyperspectral image classification method based on a novel lightweight LCTNet model, which specifically includes the following steps:

[0068] Step S1: Process the original hyperspectral image to make it conform to the tensor shape required for the classification task and model input. The task objective of hyperspectral image classification in Step S1 is to accurately determine the ground object category corresponding to each pixel point based on the spectral and spatial information in the dataset. To achieve this goal, the dataset needs to be divided into several small cubes centered on pixel points, each containing both spectral features and spatial features.

[0069] Specifically: Process the dataset \(X\in\mathbb{R}^{}\) L×H×W To enable the use of 3D convolution on \(X\), add a dimension to \(X\) so that \(X\in\mathbb{R}^{}\) 1×L×H×W , and then divide the data into several overlapping spatio-spectral small cubes, each cube containing a pixel region of a fixed size. These small cubes are divided in a way that the pixel point with samples is the center, and expand outward to a three-dimensional region of \(L\times11\times11\) (where \(11\times11\) represents the expansion in the spatial dimension and \(L\) is the expansion in the spectral dimension). The label of the small cube is the same as the label of its central pixel point, and the size of the small cube is \(X\in\mathbb{R}^{}\) 1×L×11×11 .

[0070] Step S2: Extract key spectral features from the original hyperspectral data through a spectral extraction - dimensionality reduction module and reduce the dimensionality of the data, thereby reducing the computational burden of subsequent processing. In Step S2, there is dimensional redundancy in the spectral features of the hyperspectral image, which affects the analysis and calculation efficiency. Traditional PCA dimensionality reduction may lose spectral information. Therefore, a spectral extraction - dimensionality reduction module is designed to extract more spectral features and reduce the spectral dimension using 3D convolution technology, improving the analysis accuracy and efficiency.

[0071] Specifically: Use two 3D convolutions with different kernels and a 1D convolution to form the spectral extraction - dimensionality reduction module, as shown in Module 1 in Figure 1 . The specific process is as follows:

[0072] \(X^{}\) Conv \(=\text{Conv}(\text{Reshape}(\text{Conv3D}2(\text{Conv3D}1(X^{}\) HSI ))))

[0073] where \(X^{}\) HSI is the processed hyperspectral image, \(X^{}\) Conv is the output of the spectral extraction - dimensionality reduction module. \(\text{Conv3D}1(\cdot)\) is a convolution with a kernel of \((7\times1\times1)\), 8 output channels, and a stride of 3. \(\text{Conv3D}2(\cdot)\) is a convolution with a kernel of \((3\times1\times1)\), 8 output channels, and a stride of 2. \(\text{Reshape}(\cdot)\) means concatenating the feature maps along the spectral dimension.

[0074] Step S3: While effectively reducing the spectral dimension through the spatial-spectral feature marking module, dual feature annotation of spatial and spectral information is performed. To alleviate the complexity and multi-source nature of hyperspectral images, a spatial-spectral dual marking module is designed to emphasize the important feature regions of space and spectrum in HIS. This module significantly enhances the model's sensitivity to important information while suppressing the influence of redundant and noisy data. Through this method, the model can fully retain key feature information while reducing the number of parameters, thereby significantly improving the classification accuracy. The spatial-spectral feature marking module is mainly composed of receptive field attention convolution (RFAConv), efficient channel attention (ECA), and 1×1 convolution. X Conv They are respectively sent into RFAConv and ECA, and the obtained results are concatenated along the channels. Finally, a convolution with a kernel of 1×1 is used to reduce the dimension and perform channel interaction, as Figure 1 shown in Module 2. The specific process can be expressed by the formula as

[0075] X s = Conv(Cat(RFA(X Com ), ECA(X Com )))

[0076] = Conv(Cat(X RFA , X ECA ))

[0077] where, X S is the output of the spatial-feature marking module, Conv(·) is a 2D convolution with a kernel of 1×1 and an output channel of 64, RFA(·) is to perform receptive field attention convolution, and ECA(·) is to perform efficient channel attention operation.

[0078] Furthermore, the implementation processes of RFAConv and ECA will be described in detail below.

[0079] 1) Implementation of RFAConv: RFAConv generates a receptive field feature map by combining spatial features and attention weights, thus solving the limitation of convolutional kernel parameter sharing, ensuring that the features in each receptive field slider can be fully expressed, and thereby improving the performance of the model. The module structure is as Figure 2 shown, and its implementation steps are as follows:

[0080] Step 1: Use grouped convolution to obtain the spatial features of the original feature map. To avoid the problem of partial parameter sharing after the convolutional kernel is translated, the output channel of the grouped convolution is set to the square of the convolutional kernel. The specific process can be expressed as:

[0081] F rf = σ(g C (X Conv , C, k 2C, k×k))

[0082] where F rf is the spatial feature, σ(·) is the activation function, gC(·) is the grouped convolution, C is the number of channels of the original feature map, k 2 C is the number of channels of the feature map after convolution, and k is the kernel size. g C (X Conv , C, k 2 C, k×k) represents performing grouped convolution with a kernel of k×k, a group of C, an input channel of C, and an output channel of k 2 C on X Conv for convolution.

[0083] Step 2: While obtaining the spatial feature, obtain the attention weight. After aggregating the global information of each receptive field feature through average pooling. Then use 1×1 grouped convolution operation to interact the information to obtain the attention weight. Finally, use Softmax to emphasize the importance of each feature in the receptive field feature. The specific process can be expressed as:

[0084] A rf = Softmax(g(AvgPool(X Conv ), C, k 2 C, 1×1))

[0085] where, is the attention weight.

[0086] Step 3: After obtaining the spatial feature F rf and the attention weight A rf , multiply them point by point, and reorganize the extra channel features into the space through feature recombination, so that the new feature map contains all the features of the receptive field and there is no duplication. The specific process can be expressed as:

[0087] F = Reshape(F rf ×A rf , (C, kH, kW))

[0088] where F is the receptive field feature map, each k×k region in F represents the feature of a receptive field slider, ensuring that there is no partial parameter sharing after convolution, and H and W are the length and width of the feature map respectively.

[0089] Step 4: Finally, use convolution to extract features and solve the problem of feature sharing within the receptive field slider.

[0090] The convolution process can be expressed as

[0091] X RFA = Conv(F)

[0092] Among them, Conv(·) is a convolution with equal stride and convolution size, and X RFA is the output feature.

[0093] (2) ECA: The ECA module cleverly uses global average pooling and one-dimensional convolution techniques to accurately extract channel features, laying a solid foundation for image classification. By learning the feature relationships between channels, the module generates weights to strengthen key features and weaken redundant information, significantly enhancing the feature representation ability and classification efficiency. The module structure is as Figure 3 shown, and the steps are as follows:

[0094] Step 1: Perform global average pooling on the input feature map to obtain the global information of each channel. This step compresses the spatial dimensions (width and height) of the feature map to 1. Specifically, it can be expressed as

[0095]

[0096] where X pool is the value after global average pooling of X Conv , H and W are the spatial sizes of the feature map, is the output X of the preliminary extraction module Conv in which the element value at the h-th row and w-th column.

[0097] Step 2: To achieve cross-channel information interaction, generate the weights of each channel. Perform dimensional transformation on the globally average-pooled X pool and then apply one-dimensional convolution. Specifically, it can be expressed as

[0098] X’ pool = Conv1D k (X pool )

[0099] where K is the kernel size of one-dimensional convolution. To perform effective cross-channel interaction in the case of different numbers of channels, K is determined by an adaptive method. The calculation formula for the kernel size k is as follows

[0100]

[0101] where C is the number of channels, and γ and b are hyperparameters (set as γ = 2, b = 1).

[0102] Step 3: Finally, apply the Sigmoid activation function to the output of one-dimensional convolution to obtain the attention weights of each channel. Then multiply the attention weight tensor element-wise with the feature map X Conv . Specifically, it can be expressed as

[0103]

[0104] Among them, X ECA is the output feature, X Conv is the output of the preliminary extraction module, and Sigmoid(X' pool ) is the channel attention weight.

[0105] Step S4: Introduce Top-k sparse attention (TKSA) into the Transformer module to construct the TKformer module, which improves the accuracy while maintaining lightweight. To enhance the long-distance dependence of the model, the Transformer structure is used, and TKSA is introduced to form the TKformer. TKSA reduces the influence between irrelevant features by intelligently screening and focusing on the most relevant feature interactions, achieving a reduction in computational complexity while improving the accuracy. Combining multi-head attention and positional encoding, it accurately analyzes the complex characteristics of hyperspectral images and generates rich feature representations.

[0106] TKformer consists of positional encoding, TKSA, and MLP. The specific process can be expressed as:

[0107] X T = MLP(Norm(TKSA(Norm(X S ))))

[0108] Among them, X T is the output of the Transformer, MLP(·) is the multi-layer perceptron, TKSA(·) performs Top-k Sparse Attention, and X S is the output of the double-label module.

[0109] TKSA: Since the standard Transformer globally calculates self-attention for all Tokens, this may involve noisy interactions between irrelevant features, and the efficiency and accuracy will be affected. To address these limitations, we introduce TKSA, whose structure is as Figure 4 shown, and the implementation steps are as follows:

[0110] Step 1: Implement channel information interaction by a 1×1 convolutional kernel, and then use a 3×3 depth convolutional kernel to obtain the Q, K, V matrices. The purpose of using a 1×1 convolutional kernel is to apply self-attention to multiple channels rather than the spatial dimension to reduce time and memory complexity. Specifically, it can be expressed as

[0111] QKV = Conv3(Conv1(X s ))

[0112] Step 2: Calculate the pixel pair similarity between all rearranged queries and keys, and then mask out unnecessary elements in the transposed attention matrix M, where these elements correspond to lower attention weights. Different from the Dropout strategy that randomly discards scores, TKSA adopts an adaptive selection strategy, retaining the top K scores (Top-k) in the attention matrix M that contribute the most to the result to retain the most significant components and remove the useless parts. Here, K is an adjustable parameter, and the sparsity is dynamically controlled by weighted averaging of several appropriate ratios (such as 2 / 3). And only the Top-k values within the range of each row of M are normalized, and then Softmax is calculated. This method ensures that the retained elements are the most important, thus improving the efficiency and effectiveness of the model. For other elements with scores less than Top-k, we use a scattering function to replace their probabilities at a given exponent with 0. This dynamic selection makes the attention change from dense to sparse, which is obtained in the following way:

[0113]

[0114] where λ is an optional temperature factor, T k (·) is a learnable top-k selection operator. T k (·) can be expressed as

[0115]

[0116] where t i is the k-th maximum value in the j-th row.

[0117] Step 3: Use the multi-head strategy to concatenate all the outputs of the multi-head attention, and then obtain the final result through linear projection. This process can be expressed as

[0118] TKSA(Q,K,V) = Concat(SA1,SA2,…,SA h )

[0119] where h is the number of attention heads, and SA (SparseAttention) is the attention score of a single head.

[0120] Step S5: Perform global average pooling on the feature map and flatten it into a one-dimensional vector, input it into the fully connected layer, classify it, and output the type of ground cover. Specifically: Perform global average pooling on the feature map learned in the previous convolution operation. These feature maps contain the feature information extracted from the hyperspectral image. By calculating the global average value on each channel of the feature map, the spatial dimension is compressed to 1×1, effectively reducing the data volume and extracting the key channel feature information. Flatten the feature vector after global average pooling into a one-dimensional vector. Finally, input the flattened feature vector into a fully connected layer. This fully connected layer performs weighted summation and activation on the input features according to the association relationship between the input features and various category labels, thereby outputting the classification of each category and completing the classification task of the ground cover.

[0121] Taking the Pavia University (PU) dataset (tensor shape 103×610×340) as an example, the specific implementation process of the LCTNet model is described.

[0122] X ∈ R 103×610×340 is the initial shape of the PU dataset, where the numbers from right to left represent the spectral dimension, spectral and spatial sizes respectively. First, process the PU dataset. To be able to use 3DConv on X, add a dimension to X so that X ∈ R 1×103×610×340 . Then, divide the data into several overlapping spatio-spectral small cubes, and each cube contains a pixel region of a fixed size. These small cubes are divided in such a way that the pixel points with samples are centered, and extend outward to a three-dimensional region of 103×11×11 (where 11×11 represents the extension of the spatial dimension and 103 is the extension of the spectral dimension). The label of the small cube is the same as the label of its central pixel point. The size of the small cube is X ∈ R 1×103×11×11 .

[0123] Then, send the data into the LCTNet model. After processing the PU dataset, first, the spectral extraction - dimensionality reduction module extracts spectral features from X and reduces the spectral dimension. At this time, X ∈ R 64×11×11 . Then, send X into the designed spatio-spectral feature marking module to highlight the spatial and spectral information at different positions, so as to better extract the key local information and increase the classification accuracy. At this time, X ∈ R 32×11×11 . Then, while extracting global features through the TKformer module, shield the influence between irrelevant features. At this time, X ∈ R 32×11×11

[0124] Finally, classify the output of the TKformer. X performs global average pooling to achieve the aggregation of spatial information. Then, perform flattening processing. At this time, X ∈ R 32 , and then use the fully connected layer to achieve the classification of the ground cover.

[0125] The present invention covers any alternatives, modifications, equivalent methods and solutions made to the essence and scope of the present invention. In order to enable the public to thoroughly understand the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes and procedures are not described in detail in order to avoid unnecessary confusion to the essence of the present invention.

[0126] Although the disclosed embodiments of the present invention are as above, the content is only an embodiment adopted for the convenience of understanding the present invention and is not used to limit the present invention. Any person skilled in the art within the scope of the present invention can make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed by the present invention. However, the scope of patent protection of the present invention shall still be subject to the scope defined by the appended claims.

Claims

1. A hyperspectral image classification method based on a new lightweight model of LCTNet, characterized in that: It includes the following steps: S1: Process the original hyperspectral image to make it conform to the tensor shape required for the classification task and model input; S2: Extract key spectral features from the original hyperspectral data through a spectral extraction-dimension reduction module and reduce the data dimension; S3: While effectively reducing the spectral dimension through a spatial-spectral double-labeling module, perform double feature labeling on spatial and spectral information; S4: Introduce TKSA into the Transformer module to construct the TKformer module; S5: Perform global average pooling on the feature map obtained in the previous step to compress the spatial dimension and extract key channel features. Subsequently, flatten the pooled feature vector into a one-dimensional vector, input it into the fully connected layer, and perform classification to output the type of ground cover.

2. A hyperspectral image classification method based on the novel lightweight LCTNet model according to claim 1, characterized in that: In S1, the processing of the original hyperspectral image is specifically as follows: Divide the dataset into several small cubes centered on pixel points that contain both spectral features and spatial features.

3. A hyperspectral image classification method based on the novel lightweight LCTNet model according to claim 1, characterized in that: In S2, the spectral extraction-dimension reduction module consists of three-dimensional convolutions with two different kernels and a one-dimensional convolution. The specific process is as follows: X Conv = Conv(Reshape(Conv3D2(Conv3D1(X HSI )))) Among them, X HSI is the processed hyperspectral image, and X Conv is the output of the spectral extraction-dimension reduction module. Conv3D1(·) is a convolution with a kernel of (7×1×1), 8 output channels, and a stride of 3; Conv3D2(·) is a convolution with a kernel of (3×1×1), 8 output channels, and a stride of 2; Reshape(·) means splicing the feature maps along the spectral dimension.

4. A hyperspectral image classification method based on the novel lightweight LCTNet model according to claim 3, characterized in that: In the step S3, the spatial-spectral feature marking module is composed of a receptive field attention convolution (RFAConv), an efficient channel attention (ECA), and a 1×1 convolution. X Conv They are respectively sent into the RFAConv and the ECA, and the obtained results are concatenated along the channels. Finally, a convolution with a kernel of 1×1 is used to reduce the dimension and perform channel interaction. The specific process can be expressed as: X s = Conv(Cat(RFA(X Com ), ECA(X Com ))) = Conv(Cat(X RFA ,X ECA )) Among them, X S is the output of the spatial-feature marking module, Conv(·) is a 2D convolution with a kernel of 1×1 and 64 output channels, RFA(·) is a receptive field attention convolution, and ECA(·) is an efficient channel attention operation.

5. A hyperspectral image classification method based on the novel lightweight LCTNet model according to claim 4, characterized in that: The implementation processes of RFAConv and ECA are as follows: 1) RFAConv implementation: RFAConv generates a receptive field feature map by combining spatial features and attention weights. Its implementation steps are as follows: Step 1: Use grouped convolution to obtain the spatial features of the original feature map. The output channels of the grouped convolution are set to the square of the convolution kernel. The specific process can be expressed as: F rf = σ(g C (X Conv , C, k 2 , C, k×k)) Among them, F rf is a spatial feature, σ(·) is an activation function, and g C (·) is a grouped convolution, C is the number of channels of the original feature map, and k 2 C is the number of channels of the feature map after convolution, k is the kernel size, and g C (X Conv , C, k 2 C, k×k) represents performing a grouped convolution with a kernel of k×k, a group of C, an input channel of C, and an output channel of k 2 C on X Conv for convolution; Step 2: While obtaining the spatial features, obtain the attention weights. After aggregating the global information of each receptive field feature through average pooling, then use 1×1 grouped convolution operations to interact the information to obtain the attention weights. Finally, use Softmax to emphasize the importance of each feature in the receptive field feature. The specific process can be expressed as: A rf = Softmax(g(AvgPool(X Conv ), C, k 2 C, 1×1)) Among them, is the attention weight; Step 3: After obtaining the spatial feature F rf and the attention weight A rf perform a dot product on them, and through feature recombination, reorganize the extra channel features into the space, so that the new feature map contains the features of all receptive fields and there is no duplicate part. The specific process can be expressed as: F = Reshape(F rf × A rf , (C, kH, kW)) where F is the receptive field feature map, each k×k region in F represents the feature of a receptive field slider, and H and W are the length and width of the feature map respectively; Step 4: Finally, use convolution to extract features. The convolution process can be expressed as: X RFA = Conv(F) Among them, Conv(·) is a convolution with equal stride and convolution size, and X RFA is the output feature; 2) ECA implementation: The ECA module uses global average pooling and one-dimensional convolution techniques to accurately extract channel features. This module generates weights by learning the feature relationships between channels, thereby strengthening key features and weakening redundant information. Its steps are as follows: Step 1: Perform global average pooling on the input feature map to obtain the global information of each channel, and compress the spatial dimension of the feature map to 1. It can be specifically expressed as: Among them, X pool is the value after global average pooling, H and W are the spatial sizes of the feature map, Conv and is the output X of the preliminary extraction module Conv in the element value at the h-th row and the w-th column; Step 2: For X after global average pooling pool perform dimensional transformation and then apply one-dimensional convolution, which can be specifically expressed as: X′ pool = Conv1D k (X pool ) where K is the kernel size of the one-dimensional convolution, which is determined by an adaptive method. The calculation formula for the kernel size k is as follows: where C is the number of channels, and γ and b are hyperparameters (set as γ = 2, b = 1); Step 3: Finally, apply the Sigmoid activation function to the output of the one-dimensional convolution to obtain the attention weights for each channel, and then element-wise multiply the attention weight tensor with the feature map X Conv as follows: Among them, X ECA is the output feature, and X Conv is the output of the preliminary extraction module. Sigmoid(X' pool ) is the channel attention weight.

6. A hyperspectral image classification method based on the novel lightweight LCTNet model according to claim 1, characterized in that: In step S4, the TKformer module consists of a positional encoding, TKSA, and a multi-layer perceptron layer MLP. It can be specifically expressed as: X T = MLP(Norm(TKSA(Norm(X S )))) Among them, X T is the output of the Transformer, MLP(·) is the multi-layer perceptron layer, TKSA(·) performs Top-k Sparse Attention, and X S is the output of the dual-tagging module.

7. A hyperspectral image classification method based on the novel lightweight LCTNet model according to claim 6, characterized in that: The implementation steps for introducing TKSA are as follows: Step 1: Implement channel information interaction through a convolution with a kernel of 1×1, and then use a depth convolution with a kernel of 3×3 to obtain the Q, K, and V matrices. It can be specifically expressed as: QKV = Conv3(Conv1(X s )) Step 2: Calculate the pixel-pair similarity between all rearranged queries and keys, then mask out unnecessary elements in the transposed attention matrix M, and retain the top K scores that contribute the most to the result in the attention matrix M to retain the most significant components and remove the useless parts. K is an adjustable parameter, and the sparsity size is dynamically controlled by weighted averaging of several appropriate ratios. Moreover, only the Top-k values within the range of each row of M are normalized, and then Softmax is calculated. For other elements with scores less than the Top-k scores, we use a scattering function to replace their probabilities at a given exponent with 0. This dynamic selection makes the attention change from dense to sparse, which is obtained in the following way: where λ is an optional temperature factor, T k (·) is a learnable top-k selection operator, T k (·) can be expressed as: where t i is the k-th maximum value in the j-th row; Step 3: Use the multi-head strategy, concatenate all the outputs of the multi-head attention, and then obtain the final result through linear projection. This process can be expressed as: TKSA(Q, K, V) = Concat(SA1, SA2, …, SA h ) where h is the number of attention heads, and SA is the attention score of a single head.

8. A hyperspectral image classification method based on the novel lightweight LCTNet model according to claim 1, characterized in that: The specific content of the step S5 is as follows: perform global average pooling on the feature map learned in the previous convolution operation to compress the spatial dimension to 1×1. The feature vector after global average pooling is flattened into a one-dimensional vector. Finally, the flattened feature vector is input into a fully connected layer to perform weighted summation and activation on the input features, so as to output the classification of each category and complete the classification task of the ground cover.