Hyperspectral image classification method

By combining 3-D-2-D hybrid convolutional neural network, large-selective kernel LSK network and CASMamba module, the feature extraction problem in hyperspectral image classification is solved, and more efficient feature extraction and classification accuracy is achieved, adapting to hyperspectral remote sensing image data sets with different spatial resolutions.

CN120279313AActive Publication Date: 2025-07-08HENGYANG NORMAL UNIV

Patent Information

Application Number
CN202510346201.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Existing hyperspectral image classification methods are difficult to effectively utilize spatial-spectral joint features. Traditional methods such as LDA and PCA ignore spatial correlations. Shallow models such as SVM are difficult to adaptively extract deep features, and the high complexity of the Transformer architecture limits computing efficiency.

Method used

The 3-D-2-D hybrid convolutional neural network is combined with the large-selective core LSK network, and the CASMamba module is introduced to integrate the convolution addition self-attention and visual state space sequence model. The KANLinear module replaces the traditional linear layer to build a hyperspectral image classification model.

Benefits of technology

It improves the performance and robustness of hyperspectral image classification, effectively captures spectral-space features, reduces calculation complexity, and improves classification accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279313A_ABST
    Figure CN120279313A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image classification, and discloses a hyperspectral image classification method, which comprises the following steps: obtaining training data including hyperspectral image training data and corresponding image classification labels; constructing an initial image classification model, wherein the initial image classification model comprises a spectrum-space feature extraction module, a local feature extraction module and a KANLinear module which are connected in sequence; training the initial image classification model based on the training data to obtain a trained image classification model; and executing an image classification task of the to-be-classified hyperspectral image based on the trained image classification model. According to the technical scheme, the hyperspectral image classification task can be efficiently processed, the classification performance is improved, and computing resources are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image classification, and particularly relates to a hyperspectral image classification method. Background Art

[0002] Hyperspectral imaging technology, by improving spatial resolution and spectral coverage (hundreds of bands in the range of 400 - 2500 nanometers), fuses spatial-spectral multi-dimensional data and demonstrates important value in fields such as agricultural monitoring and urban planning. The core challenge of the classification task lies in analyzing complex spectral features. Traditional methods such as LDA, PCA and other spectral analysis techniques are limited because they ignore spatial correlations. Although spatial-spectral joint methods improve the effect, shallow models (such as SVM) are difficult to adaptively extract deep features. In recent years, the Transformer architecture has made breakthroughs with its long-range modeling ability, but the high complexity of its self-attention mechanism restricts the computational efficiency. In summary, hyperspectral image classification is of great significance in remote sensing applications, but traditional methods and existing technologies still have deficiencies. Future research should focus on developing more efficient and more adaptive feature extraction methods to make full use of the rich information of hyperspectral images and overcome the limitations of existing technologies. Summary of the Invention

[0003] The purpose of the present invention is to provide a hyperspectral image classification method to solve the problems existing in the above-mentioned prior art.

[0004] To achieve the above object, the present invention provides a hyperspectral image classification method, including: obtaining training data, where the training data includes hyperspectral image training data and corresponding image classification labels; constructing an initial image classification model, where the initial image classification model includes a spectral-spatial feature extraction module, a local feature extraction module, and a KANLinear module connected in sequence; training the initial image classification model based on the training data to obtain a trained image classification model; and performing an image classification task on the to-be-classified hyperspectral image based on the trained image classification model.

[0005] Optionally, the training process of the image classification model specifically includes: inputting the training data into the image classification model for classification prediction, and training according to the target loss function to obtain a trained image classification model.

[0006] Optionally, the processing process of the image classification model specifically includes: dividing the hyperspectral image data into multiple overlapping 3-D patches; using a spectral-spatial feature extraction module to extract features from each 3-D patch to obtain spectral-spatial features; inputting the extracted spatial features into a local feature extraction module, and performing feature extraction and fusion through convolutional addition self-attention and a visual state space sequence model to obtain fused features; inputting the fused features into a KANLinear module, and transforming the features through a learnable activation function parameterized by B-spline to output the final image classification result.

[0007] Optionally, the using of the spectral-spatial feature extraction module to extract features from each 3-D patch specifically includes: inputting the 3-D patch into a 3-D convolutional layer, and the activation value at the spatial position (x, y, z) in the feature map of the i th -th layer and the j th -th feature map is calculated as:

[0008]

[0009] where φ represents the activation function, B i,j represents the bias term, the dimension of the 3-D convolutional kernel is specified by H i , W i and R i respectively representing its height, width and depth. In addition, represents the weight corresponding to the position (h', w', r') in the θ th -th feature map; ω is the weight parameter of the convolutional kernel, and v represents the activation value of the previous layer;

[0010] Set k0 as the number of 3-D convolutional kernels used in the 3-D convolutional layer, the size of each 3-D kernel is k1×k2×k3, and the size of the output feature cube is k0@(s - k1 + 1)×(s - k2 + 1)×(B - k3 + 1). Further, input the output of the 3-D convolution into a 2-D convolutional layer, where the input size is (s - k1 + 1)×(s - k2 + 1)×k0(B - k3 + 1); in the 2-D convolutional layer, the activation value at the spatial position (x, y) is calculated by the following formula:

[0011]

[0012] where H' i and W' i respectively represent the height and width of the 2-D convolutional kernel, the parameter represents the weight corresponding to the position (h', w') related to the θ th -th feature map, B represents the spectral dimension, and s is the window length;

[0013] Furthermore, for the input X', a series of depth convolutions with different receptive fields are adopted to extract rich context features spanning different spatial ranges, as shown in the following formula:

[0014]

[0015] where represents the depth convolution operation. Suppose there are D deconstructed convolutional kernels, and each convolutional kernel is then processed by a 1×1 convolutional layer, denoted as

[0016]

[0017] The features obtained from multiple convolutional kernels with different receptive field ranges are merged, and channel maximum pooling and average pooling operations are performed on , denoted as P max (·) and P avg (·), respectively, in order to effectively capture spatial relationships. The specific calculation formulas include:

[0018]

[0019] where SA avg and SA max represent the average and maximum values of the pooled spatial descriptors respectively, represents the concatenated feature tensor obtained by merging the multi-scale features extracted from the convolutional kernels with different receptive fields. Further, a convolutional layer is used to connect and transform the pooled features to generate N spatial attention maps, thereby enhancing information interaction:

[0020]

[0021] In the formula, is the spatial attention map;

[0022] For each spatial attention map, the sigmoid activation function σ(·) is adopted, that is By performing a weighted operation on the deconstructed convolutional kernel features and the spatial mask, and then integrating through the convolutional layer F(·), the attention feature S is obtained:

[0023]

[0024] Element-wise multiplication of X' and S is performed, denoted as: Y' = X'·S. In the formula, Y' is the spectral-spatial feature.

[0025] Optionally, the process of obtaining the fusion feature specifically includes: dividing the input hyperspectral image data into a number of non-overlapping image patches through a patch embedding layer to obtain an embedded representation; inputting the embedded representation into parallel convolutional addition self-attention branch and visual state space sequence module branch for feature extraction; performing element-wise multiplication and residual connection processing on the features of the convolutional addition self-attention branch and the visual state space sequence module branch to fuse local features and long-range dependence features to obtain a fusion feature.

[0026] Optionally, the processing process of the convolutional addition self-attention branch specifically includes: extracting local features of the embedded representation through a convolutional operation and capturing global context information by combining additive self-attention.

[0027] Optionally, the processing process of the visual state space sequence module branch specifically includes: performing layer normalization processing on the input embedded representation, performing dual-path feature extraction and state space modeling on the normalized data to capture long-range dependence relationships; wherein, the first path captures linear patterns in the data through a linear transformation and a non-linear activation function, and the second path combines a linear transformation, a depthwise separable convolution and an activation function to improve computational efficiency and capture fine-grained spatial features; merging the outputs of the first path and the second path along the channel dimension to obtain long-range dependence features.

[0028] Optionally, the transformation of the feature by the learnable activation function parameterized by B-spline specifically includes: defining a regular grid and calculating the grid step size h:

[0029]

[0030] where G r is the range of the grid, defaulting to [-1, 1], generating the grid r is the dimension index of the input feature, G s is a parameter for controlling the grid size, D is the dimension of the input feature and S is the order of the B-spline function;

[0031] For each input feature calculate its relationship with the grid points and calculate the function value according to the recursive definition of the B-spline function; for high-order splines, the B-spline function is recursively calculated as:

[0032]

[0033] where is the k-th order B-spline function, g is a tensor composed of grid points for defining discrete points in each feature space, g k is the k-th node value in the grid point sequence, belonging to the discrete grid point tensor g that defines the support domain of the B-spline basis function, is the input eigenvalue;

[0034] adopt a weight W represented by a piecewise polynomial spline perform a further transpose transformation on the input obtain the final output y out :

[0035]

[0036] wherein, is the output obtained by calculating through the base weight, is the output obtained by weighting through the piecewise polynomial, W base is the weight matrix of a base linear transformation, is W base is the transpose of, and the final output is y out ∈R N×O , and O represents the output feature dimension.

[0037] The technical effects of the present invention are as follows:

[0038] The present invention combines the 3-D—2-D hybrid convolutional neural network with the large selection kernel LSK network, which can gradually construct feature representations and ensure the gradual feature abstraction from shallow to deep layers, thus effectively dealing with the complexity of hyperspectral images. The present invention proposes the CASMamba module, which synergistically integrates the local feature extraction ability of the convolutional layer, the ability of the self-attention mechanism to focus on key information, and the long-range dependence modeling ability of the visual state space sequence model VSSM. The combination of these three provides a more comprehensive and accurate feature extraction ability, thereby improving the classification performance. In the present invention, the KANLinear module is used to replace the traditional linear layer and provides stronger interpretability and flexibility. Through this alternative solution, the model can better capture the potential spatial regularity and spectral features, and this advantage helps to improve the performance of the model in hyperspectral images, especially in effectively dealing with high complexity when processing high-dimensional data. Description of the Drawings

[0039] Figure 1 is the flowchart of the hyperspectral image classification method based on the large selection kernel and convolutional addition self-attention Mamba in the embodiment of the present invention;

[0040] Figure 2 is the framework diagram of the hyperspectral image classification method based on the large selection kernel and convolutional addition self-attention Mamba in the embodiment of the present invention;

[0041] Figure 3 is the framework diagram of the CAS-VSSM module in the embodiment of the present invention;

[0042] Figure 4 This is the overall classification accuracy graph of each algorithm in the Botswana dataset in the embodiments of the present invention; among them, Figure 4 in (a) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm HybridSN in the embodiments of the present invention; Figure 4 in (b) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm LSFAT in the embodiments of the present invention; Figure 4 in (c) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm DBCT in the embodiments of the present invention; Figure 4 in (d) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm DCTN in the embodiments of the present invention; Figure 4 in (e) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm LSGA in the embodiments of the present invention; Figure 4 in (f) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm SSFTT in the embodiments of the present invention; Figure 4 in (g) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm MASSFormer in the embodiments of the present invention; Figure 4 in (h) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm CVSSN in the embodiments of the present invention; Figure 4 in (i) is the overall classification accuracy graph of the Botswana dataset based on the comparative algorithm SSmamba in the embodiments of the present invention; Figure 4 in (j) is the overall classification accuracy graph of the Botswana dataset based on HLSK-CASMamba proposed in this embodiment in the embodiments of the present invention;

[0043] Figure 5 This is the overall classification accuracy graph of each algorithm in the Houston2013 dataset in the embodiments of the present invention; among them, Figure 5 in (a) is the overall classification accuracy graph of the Houston2013 dataset based on the comparative algorithm HybridSN in the embodiments of the present invention; Figure 5 in (b) is the overall classification accuracy graph of the Houston2013 dataset based on the comparative algorithm LSFAT in the embodiments of the present invention; Figure 5 in (c) is the overall classification accuracy graph of the Houston2013 dataset based on the comparative algorithm DBCT in the embodiments of the present invention; Figure 5 in (d) is the overall classification accuracy graph of the Houston2013 dataset based on the comparative algorithm DCTN in the embodiments of the present invention; Figure 5Figure (e) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm LSGA in the embodiments of the present invention; Figure 5 Figure (f) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm SSFTT in the embodiments of the present invention; Figure 5 Figure (g) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm MASSFormer in the embodiments of the present invention; Figure 5 Figure (h) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm CVSSN in the embodiments of the present invention; Figure 5 Figure (i) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm SSmamba in the embodiments of the present invention; Figure 5 Figure (j) is the overall classification accuracy graph of the Houston2013 dataset based on HLSK-CASMamba proposed in the embodiments of the present invention;

[0044] Figure 6 is the overall classification accuracy graph of each algorithm in the Pavia University dataset in the implementation of the present invention; among them, Figure 6 Figure (a) is the overall classification accuracy graph of the Pavia University dataset based on the comparison algorithm HybridSN in the embodiments of the present invention; Figure 6 Figure (b) is the overall classification accuracy graph of the Pavia University dataset based on the comparison algorithm LSFAT in the embodiments of the present invention; Figure 6 Figure (c) is the overall classification accuracy graph of the Pavia University dataset based on the comparison algorithm DBCT in the embodiments of the present invention; Figure 6 Figure (d) is the overall classification accuracy graph of the Pavia University dataset based on the comparison algorithm DCTN in the embodiments of the present invention; Figure 6 Figure (e) is the overall classification accuracy graph of the Pavia University dataset based on the comparison algorithm LSGA in the embodiments of the present invention; Figure 6 Figure (f) is the overall classification accuracy graph of the Pavia University dataset based on the comparison algorithm SSFTT in the embodiments of the present invention; Figure 6 Figure (g) is the overall classification accuracy graph of the Pavia University dataset based on the comparison algorithm MASSFormer in the embodiments of the present invention; Figure 6 Figure (h) is the overall classification accuracy graph of the Pavia University dataset based on the comparison algorithm CVSSN in the embodiments of the present invention; Figure 6Figure (i) is the overall classification accuracy graph of the Pavia University dataset based on the contrast algorithm SSmamba in the embodiment of the present invention; Figure 6 Figure (j) is the overall classification accuracy graph of the Pavia University dataset based on the proposed HLSK-CASMamba in the embodiment of the present invention;

[0045] Figure 7 is the overall classification accuracy of 10 algorithms in each dataset under the training sample ratio of 2% - 10% in the embodiment of the present invention; among them, Figure 7 Figure (a) is the overall classification accuracy of 10 algorithms in the Botswana dataset under the training sample ratio of 2% - 10%; Figure 7 Figure (b) is the overall classification accuracy of 10 algorithms in the Houston2013 dataset under the training sample ratio of 2% - 10%; Figure 7 Figure (c) is the overall classification accuracy of 10 algorithms in the Pavia University dataset under the training sample ratio of 2% - 10%;

[0046] Figure 8 is the classification accuracy of 10 algorithms in each dataset under different Patch Size in the embodiment of the present invention; among them, Figure 8 Figure (a) is the classification accuracy of 10 algorithms in the Botswana dataset under different Patch Size; Figure 8 Figure (b) is the classification accuracy of 10 algorithms in the Houston2013 dataset under different Patch Size; Figure 8 Figure (c) is the classification accuracy of 10 algorithms in the Pavia University dataset under different Patch Size;

[0047] Figure 9 is the classification accuracy of 10 algorithms in each dataset under different Batch Size in the embodiment of the present invention; among them, Figure 9 Figure (a) is the classification accuracy of 10 algorithms in the Botswana dataset under different Batch Size; Figure 9 Figure (b) is the classification accuracy of 10 algorithms in the Houston2013 dataset under different Batch Size; Figure 9In (c), it is the classification accuracy of 10 algorithms in the embodiments of the present invention on the Pavia University dataset under different Batch Size values;

[0048] Figure 10 It is the classification accuracy of 10 algorithms in the embodiments of the present invention on different datasets under different learning rates; among them, Figure 10 In (a), it is the classification accuracy of 10 algorithms in the embodiments of the present invention on the Botswana dataset under different learning rates; Figure 10 In (b), it is the classification accuracy of 10 algorithms in the embodiments of the present invention on the Houston2013 dataset under different learning rates; Figure 10 In (c), it is the classification accuracy of 10 algorithms in the embodiments of the present invention on the Pavia University dataset under different learning rates. Detailed implementation manners

[0049] As Figure 1 - Figure 10 shown, in this embodiment, a hyperspectral image classification method is provided, including: obtaining training data, where the training data includes hyperspectral image training data and corresponding image classification labels; constructing an initial image classification model, where the initial image classification model includes a spectral-spatial feature extraction module, a local feature extraction module, and a KANLinear module connected in sequence; training the initial image classification model based on the training data to obtain a trained image classification model; and performing an image classification task on the to-be-classified hyperspectral image based on the trained image classification model.

[0050] The present embodiment provides a hyperspectral image classification method based on large selective kernel and convolutional additive self-attention Mamba. First, a spectral-spatial feature extraction module is constructed, which integrates 3-D convolutional layer, 2-D convolutional layer and large selective kernel network to efficiently extract abstract spectral and spatial features of hyperspectral images. Secondly, a new CASMamba model (Conv Additive Self-Attention Mam ba Module convolution incremental self-attention state space module) is proposed, and the core module CAS-VSSM combines convolutional additive self-attention and visual state space sequence model. This combination can not only extract local features using convolutional layers, but also model spatial dependencies through self-attention mechanism, and capture long-range dependencies through VSSM model. Finally, this embodiment introduces KANLinear module (Kolmogorov-Arnold Network Linear Module Kolmogorov-Arnold Network Linear Module), which replaces the traditional linear layer, so that the model can better cope with the high dimension and complexity of hyperspectral images, thereby obtaining higher classification accuracy and robustness. The method of this embodiment is based on a deep learning framework, making full use of the large selective kernel and convolution addition self-attention Mamba module to effectively capture the spectral-spatial features and deep semantic features of hyperspectral images. It constructs a hyperspectral remote sensing image dataset that adapts to different spatial resolutions and has strong portability, which can better meet the needs of image classification.

[0051] First, this embodiment constructs a spectral-spatial feature extraction module that integrates 3-D convolutional layers, 2-D convolutional layers, and large selective kernel networks to efficiently extract abstract spectral and spatial features of hyperspectral images. Secondly, a new CASMamba model is proposed. The core module CAS-VSSM combines convolutional addition self-attention and visual state space sequence model. This combination can not only extract local features using convolutional layers, but also model spatial dependencies through self-attention mechanisms and capture long-range dependencies through VSSM models. Finally, this embodiment introduces the KANLinear module to replace the traditional linear layer, further optimizing the sample label acquisition process.

[0052] The spectral-spatial feature extraction module specifically includes: The hyperspectral data cube is represented as I∈R M×N×B , where M represents the width of the image, N represents the height of the image, and B represents the spectral dimension. Each pixel in I corresponds to a label vector Y = (y1, y2, ..., y C), where C represents the land cover classification category. Hyperspectral images contain B bands, which can capture rich spatial-spectral features, but also bring significant processing complexity and redundant information. To address this issue, principal component analysis (PCA) is usually applied in the preprocessing stage to reduce the number of bands from B to K while preserving the original spatial structure. Therefore, the adjusted input is represented as X ∈ R M×N×K .

[0053] The input data is divided into multiple overlapping 3-D patches P ∈ R s×s×K , centered at spatial coordinates (m, n), where s × s represents the window size. The resulting 3-D patches are organized into a grid of size (M - s + 1) × (N - s + 1). Each 3-D patch P m,n , centered at (m, n), has a width range from m - (s - 1) / 2 to m + (s - 1) / 2 and a height range from n - (s - 1) / 2 to n + (s - 1) / 2.

[0054] To fully utilize the abstract spatial-spectral features of each sample patch, a hybrid convolutional layer is designed, which consists of a 3-D convolutional layer and a 2-D convolutional layer. Each training sample patch, of size s × s × K, is input into the 3-D convolutional layer. Based on this, the activation value at the spatial position (x, y, z) in the i th -th layer and j th -th feature map is calculated as follows:

[0055]

[0056] where φ represents the activation function, B i,j represents the bias term. The dimensions of the 3-D convolutional kernel are specified by H i , W i and R i , representing its height, width, and depth, respectively. Additionally, represents the weight corresponding to the position (h', w', r') in the θ th -th feature map, ω is the weight parameter of the convolutional kernel, and v represents the activation value of the previous layer.

[0057] Let k0 be the number of 3-D convolutional kernels used in the 3-D convolutional layer. The size of each 3-D kernel is k1 × k2 × k3, and the size of the output feature cube is k0 @ (s - k1 + 1) × (s - k2 + 1) × (B - k3 + 1). Further, the output of the 3-D convolution is input into the 2-D convolutional layer, where the input size is (s - k1 + 1) × (s - k2 + 1) × k0(B - k3 + 1). In the 2-D convolutional layer, the activation value at the spatial position (x, y) is calculated as follows:

[0058]

[0059] Among them, H' i and W' i represent the height and width of the 2-D convolutional kernel respectively. The parameter represents the weight corresponding to the position (h', w') related to the feature map of θ th .

[0060] Furthermore, for the input X', a series of depth convolutions with different receptive fields are adopted, aiming to extract rich context features spanning different spatial ranges. The formula is as follows: Among them, represents the depth convolution operation. Suppose there are D deconstructed convolutional kernels, and each convolutional kernel is subsequently processed by a 1×1 convolutional layer, denoted as In this part, the feature maps generated by large convolutional kernels of multiple scales are used to enhance the classification performance by focusing on the most relevant spatial context to distinguish different ground object categories. The specific implementation process is as follows:

[0061] First, the features obtained from multiple convolutional kernels with different receptive field ranges are merged. Then, for , channel maximum pooling and average pooling operations are performed, denoted as P max (·) and P avg (·) respectively, in order to effectively capture the spatial relationship.

[0062]

[0063] Among them, SA avg and SA max represent the average value and the maximum value of the pooled spatial descriptors respectively, represents the concatenated feature tensor obtained after extracting multi-scale features by merging convolutional kernels with different receptive fields. In addition, a convolutional layer is further adopted to connect and transform the pooled features to generate N spatial attention maps, thereby enhancing the information interaction.

[0064]

[0065] In order to generate different spatial selection masks for the deconstructed large convolutional kernels, the sigmoid activation function σ(·) is adopted for each spatial attention map, that is In addition, the attention feature S is obtained by weighting the deconstructed convolutional kernel features with the spatial mask and then integrating them through the convolutional layer F(·).

[0066]

[0067] Finally, perform an element-wise multiplication of X' and S, denoted as: Y' = X'·S, where Y' is the spectral-spatial feature.

[0068] The proposed spectral-spatial feature extraction module includes the following stages: First, use a 3-D convolutional layer to extract shallow spectral feature representations from the principal components. Then, apply a 2-D convolutional layer to capture more abstract spectral and spatial features. Finally, by dynamically adjusting a larger spatial receptive field, use the large selection kernel module to enhance the extraction of deep context information, so as to adapt to the changes of different land cover classes.

[0069] The convolutional addition self-attention Mamba module specifically includes: The CASMamba network first divides the input x ∈ R H×W×3 into non-overlapping 4×4 image patches through a patch embedding layer to obtain an embedded representation Then, the embedded data undergoes layer normalization to standardize the feature distribution, and then is processed by the CA SMamba backbone network. The CAS-VSSM module is the core module of the CASMamba network. This module adopts a dual-branch structure, that is, it combines the convolutional addition self-attention module and the visual state space sequence module.

[0070] In the convolutional addition self-attention branch, the CAS mechanism extracts local features from the input by combining convolutional operations with additive self-attention, innovatively capturing both local spatial dependencies and global context information at the same time. This mechanism optimizes efficiency by reducing computational complexity and has lower computational overhead compared to traditional attention mechanisms. The CAS mechanism dynamically weights local features, enhances the network's focusing ability in key regions, improves the feature extraction effect, and at the same time avoids a significant increase in computational cost.

[0071] In the visual state space sequence module (VSSM) branch, the input is first normalized by layer normalization to stabilize the learning process, and then divided into two processing paths, which focus on different feature extraction strategies respectively. The first path captures the linear patterns in the data through linear transformation and non-linear activation functions. The second path combines linear transformation, depthwise separable convolution and activation functions, uses depthwise separable convolution to improve computational efficiency, reduce the number of parameters, and at the same time maintain the capture of fine-grained spatial features. The VSSM promotes high-level feature extraction by dynamically modeling the state space sequence and enhances the network's understanding of multi-scale spatio-temporal dependencies. The integration of the two paths balances computational efficiency and the extraction capabilities of low-level and high-level features.

[0072] The extracted features are normalized across layers and multiplied element-wise with the output of the first branch to achieve feature fusion. The fused features are processed through a linear layer and a residual connection is introduced to generate the final output of the CAS-VSS M module. In the VSSM branch, the default activation function is SiLU. At the final stage of the network, the outputs from the two branches are merged along the channel dimension and fused through a 1×1 convolutional layer to facilitate information interaction between channels.

[0073] Finally, by combining the CAS mechanism in the CAS branch with the state space modeling in the VSSM branch, this architecture can effectively model the complex interactions in the input, thereby improving performance in tasks that require fine-grained recognition and temporal awareness while maintaining high computational efficiency. The specific implementation process is as follows:

[0074] In the CAS module, the similarity function is calculated as:

[0075] Sim(Q,K) = φ(Q) + φ(K) s.t. φ(Q) = C(S(Q))

[0076] where the function φ(Q) represents the context mapping function, and φ(K) includes the sigmoid-based channel attention C(·) ∈ R N×d and the spatial attention S(·) ∈ R N×d . Based on this, the output of CAS is expressed as:

[0077] O = τ(φ(Q) + φ(K)) × V

[0078] where τ(·) ∈ R N×d represents the linear transformation for integrating context information, and O(N) represents the complexity. The traditional state space model can be conceptualized as a linear time-invariant system that transforms the input sequence x(t) ∈ R d×L into the output response y(t) ∈ R through the hidden state h(t) ∈ R. d represents the dimension of the state and L represents the sequence length.

[0079] This transformation is calculated as:

[0080]

[0081] where t represents the time step of the current input, h'(t) represents the hidden state corresponding to the current input x(t), and h(t) represents the hidden state of the previous time step. The matrix A ∈ R d×d represents the evolution parameter, B ∈ R d×L and C ∈ R d×L represent the projection parameters respectively.

[0082] Traditional state - space models have limited ability to adapt to different inputs when training parameters, which restricts their ability to effectively model dynamic systems. To address this limitation, the VSSM model is introduced as a discretized version of the continuous SSM. In this module, the continuous parameters A and B are converted into discrete parameters and using the step - size parameter △. Zero - order hold (ZOH) is the most commonly used discretization method, and the calculation formula is as follows:

[0083]

[0084] where I represents the identity matrix. The equation of the discretized state - space model is as follows:

[0085]

[0086] where, h p represents the hidden state at time step p, X p and y p represent the input sequence and output sequence at time step p, respectively.

[0087] The KANLinear module specifically includes: When determining labels for each pixel category, in this embodiment, the KANLinear layer is used to replace the traditional linear layer. Different from traditional neural network architectures that use fixed activation functions at the neuron level, the KANLinear layer applies learnable activation functions at the edges (connections) of the network. These activation functions are parameterized by B - splines, and B - splines provide a flexible and stable method for modeling complex relationships between features. This replacement not only enhances the model's representation ability but also improves computational efficiency by reducing memory usage. This flexibility and efficiency make the KANLinear layer have stronger advantages in hyperspectral image classification because hyperspectral image classification requires effective processing and classification of complex high - dimensional data.

[0088] The input tensor is where N is the batch size and D is the dimension of the input features, which is calculated as:

[0089]

[0090] where, is the feature vector of the i - th sample. This module uses a regular grid for interpolation, and the size of the grid is controlled by the G s parameter, and the interval of the grid is controlled by the G r parameter. For each feature in the input data, a grid is established within the defined interval of the feature.

[0091] First, calculate the grid step size h:

[0092]

[0093] Among them, G r is the range of the grid, defaulting to [-1, 1]. Then, the grid is generated r is the dimension index of the input feature, and S is the order of the B-spline function, indicating that a certain number of piecewise spline functions are used for interpolation on each feature.

[0094] For each input feature Calculate its relationship with the grid points and calculate the function value according to the recursive definition of the B-spline function. B(X) is the B-spline function. First, define an exponential function on each grid to mark the grid interval where the input feature is located. That is, if X i is between two consecutive grid points, then B(X) is 1, otherwise it is 0.

[0095] For high-order splines, the B-spline function is recursively calculated as follows:

[0096]

[0097] Among them, is the k-th order B-spline function, g is a tensor composed of grid points, used to define the discrete points in each feature space. g k is the k-th node value in the grid point sequence, belonging to the discrete grid point tensor g that defines the support domain of the B-spline basis function, is the input feature value;

[0098] Finally, for the accurate modeling of hyperspectral images, a weight W represented by a piecewise polynomial spline is used to perform a further transpose transformation on the input to obtain the final output y out :

[0099]

[0100] Among them, is the output obtained by calculating through the base weight, is the output obtained after weighted by the piecewise polynomial. W base is the weight matrix of a base linear transformation, is the transpose of W base , and the final output is y out ∈R N×O , where O represents the output feature dimension.

[0101] This embodiment proposes a new algorithm, a hyperspectral image classification method based on large selective kernels and convolutional addition self-attention Mamba, for hyperspectral image classification tasks. Although convolutional neural networks (CNNs) are effective in hyperspectral image classification tasks, they often struggle to fully capture complex semantic features, and as the network depth increases, the computational cost rises significantly. In contrast, although Transformer has good results in modeling spectral-spatial dependencies, due to its complexity, it also brings significant computational overhead. The Mamba model uses a state space model to provide an effective alternative, which efficiently captures long-range dependencies in HSI with linear complexity while ensuring computational efficiency. This feature not only improves classification performance but also optimizes computational resources, making it a promising method for hyperspectral image applications. This embodiment proposes the HLSK-Mamba model, which introduces three key modules to improve performance and efficiency. First, a spectral-spatial feature extraction module is constructed, combining 3-D convolutional layers, 2-D convolutional layers, and a large selective kernel network to effectively extract abstract spectral and spatial features in hyperspectral images. Second, this embodiment proposes a new CASMamba model, whose core module CAS-VSSM combines convolutional additive self-attention and a visual state space sequence model. This integrated method effectively utilizes the local feature extraction ability of convolutional layers, the ability of the self-attention mechanism to focus on key information, and the long-range dependency modeling ability of VSSM. Finally, the KANLinear module is used to better handle the high dimensionality and complexity of hyperspectral images, thereby enhancing the ability to obtain accurate sample labels.

[0102] To effectively capture the abstract spatio-spectral feature representation and the inherent unique prior knowledge in hyperspectral images, this embodiment proposes an architecture that combines a 3-D - 2-D hybrid convolutional neural network with a large selection kernel (LSK) network. Such an architecture can gradually construct feature representations, ensuring gradual feature abstraction from shallow to deep layers, thus effectively handling the complexity of hyperspectral images.

[0103] This embodiment proposes the CASMamba module, which synergistically integrates the local feature extraction ability of convolutional layers, the ability of the self-attention mechanism to focus on key information, and the long-range dependency modeling ability of the visual state space sequence model (VSSM). The combination of these three provides a more comprehensive and accurate feature extraction ability, thus enhancing the classification performance.

[0104] In this embodiment, the KANLinear module is used to replace the traditional linear layer, providing stronger interpretability and flexibility. Through this replacement, the model can better capture potential spatial regularities and spectral features, which helps improve the model's performance in hyperspectral images, especially in effectively handling high complexity when dealing with high-dimensional data.

[0105] This embodiment conducted a series of extensive experiments on three publicly available benchmark hyperspectral image datasets: Houston2013, Botswana, and Pavia University. The experimental results show that the network proposed in this embodiment outperforms some of the state-of-the-art methods in this field.

[0106] This embodiment proposes a new algorithm, a hyperspectral image classification method based on large selective kernel and convolutional addition self-attention Mam ba, for hyperspectral image classification tasks. Although convolutional neural networks (CNNs) are effective in hyperspectral image classification tasks, they often struggle to fully capture complex semantic features, and with the increase in network depth, the computational cost increases significantly. In contrast, although Transformer has good results in modeling spectral-spatial dependence relationships, due to its complexity, it also brings significant computational overhead. The Mamba model uses a state space model to provide an effective alternative. While ensuring computational efficiency, it can efficiently capture long-range dependence relationships in HSI with linear complexity. This feature not only improves classification performance but also optimizes computational resources, making it a promising method for hyperspectral image applications. This embodiment proposes the HLSK-Mamba model, which introduces three key modules to improve performance and efficiency. First, a spectral-spatial feature extraction module is constructed, combining 3-D convolutional layers, 2-D convolutional layers, and a large selective kernel network to effectively extract abstract spectral and spatial features in hyperspectral images. Second, this embodiment proposes a new CASMamba model, whose core module CAS-VSSM combines convolutional additive self-attention and a visual state space sequence model. This integrated method effectively utilizes the local feature extraction ability of convolutional layers, the ability of the self-attention mechanism to focus on key information, and the long-range dependence modeling ability of VSSM. Finally, the KANLinear module is used to better handle the high dimensionality and complexity of hyperspectral images, thereby enhancing the ability to obtain accurate sample labels.

[0107] Figure 1It is a flowchart of a hyperspectral image classification method based on a large selective kernel and convolutional addition self-attention Mamba. The process mainly includes: First, a spectral-spatial feature extraction module is constructed, integrating 3-D convolutional layers, 2-D convolutional layers, and a large selective kernel network to efficiently extract the abstract spectral and spatial features of hyperspectral images. Second, a novel CASMamba module is proposed, and the core module CAS-VSSM combines convolutional addition self-attention and a visual state space sequence model. This combination can not only use convolutional layers to extract local features, but also use the self-attention mechanism to focus on key information, and can also use the VSSM model to capture long-range dependencies. Finally, the KANLinear module is used to further optimize the process of obtaining sample labels.

[0108] Figure 2 It is a framework diagram of a hyperspectral image classification method based on a large selective kernel and convolutional addition self-attention Mamba. The basic framework includes a spectral-spatial feature extraction module, a CASMamba module, and a KANLinear module.

[0109] Figure 3 It is a framework diagram of the CAS-VSSM module, including a CAS branch module and a VSSM branch module.

[0110] Figure 4 (a)- Figure 4 (j) shows the classification results of 10 algorithms on the Botswana dataset. These algorithms include HybridSN, LSFAT, DBCT, DCTN, LSGA, SSFTT, MASSFormer, CVSSN, SSmamba, and HLSK-CASMamba. By comparing these classification results, it can be seen that the proposed HLSK-CASMamba algorithm (i.e., Figure 4 (j)) shows the best classification effect.

[0111] Figure 5 (a)- Figure 5 (j) shows the classification results of 10 algorithms on the Houston2013 dataset. These algorithms include HybridSN, LSFAT, DBCT, DCTN, LSGA, SSFTT, MASSFormer, CVSSN, SSmamba, and HLSK-CASMamba. By comparing these classification results, it can be seen that the proposed HLSK-CASMamba algorithm (i.e., Figure 5 (j)) shows the best classification effect.

[0112] Figure 6 (a)- Figure 6(j) shows the classification results of 10 algorithms on the Pavia University dataset. These algorithms include HybridSN, LSFAT, DBCT, DCTN, LSGA, SSFTT, MAS SFormer, CVSSN, SSmamba, and HLSK-CASMamba. By comparing these classification results, it can be seen that the proposed HLSK-CASMamba algorithm (i.e., Figure 6 (j)) exhibits the best classification performance.

[0113] Figure 7 (a)- Figure 7 (c) show the changes in the classification accuracy of ten algorithms when using different proportions of training samples on three different datasets.

[0114] Refer to Figure 8 (a)- Figure 8 (c) show the changes in the classification accuracy under different Patch Sizes. In Figure 8 (a), for the Botswana dataset, the best classification performance is obtained when the Patch Size is 11×11. In Figure 8 (b), for the Houston2013 dataset, the best classification performance is obtained when the Patch Size is 9×9. In Figure 8 (c), for the Pavia University dataset, the best classification performance is obtained when the Patch Size is 13×13.

[0115] Refer to Figure 9 (a)- Figure 9 (c) show the changes in the classification accuracy under different Batch Sizes. In Figure 9 (a), for the Botswana dataset, the best classification performance is obtained when the Batch Size is 64. In Figure 9 (b), for the Houston2013 dataset, the best classification performance is obtained when the Batch Size is 64. In Figure 9 (c), for the Pavia University dataset, the best classification performance is obtained when the Batch Size is 64.

[0116] Refer to Figure 10 (a)- Figure 10 (c) show the changes in the classification accuracy under different learning rates. In Figure 10 (a), for the Botswana dataset, the best classification performance is obtained when the learning rate is 2 e-3 When. InFigure 10 In (b), for the Houston2013 dataset, the best classification effect was obtained when the learning rate was 2 e-3 . In Figure 10 (c), for the PaviaUniversity dataset, the best classification effect was obtained when the learning rate was 2 e-3 .

[0117] The above are specific examples of this embodiment and are not used to limit this embodiment. The hyperspectral image classification method based on the large selection kernel and convolutional addition self-attention Mamba provided in this embodiment is also applicable to classifying other non-hyperspectral images. Without departing from the essence and scope of this embodiment, some adjustments and optimizations can be made, and the protection scope of this embodiment shall be subject to the claims.

Claims

1. A hyperspectral image classification method, characterized in that, Including: Obtain training data, where the training data includes hyperspectral image training data and corresponding image classification labels; Construct an initial image classification model, where the initial image classification model includes a spectral-spatial feature extraction module, a local feature extraction module, and a KANLinear module connected in sequence; Train the initial image classification model based on the training data to obtain a trained image classification model; Perform an image classification task on the hyperspectral image to be classified based on the trained image classification model.

2. The hyperspectral image classification method according to claim 1, characterized in that The training process of the image classification model specifically includes: Input the training data into the image classification model for classification prediction, and train according to the target loss function to obtain a trained image classification model.

3. A hyperspectral image classification method according to claim 1, characterized in that The processing process of the image classification model specifically includes: Divide the hyperspectral image data into multiple overlapping 3-D patches; Use the spectral-spatial feature extraction module to extract features from each 3-D patch to obtain spectral-spatial features; Input the extracted spatial features into the local feature extraction module, and perform feature extraction and fusion through convolutional addition self-attention and visual state space sequence models to obtain fused features; Input the fused features into the KANLinear module, transform the features through a learnable activation function parameterized by B-spline, and output the final image classification result.

4. A hyperspectral image classification method according to claim 3, characterized in that, The process of using the spectral-spatial feature extraction module to extract features from each 3-D patch specifically includes: Input the 3-D patch into the 3-D convolutional layer at the i-th th layer and the j-th th activation value at the spatial location (x, y, z) in the feature map is calculated as: Among them, φ represents the activation function, and B i,j represents the bias term. The dimensions of the 3-D convolutional kernel are specified by H i , W i and R i respectively representing its height, width, and depth. In addition, represents the weight corresponding to the position (h', w', r') in the θ th feature map; ω is the weight parameter of the convolutional kernel, and v represents the activation value of the previous layer; Set k0 as the number of 3-D convolution kernels used in the 3-D convolution layer, the size of each 3-D kernel is k1×k2×k3, and the size of the output feature cube is k0@(s - k1 + 1)×(s - k2 + 1)×(B - k3 + 1). Further, input the output of the 3-D convolution into the 2-D convolution layer, where the input size is (s - k1 + 1)×(s - k2 + 1)×k0(B - k3 + 1); in the 2-D convolution layer, the activation value at the spatial position (x, y) is calculated by the following formula: Among them, H' i and W' i represent the height and width of the 2-D convolution kernel respectively, and the parameter represents the weight corresponding to the position (h', w') related to the feature map of θ th , B represents the spectral dimension, and s is the window length; Further, for the input X', a series of depth convolutions with different receptive fields are used to extract rich context features spanning different spatial ranges. The formula is as follows: Among them, represents a depth convolution operation. Suppose there are D deconstructed convolutional kernels, and each convolutional kernel is subsequently processed by a 1×1 convolutional layer, denoted as Merge the features obtained from multiple convolutional kernels with different receptive field ranges, and perform channel maximum pooling and average pooling operations, denoted as P max (·) and P avg (·), respectively, in order to effectively capture spatial relationships. The specific calculation formulas are as follows: Among them, SA avg and SA max respectively represent the average value and the maximum value of the spatial descriptors after pooling. represents the concatenated feature tensor obtained by merging the multi-scale features extracted by the convolution kernels with different receptive fields, and further uses a convolutional layer to connect and transform the features after pooling, generating N spatial attention maps to enhance the interaction of information: In the formula, is the spatial attention map; Apply the sigmoid activation function σ(·) to each spatial attention map, i.e., By performing a weighted operation on the deconstructed convolution kernel features and the spatial mask, and then integrating through the convolution layer F(·), the attention feature S is obtained: Perform an element-wise multiplication of X' and S, denoted as: Y′ = X′·S In the formula, Y' is the spectral-spatial feature.

5. A hyperspectral image classification method according to claim 3, characterized in that, The process of obtaining the fused features specifically includes: Divide the input hyperspectral image data into several non-overlapping image patches through a patch embedding layer to obtain an embedded representation; Input the embedded representation into parallel convolutional addition self-attention branches and visual state space sequence module branches for feature extraction; Perform element-wise multiplication and residual connection processing on the features of the convolutional addition self-attention branch and the visual state space sequence module branch to fuse local features and long-range dependence features to obtain fused features.

6. A hyperspectral image classification method according to claim 5, characterized in that, The processing process of the convolutional addition self-attention branch specifically includes: Extract the local features of the embedded representation through convolution operations, and capture the global context information by combining additive self-attention.

7. A hyperspectral image classification method according to claim 5, characterized in that, The processing process of the visual state space sequence module branch specifically includes: Perform layer normalization on the input embedded representation, and perform dual-path feature extraction and state space modeling on the normalized data to capture long-range dependencies; among them, the first path captures the linear patterns in the data through linear transformation and non-linear activation functions, and the second path combines linear transformation, depthwise separable convolution and activation functions to improve computational efficiency and capture fine-grained spatial features; Merge the outputs of the first path and the second path along the channel dimension to obtain long-range dependence features.

8. A hyperspectral image classification method according to claim 3, wherein The transformation of the features by the learnable activation function parameterized by B-spline specifically includes: Define a regular grid and calculate the grid step size h: where G r is the range of the grid, defaulting to [-1, 1], and generating the grid r is the dimensional index of the input feature, G s is the parameter for controlling the grid size, D is the input feature dimension, and S is the order of the B-spline function; For each input feature calculate its relationship with the grid points and calculate the function values according to the recursive definition of the B-spline function; For high-order splines, the B-spline function is recursively calculated as: Among them, is the k-th order B-spline function, and g is a tensor composed of grid points used to define the discrete points of each feature space. g k is the k-th node value in the grid point sequence and belongs to the discrete grid point tensor g that defines the support domain of the B-spline basis function. is the input eigenvalue; A weight W represented by a piecewise polynomial spline Perform a further transpose transformation on the input Obtain the final output y out : Among them, is the output obtained by calculating the basic weight, is the output obtained after weighted by piecewise polynomial, and W base is the weight matrix of a basic linear transformation, is the transpose of W base , and the final output is y out ∈R N×O , where O represents the output feature dimension.

Citation Information

Patent Citations

  • Hyperspectral image classification method and device based on spatial-spectral double-branch convolutional network

    CN115249332A

  • Hyperspectral remote sensing image classification method based on hybrid convolutional neural network

    CN115909052A

Cited By

  • Intelligent urine component detection method based on spectrum identification and deep learning

    CN121190854A

  • Medical image classification method based on MedConvMama model

    CN121904444A