Hyperspectral classification method and system based on wavelet spatial spectrum Mama
Through the wavelet spatial spectral Mamba method, combined with Haar wavelet transformation and Mamba network, the problems of long-range dependence and computational complexity in hyperspectral image classification are solved, and efficient multi-scale feature extraction and global modeling are realized, which improves classification accuracy and reduces computational cost.
Patent Information
- Application Number
- CN202510472829.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
The existing hyperspectral image classification methods are difficult to effectively handle long-range dependencies, and the computational complexity is high, resulting in inefficient classification, especially on large-scale hyperspectral image datasets.
The wavelet spatial spectral Mamba method is used to decompose the hyperspectral image into overlapping 3D blocks, combine the Haar wavelet transformation and the Mamba network of the state space model to perform multi-scale feature extraction and long-range dependency capture, and use the deep separable convolutional network for spatial and spectral feature learning, and control the model complexity through L2 regularization.
It significantly improves classification accuracy, reduces computing resource consumption, has superior robustness and generalization capabilities, and is suitable for environmental monitoring and land use management and other fields.
Smart Images

Figure CN120339710A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a hyperspectral classification method and system based on wavelet spatial spectrum Mamba, belonging to the technical fields of hyperspectral image processing and deep learning. Background Art
[0002] The hyperspectral image classification method aims to solve the problem of accurately distinguishing different ground object categories from hyperspectral data. The hyperspectral imaging technology captures surface detail information through continuous spectral bands, providing rich data for fine classification. This technology has been widely applied in fields such as urban expansion monitoring, land cover change analysis, and disaster assessment.
[0003] The hyperspectral image classification task performs pixel-level classification by learning the spectral and spatial features of pixels. At the same time, hyperspectral images contain a large amount of spectral information, and each pixel has hundreds of bands. Therefore, this places higher requirements on the information integration and learning capabilities of the model. Traditional machine learning techniques and deep learning techniques are both applied in the hyperspectral image classification task. Convolutional neural networks show good performance in processing multi-modal data, but it is difficult for them to handle the long-term dependency of categories with similar spectra. Recurrent neural networks can simulate these long-term dependencies, but lack the ability of model synchronous training, which is a challenge when dealing with large-scale hyperspectral image datasets. To address this problem, the network model Vision Transformer (ViT) based on the Transformer architecture uses the self-attention mechanism to directly capture the long-term dependencies of images and has the ability of parallel training. However, the learning of long-term dependencies by ViT limits the learning of fine-grained texture detail features, and its time complexity of O(N 2 ) leads to low training and inference efficiency of the hyperspectral image classification model. Therefore, there is an urgent need for a hyperspectral classification method that can simultaneously take into account multi-scale feature extraction, have global modeling capabilities, and low time complexity. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above deficiencies and provide a hyperspectral classification method based on wavelet spatial spectrum Mamba. This method can significantly improve the classification accuracy while reducing the consumption of computing resources, has excellent robustness and generalization capabilities, and has broad application prospects in fields such as land use monitoring and environmental management.
[0005] The technical solution adopted by the present invention is as follows: The hyperspectral classification method based on wavelet spatial spectrum Mamba includes the following steps: S1. Preprocess the input hyperspectral image and divide the image into 3D small blocks with overlapping regions; S2. Perform spatial and spectral feature learning on each 3D patch image respectively to generate spatial feature S and spectral feature F; S3. Perform Haar wavelet transform on both the initially extracted spatial feature S and spectral feature F to generate multiple wavelet subbands, and concatenate the multiple wavelet subbands obtained by transforming S and F along the channels to obtain a new 3D representation, i.e., the feature sequence; S4. Feed the wavelet-transformed feature sequence into the Mamba network based on the state space model, and capture the spatio-temporal dependence relationship through the state transition and update mechanism with linear time complexity; take the feature sequence through a linear transformation as the input of the Mamba network, and let the input be Let the state vector h t represent the hidden state at time t , and its state transition and update formulas are as follows: , where, W transition and W update are learnable weights of the network, ReLU(•) is the activation function, h t-1 is the hidden state at the previous moment, E t is the wavelet feature after the current linear transformation, representing the feature at a certain time step (or sequence position); S5. Input the final hidden state h t obtained by the network learning of the 3D patch into a linear classifier for classification to obtain the classification result corresponding to the central pixel of the 3D patch. Perform the same calculation for each 3D patch, and assign the classification results of each 3D patch according to their positions in the image to obtain the classification result of the entire input image.
[0006] In the above method, in step S2, a depthwise separable convolutional network and a classical convolutional network are respectively used for spatial and spectral feature learning.
[0007] In step S3, the Haar wavelet transform uses low-pass and high-pass filters to decompose step by step along the row and column directions to generate multiple subbands. The specific process of the Haar wavelet transform is to first apply the low-pass filter and the high-pass filter along the "row" direction to generate two subbands and , and then apply the same filters to these two subbands along the "column" direction to finally obtain four wavelet subbands , , . , these four wavelet subbands respectively capture the low-frequency approximation information, horizontal high-frequency information, vertical high-frequency information, and diagonal high-frequency information in the input features.
[0008] The final hidden state in step S5 h t is input into a linear classifier to obtain the classification result: , where σ(•) is the sigmoid function, W classifier is the weight of the classification layer, y (α,β) represents the classification result at position ( α , β ), and at the same time, L2 regularization is used on the weights of the classifier L to control the model complexity and prevent overfitting and improve the generalization ability of the model.
[0009] Among them L the specific steps of L2 regularization are as follows: In the training stage, the L2 norm (i.e., the sum of the squares of all weights) of the weight matrix W classifier of the classifier (i.e., the fully connected layer) L is added to the loss function to form a composite optimization objective: , where, L CE is the cross-entropy loss, λ is the regularization coefficient (the default value is set to 0.01) used to balance the model fitting ability and complexity.
[0010] Another object of the present invention is to provide a hyperspectral classification system based on wavelet spatial spectral Mamba, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the hyperspectral classification method based on wavelet spatial spectral Mamba as described above.
[0011] The beneficial effects of the present invention are: The present invention takes a hyperspectral image cube as the input, divides the image into overlapping 3D patches, then decomposes the spatial and spectral information of each patch simultaneously, performs multi-resolution analysis on the reconstructed features using the classical Haar wavelet, and captures long-range dependencies through the Mamba network of the state space model. The present invention effectively balances the relationship between the network model performance and the computational complexity, and at the same time has the ability of multi-scale feature extraction and global modeling, can significantly improve the classification accuracy while reducing the consumption of computing resources, and has excellent robustness and generalization ability.
[0012] By combining spatial-spectral wavelet feature extraction with the state-space Mamba network, the present invention realizes the efficient capture of spectral and spatial information in hyperspectral images. Using wavelet transform for multi-scale decomposition, it effectively extracts image details and suppresses noise. Combined with the linear state modeling of the Mamba network, it effectively captures local texture and global semantic information in the image, realizing fast and accurate hyperspectral data processing. Compared with traditional CNN or Transformer methods and other vision Mamba methods, considering the particularity of hyperspectral images, the present invention innovatively combines wavelet transform with the Mamba network. The two perform their respective functions and cooperate closely. The Haar wavelet transform decomposes the hyperspectral image into multi-scale sub-bands to extract multi-scale features. The Mamba network is based on the state-space model and models the long-range dependence of sequence data with linear complexity, overcoming the defect of the quadratic computational complexity of traditional Transformers. At the same time, the local detail extraction of wavelet transform and the global sequence modeling of Mamba complement each other. The former strengthens the fine-grained representation of space-spectra, and the latter integrates the dependencies across bands and pixels. Wavelet decomposition reduces the input dimension and alleviates the pressure of the Mamba to process high-dimensional data. The linear complexity of Mamba ensures the efficient fusion of multi-scale features. In this way, while maintaining a low computational cost, it has stronger generalization ability and application adaptability, and can be widely applied to multiple fields such as environmental monitoring, land use management, and resource investigation. Description of the Drawings
[0013] Figure 1 It is a network model architecture diagram of the method of the present invention. Detailed Embodiments
[0014] The present invention will be further described below in conjunction with the detailed embodiments.
[0015] Embodiment 1 A hyperspectral classification method based on wavelet spatial-spectral Mamba, including the following steps: S1. Preprocess the input hyperspectral image and divide the image into 3D small blocks with overlapping regions: Let the input be a M × N × B hyperspectral image cube X , divide X into overlapping 3D blocks of size P×P×B. Each block P (α,β) is extracted centered on specific spatial coordinates ( α , β ). Specifically, the width sampling range of each block is , and the height value range is , covering all data in this area at the same time B in the meantime. The class labels of these 3D blocks are determined by the label at the central pixel ( α , β ), and the final output is ([[]] M - P + 1) × (N - P + 1) such 3D blocks, providing input for subsequent wavelet transform and feature extraction. Through this block division method, spatial and spectral information in the hyperspectral image can be better captured.
[0016] S2. Perform spatial and spectral feature learning on each 3D small-block image respectively to generate spatial feature S and spectral feature F: Taking one of the 3D blocks as an example (the following operations will be performed on each 3D image block divided from the hyperspectral image in the present invention), in order to represent the spatial features of the hyperspectral image, depthwise separable convolution is used to preliminarily learn its spatial features: , where represents depthwise separable convolution with a convolution kernel size of 3×3, C represents the feature dimension after convolution. Thus, S represents the preliminarily extracted spatial features; For spectral features, the focus is on the correlation between each spectral band. Therefore, in the present invention, convolution with a convolution kernel of 1×1 is used for feature learning between spectral channels: , where represents channel convolution with a convolution kernel size of 1×1, C represents the feature dimension after convolution. Thus, F represents the preliminarily extracted spectral features.
[0017] S3. Perform Haar wavelet transform on both the preliminarily extracted spatial feature S and spectral feature F to generate multiple wavelet subbands, and splice the multiple wavelet subbands obtained by transforming S and F along the channels to obtain a new 3D representation, that is, the feature sequence: Perform the same Haar wavelet transform on the preliminarily extracted spatial feature S and spectral feature F . For unified description, let the input of the wavelet transform be .
[0018] The specific process of the Haar wavelet transform is to first apply a low-pass filter and a high-pass filter along the "row" direction to generate two subbands and ; Then, the same filter is applied to these two sub-bands along the "column" direction, and finally four wavelet sub-bands are obtained , , , .
[0019] These four wavelet sub-bands respectively capture the low-frequency approximation information, horizontal high-frequency information, vertical high-frequency information, and diagonal high-frequency information in the input features. Each wavelet sub-band can be regarded as a downsampling of the input features . This multi-scale decomposition can effectively extract different frequency components and spatial features of the features, providing a rich multi-scale representation for subsequent feature learning. The whole process realizes spatial downsampling while retaining all input information, which helps to reduce the computational complexity and highlight important features. That is, it can achieve multi-scale detail capture and noise reduction while retaining the key information of the image.
[0020] As Figure 1 shown, in the present invention, for spatial features S and spectral features F , a total of two groups of eight wavelet sub-bands are generated, namely the four wavelet sub-bands S of spatial features SLL , SLH , SHL and SHH , and the four wavelet sub-bands F of spectral features FLL , FLH , FHL and FHH , fully capturing and utilizing the spatial and spectral information of the hyperspectral image.
[0021] Subsequently, the present invention splices the eight wavelet sub-bands along the channels to obtain a new 3D representation .
[0022] This splicing operation retains the multi-scale feature information extracted by the wavelet transform, and at the same time organizes the spatial features and spectral features into a unified data structure, which is convenient for subsequent Mamba state space modeling. In this way, the present invention can simultaneously utilize the information of different spectral and spatial scales, thereby enhancing the richness of feature representation while maintaining the computational efficiency, and providing a comprehensive input containing multi-source and multi-scale information for subsequent state space modeling.
[0023] S4. Feed the feature sequence after wavelet transform into the Mamba network based on the state space model, and capture the spatio-temporal dependence relationship through the state transition and update mechanism with linear time complexity: The Mamba architecture is a novel sequence modeling method based on the State Space Model (SSM), aiming to address the quadratic computational complexity caused by the self-attention mechanism in traditional Transformer architectures. The problem and its calculation process are as follows: Mamba models the dynamic evolution process of sequence data through state space equations, taking the multi-scale feature sequence obtained by Haar wavelet transform X’ After linear transformation, it serves as the input to the Mamba network. Let the input be , and let the state vector h t represent the hidden state at time t . Its state transition and update formulas are as follows: , where W transition and W update are learnable weights of the network, ReLU(•) is the activation function, h t-1 is the hidden state at the previous moment, E t is the wavelet feature after the current linear transformation, representing the feature at a certain time step (or sequence position). In hyperspectral images, "time" can be understood as the feature expression along a certain sequence, used to capture the dependencies between space and spectrum.
[0024] Considering the massive space-spectral information in hyperspectral images, the Mamba network based on state space modeling is combined with Haar wavelet transform. It innovatively extracts multi-resolution features (low-frequency global information and high-frequency details) explicitly through wavelet decomposition, and uses the linear complexity of Mamba to model long-range dependencies. Compared with other vision Mambas that rely on multi-channel and residual structures for feature learning, it is more suitable for the complex space-spectral relationships of hyperspectral data, reducing the computational burden when processing high-dimensional data.
[0025] S5. Input the final hidden state h t learned by the network for the 3D patch into a linear classifier for classification, obtaining the classification result corresponding to the central pixel of the 3D patch. Perform the same calculation for each 3D patch, and assign the classification results of each 3D patch according to their positions in the image to obtain the classification result of the entire input image: Input the final hidden state h t into a linear classifier to obtain the classification result: , where σ(•) is the sigmoid function, W classifier is the weight of the classification layer, y (α,β) represents the classification result at position ( α , β ), and at the same time, L2 regularization (regularization coefficient L = 0.01) is used on the weights of the classifier to control the model complexity, prevent overfitting and improve the generalization ability of the model. λ = 0.01) to control the model complexity, prevent overfitting and improve the generalization ability of the model.
[0026] where L the specific implementation of L2 regularization is as follows: In the training stage, the L2 norm (i.e., the sum of the squares of all weights) of the weight matrix W classifier of the classifier (i.e., the fully connected layer) is added to the loss function to form a composite optimization objective: , where L CE is the cross-entropy loss, λ is the regularization coefficient (default value is set to 0.01), which is used to balance the model fitting ability and complexity. L2 regularization forces the classifier weights to tend to a smaller numerical distribution by penalizing large weight values, avoids being overly sensitive to noise or accidental patterns in the training data, and at the same time reduces the model degrees of freedom by constraining the weight space, preventing the complex model from falling into over-parameterization in the small-sample scenario, thereby improving the network robustness.
[0027] Embodiment 2 A hyperspectral classification system based on wavelet space spectrum Mamba, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the hyperspectral classification method based on wavelet space spectrum Mamba as described in Embodiment 1 above.
[0028] The above is a further description of the present invention in combination with the embodiments, and the protection scope of the present invention is not limited thereto.
Claims
1. A hyperspectral classification method based on wavelet spatial spectrum Mamba, characterized in that, The steps are as follows: S1. Preprocess the input hyperspectral image and divide the image into 3D patches with overlapping regions; S2. Perform spatial and spectral feature learning on each 3D patch image respectively to generate spatial feature S and spectral feature F; S3. Perform Haar wavelet transform on both the preliminarily extracted spatial feature S and spectral feature F to generate multiple wavelet subbands, and concatenate the multiple wavelet subbands obtained by transforming S and F along the channels to obtain a new 3D representation, i.e., the feature sequence; S4. Feed the wavelet-transformed feature sequence into the Mamba network based on the state space model, and capture spatio-temporal dependencies through the state transition and update mechanism with linear time complexity; The feature sequence is linearly transformed as the input of the Mamba network. Let the input be , and let the state vector h t represent the hidden state at time t . Its state transition and update formulas are as follows: , Among them, W transition and W update are network learnable weights, and ReLU(•) is the activation function, h t-1 is the hidden state at the previous moment, E t is the wavelet feature after the current linear transformation, representing the feature at a certain sequence position; S5. Input the final hidden state obtained by network learning of the 3D patch h t into a linear classifier for classification to obtain the classification result corresponding to the central pixel of the 3D patch. Perform the same calculation for each 3D patch, and assign the classification results of each 3D patch according to their positions in the image, then the classification result of the entire input image can be obtained.
2. The hyperspectral classification method based on wavelet spatial spectrum Mamba according to claim 1, wherein, In step S2, a depthwise separable convolutional network and a classical convolutional network are respectively used for spatial and spectral feature learning.
3. The hyperspectral classification method based on wavelet spatial spectrum Mamba according to claim 1, characterized in that, In step S3, the Haar wavelet transform uses low-pass and high-pass filters to decompose step by step along the row and column directions, generating multiple sub-bands. The specific process of the Haar wavelet transform is to first apply the low-pass filter along the "row" direction and the high-pass filter , generating two sub-bands and ; Then, the same filter is applied to these two sub-bands along the "column" direction, and finally four wavelet sub-bands are obtained. , , , , and these four wavelet sub-bands capture the low-frequency approximation information, horizontal high-frequency information, vertical high-frequency information, and diagonal high-frequency information in the input features respectively.
4. The hyperspectral classification method based on wavelet spatial spectrum Mamba according to claim 1, characterized in that the steps Final hidden state in S5 h t Input into a linear classifier to obtain a classification result: , where σ(•) is the sigmoid function, W classifier is the weight of the classification layer, y (α,β) represents the classification result at the position ([ α , β ), and at the same time, L2 regularization is used on the weights of the classifier to control the model complexity. L 5. The hyperspectral classification method based on wavelet spatial spectrum Mamba according to claim 4, characterized in that L The specific steps of 2-regularization are as follows: During the training phase, the weight matrix of the classifier W classifier of L the L2 norm is added to the loss function to form a composite optimization objective: , Among them, L CE is the cross-entropy loss, λ is the regularization coefficient, which is used to balance the model fitting ability and complexity.
6. A hyperspectral classification system based on wavelet spatial spectrum Mamba, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the hyperspectral classification method based on wavelet spatial-spectral Mamba according to any one of claims 1-5.
Citation Information
Cited By
Hyperspectral anomaly detection method based on dual-domain consistency reconstruction network
CN122244624A
Hyperspectral anomaly detection method based on dual-domain consistent reconstruction network
CN122244624B