Pyramid structure-based space-spectrum Mama hyperspectral image classification method
By adopting the spatial-spectral Mamba hyperspectral image classification method based on the pyramid structure, the problem of insufficient fusion of spatial and spectral information in hyperspectral images is solved, achieving high-precision classification and low computational complexity in complex backgrounds, thus breaking through the limitations of traditional methods.
Patent Information
- Application Number
- CN202511066784.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
Existing hyperspectral image classification methods have limitations in the fusion of spatial and spectral information, making it difficult to fully explore the interaction relationships between different scales. They are particularly inaccurate in complex backgrounds and fine-grained classification tasks, and have high computational complexity.
We employ a pyramid-structure-based spatial-spectral Mamba hyperspectral image classification method. This method extracts multi-scale spatial spectral features through the pyramid structure, combines them with the Mamba architecture for feature modeling and fusion, and uses a bidirectional guidance mechanism to achieve deep cross-modal interaction, thereby reducing computational complexity and improving classification accuracy.
It significantly improves classification accuracy on a very small training set, effectively integrates spatial and spectral information, and performs particularly well in fine-grained classification tasks with complex backgrounds, improving discrimination ability and computational efficiency.
Smart Images

Figure CN120976624A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computer vision and machine learning, and specifically relates to hyperspectral image classification, and proposes a spatial-spectral Mamba hyperspectral image classification method based on a pyramid structure, which can be used in environmental monitoring, military, civil and other fields. BACKGROUND
[0002] Hyperspectral imaging technology (HSI) can provide more details than traditional RGB imaging by capturing image information in different spectral bands. Hyperspectral images can effectively express the spectral characteristics of ground objects and are widely used in environmental monitoring, agriculture, mineral exploration, military reconnaissance and other fields. However, there is a large amount of redundant information and high-dimensional characteristics in hyperspectral images, which makes the classification task challenging.
[0003] Traditional hyperspectral image classification methods mostly rely on classic feature extraction techniques such as principal component analysis (PCA) and linear discriminant analysis (LDA). These methods lose some key information in the dimensionality reduction process, resulting in a decrease in classification accuracy. In recent years, deep learning methods, especially convolutional neural networks (CNN) and deep convolutional networks (DCN), have made significant progress in hyperspectral image classification. These methods can automatically extract high-dimensional features, but they still have limitations in the effectiveness of spatial and spectral information fusion. Currently, although some methods attempt to fuse spatial and spectral information, such as the method based on spatial-spectral convolutional network (SSCNN), these methods still cannot completely solve the complex interaction problem between the two. Traditional feature extraction methods often cannot fully consider the multi-scale and multi-directional features of spatial and spectral information, making it difficult to further improve the classification performance.
[0004] In the prior art document ASSMN [Wang D, Du B, et al. Adaptive Spectral-Spatial Multiscale Contextual Feature Extraction for Hyperspectral Image Classification [J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 59: 2461-2477], a multiscale adaptive spectral-spatial feature extraction model is designed, which uses a long short-term memory network (LSTM) to extract and fuse spectral and spatial information in hyperspectral images. LSTM is used to capture long-term dependencies in hyperspectral data and effectively model the dynamic interaction between spatial context information and spectral information. However, although ASSMN exhibits high performance in processing hyperspectral images, there are still the following challenges: first, the training and computational overhead of LSTM is large, especially when training on large-scale datasets, the computational efficiency may be limited; second, the ability of LSTM to model the interaction between spectral and spatial information is limited, and it is difficult to fully capture the deep-level connections between spatial and spectral features in complex scenarios. Currently, Mamba has a strong sequence modeling and low computational complexity advantage, and has achieved significant precision improvement effect in hyperspectral image classification tasks. For example, Zhang et al. [Zhang, Y., Liu, F., & Wang, H., "Hyperspectral Image Classification with Mamba," IEEE Transactions on Geoscience and Remote Sensing, vol. 63, no. 6, pp. 3590-3603, 2024.] applied the Mamba architecture to extract local and global spatial-spectral feature information and achieved excellent classification results. However, this method still has certain limitations in deep-level mining of information at different scales, especially in complex background and fine-grained classification tasks, it has not been able to fully mine the interaction between scales. SUMMARY
[0005] The present application aims at the deficiencies of the prior art, and proposes a spatial-spectral Mamba hyperspectral image classification method based on a pyramid structure, which is used to solve the problems of poor calculation accuracy in spatial and spectral information fusion, scale deep mining, and limited samples in the prior art. First, multi-scale spatial-spectral features are extracted through the pyramid structure, fully considering the information changes in different scales and directions in the hyperspectral image. Then, the Mamba architecture is used for efficient feature modeling and fusion, thereby strengthening the interaction between spatial-spectral features and solving the problem of insufficient spatial and spectral information fusion in traditional methods. At the same time, high classification accuracy is achieved on a very small sample training set, solving the difficult problem of sample labeling. The present application can achieve breakthroughs in multi-scale feature extraction and spatial-spectral interactive modeling, effectively fusing spatial and spectral information, and significantly improving the classification accuracy on a very small training set, especially in fine-grained classification tasks under complex backgrounds.
[0006] To achieve the above-mentioned purposes, the technical scheme of the present application includes the following:
[0007] (1) Obtain a hyperspectral image sample set from a plurality of public data sets, and randomly divide it into a training set and a test set; the public data set covers a plurality of ground object types and complex scenes, including Indian Pines, Salinas, and Pavia University;
[0008] (2) Represent the hyperspectral image as where H, W and C represent the height, width and band number of the image respectively; each pixel vector is regarded as an independent sample, and the training set is represented as Q train , the test set is Q test , and a hyperspectral image block is generated around each center pixel;
[0009] (3) Construct a pyramid-based spectral Mamba and spatial Mamba module, denoted as PspeM and PspaM modules, respectively, wherein the PspeM module is used to extract discriminative features from different scale spectral subspaces, and the PspaM module is used to capture spatial sequence features across multiple receptive fields;
[0010] (4) A bidirectional guiding mechanism, i.e., spectral-aware spatial weighting and spatial-aware spectral weighting, is used to construct a spatial-spectral information fusion module SSIF, which is used to realize deep cross-modal interaction and integration;
[0011] (5) Construct a pyramid-based spatial-spectral Mamba model PS 2 Mamba:
[0012] The image block is input into the parallel double-branch structure composed of the PspeM and PspaM modules, is subjected to spectral and spatial coding in the double-branch structure, and spectral and spatial features are acquired; the two kinds of features are fused through a spatial-spectral information fusion module SSIF to obtain fused features; and the fused features are sent into a linear classifier to generate a final classification prediction result;
[0013] (6) The PS 2 Mamba model is trained by using the training set data, adopting cross entropy as a loss function, and through back propagation until the loss function of the model converges or a preset stop condition is met, so that a trained model is obtained;
[0014] (7) The trained PS 2 Mamba model is used to classify the test set to obtain a final classification result.
[0015] Compared with the prior art, the present application has the following advantages:
[0016] Firstly, the bidirectional-multidirectional Mamba scanning mechanism is introduced to replace the traditional self-attention mechanism of the Transformer, so that the computational complexity is controlled at a low level, thereby avoiding the computational bottleneck in processing spatial-spectral information in a hyperspectral image, and the running efficiency of the model is significantly improved; the mechanism has linear complexity and can efficiently process high-dimensional data, avoiding the exponential growth of the computational burden as the length of the input sequence increases.
[0017] Secondly, in the present application, the PSpeM module and the PspaM module are designed, wherein the PSpeM module is used to capture spectral dependence under different receptive fields; the module first enhances the target pixel representation by using neighborhood spectral information, and then extracts deep spectral features of different scales through bidirectional Mamba scanning, so as to more accurately model the complex relationship between spectra; the PspaM module is used to extract multi-scale spatial sequence features, which adopts an eight-direction scanning strategy at each scale to effectively model long-range spatial dependence relationships, and integrates the spatial features of each scale through adaptive fusion to obtain rich spatial context information; the multi-scale extraction strategy is adopted, and the features are fused through an adaptive weighting mechanism, so that the model can effectively capture multi-level semantic information in the hyperspectral image, especially in the classification task of fine-grained objects; this multi-scale feature extraction capability can better process details and complex backgrounds in the image, and improve the discrimination ability.
[0018] Thirdly, the present application can systematically consider the synergistic effect of spatial and spectral dimensions by performing multiscale modeling of spectral and spatial features through the PSpeM and PSpaM modules respectively and realizing interactive fusion of spatial and spectral information through the SSIF module, thereby overcoming the bias problem of existing methods when modeling spatial or spectral dimensions; this synergistic modeling strategy makes the utilization of spatial-spectral information more sufficient, thereby significantly improving the classification accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 Figure 1 is a schematic diagram of the implementation process of the method of the present application; wherein a is the PSpeM module of the present application, b is the PSpaM module of the present application, c is the SSIF spatial-spectral feature cross-fusion module of the present application, d is the Mamba overall architecture of the present application, e is the spectral Mamba component, and f is the spatial Mamba component. 2 Figure 1 is a schematic diagram of the implementation process of the method of the present application; wherein a is the PSpeM module of the present application, b is the PSpaM module of the present application, c is the SSIF spatial-spectral feature cross-fusion module of the present application, d is the Mamba overall architecture of the present application, e is the spectral Mamba component, and f is the spatial Mamba component.
[0020] Figure 2 Figure 2 is a diagram of classification results on different hyperspectral data sets using the method of the present application and existing methods; wherein (a) shows the classification results on the Indian Pines data set; (b) shows the classification results on the Salinas data set; and (c) shows the classification results on the Pavia University data set. DETAILED DESCRIPTION
[0021] In order to make the objectives and advantages of the present application more clear and apparent, the technical content of the present application is described in detail below with reference to the drawings.
[0022] Embodiment 1: Referring to Figure 1 The present application proposes a spatial-spectral Mamba hyperspectral image classification method based on a pyramid structure, which specifically comprises the following steps:
[0023] Step 1) Obtain a hyperspectral image from a plurality of public data sets to form a sample set, and randomly divide it into a training set and a test set; the public data sets cover a plurality of ground object types and complex scenes, including Indian Pines, Salinas and Pavia University; these data sets are widely used in hyperspectral image classification research and are representative; in order to ensure the comprehensiveness of the experiment, the samples in the data set include different spectral bands and have high resolution, and are suitable for verifying the effectiveness of the method of the present application under different environments and conditions. In this embodiment, the data set is randomly divided according to a certain proportion (10 pixel points in each class are selected as the training set, and the rest are the test set), to ensure that the divided data set can represent the entire data distribution. During the division of the training set and the test set, the number of samples in each class is balanced to avoid the influence of class imbalance on the classification performance.
[0024] Step 2) Representing hyperspectral image as where H, W and C represent the height, width and number of bands of the image respectively; where each pixel vector is regarded as an independent sample, the training set is represented as Q train , the test set is Q test , and a hyperspectral image block is generated around each center pixel;
[0025] Step 3) Constructing pyramid-based spectral Mamba and spatial Mamba modules, denoted as PspeM and PspaM modules respectively, where the PspeM module is used to extract discriminative features from different scale spectral subspaces, and the PspaM module is used to capture spatial sequential features across multiple receptive fields. The PspeM module includes a spectral enhancement component, a spectral pyramid structure, multiple spectral Mamba feature extraction components, and a spectral adaptive fusion component; where the spectral pyramid structure is realized by layer-by-layer downsampling operation; the spectral Mamba feature extraction component is based on a forward and backward bidirectional scanning mechanism, and is used to extract multi-scale deep spectral feature information; the module takes the hyperspectral image block as the input of the spectral enhancement component, generates the spectral enhancement vector of the center pixel using the spatial context information around the center pixel, and then dynamically adjusts the spectral response of the center pixel using the local spatial context; then, deep spectral dimension features are extracted through the multi-scale downsampling operation of the spectral pyramid structure and the forward and backward bidirectional scanning of the spectral Mamba feature extraction component, and finally the final spectral features are obtained through the adaptive feature fusion component; the specific implementation is as follows:
[0026] (3a1) Let the input image block be where H1x W1 represents the size of the spatial neighborhood, and C represents the number of spectral bands; let the center pixel position of the image block be the adjacent spatial (i,j) pixel position be
[0027] (3a2) Calculate the Euclidean distance, and convert the distance into a similarity score through the sigmoid activation function σ; reweight each pixel in the image block according to its similarity with the center pixel to obtain the weighted image block P i,j , which is represented as follows:
[0028] P i,j '=σ(-||P i,j -P C ||2)×P i,j ;
[0029] (3a3) Process the weighted image block P i,jneighborhood enhanced spectral vector
[0030]
[0031] (3a4) the neighborhood enhanced spectral vector Q Spe is subjected to principal component analysis (PCA) to obtain a reduced dimension feature matrix
[0032] (3a5) a spectral pyramid structure is designed, and the reduced dimension feature matrix is subjected to a convolution-based downsampling operation at each layer to halve the spectral dimension step by step:
[0033] Z i = Conv 1×1 (Z i-1 ),
[0034] wherein represents the spectral feature of the i-th layer of the pyramid, i = 1, 2, …, L, and L represents the number of layers of the pyramid; each layer of the spectral pyramid is subjected to feature transformation by a spectral Mamba module; the features in the pyramid are in the form of one-dimensional sequences
[0035] (3a6) deep multi-scale features are extracted layer by layer by using a spectral Mamba component based on a forward and backward bidirectional scanning mechanism:
[0036]
[0037] wherein and C Spe represent a state transition matrix, an input matrix, and an observation matrix, respectively; the state transition probability matrix represents the state transition relationship from time t-1 to time t, the input matrix is used to describe how to convert the input x i into the influence of the state, and the observation matrix is used to convert the hidden state h t into the output y t ; a group of output sequences
[0038]
[0039] (3a7) the multi-scale extracted spectral features are aggregated by a spectral adaptive fusion component to obtain the final spectral feature F Spe :
[0040]
[0041] wherein β i represents a learnable parameter used to adaptively control the contribution degree of each scale feature.
[0042] The PSpaM module includes a feature encoding component, a spatial pyramid structure, a plurality of spatial Mamba feature extraction components, and a spatial adaptive fusion component; wherein the spatial pyramid structure is realized by layer-by-layer downsampling operation; the spatial Mamba feature extraction component is realized based on the SS2D multi-directional scanning mechanism, and is used for extracting multi-scale deep spatial feature information; the specific implementation is as follows:
[0043] (3b1) assuming that the input image block is wherein H1xW1 represents the size of the spatial neighborhood, and C represents the number of spectral bands; the input image block is projected into an embedding space; after patch encoding of the original hyperspectral image block, an output tensor Y is obtained The multi-scale structure is realized by downsampling operation of a plurality of 3x3 convolution kernels, and the process is defined as follows:
[0044] X i =Conv 3×3 (X i-1 ),i=1,2,...,L
[0045] wherein i=1, 2...L, and L represents the total number of layers of the pyramid;
[0046] (3b2) the spatial Mamba component based on the SS2D multi-directional scanning mechanism is adopted to extract deep spatial feature information under the multi-scale structure;
[0047] X1=σ(Linear(X i ))
[0048] X2=LN(SS2D(σ(Conv(Linear(X i ))) )
[0049] Y i =Linear(X1⊙X2)
[0050] wherein X1, λ represents the inflation coefficient; and is the Hadamard product, and σ(·) is the SILU activation function;
[0051] (3b3) the selected spatial features of different scales are standardized to a unified representation by using the global average pooling GAP, so as to prepare for the subsequent fusion of multi-scale feature information;
[0052] (3b4) the multi-scale spatial features are optimized by the spatial adaptive fusion component, and the final aggregated spatial feature F Spe is obtained:
[0053]
[0054] where α i is a learnable parameter.
[0055] Step 4) A bidirectional guidance mechanism, i.e., spatial-aware spectral weighting and spectral-aware spatial weighting, is adopted to construct a spatial-spectral information fusion module SSIF for realizing deep cross-modal interaction and integration. In this embodiment, the spatial-spectral information fusion module SSIF simultaneously adopts the spatial-aware spectral weighting and the spectral-aware spatial weighting mechanism to realize mutual guidance and optimization, and is expressed as follows:
[0056] δ1=F spe ×Softmax(F spa )
[0057] δ2=F spa ×Softmax(F spe )
[0058] F joint =Concat(δ1,δ2)
[0059] where δ1 represents the spatial-aware weight applied to the spectral feature, δ2 represents the spectral-aware weight applied to the spatial feature, and Concat(·) represents mapping and connecting two features along the channel dimension to form a unified joint representation.
[0060] Step 5) A spatial-spectral Mamba model PS 2 Mamba based on a pyramid structure is constructed.
[0061] The PspeM and PspaM modules are used to form a parallel double-branch structure, the image blocks are input into the double-branch structure for spectral and spatial encoding to obtain spectral and spatial features, the spatial-spectral information fusion module SSIF is used to fuse the two kinds of features to obtain fused features, and the fused features are sent into a linear classifier to generate a final classification prediction result.
[0062] Step 6) The training set data is used, cross-entropy is used as a loss function, and the PS 2 Mamba model is trained through back propagation until the loss function of the model converges or a preset stopping condition is met, and a trained model is obtained. In this embodiment, the loss function of the model is expressed as follows:
[0063]
[0064] where B represents the batch size, x b is the input data of the b-th batch, y b is the corresponding label, and f(x b) is the output of the model prediction, and H(·) represents the cross-entropy loss function.
[0065] The model training process is as follows:
[0066] (6.1) The training set data is input into the model in batches, and each batch of input data includes spatial spectral features of multiple hyperspectral images, and multi-layer feature maps with different scales; the label set corresponding to each input image contains the land cover class to which each pixel point belongs;
[0067] (6.2) The model extracts and fuses the input spatial spectral information layer by layer through each network layer, learns the spatial spectral features at different scales, and maps them to the high-dimensional feature space of the current layer;
[0068] (6.3) The cross-entropy loss function is used as the training target of the model, and for each pixel point, the cross-entropy loss function is used to calculate the difference between the predicted class and the true class of the point;
[0069] (6.4) After each forward propagation, the error calculated based on the loss function is used to calculate the gradient through the back propagation algorithm; during the back propagation process, the gradient of each layer weight is calculated through the chain rule, and the parameters of the model are updated using the gradient descent algorithm;
[0070] (6.5) During the training process, an L2 regularization term is added to constrain the parameters, and a Dropout technique is used to randomly discard a portion of neurons during each training, which is used to prevent overfitting;
[0071] (6.6) The small batch gradient descent method is used to iteratively optimize the model and update the weights of the model; until the loss function of the model converges or the preset stopping condition is met; the preset stopping condition described in this embodiment is to reach the maximum number of training rounds or the accuracy of the validation set reaches the preset level.
[0072] Step 7) The trained PS 2 The Mamba model classifies the test set to obtain the final classification result.
[0073] Embodiment Two: The overall implementation steps of the classification method proposed in this embodiment are the same as those of Embodiment One, and the parameter settings are given, and the implementation process of the present application is further described in detail with specific examples:
[0074] Step One: Overall flow of the model
[0075] This embodiment takes the Indian Pines data set as an example; the hyperspectral image is represented as Where the height, width and band number of the image are 145, 145 and 200 respectively. Each pixel vector The whole dataset is divided into training set Q train and test set Q test , and 10 pixels are randomly selected from each class as the training set, and the rest of the pixels are the test set. Around each center pixel, an image block is generated, which is then input into two parallel branches, a pyramid-based spectral Mamba module (PSpeM) for extracting discriminative features from spectral subspaces at different scales to enhance the modeling of spectral variation trends, and a pyramid-based spatial Mamba module (PSpaM) for capturing spatial sequential features across multiple receptive fields to improve the representation of features at different spatial scales, for spectral and spatial encoding. In addition, in order to effectively integrate spatial and spectral features, a spatial-spectral interaction fusion module (SSIF) is introduced, which adopts a bidirectional guiding mechanism, i.e. spectral-aware spatial weighting and spatial-aware spectral weighting, for deep cross-modal interaction and integration. This mechanism enhances the complementarity of features while suppressing redundant information. Finally, the fused features are sent to a simple linear classifier to generate the final classification prediction. During the training process, the cross-entropy is used as the loss function, and the model parameters are updated by calculating the loss between the training samples and their labels, as shown in the following formula:
[0076]
[0077] where H(q, p) represents the cross-entropy between distributions q and p, f(·) represents the proposed model parameters, and 64 represents the batch size.
[0078] Step two: pyramid-based spectral Mamba module PSpeM:
[0079] Referring to part b in Figure 1 , the PSpeM module receives a hyperspectral image block as input, which is first sent to a spectral enhancement component that generates a spectral enhancement vector for the center pixel using the spatial context information around the center pixel. This component dynamically adjusts the spectral response of the center pixel using local spatial context, effectively alleviating the limitations of directly performing global average pooling on the input image block. Specifically, let the input image block be represented as P∈R 28 ×28×200 , where 28x28 represents the size of the spatial neighborhood and 200 represents the number of spectral bands. Let the center pixel position of the image block be and the adjacent spatial (i,j) pixel position be First, the Euclidean distance is calculated to measure spectral similarity. This distance is then converted into a similarity score, a conversion achieved using the sigmoid activation function σ. Each pixel in the image patch is reweighted based on its similarity to the center pixel. This process can be expressed by the following formula:
[0080] P i,j '=σ(-||P i,j -P C ||2)×P i,j
[0081] The enhanced image patch is processed by global average pooling (GAP) to obtain the spectral vector:
[0082]
[0083] in This is the spectral vector enhanced by neighborhood. Furthermore, to reduce computational complexity and eliminate redundant information between spectral dimensions, this invention modifies the feature matrix Q... Spe Principal component analysis (PCA) was performed to obtain the dimensionality-reduced feature matrix. To obtain multi-scale spectral characterization, this paper proposes a spectral pyramid structure in which the spectral dimension is halved at each layer through a convolution-based downsampling operation.
[0084] Z i =Conv 1×1 (Z i-1 ), i = 1, 2... L
[0085] in This represents the spectral characteristics of the i-th layer of the pyramid. Each level of the spectral pyramid is accessed via the spectral Mamba module ( Figure 1 Feature transformations are performed (as shown in e) to enhance the correlation between spectral dimensions and extract long-range dependencies. All features in the pyramid are in one-dimensional sequence form. Sequence features are extracted using a bidirectional scanning mechanism:
[0086]
[0087] in and C Spe These represent the trainable parameters. A set of output sequences can be obtained:
[0088]
[0089] Finally, the spectral characteristics F after aggregation Spe Further optimization is achieved through an adaptive fusion module to improve the quality of feature representation.
[0090]
[0091] where β i is a learnable parameter to adaptively control the contribution of each scale feature.
[0092] Step three: Pyramid-based spatial Mamba module (PSpaM)
[0093] The detailed structure of the PSpaM module is shown in Figure 1 Part c. Specifically, the input of this module is the segmented image patch, which is projected into an embedding space. The process can be represented as:
[0094]
[0095] After Patch encoding of the original hyperspectral image patch, the output tensor is denoted as This feature tensor is then passed through a series of down-sampling operations using 3x3 convolution kernels, which is defined as:
[0096] X i = Conv 3×3 (X i-1 ), i = 1, 2,..., L
[0097] where i represents the number of layers of the pyramid. All down-sampled feature maps will be input into the spatial Mamba module for processing. To bridge the gap between the sequential nature of one-dimensional selective scanning and the non-sequential structure of two-dimensional visual data, the present application adopts the SS2D scanning mechanism. As shown in Figure 1 f, compared with the traditional four-path scanning method (which starts from positions 1 and 9 and scans in two directions respectively), the present application further increases four scanning paths starting from positions 3 and 7. This scanning enhancement design enables the model to converge context information from more directions and perspectives, thereby improving the integrity and expressiveness of spatial context modeling.
[0098] X1 = σ(Linear(X i ))
[0099] X2 = LN(SS2D(σ(Conv(Linear(X i ))))
[0100] Y i = Linear(X1 ⊙ X2)
[0101] where X1, λ represents the inflation coefficient. ⊙ is the Hadamard product, and σ(·) is the SILU activation function.
[0102] To achieve efficient batch parallel processing, global average pooling (GAP) is used to aggregate selected spatial features from different scales into a unified representation. Subsequently, an adaptive feature fusion module is introduced to optimize the aggregated spatial features to enhance their discriminative and expressive capabilities:
[0103]
[0104] where alpha i is a learnable parameter that adaptively controls the contribution of each scale. It ensures a balance between fine-grained local details and global contextual information.
[0105] Step four: spatial-spectral interactive fusion module (SSIF)
[0106] To fully exploit the complementarity between spectral and spatial representations in hyperspectral images, as shown in part d of Figure 1 , the present invention designs an SSIF module. Specifically, the module simultaneously uses "spatial-aware spectral weighting" and "spectral-aware spatial weighting" mechanisms to achieve mutual guidance and optimization, effectively suppressing redundant features and strengthening the expression of key features. The process can be expressed as follows:
[0107] delta 1 = F spe * Softmax (F spa )
[0108] delta 2 = F spa * Softmax (F spe )
[0109] F joint = Concat (delta 1, delta 2)
[0110] delta 1 represents spatial-aware weighting applied to spectral features, and delta 2 represents spectral-aware weighting applied to spatial features. Concat(·) concatenates the two features along the channel dimension to form a unified joint representation.
[0111] Example three: The overall implementation steps of the hyperspectral image classification method proposed in this embodiment are the same as those in example one. Now the training process of the model in the present invention will be described in further detail:
[0112] a. Data input:
[0113] After the training data is prepared through the preprocessing step, it is input into the PS 2 Mamba model in batches. Each batch of input data includes spatial-spectral features of multiple hyperspectral images with multiple layers of features of different scales. The label set corresponding to each input image contains the land cover class to which each pixel belongs.
[0114] b. Forward propagation:
[0115] PS 2 The Mamba model performs forward propagation. Each layer of the network performs convolution operations on the input features, combining the multi-scale convolution characteristics of the pyramid structure to extract and fuse spatial-spectral information layer by layer. In this process, the model automatically learns spatial-spectral features at different scales and maps these features to a higher-dimensional feature space.
[0116] c. Loss function:
[0117] The present application adopts the Cross-Entropy Loss as the training target of the model, which is used to measure the error between the predicted result and the actual label. Cross-Entropy Loss performs well in multi-classification tasks and can effectively optimize classification accuracy. For each pixel, Cross-Entropy Loss calculates the difference between the predicted class and the true class.
[0118] d. Backpropagation and gradient update:
[0119] After each forward propagation, based on the error calculated by the loss function, the gradient is calculated through the Backpropagation algorithm. During the backpropagation process, the gradient of each layer weight is calculated through the chain rule, and the parameters of the model are updated using the gradient descent algorithm. In order to accelerate convergence and prevent overfitting, the present application adopts the Adam optimization algorithm, which combines momentum and adaptive learning rate adjustment, and can efficiently update network weights.
[0120] e. Regularization and prevention of overfitting:
[0121] In order to improve the generalization ability of the model and prevent overfitting, the present application adds an L2 regularization term to constrain the network parameters during training. In addition, the Dropout technique is used to randomly discard a portion of neurons during each training to prevent the model from overfitting to the training data.
[0122] f. Training parameters and iterations:
[0123] During training, the Mini-Batch Gradient Descent method is used to iteratively optimize the model. Each iteration uses a small batch of data from the training set to train and update the model's weights. The training process continues for several rounds (Epochs) until the model's loss function converges to a small value, or the preset stopping condition is met (such as the maximum number of training rounds or the validation set accuracy reaches a certain level).
[0124] g. Model verification and parameter adjustment:
[0125] Through the trained PS 2 The Mamba model classifies the test set and evaluates according to the classification result. The evaluation indexes include classification accuracy (Overall Accuracy, OA), average accuracy (Average Accuracy, AA), etc. Further, the confusion matrix and F1-score are used for fine-grained analysis to verify the classification performance of the method on different data sets. At the same time, compared with the existing classical method, the advantages of the method in the aspects of spatial-spectral information fusion and precision improvement are proved.
[0126] The effect of the present application will be further described below in combination with experiments.
[0127] 1. Experimental conditions:
[0128] The experiment of the present application is carried out under the hardware environment of CPU main frequency 3.00GHz, memory 48GB, and the software environment of Windows10 operating system and Python 3.7.
[0129] 2. Experimental content:
[0130] The present application is qualitatively and quantitatively compared with a variety of popular algorithms on three published hyperspectral anomaly detection datasets. The published hyperspectral anomaly detection datasets used in the experiment include Indian Pines, Salinas and Pavia University datasets; the ten mainstream algorithms compared are respectively: RF: [Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5-32], SVM: [Classification of hyperspectral remote sens. images with support vector machines], 2D-CNN: [H. Zhang, Y. Li, Y. Zhang, and Q. Shen, “Spectral-spatial classification of hyperspectral imagery using a dual-channel convolutional neural network,” Remote sensing letters, vol. 8, no. 5, pp. 438-447, 2017.], 3D-CNN: [H. Zhang, Y. Li, Y. Jiang, P. Wang, Q. Shen, and C. Shen, “Hyperspectral classification based on lightweight 3-d-cnn with transfer learning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 57, no. 8, pp. 5813-5828, 2019.], A 2 S 2KResNet: S. K. Roy, S. Manna, T. Song, and L. Bruzzone, “Attention-based adaptive spectral-spatial kernel resnet for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 9, pp. 7831-7843, 2021.], SSTN: [Z. Zhong, Y. Li, L. Ma, J. Li, and W.-S. Zheng, “Spectral-spatial transformer network for hyperspectral image classification: A factorized architecture search framework,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-15, 2022.], SSFTT: [L. Sun, G. Zhao, Y. Zheng, and Z. Wu, “Spectral-spatial feature tokenization transformer for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-14, 2022], SS-MTR: [L. Huang, Y. Chen, and X. He, “Spectral-spatial masked transformer with supervised and contrastive learning for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1-18, 2023.], HyperMamba: [Q. Liu, J. Yue, Y. Fang, S. Xia, and L.Fang, “Hypermamba: Aspectral-spatial adaptive mamba for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1-14, 2024.], IGroupSS-Mamba: [Y. He, B. Tu, P. Jiang, B. Liu, J. Li, and A. Plaza, “Igroupss mamba: Interval groups spatial-spectral mamba for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1-17, 2024.].
[0131] Table 1: Quantitative comparison of the present application and existing methods
[0132]
[0133] The present application first conducts a qualitative comparison with ten popular comparison algorithms on three public hyperspectral anomaly detection datasets; from Figure 2 It can be seen that the detection results of the present application are better than those of existing algorithms, wherein:
[0134] Figure 2 (a) in the above table: shows the classification results on the Indian Pines dataset. This dataset contains hyperspectral images of different ground object categories, with complex spectral characteristics and background information; the classification results in the figure show that, compared with existing methods, the present application can better identify fine-grained ground objects, and the classification accuracy is significantly improved, especially in complex backgrounds.
[0135] Figure 2 (b) in the above table: shows the classification results on the Salinas dataset. This dataset is mainly used for ground object classification in the field of agriculture, and contains rich crop and natural ground object categories. The results in the figure show that the present application can better capture the subtle differences in spectral features and improve the expression ability of spatial features. Compared with traditional methods, the present application can more accurately distinguish between different ground objects, especially in complex scenes, and more accurately identify crop categories.
[0136] Figure 2(c) in the figure: shows the classification results on Pavia University dataset. This dataset is commonly used for hyperspectral image classification in urban environment, with high spatial resolution and rich ground object categories. The results in the figure show that the classification accuracy between different ground object categories is significantly improved by using the method of the application, especially in those ground object categories that are easily confused by traditional methods (such as between buildings and roads), the method of the application shows stronger distinguishing ability.
[0137] At the same time, the application is compared with ten popular contrast algorithms on three datasets, as shown in Table 1, and the application has the optimal classification accuracy.
[0138] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant national and regional laws, regulations and standards, and provide corresponding operation entrances for users to choose authorization or refusal.
[0139] The above simulation analysis proves the correctness and effectiveness of the method of the application.
[0140] The part of the application not described in detail belongs to the common knowledge of those skilled in the art.
[0141] The above only describes the preferred embodiments of the application and does not limit the application. Obviously, for those skilled in the art, after understanding the content and principles of the application, various modifications and changes in form and details can be made without departing from the principles and structures of the application, but these modifications and changes based on the idea of the application are still within the protection scope of the claims of the application.
Claims
1. A pyramid-based spatial-spectral Mamba hyperspectral image classification method, characterized in that, Comprising the following steps: (1) Obtain a hyperspectral image sample set from a plurality of public data sets, and randomly divide it into a training set and a test set; the public data set covers a plurality of ground object types and complex scenes, including IndianPines, Salinas and PaviaUniversity; (2) Representing the hyperspectral image as where H, W and C represent the height, width and number of bands of the image respectively; where each pixel vector is considered as an independent sample, the training set is represented as Q train , the test set as Q test , and a hyperspectral image block is generated around each center pixel; (3) Construct a pyramid-based spectral Mamba and spatial Mamba module, denoted as PspeM and PspaM module, respectively, wherein the PspeM module is used to extract discriminative features from different scale spectral subspaces, and the PspaM module is used to capture spatial sequence features across multiple receptive fields; (4) A bidirectional guiding mechanism, i.e., spectral-aware spatial weighting and spatial-aware spectral weighting, is used to construct a spatial-spectral information fusion module SSIF for realizing deep cross-modal interaction and integration; (5) Constructing a spatial-spectral Mamba model based on a pyramid structure PS 2 Mamba: The parallel double-branch structure is composed of the PspeM and PspaM modules, and the image blocks are input into the double-branch structure for spectral and spatial coding to obtain spectral and spatial features; The two kinds of features are fused by the spatial-spectral information fusion module SSIF to obtain fused features; And send it to a linear classifier to generate the final classification prediction result; (6) Using the training set data, adopting cross-entropy as the loss function, and training the PS 2 Mamba model until the loss function of the model converges or the preset stopping condition is met, to obtain a trained model. (7) by the trained PS 2 The Mamba model classifies the test set to obtain the final classification result.
2. The method of claim 1, wherein: The PspeM module of step (3) comprises a spectral enhancement component, a spectral pyramid structure, a plurality of spectral Mamba feature extraction components, and a spectral adaptive fusion component; The spectral pyramid structure is realized by layer-by-layer downsampling operation; the spectral Mamba feature extraction component is based on a forward and backward bidirectional scanning mechanism and is used to extract multi-scale deep spectral feature information; the module takes a hyperspectral image block as the input of the spectral enhancement component, generates a spectral enhancement vector for the center pixel using the spatial context information around the center pixel, and then dynamically adjusts the spectral response of the center pixel using the local spatial context; then, deep spectral dimension features are extracted through the multi-scale downsampling operation of the spectral pyramid structure and the forward and backward bidirectional scanning of the spectral Mamba feature extraction component, and finally the final spectral features are obtained through the adaptive feature fusion component; the specific implementation is as follows: (3al) Let the input image block be where H1xW1denotes the size of the spatial neighborhood, and C denotes the number of spectral bands; let the center pixel position of the image block be Let the neighboring spatial (i,j) pixel positions be (3a2) Calculate the Euclidean distance and convert the distance into a similarity score by a sigmoid activation function σ, and reweight each pixel in the image patch P according to its similarity to the center pixel i,j is represented as follows: (3a3) processing the weighted image patches P by global average pooling, GAP i,j , obtaining neighborhood enhanced spectral vectors (3a4) the spectrum vector Q after neighborhood enhancement Spe perform principal component analysis (PCA) to obtain a feature matrix after dimension reduction (3a5) The spectral pyramid structure is designed, and the feature matrix after dimension reduction is subjected to convolution-based downsampling operation at each layer to halve the spectral dimension step by step: Z i = Conv 1×1 (Z i-1 ), wherein, represents the spectral feature of the i-th layer of the pyramid, i = 1, 2, …, L, L represents the number of layers of the pyramid; each layer of the spectral pyramid is subjected to feature transformation by a spectral Mamba module; the features in the pyramid are in the form of one-dimensional sequences (3a6) Deep multi-scale features are extracted layer by layer by using the spectral Mamba component based on the forward and backward bidirectional scanning mechanism: wherein and C Spe respectively represent a state transition matrix, an input matrix and an observation matrix; the state transition probability matrix represents the state transition relationship from time t-1 to time t, the input matrix is used to describe how to convert the input x i into the influence of the state, and the observation matrix is used to convert the hidden state h t into the output y t ; and a set of output sequences (3a7) aggregating the multi-scale extracted spectral features by a spectral self-adaptive fusion component to obtain final spectral features F Spe : where β i denotes a learnable parameter for adaptively controlling the contribution degree of each scale feature.
3. The method of claim 1, wherein: The PSpaM module of step (3) comprises a feature encoding component, a spatial pyramid structure, a plurality of spatial Mamba feature extraction components, and a spatial adaptive fusion component; The spatial pyramid structure is realized by layer-by-layer downsampling operation; the spatial Mamba feature extraction component is realized based on the SS2D multi-directional scanning mechanism and is used to extract multi-scale deep spatial feature information; the specific implementation is as follows: (3b1) Let the input image block be where H1xW1denotes the size of the spatial neighborhood, C denotes the number of spectral bands; project it into an embedding space; After patch encoding of the original hyperspectral image block, the output tensor is obtained The multi-scale structure is realized by downsampling operation of multiple 3x3 convolution kernels, and the process is defined as: X i = Conv 3×3 (X i-1 ),i = 1,2,...,L Wherein, i=1, 2...L, L represents the total number of pyramid layers; (3b2) The spatial Mamba component based on the SS2D multi-directional scanning mechanism is used to extract deep spatial feature information under multi-scale structure; X1 = σ(Linear(X i )) X2 = LN(SS2D(σ(Conv(Linear(X i )))))) Y i = Linear(X1 ⊙ X2) wherein X1, λ denotes the expansion coefficient; is the Hadamard product, and σ(·) is the SILU activation function. (3b3) The selected spatial features of different scales are normalized to a unified representation by using global average pooling (GAP) to prepare for the subsequent fusion of multi-scale feature information. (3b4) optimizing the multi-scale spatial features by a spatial adaptive fusion component to obtain final aggregated spatial features F Spe : where a i is a learnable parameter.
4. The method of claim 1, wherein: The spatial-spectral information fusion module SSIF in step (4) simultaneously uses spatial-aware spectral weighting and spectral-aware spatial weighting mechanisms to achieve mutual guidance and optimization, as shown below: δ1 = F spe x Softmax(F spa ) δ2 = F spa x Softmax(F spe ) F joint = Concat(δ1,δ2) where δ1 represents the spatial-aware weight applied to the spectral feature, δ2 represents the spectral-aware weight applied to the spatial feature; Concat(·) represents mapping and connecting two features along the channel dimension to form a unified joint representation.
5. The method of claim 1, wherein: The PS described in step (6) is subjected to 2 The Mamba model is trained, the implementation process is as follows: (6.1) The training set data is input into the model in batches. Each batch of input data includes the spatial-spectral features of multiple hyperspectral images, with multi-layer feature maps of different scales. The label set corresponding to each input image contains the ground object category to which each pixel point belongs. (6.2) The model extracts and fuses the input spatial-spectral information layer by layer through each network, learns the spatial-spectral features at different scales, and maps them to the high-dimensional feature space of the current layer. (6.3) The cross-entropy loss function is used as the training target of the model. For each pixel point, the cross-entropy loss function is used to calculate the difference between the predicted class and the true class. (6.4) After each forward propagation, the error calculated based on the loss function is used to calculate the gradient through the backpropagation algorithm. During the backpropagation process, the gradient of each layer weight is calculated through the chain rule, and the model parameters are updated using the gradient descent algorithm. (6.5) During the training process, L2 regularization is added to constrain the parameters, and Dropout technology is used to randomly discard a portion of neurons during each training to prevent overfitting. (6.6) Small batch gradient descent method is used to iteratively optimize the model and update the model weights. The training process continues until the model loss function converges or the preset stopping condition is met.
6. The method of hyperspectral image classification according to claim 1 or 5, characterized in that: The loss function of the model in step (6) is shown below: where B denotes batch size, x b is the input data of the b-th batch, y b is the corresponding label, f(x b ) is the output predicted by the model, and H(·) denotes the cross-entropy loss function.
7. The method of hyperspectral image classification according to claim 5, characterized in that: The preset stopping condition in step (6.6) is to reach the maximum number of training rounds or the accuracy of the validation set reaches the preset level.
Citation Information
Cited By
Signal modulation identification method and system based on feature pyramid and Mama hybrid model, computer equipment and readable storage medium
CN121547327A
A signal modulation recognition method, system, computer device and readable storage medium based on a feature pyramid and Mamba hybrid model
CN121547327B