Frequency domain Mama multi-scale hyperspectral image classification method and device
Through the multi-scale hyperspectral image classification method of frequency domain Mamba, combined with spatial spectral interactive feature extraction, frequency domain perturbation embedding and adaptive bidirectional Mamba module, the problem of insufficient information integration in hyperspectral image classification is solved, and high-precision and stable classification effect is achieved.
Patent Information
- Application Number
- CN202510298941.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing hyperspectral image classification methods are difficult to effectively integrate information at different scales, resulting in loss of details and poor classification results, especially when dealing with detailed features such as textures and boundaries, the classification accuracy is affected.
The multi-scale hyperspectral image classification method of frequency domain Mamba is adopted. Through the spatial spectrum interactive feature extraction module, the frequency domain perturbation embedding module and the adaptive bidirectional Mamba module, combined with the multi-layer feature integration mechanism, high-frequency detailed features and low-frequency global information are extracted, the comprehensive utilization ability of spatial information is enhanced, and the long-distance dependence relationship is captured through adaptive bidirectional spectral analysis.
It significantly improves the accuracy and robustness of hyperspectral image classification, improves the clarity of classification boundaries and class distinction ability, and enhances the generalization ability and stability of the model in complex scenarios.
Smart Images

Figure CN120259735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image classification, and particularly to a multi-scale hyperspectral image classification method and device based on frequency-domain Mamba. Background Art
[0002] Hyperspectral images are high-dimensional data composed of a large number of continuous spectral bands, and can provide rich spatial and spectral details at the same time. This characteristic has attracted wide attention in the fields of geological exploration, anomaly detection, precision agriculture, atmospheric science, and military reconnaissance. However, due to its complex spectral-spatial relationship, how to achieve high-precision hyperspectral image classification is a challenging research topic.
[0003] Traditional hyperspectral image classification methods usually use techniques such as support vector machines, sparse representation models, linear regression, and maximum a posteriori estimation. These methods mainly rely on the extraction of shallow features and are difficult to obtain discriminative and informative features in complex environments. For example, these methods usually rely on single-scale feature processing and cannot effectively integrate information at different scales, resulting in the loss of details. Due to the neglect of the complementarity at different levels, the context information is prone to be inaccurate, affecting the image quality. The lack of a hierarchical structure makes it difficult for the model to adjust the feature importance, resulting in inaccurate information transmission and affecting the classification effect.
[0004] In recent years, deep learning techniques have significantly improved the performance of hyperspectral image classification. Among them, convolutional neural networks have been widely used due to their powerful feature extraction capabilities. However, the sliding window-based architecture of convolutional neural networks has limitations in capturing global context and long-range dependencies. With the development of the attention mechanism, Transformer-based models have performed excellently in hyperspectral image classification tasks, capturing the dependencies between remote positions and spectral bands through global modeling capabilities. However, although the self-attention mechanism based on Transformer can effectively model global relationships, due to its quadratic computational complexity, it significantly increases the complexity of the model, resulting in increased computational overhead and resource consumption.
[0005] In this context, the state space model (SSM) has become an emerging method, which demonstrates powerful capabilities in modeling long-range dependencies through state transition mechanisms. The Mamba model is a typical representative of SSM, showing superiority in multiple benchmark tasks through linear scalability and maintaining linear complexity. However, existing Mamba models still have limitations in hyperspectral image classification. For example, flattening a two-dimensional image into a one-dimensional sequence will destroy local spatial dependencies, resulting in the loss of spatial information. At the same time, single-scale spectral scanning limits the full utilization of spectral context information, especially when dealing with detailed features such as textures and boundaries, and the classification accuracy is affected. Summary of the Invention
[0006] In view of the above problems, the object of the present invention is to provide a multi-scale hyperspectral image classification method and device for frequency-domain Mamba. By introducing a frequency-domain perturbation embedding module, high-frequency detail features and low-frequency global information are extracted by using two-dimensional Fourier transform, and efficient extraction of local and global features is realized through grouped convolution embedding, enhancing the comprehensive utilization ability of spatial information; an adaptive bidirectional Mamba module is designed, and through a bidirectional spectral analysis mechanism and remote modeling ability, long-distance dependence relationships between spectral bands are effectively captured, and the balance between global information and local details is optimized in combination with frequency-domain characteristics, thereby improving the classification accuracy; a feature integration mechanism is proposed, and through adaptive cross-layer feature recalibration, deep fusion of multi-layer features is realized, effectively improving the problems of fuzzy classification boundaries and unclear regional differentiation; at the same time, spectral multi-scale features are gradually extracted through hierarchical structure design, and spatial multi-scale feature integration is carried out, significantly improving the generalization ability and performance of the classification model in complex scenarios.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] On the one hand, a multi-scale hyperspectral image classification method for frequency-domain Mamba is provided, and the method includes the following steps:
[0009] S1. Collect hyperspectral image data and establish a sample data set;
[0010] S2. Preprocess the sample data set, divide it into a training set and a test set, and divide the input image into multiple image patches;
[0011] S3. Input the image patches into a spatial-spectral interaction feature extraction module, extract local spatial features through a spatial branch, extract spectral-spatial features through a spectral-spatial fusion branch, and combine the local spatial features with the spectral-spatial features to obtain interaction features;
[0012] S4. Input the interaction features into a frequency-domain perturbation embedding module, convert them to the frequency domain through fast Fourier transform, extract the frequency-domain features after adding noise perturbation, and then restore them through inverse fast Fourier transform to obtain the features after frequency-domain processing;
[0013] S5. Input the features after frequency-domain processing into an adaptive bidirectional Mamba module to perform forward and backward extraction of spatial features and spectral features;
[0014] S6. Use a spectral feature integration module to integrate the spectral features at different stages to obtain the integrated spectral features;
[0015] S7. Use a spatial feature integration module to integrate the spatial features at different stages to obtain the integrated spatial features;
[0016] S8. Further integrate the integrated spectral features and the integrated spatial features to obtain global features and achieve the final classification.
[0017] Optionally, step S1 specifically includes:
[0018] Collect hyperspectral image data from publicly available hyperspectral image datasets to establish a sample dataset;
[0019] The publicly available hyperspectral image datasets include: Indian Pines, Pavia University, Trento, and LongKou.
[0020] Optionally, step S2 specifically includes:
[0021] Preprocess the sample dataset, including: removing noisy bands, normalization, and data augmentation;
[0022] Divide the preprocessed data into a training set and a test set in a ratio of 9:1;
[0023] Input image , where B represents the batch size, C in is the number of spectral channels of the input, and H and W are the height and width of the input image respectively; the input image is sliced into N image patches, denoted as , where represents the i-th image patch.
[0024] Optionally, the spatio-spectral interaction feature extraction module includes a spatial branch and a spectral-spatial fusion branch. Step S3 specifically includes:
[0025] Input the segmented image patches into the spatial branch and the spectral-spatial fusion branch simultaneously;
[0026] In the spatial branch, the image patches first perform channel interaction through 1×1 convolution while maintaining the spatial resolution, and then use 3×3 convolution to capture local spatial features; both of the above convolutions are accompanied by batch normalization and the GELU activation function. The specific formulas are as follows:
[0027]
[0028] Among them, Conv 2d represents 2D convolution, BN represents batch normalization, GELU represents the activation function, and x spa represents the local spatial features extracted through the spatial branch;
[0029] In the spectral-spatial fusion branch, extraction is performed through 1×1×1 3D convolution, and the formula is as follows:
[0030]
[0031] where Conv 3d represents 3D convolution, and x spe represents the spectral-spatial features extracted through the spectral-spatial fusion branch;
[0032] The local spatial features and the spectral-spatial features are combined by element-wise addition, and an attention map is generated through dynamic gate embedding. The combined features are multiplied element-wise with the attention map, and the formula is as follows:
[0033]
[0034] where x fus represents the combined features, σ represents the activation function, and x SSIB represents the finally obtained interaction features.
[0035] Optionally, step S4 specifically includes:
[0036] Performing a fast Fourier transform on the interaction feature x SSIB :
[0037]
[0038] where F is the fast Fourier transform operation, x SSIB is the interaction feature extracted by the spatio-spectral interaction feature extraction module, t is the feature map after the fast Fourier transform, t real is the real part feature of the input, and t imag is the imaginary part feature of the input;
[0039]
[0040] t real N is the real part feature after adding noise perturbation, and t imag N is the imaginary part feature after adding noise perturbation. α1 and α2 respectively represent multiplicative noise, sampled from the uniform distribution U(-1,1), and β1 and β2 respectively represent additive noise, sampled from the normal distribution N(0,1);
[0041] Performing grouped convolution embedding processing on the real part feature and the imaginary part feature after adding noise perturbation, and then respectively performing GELU non-linear activation function processing:
[0042]
[0043] Conv g [Group convolution] g Indicates group convolution, [Real part output after group convolution] Indicates the real part output after group convolution processing, [Imaginary part output after group convolution] Indicates the imaginary part output after group convolution processing; then, information fusion is performed through complex transformation:
[0044]
[0045] where t out [Complex feature map] out is the output complex feature map, composed of the real part and the imaginary part; finally, the feature after frequency domain processing is obtained through the inverse fast Fourier transform:
[0046]
[0047] where F -1 [Inverse fast Fourier transform] -1 Indicates the inverse fast Fourier transform, and T L [Feature after frequency domain processing] L is the feature after frequency domain processing.
[0048] Optionally, the step S5 specifically includes:
[0049] Performing rotational position encoding on the feature T L [Feature after frequency domain processing] L after frequency domain processing to obtain the embedded sequence representation T L ROPE , and normalizing and linearly projecting the embedded sequence representation T L ROPE into a high-dimensional space to obtain the forward feature [Forward feature] , denoted as z:
[0050]
[0051] Reversing the forward feature [Forward feature] to obtain [Backward feature] as the backward feature, and then using [Backward feature] and [Forward feature] as inputs to perform forward and backward spatial feature extraction; the forward feature [Forward feature] and the backward feature [Backward feature] are respectively passed through a 1D convolution and the activation function SiLU, and are respectively input into the state space equation SSM; the SSM models the conversion from the input x(t) to the output y(t) through the hidden state h(t), and the state space equation SSM is defined as follows:
[0052]
[0053] where A is the state transition matrix, B maps the input to the state, and C is the conversion from the state to the output; the hyperspectral data is a discrete signal, and the zero-order hold method is used for discretization so that the continuous state space model can be applied to discrete sequence modeling, and the conversion formula is as follows:
[0054]
[0055] Among them, is the discretized parameter, and ΔA, ΔB, and ΔC are learnable parameters introduced by the Mamba structure to dynamically adjust the context representation of the model and improve the spectral feature adaptability; x t is the state at time t, and h t is the hidden state at time t, and h t-1 is the hidden state at time t-1, and I is the unit eigenvector;
[0056] After passing through the state space equation SSM, it is then multiplied by the gating activation function G respectively to obtain the output features F f and F b , and the specific process is as follows:
[0057]
[0058] Among them, represents the flipping operation of the data; the output features of each branch are passed through the average pooling layer AP, convolutional layer Conv 2d , activation function layer RELU in the subsequent stage, and finally, after calculating the weights through the Sigmoid activation function, different features are weighted and fused to obtain the final output:
[0059]
[0060] Among them, W represents the weight calculated by the Sigmoid activation function, and F fuse represents the fused feature.
[0061] Optionally, the step S6 specifically includes:
[0062] Extract the spectral features of three different stages and input them into different convolutional layers C1, C2, and C3 respectively for channel dimension alignment operation;
[0063] After completing the channel dimension alignment, the spectral features enter the channel attention block to further adjust the weights θ1, θ2, and θ3 of each channel;
[0064] The adjusted spectral features are integrated through element-wise addition to obtain the integrated spectral features.
[0065] Optionally, the step S7 specifically includes:
[0066] Extract the spatial features of three different stages and construct cross-resolution feature representations through upsampling by 8 times, 4 times, and 2 times respectively;
[0067] The spatial features of the first stage are fused with the spatial features of the second and third stages after upsampling through a symmetric structure of 2x downsampling and 4x downsampling, respectively;
[0068] A linear layer and a Sigmoid activation function are introduced in each stage, and then passed through convolutional layers S1, S2, and S3 respectively, where the weights of each stage are dynamically adjusted by learnable parameters δ1, δ2, and δ3;
[0069] The adjusted spatial features are integrated through element-wise addition to obtain the integrated spatial features.
[0070] Optionally, the step S8 specifically includes:
[0071] The integrated spectral features and the integrated spatial features are sequentially passed through a global average pooling layer and a fully connected layer to obtain global features;
[0072] The Softmax classifier is used to calculate the class probabilities for the global features and output the classification results.
[0073] On the other hand, a multi-scale hyperspectral image classification device for frequency-domain Mamba is provided to implement the method described in any one of the above, and the device includes:
[0074] A data collection module for collecting hyperspectral image data and establishing a sample data set;
[0075] A preprocessing module for preprocessing the sample data set, dividing it into a training set and a test set, and splitting the input image into multiple image patches;
[0076] A spatial-spectral interaction feature extraction module, including a spatial branch and a spectral-spatial fusion branch; for the input image patches, local spatial features are extracted through the spatial branch, spectral-spatial features are extracted through the spectral-spatial fusion branch, and the local spatial features are combined with the spectral-spatial features to obtain interaction features;
[0077] A frequency-domain perturbation embedding module for converting the interaction features to the frequency domain through the fast Fourier transform, extracting the frequency-domain features after adding noise perturbations, and then restoring them through the inverse fast Fourier transform to obtain the features after frequency-domain processing;
[0078] An adaptive bidirectional Mamba module for forward and backward extraction of spatial features and spectral features from the features after frequency-domain processing;
[0079] A spectral feature integration module for integrating the spectral features of different stages to obtain the integrated spectral features;
[0080] A spatial feature integration module for integrating spatial features at different stages to obtain integrated spatial features;
[0081] A global feature integration and classification module for further integrating the integrated spectral features and the integrated spatial features to obtain global features and achieve final classification.
[0082] On the other hand, an electronic device is provided, and the electronic device includes:
[0083] A processor;
[0084] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the multi-scale hyperspectral image classification method of the frequency-domain Mamba as described above are implemented.
[0085] On the other hand, a computer-readable storage medium is provided, in which program code is stored. The program code can be called by a processor to execute the steps of the multi-scale hyperspectral image classification method of the frequency-domain Mamba as described above.
[0086] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0087] (1) Through the spatial-spectral interaction feature extraction module of the present invention, the problem of insufficiently mining spatial information in hyperspectral image classification is effectively solved. Multi-scale convolution and 3D convolution are used to extract local and global features, and the gating enhancement mechanism is combined to optimize feature expression, improving the fusion ability of spatial and spectral features, thereby enhancing the classification accuracy.
[0088] (2) The present invention innovatively introduces a frequency-domain perturbation embedding module, uses two-dimensional Fourier transform to extract high-frequency detail features and low-frequency global information, improves the feature representation ability by adding noise perturbation, and combines group convolution to optimize feature extraction, enabling the model to more accurately separate edge details and global background information, and improving the clarity of the classification boundary and the category discrimination ability.
[0089] (3) Aiming at the limitations of the existing methods in modeling long-range dependencies and insufficient utilization of spectral information, the present invention designs an adaptive bidirectional Mamba module, improves the dependence modeling ability of the spectral dimension through a bidirectional spectral analysis mechanism and rotational position encoding, ensures that the model can more comprehensively understand the spectral information of hyperspectral data, and improves the robustness of classification.
[0090] (4) The present invention adopts a spectral feature and spatial feature integration mechanism. Through cross-layer feature recalibration, it realizes the deep fusion of spectral features and spatial features at different levels, and combines hierarchical architecture design to gradually extract multi-level features, enabling the model to exhibit stronger generalization ability and stability in different datasets and complex scenarios, thereby effectively improving the accuracy and reliability of hyperspectral image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0092] Figure 1 is the overall block diagram of a multi-scale hyperspectral image classification method of frequency-domain Mamba provided by an embodiment of the present invention;
[0093] Figure 2 is the schematic diagram of image preprocessing provided by an embodiment of the present invention;
[0094] Figure 3 is the schematic diagram of the spatial-spectral interaction feature extraction module provided by an embodiment of the present invention;
[0095] Figure 4 is the schematic diagram of the frequency-domain perturbation embedding module provided by an embodiment of the present invention;
[0096] Figure 5 is the schematic diagram of the adaptive bidirectional Mamba module provided by an embodiment of the present invention;
[0097] Figure 6 is the schematic diagram of the spectral feature integration module provided by an embodiment of the present invention;
[0098] Figure 7 is the schematic diagram of the spatial feature integration module provided by an embodiment of the present invention;
[0099] Figure 8 is the schematic diagram of global information integration and classification provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0100] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0101] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner.
[0102] The embodiments of the present invention provide a multi-scale hyperspectral image classification method for frequency-domain Mamba. This method can be implemented by an electronic device, which can be a terminal or a server. As Figure 1 shown, the processing flow of this method includes the following steps:
[0103] S1. Collect hyperspectral image data and establish a sample data set.
[0104] In the embodiments of the present invention, hyperspectral image data is collected from publicly available hyperspectral image data sets to establish a sample data set. The publicly available hyperspectral image data sets include: Indian Pines, Pavia University, Trento, and LongKou, etc., which contain rich spectral information and spatial information.
[0105] S2. Preprocess the sample data set, divide it into a training set and a test set, and segment the input image into multiple image patches.
[0106] As Figure 2 shown, preprocessing the sample data set includes: removing noisy bands, normalizing, and data augmentation. Then, the preprocessed data is divided into a training set and a test set in a ratio of 9:1 to ensure data quality and the generalization ability of the model.
[0107] The preprocessed image is used as the input image of the model, denoted as , where B represents the batch size, C in is the number of input spectral channels, and H and W are the height and width of the input image respectively. The input image is cut into N image patches, denoted as , where represents the i-th image patch.
[0108] S3. Input the image patches into the spatio-spectral interaction feature extraction module, extract local spatial features through the spatial branch, extract spatio-spectral features through the spectral-spatial fusion branch, and combine the local spatial features with the spatio-spectral features to obtain interaction features.
[0109] As Figure 3 shown, the spatio-spectral interaction feature extraction module includes a spatial branch and a spectral-spatial fusion branch. The segmented image patches Input the spatial branch and the spectral-spatial fusion branch simultaneously.
[0110] In the spatial branch, the image patch first undergoes channel interaction through a 1×1 convolution while maintaining the spatial resolution, and then a 3×3 convolution is used to capture local spatial features; both of the above convolutions are accompanied by batch normalization and the GELU activation function to facilitate stable training and introduce non-linearity. The specific formulas are as follows:
[0111]
[0112] Among them, Conv 2d represents 2D convolution, BN represents batch normalization, GELU represents the activation function, and x spa represents the local spatial features extracted through the spatial branch.
[0113] In the spectral-spatial fusion branch, it is extracted through a 1×1×1 3D convolution. The formula is as follows:
[0114]
[0115] Among them, Conv 3d represents 3D convolution, and x spe represents the spectral-spatial features extracted through the spectral-spatial fusion branch.
[0116] The local spatial features and the spectral-spatial features are combined by element-wise addition, and an attention map is generated through dynamic gate embedding. The combined features are multiplied element-wise with this attention map. The formula is as follows:
[0117]
[0118] Among them, x fus represents the combined features, σ represents the activation function, and x SSIB represents the finally obtained interaction features.
[0119] S4. Input the interaction features into the frequency-domain perturbation embedding module, convert them to the frequency domain through the fast Fourier transform, extract the frequency-domain features with added noise perturbations, and then restore them through the inverse fast Fourier transform to obtain the features after frequency-domain processing.
[0120] As Figure 4 shown, the frequency-domain perturbation embedding module uses the two-dimensional fast Fourier transform to convert the input image to the frequency domain to enhance the separability of features. The low-frequency part retains the global information, and the high-frequency part extracts the local texture information. Noise perturbations are added to both the high and low frequencies, thereby improving the recognition ability for boundary regions. Group convolution is used to reduce the computational complexity while optimizing the feature expression ability of different frequency-domain components. Combining the frequency-domain extracted feature information and removing redundant information ensures that the model can learn more discriminative features.
[0121] The Fourier transform is used to convert a time-domain signal into the frequency domain to help analyze the frequency-domain characteristics of a hyperspectral image. The mathematical representation of the two-dimensional discrete Fourier transform is as follows:
[0122]
[0123] where F(h, w) represents the value of the signal at position (h, w), and u and v are the horizontal and vertical spatial frequency indices in the frequency domain, respectively. The corresponding inverse Fourier transform is used to recover the time-domain information:
[0124]
[0125] In the embodiments of the present invention, for the interaction feature x SSIB a fast Fourier transform is performed:
[0126]
[0127]
[0128] where F is the fast Fourier transform operation, x SSIB is the interaction feature extracted by the spatial spectral interaction feature extraction module, t is the feature map after the fast Fourier transform, t real is the real part feature of the input, and t imag is the imaginary part feature of the input.
[0129]
[0130] t real N is the real part feature after adding noise perturbation, and t imag N is the imaginary part feature after adding noise perturbation. α1 and α2 respectively represent multiplicative noise, sampled from the uniform distribution U(-1, 1), and β1 and β2 respectively represent additive noise, sampled from the normal distribution N(0, 1).
[0131] Performing grouped convolution embedding processing on the real part feature and the imaginary part feature after adding noise perturbation can effectively integrate spectral channel information and reduce complexity, and then GELU non-linear activation function processing is performed respectively to enhance non-linear feature expression:
[0132]
[0133] Conv g represents grouped convolution, represents the real part output after grouped convolution processing, represents the imaginary part output after grouped convolution processing. Then, information fusion is performed through complex number transformation:
[0134]
[0135] Among them, t out is the output complex feature map, which consists of a real part and an imaginary part. Finally, the features processed in the frequency domain are obtained through the inverse fast Fourier transform:
[0136]
[0137] Among them, F -1 represents the inverse fast Fourier transform, and T L are the features processed in the frequency domain.
[0138] In order to be able to process according to the hierarchical structure, the obtained T L spectral dimension is mapped to a higher feature space. The design of this module ensures that the model can efficiently extract feature information in the frequency domain and reconstruct and optimize the data in the spatial domain, improving the classification ability of the model for hyperspectral images.
[0139] S5. Input the features processed in the frequency domain into the adaptive bidirectional Mamba module to extract forward and backward spatial features and spectral features.
[0140] As Figure 5 shown, the features T L processed in the frequency domain are subjected to rotational position encoding to obtain the embedded sequence representation T L ROPE . The embedded sequence representation T L ROPE is normalized and linearly projected into a high-dimensional space to obtain the forward feature , denoted as z:
[0141]
[0142] The forward feature is reversed to obtain as the backward feature. Then, and are used as inputs to perform forward and backward spatial feature extraction. The forward feature and the backward feature are respectively passed through a 1D convolution and the activation function SiLU, and are respectively input into the state space equation SSM. This module uses the state space model SSM to process spectral sequence data, and describes the state evolution with a linear ordinary differential equation (ODE). SSM models the conversion from the input x(t) to the output y(t) through the hidden state h(t), and the state space equation SSM is defined as follows:
[0143]
[0144] Among them, A is the state transition matrix, B maps the input to the state, and C is the conversion from the state to the output; the hyperspectral data is a discrete signal, and the zero-order hold method is used for discretization so that the continuous state space model can be applied to discrete sequence modeling. The conversion formula is as follows:
[0145]
[0146] Among them, are the parameters after discretization, ΔA, ΔB, and ΔC are the learnable parameters introduced by the Mamba structure to dynamically adjust the context representation of the model and improve the spectral feature adaptability. x t is the state at time t, h t is the hidden state at time t, h t-1 is the hidden state at time t - 1, and I is the unit eigenvector.
[0147] After passing through the state space equation SSM, it is then multiplied by the gated activation function G respectively to obtain the output features F f and F b of the forward and backward branches. The specific process is as follows:
[0148]
[0149] Among them, represents the flipping operation of the data. The output features of each branch are passed through the average pooling layer AP, convolutional layer Conv 2d , and activation function layer RELU in the subsequent stage. Among them, the global feature integration is first performed through the average pooling layer, then the spectral information is extracted through multiple convolutional layers in sequence, and the nonlinear transformation is performed through the activation function layer RELU to enhance the expression ability of the features.
[0150] Finally, after calculating the weights through the Sigmoid activation function, different features are weighted and fused to obtain the final output:
[0151]
[0152] Among them, W represents the weight calculated by the Sigmoid activation function, and F fuse represents the fused features. This module enables the model to adaptively capture spectral information and improve the discriminative ability of the final features through multi-branch convolution, state modeling, and feature fusion.
[0153] S6. Use the spectral feature integration module to integrate the spectral features at different stages to obtain the integrated spectral features.
[0154] Such as Figure 6As shown, this module processes the spectral features in three stages, extracts the spectral features of different stages, and aligns them. Specifically, it includes: extracting the spectral features of three different stages, respectively inputting them into different convolutional layers C1, C2, and C3 for channel dimension alignment operations; after completing the channel dimension alignment, the spectral features enter the channel attention block to further adjust the weights θ1, θ2, and θ3 of each channel, enhancing the model's attention to important features; the adjusted spectral features are integrated through element-wise addition to achieve effective fusion and refinement of features, and finally the integrated spectral features are obtained.
[0155] S7. Use the spatial feature integration module to integrate the spatial features of different stages to obtain the integrated spatial features.
[0156] As Figure 7 shown, this module adopts a multi-scale spatial feature interaction mechanism to achieve in-depth integration of spatial context information through hierarchical sampling and non-linear transformation. Specifically, it includes: extracting the spatial features of three different stages, respectively constructing cross-resolution feature representations through 8-fold, 4-fold, and 2-fold upsampling; then, the spatial features of the first stage are fused with the upsampled spatial features of the second and third stages respectively through the symmetric structure of 2-fold downsampling and 4-fold downsampling, further strengthening the coupling relationship between local and global spatial features; a linear layer and a Sigmoid activation function are introduced in each stage to compress the dimension and perform non-linear calibration on the features respectively, and then through convolutional layers S1, S2, and S3, spatial-sensitive feature responses are extracted. Among them, the weights of each stage are dynamically adjusted through learnable parameters δ1, δ2, and δ3 to balance the semantic differences of multi-scale features in an adaptive manner; the adjusted spatial features are integrated through element-wise addition to obtain the integrated spatial features.
[0157] In the embodiments of the present invention, through cascaded multi-stage processing, the module can effectively capture the spatial dependencies under different receptive fields, suppress redundant information during the feature fusion process, and finally output strongly discriminative spatial enhanced features. This design significantly improves the model's ability to model complex spatial structures, especially in terms of texture detail restoration and geometric consistency maintenance.
[0158] S8. Further integrate the integrated spectral features and the integrated spatial features to obtain global features and achieve final classification.
[0159] As Figure 8 shown, the integrated spectral features and the integrated spatial features are sequentially passed through the global average pooling layer and the fully connected layer to obtain global features; then, a Softmax classifier is used to calculate the class probabilities for the global features and output the classification results.
[0160] Among them, the spatial information is compressed through global average pooling to reduce the computational burden and improve the model inference speed. The fully connected layer is used for final classification, and Softmax is combined to calculate the class probabilities to achieve efficient classification decisions. The goal of this step is to ensure that the classifier can maintain stable performance on different scenarios and datasets.
[0161] Correspondingly, an embodiment of the present invention further provides a multi-scale hyperspectral image classification device for frequency-domain Mamba. The device is used to implement the above multi-scale hyperspectral image classification method, and the device includes:
[0162] A data collection module, configured to collect hyperspectral image data and establish a sample data set;
[0163] A preprocessing module, configured to preprocess the sample data set, divide it into a training set and a test set, and segment the input image into multiple image patches;
[0164] A spatial-spectral interaction feature extraction module, including a spatial branch and a spectral-spatial fusion branch; for the input image patch, local spatial features are extracted through the spatial branch, spectral-spatial features are extracted through the spectral-spatial fusion branch, and the local spatial features are combined with the spectral-spatial features to obtain interaction features;
[0165] A frequency-domain perturbation embedding module, configured to convert the interaction features to the frequency domain through fast Fourier transform, extract the frequency-domain features after adding noise perturbations, and then restore them through inverse fast Fourier transform to obtain the features after frequency-domain processing;
[0166] An adaptive bidirectional Mamba module, configured to perform forward and backward extraction of spatial features and spectral features on the features after frequency-domain processing;
[0167] A spectral feature integration module, configured to integrate the spectral features at different stages to obtain the integrated spectral features;
[0168] A spatial feature integration module, configured to integrate the spatial features at different stages to obtain the integrated spatial features;
[0169] A global feature integration and classification module, configured to further integrate the integrated spectral features and the integrated spatial features to obtain global features and achieve final classification.
[0170] The device of this embodiment can be used to execute Figure 1 the technical solutions of the method embodiments shown, and its implementation principle and technical effects are similar, so they will not be elaborated here.
[0171] Compared with the prior art, the present invention has the following advantages:
[0172] (1) Solve the problem of insufficient exploration of spatial information: Existing methods process images by flattening them into one-dimensional sequences, resulting in the loss of spatial information. To address this issue, the present invention designs a spatial-spectral interaction module that effectively integrates spatial and spectral information. This module uses two-dimensional convolution to extract the spatial features of the image, while using three-dimensional convolution for the interaction of spectral and spatial information, preserving the spatial local dependencies in the image and enhancing the ability to capture spatial information, thereby improving the classification accuracy. Additionally, a frequency-domain perturbation embedding module is introduced. The data is transformed into the frequency domain through two-dimensional fast Fourier transform, and high-frequency details and low-frequency global information are extracted respectively. By adding noise perturbations, the feature representation ability is enhanced. Group convolution operations are used to extract features in different frequency domains, enabling information fusion at the frequency-domain level, effectively capturing the detailed features and global structure in the image, reducing the problem of blurred region boundaries, and improving the accuracy of image classification.
[0173] (2) Solve the problems of long-range dependence and insufficient utilization of spectral information: Traditional convolutional neural networks have difficulty capturing global context and long-range dependence through the sliding window mechanism. Although the self-attention mechanism based on Transformer can effectively model long-range dependence, it has high computational complexity and large resource consumption. Therefore, the present invention introduces an adaptive bidirectional Mamba module that captures long-range dependence by bidirectionally scanning spectral information. Combining the linear state-space modeling ability of the Mamba model, the computational efficiency is effectively improved. The linear transfer characteristics of Mamba result in a lower computational cost when dealing with long-range dependence, avoiding the high computational overhead of traditional methods, while retaining the sensitivity to local details, thereby improving the processing efficiency of the model under large-scale data and the ability to capture global information.
[0174] (3) Solve the problem of insufficient generalization ability: Existing methods usually rely on a single feature processing method and lack effective hierarchical processing, resulting in insufficient information integration. Especially in complex data scenarios, the classification effect is unstable. To solve this problem, the present invention adopts a hierarchical structure. This structure processes the image layer by layer through multiple feature extraction stages. Each layer extracts feature information at different levels, and the spectral and spatial feature integration module fuses and weights the features at different levels. Through this hierarchical processing, the model can adjust the importance of features at different levels, ensuring the effective integration of multi-scale information, and improving the generalization ability and classification accuracy of the model.
[0175] In summary, the present invention combines spatial-spectral interaction feature extraction, frequency-domain perturbation embedding, adaptive bidirectional Mamba, and hierarchical multi-scale feature fusion to break through the limitations of traditional hyperspectral image classification methods and achieve efficient and accurate hyperspectral image classification. After the hyperspectral image data is input, the spatial-spectral interaction feature extraction module is used to extract interaction features, and then the frequency-domain perturbation embedding module and the adaptive bidirectional Mamba module are used to model long-range dependencies. Finally, the hierarchical feature fusion module optimizes the classification results and outputs a high-precision classification map. Experimental verification shows that this method has achieved leading classification accuracy on multiple hyperspectral datasets and is superior to existing methods in terms of boundary clarity, computational efficiency, and generalization ability. By further optimizing the computational efficiency, the present invention can be applied to larger-scale hyperspectral data processing tasks and extended to multi-modal data fusion and remote sensing application fields.
[0176] In an exemplary embodiment, the present invention also provides an electronic device, which includes:
[0177] A processor;
[0178] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the multi-scale hyperspectral image classification method of frequency-domain Mamba as described above are implemented.
[0179] In an exemplary embodiment, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by the processor to implement the steps of the multi-scale hyperspectral image classification method of frequency-domain Mamba as described above. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0180] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or terminal device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or terminal device including the element.
[0181] The specification mentions "one embodiment", "an embodiment", "an exemplary embodiment", "some embodiments", etc., indicating that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when describing a specific feature, structure, or characteristic in combination with an embodiment, implementing such feature, structure, or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0182] It should be understood that the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Herein, A and B can be singular or plural. Additionally, the character " / " herein generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0183] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0184] It should be understood that in various embodiments of the present invention, the magnitude of the sequence numbers of the above processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0185] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0186] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, may exist independently physically for each unit, or two or more units may be integrated in one unit.
[0188] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0189] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0190] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-scale hyperspectral image classification method for frequency-domain Mamba, characterized in that, It includes the following steps: S1. Collect hyperspectral image data and establish a sample data set; S2. Preprocess the sample data set, divide it into a training set and a test set, and segment the input image into multiple image patches; S3. Input the image patches into the spatial-spectral interaction feature extraction module, extract local spatial features through the spatial branch, extract spectral-spatial features through the spectral-spatial fusion branch, and combine the local spatial features with the spectral-spatial features to obtain interaction features; S4. Input the interaction features into the frequency-domain perturbation embedding module, convert them to the frequency domain through the fast Fourier transform, extract the frequency-domain features after adding noise perturbations, and then restore them through the inverse fast Fourier transform to obtain the features after frequency-domain processing; S5. Input the features after frequency-domain processing into the adaptive bidirectional Mamba module to perform forward and backward spatial feature and spectral feature extraction; S6. Use the spectral feature integration module to integrate the spectral features at different stages to obtain the integrated spectral features; S7. Use the spatial feature integration module to integrate the spatial features at different stages to obtain the integrated spatial features; S8. Further integrate the integrated spectral features and the integrated spatial features to obtain global features and achieve final classification.
2. The multi-scale hyperspectral image classification method according to claim 1, wherein The specific steps of S1 include: Collect hyperspectral image data from the publicly available hyperspectral image data set and establish a sample data set; The publicly available hyperspectral image data set includes: Indian Pines, Pavia University, Trento, and LongKou.
3. The multi-scale hyperspectral image classification method according to claim 1, characterized in that The specific steps of S2 include: Preprocess the sample data set, including: removing noisy bands, normalization processing, and data augmentation; Divide the preprocessed data into a training set and a test set according to a ratio of 9:1; Input image , where B represents the batch size, and C in is the number of spectral channels of the input, and H and W are the height and width of the input image respectively; the input image is sliced into N image patches, denoted as , where represents the i-th image patch.
4. The multi-scale hyperspectral image classification method according to claim 1, wherein The spatial-spectral interaction feature extraction module includes a spatial branch and a spectral-spatial fusion branch. The specific steps of S3 include: The segmented image patches are simultaneously input into the spatial branch and the spectral-spatial fusion branch; In the spatial branch, the image patches first perform channel interaction through 1×1 convolution while maintaining the spatial resolution, and then use 3×3 convolution to capture local spatial features; both of the above convolutions are accompanied by batch normalization and the GELU activation function. The specific formulas are as follows: ; Among them, Conv 2d represents 2D convolution, BN represents batch normalization, GELU represents the activation function, and x spa represents the local spatial features extracted by the spatial branch; In the spectral-spatial fusion branch, it is extracted through 3D convolution of 1×1×1, and the formula is as follows: ; Among them, Conv 3d represents 3D convolution, and x spe represents the spectral-spatial features extracted by the spectral-spatial fusion branch; Add the local spatial features and the spectral-spatial features element-wise for combination, generate an attention map through dynamic gate embedding, and multiply the combined features with the attention map element-wise. The formula is as follows: ; Among them, x fus represents the combined feature, σ represents the activation function, and x SSIB represents the finally obtained interaction feature.
5. The multi-scale hyperspectral image classification method according to claim 1, characterized in that The specific steps of S4 include: For the interaction feature x SSIB perform a fast Fourier transform: ; Among them, F is the fast Fourier transform operation, and x SSIB is the interaction feature extracted by the spatial spectral interaction feature extraction module, t is the feature map after the fast Fourier transform, and t real is the real part feature of the input, and t imag is the imaginary part feature of the input; ; t real N is the real part feature after adding noise perturbation, t imag N is the imaginary part feature after adding noise perturbation, α1 and α2 respectively represent multiplicative noise, sampled from the uniform distribution U(-1, 1), and β1 and β2 respectively represent additive noise, sampled from the normal distribution N(0, 1); Perform group convolution embedding processing on the real part features and imaginary part features after adding noise perturbations, and then perform GELU non-linear activation function processing respectively: ; Conv g Indicates group convolution, Indicates the real part output after group convolution processing, Indicates the imaginary part output after group convolution processing; then, information fusion is performed through complex number transformation: ; where t out is the output complex feature map, which consists of a real part and an imaginary part; finally, the features processed in the frequency domain are obtained through inverse fast Fourier transform: ; Among them, F -1 represents the inverse fast Fourier transform, and T L is the feature after frequency domain processing.
6. The multi-scale hyperspectral image classification method according to claim 1, characterized in that The specific steps of S5 include: The feature T after frequency-domain processing L Perform rotational position encoding to obtain the embedded sequence representation T L ROPE , and normalize and linearly project the embedded sequence representation T L ROPE onto a high-dimensional space to obtain the forward feature , denoted as z: ; Reverse the forward feature to obtain as the backward feature. Then, use and as inputs to perform forward and backward spatial feature extraction. The forward feature and the backward feature are respectively passed through 1D convolution and the activation function SiLU, and then are respectively input into the state space equation SSM. The SSM models the conversion from the input x(t) to the output y(t) through the hidden state h(t). The state space equation SSM is defined as follows: ; Among them, A is the state transition matrix, B maps the input to the state, and C is the conversion from the state to the output; the hyperspectral data is a discrete signal, and the zero-order hold method is used for discretization so that the continuous state space model can be applied to discrete sequence modeling. The conversion formula is as follows: ; Among them, is the parameter after discretization, and ΔA, ΔB, and ΔC are learnable parameters introduced by the Mamba structure to dynamically adjust the context representation of the model and improve the spectral feature adaptability; x t is the state at time t, and h t is the hidden state at time t, and h t-1 is the hidden state at time t-1, and I is the unit eigenvector; After passing through the state space equation SSM, multiply by the gated activation function G respectively to obtain the output features F of the forward branch and the backward branch f and F b , and the specific process is as follows: ; Among them, represents the flipping operation of the data; the output features of each branch are passed through the average pooling layer AP and the convolutional layer Conv 2d , the activation function layer RELU, and finally, after calculating the weights through the Sigmoid activation function, different features are weighted and fused to obtain the final output: ; Among them, W represents the weight calculated by the Sigmoid activation function, and F fuse represents the fused feature.
7. The multi-scale hyperspectral image classification method according to claim 1, wherein The specific steps of S6 include: Extract the spectral features of three different stages and input them into different convolutional layers C1, C2, and C3 respectively for channel dimension alignment operation; After completing the channel dimension alignment, the spectral features enter the channel attention block to further adjust the weights θ1, θ2, and θ3 of each channel; The adjusted spectral features are integrated through element-wise addition to obtain the integrated spectral features.
8. The multi-scale hyperspectral image classification method according to claim 1, characterized in that The specific steps of step S7 include: Extract the spatial features of three different stages and construct cross-resolution feature representations through 8x, 4x, and 2x upsampling respectively; The spatial features of the first stage are fused with the upsampled spatial features of the second and third stages respectively through the symmetric structure of 2x downsampling and 4x downsampling; A linear layer and a Sigmoid activation function are introduced in each stage, and then they pass through convolutional layers S1, S2, and S3 respectively, where the weights of each stage are dynamically adjusted through learnable parameters δ1, δ2, and δ3; The adjusted spatial features are integrated through element-wise addition to obtain the integrated spatial features.
9. The multi-scale hyperspectral image classification method according to claim 1, wherein The specific steps of step S8 include: The integrated spectral features and the integrated spatial features are sequentially passed through the global average pooling layer and the fully connected layer to obtain the global features; Use the Softmax classifier to calculate the class probabilities for the global features and output the classification results.
10. A multi-scale hyperspectral image classification device for frequency-domain Mamba, the device being used to implement the method according to any one of claims 1 to 9, characterized in that, The device includes: A data collection module for collecting hyperspectral image data and establishing a sample data set; A preprocessing module for preprocessing the sample data set, dividing it into a training set and a test set, and splitting the input image into multiple image patches; A spatial-spectral interaction feature extraction module, including a spatial branch and a spectral-spatial fusion branch; for the input image patches, extract local spatial features through the spatial branch, extract spectral-spatial features through the spectral-spatial fusion branch, and combine the local spatial features with the spectral-spatial features to obtain interaction features; A frequency-domain perturbation embedding module for converting the interaction features to the frequency domain through the fast Fourier transform, extracting the frequency-domain features with added noise perturbations, and then restoring them through the inverse fast Fourier transform to obtain the frequency-domain processed features; An adaptive bidirectional Mamba module for forward and backward extraction of spatial features and spectral features from the frequency-domain processed features; A spectral feature integration module for integrating the spectral features of different stages to obtain the integrated spectral features; A spatial feature integration module for integrating the spatial features of different stages to obtain the integrated spatial features; A global feature integration and classification module for further integrating the integrated spectral features and the integrated spatial features to obtain the global features and achieve the final classification.
Citation Information
Patent Citations
High-resolution remote sensing image target detection method based on multi-scale network
CN118485927A
Hyperspectral image classification method and system based on local-global feature extraction
CN118628788A
Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model
CN119152205A
Face sketch synthesis method and system based on perception loss
CN119228938A
Medical image segmentation method based on residual axial attention
CN119229127A
Cited By
Medical hyperspectral image enhancement method and system based on multi-domain fusion
CN121582077A
Hyperspectral image generalization classification method and device based on dynamic frequency domain mixing and progressive decoupling
CN121861492A
Multi-head attention-based LIBS spectrum grade prediction method
CN121958991A
Hyperspectral image denoising method based on spatial spectral feature fusion
CN121961915A