A Hyperspectral Image Classification Method Combining Mamba and Chebyshev Graph Convolution
By fusing the convolution method of Mamba and Chebishev, the spectral features and spatial information in hyperspectral images are effectively fused, and the problems of fusion difficulties and high computational complexity in the prior art are solved, thereby achieving high-precision hyperspectral image classification and efficient computing performance.
Patent Information
- Application Number
- CN202510241315.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Existing hyperspectral image classification methods are difficult to effectively integrate spectral features and spatial information, and the computational complexity and model complexity are high, making it difficult to process large-scale hyperspectral data.
The hyperspectral image classification method that integrates Mamba and Chebischev graph convolution is adopted, and the bidirectional Mamba model captures the long-range dependence between space and spectrum through band selection, and uses the feature propagation ability of reparameterized Chebischev graph convolution in non-Euclidean space to deeply explore the complementarity between spatial features and spectral features.
It significantly improves the accuracy of hyperspectral image classification, reduces calculation costs, and improves the efficiency and stability of processing large-scale hyperspectral data.
Smart Images

Figure CN119741558B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image detection, and particularly to a hyperspectral image classification method that fuses Mamba and Chebyshev graph convolutions. Background Art
[0002] Hyperspectral image detection technology has a wide range of applications in multiple fields, such as environmental monitoring, agricultural remote sensing, etc. Hyperspectral imaging technology can provide rich spectral information in the range from ultraviolet to near-infrared, and can analyze objects and scenes in the image more accurately. However, traditional hyperspectral image classification methods usually face many challenges, such as the effective fusion of spectral features and spatial information, high computational complexity when dealing with large-scale data, etc.
[0003] Currently, hyperspectral image classification methods based on basic spectral information include support vector machines, nearest neighbors, decision trees, extreme learning machines, etc. These methods mainly rely on spectral features for pixel-level classification, but often ignore the spatial information of the image. Therefore, for complex image scenes, the classification accuracy is usually not high. To improve the classification accuracy, in recent years, researchers have tried to combine spatial information with spectral information and proposed hyperspectral image classification models based on convolutional neural networks or graph convolutional networks. Usually, spatial information is processed through multiple layers of convolutions, and spectral information is used for feature extraction. However, the above methods still have the following problems:
[0004] (1) Most current hyperspectral image classification methods use independent network architectures to extract spatial features and spectral features respectively, and usually ignore the deep association between the two during processing. Although these methods can ensure a certain classification accuracy in some cases, when fusing spatial scale and spectral scale information, they often use simple linear fusion methods or independent processing, and it is difficult to effectively capture the complex non-linear relationship between the two and the feature differences at different scales, resulting in the model being unable to fully explore the complementarity between spatial information and spectral features.
[0005] (2) Although traditional network models such as convolutional neural networks and transformer models have significantly improved classification accuracy compared to machine learning methods, their computational complexity and model complexity are still problems, especially in the processing of large-scale hyperspectral datasets. Since the resolution of hyperspectral images is generally relatively high, when the dataset scale increases, it is difficult to improve the computational efficiency of traditional models. Therefore, when processing large-scale hyperspectral images, the required resource and time costs increase significantly. And due to the limitation of the network sampling structure of general convolutional neural network models, they are unable to process information in non-Euclidean spaces and are difficult to effectively establish the similarity relationship between pixels in the image.
[0006] (3) Existing hyperspectral classification models often adopt a single network architecture, which makes it difficult to balance the processing of spectral features and spatial features while capturing the similarity relationships between pixels in the image. It is very difficult for a single network architecture to achieve a balance among spectral features, spatial features, and pixel relationship modeling. Convolutional neural networks mainly focus on the extraction of spatial features in their design, so the processing of spectral features is usually relatively simple, which may lead to the weakening of the importance of spectral features. In addition, there may be cross-region spectral similarities between pixels in complex hyperspectral scenarios, and these pixels may be ignored by convolutional neural networks. If the graph convolutional neural network does not combine spatial features, it will be difficult to distinguish images with similar spectra but different spatial distributions.
[0007] Therefore, there is a problem of difficulty in fusing spectral features and spatial features in the prior art. Summary of the Invention
[0008] Aiming at the above problems, the purpose of the present invention is to provide a hyperspectral image classification method that fuses Mamba and Chebyshev graph convolution. By band selection, the bidirectional Mamba model is enhanced to capture the long-range dependence relationships in the spatial and spectral dimensions. At the same time, the feature propagation ability of the reparameterized Chebyshev graph convolution in the non-Euclidean space is utilized to deeply explore the complementarity between spatial features and spectral features, significantly improving the classification accuracy, reducing the computational cost, and enhancing the efficiency and stability when processing large-scale hyperspectral data.
[0009] To solve the above technical problems, the present invention provides the following technical solutions:
[0010] On the one hand, a hyperspectral image classification method that fuses Mamba and Chebyshev graph convolution is provided. The method includes the following steps:
[0011] S1. Collect a hyperspectral image dataset and divide it into a training set and a validation set according to a certain proportion;
[0012] S2. Preprocess the image data, and divide each image into a predetermined number of image patches X centered on each pixel;
[0013] S3. Input the image patch X into a heterogeneous space convolution block for processing to obtain output data X with different receptive fields; out ;
[0014] S4. Divide X out into two branches. One branch is processed by a band selection enhanced bidirectional Mamba branch to obtain the long-range dependence relationships between spatial features and spectral features in the image;
[0015] S5. Use the band selection Mamba model to generate channel importance scores S for different spectral channels, and perform weighted processing on the output of the band selection enhanced bidirectional Mamba branch to obtain the first output YM ;
[0016] S6. Input another branch of X out into the reparameterized Chebyshev graph convolutional branch for processing to obtain the similarity relationship between different pixels in the image, and obtain the second output Y;
[0017] S7. Use the dual-branch integration module to fuse the first output Y M and the second output Y to obtain the dual-branch fusion feature, and input it into the classifier to obtain the final classification result.
[0018] Optionally, step S1 specifically includes:
[0019] Collect multiple different publicly available hyperspectral image datasets with different spectral band numbers, coverage ranges, image sizes, and ground sampling distances between different datasets;
[0020] Divide the collected hyperspectral image datasets into a training set and a validation set at a ratio of 11:1; among them, the training set is used for model training, and the validation set is used to verify the model effect; save the model with the best performance on the validation set as the final model.
[0021] Optionally, step S2 specifically includes:
[0022] Divide each image into a predetermined number of image patches X centered on each pixel point as the input of the model; input image patch X ∈ R B×Cin×H×W , where B represents the batch size, C in is the number of spectral channels of the input, and H and W are the height and width of the image patch.
[0023] Optionally, in step S3, the processing process of the heterogeneous spatial convolution block includes:
[0024] Use a 1×1 convolution to map the input image patch X to a high-dimensional embedding space to generate a high-dimensional feature X emb :
[0025]
[0026] where X emb ∈ R B×Cemb×H×W , C emb is set to 256, indicating the number of spectral channels after embedding;
[0027] Perform 1×1 convolution operation and 3×3 convolution operation on the high-dimensional feature X emb , fuse the two convolution operation results element-wise to obtain a multi-scale feature combination, and then use a 1×1 convolution to restore the number of spectral channels to the original input spectral channel number C in , the operation is as follows:
[0028]
[0029] where X out ∈R B×Cin×H×W , the 1×1 convolution operation is used to extract the local fine-grained features of each pixel point, and the 3×3 convolution operation is used to extract the coarse-grained features larger than the local receptive field.
[0030] Optionally, in the step S4, the processing process of the band selection enhanced bidirectional Mamba branch includes:
[0031] Input the output data X of the heterogeneous spatial convolution block out into the position information encoder, and add position information to the input data through position information encoding to provide the context information of the pixel arrangement in space;
[0032] Input the encoded data into the depthwise separable convolution layer for feature extraction, and extract the spatial feature X of the image c ; the formula expression of the band selection enhanced bidirectional Mamba branch is:
[0033]
[0034] where Mamba spaf (X c ) and Mamba spab (X c ) represent the operation processes of the forward spatial Mamba and the backward spatial Mamba respectively, Flip(X c ) represents the flipping operation of the data, X spaf is the output of the forward spatial Mamba, and X spab is the output of the backward spatial Mamba.
[0035] Optionally, in the step S5, the processing process of the band selection Mamba model includes:
[0036] Exchange the spatial dimension and the spectral dimension of the output result X of the depthwise separable convolution layer c to obtain X r ,
[0037]
[0038] After X r passes through the forward spectral Mamba and the backward spectral Mamba operations respectively, add them to obtain the spectral feature X b :
[0039]
[0040] For the spectral feature X bPerform pooling, fully connected layer, and sigmoid activation function operations to generate the channel importance score S, where S represents the importance of different spectral channels:
[0041]
[0042] Multiply the output of the band selection enhanced bidirectional mamba branch by the channel importance score S to obtain the weighted feature Y spaf and Y spab :
[0043]
[0044] Fuse the weighted features through a convolutional layer and a concatenation operation to obtain the first output Y M :
[0045]
[0046] Y M is a feature representation that includes the long-range dependence relationship of the spectral and spatial information in the forward and backward directions in the image.
[0047] Optionally, in step S6, the processing process of the reparameterized Chebyshev graph convolutional branch includes:
[0048] Normalize the output data X of the heterogeneous spatial convolutional block, and regard each pixel as a node in the graph structure; calculate the similarity between all nodes using cosine similarity to obtain the vector similarity matrix M; out Set the threshold T for the matching indicator, filter the vector similarity matrix M to generate the adjacency matrix A; perform degree mapping on the adjacency matrix to obtain the degree matrix D, and construct the graph Laplacian matrix L according to the degree matrix D and the adjacency matrix A;
[0049] To enable the nodes to retain their own features, add self-loops to the graph Laplacian matrix L to obtain the updated graph Laplacian matrix
[0050] ; ;
[0051] Use the Chebyshev iteration method to calculate the updated graph Laplacian matrix to obtain Chebyshev polynomials of multiple orders; represents the k-th order Chebyshev polynomial, which is calculated through a recursive relationship; input the calculation result into the reparameterized weight calculator to obtain the reparameterized Chebyshev coefficient w k :
[0052] The output of the reparameterized Chebyshev graph convolutional branch, that is, the second output Y, is expressed as:
[0053]
[0054] Among them, X in is the input feature matrix, which contains the features of the nodes in the graph; Y ∈ R N×C , N is the number of pixels, C is the dimension of the output features, and K is the order of the Chebyshev polynomial.
[0055] Optionally, the step S7 specifically includes:
[0056] Perform pooling, first-layer convolution, first activation function, second-layer convolution, and second activation function operations on the first output Y of the band selection enhanced bidirectional mamba branch M and the second output Y of the reparameterized Chebyshev graph convolutional branch respectively to generate channel attention weights;
[0057] Perform element-wise multiplication of the output features of each branch with the corresponding channel attention weights, and then fuse the weighted features to obtain the double-branch fusion features;
[0058] Input the double-branch fusion features into a classifier to perform a classification task to obtain the final classification result.
[0059] On the other hand, an electronic device is provided, and the electronic device includes:
[0060] A processor;
[0061] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are loaded and executed by the processor, the steps of the above hyperspectral image classification method are implemented.
[0062] On the other hand, a computer-readable storage medium is provided, in which program code is stored, and the program code can be called by a processor to execute the steps of the above hyperspectral image classification method.
[0063] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0064] 1. The present invention provides a double-branch structure combining a band selection enhanced bidirectional mamba network and a reparameterized Chebyshev graph convolutional network, which can process spectral information and spatial information simultaneously, effectively fuse the features of both, accurately capture the complex non-linear relationship between spectral features and spatial features, and at the same time can comprehensively capture the similarity relationship between image pixels, improving the accuracy of hyperspectral image classification.
[0065] 2. The graph structure modeling of graph convolution in the present invention can effectively process the similarity relationship between pixels in non-Euclidean space, use the mamba model to enhance the selection of band information and construct the long-range dependence relationship of pixels, providing a new solution for exploring complex spatial relationships in hyperspectral images.
[0066] 3. Relying on the linear complexity of the Mamba model and the efficient graph structure modeling advantages of the graph convolutional network, the processing efficiency of the present invention on large-scale datasets is superior to that of convolutional neural networks and transformer models, and it has broad application potential in environments with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0068] Figure 1 is a schematic overall flow diagram of a hyperspectral image classification method that integrates Mamba and Chebyshev graph convolution provided by an embodiment of the present invention;
[0069] Figure 2 is a schematic diagram of a heterogeneous spatial convolution block provided by an embodiment of the present invention;
[0070] Figure 3 is a schematic diagram of a band selection enhanced bidirectional Mamba branch provided by an embodiment of the present invention;
[0071] Figure 4 is a schematic diagram of a band selection Mamba model provided by an embodiment of the present invention;
[0072] Figure 5 is a schematic diagram of a reparameterized Chebyshev graph convolution branch provided by an embodiment of the present invention;
[0073] Figure 6 is a schematic diagram of a dual-branch integration module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present invention in conjunction with the drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0075] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "exemplary" is intended to present concepts in a specific manner.
[0076] An embodiment of the present invention provides a hyperspectral image classification method that fuses Mamba and Chebyshev graph convolutions. As Figure 1 shown, the overall process of this method includes: The original hyperspectral image is first divided into image patches. After these image patches are processed by the heterogeneous spatial convolution block, they are divided into two branches. One branch is processed by the band selection enhanced bidirectional Mamba branch to obtain the long-range dependence relationship between the spatial features and spectral features in the hyperspectral image. The other branch is processed by the reparameterized Chebyshev graph convolution branch to obtain the similarity relationship between different pixels in the image. Finally, the outputs of the two branches are effectively fused through the dual-branch integration module, and then the fused features are processed through the pooling operation and the classifier to output the final classification result.
[0077] In the training stage, the cross-entropy loss function (Cross Entropy Loss) is used to measure the outputs of the band selection enhanced bidirectional Mamba branch and the reparameterized Chebyshev graph convolution branch and the final output prediction error. At the same time, in order to ensure that the decision results of the two branches are consistent, the KL divergence loss function (Kullback-Leibler Divergence Loss) is introduced for the outputs of the two branches.
[0078] Specifically, the method includes the following steps:
[0079] S1. Collect a hyperspectral image dataset and divide it into a training set and a validation set according to a ratio.
[0080] In this step, four different widely used public hyperspectral image datasets are collected. The number of spectral bands, coverage range, image size, and ground sampling distance are different among different datasets.
[0081] The collected hyperspectral image dataset is divided into a training set and a validation set according to a ratio of 11:1 to form a complete dataset. Among them, the training set is used for model training, and the validation set is used to verify the model effect. The model with the best performance on the validation set is saved as the final model.
[0082] S2. Preprocess the image data and divide each image into a predetermined number of image patches X centered on each pixel.
[0083] In this step, each image is divided into a predetermined number of image patches X centered on each pixel. Each pixel and its surrounding pixels together form an image patch. In this way, as many image patches as there are pixels in the image can be divided, which serves as the input of the model.
[0084] The input image patch X ∈ R B×Cin×H×W , where B represents the batch size, C inis the number of input spectral channels, and H and W are the height and width of the image patch.
[0085] S3. Input the image patch X into the heterogeneous spatial convolution block for processing to obtain output data X with different receptive fields. out 。
[0086] In this step, as Figure 2 shown, the processing process of the heterogeneous spatial convolution block includes:
[0087] Use a 1×1 convolution to map the input image patch X to a high-dimensional embedding space to generate high-dimensional feature X emb :
[0088]
[0089] where X emb ∈R B×Cemb×H×W , and C emb is set to 256, representing the number of spectral channels after embedding.
[0090] This operation maps the spectral dimension of the input data to a higher feature space to prepare for subsequent operations such as convolution.
[0091] Perform 1×1 convolution operation and 3×3 convolution operation on the high-dimensional feature X emb , fuse the results of the two convolution operations element-wise to obtain a multi-scale feature combination, and then use a 1×1 convolution to restore the number of spectral channels to the original input spectral channel number C in , and the operation is as follows:
[0092]
[0093] where X out ∈R B×Cin×H×W . The 1×1 convolution operation is used to extract the local fine-grained features of each pixel point, and the 3×3 convolution operation is used to extract the coarse-grained features larger than the local receptive field.
[0094] S4. Divide X out into two branches, and one of the branches is processed through the band selection enhanced bidirectional mamba branch to obtain the long-range dependence relationship between the spatial features and spectral features in the image.
[0095] In this step, as Figure 3 shown, the processing process of the band selection enhanced bidirectional mamba branch includes:
[0096] Input the output data X out of the heterogeneous spatial convolution block into the position information encoder, and add position information to the input data through position information encoding to provide the context information of the pixel arrangement in space;
[0097] The encoded data is input into the depthwise separable convolutional layer for feature extraction to extract the spatial feature X of the image. c After that, X c is used as the input to perform forward and backward spatial feature information extraction.
[0098] The general state space model equation is:
[0099]
[0100] where h(t) ∈ R N represents the latent state, x(t) ∈ R represents the input, y(t) ∈ R represents the output, A ∈ R N×N is the state transition matrix, B ∈ R N and C ∈ R N are the projection matrices that map the input and the latent state to the output, respectively. The latent state h(t) encapsulates the internal state of the system, and these internal states evolve over time according to the input x(t) and the system dynamics determined by A. To meet the requirements of practical applications, the continuous state space model equation is discretized:
[0101]
[0102] where represents the discrete time step, represents the discrete state transition matrix, represents the discrete input matrix, and I is the identity matrix. This discretization can accurately capture the system dynamics at each time step, allowing the state to be updated at discrete intervals. The discretized state space model equation can be expressed as:
[0103]
[0104] where h t represents the latent state at time t. The output at time t is denoted by y t , and C is the output matrix. This discrete model ensures the same dynamic retention performance as the original continuous system model and can be applied to forward and backward discrete state space modeling.
[0105] Applied to the Mamba branch, the formula for the band selection enhanced bidirectional Mamba branch is:
[0106]
[0107] where, Mamba spaf (X c ) and Mamba spab (X c ) represent the operation processes of the forward spatial Mamba and the backward spatial Mamba respectively, Flip(Xc ) represents the flipping operation of the data, X spaf is the output of the forward spatial mamba. After that, the input sequence is flipped for the backward spatial mamba, X spab is the output of the backward spatial mamba. The forward spatial mamba extracts the forward dependencies from the start point to the end point of the data, and the backward spatial mamba extracts the backward dependencies from the end point to the start point of the data.
[0108] S5. Use the band selection mamba model to generate the channel importance score S for different spectral channels, and perform weighted processing on the output of the band selection enhanced bidirectional mamba branch to obtain the first output Y M .
[0109] In this step, as Figure 4 shown, the processing process of the band selection mamba model includes:
[0110] Exchange the spatial dimension and spectral dimension of the output result X c of the depthwise separable convolutional layer to obtain X r ,
[0111]
[0112] Pass X r through the forward spectral mamba and backward spectral mamba operations respectively, and then add them to obtain the spectral feature X b :
[0113]
[0114] Perform pooling, fully connected layer, and sigmoid activation function operations on the spectral feature X b to generate the channel importance score S, and S represents the importance of different spectral channels:
[0115]
[0116] Multiply the output of the band selection enhanced bidirectional mamba branch by the channel importance score S to obtain the weighted feature Y spaf and Y spab :
[0117]
[0118] Fuse the weighted features through convolutional layer and concatenation operations to obtain the first output Y M :
[0119]
[0120] Y M is a feature representation that contains the long-range dependencies of the forward and backward spectral and spatial information in the image.
[0121] S6. Input another branch of X out into the reparameterized Chebyshev graph convolutional branch for processing to obtain the similarity relationship between different pixels in the image, and obtain the second output Y.
[0122] In this step, as Figure 5 shown, the processing process of the reparameterized Chebyshev graph convolutional branch includes:
[0123] Normalize the output data X out of the heterogeneous spatial convolutional block. Treat each pixel as a node in the graph structure; calculate the similarity between all nodes using cosine similarity to obtain the vector similarity matrix M. Let x i and x j be the feature vectors of pixels i and j respectively. Each element of the vector similarity matrix M is:
[0124]
[0125] Set a threshold T for the matching indicator to filter the vector similarity matrix M and generate the adjacency matrix A. Each element of A is expressed as:
[0126]
[0127] Perform degree mapping on the adjacency matrix to obtain the degree matrix D, , and construct the graph Laplacian matrix L according to the degree matrix D and the adjacency matrix A:
[0128]
[0129] To enable the nodes to retain their own features, add self-loops to obtain the updated graph Laplacian matrix :
[0130]
[0131] Use the Chebyshev iteration method to calculate the updated graph Laplacian matrix to obtain Chebyshev polynomials of multiple orders:
[0132]
[0133] where T 0 represents the zero-order Chebyshev polynomial, and its value is fixed at 1; represents the first-order Chebyshev polynomial, which is equal to the Laplacian matrix ; represents the k-order Chebyshev polynomial, which is calculated through a recursive relationship; input the iterative calculation result into the reparameterized weight calculator to obtain the reparameterized Chebyshev coefficient wk :
[0134]
[0135] where x j is the value of the Chebyshev node calculated during initialization, and λ j is a learnable parameter; finally, the output of the reparameterized Chebyshev graph convolutional branch is obtained, which is the second output Y:
[0136]
[0137] where X in is the input feature matrix, which contains the features of the nodes in the graph; Y ∈ R N×C , N is the number of pixels, and C is the dimension of the output features. Y undergoes multi-order feature propagation, providing rich feature representations for the classification task of hyperspectral images.
[0138] S7. Use the dual-branch integration module to fuse the first output Y M and the second output Y to obtain the dual-branch fusion feature, and input it into the classifier to obtain the final classification result.
[0139] In this step, as Figure 6 shown, first, perform pooling (reducing the dimension of the features while retaining key feature information), the first layer of convolution (extracting local spatial features), the first activation function (introducing non-linearity), the second layer of convolution (further refining the feature representation to ensure richer output information), and the second activation function on the first output Y M of the band selection enhanced bidirectional Mamba branch and the second output Y of the reparameterized Chebyshev graph convolutional branch respectively to generate channel attention weights.
[0140] After that, perform element-wise multiplication of the output features of each branch with the corresponding channel attention weights, and then fuse the weighted features to obtain the dual-branch fusion feature. The fused feature has complementary information from both branches and has stronger classification representation ability.
[0141] Finally, input the dual-branch fusion feature into the classifier to perform the classification task and obtain the final classification result.
[0142] In the embodiments of the present invention, by enhancing the bidirectional Mamba model through band selection to capture the long-range dependencies in the spatial and spectral dimensions, and at the same time utilizing the feature propagation ability of reparameterized Chebyshev graph convolution in the non-Euclidean space, the complementarity between spatial features and spectral features is deeply mined, significantly improving the classification accuracy. Moreover, the present invention significantly reduces the computational cost by leveraging the linear complexity of the Mamba model and the efficient graph structure modeling of reparameterized Chebyshev graph convolution. Compared with the quadratic complexity of the popular transformer model, the Mamba model can effectively process high-resolution data, and at the same time avoids explicit feature decomposition through the recursive approximation of Chebyshev polynomials, greatly reducing resource and time consumption, and improving the efficiency and stability when processing large-scale hyperspectral data.
[0143] Compared with the prior art, the present invention has the following advantages:
[0144] 1. When existing methods process hyperspectral images, due to the independent processing of the spatial and spectral features of the images, the complex relationship between the two cannot be deeply explored, and the fusion of the spatial scale and spectral scale usually adopts simple linear methods, unable to capture the deep association between the two. To address this problem, the present invention proposes a classification method for hyperspectral images based on a dual-branch structure of band selection-enhanced bidirectional Mamba and reparameterized Chebyshev graph convolution. By enhancing the long-range dependence modeling ability of the bidirectional Mamba model for spatial and spectral information and the efficient modeling ability of the reparameterized Chebyshev graph convolution model for non-Euclidean space information through band selection, the spatial information and spectral features are deeply fused, effectively capturing the complex non-linear relationship between the two, and improving the accuracy of hyperspectral image classification.
[0145] 2. Existing methods often have low computational efficiency and high computational complexity when processing high-resolution and large-scale hyperspectral image data. To address this problem, the present invention effectively reduces the computational complexity and improves the computational efficiency by leveraging the linear complexity of the Mamba model and the efficient graph structure modeling of reparameterized Chebyshev graph convolution, expanding the applicability of the model on large-scale datasets.
[0146] 3. Existing hyperspectral classification models, due to adopting a single network architecture, are difficult to capture the similarity relationship between pixels in the image while processing spectral and spatial features at the same time, resulting in insufficient fusion ability of spectral and spatial information and poor modeling effect of pixel relationships. To address this problem, the present invention designs a dual-branch structure of band selection-enhanced bidirectional Mamba and reparameterized Chebyshev graph convolution to solve the problems of spectral feature and spatial feature extraction and image pixel similarity relationship modeling respectively. This dual-branch structure makes full use of the complementarity of the two models in the fusion stage, and can not only extract spectral and spatial features, but also comprehensively capture the complex relationships between image pixels.
[0147] In an exemplary embodiment, the present invention further provides an electronic device, which includes:
[0148] A processor;
[0149] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the above-mentioned hyperspectral image classification method are implemented.
[0150] In an exemplary embodiment, the present invention further provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the steps of the above-mentioned hyperspectral image classification method. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0151] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal device including the said element.
[0152] When referring to "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. in the specification, it indicates that the said embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such a feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge scope of those skilled in the relevant art.
[0153] It should be understood that the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0154] In the present invention, "at least one" means one or more, and "a plurality of" means two or more. "At least one of the following" or a similar expression means any combination of these items, including any combination of single item or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0155] It should be understood that in various embodiments of the present invention, the sequence numbers of the above - mentioned processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0156] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0157] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0158] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0159] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0160] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can also fully understand the present invention without the description of these details. Additionally, to avoid unnecessary confusion to the essence of the present invention, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail.
[0161] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A hyperspectral image classification method integrating Mamba and Chebyshev graph convolution, characterized in that: The following steps are involved: S1. Collect a hyperspectral image dataset and divide it into a training set and a validation set in proportion. S2. Preprocess the image data and divide each image into a predetermined number of image blocks with each pixel as the center. X ; S3, block the image X Input heterogeneous spatial convolution blocks for processing to obtain output data with different receptive fields X out ; S4. X out It is divided into two branches, one of which is processed by band-selectively enhancing the bidirectional Mamba branch to obtain the long-range dependency between the spatial and spectral features in the image; In step S4, the process of band selection enhancing bidirectional Mamba branch includes: The output data of the heterogeneous spatial convolution block X out An input position information encoder is used to add position information to the input data through position information encoding to provide context information of the spatial arrangement of pixels; The encoded data is input into the depth-separable convolutional layer for feature extraction to extract the spatial features of the image. X c ; The formula of band selection enhanced bidirectional Mamba branch is expressed as: ; Among them, Mamba spaf ( X c ) and Mamba spab ( X c ) represent the operation process of forward space Mamba and backward space Mamba respectively, Flip( X c ) represents the flip operation of data, X spaf is the output of the forward spatial Mamba, X spab The output of the backward space Mamba; S5. Generate channel importance scores for different spectral channels using the band-selective Mamba model S , and weighted processing is performed with the output of the band-selective enhanced bidirectional Mamba branch to obtain the first output Y M ; In step S5, the processing process of the band selection Mamba model includes: Swapping the output of depthwise separable convolutional layers X c The spatial and spectral dimensions of X r , ; Will X r After the forward spectral Mamba and the backward spectral Mamba operations, the spectral features are added together. X b : ; Spectral characteristics X b Perform pooling, fully connected layers, and sigmoid activation functions to generate channel importance scores S , S Indicates the importance of different spectral channels: ; The output and channel importance scores of the bidirectional Mamba branch are enhanced by band selection S Multiply them together to get the weighted features Y spaf and Y spab : ; ; The weighted features are fused through convolutional layers and concatenation operations to obtain the first output Y M : ; Y M It is a feature representation that contains the long-range dependency between the forward and backward spectral information and the spatial information in the image; S6. X out The other branch inputs the re-parameterized Chebyshev graph convolution branch for processing to obtain the similarity relationship between different pixels in the image and obtain the second output Y ; In step S6, the process of re-parameterizing the Chebyshev graph convolution branch includes: The output data of the heterogeneous spatial convolution block X out Perform layer normalization, treat each pixel as a node in the graph structure; use cosine similarity to calculate the similarity between all nodes to obtain the vector similarity matrix M ; Match indicator setting threshold T , the vector similarity matrix M Filter and generate an adjacency matrix A ; Perform degree mapping on the adjacency matrix to obtain the degree matrix D , according to the degree matrix D and the adjacency matrix A Constructing the graph Laplacian matrix L ; In order to make the nodes retain their own characteristics, the graph Laplacian matrix L Add self-loops to get the updated graph Laplacian matrix ; Use Chebyshev iteration to update the graph Laplacian matrix Perform calculations to obtain Chebyshev polynomials of multiple orders; express k The Chebyshev polynomial of order is calculated through recursive relations; the calculation result is input into the reparameterized weighter to obtain the reparameterized Chebyshev coefficient w k : The output of the reparameterized Chebyshev graph convolution branch, i.e. the second output Y It is expressed as: ; in X in is the input feature matrix, which contains the features of the nodes in the graph; Y ∈ R N×C , N is the number of pixels, C is the dimension of the output feature, K is the order of the Chebyshev polynomial; S7, using the dual-branch integrated module to output the first Y M and the second output Y Fusion is performed to obtain dual-branch fusion features, which are input into the classifier to obtain the final classification result; The step S7 specifically includes: Enhanced bidirectional Mamba branch first output for band selection Y M and the second output of the reparameterized Chebyshev graph convolution branch Y Perform pooling, first-layer convolution, first activation function, second-layer convolution, and second activation function operations respectively to generate channel attention weights; The output feature of each branch is multiplied element-by-element by the corresponding channel attention weight, and then the weighted features are fused to obtain the dual-branch fusion feature; The dual-branch fusion features are input into the classifier, the classification task is performed, and the final classification result is obtained.
2. The hyperspectral image classification method according to claim 1, characterized in that: The step S1 specifically includes: Collect multiple different public hyperspectral image datasets, with different spectral band numbers, coverage, image sizes, and ground sampling distances between different datasets; The collected hyperspectral image dataset is divided into a training set and a validation set in a ratio of 11:1; the training set is used for model training, and the validation set is used to verify the model effect; the model with the best performance on the validation set is saved as the final model.
3. The hyperspectral image classification method according to claim 1, characterized in that: The step S2 specifically includes: Divide each image into a predetermined number of image blocks centered on each pixel X , as the input of the model; input image block X ∈ R B×Cin×H×W ,in B Indicates the batch size, C in is the number of input spectral channels, H , W are the height and width of the image block.
4. The hyperspectral image classification method according to claim 1, characterized in that: In step S3, the processing process of the heterogeneous spatial convolution block includes: Use a 1×1 convolution to patch the input image. X Mapping to high-dimensional embedding space to generate high-dimensional features X emb : ; in X emb ∈ R B×Cemb×H×W , C emb Set to 256, indicating the number of spectral channels after embedding; For high-dimensional features X emb Perform 1×1 convolution and 3×3 convolution operations, add and fuse the two convolution results element by element to obtain a multi-scale feature combination, and then use a 1×1 convolution to restore the number of spectral channels to the number of spectral channels of the original input. C in , the operation is as follows: ; in X out ∈ R B×Cin×H×W , 1×1 convolution operation is used to extract local fine-grained features of each pixel, and 3×3 convolution operation is used to extract coarse-grained features larger than the local receptive field.
5. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when loaded and executed by the processor, implement the method according to any one of claims 1 to 4.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Hyperspectral image classification method and device based on spatial-spectral double-branch convolutional network
CN115249332A
Multimode self-supervised mixed Mangbar hyperspectral image classification method
CN119007024A