Hyperspectral image classification method and system, electronic equipment and storage medium
By building a pre-trained network, combining the space and channel reconstruction convolution module, the cross-scale aggregation module and the Kansformer module, the problem of redundant feature extraction and calculation in hyperspectral image classification is solved, and high-precision and low-cost image classification effect is achieved.
Patent Information
- Application Number
- CN202510114659.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing hyperspectral image classification methods may extract redundant features during feature extraction, increasing calculation costs and reducing classification accuracy. In addition, the traditional Transformer model has a large amount of calculation during high-dimensional data processing, which affects classification performance.
A hyperspectral image classification method is adopted to build a pre-trained network, including the spatial and channel reconstruction convolution module, the cross-scale aggregation module and the Kansformer module, which reduces redundant features, extracts abstract space-spectral features, and enhances feature representation through attention mechanisms and KAN networks.
It improves the classification accuracy of hyperspectral images, reduces the computational cost, and enhances the generalization ability and feature learning efficiency of the model.
Smart Images

Figure CN120014356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification, and in particular to a hyperspectral image classification method, system, electronic equipment and storage medium. Background Art
[0002] Hyperspectral image classification is crucial in the field of remote sensing because it has rich spectral and spatial information. Convolutional neural networks are capable of adaptive feature extraction, but have difficulties in deep semantic representation and have high computational costs as the number of layers increases. Although Transformers perform well in capturing high-level semantics, they face significant computational overhead in hyperspectral image classification, especially when faced with high-dimensional data. In addition, during model training, it is difficult to accurately identify different land cover categories from limited labeled samples, which affects the classification accuracy of the model.
[0003] In recent years, a variety of methods have been proposed to address the challenges in hyperspectral image classification. Traditional techniques such as support vector machines, k-nearest neighbor algorithms, and decision trees have shown some potential in hyperspectral data classification. However, these methods are often unable to adaptively extract abstract intrinsic features in the scene, resulting in reduced accuracy when classifying high-dimensional hyperspectral data. To overcome these limitations, more and more deep learning models have been widely used in hyperspectral image classification tasks.
[0004] At present, for hyperspectral image classification tasks, some models may extract many redundant features during the feature extraction process, which not only increases the computational cost but also reduces the overall classification accuracy. While other models may also affect the classification performance by enhancing the spectral-spatial features when extracting them. In addition, traditional Transformer models usually consume a lot of computing resources, which not only slows down the processing speed but also complicates the task of accurate classification. Summary of the invention
[0005] The purpose of the present invention is to provide a hyperspectral image classification method, system, electronic device and storage medium, which can improve the classification accuracy of hyperspectral images.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A hyperspectral image classification method, comprising:
[0008] Acquire a hyperspectral training data set; the training data set includes training images and corresponding labels;
[0009] Constructing a pre-trained network, and inputting the training data set into the pre-trained network for parameter optimization to obtain a trained image classification model; the pre-trained network includes a space and channel reconstruction convolution module, a cross-scale aggregation module and a Kansformer module; the space and channel reconstruction convolution module includes a space redundant unit and a channel redundant unit; the cross-scale aggregation module includes a channel attention module, a fusion convolution module and a space attention module; the Kansformer module includes a KAN network;
[0010] The image classification model is used to identify the hyperspectral image to be detected, and a hyperspectral image classification result is obtained.
[0011] Optionally, the process of inputting the training data set into the pre-training network for parameter optimization specifically includes:
[0012] Processing the training image input into a spatial and channel reconstruction convolution module to reduce redundant features and extract abstract spatial-spectral features;
[0013] The spatial-spectral features are input into a cross-scale aggregation module for processing, the spatial and spectral features of key information are emphasized, and key feature information is obtained;
[0014] The key feature information is input into the Kansformer module for processing to obtain feature enhancement data;
[0015] Parameters are adjusted based on the labels corresponding to the training images to obtain a trained image classification model.
[0016] Optionally, the training image input space and channel reconstruction convolution module are processed to reduce redundant features and extract abstract spatial-spectral features, specifically including:
[0017] In the spatial redundancy unit, the separation-reconstruction strategy is used to process the training image. First, the extracted feature map is set to X∈R B×H×W×C , where B is the batch size, C is the number of channels, H and W are the height and width of the space respectively, and the feature map X is normalized by subtracting the mean μ and dividing by the standard deviation ξ, as follows:
[0018]
[0019] Where μ and ξ are the mean and standard deviation of the feature map X, respectively. represents a set positive constant used to ensure the stability of the value during the division process, and γ and β are learnable affine transformation parameters;
[0020] Among them, the training parameter γ is used to capture the pixel variance in each batch and channel space, and the normalized correlation weight Wγ It is used to reflect the importance of different feature maps and is derived from the following formula:
[0021]
[0022] After getting the relevant weight W γ Then it is mapped to the range of 0 to 1 through the Sigmoid function. The formula is:
[0023] X γ = Gate(Sigmoid(W γ (GN(X))))
[0024] The mapped weight values are subjected to a gating operation, and the gating threshold is set to 0.5. The weights exceeding the threshold are assigned a value of 1 to form an informative weight W1, and the weights not exceeding the threshold are assigned a value of 0 to form an uninformative weight W2.
[0025] Then, let the feature map X be multiplied by the two weighted features, weight W1 and weight W2, to obtain the key features with information and redundant features with less information
[0026] Finally, using the crossover reconfiguration operation, the key features with information and redundant features with less information Fully integrate and finally generate a spatially refined feature map X ω ;
[0027] In the channel redundancy unit, the separation-conversion-fusion strategy is used to transform the feature map X ω To process, first perform a separation operation, and convert the feature map X ω It is divided into two parts: one contains αC channels and the other contains (1-α)C channels, where α is an experimentally set hyperparameter with a value of 0.5. Then a 1×1 convolution kernel is applied to compress the number of channels in each part to obtain the first separation feature X up and the second separation feature X low ;
[0028] Secondly, the conversion operation is performed, using group convolution and point convolution to transform the first separation feature X up Perform conversion processing, add and merge the output results of group convolution and point convolution respectively to obtain the first output feature Y1; use point convolution to convert the second separation feature X low Perform the conversion and compare the output result with X low Perform splicing to obtain the second output feature Y2;
[0029] Finally, a fusion operation is performed, and a simplified SKNet method is used to adaptively combine Y1 and Y2; the simplified SKNet method first uses global average pooling to integrate global spatial information and channel-level statistical information to obtain the first pooling feature S1 and the second pooling feature S2; then, the Softmax function is applied to S1 and S2 to calculate the first feature weight vector β1 and the second feature weight vector β2; finally, the fused output is calculated as: Y=β1Y1+β2Y2, where Y is the feature after channel refinement, which is used to represent the abstract spatial-spectral feature.
[0030] Optionally, the spatial-spectral features are input into a cross-scale aggregation module for processing, emphasizing the spatial and spectral features of key information to obtain key feature information, specifically including:
[0031] First, Y is dimensionalized by a 1×1 convolution to reduce the number of channels:
[0032] X 1×1 =Conv2D 1×1 (Y)∈R d×H×W
[0033] Among them, d represents the number of channels after 1×1 convolution;
[0034] In the channel attention module, the mean and maximum of each channel are first calculated by global average pooling and global maximum pooling, which are expressed as:
[0035]
[0036] Where i,j=1,2,...,C, C is the number of channels, H and W are the height and width of the space respectively;
[0037] Then, the two pooling results are transformed through two convolutional layers respectively to obtain the channel attention weights. First, the number of channels is reduced, and then nonlinear transformation is performed through activation functions ReLU and Sigmoid to generate a channel-level attention map. Finally, the number of channels is restored to the original value, specifically:
[0038] A avg =Conv2D 1×1 (ReLU(Conv2D 1×1 (X avg )))
[0039] A max =Conv2D 1×1 (ReLU(Conv2D 1×1 (X max )))
[0040] Among them, the generated A avg and Amax Represents the attention weight of each channel;
[0041] Finally, the calculated attention value is equal to X 1×1 Multiply element by element to get the weighted output feature map, expressed as:
[0042] X CAM =X 1×1 ⊙σ(A avg +A max )
[0043] In the formula, ⊙ represents element-by-element multiplication, σ represents the Sigmoid activation function, and X CAM Represents the output value of the channel attention mechanism;
[0044] In the fused convolution module, convolution kernels of sizes 3×3, 5×5, and 7×7 are used for superposition, which is expressed as:
[0045] X FCM =X 3×3 +X 5×5 +X 7×7
[0046] In the spatial attention module, the two feature maps obtained by the fused convolution module are weighted according to the spatial dimension. The shape of the two feature maps obtained by the fused convolution module in the spatial position is 1×H×W, which is expressed as:
[0047]
[0048] The two feature maps obtained by the fused convolution module are connected to form a new feature map:
[0049] X concat =[X avg ,X max ]∈R 2×H×W
[0050] Then, the spatial attention is calculated through the convolution operation:
[0051] A spatial =σ(Conv2D 7×7 (X concat ))
[0052] Among them, σ is the Sigmoid activation function, output A spatial ∈R 1×H×W represents the attention weight of each spatial position;
[0053] Finally, the output after spatial attention weighting is:
[0054] X SAM=X⊙A spatial
[0055] Therefore, the channel attention output and spatial attention output are weighted, and finally the output result is obtained through 1×1 convolution: CSAM =Conv2D 1×1 (X CAM +X SAM ).
[0056] Optionally, the key feature information is input into a Kansformer module for processing to obtain feature enhancement data, specifically including:
[0057] In the Kansformer module, batch normalization is used to replace layer normalization, and the KAN network is used to capture the complex differences between different ground object categories to obtain feature enhanced data.
[0058] The present invention also provides a hyperspectral image classification system, comprising:
[0059] A data acquisition module is used to obtain a hyperspectral training data set; the training data set includes training images and corresponding labels;
[0060] A pre-training module, used to construct a pre-training network, and input the training data set into the pre-training network for parameter optimization to obtain a trained image classification model; the pre-training network includes a space and channel reconstruction convolution module, a cross-scale aggregation module and a Kansformer module; the space and channel reconstruction convolution module includes a space redundancy unit and a channel redundancy unit; the cross-scale aggregation module includes a channel attention module, a fusion convolution module and a space attention module; the Kansformer module includes a KAN network;
[0061] The image classification module is used to use the image classification model to identify the hyperspectral image to be detected and obtain a hyperspectral image classification result.
[0062] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the hyperspectral image classification method described above.
[0063] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the hyperspectral image classification method as described above when executed by a processor.
[0064] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0065] The present invention discloses a hyperspectral image classification method, system, electronic device and storage medium, the method comprising obtaining a hyperspectral training data set; the training data set comprises training images and corresponding labels; constructing a pre-trained network, and inputting the training data set into the pre-trained network for parameter optimization to obtain a trained image classification model; using the image classification model to identify the hyperspectral image to be detected to obtain a hyperspectral image classification result. The present invention can improve the classification accuracy of hyperspectral images by utilizing the abstract feature extraction capability of convolutional neural networks, the key information focusing capability of multi-branch attention mechanisms and the feature difference enhancement capability of Kansformer. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0067] Figure 1 This is a flow chart of the hyperspectral image classification method based on cross-scale aggregation and Kansformer;
[0068] Figure 2 This is a framework diagram of the hyperspectral image classification method based on cross-scale aggregation and Kansformer;
[0069] Figure 3 It is a spatial redundant unit framework diagram;
[0070] Figure 4 It is a framework diagram of a channel redundancy unit;
[0071] Figure 5 It is a diagram of the cross-scale aggregation module framework;
[0072] Figure 6 This is the framework diagram of the MLP and KAN deep learning model;
[0073] Figure 7 This is the framework diagram of the multi-head attention mechanism module;
[0074] Figure 8(a) is the actual distribution map of 14 types of features in the Botswana dataset;
[0075] Figure 8(b) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm SSFTT;
[0076] Figure 8(c) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm GAHT;
[0077] Figure 8(d) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm morphFormer;
[0078] Figure 8(e) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm DBCT;
[0079] Figure 8(f) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm MSSTT;
[0080] Figure 8(g) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm DCTN;
[0081] Figure 8(h) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm MASSFormer;
[0082] Figure 8(i) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm RDTN;
[0083] Figure 8(j) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm LSFAT;
[0084] Figure 8(k) is the overall classification accuracy graph of the Botswana dataset based on the comparison algorithm CSA-Kanformer;
[0085] Figure 9(a) is the real distribution map of 15 types of features in the Houston2013 dataset;
[0086] Figure 9(b) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm SSFTT;
[0087] Figure 9(c) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm GAHT;
[0088] Figure 9(d) is the overall classification accuracy of the Houston2013 dataset based on the comparison algorithm morphFormer;
[0089] Figure 9(e) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm DBCT;
[0090] Figure 9(f) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm MSSTT;
[0091] Figure 9(g) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm DCTN;
[0092] Figure 9(h) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm MASSFormer;
[0093] Figure 9(i) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm RDTN;
[0094] Figure 9(j) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm LSFAT;
[0095] Figure 9(k) is the overall classification accuracy graph of the Houston2013 dataset based on the comparison algorithm CSA-Kanformer;
[0096] Figure 10(a) is the real distribution map of 16 types of objects in the WHU-Hi-HanChuan dataset;
[0097] Figure 10(b) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm SSFTT;
[0098] Figure 10(c) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm GAHT;
[0099] Figure 10(d) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm morphFormer;
[0100] Figure 10(e) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm DBCT;
[0101] Figure 10(f) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm MSSTT;
[0102] Figure 10(g) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm DCTN;
[0103] Figure 10(h) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm MASSFormer;
[0104] Figure 10(i) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm RDTN;
[0105] Figure 10(j) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm LSFAT;
[0106] Figure 10(k) is the overall classification accuracy graph of the WHU-Hi-HanChuan dataset based on the comparison algorithm CSA-Kanformer;
[0107] Figure 11(a) is the real distribution map of 22 types of features in the WHU-Hi-HongHu dataset;
[0108] Figure 11(b) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm SSFTT;
[0109] Figure 11(c) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm GAHT;
[0110] Figure 11(d) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm morphFormer;
[0111] Figure 11(e) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm DBCT;
[0112] Figure 11(f) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm MSSTT;
[0113] Figure 11(g) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm DCTN;
[0114] Figure 11(h) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm MASSFormer;
[0115] Figure 11(i) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm RDTN;
[0116] Figure 11(j) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm LSFAT;
[0117] Figure 11(k) is the overall classification accuracy graph of the WHU-Hi-HongHu dataset based on the comparison algorithm CSA-Kanformer;
[0118] Figure 12(a) is a schematic diagram of OA of 10 algorithms using 2%-10% training sample ratio on the Botswana dataset;
[0119] Figure 12(b) is a schematic diagram of OA of 10 algorithms using 2%-10% training sample ratio on the Houston2013 dataset;
[0120] Figure 12(c) is a schematic diagram of OA of 10 algorithms using 2%-10% training sample ratio on the WHU-Hi-HanChuan dataset;
[0121] Figure 12(d) is a schematic diagram of OA of the 10 algorithms using 2%-10% training sample ratio on the WHU-Hi-HongHu dataset;
[0122] Figure 13(a) is a schematic diagram of the overall classification accuracy of the CSA-Kanformer algorithm using different block sizes on the Botswana dataset;
[0123] Figure 13(b) is a schematic diagram of the overall classification accuracy of the CSA-Kanformer algorithm using different block sizes on the Houston 2013 dataset;
[0124] Figure 13(c) is a schematic diagram of the overall classification accuracy of the CSA-Kanformer algorithm using different block sizes on the WHU-Hi-HanChuan dataset;
[0125] Figure 13(d) is a schematic diagram of the overall classification accuracy of the CSA-Kanformer algorithm using different block sizes on the WHU-Hi-HongHu dataset;
[0126] Figure 14(a) is a schematic diagram of the overall classification accuracy of the CSA-Kanformer algorithm using different learning rates on the Botswana dataset;
[0127] Figure 14(b) is a schematic diagram of the overall classification accuracy of the CSA-Kanformer algorithm using different learning rates on the Houston 2013 dataset;
[0128] Figure 14(c) is a schematic diagram of the overall classification accuracy of the CSA-Kanformer algorithm using different learning rates on the WHU-Hi-HanChuan dataset;
[0129] Figure 14(d) is a schematic diagram of the overall classification accuracy of the CSA-Kanformer algorithm using different learning rates on the WHU-Hi-HongHu dataset. DETAILED DESCRIPTION
[0130] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0131] The purpose of the present invention is to provide a hyperspectral image classification method, system, electronic device and storage medium, which can improve the classification accuracy of hyperspectral images.
[0132] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0133] like Figure 1 As shown, the present invention provides a hyperspectral image classification method, comprising:
[0134] Step 100: Obtain a hyperspectral training data set; the training data set includes training images and corresponding labels.
[0135] Step 200: construct a pre-trained network, and input the training data set into the pre-trained network for parameter optimization to obtain a trained image classification model; the pre-trained network includes a space and channel reconstruction convolution module, a cross-scale aggregation module and a Kansformer module; the space and channel reconstruction convolution module includes a space redundant unit and a channel redundant unit; the cross-scale aggregation module includes a channel attention module, a fusion convolution module and a space attention module; the Kansformer module includes a KAN network.
[0136] Step 300: using the image classification model to identify the hyperspectral image to be detected, and obtaining a hyperspectral image classification result.
[0137] In the above steps: first, the spatial and channel reconstruction convolution module is used to reduce redundant features by adopting spatial redundant units and channel redundant units, thereby reducing computational costs and extracting abstract spatial-spectral features; second, the cross-scale aggregation module is used to combine the fusion convolution module, channel attention mechanism and spatial attention mechanism. This novel design achieves efficient cross-scale feature aggregation, which can selectively emphasize the spatial and spectral features of key information, thereby enhancing feature representation and improving the accuracy of hyperspectral image classification; finally, the Kansformer module is used to replace layer normalization with batch normalization to enhance training stability, accelerate the convergence process and improve the generalization ability of the model. In addition, the KAN network is used to further capture the complex differences between different categories and optimize feature learning, thereby improving training efficiency and model performance.
[0138] As a specific implementation method, the specific description of each step is as follows:
[0139] 1. The image is a hyperspectral image.
[0140] 2. The spatial and channel reconstruction convolution module specifically includes:
[0141] The spatial and channel reconstruction convolutional module includes two key components: spatial redundancy unit and channel redundancy unit. In order to capture the spatial redundancy in the features, the spatial redundancy unit adopts a separation-reconstruction approach. The separation operation aims to distinguish useful feature maps from less useful feature maps based on the spatial structure, and the scaling factor of the group normalization layer helps to evaluate the information content of each feature map. Specifically, suppose there is an intermediate feature map X∈R H×W×C , where B is the batch size, C is the number of channels, H and W are the height and width dimensions of the space, respectively. The feature map X is normalized by subtracting the mean μ and dividing by the standard deviation ξ, as follows:
[0142]
[0143] Among them, μ and ξ are the mean and standard deviation of the feature map X, respectively. is a small positive constant used to ensure numerical stability during the division process, and γ and β are learnable affine transformation parameters.
[0144] The trainable parameter γ is used to capture the pixel variance in each batch and channel space, and the normalized correlation weight W γ It is derived from the following formula, which reflects the importance of different feature maps.
[0145]
[0146] The obtained relevant weight W γ It is then mapped to the range of 0 to 1 through the Sigmoid function, with the following formula:
[0147] X γ = Gate(Sigmoid(W γ (GN(X))))
[0148] The obtained correlation weights are then subjected to a gating operation with a threshold of 0.5. Weights above the threshold are assigned a value of 1, forming informative weights W1, while weights below the threshold are assigned a value of 0, forming uninformative weights W2. The input feature X is then multiplied by these two weighted features to obtain the informative key features and redundant features with less information
[0149] Finally, a cross reconfiguration operation is applied to fully fuse the two weighted informative features and promote the information flow between them, thus ultimately generating a spatially refined feature map X ω .
[0150] Channel Redundancy Unit In order to take into account the possibility of spatial redundancy during the spatial refinement process, this module adopts a separation-transformation-fusion strategy to replace the standard convolution operation.
[0151] The separation operation refines the input feature X ω It is divided into two parts: one contains αC channels and the other contains (1-α)C channels, where α is an experimentally set hyperparameter with a value of 0.5. Then, a 1×1 convolution kernel is applied to compress the number of channels in each part, and the result is X up and X low .
[0152] The transformation operation uses a rich feature extractor on X up The outputs of these operations are combined by addition to form Y1. At the same time, point convolution is also applied to X. low , as a complementary operation, the result is concatenated with the original input to form Y2.
[0153] The fusion operation uses a simplified SKNet method to adaptively combine Y1 and Y2. First, the method uses global average pooling to integrate global spatial information and channel-level statistical information to obtain pooled features S1 and S2. Then, the Softmax function is applied to S1 and S2 to calculate the feature weight vectors β1 and β2. Finally, the fused output is calculated as Y = β1Y1 + β2Y2, where Y represents the channel-refined features.
[0154] 3. The cross-scale aggregation module specifically includes:
[0155] The cross-scale aggregation module consists of three modules: channel attention module, fusion convolution module and spatial attention module. Its main purpose is to optimize feature selection by fusing multi-scale information and utilizing the attention mechanism.
[0156] The core task of the channel attention module is to assign different weights to different channels according to the global information of each channel (calculated by average pooling and max pooling). In this way, the module is able to focus on important channel features while suppressing less important channels.
[0157] Assume that the input feature map X∈R C×H×W , where C is the number of channels, H and W are the height and width respectively. First, the global description (mean and maximum) of each channel is calculated by global average pooling and global maximum pooling, as shown below:
[0158]
[0159] Then, the two pooling results are transformed through two convolutional layers respectively to obtain the channel attention weights. First, the number of channels is reduced, and then nonlinear transformation is performed through activation functions ReLU and Sigmoid to generate a channel-level attention map, and finally the number of channels is restored to the original value. The formula is as follows:
[0160] A avg =Conv2D 1×1 (ReLU(Conv2D 1×1 (X avg )))
[0161] A max =Conv2D 1×1 (ReLU(Conv2D 1×1 (X max )))
[0162] The generated A avg and A max Represents the attention weight of each channel.
[0163] Finally, the calculated attention value is multiplied element by element with the original feature map to obtain the weighted output feature map. The formula is as follows:
[0164] X CAM =X⊙σ(A avg +A max )
[0165] Where ⊙ represents element-by-element multiplication, σ represents the Sigmoid activation function, and X CAM Represents the output value of the channel attention mechanism.
[0166] The fused convolution module uses convolution kernels of different sizes (3×3, 5×5, 7×7) to process the input and fuse features of different scales. In this way, the fused convolution module can capture feature information of different spectral-spatial scales. Input feature map X∈R C×H×W Dimensionality is performed through 1×1 convolution to reduce the number of channels:
[0167] X 1×1 =Conv2D 1×1 (X)∈R d×H×W
[0168] Then, these multi-scale features are superimposed through convolutions of different sizes. The formula is as follows:
[0169] X FCM =X 3×3 +X 5×5 +X 7×7
[0170] The purpose of the spatial attention module is to weight the feature map according to the spatial dimension, emphasizing the locations with important spatial information. By calculating the weight of each spatial location, the spatial attention module enhances the model's attention to key spatial regions. Input feature map X FCM ∈R C×H×W After average pooling and maximum pooling, two different feature maps are generated. The shape of these two feature maps in space is 1×H×W. The formula is as follows:
[0171]
[0172] These two feature maps are then concatenated together to form a new feature map:
[0173] X concat =[X avg ,X max ]∈R 2×H×W
[0174] Then, the spatial attention is calculated through the convolution operation:
[0175] A spatial =σ(Conv2D 7×7 (X concat ))
[0176] Among them, σ is the Sigmoid activation function, output A spatial ∈R 1×H×W represents the attention weight at each spatial position.
[0177] Finally, the output after spatial attention weighting is:
[0178] X SAM =X⊙A spatial
[0179] In summary, the channel attention output and the spatial attention output are weighted and finally the final output is obtained through a 1×1 convolution.
[0180] X CSAM =Conv2D 1×1 (X CAM +X SAM )
[0181] 4. The Kansformer module specifically includes:
[0182] The Kansformer module uses batch normalization instead of layer normalization to enhance the stability of training, accelerate the convergence process and improve the generalization ability of the model. In addition, the KAN network is used to further capture the complex differences between different categories and optimize feature learning, thereby improving training efficiency and model performance.
[0183] Multilayer Perceptrons, also known as fully connected feedforward neural networks, are a key building block in the foundational components. Multilayer Perceptrons play a vital role in the operation of modern deep learning models and are able to effectively approximate nonlinear functions, and KANs provide an alternative approach based on the Kolmogorov-Arnold theorem. Unlike multilayer perceptrons that use fixed activation functions, KANs apply learnable activation functions on the edges and use spline functions instead of linear weight matrices. A multilayer perceptron with N layers can be represented as a series of transformations involving an affine transformation matrix W and a nonlinear activation function σ. This relationship can be represented by the following mathematical formula:
[0184]
[0185] On the other hand, an N-layer KAN can be represented as follows, and the output mapping can be defined as:
[0186]
[0187] Among them, i represents the i-th layer of the KAN model. Let l in and l out denote the input and output dimensions of each KAN layer respectively, then ψ is given by l in × out It consists of a one-dimensional learnable activation function φ:
[0188] Ψ={φ i,j},i=1,2,...,l in ,j=1,2,...,l out
[0189] In the KAN model, the calculation results from the kth layer to the k+1th layer can be expressed in matrix form as follows:
[0190]
[0191] The Transformer architecture performs well due to its core multi-head self-attention module. The self-attention mechanism used in this module can effectively capture the relationship between feature sequences. In order to learn multiple meanings, three learnable weight matrices W are predefined Q , W K and W V . These matrices linearly map the tokens to form a three-dimensional invariant matrix of query Q, key K, and value V. The attention score is calculated over all Q and K elements, and the weights of these scores are determined by the Softmax function. In summary, the self-attention mechanism can be expressed as the following formula:
[0192]
[0193] Among them, d K is the dimension of K.
[0194] The multi-head self-attention module consists of multiple sets of weight matrices to map Q, K, and V, and applies the same operation to calculate the multi-head attention value. Then, the attention output of each head is concatenated. This process can be expressed by the following formula:
[0195] MSA(Q,K,V)=Concat(SA1,SA2,...,SA h )H
[0196] Where h represents the number of heads, H represents the parameter matrix, and where d W Indicates the number of tokens.
[0197] The weight matrix obtained from the previous step is then fed into the KAN layer. After the KAN layer, batch normalization is applied to alleviate the problems associated with gradient explosion and gradient vanishing, leading to a more efficient training process.
[0198] A new algorithm proposed in the present invention is a hyperspectral image classification method based on cross-scale aggregation and Kansformer for hyperspectral remote sensing image classification tasks. Convolutional neural networks can improve feature extraction, but have difficulties in deep semantic representation and high computational cost. Although Transformers perform well in capturing high-level semantics, they face significant overhead in hyperspectral image classification, especially when facing high-dimensional data. In order to combine the advantages of convolutional neural networks in local feature extraction with the advantages of Transformers in long-distance modeling, a CSA-Kansformer model is proposed in this embodiment, which introduces three key modules to improve performance and efficiency. First, a spatial and channel reconstruction convolution module is proposed to reduce redundant features by using spatial redundant units and channel redundant units, thereby reducing computational cost and extracting abstract spatial-spectral features. Secondly, an innovative cross-scale aggregation module is constructed, combining a fusion convolution module, a channel attention mechanism and a spatial attention mechanism. This novel design can efficiently perform multi-scale feature aggregation, selectively emphasize key spatial and spectral information, thereby enhancing feature representation and improving the accuracy of hyperspectral image classification tasks. Finally, a Kansformer model was constructed, replacing layer normalization with batch normalization to improve training stability, accelerate convergence, and improve generalization in hyperspectral image classification. In addition, the KAN network was used to further capture the complex differences between different categories and optimize feature learning, thereby improving training efficiency and model performance.
[0199] Based on the above technical solution, the following is provided: Figure 1-Figure 1 4 is an embodiment shown in FIG.
[0200] Figure 1 This is a flow chart of the hyperspectral image classification method based on cross-scale aggregation and Kansformer. The process includes first using the spatial and channel reconstruction convolution module to reduce redundant features, thereby reducing computational costs and extracting abstract spatial-spectral features; secondly, the use of the cross-scale aggregation module can selectively emphasize the spatial and spectral features of key information, thereby enhancing feature representation; finally, the Kansformer module uses batch normalization instead of layer normalization to enhance the stability of training, accelerate the convergence process and improve the generalization ability of the model. In addition, the KAN network is further used to capture the complex differences between different types of objects, enhance feature representation, and thus improve training efficiency and the generalization performance of the model.
[0201] Figure 2 This is a framework diagram of a hyperspectral image classification method based on cross-scale aggregation and Kansformer. The basic framework includes a spatial and channel reconstruction convolution module, a cross-scale aggregation module, and a Kansformer module.
[0202] Figure 3 It is a spatial redundant unit structure diagram, which adopts the separation and reconstruction method. The separation operation aims to distinguish the feature map with rich information and the less useful feature map according to the spatial structure. The reconstruction operation adds the features with rich information to the features with less information to generate more informative features, thereby saving space.
[0203] Figure 4 This is a diagram of the channel redundant unit structure. Taking into account that spatial redundancy may still exist during the spatial refinement process, this module adopts a separation-conversion-fusion strategy, replacing the standard convolution operation with three operations: separation, conversion, and fusion.
[0204] Figure 5 It is a structural diagram of the cross-scale aggregation module. Its main purpose is to optimize feature selection by fusing multi-scale information and utilizing the attention mechanism. It contains three small modules: channel attention module, fusion convolution module and spatial attention module.
[0205] Figure 6 It is the structure diagram of MLP and KAN deep learning models, showing their different algorithm structures.
[0206] Figure 7 It is a diagram of the multi-head attention module structure. Q, K, and V represent query, key, and value matrices, respectively. MatMul represents matrix multiplication.
[0207] Figure 8(a) is a real hyperspectral image data. The Botswana dataset was acquired by NASA's EO-1 satellite over the Okavango Delta in Botswana. The dataset consists of a 7.7 km long data strip collected by the Hyperion sensor at a resolution of 30 million pixels in 242 bands, covering the 400-2500nm spectral part of the 10nm window. The absorption band is removed, leaving 145 spectral bands. The dataset can be divided into 14 categories. Figure 8(b)-Figure 8(k) The classification results of 10 other algorithms on the Botswana dataset are shown, including SSFTT, GAHT, morphFormer, DBCT, MSSTT, DCTN, MASSFormer, RDTN, LSFAT and CSA-Kansformer. By comparing these classification results, it can be seen that the proposed CSA-Kansformer algorithm (i.e., Figure 8(k)) shows the best classification effect.
[0208] Figure 9(a) is a real hyperspectral image data. The Houston2013 dataset was collected using the CASI1500 sensor in and around the University of Houston, Texas, USA. The image resolution of this dataset is 349×1905 pixels, containing 144 spectral bands with a wavelength range from 0.38 to 1.05 microns (μm). It includes data from 15 different ground feature categories. Figure 9(b)-Figure 9(k) The classification results of 10 other algorithms on the Houston2013 dataset are shown, including SSFTT, GAHT, morphFormer, DBCT, MSSTT, DCTN, MASSFormer, RDTN, LSFAT and CSA-Kansformer. By comparing these classification results, it can be seen that the proposed CSA-Kansformer algorithm (i.e., Figure 9(k)) shows the best classification effect.
[0209] Figure 10(a) is a real hyperspectral image data, the WHU-Hi-HanChuan dataset was collected in Hanchuan City, Hubei Province, China. The data was captured by a 17mm focal length Headwall Nano-Hyperspec sensor mounted on a Leica Aibot X6 drone V1. The weather was clear and sunny, with a temperature of about 30 degrees Celsius and a relative humidity of 70%. The surveyed area is an urban-rural junction where buildings, water and cultivated land are mixed, with seven main crops: strawberry, cowpea, soybean, sorghum, water spinach, watermelon and vegetables. The drone operated at an altitude of 250 meters, producing an image with a size of 1217×303 pixels, covering 274 spectral bands from 400 to 1000nm, with a spatial resolution of about 0.109 meters. Due to the low elevation angle of the sun in the afternoon, many shadow areas are included in the image. Figure 10(b)-Figure 10(k) The classification results of 10 other algorithms on the WHU-Hi-HanChuan dataset are shown, including SSFTT, GAHT, morphFormer, DBCT, MSSTT, DCTN, MASSFormer, RDTN, LSFAT and CSA-Kansformer. By comparing these classification results, it can be seen that the proposed CSA-Kansformer algorithm (i.e., Figure 10(k)) exhibits the best classification effect.
[0210] Figure 11(a) is a real hyperspectral image data. The WHU-Hi-HongHu dataset was collected from 16:23 to 17:37 on November 20, 2017 in Honghu City, Hubei Province, China, using a 17mm focal length Headwall Nano-Hyperspec sensor mounted on a DJI matrix600Pro drone. The weather during data collection was cloudy, with a temperature of about 8°C and a relative humidity of 55%. The studied area is a diverse agricultural landscape featuring a variety of crops, including different varieties of Chinese cabbage, Chinese cabbage, Brassica rapa, and Brassica rapa. The drone was operated at an altitude of 100 meters and captured images at a resolution of 940×475 pixels in 270 spectral bands from 400 to 1000nm, with a spatial resolution of about 0.043 meters. Figure 11(b)-Figure 11(k) The classification results of 10 other algorithms on the WHU-Hi-HongHu dataset are shown, including SSFTT, GAHT, morphFormer, DBCT, MSSTT, DCTN, MASSFormer, RDTN, LSFAT and CSA-Kansformer. By comparing these classification results, it can be seen that the proposed CSA-Kansformer algorithm (i.e., Figure 11(k)) exhibits the best classification effect.
[0211] Figure 12(a)-Figure 12(d)The changes in classification accuracy of ten algorithms are shown when using different proportions of training samples on four different datasets.
[0212] Reference Figure 13(a)-Figure 13(d) , the variation of overall classification accuracy under different block sizes is studied. In Figure 13(a), for the Botswana dataset, the best overall classification accuracy is obtained when the block size is 13×13. Figure 13(b) shows that for the Houston2013 dataset, the best classification effect is achieved when the block size is 11×11. In Figure 13(c), for the WHU-Hi-HanChuan dataset, the best classification accuracy is shown when the block size is 19×19. Finally, in Figure 13(d), for the WHU-Hi-HongHu dataset, the best classification accuracy is achieved when the block size is 17×17.
[0213] Reference Figure 14(a)-Figure 14(d) , the variation of overall classification accuracy under different learning rates is studied. In Figure 14(a), for the Botswana dataset, the best overall classification accuracy is achieved when the learning rate is 0.0004. In Figure 14(b), for the Houston2013 dataset, the best classification accuracy is achieved when the learning rate is 0.0003. Figure 14(c) shows that for the WHU-Hi-HanChuan dataset, the highest overall classification accuracy is achieved when the learning rate is 0.0002. In Figure 14(d), for the WHU-Hi-HongHu dataset, the best classification effect is also achieved when the learning rate is 0.0002.
[0214] It can be seen that this embodiment has the following beneficial effects:
[0215] (1) A spatial and channel reconstruction convolution module is introduced to reduce redundant information. The module consists of two key components: spatial redundancy unit and channel redundancy unit. First, the spatial-spectral features of the hyperspectral image are extracted through 3D convolution. Then, the spatial redundancy unit uses the "separation and reconstruction" technique to deal with spatial redundancy, while the channel redundancy unit uses the "separation-transformation-fusion" method to deal with channel redundancy. Finally, the extracted spatial-spectral features are further processed by 2D convolution.
[0216] (2) A cross-scale aggregation module is proposed, which consists of three parts: a fused convolution module, a channel attention module, and a spatial attention module. The fused convolution module extracts multi-scale information through three convolution operations of different sizes. Channel attention adaptively learns the importance of each channel and suppresses redundant channel information. Spatial attention improves the model's perception of important features by focusing on the key spatial regions of the input. Through this dual attention mechanism of channels and space, the cross-scale aggregation module can effectively reduce redundant information and improve the model's expressiveness and performance.
[0217] (3) A Kansformer module is proposed to improve the traditional Transformer model by replacing layer normalization with batch normalization. Batch normalization normalizes each feature in a batch of samples, thereby achieving more comprehensive feature extraction. In addition, the KAN network is used to further capture the complex differences between different categories and optimize feature learning, thereby improving training efficiency and the generalization performance of the model.
[0218] (4) An extensive series of experiments are conducted on four widely recognized benchmark hyperspectral image datasets: Botswana, Houston2013, WHU-Hi-HanChuan, and WHU-Hi-HongHu. The experimental results show that the proposed network performs comparable to some of the most advanced methods in this field.
[0219] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0220] This article uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only used to help understand the core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A hyperspectral image classification method, characterized in that: include: Obtain a hyperspectral training dataset; The training data set includes training images and corresponding labels; Constructing a pre-trained network, and inputting the training data set into the pre-trained network for parameter optimization to obtain a trained image classification model; the pre-trained network includes a space and channel reconstruction convolution module, a cross-scale aggregation module and a Kansformer module; the space and channel reconstruction convolution module includes a space redundancy unit and a channel redundancy unit; the cross-scale aggregation module includes a channel attention module, a fusion convolution module and a space attention module; The Kansformer module includes a KAN network; The image classification model is used to identify the hyperspectral image to be detected, and a hyperspectral image classification result is obtained.
2. The hyperspectral image classification method according to claim 1, characterized in that: The process of inputting the training data set into the pre-training network for parameter optimization specifically includes: Processing the training image input into a spatial and channel reconstruction convolution module to reduce redundant features and extract abstract spatial-spectral features; The spatial-spectral features are input into a cross-scale aggregation module for processing, the spatial and spectral features of key information are emphasized, and key feature information is obtained; The key feature information is input into the Kansformer module for processing to obtain feature enhancement data; Parameters are adjusted based on the labels corresponding to the training images to obtain a trained image classification model.
3. The hyperspectral image classification method according to claim 2, characterized in that: The training image is input into the spatial and channel reconstruction convolution module to reduce redundant features and extract abstract spatial-spectral features, specifically including: In the spatial redundancy unit, the separation-reconstruction strategy is used to process the training image. First, the extracted feature map is set to X∈R B×H×W×C , where B is the batch size, C is the number of channels, H and W are the height and width of the space respectively, and the feature map X is normalized by subtracting the mean μ and dividing by the standard deviation ξ, as follows: Where μ and ξ are the mean and standard deviation of the feature map X, respectively. represents a set positive constant used to ensure the stability of the value during the division process, and γ and β are learnable affine transformation parameters; Among them, the training parameter γ is used to capture the pixel variance in each batch and channel space, and the normalized correlation weight W γ It is used to reflect the importance of different feature maps and is derived from the following formula: After getting the relevant weight W γ Then it is mapped to the range of 0 to 1 through the Sigmoid function. The formula is: X γ =Gate(Sigmoid(W γ (GN(X)))) The mapped weight values are subjected to a gating operation, and the gating threshold is set to 0.
5. The weights exceeding the threshold are assigned a value of 1 to form an informative weight W1, and the weights not exceeding the threshold are assigned a value of 0 to form an uninformative weight W2. Then, let the feature map X be multiplied by the two weighted features, weight W1 and weight W2, to obtain the key features with information and redundant features with less information Finally, using the crossover reconfiguration operation, the key features with information and redundant features with less information Fully integrate and finally generate a spatially refined feature map X ω ; In the channel redundancy unit, the separation-conversion-fusion strategy is used to transform the feature map X ω To process, first perform a separation operation, and convert the feature map X ω It is divided into two parts: one contains αC channels and the other contains (1-α)C channels, where α is an experimentally set hyperparameter with a value of 0.
5. Then a 1×1 convolution kernel is applied to compress the number of channels in each part to obtain the first separation feature X up and the second separating feature X low ; Secondly, the conversion operation is performed, using group convolution and point convolution to transform the first separation feature X up Perform conversion processing, add and merge the output results of group convolution and point convolution respectively to obtain the first output feature Y1; use point convolution to convert the second separation feature X low Perform the conversion and compare the output result with X low Perform splicing to obtain the second output feature Y2; Finally, a fusion operation is performed, and a simplified SKNet method is used to adaptively combine Y1 and Y2; the simplified SKNet method first uses global average pooling to integrate global spatial information and channel-level statistical information to obtain the first pooling feature S1 and the second pooling feature S2; then, the Softmax function is applied to S1 and S2 to calculate the first feature weight vector β1 and the second feature weight vector β2; finally, the fused output is calculated as: Y=β1Y1+β2Y2, where Y is the feature after channel refinement, which is used to represent the abstract spatial-spectral feature.
4. The hyperspectral image classification method according to claim 2, characterized in that: The spatial-spectral features are input into the cross-scale aggregation module for processing, emphasizing the spatial and spectral features of key information, and obtaining key feature information, specifically including: First, Y is dimensionalized by a 1×1 convolution to reduce the number of channels: X 1×1 =Conv2D 1×1 (Y)∈R d×H×W Among them, d represents the number of channels after 1×1 convolution; In the channel attention module, the mean and maximum of each channel are first calculated by global average pooling and global maximum pooling, which are expressed as: Where i,j=1,2,...,C, X is the number of channels, H and W are the height and width of the space respectively; Then, the two pooling results are transformed through two convolutional layers respectively to obtain the channel attention weights. First, the number of channels is reduced, and then nonlinear transformation is performed through activation functions ReLU and Sigmoid to generate a channel-level attention map. Finally, the number of channels is restored to the original value, specifically: A avg =Conv2D 1×1 (ReLU(Conv2D 1×1 (X avg ))) A max =Conv2D 1×1 (ReLU(Conv2D 1×1 (X max ))) Among them, the generated A avg and A max Represents the attention weight of each channel; Finally, the calculated attention value is equal to X 1×1 Multiply element by element to get the weighted output feature map, expressed as: X CAM =X 1×1 ⊙σ(A avg +A max ) In the formula, ⊙ represents element-by-element multiplication, σ represents the Sigmoid activation function, and X CAM Represents the output value of the channel attention mechanism; In the fused convolution module, convolution kernels of sizes 3×3, 5×5, and 7×7 are used for superposition, which is expressed as: X FCM =X 3×3 +X 5×5 +X 7×7 In the spatial attention module, the two feature maps obtained by the fused convolution module are weighted according to the spatial dimension. The shape of the two feature maps obtained by the fused convolution module in the spatial position is 1×H×W, which is expressed as: The two feature maps obtained by the fused convolution module are connected to form a new feature map: X concat =[X avg ,X max ]∈R 2×H×W Then, the spatial attention is calculated through the convolution operation: A spatial =σ(Conv2D 7×7 (X concat )) Among them, σ is the Sigmoid activation function, output A spatial ∈R 1×H×W represents the attention weight of each spatial position; Finally, the output after spatial attention weighting is: X SAM =X⊙A spatial Therefore, the channel attention output and spatial attention output are weighted, and finally the output result is obtained through 1×1 convolution: CSAM =Conv2D 1×1 (X CAM +X SAM ).
5. The hyperspectral image classification method according to claim 2, characterized in that: The key feature information is input into the Kansformer module for processing to obtain feature enhancement data, specifically including: In the Kansformer module, batch normalization is used to replace layer normalization, and the KAN network is used to capture the complex differences between different ground object categories to obtain feature enhanced data.
6. A hyperspectral image classification system, characterized in that: include: A data acquisition module, used to obtain a hyperspectral training data set; The training data set includes training images and corresponding labels; A pre-training module, used to construct a pre-training network, and input the training data set into the pre-training network for parameter optimization to obtain a trained image classification model; the pre-training network includes a space and channel reconstruction convolution module, a cross-scale aggregation module and a Kansformer module; the space and channel reconstruction convolution module includes a space redundancy unit and a channel redundancy unit; the cross-scale aggregation module includes a channel attention module, a fusion convolution module and a space attention module; The Kansformer module includes a KAN network; The image classification module is used to use the image classification model to identify the hyperspectral image to be detected and obtain a hyperspectral image classification result.
7. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the hyperspectral image classification method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The device stores a computer program, which, when executed by a processor, implements the hyperspectral image classification method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Space-spectrum combined hyperspectral image classification method based on channel and space attention
CN117333713A
Hyperspectral image classification method based on multi-scale cavity convolution and attention mechanism
CN118537727A
Bone marrow cell classification and identification method based on SCKannform neural network
CN119049042A
Hyperspectral image classification method
CN119169399A
Cited By
Image multi-target rapid segmentation method based on improved convolutional neural network
CN121074388A