Hyperspectral image classification method

By designing a dual-branch multi-scale feature extraction method in hyperspectral image classification, combining a lightweight transformer encoder and a multi-scale spectral feature extraction module, the overfitting problem caused by the neglected global context information and large number of model parameters in the prior art is solved, and higher classification accuracy and generalization are achieved.

CN120014344AActive Publication Date: 2025-05-16CHANGCHUN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510088984.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Due to the local receptive field limitation, the existing hyperspectral image classification method ignores global context information. The traditional convolutional neural network model has large parameters, which is prone to problems of overfitting and poor generalization.

Method used

A two-branch multi-scale feature extraction classification method is designed, including a lightweight transformer encoder module and a multi-scale spectral feature extraction module, as well as a convolutional nonlinear module, a similarity self-attention module and a multi-scale spatial feature extraction module, which adaptively fuses spectral and spatial features through a dynamic feature fusion module.

Benefits of technology

It effectively improves classification accuracy and generalization, improves the flexibility, scalability and computing efficiency of the model, and solves the problems of insufficient feature extraction and poor fusion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014344A_ABST
    Figure CN120014344A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral image classification method, and relates to the technical field of image processing, and the method comprises the following steps: preparing and preprocessing a data set; constructing a network model; selecting a loss function and an evaluation index; training a network model; according to the invention, the spectral feature extraction module and the spatial feature extraction module are processed in parallel, the spectral and spatial features of different receptive fields are obtained by adopting a multi-scale feature extraction mode, and the spectral and spatial features are adaptively fused by adopting the dynamic feature fusion module. The method can effectively improve the classification precision, training efficiency and generalization of the model, and has good flexibility and expandability. A proposed dynamic feature fusion module can adaptively perform dynamic cross-feature interaction based on a cross-attention module, the cross-attention module can realize efficient interaction between spectrum and spatial features based on a dynamic mechanism, and meanwhile, the generalization ability of the model to complex scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a hyperspectral image classification method. Background Art

[0002] Hyperspectral image classification aims to classify objects by analyzing the rich spectral information and spatial structure in pixels. It has a wide range of application prospects in many fields such as crop detection, geological exploration, environmental monitoring and military reconnaissance. The networks that can achieve good performance in hyperspectral image classification mainly include artificial neural networks, deep belief networks, graph neural networks and convolutional neural networks. Since convolutional neural networks have powerful automatic feature extraction capabilities and can effectively extract spectral and spatial information, they have been widely used in hyperspectral image classification methods. However, due to the limitation of local receptive fields, the classification method of convolutional neural networks focuses more on feature extraction in local areas and ignores global context information. In addition, traditional convolutional neural network models require a large number of labeled samples, the network parameters are large, and they are prone to problems such as overfitting and poor generalization. Therefore, it is of great practical significance in hyperspectral image classification to reduce the number of parameters and improve training efficiency while ensuring classification accuracy.

[0003] The Chinese patent publication number is "CN116563606A", and the name is "A hyperspectral image classification method based on a dual-branch spatial spectrum global feature extraction network". This method first preprocesses the data and marks the sample index. Then the preprocessed data is input into the encoding and decoding structure of the spatial sub-network for spatial feature extraction, the output pixels of the spatial sub-network are recorded according to the index value, and then the pixels selected by the training index are input into the progressive feature learning of the spectral sub-network for spectral feature extraction. Finally, the extracted spatial spectral features are classified by adaptive weighting. This method has the problems of low accuracy, poor generalization performance, insufficient feature extraction and poor feature fusion effect in hyperspectral image classification. Summary of the invention

[0004] The main purpose of the embodiment of the present invention is to propose a hyperspectral image classification method, comprising the following steps:

[0005] Step 1, prepare and preprocess the data set: prepare four hyperspectral data sets, preprocess each data set, and divide it into training set and test set;

[0006] Step 2, constructing a network model: constructing a network model including a spectral calibration block, a spectral feature extraction module, a spatial feature extraction module, a dynamic feature fusion module and an M-type classifier; wherein the spectral feature extraction module includes a lightweight transformer encoder module and a multi-scale spectral feature extraction module for extracting global spectral features and multi-scale local spectral features; the spatial feature extraction module includes a convolutional nonlinear module, a similarity self-attention module, a multi-scale spatial feature extraction module and a convolutional layer for extracting global spatial features and multi-scale local spatial features; the dynamic feature fusion module adaptively fuses spectral and spatial features based on a cross mutual attention module;

[0007] Step 3, select loss function and evaluation index: use cross entropy loss function as loss function, select overall classification accuracy, average classification accuracy, Kappa coefficient and confusion matrix as evaluation index;

[0008] Step 4, train the network model: set the maximum number of training rounds, select the Adam optimizer and use the warm-up strategy to gradually increase the learning rate, and train the network model until the value of the loss function is less than the set threshold.

[0009] Further, in step 1, the four hyperspectral datasets include the Indian pine dataset, the Salinas Valley dataset, the University of Pavia dataset, and the Houston 2013 dataset;

[0010] The preprocessing steps for each dataset include standardizing the hyperspectral image data, eliminating the differences between different spectral bands, and dividing the standardized hyperspectral image into fixed-size cubic blocks using a sliding window method.

[0011] Further, in step 2, the lightweight transformer encoder module includes layer normalization, lightweight multi-head self-attention, feedforward network and residual connection;

[0012] The steps of the lightweight multi-head self-attention include: linearly mapping the input into three matrices Q, K and V, applying a maximum pooling operation with a window size of 2 and a stride of 2 to matrices K and V before the attention operation, transforming the dimensions of matrices Q, K and V, concatenating the transformed K and V, and then passing Q, K and V through a linear layer and then entering the multi-head self-attention;

[0013] The steps of the feedforward network include: performing channel expansion through 1×1×1 convolution and R-type activation function, using channel segmentation technology to divide the features into two groups, W1 and W2, wherein one group W1 uses 1×1×7 convolution and R-type activation function to encode the local context, and then splicing the processed W1 and W2, and inputting 1×1×1 convolution for channel adjustment.

[0014] Furthermore, in step 2, the multi-scale spectral feature extraction module includes group convolution, convolution block one, convolution block two, convolution block three and convolution block four.

[0015] Further, in step 2, the convolutional nonlinear module includes a reshaping, a two-dimensional convolutional layer, a batch normalization layer and an M-type activation function;

[0016] The similarity self-attention module includes cosine similarity, Gaussian Euclidean similarity, S-type function and residual connection;

[0017] The multi-scale spatial feature extraction module includes six convolution blocks, global average pooling, a fully connected layer, a batch normalization layer and an R-type activation function.

[0018] Furthermore, in step 2, the dynamic feature fusion module includes a two-dimensional adaptive average pooling layer, a three-dimensional adaptive average pooling layer, a flattening layer and a cross attention module.

[0019] Furthermore, in step 2, the spectral calibration block uses a three-dimensional convolutional layer to reduce the spectral dimension.

[0020] Furthermore, the M-type classifier uses a multi-layer perceptron to process complex nonlinear relationships in spectral information and closely link features with classification outputs.

[0021] Furthermore, four activation functions, namely R-type activation function, L-type function, S-type function and M-type function, are used to increase the nonlinear factor of the network and fully extract image features.

[0022] Furthermore, the hyperspectral image classification method also includes: step 5, saving the model.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1. The present invention provides a dual-branch multi-scale feature extraction and classification method. First, a spectral feature extraction module is designed, including a lightweight transformer encoder module and a multi-scale spectral feature extraction module. Then, a spatial feature extraction module is designed, including a convolutional nonlinear module, a similarity self-attention module, a multi-scale spatial feature extraction module and a convolutional layer. A dynamic feature fusion module is also designed to fully fuse the extracted spectral and spatial features. The method can capture more comprehensive spatial and spectral features, thereby effectively improving classification accuracy and generalization, and at the same time has good flexibility, scalability and computational efficiency.

[0025] 2. The dynamic feature fusion module proposed in the present invention can adaptively perform dynamic cross-feature interaction based on the cross-attention module. The cross-attention module can realize efficient interaction between spectral and spatial features based on a dynamic mechanism, while improving the generalization ability of the model for complex scenes. By adaptively adjusting the weight distribution, the extracted spectral features and spatial features can be fully integrated, thereby effectively improving the accuracy, robustness and generalization of classification.

[0026] 3. The present invention proposes a similarity self-attention module in the spatial feature extraction module, which extracts global spatial features by exploring the relationship between pixels. It also proposes a multi-scale spatial feature extraction module, which extracts multi-scale local spatial features through receptive fields of different scales, and uses spectral information for adaptive adjustment to make the feature representation more accurate and effective. That is, the module can comprehensively capture spatial features, adapt to complex image patterns, and improve classification accuracy.

[0027] 4. The present invention proposes a lightweight transformer encoder module in the spectral feature extraction module, which extracts global spectral features by constructing long-range dependencies in the spectral dimension. It also proposes a multi-scale spectral feature extraction module, which extracts multi-scale local spectral features by applying receptive fields of different scales on grouped convolutions. That is, this module reduces the calculation complexity of the model while extracting rich spectral features, thereby improving training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.

[0029] Figure 1 A flowchart of the steps of the hyperspectral image classification method provided by the present invention;

[0030] Figure 2 The overall framework diagram of the network model provided by the present invention;

[0031] Figure 3 A framework diagram of a lightweight transformer encoder module provided by the present invention;

[0032] Figure 4 A framework diagram of the lightweight multi-head self-attention provided by the present invention;

[0033] Figure 5 A framework diagram of a feedforward network provided by the present invention;

[0034] Figure 6 A framework diagram of the multi-scale spectral feature extraction module provided by the present invention;

[0035] Figure 7 A framework diagram of the convolutional nonlinear module provided by the present invention;

[0036] Figure 8 A framework diagram of the similarity self-attention module provided by the present invention;

[0037] Fig. 9 This is a framework diagram of the multi-scale spatial feature extraction module provided by the present invention. DETAILED DESCRIPTION

[0038] The present invention proposes a hyperspectral image classification method, aiming to solve the problems of insufficient feature extraction, poor fusion effect and low classification accuracy caused by low spatial resolution, large data volume and poor intra-class spectral heterogeneity in existing hyperspectral image classification methods.

[0039] The hyperspectral image classification method proposed by the present invention will be described in the following specific embodiments:

[0040] Embodiment 1:

[0041] A hyperspectral image classification method comprises the following steps:

[0042] Step 1, prepare and preprocess the data set: prepare four hyperspectral data sets, preprocess each data set, and divide it into training set and test set;

[0043] Step 2, constructing a network model: constructing a network model including a spectral calibration block, a spectral feature extraction module, a spatial feature extraction module, a dynamic feature fusion module and an M-type classifier; wherein the spectral feature extraction module includes a lightweight transformer encoder module and a multi-scale spectral feature extraction module for extracting global spectral features and multi-scale local spectral features; the spatial feature extraction module includes a convolutional nonlinear module, a similarity self-attention module, a multi-scale spatial feature extraction module and a convolutional layer for extracting global spatial features and multi-scale local spatial features; the dynamic feature fusion module adaptively fuses spectral and spatial features based on a cross mutual attention module;

[0044] Specifically, the spectral calibration block uses a three-dimensional convolutional layer of size (1, 1, C) to reduce the spectral dimension while keeping the spatial dimension unchanged;

[0045] The spectral feature extraction module is divided into two branches: a lightweight transformer encoder module and a multi-scale spectral feature extraction module, which are used to extract global spectral features and multi-scale local spectral features respectively; the lightweight transformer encoder module consists of layer normalization, lightweight multi-head self-attention, feedforward network and residual connection. This module extracts global spectral features by constructing long-range dependencies in the spectral dimension; the multi-scale spectral feature extraction module consists of grouped convolution and four convolution blocks with different convolution kernel sizes, and applies different receptive fields to different channel dimensions to extract multi-scale local spectral features;

[0046] The spatial feature extraction module consists of a convolutional nonlinear module, a similarity self-attention module, a multi-scale spatial feature extraction module and a convolutional layer; the convolutional nonlinear module consists of a reshaping, a two-dimensional convolutional layer, a batch normalization layer and an M-type activation function to obtain sufficient input information; the similarity self-attention module consists of cosine similarity, Gaussian Euclidean similarity, an S-type function and a residual connection to extract global spatial features by exploring the relationship between the central pixel and adjacent pixels; the multi-scale spatial feature extraction module consists of six convolutional blocks, global average pooling, a fully connected layer, a batch normalization layer and an R-type activation function to extract multi-scale local spatial features through receptive fields of different scales, and at the same time use spectral information for adaptive adjustment to make feature representation more accurate and effective; the convolutional layer integrates global spatial features and multi-scale local spatial features;

[0047] The dynamic feature fusion module includes a two-dimensional adaptive average pooling layer, a three-dimensional adaptive average pooling layer, a flattening layer and a cross attention module. This module can fully fuse the extracted spectral features and spatial features by adaptively adjusting the weight distribution through a dynamic mechanism, which helps to achieve better classification accuracy.

[0048] Step 3, select loss function and evaluation index: use cross entropy loss function as loss function, select overall classification accuracy, average classification accuracy, Kappa coefficient and confusion matrix as evaluation index;

[0049] Step 4: Train the network model: Set the maximum number of training rounds, select the Adam optimizer, and use the warm-up strategy to gradually increase the learning rate. Train the network model until the value of the loss function is less than the set threshold.

[0050] Step 5, save the model: solidify the model parameters with the best performance according to the evaluation indicators and determine the final hyperspectral image classification model.

[0051] Further, in step 1, the four hyperspectral datasets include the Indian pine dataset, the Salinas Valley dataset, the University of Pavia dataset, and the Houston 2013 dataset;

[0052] The preprocessing steps for each dataset include standardizing the hyperspectral image data, eliminating the differences between different spectral bands, and dividing the standardized hyperspectral image into fixed-size cubic blocks using a sliding window method.

[0053] Further, in step 2, the lightweight transformer encoder module includes layer normalization, lightweight multi-head self-attention, feedforward network and residual connection;

[0054] The steps of the lightweight multi-head self-attention include: linearly mapping the input into three matrices Q, K and V, applying a maximum pooling operation with a window size of 2 and a stride of 2 to matrices K and V before the attention operation, transforming the dimensions of matrices Q, K and V, concatenating the transformed K and V, and then passing Q, K and V through a linear layer and then entering the multi-head self-attention;

[0055] The steps of the feedforward network include: performing channel expansion through 1×1×1 convolution and R-type activation function, using channel segmentation technology to divide the features into two groups, W1 and W2, wherein one group W1 uses 1×1×7 convolution and R-type activation function to encode the local context, and then splicing the processed W1 and W2, and inputting 1×1×1 convolution for channel adjustment.

[0056] Furthermore, in step 2, the multi-scale spectral feature extraction module includes group convolution, convolution block one, convolution block two, convolution block three and convolution block four.

[0057] Further, in step 2, the convolutional nonlinear module includes a reshaping, a two-dimensional convolutional layer, a batch normalization layer and an M-type activation function;

[0058] The similarity self-attention module includes cosine similarity, Gaussian Euclidean similarity, S-type function and residual connection;

[0059] The multi-scale spatial feature extraction module includes six convolution blocks, global average pooling, a fully connected layer, a batch normalization layer and an R-type activation function.

[0060] Furthermore, in step 2, the dynamic feature fusion module includes a two-dimensional adaptive average pooling layer, a three-dimensional adaptive average pooling layer, a flattening layer and a cross attention module.

[0061] Furthermore, in step 2, the spectral calibration block uses a three-dimensional convolutional layer to reduce the spectral dimension.

[0062] Furthermore, the M-type classifier uses a multi-layer perceptron to process complex nonlinear relationships in spectral information and closely link features with classification outputs.

[0063] Furthermore, four activation functions, namely R-type activation function, L-type function, S-type function and M-type function, are used to increase the nonlinear factor of the network and fully extract image features.

[0064] Embodiment 2:

[0065] A hyperspectral image classification method specifically comprises the following steps:

[0066] Step 1, prepare the dataset:

[0067] Four public hyperspectral datasets are prepared, namely the Indian pine dataset, the Salinas Valley dataset, the University of Pavia dataset, and the Houston 2013 dataset; the Indian pine dataset was acquired at the Indian pine agricultural experimental field using an infrared imaging spectrometer sensor. The wavelength range of this dataset is 0.4-2.5um, the spatial resolution is 20m, and it contains 220 bands. 20 bands are removed, and 200 bands with a spatial dimension of 145×145 are retained and divided into 16 land cover categories, with a total of 10249 labeled pixels; the Salinas Valley dataset was acquired over the Salinas Valley in California, USA using an infrared imaging spectrometer sensor. The wavelength range of this dataset is 0.4-2.5um, the spatial resolution is 3.7m, and it contains 224 bands. 20 bands are removed, and 204 bands with a size of 512×217 are retained and divided into 16 land cover categories; the University of Pavia dataset was acquired using a reflective optical system imaging spectrometer sensor. The sensor was acquired over the University of Pavia in Italy. The wavelength range of this dataset is 0.43-0.86um, the spatial resolution is 1.3m, and it contains 115 bands. 12 bands are removed, and 103 bands with a size of 610×340 are retained, which are divided into 9 land cover categories, with a total of 42,776 labeled pixels; the Houston 2013 dataset is a dataset obtained by the hyperspectral image analysis team and the Airborne Laser Mapping Center of the University of Houston in the United States. The wavelength range of this dataset is 0.38-1.05um, the spatial resolution is 2.5m, and it contains 144 bands. 48 bands are removed, and 96 bands with a size of 349×1905 are retained, which are divided into 15 land cover categories, with a total of 14,278 labeled pixels; the Indian pine dataset uses 3% as the training set, the Salinas dataset and the University of Pavia dataset use 0.5% of the sample data as the training set, and the Houston 2013 dataset uses 5% of the sample data as the training set.

[0068] Data preprocessing:

[0069] The purpose of standardization is to eliminate the differences between different spectral bands and keep the data distribution of each band consistent, thereby improving the classification performance of the model; the standardization operation needs to read the original hyperspectral image data first, calculate the normalized value of the pixel value of each band, and then use the sliding window method to divide the standardized hyperspectral image into fixed-size cube blocks. The label of each cube block is determined by the label of the central pixel. Taking the Indian pine dataset with a size of (145, 145, 200) as an example, the size of the input image block is fixed to (9, 9, 200). The standardized expression is as follows:

[0070]

[0071] Among them, x i,j represents the original value of the i-th pixel in the j-th band, x' i,j It represents the normalized value of the ith pixel in the jth band, and N is the total number of pixels in the band.

[0072] Step 2: Build a network model: The network model is as follows: Figure 2 As shown, it specifically includes a spectral calibration block, a spectral feature extraction module, a spatial feature extraction module, a dynamic feature fusion module and an M-type classifier.

[0073] The spectral calibration block uses a 3D convolutional layer with a kernel size of 1×1×7, a stride of 1×1×2, no padding, and a kernel number of 1 to reduce the spectral dimension.

[0074] The spectral feature extraction module consists of a lightweight transformer encoder module and a multi-scale spectral feature extraction module, which are used to fully extract global spectral features and multi-scale local spectral features respectively;

[0075] The lightweight transformer encoder module is specifically composed of Figure 3 As shown in the figure, it consists of layer normalization, lightweight multi-head self-attention, feedforward network and residual connection. This module extracts global spectral features by building long-range dependencies on the spectral dimension. The input is applied with layer normalization before lightweight multi-head self-attention and feedforward network, and residual connection is applied after both. Layer normalization is used to reduce the difference between features, and residual connection is used to alleviate the gradient disappearance problem.

[0076] The specific composition of lightweight multi-head self-attention is as follows Figure 4As shown in the figure, the input is first linearly mapped into three matrices Q, K and V. In order to reduce the channel dimension and computational complexity, the matrices K and V need to be subjected to a maximum pooling operation with a window size of 2 and a stride of 2 before the attention operation. Then, the matrices Q, K and V are dimensionally transformed, and the transformed K and V are concatenated. Then, Q, K and V pass through the linear layer and enter the multi-head self-attention.

[0077] The specific composition of the feedforward network is as follows Figure 5 As shown in the figure, the input first undergoes 1×1×1 convolution and R-type activation function for channel expansion, and then the channel segmentation technique is used to divide the features into two groups, W1 and W2. One group W1 uses 1×1×7 convolution and R-type activation function to encode the local context, and then the processed W1 and W2 are spliced ​​and input into 1×1×1 convolution for channel adjustment.

[0078] The specific components of the multi-scale spectral feature extraction module are as follows: Figure 6 As shown in the figure, it consists of group convolution, convolution block 1, convolution block 2, convolution block 3 and convolution block 4. This module extracts multi-scale local spectral features through different receptive fields of group convolution; each convolution block consists of a three-dimensional convolution layer, a batch normalization layer and an L-type activation function. The input size of this module is 9×9×48 and the channel dimension is 32. First, it is divided into four groups along the channel dimension. The image size of each group is 9×9×48 and the channel dimension is 8. Each group passes through different convolution blocks and obtains the same size output. The first group is input to the convolution size of 1×1× 1, the stride is 1×1×1, and the padding is 0×0×0. The second group is input to the convolution size of 1×1×3, the stride is 1×1×1, and the padding is 0×0×1. The third group is input to the convolution size of 1×1×7, the stride is 1×1×1, and the padding is 0×0×3. The fourth group is input to the convolution size of 1×1×11, the stride is 1×1×1, and the padding is 0×0×5. The output images of each group with a size of 9×9×48 and a channel dimension of 8 are channel-joined to keep the final output size consistent with the input size.

[0079] The spatial feature extraction module includes a convolutional nonlinear module, a similarity self-attention module, a multi-scale spatial feature extraction module and a convolutional layer. The size of the convolutional layer is 1×1, which is used to adjust the feature channel and increase information interaction. The spatial feature extraction module uses a residual connection between the output of the convolutional nonlinear module and the output of the convolutional layer to prevent overfitting.

[0080] The specific composition of the convolutional nonlinear module is as follows Figure 7As shown in the figure, it consists of a reshape, a two-dimensional convolution layer, a batch normalization layer and an M-type activation function, which are used to fully extract the information in the input data. The reshape is used to convert the 3D image block into a 2D image block, the convolution kernel size of the two-dimensional convolution layer is 3×3, and the padding is 1×1. The batch normalization layer is used to standardize the results, and the M-type activation function is used to enhance the nonlinear expression ability.

[0081] The similarity self-attention module is specifically composed of Figure 8 As shown in the figure, it consists of cosine similarity, Gaussian Euclidean similarity, S-type function and residual connection. This module is used to explore the relationship between the central pixel and the adjacent pixels to extract global spatial features; the input X size is 9×9×48, the central pixel X i is 1×1×48, and the adjacent pixels are X i,t =[X i,1 ,X i,2 ,X i,3 ,...,X i,9 ], the expressions of cosine similarity and Gaussian Euclidean similarity between the central pixel and the adjacent pixels are as follows:

[0082]

[0083] Among them, cosine similarity is represented by C i,t , Gaussian Euclidean similarity is denoted by G i,t , σ represents the attenuation rate.

[0084] Next, the similarity is normalized using the S-type function, and the cosine similarity self-attention map and the Gaussian Euclidean similarity self-attention map are obtained from the similarity matrix in turn. The two similarity self-attention maps are added by adaptive weights to obtain the fused similarity self-attention map. The fused similarity self-attention map is multiplied element by element with the input data X and then a residual connection is performed. The expression of the fused similarity self-attention map and the final output is as follows:

[0085] Weighted=λGaEd+(1-λ)Cos;

[0086]

[0087] Among them, Cos represents the cosine similarity self-attention map, GaEd represents the Gaussian Euclidean similarity self-attention map, Weighted represents the fused similarity self-attention map, Y represents the output, and λ represents the weight parameter, which ranges from 0 to 1 and has an initial value of 0.5.

[0088] The specific components of the multi-scale spatial feature extraction module are as follows: Fig. 9As shown in the figure, it consists of six convolution blocks, global average pooling, fully connected layers, batch normalization layers and R-type activation functions. This module is used to capture multi-scale local spatial features obtained by adaptively guiding spatial information through spectral information; each convolution block consists of a two-dimensional convolution layer, a batch normalization layer and an L-type activation function; the input data X first passes through three parallel branched convolution blocks to obtain feature maps of three different receptive fields, X1, X2 and X3, where the convolution kernel size of convolution block five is 3×3, the convolution kernel size of convolution block six is ​​5×5, and the convolution kernel size of convolution block seven is 7×7. The feature map obtained by adding X2 and X3 is then subjected to global average pooling, fully connected layers, batch normalization layers, R-type activation functions and fully connected layers to obtain a new feature map, and the new feature map is subjected to the R-type activation function to obtain the channel attention map α and in And α and The ratio of represents the importance of each element channel information in the ratio of feature map X2 to X3, and the channel attention map α and It is possible to adaptively extract more important information from feature maps X2 and X3 to obtain feature map X4. The expressions of channel attention map α and feature map X4 are as follows:

[0089]

[0090] X4=αX2+(1-α)X3;

[0091] Among them, α represents the channel attention map, σ represents the activation function, and W BN represents the batch normalization layer, W FC represents the fully connected layer, W GAP represents the global average pooling layer, represents the element addition operation, and X4 represents the output feature map;

[0092] In order to refine the target position information, the feature maps X1 and X4 need to be added after passing through convolution block eight with a convolution kernel size of 1×1 to obtain the feature map X5. In order to fully extract the multi-scale local spatial features obtained by adaptively guiding the spatial information by spectral information, the feature maps X1 and X5 need to be added after passing through convolution block eight to obtain the feature map Y.

[0093] The dynamic feature fusion module consists of a two-dimensional adaptive average pooling layer, a three-dimensional adaptive average pooling layer, a flattening layer and a cross mutual attention module. The module adaptively fuses spectral and spatial features based on the cross mutual attention module. The spectral branch passes through the three-dimensional adaptive average pooling layer and the flattening layer in sequence, and the spatial branch passes through the two-dimensional adaptive average pooling layer and the flattening layer in sequence to keep the dimensions consistent for subsequent processing. Then the spectral features are mapped to the query matrix Q1, the value matrix V1 and the key matrix K1, and the spatial features are mapped to the query matrix Q2, the value matrix V2 and the key matrix K2. The matrix obtained by concatenating the key matrix K1 and the key matrix K2 passes through the fully connected layer and the activation function to obtain the key matrix dynamically adjusted by the weight coefficient. In the spectral branch, the cross-feature attention score is obtained by multiplying the query matrix Q1 containing spectral information with the dynamically adjusted key matrix containing spectral and spatial information. The score is obtained by the S-type function to obtain the score weight, and then multiplied with the value matrix V1 containing spectral features to further perform dynamic cross-feature interactive fusion. In the spatial branch, the specific process of interactive fusion with spectral features is similar to the above steps. The specific flow of the whole process is as follows:

[0094]

[0095] M = M1 + M2;

[0096] Among them, ∥ represents the splicing operation, K spa∥spe represents the key matrix that integrates spectral and spatial information, σ represents the activation function, and W FC represents the fully connected layer, X spe represents the input of the spectral branch, X spa represents the input of the spatial branch, M1 represents the spectral feature embedded with spatial information, M2 represents the spatial feature embedded with spectral information, and M represents the fusion feature of spectral and spatial information.

[0097] The M-type classifier uses a multi-layer perceptron, which can effectively process the complex nonlinear relationship in spectral information. It has a relatively simple structure and is easy to implement.

[0098] In order to ensure the robustness of the network, more nonlinear factors are introduced to fully extract image features. The present invention uses four activation functions, namely R-type activation function, L-type function, S-type function and M-type function, which are defined as follows:

[0099]

[0100] f(x) Mish =x×tanh(ln(1+e x ));

[0101] Step 3, select the appropriate loss function and evaluation index: During the training process, the loss function selected is the cross entropy loss function often used in classification tasks. The cross entropy loss function calculates the loss value and uses the gradient back propagation algorithm to adjust the model parameters. The expression of the cross entropy loss function is as follows:

[0102]

[0103] Among them, Loss represents the cross entropy loss value, N represents the total number of samples, C represents the total number of categories, and y i,j represents the jth element of the true label of sample i, It represents the predicted probability that sample i belongs to category j.

[0104] The evaluation indicators selected are overall classification accuracy, average classification accuracy, Kappa coefficient and confusion matrix to jointly evaluate the classification performance of the model; the overall classification accuracy reflects the overall classification performance of the model, the average accuracy reflects the classification performance of each category of the model, the Kappa coefficient reflects the accuracy of the model classification, and the confusion matrix intuitively reflects the classification results of the model category. The expressions of classification accuracy, average classification accuracy, Kappa coefficient and confusion matrix are as follows:

[0105]

[0106] Among them, N represents the total number of samples, L represents the total number of categories, and C ii represents the number of correctly classified samples, C ij represents the number of samples of category i classified into category j, C(i,j) represents the number of samples of category j classified into category i, N i Represents the total number of category i.

[0107] Step 4, training the network model: In the training of the hyperspectral classification network model, the maximum number of training rounds of the training model is set to 200. The back propagation algorithm selects the Adam optimizer and uses the warm-up strategy to gradually increase the learning rate to ensure stable convergence of the model. The original learning rate is set to 0.001, and the threshold of the loss function is set to 0.0005. When the function value of the loss function is less than the threshold, the entire network training can be considered to be completed. The warm-up strategy expression is as follows:

[0108]

[0109] Among them, η represents the maximum achievable learning rate, t represents the current step, T represents the total number of training steps, ω represents the number of warm-up steps, and the learning rate gradually increases from zero to the learning rate η and then decreases to zero.

[0110] Step 5, save the model: solidify the model parameters with the best performance according to the evaluation indicators and determine the final hyperspectral image classification model.

[0111] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A hyperspectral image classification method, characterized in that: The following steps are involved: Step 1, prepare and preprocess the data set: prepare four hyperspectral data sets, preprocess each data set, and divide it into training set and test set; Step 2, constructing a network model: constructing a network model including a spectral calibration block, a spectral feature extraction module, a spatial feature extraction module, a dynamic feature fusion module and an M-type classifier; wherein the spectral feature extraction module includes a lightweight transformer encoder module and a multi-scale spectral feature extraction module for extracting global spectral features and multi-scale local spectral features; the spatial feature extraction module includes a convolutional nonlinear module, a similarity self-attention module, a multi-scale spatial feature extraction module and a convolutional layer for extracting global spatial features and multi-scale local spatial features; the dynamic feature fusion module adaptively fuses spectral and spatial features based on a cross mutual attention module; Step 3, select loss function and evaluation index: use cross entropy loss function as loss function, select overall classification accuracy, average classification accuracy, Kappa coefficient and confusion matrix as evaluation index; Step 4, train the network model: set the maximum number of training rounds, select the Adam optimizer and use the warm-up strategy to gradually increase the learning rate, and train the network model until the value of the loss function is less than the set threshold.

2. The hyperspectral image classification method according to claim 1, characterized in that: In step 1, the four hyperspectral datasets include the Indian pine dataset, the Salinas Valley dataset, the University of Pavia dataset, and the Houston 2013 dataset; The preprocessing steps for each dataset include standardizing the hyperspectral image data, eliminating the differences between different spectral bands, and dividing the standardized hyperspectral image into fixed-size cubic blocks using a sliding window method.

3. The hyperspectral image classification method according to claim 1, characterized in that: In step 2, the lightweight transformer encoder module includes layer normalization, lightweight multi-head self-attention, feedforward network and residual connection; The steps of the lightweight multi-head self-attention include: linearly mapping the input into three matrices Q, K and V, applying a maximum pooling operation with a window size of 2 and a stride of 2 to matrices K and V before the attention operation, transforming the dimensions of matrices Q, K and V, concatenating the transformed K and V, and then passing Q, K and V through a linear layer and then entering the multi-head self-attention; The steps of the feedforward network include: performing channel expansion through 1×1×1 convolution and R-type activation function, using channel segmentation technology to divide the features into two groups, W1 and W2, wherein one group W1 uses 1×1×7 convolution and R-type activation function to encode the local context, and then splicing the processed W1 and W2, and inputting 1×1×1 convolution for channel adjustment.

4. The hyperspectral image classification method according to claim 1, characterized in that: In step 2, the multi-scale spectral feature extraction module includes group convolution, convolution block one, convolution block two, convolution block three and convolution block four.

5. The hyperspectral image classification method according to claim 1, characterized in that: In step 2, the convolutional nonlinear module includes reshaping, a two-dimensional convolutional layer, a batch normalization layer and an M-type activation function; The similarity self-attention module includes cosine similarity, Gaussian Euclidean similarity, S-type function and residual connection; The multi-scale spatial feature extraction module includes six convolution blocks, global average pooling, a fully connected layer, a batch normalization layer and an R-type activation function.

6. The hyperspectral image classification method according to claim 1, characterized in that: In step 2, the dynamic feature fusion module includes a two-dimensional adaptive average pooling layer, a three-dimensional adaptive average pooling layer, a flattening layer and a cross attention module.

7. The hyperspectral image classification method according to claim 1, characterized in that: In step 2, the spectral calibration block uses a 3D convolutional layer to reduce the spectral dimension.

8. The hyperspectral image classification method according to claim 1, characterized in that: In step 2, the M-type classifier uses a multi-layer perceptron to process the complex nonlinear relationship in the spectral information and link the features with the classification output.

9. The hyperspectral image classification method according to claim 1, characterized in that: In step 3, four activation functions are selected, namely R-type activation function, L-type function, S-type function and M-type function, which are used to increase the nonlinear factor of the network and fully extract image features.

10. The hyperspectral image classification method according to claim 1, characterized in that: It also includes: Step 5, saving the model: solidifying and saving the model parameters with the best evaluation index performance.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on double-branch spatial-spectral global feature extraction network

    CN116563606A

  • Hyperspectral image classification method based on multi-scale hybrid convolutional network

    CN116524265A

  • Global interaction hyperspectral multispectral cross-modal fusion method with spectral fidelity

    CN117911830A

  • Method for classifying hyperspectral images on basis of adaptive multi-scale feature extraction model

    WO2022160771A1

Cited By

  • Hyperspectral imaging-based rapid nondestructive detection method for mixed planting of rice grains

    CN120931646A

  • A rapid nondestructive detection method for rice grain mixed seeds based on hyperspectral imaging

    CN120931646B

  • Method for distinguishing producing areas of two artemisia plants based on double-sided hyperspectral information in combination with double-flow collaborative attention classification model and application

    CN122049656A