Hyperspectral image classification method and system based on multi-scale layered mark feature fusion
By employing a multi-scale hierarchical label feature fusion method, combined with spatial-spectral joint embedding and the Transformer module, the problems of high computational resource consumption and poor training stability in hyperspectral image classification are solved, achieving more efficient and accurate feature extraction and classification.
Patent Information
- Application Number
- CN202511113977.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-28
AI Technical Summary
Existing hyperspectral image classification methods suffer from high computational resource consumption, poor training stability, and difficulty in effectively extracting rich and discriminative features.
A multi-scale hierarchical labeling feature fusion method is adopted. A hierarchical fusion framework is constructed through a spatial-spectral joint embedding module, a multi-head feature labeler, and a Transformer module. By combining multi-scale grouped convolution, spectral feature extraction, and residual attention mechanism, the sample distribution features are captured and feature semantic annotation is performed.
While maintaining computational efficiency, it extracts richer and more discriminative features, enabling it to perceive the global context while preserving local details, thereby improving classification accuracy and stability.
Smart Images

Figure CN121033501A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of hyperspectral image classification, and particularly relates to a hyperspectral image classification method and system based on multi-scale hierarchical labeled feature fusion. BACKGROUND
[0002] Hyperspectral image (HSI) is composed of hundreds of continuous and narrow spectral bands. By capturing the spectral features of target substances, accurate target recognition and material classification can be achieved. Hyperspectral image classification (HSIC) refers to the process of assigning a class label to each pixel in an image. It has been widely applied in various fields such as precision agriculture, environmental monitoring, biomedical imaging and mineral analysis. In order to solve the challenges of feature redundancy and high dimensionality, researchers have proposed a series of dimensionality reduction techniques, such as band selection and subspace learning methods. In addition, a lot of research efforts have been devoted to image denoising, spectral decomposition, target detection and classification. In the early stage of hyperspectral image classification research, the classification task mainly relied on traditional machine learning methods and manually extracted features. These methods include k-nearest neighbor (KNN) classifier, support vector machine (SVM), Bayesian estimation, multinomial logistic regression and sparse representation classification (SRC). In the scenario of small data sets or limited complexity, these methods are simple, easy to implement and computationally efficient, and can achieve relatively satisfactory classification results. With the progress of deep learning, hyperspectral image classification has entered a new era. Deep neural networks can automatically extract discriminative features, thereby reducing the dependence on manual feature engineering, and have been widely applied in hyperspectral classification tasks.
[0003] Existing techniques can make up for the deficiency of local spatial feature extraction and show good potential in capturing long-distance dependencies. However, they still face major challenges such as high consumption of computing resources and poor training stability. SUMMARY
[0004] To solve the technical problems in the above background, the present application provides a hyperspectral image classification method based on multi-scale hierarchical labeled feature fusion, the steps of which include:
[0005] Collecting a hyperspectral data set and processing it to obtain sample data;
[0006] Using the sample data to train the constructed image classification model;
[0007] Using the image classification model to complete the classification of the hyperspectral image.
[0008] Preferably, the method for processing the hyperspectral data set includes:
[0009] The obtained hyperspectral data set is extracted by 7*7 image blocks, and the class of the center pixel of each image block is taken as the class of the image block;
[0010] The extracted image blocks are divided into training samples, validation samples and validation samples;
[0011] The divided image blocks are normalized to obtain the sample data.
[0012] Preferably, the image classification model constructed by the hierarchical fusion structure is divided into four stages, each stage is composed of a spatial-spectral joint embedding module and a multi-head feature marker; Finally, a Transformer module is used, and the output of the Transformer module is used for final classification.
[0013] Preferably, the method for training the image classification model comprises:
[0014] The spatial-spectral joint embedding module is used to input the training samples into the image classification model by using spatial-spectral joint embedding;
[0015] The multi-head feature marker is used to capture sample distribution features and fuse feature semantic markers and original features;
[0016] Each stage of the image classification model performs spectral dimension compression and fuses the output of each stage;
[0017] The fused output is taken as the input of the Transformer module, and the final classification result is output by the Transformer module.
[0018] Preferably, the spatial-spectral joint embedding module comprises a multi-scale grouped convolution module, a spectral feature extraction module and a spectral residual module;
[0019] The multi-scale grouped convolution module is composed of a multi-scale grouped convolution layer, a batch normalization layer and an activation function RELU;
[0020] The spectral feature extraction module uses 3D convolution to extract spectral dimensions;
[0021] The residual spectral attention module is used for adaptive feature recalibration of the feature map.
[0022] Preferably, the multi-head feature marker uses a Gaussian distribution to initialize the weight matrix to convert shallow features into labeled semantic features, divides the feature map into multiple subspaces, and lets different heads learn independent projection and weighting coefficients in different subspaces, and then splices and fuses or sums the outputs of these subspaces.
[0023] Preferably, the image classification model utilizes the characteristics of a hierarchical fusion structure to perform different scale down-sampling on feature maps, generate feature semantic labels at different stages of the network respectively, and then fuse the multi-scale feature semantic labels for input into a subsequent network.
[0024] Preferably, in the Transformer module, GAP and Softmax are used to classify the finally extracted features to obtain the prediction probability of the sample corresponding to the features.
[0025] The application also provides a hyperspectral image classification system based on multi-scale hierarchical labeled feature fusion, which is used to implement the above method and comprises an acquisition module, a training module and a classification module.
[0026] The acquisition module is used to acquire and process a hyperspectral dataset to obtain sample data.
[0027] The training module is used to train the constructed image classification model by using the sample data.
[0028] The classification module is used to complete the classification of the hyperspectral image by using the image classification model.
[0029] Compared with the prior art, the application has the following beneficial effects:
[0030] 1. The application designs a space-spectrum joint embedding module, which combines multi-scale spatial feature extraction and spectral information capture and enhances them through residual connection and attention mechanism. This architecture can extract more rich and discriminative features while maintaining computational efficiency.
[0031] 2. The application designs a multi-head feature labeler. Inspired by the multi-head attention mechanism of Transformer, we propose a multi-head feature labeling strategy. We divide the feature map into multiple subspaces. Each head of this module performs independent projection and weighting to capture sample distribution features from multiple perspectives. This design enhances feature representation. In addition, feature semantic labeling is fused with the original features. This method can utilize both deep semantic information and shallow detail features.
[0032] 3. The application designs a hierarchical fusion framework that performs channel compression on the feature maps output by each stage. It divides the network into shallow, middle and deep layers. Feature semantic labels are generated at each layer. Then, multi-scale feature semantic labels are fused and input into subsequent modules. This design can preserve local details while perceiving global context. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to more clearly illustrate the technical solutions of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings described in the following embodiments are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0034] Figure 1 Network principle schematic diagram of the embodiment of the present application;
[0035] Figure 2 Overall structure schematic diagram of the spectrum joint embedding of the embodiment of the present application;
[0036] Figure 3 Difference schematic diagram of the ordinary convolution and the grouped convolution of the embodiment of the present application; wherein (a) represents the ordinary convolution; (b) represents the grouped convolution;
[0037] Figure 4 Spectrum attention module schematic diagram of the embodiment of the present application;
[0038] Figure 5 Multi-head feature marker of the embodiment of the present application;
[0039] Figure 6 Overall architecture schematic diagram of the Transformer of the embodiment of the present application;
[0040] Figure 7 Skip residual schematic diagram of the embodiment of the present application; wherein (a) represents the self-attention module; (b) represents the multi-head self-attention module module. DETAILED DESCRIPTION
[0041] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0042] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail with reference to the drawings and specific embodiments.
[0043] For convenience of description, the present application directly uses existing data for processing, which are three public data sets of Salinas, Pavia University and Indian Pines.
[0044] (1) Salinas: The Salinas dataset was collected by the Airborne Visible / Infrared Imaging Spectrometer (AVIRIS) sensor in 1998. The original image contains 224 bands, with wavelengths ranging from 400 nm to 2500 nm. After removing water absorption bands, 204 bands were selected for evaluation. The data has a height of 512 pixels and a width of 217 pixels. The dataset contains a total of 54129 labeled samples, covering 16 ground object categories.
[0045] (2) Pavia University: The Pavia University dataset was collected by the Reflective Optics System Imaging Spectrometer (ROSIS) sensor at the University of Pavia in northern Italy in 2001. The uncorrected dataset contains 115 spectral bands, with a band range of 0.43 to 0.86 μm. The image size is 610 x 340 pixels, with a spatial resolution of 1.3 m. The dataset covers 9 land cover categories. In the experiment, 12 noisy bands were removed, and finally 103 bands were used.
[0046] (3) Indian Pines: The Indian Pines dataset was acquired in 1992 in the northwestern United States using the Airborne Visible / Infrared Imaging Spectrometer (AVIRIS) sensor. The uncorrected dataset contains 224 spectral bands, with a wavelength range of 0.4 μm to 2.5 μm, consisting of 145 x 145 pixels, with a spatial resolution of 20 m, containing 16 land cover categories. In the experiment, 24 water absorption bands and noisy bands were removed, and finally 200 bands were selected.
[0047] Example 1
[0048] The present embodiment provides a hyperspectral image classification method based on multi-scale hierarchical labeled feature fusion, the steps comprising:
[0049] S1. Collect the hyperspectral dataset and process it to obtain sample data.
[0050] Extract 7 x 7 image blocks from the obtained hyperspectral dataset, and take the class of the center pixel of each image block as the class of the image block. This step is as follows:
[0051] For the original HSI data where M x N is the spatial size and B is the number of spectral bands. Each pixel in I has a B-dimensional spectral dimension, where each pixel forms a one hot class vector where C is the number of land cover categories. Next, perform 3D patch extraction on the HSI data I, and each 3D adjacent patch is created from I. The center pixel position of each patch is set to (x i , x j), where 0≤iM, 0≤jN. The ground truth label of each patch is determined by the label of the center pixel. When extracting the patches around a single pixel, the edge pixels cannot be retrieved, so a padding operation is performed on these pixels. The width of padding is (S-1) / 2. Therefore, the total number of 3D patches generated from I is MxN. The coverage of each patch ranges from (x i -(S-1) / 2) to (x i +(S-1) / 2) in width, and from (x j -(S-1) / 2) to (x j +(S-1) / 2) in height, and covers all B spectral bands. After removing the pixel patches with zero labels, the remaining sample patches are divided into a training sample patch set, a validation sample patch set, and a test sample patch set.
[0052] Finally, the divided image block data is normalized to obtain sample data.
[0053] S2. Using the sample data, the constructed image classification model is trained.
[0054] The model constructed in this embodiment is designed as a four-stage hierarchical framework, each stage consisting of a spatial-spectral joint embedding module and a multi-head feature marker. Finally, a Transformer module is passed through, and the output of the Transformer module is used as the final classification. The image classification model is as shown in Figure 1 .
[0055] S201. The training sample is input into the image classification model using the spatial-spectral joint embedding module in a spatial-spectral joint embedding manner.
[0056] The spatial-spectral joint embedding module proposed in this embodiment is as shown in Figure 2As shown, the module consists of a multi-scale grouped convolution layer, a batch normalization layer (BN), and an activation function RELU. Among them, the convolution kernel size used by the multi-scale grouped convolution layer of the four stages is different. The convolution kernel size of the first stage is 1x1, the second stage is 3x3, the third stage is 5x5, and the fourth stage is 7x7. Compared with using a single scale convolution kernel, such as (3x3), first of all, it will have fewer trainable parameters, and it can effectively reduce the amount of calculation required for subsequent larger convolution kernels, because each stage of the model reduces the number of channels, the first stage has the most channels, so the first stage uses the smallest convolution kernel size, and the subsequent,, and convolution kernels run on lower dimensions, thereby reducing the overall computational complexity, while the receptive field is also increasing. Secondly, different sizes of convolution kernel size can capture different scale features. Convolution focuses on the information fusion between channels, while larger convolution kernels can capture spatial context information in a larger range, thereby better extracting global features. The multi-scale design makes the model more flexible in feature extraction when facing targets of different sizes and shapes, thereby improving the model's adaptability to different scenes and data distributions. Grouped convolution has recently been introduced into the field of HSI classification. Grouped convolution has fewer parameters and lower computational complexity than ordinary convolution. By dividing the input channels into multiple groups, each group is individually convolved, so each convolution kernel only needs to process part of the channels, thereby greatly reducing the total parameter amount, reducing the computational complexity (FLOPs), and improving the computational efficiency. In addition, grouping channels can model local features for different groups, and combining with the subsequent Transformer module can model local-global spatial-spectral features, which reduces the training parameters while maintaining competitive classification results. Figure 3 The difference between ordinary convolution and grouped convolution is shown. For the input feature map First, divide along the spectral dimension into n non-overlapping feature maps:
[0057] X = {X1, X2, …, X i ,…,X n} where n represents the number of groups; For each X i Use different convolution kernels:
[0058] X′ i = W i X i +b i , i = 1, 2, …, n
[0059] where X′ i is the output of the i-th convolution kernel, W i represents the weight matrix of the i-th convolution kernel of the four stages of different size convolution kernels (k1xk1, k2xk2, k3xk3, k4xk4), and bi is the bias of the i-th convolution kernel. Finally, the output feature maps of each convolution kernel are spliced to obtain the final output X' of the grouped convolution, i.e.,
[0060] X' = Concat(X'1, X'2, …, X'N) i , … X' n )
[0061] Batch Normalization (BN) and activation function (ReLU) are used after the convolution layer to standardize the learning process and improve classification accuracy. Therefore, the specific process of the MSGE module can be represented by the following expression:
[0062]
[0063] ReLU(X) = max(0, X)
[0064]
[0065] wherein denotes the multi-scale grouped convolution.
[0066] The above module is mainly used to extract spatial information and can capture local spectral features of the spatial structure of HSI data while reducing the number of parameters. Subsequently, the spectral feature extraction module focuses on the spectral dimension using 3D convolution (kernel size = (1, 1, 7)) to avoid additional calculations in the spatial dimension, thereby effectively capturing spectral continuity and correlation. The 3D convolution kernel is only expanded in the spectral direction, which greatly reduces the computational complexity and fully utilizes the spectral information, which is beneficial to the network to better integrate the two aspects of information. For the input feature map wherein C in denotes the number of input channels, and the convolution kernel wherein k s , k s , k b respectively denote the size of the convolution kernel in the height, width, and depth directions, and C out is the number of output channels. The bias term is represented as Under the same padding strategy, let then for the output tensor any position (i, j, d) and output channel n (n = 1, 2, …, C out ), it can be represented as:
[0067]
[0068] wherein respectively denote the summation of the convolution window in the height and width directions; denotes summing over the depth direction; W (μ + p s , v + p s , ω + p b , c, n) denotes the corresponding weight in the indexed convolution kernel, ensuring that the index of the convolution kernel starts from 0. After 3D convolution, batch normalization (BN) and activation function (RELU) are added to improve classification accuracy, as described earlier.
[0069] After extracting the spatial-spectral fusion features, the feature map is output The residual spectral attention module is applied to adaptively recalibrate the features. The attention module can adaptively weight each channel according to the input data, so that more discriminative spectral features are strengthened, while redundant or noisy information is suppressed. The training parameters and computational complexity of this module can be ignored. At the same time, adding the original feature A and the feature B after attention weighting (residual connection) helps to retain the original information and alleviate the gradient vanishing problem. The specific process is shown in Figure 4 , global average pooling is used to extract the global statistical information of each channel, providing a compact description for subsequent attention to spectral features. This channel descriptor y can be represented as:
[0070]
[0071] where denotes summing over the cth channel, G denotes the GAP function, denotes the channel descriptor. Then, through 1D convolution, further learn the relationship between different channels, which can better capture the spectral correlation and improve the expression ability of spectral information. Let the 1D convolution operation be f 1D (·) (its parameters include convolution kernel and bias), and then pass through the nonlinear activation function sigmoid (denoted by σ):
[0072] s = σ (f 1D (y))
[0073] where is the attention weight of each channel.
[0074] The attention weight s is expanded to the shape of the tensor X through broadcasting, and then multiplied element-wise with X to obtain the recalibrated feature X':
[0075]
[0076] Finally, X' and the original input X are fused in residual, to obtain the final output:
[0077] X = X + X'
[0078] The global pooling and 1D convolution used in the above steps achieve effective spectral information extraction with lower computational cost and parameter overhead, which is very advantageous for processing HSI.
[0079] S202. Capture sample distribution features using a multi-head feature marker, and fuse the feature semantic labels with the original features.
[0080] This step can improve the expression ability and play a regularization effect. Using a Gaussian distribution to initialize the weight matrix converts the shallow features into labeled semantic features, which makes the semantic features expressed by the labels more consistent with the features of the sample distribution, so that the samples are easier to separate. The proposed multi-head feature labeling mechanism, as shown in Figure 5 , divides the feature map into multiple subspaces, allowing different heads to learn relatively independent projection and weighting coefficients in different subspaces, and then concatenating or summing the outputs of these subspaces. Multi-head can capture distribution features from different angles, improve expression flexibility, and enrich representation ability. Each head only needs to learn a part of the feature mapping in the subspace, which can to some extent play a regularization effect and reduce the risk of overfitting. And appropriately fusing the labeled results with the original features can allow the network to capture deep semantics while retaining the details in the shallow features.
[0081] The specific implementation process is as follows: first, the input feature map is flattened to obtain where N=SxS represents the number of spatial positions after flattening, and B represents the number of channels. Let the number of heads be H (i.e. numheads), and the dimension of each head be The parameter matrix is a Gaussian distribution initialized weight matrix, where L is the number of feature semantic labels, and after rearrangement we get The parameter matrix is a linear projection matrix. Rearrange to multi-head representation to get Then rearrange to get semantic features A:
[0082]
[0083] Next, A will rearrange the dimensions so that it becomes and perform scaling and softmax normalization on the last dimension:
[0084]
[0085] The scaling factor plays a key role in controlling the numerical range, balancing the gradient, and improving the stability of training. Next, we use W V to project the value of X' to the semantic feature space:
[0086]
[0087] Finally, the semantic feature and the multiplication are used to obtain the feature semantic label T:
[0088]
[0089] After obtaining the multi-head feature semantic label T, it is rearranged and merged into each head to obtain Then, the input is fused with the input A residual connection is made. In addition, the output of the multi-head feature labeling mechanism generates the feature semantic labels of the shallow, middle and deep networks. Since the network is a hierarchical framework, the output of the four stages is gradually reduced in channel number, resulting in different channel numbers of the feature semantic labels generated by each stage. Therefore, the feature semantic labels of the shallow, middle and deep networks need to be fused after being mapped by a fully connected layer. Then, the final output T is obtained, i.e.:
[0090]
[0091] where f(·) can be a linear transformation or a simple mapping such as convolution to ensure consistency in dimensions. The fully connected layer is out_dim represents the unified dimension of the final output of all feature semantic labels.
[0092] S203. Each stage of the image classification model compresses the spectral dimension and fuses the output of each stage.
[0093] Using the hierarchical fusion structure characteristics of the model, each stage compresses the spectral dimension and fuses the output of each stage;
[0094] In different stages of the network (such as shallow, middle and deep), feature semantic labels are generated respectively, and then the multi-scale feature semantic labels are merged and input into the subsequent network. In order to fuse the feature semantic labels generated by each stage, the dimensions of the output of each stage need to be unified. Here, the dimension of the labeled feature output by the last stage is selected as the output dimension of each stage to realize addition fusion.
[0095] S204. The fused output is input into the Transformer module, and the final classification result is output by the Transformer module.
[0096] The feature semantic labels of the shallow, middle and deep networks obtained by the previous hierarchical framework are fused as the input of the final Transformer module, and the structure is as follows Figure 6As shown. Before entering the module, the positional information of each feature semantic tag needs to be encoded. Without positional information, the sequence information cannot be used. Each feature semantic tag is represented as [T1, T2, ..., T]. L The positional embedding is then encoded and added to the feature semantic markers. The resulting sequence of embedded feature semantic markers is represented as follows:
[0097] T in = [T1, T2, ..., T L ]+PE pos
[0098] The Transformer module mainly consists of a Multi-head Self-Attention (MSA) module, a Multi-layer Perceptron (MLP) layer, and two Layer Normalization (LN) layers. It also includes two residual skip connections.
[0099] The self-attention (SA) mechanism used in this module establishes global long-distance dependencies, effectively capturing the correlations between feature sequences, such as... Figure 7 As shown in (a). For an input two-dimensional matrix. Then three learnable weight matrices need to be defined, namely: and The input matrix T is linearly mapped to the query using these three learnable matrices. key Sum Q, K, and V are divided into h parts along the dimension axis, as shown below:
[0100] Q = {Q1, Q2, ..., Q} i Q h},
[0101] K = {K1, K2, ..., K} i , ..., K h},
[0102] V = {V1, V2, ..., V} i , ..., V h}
[0103] Where h is the number of heads. The attention score is derived from the relationship between Q and K, and the weights of the score are calculated using the softmax function. The calculation formula is as follows:
[0104]
[0105] The final output of MHSA is obtained by stitching together all the data and then performing a linear projection. Figure 7(b) shown, can be represented as:
[0106] MHSA(Q, K, V) = Concat(SA1, SA2,..., SA i ,..., SA h )W
[0107] where, denotes the output projection matrix.
[0108] Next, the above output is input to the MLP layer, which is a module consisting of two linear projection layers with a Gaussian Error Linear Unit (GELU) nonlinear activation function added in between. The linear projection is realized through a fully connected (FC) layer. GELU is an activation function widely used in visual Transformers, defined as follows:
[0109]
[0110] where, Φ(·) denotes the standard Gaussian cumulative distribution function. The overall MLP module can be represented as:
[0111] MLP(X) = FC2(GELU(FC1(X)))
[0112] In the Transformer module, LayerNorm is applied before both the MHSA module and the MLP layer. In addition, a residual connection is used to alleviate the overfitting problem commonly seen in deep learning methods. The overall calculation of the Transformer encoder block can be summarized as:
[0113] T (l+1) = MHSA(LN(T)) + T
[0114] T (l+2) = MLP(LN(T (l+1) ) + T (l+1)
[0115] where, T (l+1) T (l+2) and T GAP denote the output features of the MHSA module and the MLP module, respectively.
[0116] The input and output dimensions of the Transformer module remain unchanged. Its output result will be output as a vector through a global pooling layer (GAP). Then the dimension of T GAP is changed to the number of ground object classes C through a linear layer, and finally the final classification class is calculated through a softmax function.
[0117] The final extracted features are classified using GAP and Softmax to obtain the prediction probability of the sample corresponding to the features.
[0118] S3. Using the image classification model, the hyperspectral image classification is completed.
[0119] Embodiment Two
[0120] To verify the advancement of the present application compared to the prior art, this embodiment is set as a comparative verification.
[0121] The data set is shown in Table 1, Table 2, and Table 3. In Table 1, the training and verification sample numbers of 1% labeled samples and the test sample numbers of 98% labeled samples in the SA data set.
[0122] Table 1
[0123]
[0124] Table 2, the training and verification sample numbers of 1% labeled samples and the test sample numbers of 98% labeled samples in the PU data set.
[0125] Table 2
[0126]
[0127] Table 3, the training and verification sample numbers of 10% labeled samples and the test sample numbers of 80% labeled samples in the IP data set.
[0128] Table 3
[0129]
[0130] Experimental setup: For fairness, all HSI classification methods are implemented on the PyTorch 2.3.0 deep learning framework. The number of iterations for each training is set to 100, and the batch size of training samples for each iteration is set to 64. The optimizer and learning rate adjuster of other HSI classification methods are the same as those set in the original article. The optimizer of the method proposed in this embodiment is selected as the stochastic gradient descent (SGD), the momentum is set to 0.9, and the weight decay is set to 0.00001. The learning rate of this embodiment is fixed at 0.001. In addition, all experiments are carried out in a server equipped with an NVIDIA RTX4090 24GB GPU and 128GB RAM.
[0131] In order to make each training process more stable, a preprocessing method is adopted. Let the input data HSI three-dimensional cube be where B represents the spectral dimension, and M and N represent the spatial dimension. First, normalization is performed to fix all numerical values between 0 and 1, and the formula is as follows:
[0132]
[0133] Next, normalization is performed on each band, i.e., subtracting the average value of the corresponding channel:
[0134] X"(b, :, :) = X'(b, :, :) - mean(X'(b, :, :))
[0135] where mean(X'(b, :, :)) represents the average value of the b-th band in X'. In addition, data augmentation techniques are used to address the problem of insufficient samples. For each batch of input data, there is the same probability of vertical flipping, horizontal flipping, 90-degree rotation, 180-degree rotation, and 270-degree rotation.
[0136] To quantitatively evaluate the HSI classification performance, the overall accuracy (OA), average accuracy (AA), kappa coefficient, and per-class accuracy are used as evaluation indicators. OA represents the proportion of the number of correctly predicted samples to the total number of samples, AA is the average of the classification accuracy of each class, and kappa coefficient is a statistical quantity for evaluating the consistency between the true label map and the classification result. In this embodiment, 5 experiments are performed for all HSI classification methods. In order to reduce the influence of experimental randomness, 5 random seeds are used, and the same results are obtained when the sample is divided 5 times using different HSI classification methods. Finally, their average values are reported. In addition, the results obtained by different methods can be qualitatively compared by visualizing the classification color map. Figure 1
[0137] The results of the proposed method are compared with the latest proposed and widely used Transformer method and representative and widely used deep learning method: Group Attention Hierarchical Transformer for Hyperspectral Image Classification (GAHT), Spectral-Spatial Morphological Attention Transformer for Hyperspectral Image Classification (MorphFormer), Spectral-Spatial Feature Tagging Transformer for Hyperspectral Image Classification (SSFTT), Memory Enhanced Spectral-Spatial Transformer for Hyperspectral Image Classification (MassFormer), Re-thinking Hyperspectral Image Classification with Transformer (SpeFormer), Multi-Scale 3D Deep Convolutional Neural Network (M3D-DCNN), Deep Feature Fusion Network (DFFN), and Residual Spectral-Spatial Attention Network for Hyperspectral Image Classification (RSSAN).
[0138] The comparative experimental results are analyzed as follows:
[0139] Table 4
[0140]
[0141] Table 5
[0142]
[0143] Table 6
[0144]
[0145] Table 4 is the classification result of different classification methods on the SA data set with 1% labeled samples of the training samples; Table 5 is the classification result of different classification methods on the PU data set with 1% labeled samples of the training samples; and Table 6 is the classification result of different classification methods on the IP data set with 10% labeled samples of the training samples.
[0146] Overall, the classification performance of the scheme on the three data sets is significantly better than that of the hyperspectral image classification method based on deep learning, and is also very competitive compared with the current hyperspectral image classification method based on Transformer. The classification accuracy of the proposed model for a small number of classes is slightly lower than the optimal value, and good classification results can be maintained in each class, and the classification results are very balanced. In addition, the evaluation indicators obtained from the three data sets show that the variance of the proposed method is much smaller than that of other methods, and it can be concluded that the model of the embodiment is more stable than other models.
[0147] Embodiment three
[0148] The embodiment also provides a hyperspectral image classification system based on multi-scale hierarchical labeled feature fusion, which comprises an acquisition module, a training module and a classification module; the acquisition module is used for acquiring and processing hyperspectral data sets to obtain sample data; the training module is used for training the constructed image classification model by using the sample data; and the classification module is used for completing the classification of the hyperspectral image by using the image classification model.
[0149] The above-described embodiments only describe the preferred modes of the present application, and do not limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements to the technical solutions of the present application made by those skilled in the art shall fall within the protection scope determined by the claims of the present application.
Claims
1. A hyperspectral image classification method based on multi-scale hierarchical label feature fusion, characterized in that the steps are as follows: include: Collect and process hyperspectral datasets to obtain sample data; The constructed image classification model is trained using the sample data; The image classification model described above is used to classify hyperspectral images.
2. The hyperspectral image classification method based on multi-scale hierarchical label feature fusion according to claim 1, characterized in that, Methods for processing hyperspectral datasets include: The obtained hyperspectral dataset is processed into 7×7 image patches, and the category of the center pixel of each image patch is used as the category of that image patch. The extracted image patches are divided into training samples, validation samples, and validation samples. The image blocks are normalized to obtain the sample data.
3. The hyperspectral image classification method based on multi-scale hierarchical label feature fusion according to claim 1, characterized in that, The image classification model constructed using a hierarchical fusion structure is divided into four stages, each consisting of a spatial-spectral joint embedding module and a multi-head feature labeler; finally, it passes through a Transformer module, and the output of the Transformer module is used for the final classification.
4. The hyperspectral image classification method based on multi-scale hierarchical label feature fusion according to claim 3, characterized in that, The methods for training the image classification model include: The training samples are input into the image classification model using a spatial-spectral joint embedding module. Multi-head feature taggers are used to capture sample distribution features, and feature semantic tags are fused with the original features; The image classification model performs spectral dimension compression at each stage and fuses the outputs of each stage; The fused output is used as input to the Transformer module, which then outputs the final classification result.
5. The hyperspectral image classification method based on multi-scale hierarchical label feature fusion according to claim 3, characterized in that, The spatial-spectral joint embedding module includes: a multi-scale grouped convolution module, a spectral feature extraction module, and a spectral residual module; The multi-scale grouped convolutional module consists of a multi-scale grouped convolutional layer, a batch normalization layer, and an activation function ReLU. The spectral feature extraction module utilizes 3D convolution to extract spectral dimensions; The residual spectral attention module is used to perform adaptive feature recalibration on the feature map.
6. The hyperspectral image classification method based on multi-scale hierarchical label feature fusion according to claim 3, characterized in that, The multi-head feature labeler uses a Gaussian distribution to initialize the weight matrix, converting shallow features into labeled semantic features. It divides the feature map into multiple subspaces, allowing different heads to learn independent projections and weighting coefficients in different subspaces. Then, the outputs of these subspaces are concatenated or summed and fused.
7. The hyperspectral image classification method based on multi-scale hierarchical label feature fusion according to claim 3, characterized in that, The image classification model utilizes the characteristics of a hierarchical fusion structure to downsample the feature map at different scales, generate feature semantic labels at different stages of the network, and then fuse the multi-scale feature semantic labels into the subsequent network.
8. The hyperspectral image classification method based on multi-scale hierarchical label feature fusion according to claim 3, characterized in that, In the Transformer module, GAP and Softmax are used to classify the finally extracted features and obtain the predicted probability of the sample corresponding to the feature.
9. A hyperspectral image classification system based on multi-scale hierarchical label feature fusion, the system being used to implement the method according to any one of claims 1-8, characterized in that, include: The module consists of a data acquisition module, a training module, and a classification module. The acquisition module is used to acquire hyperspectral datasets and process them to obtain sample data; The training module is used to train the constructed image classification model using the sample data; The classification module is used to classify hyperspectral images using the image classification model.