Hyperspectral image classification method based on spectrum-space double fusion network
By introducing spectral-space dual fusion network (S2DFT) into hyperspectral image classification, combined with principal component analysis and multi-spectral spatial self-attention technologies, the problems of high-dimensionality and limited sample annotation of hyperspectral image data are solved, and higher classification accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510127490.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-06
AI Technical Summary
The high-dimensionality, redundant information, and limited sample annotation of hyperspectral image data make it challenging to identify different land cover categories during model training, affecting classification accuracy.
A hyperspectral image classification method based on spectral-space dual fusion network (S2DFT) is proposed. The shallow features are extracted through principal component analysis (PCA) dimensionality reduction, the design of spatial feature extraction module (SPAEM) and the spectral feature extraction module (SPEEM), and the multi-head spectral spatial self-attention (MHS3A) module are fused, and the spectrum and spatial information are finally classified using global average pooling and fully connected layers.
Performance evaluation experiments performed on four HSI datasets show that the S2DFT method has higher accuracy and generalization capabilities, especially when the number of training samples is limited, improving the accuracy of land cover classification.
Smart Images

Figure CN119942348A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hyperspectral image processing, and in particular relates to a hyperspectral image classification method based on a spectral-spatial dual fusion network. Background Art
[0002] With the rapid development of spectral imaging technology, modern hyperspectral images have achieved significant improvements in both spatial and spectral resolution, and can capture two-dimensional spatial data and one-dimensional spectral data containing rich information. Due to its detailed spectral characteristics, hyperspectral images are widely used in precision agriculture, biomedical imaging, urban planning, mineral exploration, environmental monitoring and other fields. Its core tasks in various applications include data denoising, unmixing, target detection and classification. Among these technologies, land cover classification is particularly critical, aiming to assign corresponding class labels to each pixel based on the spectral and spatial features in the image. However, the high dimensionality, redundant information and limited sample annotation of hyperspectral image data make it challenging to identify different land cover categories during model training, thus affecting the classification accuracy. Therefore, how to extract effective features from these high-dimensional data and overcome the problem of sample scarcity has become a key difficulty in current hyperspectral image classification technology.
[0003] In the early research of hyperspectral image classification, traditional methods such as polynomial regression, Bayesian estimation, linear regression, support vector machine and random forest were widely used. These methods usually rely on manually designed features and focus mainly on spectral information. However, they face challenges in feature extraction, high-dimensional data management and nonlinear modeling. In recent years, with the rapid development of deep learning technology, its outstanding ability in deep feature extraction has attracted widespread attention and has been favored by many researchers. Deep learning technology has made significant breakthroughs in a series of computer vision tasks, such as image classification, object detection, and image segmentation. These breakthroughs not only demonstrate the powerful potential of deep learning, but also provide new ideas and methods for research in related fields. In addition, with the continuous development of deep learning technology, a variety of innovative network architectures have emerged. For example, recurrent neural network (RNN), through its advantages in sequence data processing, plays an important role in time series problems; generative adversarial network, provides a novel framework for the training of generative models, and promotes the development of image generation, style transfer and other fields; convolutional neural network (CNN), with its outstanding performance in image processing, has become a basic model in the field of computer vision. The diversification and innovation of these deep learning architectures have made deep learning technology increasingly widely used in hyperspectral image classification tasks and promoted research progress in this field.
[0004] Due to the excellent performance of convolutional neural networks in spatial feature extraction, many CNN-based methods have been used for hyperspectral image classification and have gradually become the mainstream method for processing hyperspectral image classification tasks. Initially, researchers only used convolutional layers to solve hyperspectral image classification tasks, including 1D-CNN, 2D-CNN, 3DCNN, etc. The method based on two-dimensional convolutional neural network (CNN) effectively solves the limitations of traditional methods through automatic feature extraction. However, hyperspectral images usually cover hundreds of spectral bands, which makes the parameter scale of related algorithms expand. Then, three-dimensional convolutional neural network (3D-CNN) was introduced into the hyperspectral image classification task. The HybridSN designed by someone integrates 3D-CNN and 2D-CNN, first uses 3D-CNN to capture joint spatial and spectral features, and then 2D-CNN deepens and extracts higher-level spatial features, achieving efficient performance in hyperspectral image classification. In addition, a dual-branch dual-attention mechanism network was proposed, and two branches were designed based on 3D-CNN to capture spectral and spatial features respectively. The two branches are optimized through channel attention blocks and spatial attention blocks respectively, which improves the accuracy of feature extraction, especially when the number of training samples is limited. Although the three-dimensional convolution kernel has fewer parameters than the two-dimensional convolution kernel, these three-dimensional CNN-based schemes still have a large number of parameters. And as the depth of CNN continues to increase, it also leads to many challenges such as gradient vanishing and explosion. To address such problems, researchers introduced deep separable convolution and attention mechanisms into the field of hyperspectral image classification.
[0005] The core feature of Transformer is the self-attention mechanism, which can capture global dependencies in the input sequence. It is this feature that makes Transformer perform well in processing sequential data and has a strong ability to model long-range dependencies. Therefore, someone first applied Transformer to visual tasks and proposed Vision Transformer (ViT) as a benchmark model for image processing. ViT can effectively capture the relationship between different positions in the image through the self-attention (SA) mechanism, integrating global and local information, thereby significantly improving the perception ability of the model. The success of ViT demonstrates the wide adaptability of the Transformer model, indicating that it is not limited to text data, but can also be applied to various types of data. Hyperspectral images contain rich spectral information and have a sequential structure, so researchers have developed a variety of Transformer-based methods and successfully applied them to hyperspectral image classification tasks. For example, someone proposed a backbone network called SpectralFormer, which can extract local spectral sequence information from adjacent spectral bands by reconsidering the sequential characteristics of Transformer. Compared with the traditional ViT method, this method generates grouped spectral embeddings, which significantly improves the classification accuracy. In order to further capture more spectral and spatial information. A new dual-attention Transformer encoder was designed to fuse local spatial information with global spectral features to maximize the combined performance of spectral and spatial information.
[0006] Although Transformer-based methods can effectively model the long-range dependencies of spectral information, they usually perform poorly when processing local information due to the lack of inductive biases suitable for images (such as translation invariance), which in turn affects the classification performance. In addition, Transformer models usually have a large number of parameters and FLOPs, which requires more training data and training time. To address these problems, many studies have tried to combine the advantages of convolutional neural networks (CNNs) in local modeling with the advantages of Transformers in long-range modeling to optimize the extraction of spatial spectral information. The Spectral–Spatial Feature Tokenization Transformer (SSFTT) was proposed, which extracts shallow spectral and spatial features by combining 3D-2D convolutional layers, and then uses Gaussian weighted feature tokenizers for feature transformation, and the transformed features are input into the Transformer encoder for further processing. This method effectively overcomes the limitations of CNN-based classification methods in extracting deep semantic features. In addition, some studies combine CNN with attention mechanisms to further integrate the advantages of both. For example, some people introduced convolution operations in the multi-head attention mechanism and applied parallel residual blocks to capture local spectral spatial features in hyperspectral image patches. Some people proposed a new convolutional Transformer architecture, combining convolution with Transformer models to simultaneously model height, width, and spectral dimensions, designed a convolution-based interactive attention mechanism, and extracted spatial spectral features through parallel CNN and Transformer branches.
[0007] So far, deep learning models have shown significant advantages in hyperspectral image classification. However, existing deep learning methods related to hyperspectral image classification still face several challenges, including the following aspects:
[0008] (1) Hyperspectral images are high-dimensional and redundant. In addition, the sample annotation cost is high and the number is limited, which makes it difficult to extract effective features in model training and affects the classification accuracy.
[0009] (2) The spectral characteristics of different land cover categories are highly similar, and there is significant variability within each category due to factors such as light and season. Traditional methods are difficult to capture subtle differences.
[0010] (3) Existing deep learning models face problems such as parameter expansion and gradient vanishing when processing high-dimensional data, and the fusion of local and global features is insufficient, which restricts the improvement of classification performance. Summary of the invention
[0011] In order to solve the technical problems raised in the background technology, the present invention provides a hyperspectral image classification method based on a spectral-spatial double fusion transformer (S2DFT), and the technical solution adopted is as follows:
[0012] Step 1: Hyperspectral image dataset preparation
[0013] The method uses four hyperspectral image datasets: the Pavia University dataset was collected by the Reflection Optical System Imaging Spectrometer (ROSIS) sensor in an aerial survey over Pavia, northern Italy in 2001. The spatial resolution of the dataset is 1.3 meters and the image size is 610×340 pixels. It originally contained 115 spectral bands, but 12 noise bands were removed in the experiment and 103 bands were finally used. The dataset covers nine different land object categories; the Salinas dataset was collected in the Salinas Valley, California, USA, using advanced hyperspectral imaging technology. The spatial resolution of the dataset is 3.7 meters, contains 512×217 pixels, and covers 204 spectral bands. The dataset captures the diverse characteristics of the region and covers 16 different land object categories that showcase the region's agricultural and natural elements; the Botswana dataset was acquired by NASA's EO-1 satellite over the Okavango Delta and contains a 7.7-kilometer data strip collected by the Hyperion sensor. The spatial resolution of the dataset is 30 meters. It originally includes 242 spectral bands. After removing the absorption band, 145 spectral bands are retained. The dataset is divided into 14 different categories; the WHU-Hi-LongKou dataset was collected on July 17, 2018 in Longkou, Hubei, China, using a HeadwallNano-Hyperspec imaging sensor with an 8 mm focal length. The image size is 550×400 pixels, including 270 spectral bands, covering nine types of ground objects.
[0014] Step 2: Data Preprocessing
[0015] 2.1 Input: Hyperspectral image data X∈R H×W×S :Here H represents the image height, W represents the image width, and S represents the number of spectral bands, which is usually three-dimensional data composed of dozens to hundreds of bands. Each pixel contains a multidimensional spectral feature vector. The ground truth label Y∈R H×W: Represents the classification label of the image, which is a two-dimensional array that matches the spatial dimension of the HSI data and is used to guide the model to learn the correct classification or regression task during training. Each pixel is labeled as a certain category, such as vegetation, buildings, water bodies, etc. Number of training iterations N: This is a hyperparameter used to control the number of iterations of model training. Too many iterations may lead to overfitting, while insufficient iterations may lead to underfitting.
[0016] 2.2 Principal Component Analysis (PCA):
[0017] Hyperspectral images usually have very high spectral dimensions, but contain a lot of redundant information, which increases the computational complexity. Therefore, PCA is an effective dimensionality reduction method that can extract the principal components containing the most information while reducing the dimensionality of the data. PCA converts X into low-dimensional data X', retaining the principal components to maximize the variance distribution of the data. If S = 200, after PCA dimensionality reduction, S' = 10 principal components may be retained.
[0018] 2.3 Data segmentation:
[0019] Divide the data into training set and test set: training set X train ={x train ,y train}: used for model learning. Test set X test ={x test ,y test}: used for model performance verification. In order to evaluate the classification performance of the model under limited sample conditions, we divided the number of samples used in the experiment. 0.5% of the samples were selected for training on the Pavia University dataset and the Salinas dataset, 0.2% on the Botswana dataset, and 0.2% on the WHU-Hi-LongKou dataset. All training samples were randomly selected to ensure the stability of the algorithm. In addition to the training samples, the rest were used for testing.
[0020] Step 3: Hyperspectral image classification task: Initialize the S2DFT model
[0021] This paper proposes a S2DFT method for hyperspectral image classification. The network realizes full spectral-spatial feature weighted extraction and integrates spatial-spectral interactive features. First, after using principal component analysis (PCA) for dimensionality reduction, the spatial feature extraction module (SPAEM) and the spectral feature ectraction module (SPEEM) are designed to extract shallow spatial-spectral features. Secondly, we design a channel attention module (CAM) to combine convolution with RELU to perform weighted output of features extracted by SPAEM. Thirdly, multihead spectral-spatial self attention (MHS3A) is designed to replace multi-head self-attention (MHSA) to obtain spatial-spectral interactive features. Finally, the spatial-spectral interactive features and the spectral weighted features are fused, and classification is performed using global average pooling (GAP) and fully connected layers.
[0022] According to the nature of the task, a series of parameters of SPEEM and SPAEM are initialized to provide a basis for the efficient operation and accuracy of the model. First, the SPEEM module extracts spectral information from the input image through a series of convolution operations and maps it to a high-dimensional embedding space through feature transformation. This process not only relies on the original spectral wavelength information, but also captures subtle spectral differences in the image through different convolution kernels, which is particularly important in hyperspectral image analysis. Hyperspectral images usually contain hundreds of bands, and through the SPEEM module, the model can effectively extract valuable spectral features from these bands, reduce the interference of redundant information, and provide clear and accurate input for subsequent deep learning models. The convolution operations and nonlinear activation functions in this process can help the model discover implicit relationships between spectral data in high-dimensional space, which helps to distinguish different substances or categories.
[0023] Next, the SPAEM module is initialized, which extracts spatial features in the image by using multi-scale convolution kernels and spatial filtering operations. The SPAEM module focuses on capturing spatial information, especially when facing complex image scenes, and can effectively identify local and global structures. This module uses convolution kernels of multiple scales. Convolution kernels of different scales can extract detailed features at different spatial resolutions while maintaining an understanding of the overall structure of the image. It is especially suitable for processing complex images with different target sizes and shapes. By performing convolution operations at multiple scales, SPAEM can not only identify subtle features in local areas, but also capture macro information in the image, greatly improving the spatial perception ability of the model.
[0024] Then, the MHS3A module is initialized, which receives the feature vectors from the previous modules (such as SPEEM and SPAEM) as input. The input feature vector is first transformed through a linear layer, which can be regarded as a generator of query (Q), key (K), and value (V) in the self-attention mechanism. In addition, the input feature vector is multiplied by the spectral weights. These weights may be obtained through pre-training to emphasize the bands in the input data that are more helpful for the classification task. The linearly transformed feature vector is fed into the Scale Dot-Product Attention module. This module is the core of the self-attention mechanism and is used to capture the dependencies between different parts of the input data. In this module, the input data is first mapped into three vectors: query (Q), key (K), and value (V). Then, the dot product of the query and the key is calculated to obtain an attention score matrix. This score matrix is scaled by a scaling operation, usually divided by a scaling factor, which can be the square root of the dimension of the key vector to avoid the dot product result being too large. Next, the scaled attention scores are converted into probability distributions through the Softmax function, which represents the relative importance of different parts of the input data. Finally, this probability distribution is multiplied by the value vector (V) to obtain weighted value vectors that aggregate information from different parts of the input data. The MHS3A module contains multiple parallel self-attention heads, each of which performs the above attention calculations independently. This multi-head attention mechanism allows the model to learn features from different representation subspaces simultaneously, enhancing the expressiveness of the model. The outputs from multiple attention heads are spliced together to form a comprehensive feature representation. The spliced features are transformed through a linear layer to further integrate the features and prepare the output. Finally, the linearly transformed features are fed into the output layer to generate the final classification results.
[0025] At the same time, during the training process, the Adam optimizer is initialized and used to optimize the model's training process. By combining momentum and adaptive learning rate adjustment, the Adam optimizer can effectively deal with gradient problems in different training stages, improve convergence speed and reduce instability in training. By continuously adjusting the model parameters, the optimizer will guide the model to minimize the loss function and improve the model's performance in the task. The initialization of the loss function is also crucial. It will evaluate the model's performance and guide the model's training based on the characteristics of the task. In classification tasks, the cross entropy loss function can be optimized based on the difference between the category label and the probability distribution of the model output.
[0026] Specifically, the SPEEM module extracts spectral information through convolutional layers and performs feature transformation, mapping this information to a high-dimensional embedding space. This process not only reveals tiny spectral differences in hyperspectral images, but also compresses redundant information in the image and extracts the most representative features, which is crucial for subsequent analysis. The SPAEM module, under the action of multi-scale convolution kernels, extracts spatial features of different scales, not only focusing on the details of the local area of the image, but also comprehensively considering the overall structure of the image. The MHS3A module further strengthens the interaction between spectral and spatial information through a multi-head self-attention mechanism, and can accurately capture key features in the image at both the global and local levels. The combination of the two ensures that the model can extract detailed spectral features and capture comprehensive spatial information when processing images, greatly improving the model's ability in complex image analysis.
[0027] Step 4: Set up the optimizer and loss function
[0028] Select the appropriate optimizer and loss function according to the complexity of the model and the requirements of the hyperspectral image classification task.
[0029] Specific solutions include:
[0030] 4.1 According to the characteristics of the hyperspectral image classification task, the cross entropy loss function K is used CE To calculate Y pretrain and the true label Y train The error between , the cross entropy loss function is one of the commonly used loss functions in classification problems, and the loss function is expressed as:
[0031]
[0032] Among them, y i represents the true category label, Represents the predicted probability, M is the number of samples, and by calculating the loss function, we can quantify the accuracy of the model prediction;
[0033] 4.2 The Adam optimizer is used to update the model weights, and the exponential decay learning rate scheduler LearningRate Scheduler is used to optimize the model parameters to ensure that the model can be updated stably. The learning rate adjustment formula is as follows:
[0034] η t =η0·exp(-λ·t)
[0035] Among them, η t is the learning rate of the tth round, η0 is the initial learning rate, and λ is the decay rate hyperparameter;
[0036] Step 5: S2DFT model: training and saving phase
[0037] The goal of the training phase is to use the existing labeled image Y and the input hyperspectral image data X to train the S2DFT model so that it can learn effective features and make accurate predictions. The trained S2DFT model is saved to prepare for subsequent testing.
[0038] 5.1 Training loop:
[0039] The training cycle is repeated from i=1 to N until the model parameters reach the optimal state. The training cycle is the core of the model training process. Through continuous iterative optimization, the model can gradually learn the inherent laws and characteristics of the data. In each iteration, the model will go through a complete forward propagation and back propagation process, thereby gradually adjusting its internal parameters and improving prediction accuracy.
[0040] 5.2 Model saving:
[0041] When the training cycle is repeated N times, the model is trained. At this time, the parameters of S2DFT have been adjusted to the optimal state, the prediction performance of the model reaches the expected goal, and the optimal model is saved.
[0042] Step 6: S2DFT Model: Testing Phase
[0043] 6.1 Test data input, the test set sample x test Input into the trained S2DFT model. In this stage, only forward propagation is performed, and back propagation and parameter update are not required.
[0044] 6.2 Label prediction, the model outputs the predicted feature vector y for each test sample predict By comparing the feature vector with the classification threshold, the model generates the corresponding prediction label.
[0045] These predicted labels can be compared with the true labels y test Comparison is used to evaluate the performance of the model.
[0046] The present invention has the following advantages:
[0047] Performance evaluation experiments on 4 HSI datasets show that S2DFT has higher accuracy and generalization ability than existing advanced CNN and Transformer-based methods. With 0.5% training samples, the OA of Pavia University is improved by 4.39% on average, and the OA of Salinas is improved by 2.98% on average. With 2% training samples, the OA of Botswana is improved by 8.46% on average. With 0.2% training samples, the OA of WU-HI-LongKou is improved by 2.56% on average. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is the spatial-spectral dual fusion S2DFT model of the present invention;
[0049] Figure 2 is the spatial feature extraction module (SPAEM) of the present invention;
[0050] Figure 3 is the spectral feature extraction module (SPEEM) of the present invention;
[0051] Figure 4 is the Multi-Head Spectral Spatial Self-Attention (MHS3A) module of the present invention;
[0052] Figure 5 This is a visualization of the comparison results of the present invention on the Pavia University dataset;
[0053] Figure 6 It is a visualization diagram of the comparison results of the present invention on the Salinas dataset;
[0054] Figure 7 It is a visualization diagram of the comparison results of the present invention on the Botswana dataset;
[0055] Figure 8 This is a visualization diagram of the comparison results of the present invention on the WU-HI-LongKou dataset. DETAILED DESCRIPTION
[0056] The specific technical solutions of the present invention are further described below to help those skilled in the art further understand the present invention without limiting the rights thereof.
[0057] Example 1, hyperspectral image classification method based on spectral-spatial double fusion transformer (S2DFT):
[0058] The system for implementing the method consists of a network module based on CNN and Transformer, a feature fusion module and a classification module.
[0059] 1.1 Principal Component Analysis (PCA)
[0060] Assume that the original hyperspectral image is X, where h and w are the height and width of the hyperspectral image, respectively, and s is the number of spectral channels. Therefore, the hyperspectral image contains s spectral bands, providing useful spectral information, but with redundancy. Therefore, we use PCA to transform X into Xpca with the number of channels reduced to c, and then extract patches centered on each pixel of Xpca to obtain X1.
[0061] 1.2 Spectral spatial feature extraction module (SPAEM, SPEEM)
[0062] like Figure 2 As shown in the figure, SPAEM obtains shallow spatial features of HSI at different scales through four two-dimensional convolutional layers with different convolution kernel sizes. First, X1 after PCA dimensionality reduction is input into three spatial extraction blocks, each of which contains a convolutional layer with convolution kernel sizes of 3×3, 5×5 and 7×7, respectively. Each convolutional layer is followed by a batch normalization (BN) and a rectified linear unit activation function (RELU), and then a two-dimensional convolution with a kernel size of 1×1 is used to effectively integrate spatial features of different scales, and the spatial extraction feature Xspa is output.
[0063] like Figure 3 As shown in the figure, for SPEEM, its structure is similar to SPAEM, containing 3 spectral extraction blocks, each of which contains a one-dimensional convolution layer with a convolution kernel size of 1×1. Each convolution layer is followed by BN and RELU, and the shallow spectral features of the hyperspectral image are obtained by continuously changing the spectral dimension of X1. Then, the spectral features of different dimensions are integrated through a one-dimensional convolution with a kernel size of 1×1, and the spectral extraction feature Xspe is output. The corresponding formula is as follows:
[0064] x spa =f1(Concatf3x,f5f3x),f7f5f3x))))
[0065] f i (x) = RELU(BNConv2dx)
[0066] x spe =g4(Concatg1x,g2g1x),g3g2g1x))))
[0067] g i (x) = RELU(BNConc1dx)
[0068] Conv(x)=W·x+b
[0069]
[0070] RELU(x)=max(x,0)
[0071] Where fi represents the calculation of a two-dimensional convolutional layer with a kernel size of i, and gi represents the calculation of the i-th one-dimensional convolutional layer. W is the weight matrix, b is the bias, α is the numerical stability parameter, and γ and β are learnable parameter vectors.
[0072] 1.3 Spectral Attention Module (SAM)
[0073] To further extract spectral information from hyperspectral images, we design a spectral attention module (SAM), as Figure 1 As shown in Figure 2, SAM consists of two 2D convolutional layers and a RELU activation function. First, the spectral feature Xspe extracted by SPEEM is used to compress the spatial information into channel descriptors through global average pooling (GAP).
[0074] x gap =GAP(x spe )∈R 1×1×d
[0075]
[0076] In the formula, x gap It represents the feature map after spatial information compression, and q is the input feature map.
[0077] Next, we apply the first 2D convolution to transform x gap Compressed to x l gap ∈R1×1×l, which reduces the computational complexity. Then the RELU activation function is applied to learn x l gap The nonlinear interaction between channels in the gap. Third, the second convolution recovers x d gap ∈R1×1×d,
[0078] The spectral weight (SW) is obtained by using the Sigmoid activation function. Finally, SW is multiplied by x spe , and obtain the weighted spectral feature x sw ∈R N ×d In order to prevent network degradation and improve convergence speed, sw Add x spe , get the final spectrum feature map x sam ∈RN×d .
[0079] 1.4 Spatial and spectral feature fusion module
[0080] like Figure 1 As shown, using s pem The output features X spa As the input of FSSFM, the spatial-spectral sequence interaction is realized. First, the feature map X spa Flattened to sequence data x f spa ∈RN×c, and then map the tokens to the hidden layer x f1 spa ∈RN×d, where d represents the dimension of the hidden layer. Use the class label and use position embedding to mark the position information of each tag, which facilitates the final classification task while retaining the position information. This process can be expressed as:
[0081]
[0082] In the formula, φ1 represents plane operation, φ2 represents linear projection operation, and x cls and PE pos Represent the classification label and learnable position information respectively. In addition, MHSA in ViT lacks spectral perception, which limits its development in hyperspectral image classification. To solve this problem, we designed the MHS3A module. By iteratively applying MHS3A and MLP, the sequence features can be fully interacted and the global correlation can be effectively calculated. This process can be expressed as:
[0083] Y m =MHS3A(LNx spa ,x sw )+LN(x spa ,x sw )
[0084] Z m =MLP(LNY m )+LN(Y m )
[0085] Where Y m and Z m are the output feature maps of MHS3A and multi-layer perceptron (MLP), and LN is the layer normalization calculation. Next, we will describe the modules involved in this process in detail.
[0086] 1.5 Classification Module
[0087] By fusing the global spatial spectral weighted interactive features obtained by FSSFM and the local spectral weighted features obtained by CAM, the discriminative features between different categories can be more fully obtained, further improving the classification performance. The specific process is as follows:
[0088] x fusion =x fssm +x sam
[0089] Where xfusion is the fusion feature map of XFSSM and xsam.
[0090] Then, the cross entropy loss function is used to calculate the training loss and update the model weights. The process is as follows:
[0091]
[0092] Where N is the total number of training samples, k i and k ∈ i are the true label and predicted category probability of the i-th sample, respectively.
[0093] 1.6 Loss Function
[0094] The cross entropy loss function is one of the commonly used loss functions in classification problems. It can measure the difference between the predicted probability distribution and the true label distribution. Specifically, the loss function can be expressed as:
[0095]
[0096] Among them, y i represents the true category label, represents the predicted probability, and M is the number of samples. By calculating the loss function, we can quantify the accuracy of the model prediction and provide a basis for subsequent optimization steps. At the same time, in order to ensure the stability and generalization ability of the model, we use regularization methods to avoid overfitting.
[0097] 1.7 Model Optimization
[0098] The Adam optimizer is an adaptive learning rate optimization algorithm that combines the advantages of momentum and RMSprop. During the optimization process, the Adam optimizer adjusts the learning rate of each parameter based on the variance and mean of the historical gradient estimate, thereby ensuring that the model can converge to the optimal solution more quickly and stably. At the same time, in order to avoid the problem of gradient explosion or disappearance, we use methods such as learning rate decay to further adjust the optimization process.
[0099] Example 2, experiments and results of a hyperspectral image classification method based on a spectral-spatial dual fusion network (S2DFT) on the Pavia University dataset.
[0100] 2.1. Dataset Introduction
[0101] The Pavia University dataset was collected by the Reflection Optical System Imaging Spectrometer (ROSIS) sensor in an aerial survey over Pavia, northern Italy in 2001. The spatial resolution of the dataset is 1.3 meters and the image size is 610×340 pixels. It originally contained 115 spectral bands with a wavelength range of 0.43 to 0.86μm, but 12 noise bands were removed in the experiment and 103 bands were finally used. The dataset covers nine different ground object categories.
[0102] 2.2 Experimental Setup
[0103] 2.2.1 Hyperparameters
[0104] For our method, the Adam optimizer is used to update the network parameters, and the learning rate is set to 5e-4. For the training epoch, the training process stops when the training loss tends to be stable. The training epoch is uniformly set to 50, under which the training loss is stable, and the batch size is set to 64. All comparison methods use the code publicly provided by the original author and are re-experimented on the same hardware configuration to ensure the consistency of the experiments.
[0105] 2.2.2 Training Samples
[0106] In order to evaluate the classification performance of the model under limited sample conditions, we divided the number of samples used in the experiment. 0.5% of the samples were selected for training on the Pavia University dataset, and all training samples were randomly selected to ensure the stability of the algorithm. In addition to the training samples, the rest were used for testing.
[0107] 2.2.3 Operating platform and indicators
[0108] To ensure a fair comparison, the proposed method and all the compared methods are implemented in the PyTorch framework based on PyTorch 2.0.0 and cuda11.8. Both training and testing are performed on devices with Intel(R) Core(TM) i7-14700 and NVIDIAGeForce RTX 4060 (8GB GPU) graphics cards. In order to quantitatively evaluate the classification effect of each method, four commonly used evaluation metrics are used. Single-class accuracy represents the ratio of the number of correctly classified samples in each class to the total number of samples in that class. Overall accuracy (OA) is the ratio of all correctly classified samples to the total number of all samples. Average accuracy (AA) is the arithmetic mean of all single-class accuracies. Kappa coefficient (κ) is used to measure the consistency between the model classification results and the actual labels. For all these metrics, higher values indicate better performance of the method.
[0109] 2.3 Experimental analysis
[0110] We compare the proposed method with other advanced deep learning methods, including 5 CNN-based methods and 5 Transformer-based methods. Specifically, the CNN-based methods include 3-DCNN, deep feature fusion network (DFFN), HybridSN, double-branch dual-attention (DBDA), and compact band weighting (CBW). Transformer-based methods include SpectralFormer, Spectral–Spatial Feature Tokenization Transformer (SSFTT), Group-Aware Hierarchical Transformer (GAHT), and MorphFormer. The parameter settings of these methods are the same as those in the corresponding articles.
[0111] On the Pavia University dataset, the classification accuracy table and classification diagram of each method are shown in Table 1 and Figure 5 shown.
[0112] in Figure 5(a) 3D-CNN (83.71%), (b) DFFN (91.48%), (c) HybridSN (91.60%), (d) DBDA (92.10%), (e) CBW (85.26%), ( f) SF (83.61%), (g) GAHT (94.16%), (h) SSFTT (93.72%), (i) morphformer (94.13%), (j) S2DFT (94.37%).
[0113] From the classification results, it can be seen that the OA and Kappa coefficients of S2DFT are better than other methods, reaching 94.37% and 92.49%, and are 0.21%-12.76% and 0.21%-14.93% higher than other methods, respectively. Although AA is slightly lower than GAHT, the difference is not significant, which may be due to the poor classification effect of the 7th and 8th categories. Among the comparison methods, the results of 3D-CNN and SpectralFormer are poor. 3D-CNN mainly classifies by extracting local spatial-spectral features, and cannot capture the slight differences between categories well in complex scenes. SpectralFormer uses the Transformer architecture to capture long-range dependencies and cannot adapt well to the specific spectral feature space, resulting in unsatisfactory classification results. This also reflects that the classification effect of both CNN-based methods and transformer-based methods is not outstanding in the early stages of development. In contrast, S2DFT combines the advantages of CNN and transformer, and has stronger spectral-spatial feature extraction capabilities. It is particularly outstanding in the classification of objects with complex backgrounds and high spectral similarity.
[0114] Table 1 Comparison of the proposed model with other models on the Pavia University dataset
[0115]
[0116]
[0117] Example 3, experiments and results of a hyperspectral image classification method based on a spectral-spatial dual fusion network (S2DFT) on the Salinas dataset.
[0118] 3.1. Dataset Introduction
[0119] The Salinas dataset was collected in the Salinas Valley of California, USA, using advanced hyperspectral imaging technology. The dataset has a spatial resolution of 3.7 meters, contains 512×217 pixels, and covers 204 spectral bands. The dataset captures the diverse characteristics of the region, covering 16 different feature categories that showcase agricultural and natural elements in the region.
[0120] 3.2 Experimental Setup
[0121] 3.2.1 Hyperparameters
[0122] For our method, the Adam optimizer is used to update the network parameters, and the learning rate is set to 5e-4. For the training epoch, the training process stops when the training loss tends to be stable. The training epoch is uniformly set to 50, under which the training loss is stable, and the batch size is set to 64. All comparison methods use the code publicly provided by the original author and are re-experimented on the same hardware configuration to ensure the consistency of the experiments.
[0123] 3.2.2 Training samples
[0124] In order to evaluate the classification performance of the model under limited sample conditions, we divided the number of samples used in the experiment. 0.5% of the samples were selected for training on the Salinas dataset, and all training samples were randomly selected to ensure the stability of the algorithm. In addition to the training samples, the rest were used for testing.
[0125] 3.2.3 Operating platform and indicators
[0126] To ensure a fair comparison, the proposed method and all the compared methods are implemented in the PyTorch framework based on PyTorch 2.0.0 and cuda11.8. Both training and testing are performed on devices with Intel(R) Core(TM) i7-14700 and NVIDIAGeForce RTX 4060 (8GB GPU) graphics cards. In order to quantitatively evaluate the classification effect of each method, four commonly used evaluation metrics are used. Single-class accuracy represents the ratio of the number of correctly classified samples in each class to the total number of samples in that class. Overall accuracy (OA) is the ratio of all correctly classified samples to the total number of all samples. Average accuracy (AA) is the arithmetic mean of all single-class accuracies. Kappa coefficient (κ) is used to measure the consistency between the model classification results and the actual labels. For all these metrics, higher values indicate better performance of the method.
[0127] 3.3 Experimental analysis
[0128] We compare the proposed method with other advanced deep learning methods, including 5 CNN-based methods and 5 Transformer-based methods. Specifically, the CNN-based methods include 3-DCNN, deep feature fusion network (DFFN), HybridSN, double-branch dual-attention (DBDA), and compact band weighting (CBW). Transformer-based methods include SpectralFormer, Spectral–Spatial Feature Tokenization Transformer (SSFTT), Group-Aware Hierarchical Transformer (GAHT), and MorphFormer. The parameter settings of these methods are the same as those in the corresponding articles.
[0129] On the Salinas dataset, the classification accuracy table and classification diagram of each method are shown in Table 2 and Figure 6 As shown, Figure 6 (a) 3D-CNN (87.09%), (b) DFFN (89.19%), (c) HybridSN (95.19%), (d) DBDA (94.96%), (e) CBW (94.22%), ( f) SF (88.17%), (g) GAHT (94.90%), (h) SSFTT (95.01%), (i) morphformer (95.11%), (j) S2DFT (95.63%).
[0130] As can be seen from the table, the three indicators of S2DFT are all over 95%, reaching 95.63%, 96.09% and 95.14% respectively. Moreover, the three indicators of S2DFT have reached the best, which are 0.58%-7.97%, 1.49%-17.2% and 0.76%-10.54% higher than other methods respectively. Among the 16 types of objects contained in this dataset, S2DFT has the highest classification accuracy for 7 categories. And four categories have achieved an accuracy of 100.00%, which also verifies the advantage of the category adaptability of our proposed method. It also means that the method can accurately classify objects with obvious spectral differences, and can still maintain a high classification accuracy for objects with more complex or similar spectral features. It can be clearly seen from the classification diagram that S2DFT performs well overall in classifying label features, and also achieves the least classification noise. In contrast, the OA of DFFN and CBW are 89.19% and 94.22% respectively, indicating that the traditional CNN method has limitations in modeling spectral dimensions. Although Transformer methods such as SSFTT and morphformer perform well, their overall performance is still inferior to S2DFT. In addition, the Transformer method performs better than the CNN method on this dataset, which shows that it has advantages in modeling complex spectral features.
[0131] Table 2 Comparison of the proposed model with other models on the Salinas dataset
[0132]
[0133]
[0134] Example 4, experiments and results of a hyperspectral image classification method based on a spectral-spatial dual fusion network (S2DFT) on a Botswana dataset.
[0135] 4.1. Dataset Introduction
[0136] The Botswana dataset was acquired by NASA's EO-1 satellite over the Okavango Delta and contains a 7.7 km swath of data collected by the Hyperion sensor. The dataset has a spatial resolution of 30 m / pixel and originally included 242 spectral bands covering the 400–2500 nm range with an interval of 10 nm. After removing the absorption bands, 145 spectral bands remained. The dataset is divided into 14 different categories.
[0137] 4.2 Experimental Setup
[0138] 4.2.1 Hyperparameters
[0139] For our method, the Adam optimizer is used to update the network parameters, and the learning rate is set to 5e-4. For the training epoch, the training process stops when the training loss tends to be stable. The training epoch is uniformly set to 50, under which the training loss is stable, and the batch size is set to 64. All comparison methods use the code publicly provided by the original author and are re-experimented on the same hardware configuration to ensure the consistency of the experiments.
[0140] 4.2.2 Training Samples
[0141] In order to evaluate the classification performance of the model under limited sample conditions, we divided the number of samples used in the experiment. We selected 2% of the samples from the Botswana dataset for training, and all training samples were randomly selected to ensure the stability of the algorithm. In addition to the training samples, the rest were used for testing.
[0142] 4.2.3 Operating platform and indicators
[0143] To ensure a fair comparison, the proposed method and all the compared methods are implemented in the PyTorch framework based on PyTorch 2.0.0 and cuda11.8. Both training and testing are performed on devices with Intel(R) Core(TM) i7-14700 and NVIDIAGeForce RTX 4060 (8GB GPU) graphics cards. In order to quantitatively evaluate the classification effect of each method, four commonly used evaluation metrics are used. Single-class accuracy represents the ratio of the number of correctly classified samples in each class to the total number of samples in that class. Overall accuracy (OA) is the ratio of all correctly classified samples to the total number of all samples. Average accuracy (AA) is the arithmetic mean of all single-class accuracies. Kappa coefficient (κ) is used to measure the consistency between the model classification results and the actual labels. For all these metrics, higher values indicate better performance of the method.
[0144] 4.3 Experimental Analysis
[0145] We compare the proposed method with other advanced deep learning methods, including 5 CNN-based methods and 5 Transformer-based methods. Specifically, the CNN-based methods include 3-DCNN, deep feature fusion network (DFFN), HybridSN, double-branch dual-attention (DBDA), and compact band weighting (CBW). Transformer-based methods include SpectralFormer, Spectral–Spatial Feature Tokenization Transformer (SSFTT), Group-Aware Hierarchical Transformer (GAHT), and MorphFormer. The parameter settings of these methods are the same as those in the corresponding articles.
[0146] On the Botswana dataset, the classification accuracy table and classification diagram of each method are shown in Table 3 and Figure 7 shown.
[0147] Figure 7(a) 3D-CNN (78.86%), (b) DFFN (81.49%), (c) HybridSN (86.17%), (d) DBDA (88.72%), (e) CBW (88.00%), (f) SF (82.37%), (g) GAHT (92.39%), (h) SSFTT (91.66%), (i) morphformer (85.23%), (j) S2DFT (94.67%). The analysis of the result data shows that the S2DFT proposed in this paper is superior to other comparison methods and has achieved the best results in terms of OA, AA and Kappa coefficient. Using only 2% of the training samples, the OA, AA and Kappa coefficients of our S2DFT model reach 94.67%, 94.11%, and 94.23%, respectively, and are 1.28%-13.18%, 3.06%-19.15%, and 1.41%-17.21% higher than the other methods, respectively. Despite the limited training samples and the imbalanced distribution of sample categories, our proposed S2DFT model achieves better classification results than the other nine methods. Among the compared algorithms, the hybrid model HybridSN based on 2DCNN and 3DCNN has a strong feature extraction capability, but its OA value is only 86.17%, which is significantly lower than the other methods when the training samples are limited. The DBDA model shows significant improvement over the HybridSN by integrating an effective attention mechanism into the spectral and spatial feature extraction branches. GAHT introduces a grouped pixel embedding (GPE) module and restricts the multi-head self-attention to the local spectral space, which performs best among the compared algorithms. S2DFT performs well in the accuracy of all 7 categories, especially for the 4th and 12th categories, the accuracy of our proposed model reaches 100%.
[0148] Table 3 Comparison of the proposed model with other models on the Botswana dataset
[0149]
[0150]
[0151] Example 5, experiments and results of a hyperspectral image classification method based on a spectral-spatial dual fusion network (S2DFT) on the WHU-Hi-LongKou dataset.
[0152] 5.1. Dataset Introduction
[0153] The WHU-Hi-LongKou dataset was collected on July 17, 2018 in Longkou, Hubei, China, using a Headwall Nano-Hyperspec imaging sensor with an 8 mm focal length. The image size is 550×400 pixels, including 270 spectral bands with a wavelength range of 0.400 to 1.0 μm. The study area is a simple agricultural scene covering nine types of objects, including six types of crops.
[0154] 5.2 Experimental Setup
[0155] 5.2.1 Hyperparameters
[0156] For our method, the Adam optimizer is used to update the network parameters, and the learning rate is set to 5e-4. For the training epoch, the training process stops when the training loss tends to be stable. The training epoch is uniformly set to 50, under which the training loss is stable, and the batch size is set to 64. All comparison methods use the code publicly provided by the original author and are re-experimented on the same hardware configuration to ensure the consistency of the experiments.
[0157] 5.2.2 Training Samples
[0158] In order to evaluate the classification performance of the model under limited sample conditions, we divided the number of samples used in the experiment. 0.2% of the samples were selected for training on the WHU-Hi-LongKou dataset, and all training samples were randomly selected to ensure the stability of the algorithm. In addition to the training samples, the rest were used for testing.
[0159] 5.2.3 Operating platform and indicators
[0160] To ensure a fair comparison, the proposed method and all the compared methods are implemented in the PyTorch framework based on PyTorch 2.0.0 and cuda11.8. Both training and testing are performed on devices with Intel(R) Core(TM) i7-14700 and NVIDIAGeForce RTX 4060 (8GB GPU) graphics cards. In order to quantitatively evaluate the classification effect of each method, four commonly used evaluation metrics are used. Single-class accuracy represents the ratio of the number of correctly classified samples in each class to the total number of samples in that class. Overall accuracy (OA) is the ratio of all correctly classified samples to the total number of all samples. Average accuracy (AA) is the arithmetic mean of all single-class accuracies. Kappa coefficient (κ) is used to measure the consistency between the model classification results and the actual labels. For all these metrics, higher values indicate better performance of the method.
[0161] 5.3 Experimental Analysis
[0162] We compare the proposed method with other advanced deep learning methods, including 5 CNN-based methods and 5 Transformer-based methods. Specifically, the CNN-based methods include 3-DCNN, deep feature fusion network (DFFN), HybridSN, double-branch dual-attention (DBDA), and compact band weighting (CBW). Transformer-based methods include SpectralFormer, Spectral–Spatial Feature Tokenization Transformer (SSFTT), Group-Aware Hierarchical Transformer (GAHT), and MorphFormer. The parameter settings of these methods are the same as those in the corresponding articles.
[0163] On the WHU-Hi-LongKou dataset, the classification accuracy table and classification diagram of each method are shown in Table 4 and Figure 8 shown.
[0164] Figure 8 (a) 3DCNN (92.37%), (b) DFFN (95.85%), (c) HybridSN (96.45%), (d) DBDA (96.43%), (e) CBW (95.31%), ( f) SF (96.25%), (g) GAHT (96.85%), (h) SSFTT (97.18%), (i) morphformer (96.23%), (j) S2DFT (98.44%).
[0165] As can be seen from the table, the classification performance of the CNN-based and Transformer-based methods on this dataset is close, and all algorithms have achieved good classification performance. However, S2DFT is still competitive, with OA of 98.44%, AA of 95.25%, and Kappa of 97.97%. The OA and Kappa coefficients are 1.26%-6.07% and 0.97%-10.61% ahead, respectively, and AA is only 0.16% lower than SSFTT. In addition, the table shows that S2DFT has the lowest standard deviation, indicating that its results have higher stability and consistency. The SSFTT method combines the structural features of hybrid CNN and Transformer to achieve a joint extraction process of fine spectral spatial features in different scenarios. At the same time, as the method that performs closest to S2DFT, its OA is 97.18%, showing a strong competitive advantage in the performance of a single category.
[0166] Table 4 Comparison of the proposed model with other models on the WHU-Hi-LongKou dataset
[0167]
[0168]
Claims
1. A hyperspectral image classification method based on spectral-spatial dual fusion network, characterized in that: The following steps are involved: Step 1: Hyperspectral image dataset preparation There are four hyperspectral image datasets, namely Pavia University dataset, Salinas dataset, Botswana dataset and WHU-Hi-LongKou dataset; Step 2: Data Preprocessing Input, hyperspectral image data X∈R H×W×S :Here H represents the image height, W represents the image width, S represents the number of spectral bands, which is usually three-dimensional data composed of dozens to hundreds of bands. Each pixel contains a multidimensional spectral feature vector, and the ground truth label Y∈R H×W : Represents the classification label of the image, which is a two-dimensional array that matches the spatial dimension of the hyperspectral image data. It is used to guide the model to learn the correct classification or regression task during the training process. Each pixel is labeled as a certain category; Number of training iterations N: This is a hyperparameter used to control the number of iterations of model training; Principal component analysis, hyperspectral images usually have very high spectral dimensions, but they contain a lot of redundant information, which increases the computational complexity. PCA converts X into low-dimensional data X' and retains the principal components to maximize the variance distribution of the data; Data segmentation: Divide the data into training set and test set: training set X train ={x train ,y train }: used for model learning, test set X test ={x test ,y test }: used for model performance verification. In order to evaluate the classification performance of the model under limited sample conditions, we divided the number of samples used in the experiment and randomly selected samples from the dataset for training; Step 3: Hyperspectral image classification task After using principal component analysis to reduce the dimension, the spatial feature extraction module and the spectral feature extraction module are designed to extract shallow spatial spectral features; We designed a channel attention module that combines convolution with RELU to perform weighted output of features extracted by SPAEM; Design multi-head spectral spatial attention to replace multi-head attention and obtain spatial-spectral interaction features; The spatial-spectral interaction features and the spectral weighted features are fused and classified using global average pooling and fully connected layers; Step 4: Set up the optimizer and loss function Select the appropriate optimizer and loss function according to the complexity of the model and the requirements of the hyperspectral image classification task. The specific solution is as follows: Loss calculation, after feature extraction, uses the cross entropy loss function K CE To calculate Y pretrain and the true label Y train The error between , the cross entropy loss function is one of the commonly used loss functions in classification problems, and the loss function is expressed as: Among them, y i represents the true category label, Represents the predicted probability, M is the number of samples, and by calculating the loss function, we can quantify the accuracy of the model prediction; Optimize model parameters, use Adam optimizer to update model weights, and combine exponential decay learning rate scheduler Learning Rate Scheduler to optimize model parameters; the learning rate adjustment formula is as follows: or t =η0·exp(-λ·t Among them, η t is the learning rate of the tth round, η0 is the initial learning rate, and λ is the decay rate hyperparameter; Step 5: S2DFT model training and saving The goal of the training phase is to use the existing label image Y and the input hyperspectral image data X to train the S2DFT model so that it can learn effective features and make accurate predictions. The trained S2DFT model is saved to prepare for subsequent testing. Training cycle, iterative process, training cycle from i = 1 to N, this process will be repeated until the model parameters reach the optimal state, the training cycle is the core of the model training process, through continuous iterative optimization, the model can gradually learn the inherent laws and characteristics of the data, in each iteration, the model will go through a complete forward propagation and back propagation process, so as to gradually adjust its internal parameters and improve the prediction accuracy; Model saving, when the training cycle is repeated N times, the model training is completed. At this time, the parameters of S2DFT have been adjusted to the optimal state, the prediction performance of the model reaches the expected goal, and the optimal model is saved. Step 6: S2DFT model testing phase The test set sample x test Input into the trained S2DFT model. In this stage, only forward propagation is performed, and back propagation and parameter update are not required. The model outputs the predicted feature vector y for each test sample predict By comparing the feature vector with the classification threshold, the model generates the corresponding prediction label; These predicted labels are compared with the true labels y test Comparison is used to evaluate the performance of the model.
2. The hyperspectral image classification method based on spectral-spatial dual fusion network as claimed in claim 1, characterized in that: The step 3 initializes the parameters of the spectral feature extraction module and the spatial feature extraction module according to the nature of the task; The SPEEM module extracts spectral information from the input image through a series of convolution operations and maps it to a high-dimensional embedding space through feature transformation; Hyperspectral images usually contain hundreds of bands. Through the SPEEM module, the model can effectively extract valuable spectral features from these bands. Initialize the SPAEM module, which extracts spatial features in the image by using multi-scale convolution kernels and spatial filtering operations; The SPAEM module focuses on capturing spatial information. This module uses convolution kernels of multiple scales, which can extract detailed features at different spatial resolutions. Initialize the MHS3A module, which receives the feature vector from the previous module as input. The input feature vector is first transformed through a linear layer, which is regarded as the generator of query, key and value in the self-attention mechanism; In addition, the input feature vector is also multiplied by the spectral weights. These weights are obtained through pre-training and are used to emphasize the bands in the input data that are more helpful for the classification task. The linearly transformed feature vector is sent to the dot product attention module. This module is the core of the self-attention mechanism and is used to capture the dependencies between different parts of the input data. In this module, the input data is first mapped into three vectors: query, key, and value. Then, the dot product of the query and the key is calculated to obtain an attention score matrix. This score matrix is scaled, usually divided by a scaling factor, which is the square root of the dimension of the key vector. The scaled attention scores are converted into probability distributions through the Softmax function. This distribution represents the relative importance of different parts of the input data. Finally, this probability distribution is multiplied by the value vector to obtain weighted value vectors. These vectors aggregate the information of different parts of the input data. The MHS3A module contains multiple parallel self-attention heads, each of which performs the above attention calculations independently. The outputs from multiple attention heads are spliced together to form a comprehensive feature representation. The spliced features are transformed through a linear layer to further integrate the features and prepare the output. Finally, the linearly transformed features are sent to the output layer to generate the final classification results. During the training process, the Adam optimizer is initialized and used to optimize the model's training process. The SPEEM module extracts spectral information through convolutional layers and performs feature transformation to map this information into a high-dimensional embedding space.
3. The hyperspectral image classification method based on spectral-spatial dual fusion network as claimed in claim 1, characterized in that: Step 5 described above uses learning rate decay to further adjust the optimization process.
4. The hyperspectral image classification method based on spectral-spatial dual fusion network as claimed in claim 2, characterized in that: The Adam optimizer combines momentum and adaptive learning rate adjustment. By continuously adjusting the model parameters, the optimizer will guide the model to minimize the loss function and improve the performance of the model in the task. In the classification task, the cross entropy loss function can be optimized according to the difference between the category label and the probability distribution of the model output.
5. The hyperspectral image classification method based on spectral-spatial dual fusion network as claimed in claim 2, characterized in that: The MHS3A module further strengthens the interaction between spectral and spatial information through a multi-head self-attention mechanism, and can accurately capture key features in the image at both the global and local levels. The combination of the two ensures that the model can not only extract detailed spectral features but also capture comprehensive spatial information when processing images, greatly improving the model's ability in complex image analysis.
Citation Information
Cited By
Method and device for constructing leaf detection model based on hyperspectral imaging technology
CN120876408A
Remote sensing hyperspectral image classification method, system, equipment and medium
CN121280788A