Hyperspectral image classification method, device and equipment based on multi-attention fusion

By constructing a hyperspectral image classification method with a three-branch framework and a multi-attention fusion module, the problem of low classification accuracy of hyperspectral images is solved, and higher classification accuracy and feature extraction capability are achieved, making it suitable for hyperspectral image classification.

CN120997665APending Publication Date: 2025-11-21HUNAN UNIV OF CHINESE MEDICINE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511042621.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing hyperspectral image classification methods are insufficient in terms of detail and local feature extraction, resulting in low image classification accuracy, especially when the sample size is limited.

Method used

A hyperspectral image classification method based on multi-attention fusion is adopted, and a three-branch framework is constructed, including P, I and D branches. The multi-attention fusion module superimposes and fuses features at different scales, and different attention mechanisms are used to enhance feature extraction and alleviate the feature loss problem caused by the increase in depth.

Benefits of technology

It significantly improves the classification accuracy of hyperspectral images, especially when the number of samples is limited, and can achieve higher classification performance, enhancing the model's feature reuse capability and context information extraction capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997665A_ABST
    Figure CN120997665A_ABST
Patent Text Reader

Abstract

The invention relates to a hyperspectral image classification method, device and equipment based on multi-attention fusion, and the method comprises the steps: gradually and effectively extracting space-spectrum and cross-spectrum feature information from a plurality of scales from P, I and D branches with complementary responsibilities; meanwhile, features are simultaneously extracted through a three-branch framework to enhance the ability of the model to extract context information, and the loss of irrelevant and local features is remarkably reduced. Besides, the multi-attention branch fusion module performs feature stacking on processing results of different scales, and selects the most suitable attention mechanism to perform feature absorption, so that the feature reuse capability is effectively enhanced, and the problem of feature loss caused by increase of the depth of the model is relieved. Compared with a traditional machine learning method and an excellent deep learning method, the method can always obtain higher classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image classification, and relates to a hyperspectral image classification method, device and equipment based on multi-attention fusion. BACKGROUND

[0002] Hyperspectral technology is widely used in many fields, such as vegetation, soil properties, environmental monitoring, geophysical prospecting, etc. Hyperspectral images are rich in spectral and spatial information, and compared with natural images, they can better distinguish the physical differences between surface materials through extensive and dense spectral images, and are a field with high research activity. Hyperspectral technology has also received extensive attention in the field of remote sensing. In the field of remote sensing, hyperspectral data can provide more detailed spectral information for the description of ground objects, and use these continuous spectral information to image ground objects or urban surfaces. Therefore, hyperspectral data is widely used in ground object classification tasks and has achieved good classification results. Based on the advantages of high resolution and integration of image and spectrum of hyperspectral images, many researchers have conducted a lot of research on the classification algorithm thereof. In recent years, due to the unique spatial-spectral characteristics of hyperspectral images, they play an important role in many fields. Using deep neural networks for effective feature extraction, and then designing efficient and high-precision network algorithm structure gradually become the research hotspot of scholars. Hyperspectral images are difficult to obtain and have limited samples. Although the hyperspectral image classification method based on convolutional neural network (CNN) significantly improves the performance, there are still some deficiencies in detail and local feature extraction, which leads to low classification accuracy of images. SUMMARY

[0003] In view of the problems existing in the above-mentioned traditional method, the application provides a hyperspectral image classification method, device and equipment based on multi-attention fusion, which can improve the classification accuracy of images.

[0004] In order to achieve the above-mentioned purpose, the embodiments of the application adopt the following technical solutions: On the one hand, a hyperspectral image classification method based on multi-attention fusion is provided, which comprises the following steps: A hyperspectral image classification model based on multi-attention fusion is constructed; the hyperspectral image classification model comprises a three-branch framework inspired by a PID controller, a multi-attention fusion module and a classification module; the three-branch framework comprises a P branch, an I branch and a D branch; the P branch, the I branch and the D branch respectively adopt different sizes of convolution kernels and pooling layers to extract features of different scales from the hyperspectral image; the multi-attention fusion module is used for superimposing the features extracted by the P branch, the I branch and the D branch to obtain PI features, PID features and ID features, and then different attention mechanisms are used for feature extraction, and the extracted results are fused to obtain fused features.

[0005] The hyperspectral image is input into the I branch to obtain shallow features, deep features and I output features.

[0006] The shallow features are input into the P branch to obtain P output features.

[0007] The deep features are input into the D branch to obtain D output features.

[0008] The I output features, the P output features and the D output features are input into the multi-attention fusion module to obtain fusion features.

[0009] The fusion features are input into the classification module to obtain a hyperspectral image classification result.

[0010] In another aspect, a hyperspectral image classification device based on multi-attention fusion is also provided, which comprises: A hyperspectral image classification model construction module is configured to construct a hyperspectral image classification model based on multi-attention fusion; the hyperspectral image classification model comprises a three-branch framework, a multi-attention fusion module and a classification module; the three-branch framework, the multi-attention fusion module and the classification module are inspired by a PID controller; the three-branch framework comprises a P branch, an I branch and a D branch; the P branch, the I branch and the D branch respectively adopt different sizes of convolution kernels and pooling layers to extract features of different scales from the hyperspectral image; the multi-attention fusion module is configured to superimpose the features extracted by the P branch, the I branch and the D branch to obtain PI features, PID features and ID features, and then respectively adopt different attention mechanisms to extract features and fuse the extracted results to obtain fusion features.

[0011] A feature extraction module is configured to input the hyperspectral image into the I branch to obtain shallow features, deep features and I output features; input the shallow features into the P branch to obtain P output features; input the deep features into the D branch to obtain D output features; and input the I output features, the P output features and the D output features into the multi-attention fusion module to obtain fusion features.

[0012] A hyperspectral image classification module is configured to input the fusion features into the classification module to obtain a hyperspectral image classification result.

[0013] In another aspect, a computer device is also provided, which comprises a memory and a processor; the memory stores a computer program; and the processor implements the steps of the hyperspectral image classification method based on multi-attention fusion of any of the above aspects when executing the computer program.

[0014] One of the above technical solutions has the following advantages and beneficial effects: The hyperspectral image classification method, device and equipment based on multi-attention fusion described above effectively extracts spatial-spectral and cross-spectral feature information from multiple scales by gradually extracting from the P, I and D three branches with complementary responsibilities, simultaneously extracts features through the three-branch framework to enhance the ability of the model to extract context information, and significantly reduces the loss of irrelevant and local features. In addition, the multi-attention branch fusion module stacks the processing results of different scales and selects the most suitable attention mechanism for feature absorption, effectively enhancing the feature reuse capability and alleviating the feature loss problem caused by the increase of model depth. Compared with traditional machine learning methods and excellent deep learning methods, the method can always achieve higher classification accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0016] Figure 1 A flowchart of a hyperspectral image classification method based on multi-attention fusion in an embodiment; Figure 2 A block diagram of a hyperspectral image classification model structure based on multi-attention fusion in an embodiment; Figure 3 A schematic diagram of the P branch in an embodiment; Figure 4 A structural schematic diagram of the skip weight channel SWCA module in an embodiment; Figure 5 A schematic diagram of the D branch in an embodiment; Figure 6 A schematic diagram of the I branch in an embodiment; Figure 7 A structural schematic diagram of the multi-attention fusion module in an embodiment; Figure 8Figures showing classification results of various methods on the HT dataset in one embodiment, where (a) is the original HIS, (b) is the ground truth map, (c) is the classification results of MLR on the HT dataset, (d) is the classification results of SVM on the HT dataset, (e) is the classification results of RF on the HT dataset, (f) is the classification results of ELM on the HT dataset, (g) is the classification results of CNN2D on the HT dataset, (h) is the classification results of PPF on the HT dataset, (i) is the classification results of SF on the HT dataset, and (j) is the classification results of the present method on the HT dataset; Figure 9 Figures showing classification results of various methods on the SV dataset in one embodiment, where (a) is the original HIS, (b) is the ground truth map, (c) is the classification results of MLR on the SV dataset, (d) is the classification results of SVM on the SV dataset, (e) is the classification results of RF on the SV dataset, (f) is the classification results of ELM on the SV dataset, (g) is the classification results of CNN2D on the SV dataset, (h) is the classification results of PPF on the SV dataset, (i) is the classification results of SF on the SV dataset, and (j) is the classification results of the present method on the SV dataset; Figure 10 Figures showing classification results of various methods on the UP dataset in one embodiment, where (a) is the original HIS, (b) is the ground truth map, (c) is the classification results of MLR on the UP dataset, (d) is the classification results of SVM on the UP dataset, (e) is the classification results of RF on the UP dataset, (f) is the classification results of ELM on the UP dataset, (g) is the classification results of CNN2D on the UP dataset, (h) is the classification results of PPF on the UP dataset, (i) is the classification results of SF on the UP dataset, and (j) is the classification results of the present method on the UP dataset; Figure 11 Figures showing classification performance of various methods on the IP, SV, and UP datasets using fewer samples in one embodiment, where (a) is the classification performance of various methods on the IP dataset using fewer samples, (b) is the classification performance of various methods on the SV dataset using fewer samples, and (c) is the classification performance of various methods on the UP dataset using fewer samples; Figure 12 Classification performance of the present method on different numbers of samples in one embodiment.

[0017] Figure 13For the classification performance column chart of the combination of the three different attention mechanisms in the order of branch one, branch two and branch three of the method in one embodiment, (a) is the classification performance column chart of the combination of CBAM+GAM+NAM, (b) is the classification performance column chart of the combination of CBAM+NAM+GAM, (c) is the classification performance column chart of the combination of GAM+CBAM+NAM, (d) is the classification performance column chart of the combination of GAM+NAM+CBAM, (e) is the classification performance column chart of the combination of NAM+GAM+CBAM, and (f) is the classification performance column chart of the combination of NAM+CBAM+GAM used in this paper. DETAILED DESCRIPTION

[0018] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.

[0020] It should be noted that the term “embodiment” mentioned herein means that the specific features, structures or characteristics described in combination with the embodiments can be included in at least one embodiment of the present application. The phrase is shown at various places in the specification does not necessarily refer to the same embodiment, nor is it independent or alternative to other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments. The term “and / or” used herein refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0021] The embodiments of the present application will be described in detail below in combination with the drawings of the embodiments of the present application.

[0022] In one embodiment, as shown in Figure 1 a multi-attention fusion based hyperspectral image classification method can be provided, which can include the following processing steps 100 to 110: Step 100: Construct a hyperspectral image classification model based on multi-attention fusion; the hyperspectral image classification model comprises a PID controller inspired three-branch framework, a multi-attention fusion module, and a classification module; the three-branch framework comprises a P branch, an I branch, and a D branch; the P branch, the I branch, and the D branch respectively use different size convolution kernels and pooling layers to extract features of different scales from the hyperspectral image; the multi-attention fusion module is used to superimpose the features extracted by the P branch, the I branch, and the D branch to obtain PI features, PID features, and ID features, and then different attention mechanisms are used for feature extraction, and the extracted results are fused to obtain fused features.

[0023] Specifically, the hyperspectral image classification model based on multi-attention fusion uses a feature fusion method of first extracting features, then superimposing, then extracting features, and then fusing features, which cleverly extracts features of different branches multiple times, realizes the fusion of features of different scales, avoids the insufficient expression ability of single-scale features, and makes the feature map contain more rich information.

[0024] When fusing the features of the P branch, the I branch, and the D branch, the present application proposes a multi-attention fusion module, that is, different branches correspond to different attention, according to the appearance form of each branch, the most suitable attention mechanism is matched, and the features in each branch are fully extracted, and the image classification is more effectively performed.

[0025] The structure block diagram of the hyperspectral image classification model based on multi-attention fusion is as shown in Figure 2 .

[0026] Step 102: input the hyperspectral image into the I branch to obtain shallow features, deep features, and I output features.

[0027] Step 104: input the shallow features into the P branch to obtain P output features.

[0028] Step 106: input the deep features into the D branch to obtain D output features.

[0029] Step 108: input the I output features, the P output features, and the D output features into the multi-attention fusion module to obtain fused features.

[0030] Specifically, the features obtained by the P branch and the D branch are superimposed into the I branch, and the features of the I branch are also superimposed into the P branch and the D branch, which increases the nonlinearity of the network, helps the network to learn more complex features, and also enables the network to learn the feature representation of different branches.

[0031] The fusion process, employing three different attention mechanisms, involves one feature map stacking and one feature map connection. Using convolutional kernels and pooling layers of varying sizes in different branches to extract features at different scales effectively captures both local details and global contextual information. Furthermore, the branching structure allows the network to learn multiple feature representations in parallel, improving efficiency and performance.

[0032] Step 110: Input the fused features into the classification module to obtain the hyperspectral image classification result.

[0033] The aforementioned hyperspectral image classification method based on multi-attention fusion effectively extracts spatial-spectral and transspectral feature information from multiple scales progressively through three complementary branches (P, I, and D). Simultaneously, the three-branch framework enhances the model's ability to extract contextual information, significantly reducing the loss of irrelevant and local features. Furthermore, the multi-attention branch fusion module stacks features from processing results at different scales and selects the most suitable attention mechanism for feature absorption, effectively enhancing feature reuse and mitigating feature loss issues caused by increased model depth. Compared to traditional machine learning methods and advanced deep learning methods, this method consistently achieves higher classification accuracy.

[0034] In one embodiment, such as Figure 3 As shown, the P branch includes: two skip-weight channel SWCA modules and a first convolutional module; the first convolutional module includes convolutional layers of various scales; the skip-weight channel SWCA module includes convolutional layers, max pooling layers, and a Senet attention mechanism; in the first skip-weight channel SWCA module: shallow features are processed by convolution, batch normalization, and activation to obtain a new feature map; the shallow features and the new feature map are added together and then max pooled to obtain intermediate features; the intermediate features are input into the Senet attention mechanism to obtain weights; the Senet attention mechanism is used to perform a global average operation on the intermediate features and then normalize them to obtain normalized features; the normalized features are linearized using two learnable parameters and then processed by one-dimensional convolution and Sigmoid activation to obtain weights; the weights are multiplied element-wise with the intermediate features, and the result is processed by convolution, batch normalization, and ReLU activation to obtain the output features of the skip-weight channel SWCA module.

[0035] Specifically, in order to better extract features and enhance the attention degree of important features, the present application proposes a shortcut weight channel attention (SWCA) module, which connects features of different scales to make the features more fine. At the same time, this mechanism can make the network pay more attention to important features and improve the performance of the model. The structure of the shortcut weight channel attention (SWCA) module is shown in Figure 4 The SWCA module uses skip connection feature fusion and multiplies the weighted features with the original feature map to obtain the output feature map, thereby improving the ability of the model to extract prominent features.

[0036] The P branch analyzes and pre-saves detailed information in the high-resolution feature map.

[0037] Figure 3 is the detailed structure of the P branch, which includes two shortcut weight channel attention (SWCA) modules and a series of convolutional layers, and finally obtains the final feature map through feature superposition and NAM attention mechanism processing. The input of the P branch is the output feature (i.e. shallow feature) obtained after the first SWCA processing of the original feature map. In the P branch, there are two shortcut weight channel attention (SWCA) modules for feature extraction, and two 1x3 convolutions (CNNs) to make the feature map into a regular MxM size, and then six CNNs with convolution kernel sizes of 7x7, 7x7, 4x4, 3x3, 3x3, and 3x3 are used for multi-scale spatial joint feature extraction. The output obtained is superimposed with the output feature of the I branch, and finally the most suitable NAM attention is used to obtain the final output.

[0038] The normalized attention mechanism (NAM) is an efficient and lightweight attention mechanism that integrates the CBAM module and redesigns the channel and spatial attention submodules. For the channel attention submodule, the scaling factor in batch normalization (BN) is used, as shown in the following formula. The scaling factor measures the variance of the channels and indicates their importance.

[0039]

[0040] where and are the standard deviations of batch B, and are trainable affine transformation parameters (scale and shift).

[0041] The output feature of the channel attention submodule (Channel Attention Module) in NAM is: where is the weight.

[0042] Applying the scale factor of Batch Normalization to the spatial dimension, we name it as pixel normalization. The output feature of the spatial attention sub-module (Spatial Attention Module) in the corresponding NAM is denoted as , where is the weight, and and are the scaling factors, respectively.

[0043] In an embodiment, the D branch includes: a skip weight channel (SWCA) module and a second convolutional module; the second convolutional module includes convolutional layers of multiple different scales; the skip weight channel (SWCA) module includes a convolutional layer, a max-pooling layer, and a Senet attention mechanism.

[0044] Specifically, the D (derivative) branch extracts high-frequency features to predict the boundary region.

[0045] Figure 5 is the detailed structure of the D branch, the input of which is the output feature of the branch CBAM after two SWCA extractions, which has reached the deep stage of feature extraction, and the branch has only one SWCA module feature extraction and one 1x3 CNN. Six times of convolution kernel size of 5x5, 5x5, 5x5, 5x5, 4x4, and 3x3 CNN are used for multi-scale spatial joint feature extraction. The output obtained is stacked with the output of the I branch, and finally the most suitable global attention mechanism (GAM) for this branch is used to obtain the final output.

[0046] The GAM attention mechanism improves the performance of the network by reducing information diffusion and putting global interaction representation. It contains two modules: a channel attention module and a spatial attention module. The channel attention sub-module uses a three-dimensional arrangement to preserve information in three dimensions. Then, it amplifies the cross-dimensional channel-space dependency with a two-layer MLP. (The MLP is an encoder-decoder structure, the same as BAM, and its compression ratio is r). In the spatial attention sub-module, in order to focus on spatial information, two convolutional layers are used for spatial information fusion, and the same reduction ratio r as BAM is also used from the channel attention sub-module. At the same time, since the max-pooling operation reduces the use of information, it has a negative impact. Here, the pooling operation is deleted to further preserve the characteristic mapping, and the overall structure is as shown in Figure 2 . Given the input feature map F1∈R C×H×W , the intermediate state and are defined as: ; Represent input, will After three-dimensional arrangement and MLP operation to , after sigmoid activation , after feature multiplication . After a 7x7 convolution layer, channel reduction is performed, and then a new feature map is obtained after a 7x7 convolution, and after an activation function to Finally , after feature multiplication , the final input .

[0048] In an embodiment, the I branch includes a basic convolution module, two skip weight channel SWCA modules, a CBAM attention mechanism, and a third convolution module including convolution layers of multiple different scales; the skip weight channel SWCA module includes a convolution layer, a maximum pooling layer, and a Senet attention mechanism; and the step 102 includes: inputting the hyperspectral image into the basic convolution module of the I branch to obtain convolution features; inputting the convolution features into the first skip weight channel SWCA module of the I branch to obtain shallow layer features; inputting the shallow layer features into the second skip weight channel SWCA module of the I branch to obtain deep layer features; inputting the deep layer features after convolution processing into the CBAM attention mechanism of the I branch to obtain attention features; and inputting the attention features into the third convolution module of the I branch to obtain I output features.

[0049] Specifically, the I branch aggregates context information locally and globally to resolve long-range dependencies.

[0050] The detailed structure of the I branch is as shown in Figure 6As shown, the I branch is also the main branch of the network throughout the entire feature extraction. The I branch also undergoes twice the feature extraction of the skip weight channel SWCA module, and uses three 1x3 CNNs to make the feature map into a regular MxM size, and uses six CNNs with sizes of 7x7, 5x5, 4x4, 3x3, 3x3, and 3x3 for joint feature extraction in multiple scales. The difference is that to enhance feature propagation and greatly improve feature expression capability, the CBAM branch introduces a channel-spatial attention mechanism (CBAM) after each CNN feature extraction after being regularized. In the case of slightly increasing the amount of calculation and parameters, the model performance is greatly enhanced. Finally, as the main branch, the I branch superimposes the output features of the P branch and the I branch, and obtains the output after CBAM attention. Convolutional attention (CBAM) is a lightweight general module, including a channel attention mechanism (Channel Attention Module) and a spatial attention mechanism (Spatial Attention Module). In the feature map, CBAM will infer the attention map along two independent dimensions (channels and space), and then multiply the attention map and the input feature map to perform adaptive feature optimization.

[0051] The channel attention mechanism (Channel Attention Module) is to compress the feature map in the spatial dimension to obtain a one-dimensional vector for operation. When compressing in the spatial dimension, not only the average value pooling (Average Pooling) is considered, but also the maximum value pooling (Max Pooling). The channel attention mechanism can be expressed as:

[0052] The spatial attention mechanism (Spatial Attention Module) is to compress the channel, and average value pooling and maximum value pooling are performed in the channel dimension. The operation of MaxPool is to extract the maximum value in the channel, and the number of extraction times is high times width; the operation of AvgPool is to extract the average value in the channel, and the number of extraction times is also high times width; then the extracted feature maps (the number of channels is 1) are combined to obtain a 2-channel feature map. The spatial attention mechanism can be expressed as:

[0053] wherein, sigmoid operation, 7x7 represents the size of the convolution kernel, and the 7x7 convolution kernel is better than the 3x3 convolution kernel.

[0054] In an embodiment, as Figure 7As shown, the multi-attention fusion module includes three different attention mechanisms; step 108 includes: superimposing the P output feature and the D output feature into the I output feature to obtain the PID feature; superimposing the P output feature and the D output feature into the I output feature, respectively, to obtain the PI feature and the PD feature; processing the PI feature, the PID feature, and the PD feature through different attention mechanisms, respectively, to obtain the P branch output feature, the I branch output feature, and the D branch output feature; and performing flat connection on the P branch output feature, the I branch output feature, and the D branch output feature to obtain the fusion feature.

[0055] In one embodiment, the three different attention mechanisms are: the NAM attention mechanism, the CBAM attention mechanism, and the GAM attention mechanism; the P branch corresponds to the NAM attention mechanism; the I branch corresponds to the CBAM attention mechanism; and the D branch corresponds to the GAM attention mechanism.

[0056] Specifically, the embodiment adopts three different attention mechanisms, namely, NAM, CBAM, and GAM, and performs a plurality of different attention matching tests. Because of the different scale feature extraction and feature expression forms of the three branches, the most suitable attention mechanism for each branch is selected, and the feature expression is further enhanced. Finally, the feature maps obtained by the three branches are flatly connected, and are converted into the final classification prediction result through a fully connected layer.

[0057] In one embodiment, the classification module includes a fully connected layer.

[0058] In some embodiments, experimental examples are also provided. Two hyperspectral datasets, Houston 2013, Salinas Valley (SV), and University of Pavia, are used to evaluate the method proposed herein. For each dataset, the numerical values are first normalized to 0-1, and then for each class, 200 pixels and their surrounding 5x5 neighborhoods are randomly selected as training samples, and the remaining pixels are used as test samples.

[0059] The Houston (HT) dataset has a data size of 349x1905 and contains 144 bands with a spectral range from 364 nm to 1046 nm. The ground truth labels include 15 land cover classes such as trees and soil, with a total of 15029 pixels. The selected classes and sample numbers are shown in Table 1: Table 1. Sample selection of Houston dataset

[0060] The Salinas Valley (SV) dataset, after removing 20 water vapor and noise bands, retains 204 bands. The data size is 512×217 pixels, with a spatial resolution of 3.7 meters. Ground truth markers are divided into 16 land cover categories; the specific land cover types and pixel counts are shown in Table 2. Table 2 Sample selection for the Salinas Valley dataset

[0061] After noise removal and other band processing, the University of Pavia (UP) dataset contains 103 bands. The image size is 610×340 pixels with a spatial resolution of 1.3 meters. The spectral range is 0.43–0.86 micrometers. Approximately 20% of the pixels are labeled as ground truth, covering various urban structures, soil, natural objects, and shadows. The specific pixel counts are shown in Table 3.

[0062] Table 3 Sample selection for the University of Pavia dataset

[0063] (1) Parameter settings The model parameters were set empirically: the Adam optimizer was used, the learning rate was set to 0.00001, the batch size was set to 256, and the loss function was cross-entropy loss. To evaluate the effectiveness of this method, support vector machines (SVM), multinomial logistic regression (MLR), extreme learning machines (ELM), random forests (RF), CNN, PPF, SpectralFormer, RAMiT, MSAA, and SAFM were used. These methods were experimented on the same training and test sets. The classification accuracy of various methods was evaluated using overall accuracy (OA), average accuracy (AA), and the Kappa coefficient.

[0064] (2) Classification performance on the Houst dataset The classification results of each method on the HT dataset are shown below. Figure 8 As shown, Figure 8 (a) is the original HIS. Figure 8 (b) is the ground truth map. Figure 8 (c) shows the classification results of MLR on the HT dataset. Figure 8 (d) is a schematic diagram of the classification results of SVM on the HT dataset.Figure 8 Figure (e) is a schematic diagram of the classification results of RF on the HT dataset, Figure 8 Figure (f) is a schematic diagram of the classification results of ELM on the HT dataset, Figure 8 Figure (g) is a schematic diagram of the classification results of CNN2D on the HT dataset, Figure 8 Figure (h) is a schematic diagram of the classification results of PPF on the HT dataset, Figure 8 Figure (i) is a schematic diagram of the classification results of SF on the HT dataset, Figure 8 Figure (j) is a schematic diagram of the classification results of the method of the present application on the HT dataset. It can be seen that the traditional machine learning method performs poorly, with many classification errors, and the effect of each method on soil is also good. The deep learning-based methods all perform well, and SAFM and the method of the present application perform well in each category, and the method of the present application is the most prominent, with better classification performance.

[0065] As shown in Table 4, the classification accuracy of each method in each category, as well as the overall accuracy, average accuracy and Kappa coefficient are shown. From the table, it can be seen that the accuracy of each method in the grass_synergic category is also more than 98%. In other categories, MLR performs much worse than other methods, resulting in the lowest OA and Kappa coefficient. SF and SFAM are better, and the method of the present application performs well in each category, with an OA that is 1.09% higher than the second-ranked SAFM and nearly 40% higher than the worst MLR. From the results in the table, it can be seen that the method of the present application has higher classification performance.

[0066] Table 4 Classification accuracy of each method in each category on the HT dataset

[0067] (3) Classification performance on the SV dataset Figure 9 The performance of various methods on the SV dataset is given, wherein Figure 9 Figure (a) is the original HIS, Figure 9 Figure (b) is the ground truth map, Figure 9 Figure (c) is the classification results of MLR on the SV dataset, Figure 9 Figure (d) is the classification results of SVM on the SV dataset, Figure 9 Figure (e) is the classification results of RF on the SV dataset, Figure 9 Figure (f) is the classification results of ELM on the SV dataset, Figure 9 Figure (g) is the classification results of CNN2D on the SV dataset, Figure 9 Figure (h) is the classification results of PPF on the SV dataset, Figure 9The classification results of SF on the SV dataset, Figure 10 The classification results of the proposed method on the SV dataset. It can be seen that the traditional machine learning methods perform poorly and make many classification errors, but each method performs well on the Lettuce-romaine-5wk class. Several methods based on deep learning perform well on the Vinyard-vertical-trellis class. SF and SAFM perform well in general, but the proposed method performs particularly well and has better classification performance.

[0068] According to Table 5, the classification accuracy of each method is extremely style. Among them, the accuracy of each method in the Fallow-rough-plow class is close to 100%, and the accuracy of the Lettuce-romaine-5wk class is also more than 99%. However, on other classes, MLR and RF perform significantly worse than other methods, resulting in the lowest OA and Kappa coefficient. The accuracy of SVM and ELM is not much different, while the performance of CNN, PPF and SF is slightly better, with high accuracy and not much difference. The proposed method performs well on each class, with OA, AA and Kappa coefficient values higher than other methods. The results in the table show that the proposed method has better classification performance.

[0069] Table 5 Classification accuracy of each method on each class of the SV dataset

[0070] (4) Classification performance on the UP dataset Figure 11 The performance of various methods on the UP dataset is shown. It can be observed that in the class with the most pixels, such as Asphalt, Bare Soil and Meadows, the classification errors of each method can reflect their classification performance. MLR shows the most classification errors, followed by RF and ELM. SF performs well in classification, but has more misclassification in the Bare Soil class. CNN and PPF perform very well in misclassification, but the proposed method still has the best classification performance.

[0071] As shown in Table 6, the objective evaluation metrics demonstrate that all methods exhibit excellent classification performance in the Shadows and Painted metal sheets classes. Overall, MLR has the lowest OA (Objective Occlusion), SVM performs best among the four traditional machine learning algorithms, but is slightly inferior to deep learning-based methods. CNN and PPF achieve classification accuracies of 99.02% and 97.87% respectively, exhibiting very good classification performance. SF's classification accuracy is relatively lower, particularly in the Asphalt class. The proposed method achieves an overall accuracy of 99.65%, indicating that it possesses the highest classification performance on this dataset.

[0072] Table 6: Classification performance of each method on the SV dataset

[0073] (5) Classification performance on different training samples To compare the performance of various methods under the influence of fewer training samples, this paper adopts the same approach, setting the number of training samples for each class to 50, 100, and 200 respectively. To further compare the performance of various methods under the influence of fewer training samples, the classification results of each method are calculated, as shown below. Figure 11 As shown, Figure 11 (a) shows the classification performance on the IP dataset using fewer samples. Figure 11 (b) is a schematic diagram illustrating the classification performance on the SV dataset using a smaller number of samples. Figure 12 Figure (c) illustrates the classification performance on the UP dataset with a smaller number of samples. It is clear that as the number of samples increases, the classification performance of each method on each dataset shows a significant upward trend, which is consistent with our expectations. Traditional machine learning algorithms like SVM still exhibit relatively good classification performance because their trend is relatively stable. When there are individual samples with large biases in the training samples, the output matrix becomes ill-conditioned because the weight matrix from the input neurons to the hidden neurons and the hidden layer thresholds in the ELM method are randomly generated. This results in an unstable network structure, poor robustness, and reduced classification performance, thus leading to unstable performance on these three datasets. Deep learning-based methods significantly outperform traditional machine learning methods on all three datasets regardless of the number of training samples, and the method mentioned in this paper consistently demonstrates the best classification ability.

[0074] (6) Classification performance under different sample sizes Due to the limited number of hyperspectral image samples, this embodiment employs a random generation method, generating 100,000 samples in each class to increase the sample size and improve classification performance. Simply put, a larger sample size allows for the learning and extraction of more detailed features, but it also significantly increases computational cost. To evaluate the impact of sample size on classification performance, this paper sets four sample sizes: 10,000, 50,000, 100,000, and 200,000, using the proposed network structure to evaluate the final classification performance. The results are as follows... Figure 13 As shown, it can be seen that the classification performance of the network gets better and better as the number of samples increases. When the number of samples reaches 100,000, the classification accuracy basically tends to be at a horizontal level. In order to obtain the highest classification performance, the number of test samples in this paper is set at 100,000.

[0075] This paper uses three different attention branches with varying sizes and scales. To compare the performance of NAM, GAM, and CBAM attention on each of the three branches, a total of six different combination structures are used. Following the same approach, the classification performance of each combination structure is calculated, and a test sample of 100,000 is selected. The results are as follows: Figure 13 As shown, Figure 13 (a) is a bar chart showing the classification performance of the CBAM+GAM+NAM combination. Figure 13 (b) is a bar chart showing the classification performance of the CBAM+NAM+GAM combination. Figure 13 (c) is a bar chart showing the classification performance of the GAM+CBAM+NAM combination. Figure 13 The middle (d) is a bar chart showing the classification performance of the GAM+NAM+CBAM combination. Figure 1 The middle (e) is a bar chart showing the classification performance of the NAM+GAM+CBAM combination. Figure 1 Table (f) shows the classification performance of the NAM+CBAM+GAM combination used in this application. It can be seen that because the feature maps processed by each branch have different dimensionalities, using different attention mechanisms has a certain impact on classification performance. Comparing the OA, AA, and Kappa coefficients of the six different combinations, it is clear that the combination structure of branch 1 (NAM), branch 2 (CBAM), and branch 3 (GAM) used in this paper is significantly more stable than the other five combinations, and its classification accuracy on the three datasets is also much higher than the other five.

[0076] It should be understood that, although the above process ​ The steps in the diagram are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order in which these steps are executed; they can be performed in other orders. Furthermore, the above process... ​At least one part of the steps can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the sub-steps or stages is not necessarily sequential, but can be alternately executed with other steps or sub-steps or stages of other steps.

[0077] In one embodiment, a hyperspectral image classification device based on multi-attention fusion is also provided, and the device comprises: The hyperspectral image classification model construction module is configured to construct a hyperspectral image classification model based on multi-attention fusion; the hyperspectral image classification model comprises a three-branch framework, a multi-attention fusion module, and a classification module; the three-branch framework, the multi-attention fusion module, and the classification module are inspired by a PID controller; the three-branch framework comprises a P branch, an I branch, and a D branch; the P branch, the I branch, and the D branch respectively use convolution kernels of different sizes and pooling layers to extract features of different scales from the hyperspectral image; the multi-attention fusion module is configured to superimpose the features extracted by the P branch, the I branch, and the D branch to obtain PI features, PID features, and ID features, and then use different attention mechanisms to extract features, and fuse the extracted results to obtain fused features.

[0078] The feature extraction module is configured to input the hyperspectral image into the I branch to obtain shallow features, deep features, and I output features; input the shallow features into the P branch to obtain P output features; input the deep features into the D branch to obtain D output features; and input the I output features, the P output features, and the D output features into the multi-attention fusion module to obtain the fused features.

[0079] The hyperspectral image classification module is configured to input the fused features into the classification module to obtain a hyperspectral image classification result.

[0080] In one embodiment, the P branch includes two skip weight channel (SWCA) modules and a first convolution module; the first convolution module includes convolution layers of multiple different scales; the skip weight channel (SWCA) module includes a convolution layer, a max-pooling layer, and a Senet attention mechanism; in the first skip weight channel (SWCA) module: the shallow features are processed through convolution, batch normalization, and activation to obtain new feature maps; the shallow features and the new feature maps are added and then subjected to max-pooling to obtain intermediate features; the intermediate features are input into the Senet attention mechanism to obtain weights; the Senet attention mechanism is used to perform global average operation on the intermediate features and then perform normalization processing to obtain normalized features; the normalized features are linearly processed using two learnable parameters, subjected to one-dimensional convolution operation, and activated by a Sigmoid function to obtain the weights; and the weights are multiplied by corresponding elements of the intermediate features, and the obtained results are subjected to convolution, batch normalization, and ReLU activation to obtain the output features of the skip weight channel (SWCA) module.

[0081] In one embodiment, the D branch in the feature extraction module includes a skip weight channel (SWCA) module and a second convolution module; the second convolution module includes convolution layers of multiple different scales; and the skip weight channel (SWCA) module includes a convolution layer, a max-pooling layer, and a Senet attention mechanism.

[0082] In one embodiment, the I branch includes a base convolution module, two skip weight channel (SWCA) modules, a CBAM attention mechanism, and a third convolution module; the third convolution module includes convolution layers of multiple different scales; the skip weight channel (SWCA) module includes a convolution layer, a max-pooling layer, and a Senet attention mechanism; and the feature extraction module is further configured to input the hyperspectral image into the base convolution module of the I branch to obtain convolution features; input the convolution features into the first skip weight channel (SWCA) module of the I branch to obtain shallow features; input the shallow features into the second skip weight channel (SWCA) module of the I branch to obtain deep features; input the deep features after convolution processing into the CBAM attention mechanism of the I branch to obtain attention features; and input the attention features into the third convolution module of the I branch to obtain I output features.

[0083] In an embodiment, the multi-attention fusion module includes three different attention mechanisms; the feature extraction module is further configured to superimpose the P output feature and the D output feature into the PID feature mainly with the I output feature; superimpose the I output feature into the PI feature and the PD feature mainly with the P output feature and the D output feature respectively; process the PI feature, the PID feature, and the PD feature through different attention mechanisms respectively to obtain the P branch output feature, the I branch output feature, and the D branch output feature; and perform flat connection on the P branch output feature, the I branch output feature, and the D branch output feature to obtain the fusion feature.

[0084] In an embodiment, the three different attention mechanisms in the feature extraction module are: the NAM attention mechanism, the CBAM attention mechanism, and the GAM attention mechanism; the NAM attention mechanism corresponds to the P branch; the CBAM attention mechanism corresponds to the I branch; and the GAM attention mechanism corresponds to the D branch.

[0085] In an embodiment, the classification module in the feature extraction module includes a fully connected layer.

[0086] It can be understood that the specific explanations and descriptions of the hyperspectral image classification device based on multi-attention fusion can refer to the corresponding explanations and descriptions of the embodiments of the hyperspectral image classification method based on multi-attention fusion described above, which will not be repeated here. Each module in the above hyperspectral image classification device based on multi-attention fusion can be realized by software, hardware, and combinations thereof, in whole or in part. The above modules can be embedded in or independent of a device with data processing function in hardware form, or can be stored in the memory of the aforementioned device in software form, so that the processor can call and execute the operations corresponding to each module. The aforementioned device can be, but is not limited to, various types of data processing computer devices in the prior art.

[0087] In an embodiment, a computer device is also provided, which includes a memory and a processor, the memory stores a computer program, and the processor implements the processing steps of the above hyperspectral image classification method based on multi-attention fusion when executing the computer program.

[0088] It can be understood that the above computer device includes other software and hardware components not listed in the present specification in addition to the above-mentioned memory and processor, which can be determined according to the specific model of the image processing computer in different application scenarios, and will not be described in detail one by one in the present specification.

[0089] The technical features of the above embodiments can be combined arbitrarily, and to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present specification.

[0090] The above embodiments only express several implementation manners of the present application, which are described in a more specific and detailed manner, but cannot be understood as a limitation on the protection scope of the present application. It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, and all belong to the protection scope of the present application.

Claims

1. A hyperspectral image classification method based on multi-attention fusion, characterized in that, The method includes the following steps: Construct a hyperspectral image classification model based on multi-attention fusion; The hyperspectral image classification model includes a three-branch framework inspired by PID controllers, a multi-attention fusion module, and a classification module. The three-branch framework includes a P branch, an I branch, and a D branch. The P, I, and D branches respectively use convolutional kernels and pooling layers of different sizes to extract features of different scales from the hyperspectral image. The multi-attention fusion module is used to superimpose the features extracted by the P branch, I branch, and D branch to obtain PI features, PID features, and ID features. Then, different attention mechanisms are used to extract features, and the extracted results are fused to obtain fused features. The hyperspectral image is input into the I branch to obtain shallow features, deep features, and I output features; The shallow features are input into the P branch to obtain the P output features; The deep features are input into the D branch to obtain the D output features; The I output features, P output features, and D output features are input into the multi-attention fusion module to obtain the fused features; The fused features are input into the classification module to obtain the hyperspectral image classification result.

2. The hyperspectral image classification method based on multi-attention fusion according to claim 1, characterized in that, The P branch includes: two skip-weight channel SWCA modules and a first convolutional module; the first convolutional module includes convolutional layers of various scales; the skip-weight channel SWCA module includes convolutional layers, max pooling layers, and the Senet attention mechanism; In the first skip weight channel SWCA module: The shallow features are processed by convolution, batch normalization, and activation to obtain a new feature map; The shallow features and the new feature map are added together and then max pooling is performed to obtain the intermediate features; The intermediate features are input into the Senet attention mechanism to obtain weights; the Senet attention mechanism is used to perform a global averaging operation on the intermediate features and then normalize them to obtain normalized features; the normalized features are linearized using two learnable parameters and then activated by a one-dimensional convolution operation and the Sigmoid function to obtain weights. The weights are multiplied element-wise by the intermediate features, and the result is then subjected to convolution, batch normalization, and ReLU activation to obtain the output features of the skip weight channel (SWCA) module.

3. The hyperspectral image classification method based on multi-attention fusion according to claim 1, characterized in that, Branch D includes: a skip weight channel (SWCA) module and a second convolutional module; the second convolutional module includes convolutional layers of various scales; the skip weight channel (SWCA) module includes convolutional layers, max pooling layers, and the Senet attention mechanism.

4. The hyperspectral image classification method based on multi-attention fusion according to claim 1, characterized in that, Branch I includes: a basic convolutional module, two skip-weight channel (SWCA) modules, a CBAM attention mechanism, and a third convolutional module, which includes convolutional layers of various scales; the skip-weight channel (SWCA) modules include convolutional layers, max pooling layers, and the Senet attention mechanism. The hyperspectral image is input into the I branch to obtain shallow features, deep features, and I output features, including: The hyperspectral image is input into the basic convolutional module of the I branch to obtain convolutional features; The convolutional features are input into the first skip weight channel (SWCA) module of the I branch to obtain shallow features; The shallow features are input into the second skip weight channel SWCA module of the I branch to obtain the deep features; The deep features are processed by convolution and then input into the CBAM attention mechanism of the I branch to obtain attention features; The attention features are input into the third convolutional module of the I branch to obtain the I output features.

5. The hyperspectral image classification method based on multi-attention fusion according to claim 1, characterized in that, The multi-attention fusion module includes three different attention mechanisms; The I output features, P output features, and D output features are input into the multi-attention fusion module to obtain fused features, including: By superimposing the P and D output features with the I output feature as the main feature, the PID feature is obtained. By superimposing the I output feature with the P output feature and the D output feature respectively, we obtain the PI feature and the PD feature. The PI feature, the PID feature, and the PD feature are processed through different attention mechanisms to obtain the P-branch output feature, the I-branch output feature, and the D-branch output feature. The P-branch output features, I-branch output features, and D-branch output features are flattened and connected to obtain the fused features.

6. The hyperspectral image classification method based on multi-attention fusion according to claim 5, characterized in that, The three different attention mechanisms are: NAM attention mechanism, CBAM attention mechanism, and GAM attention mechanism; The P branch corresponds to the NAM attention mechanism; the I branch corresponds to the CBAM attention mechanism; and the D branch corresponds to the GAM attention mechanism.

7. The hyperspectral image classification method based on multi-attention fusion according to claim 5, characterized in that, The classification module includes a fully connected layer.

8. A hyperspectral image classification device based on multi-attention fusion, characterized in that, The device includes: A hyperspectral image classification model construction module is used to construct a hyperspectral image classification model based on multi-attention fusion. The hyperspectral image classification model includes a three-branch framework, a multi-attention fusion module, and a classification module. Inspired by PID controllers, the three-branch framework comprises a P branch, an I branch, and a D branch. The P, I, and D branches respectively use convolutional kernels and pooling layers of different sizes to extract features of different scales from the hyperspectral image. The multi-attention fusion module is used to superimpose the features extracted by the P, I, and D branches to obtain PI features, PID features, and ID features. Then, different attention mechanisms are used for feature extraction, and the extracted results are fused to obtain fused features. The feature extraction module is used to input the hyperspectral image into the I branch to obtain shallow features, deep features, and I output features; input the shallow features into the P branch to obtain P output features; input the deep features into the D branch to obtain D output features; and input the I output features, P output features, and D output features into the multi-attention fusion module to obtain fused features. A hyperspectral image classification module is used to input the fused features into the classification module to obtain the hyperspectral image classification result.

9. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the hyperspectral image classification method based on multi-attention fusion as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Multifunctional grabbing mechanical arm for rail inspection equipment and use method of multifunctional grabbing mechanical arm

    CN121893228A