Hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion

By combining the dual-branch feature fusion network model of CNN and AGCN, adaptive graph convolution and multi-scale feature fusion are used to solve the problem of insufficient subtle feature extraction in hyperspectral image classification, achieving more efficient feature extraction and more accurate classification.

CN120355976AActive Publication Date: 2025-07-22XIAN UNIV OF TECH

Patent Information

Application Number
CN202510348116.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-22
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing hyperspectral image classification methods cannot fully extract subtle features, resulting in insufficient classification accuracy, and traditional methods are difficult to dig up complex spectral and spatial structure information, making it easy to cause 'dimensional disasters'.

Method used

A hyperspectral image classification method based on adaptive graph convolution and multi-scale features fusion, combined with a dual-branch feature fusion network model of CNN and AGCN, the spectral-space features of the hyperspectral image are extracted through multi-scale residual convolution blocks, spatial attention mechanisms and graph attention networks, and feature fusion is performed using multi-strategy fusion.

Benefits of technology

It improves the accuracy and efficiency of hyperspectral image classification, reduces the complexity of data processing, can better mine effective information in the data, and improves the performance and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355976A_ABST
    Figure CN120355976A_ABST
Patent Text Reader

Abstract

The invention aims to provide a hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion, and the method is specifically implemented according to the following steps: 1, collecting hyperspectral images to construct a hyperspectral data set, and dividing the constructed hyperspectral data set into a training set, a verification set and a test set by the hyperspectral data set; step 2, constructing a double-branch feature fusion classification network model combining a CNN and an AGCN; 3, inputting the training set divided in the step 1 into the classification network model constructed in the step 2, setting parameters, and training; and step 4, obtaining a hyperspectral image ground feature classification result, and evaluating indexes, data and image information of the classification result. The problem that in the prior art, a hyperspectral image classification method cannot fully extract fine features, and the classification precision is limited is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing, and specifically relates to a hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion. Background Technique

[0002] As an important product of remote sensing technology, hyperspectral images are obtained by remote sensing satellite observations and contain rich spectral and spatial information. In terms of spectrum, it can capture the reflection, absorption, and emission characteristics of ground objects in different bands, forming continuous and high-precision spectral curves, which can be used as a unique basis for ground object recognition. In terms of spatial information, it can accurately present the position, shape, size, and mutual relationship of ground objects, and construct a spatial distribution picture of the earth's surface. Hyperspectral images play a key role in many fields. In military reconnaissance, it can help detect hidden military targets and provide intelligence support; in environmental monitoring, it can monitor atmospheric, water quality, and soil pollution, identify pollutants, and understand the scope and trend of pollution; in vegetation surveys, it helps to understand vegetation growth, species distribution, and health status; in the field of geological exploration, it can assist in identifying rocks and minerals and inferring the distribution of mineral resources.

[0003] However, the application of hyperspectral images faces many challenges. Its data volume is huge, resulting from high-resolution imaging and numerous spectral bands, which requires high storage and transmission equipment, increasing the complexity and time cost of data processing. At the same time, it has the characteristics of high dimensionality, with a complex data space and prominent data redundancy problems. The correlation between adjacent bands leads to repeated information, wasting storage space and interfering with subsequent processing. Traditional classification methods, such as those based on statistical models, perform poorly in processing hyperspectral images. These methods are based on simple assumptions and models, and it is difficult to mine complex spectral and spatial structure information. Facing high dimensionality and data redundancy, they are prone to the "curse of dimensionality", resulting in large computational amounts and a decline in classification accuracy. In recent years, deep learning technology has achieved remarkable results in the field of image classification, which can automatically learn effective feature representations and improve classification accuracy. However, there are still deficiencies in hyperspectral image classification. Most methods are based on pixel classification, ignoring the spatial relationship between pixels and the overall structure information of the image, lacking an overall understanding of the ground objects in the image. Some methods use simple convolutional neural networks (CNNs). Due to the complex spectral characteristics of hyperspectral images, they cannot fully extract fine features, limiting the classification accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion, which solves the problem that the existing hyperspectral image classification methods cannot fully extract fine features and limit the classification accuracy.

[0005] The technical solution adopted by the present invention is a hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion, which is specifically implemented according to the following steps:

[0006] Step 1: Collect hyperspectral images to construct a hyperspectral dataset, and divide the constructed hyperspectral dataset into a training set, a validation set, and a test set;

[0007] Step 2: Construct a dual-branch feature fusion classification network model combining CNN and AGCN;

[0008] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and perform training;

[0009] Step 4: Obtain the land cover classification results of the hyperspectral images, and evaluate the indicators, data, and image information of the classification results.

[0010] The features of the present invention also lie in that,

[0011] Step 1 is specifically implemented according to the following steps:

[0012] Taking the hyperspectral dataset as the input, first perform preprocessing, that is, divide it into a training set, a test set, and a validation set. The training set is used for model training to enable the model to learn the mapping relationship between the spectral-spatial features of the hyperspectral images and the land cover classes; the validation set is used to adjust the hyperparameters during the model training process to prevent overfitting of the model. By evaluating the model performance on the validation set, the optimal model parameters are selected; the test set is used to objectively evaluate the generalization ability and classification performance of the model after the model training is completed, and obtain the final classification accuracy index of the model.

[0013] The dual-branch feature fusion classification network model combining CNN and AGCN in Step 2 includes two branches of CNN and AGCN and a feature fusion module. The CNN branch includes three parts, namely data preprocessing, multi-scale residual convolution block, and spatial attention mechanism. The AGCN branch also includes three parts, namely data preprocessing, adaptive graph convolution, and graph attention network.

[0014] In Step 2, the CNN branch first preprocesses the input data, that is, performs denoising and spectral feature extraction on a pixel-by-pixel basis while retaining the local spatial structure information of the image. In the first layer, the input data is first normalized, the number of channels is converted using a 1x1 convolution kernel, and then passed through the LeakyReLU activation function; in the second layer, the output of the first layer is first normalized, and a 3x3 convolution kernel is used for feature transformation, and then passed through the LeakyReLU activation function;

[0015] The feature map after denoising and spectral transformation is then passed through a multi-scale convolutional layer and residual connections. The specific convolutional steps are as follows: Convolution kernels of different sizes are used, including 7x7, 5x5, 3x3, and 1x1. The 7x7 convolutional kernel can capture local spatial features in a larger range and is suitable for extracting the macroscopic structure information of the image. The 5x5 and 3x3 convolutional kernels capture local features in medium and smaller ranges respectively and can extract the detailed information of the image. The 1x1 convolutional kernel is used to adjust the channel dimension and reduce the computational amount. Convolution operations are performed on the input feature map using different-sized convolutional kernels to generate feature maps of multiple scales. Each convolutional kernel generates one feature map, and the feature maps of different scales are concatenated together in the channel dimension to form a multi-scale feature map. Normalization is added after each convolutional layer to standardize the input data, reduce internal covariance, and accelerate the convergence of the model. Then, the LeakyReLU function is used to introduce non-linearity and enhance the expressive ability of the model. At the same time, in each convolutional block, the input feature is added to the output of the convolutional layer to form a residual connection. In the convolutional block, the input is x, and the output after passing through the convolutional layer and activation function is F(x). The residual connection directly passes the input to the output in the form of x + F(x).

[0016] The specific steps of the channel attention module CAM are as follows: Max-pooling and average-pooling are performed on the feature map obtained through multi-scale convolution to obtain two feature maps. Then, convolution operations are used to process the pooled feature maps to generate two attention maps. The two attention maps are merged and normalized through the Sigmoid function to generate the final spatial attention map.

[0017] In step 2, the AGCN branch is a sub-network for processing superpixel-level features, which extracts graph signal features and utilizes the relationships between superpixels. First, preprocessing is performed, that is, the input data is preprocessed using superpixel segmentation, which is to perform superpixel segmentation of the original data with different sizes. Then, the processed superpixel data is used as the input of the AGCN branch. The input data of the AGCN branch is the feature matrix S obtained during superpixel segmentation, with a shape of (N, in-channels), where N is the number of superpixels and in-channels is the feature dimension of each superpixel. The Euclidean distance between all superpixel features is calculated to obtain the distance matrix. The formula is:

[0018]

[0019] S i and S j are the i-th row and j-th row of the feature matrix S respectively, and S i,k and S j,k are the k-th elements of the i-th row and j-th row of the feature matrix S respectively;

[0020] The generated Gaussian similarity matrix is the dynamic adjacency matrix D, which is used to represent the similarity relationship between superpixels. The formula is as follows:

[0021]

[0022] where σ is a learnable parameter;

[0023] Then, fuse the static and dynamic adjacency matrices. The matrix A is a pre-defined static adjacency matrix, which represents the fixed connection relationship between superpixels. Use the fusion weights γ and (1 - γ) to fuse the static adjacency matrix A and the dynamic adjacency matrix to obtain the fused adjacency matrix. Use the node attention mechanism to weight the node features, enhance the important features and suppress the unimportant features, calculate the attention weights, and then map the input features to a new feature space through a linear layer. Use the fused adjacency matrix for graph convolution operations. The adaptive graph convolution module needs to go through three processes from input to graph convolution;

[0024] The fused graph attention mechanism GAT first performs a linear transformation on the input node feature matrix, aiming to map the node features to a new feature space. This step is achieved through a learnable weight matrix W. The initialization of the weight matrix follows the Xavier uniform distribution. The feature matrix Wh obtained after the linear transformation is used for subsequent attention calculations. Repeat and expand each node feature vector of the linearly transformed feature matrix Wh to form all possible node pair combinations, that is, for each node i, combine its feature vector Whi with the feature vectors Whj of all other nodes j to form a new feature matrix. Then, use a learnable attention vector a to calculate the attention scores between node pairs, and then introduce non-linearity through a LeakyReLU activation function. The formula is as follows:

[0025] e ij =LeakyReLU(a T [Wh i [Wh j )(3)

[0026] The calculated attention scores e ij need to be normalized to ensure the legality and stability of the weights. In GAT, the softmax function is used to normalize the attention scores so that their sum among all neighbors of node i is 1. In addition, only when the value in the adjacency matrix is greater than 0, node j is regarded as a valid neighbor of node i. The normalized attention weights a ij are used to weight the transformed features Whj. In this way, the new feature of each node is the weighted average of the features of its neighbor nodes, and the attention weights dynamically measure the importance of neighbor nodes, and perform weighted summation to obtain the new feature vector of node i;

[0027] Step 2 is specifically as follows:

[0028] The feature fusion module adopts three feature fusion strategies, namely weighted sum fusion, concatenation fusion, and attention mechanism fusion. Then, the outputs of the three feature fusion strategies are obtained, and the three outputs are input into the adaptive feature fusion to obtain the final output;

[0029] In weighted sum fusion, the output scores of CNN and AGCN are classified through linear layers respectively to obtain two classification probability score matrices. Two learnable weight parameters (γ1 and γ2) are used to perform weighted sum on these two score matrices to generate the final classification probability, where the weight parameters are normalized to the range of [0,1] through the sigmoid function;

[0030] In the concatenation fusion strategy, the features of CNN and AGCN are concatenated together in the channel dimension to form a high-dimensional feature vector, and then a fully connected layer is used to perform nonlinear transformation on the concatenated features to generate the final classification result;

[0031] In the attention mechanism fusion strategy, the attention mechanism is used to dynamically adjust the contribution ratio of the features of the CNN branch and the AGCN branch. Specifically, first, the features of the two branches are projected to the same feature dimension through linear layers respectively, and then the projected features are concatenated together and input into a multi-layer perceptron MLP. The MLP outputs an attention weight, whose value is in the range of [0,1]. Finally, the features of the two branches are weighted and summed according to the attention weight to obtain the final fusion result;

[0032] Finally, through the adaptive fusion strategy, according to the differences of the input features, the optimal combination ratio of the three fusion strategies is automatically learned. The specific method is to concatenate the outputs Y1, Y2, and Y3 of the three fusion strategies in the feature dimension, and then input them into an adaptive fusion module. The module outputs three adaptive weights, whose values are in the range of [0,1], and the sum of the three weights is 1. Finally, the outputs of the three fusion strategies are weighted and summed according to the adaptive weights to obtain the final fusion result Y.

[0033] Step 3 is specifically implemented according to the following steps:

[0034] Set the learning rate to 0.001, Epoch to 200, and the optimizer to Adam. Randomly select 5% of the labeled pixels of the Salinas dataset as the training set and input it into the network model obtained in Step 2, and finally obtain a trained classification network model.

[0035] Step 4 is specifically implemented according to the following steps:

[0036] Select two image classification evaluation metrics, namely the average accuracy AA and the Kappa coefficient. The value range of AA is between 0 and 1, and the value range of the Kappa value is between -1 and 1. The higher the value of the metric, the better the classification effect.

[0037] The beneficial effects of the present invention are as follows. The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion is more efficient in data processing compared with traditional methods. It reduces the complexity and time cost of data processing, can better mine the effective information in the data, and improves the classification effect. Traditional classification methods are based on simple assumptions and models, making it difficult to mine the complex spectral and spatial structure information of hyperspectral images and prone to the "curse of dimensionality". Most deep learning methods are based on pixel classification, ignoring the spatial relationship between pixels and the overall structure information of the image. Some simple CNN methods cannot fully extract fine features. The double-branch of the present invention works collaboratively to comprehensively extract the spectral and spatial features of hyperspectral images. At the same time, the multi-strategy fusion method fully utilizes the advantages of different fusion strategies. Compared with a single fusion method, it can more flexibly process different types of hyperspectral image data, improve the performance and generalization ability of the model, and achieve more accurate classification in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is the overall flowchart of the hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention;

[0039] Figure 2 is the network framework diagram of the hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention;

[0040] Figure 3 is the framework diagram of each convolutional block of CNN in the hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention;

[0041] Figure 4 is the framework diagram of the CAM module in the hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention;

[0042] Figure 5 is the comparison diagram of the present invention with other methods when selecting 5% of the training samples on the Salinas dataset in the hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0043] The present invention will be described in detail below in conjunction with the drawings and specific embodiments.

[0044] The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention has a flowchart as Figure 1As shown below, it is implemented specifically according to the following steps:

[0045] Step 1: Collect hyperspectral images to construct a hyperspectral dataset, which contains spectral information of multiple continuous and narrow bands from visible light to near-infrared. Divide the constructed hyperspectral dataset into a training set, a validation set, and a test set to ensure a reasonable distribution of each category in each subset.

[0046] Step 1 is specifically implemented according to the following steps:

[0047] Take the hyperspectral dataset as the input and first perform preprocessing, that is, divide it into a training set, a test set, and a validation set. The training set is used for model training to enable the model to learn the mapping relationship between the spectral-spatial features of hyperspectral images and the ground object categories; the validation set is used to adjust hyperparameters during model training to prevent overfitting. By evaluating the model performance on the validation set, the optimal model parameters are selected; the test set is used to objectively evaluate the generalization ability and classification performance of the model after the model training is completed, and the final classification accuracy index of the model is obtained.

[0048] Step 2: Combine Figure 2 , to construct a dual-branch feature fusion classification network model that combines CNN and AGCN;

[0049] The dual-branch feature fusion classification network model that combines CNN and AGCN in Step 2 includes two branches, CNN and AGCN, and a feature fusion module. The CNN branch consists of three parts, namely data preprocessing, multi-scale residual convolutional blocks, and spatial attention mechanisms. The AGCN branch also consists of three parts, namely data preprocessing, adaptive graph convolution, and graph attention networks.

[0050] Combine Figure 3 , Figure 4 , in Step 2, the CNN branch first preprocesses the input data, that is, performs denoising and spectral feature extraction at the pixel level while retaining the local spatial structure information of the image, which can enhance the expression ability of spectral features and suppress noise at the same time. In the first layer, the input data is first normalized, the number of channels is converted using a 1x1 convolutional kernel, and then passed through the LeakyReLU activation function; in the second layer, the output of the first layer is first normalized, and a 3x3 convolutional kernel is used for feature transformation, and then passed through the LeakyReLU activation function;

[0051] The feature map after denoising and spectral transformation is then passed through multi-scale convolutional layers and residual connections to enhance the feature representation ability. Specifically, in the convolution steps, convolutional kernels of different sizes are used, including 7x7, 5x5, 3x3, and 1x1 convolutional kernels. The 7x7 convolutional kernel can capture local spatial features in a larger range and is suitable for extracting the macroscopic structure information of the image. The 5x5 and 3x3 convolutional kernels capture medium and small-range local features respectively and can extract the detailed information of the image. The 1x1 convolutional kernel is used to adjust the channel dimension, reduce the computational amount, perform convolution operations on the input feature map, generate feature maps of multiple scales using different-sized convolutional kernels, with each convolutional kernel generating one feature map, and then concatenate the feature maps of different scales in the channel dimension to form a multi-scale feature map. Normalization is added after each convolutional layer to standardize the input data, reduce internal covariance, and accelerate the convergence of the model. Then, after using the LeakyReLU function to introduce non-linearity and enhance the model's representation ability, at the same time, in each convolutional block, the input feature is added to the output of the convolutional layer to form a residual connection. In the convolutional block, the input is x, and the output after passing through the convolutional layer and activation function is F(x). The residual connection directly passes the input to the output in the way of x + F(x).

[0052] The main function of the channel attention module CAM is to adaptively adjust the importance of each channel in the feature map, enabling the network to pay more attention to those channels that contain more useful information, thereby improving the performance of the model. It dynamically adjusts the weights of each position in the multi-scale feature map output by the multi-scale CNN convolutional module and generates a spatial attention map. The specific steps of the channel attention module CAM are to perform max-pooling and average-pooling on the feature map obtained by multi-scale convolution to obtain two feature maps, then use convolution operations to process the pooled feature maps to generate two attention maps, merge the two attention maps, and normalize them through the Sigmoid function to generate the final spatial attention map.

[0053] In step 2, the AGCN branch is a sub-network for processing superpixel-level features, which extracts graph signal features and utilizes the relationships between superpixels. First, preprocessing is performed, that is, the input data is preprocessed using superpixel segmentation, which is to perform superpixel segmentation of the original data with different sizes, and then the processed superpixel data is used as the input of the AGCN branch. The input data of the AGCN branch is the feature matrix S obtained during superpixel segmentation, with a shape of (N, in-channels), where N is the number of superpixels and in-channels is the feature dimension of each superpixel. Calculate the Euclidean distance between all superpixel features to obtain the distance matrix. The formula is:

[0054]

[0055] Si and S j are the i-th row and the j-th row of the feature matrix S respectively, and S i,k and S j,k are the k-th elements of the i-th row and the j-th row of the feature matrix S respectively;

[0056] The generated Gaussian similarity matrix is the dynamic adjacency matrix D, which is used to represent the similarity relationship between superpixels. The formula is:

[0057]

[0058] where σ is a learnable parameter;

[0059] Then fuse the static and dynamic adjacency matrices. The matrix A is a pre-defined static adjacency matrix, which represents the fixed connection relationship between superpixels. Use the fusion weights γ and (1 - γ) to fuse the static adjacency matrix A and the dynamic adjacency matrix to obtain the fused adjacency matrix. Use the node attention mechanism to weight the node features, enhance the important features and suppress the unimportant features, calculate the attention weights, and then map the input features to a new feature space through a linear layer. Use the fused adjacency matrix for graph convolution operations. The adaptive graph convolution module has to go through three processes from input to graph convolution;

[0060] The fused graph attention mechanism GAT highlights the influence of important neighbor nodes by introducing the attention mechanism and dynamically assigning different weights to the neighbors of each node. Specifically, the core of the attention mechanism is to calculate the importance scores between nodes, which is achieved through the following steps. The fused graph attention mechanism GAT first performs a linear transformation on the input node feature matrix, aiming to map the node features to a new feature space. This step is achieved through a learnable weight matrix W. The initialization of the weight matrix follows the Xavier uniform distribution. The feature matrix Wh obtained after the linear transformation is used for subsequent attention calculations. The feature vector of each node in the linearly transformed feature matrix Wh is repeatedly extended to form all possible combinations of node pairs. That is, for each node i, its feature vector Whi is combined with the feature vectors Whj of all other nodes j to form a new feature matrix. Then, a learnable attention vector a is used to calculate the attention scores between node pairs, and then a LeakyReLU activation function is introduced to introduce non-linearity. The formula is:

[0061] e ij = LeakyReLU(a T [Wh i [Wh j ) (3)

[0062] The calculated attention score e ijNormalization is required to ensure the legality and stability of the weights. In GAT, the softmax function is used to normalize the attention scores so that their sum over all neighbors of node i is 1. Additionally, node j is considered a valid neighbor of node i only when the value in the adjacency matrix is greater than 0. The normalized attention weight a ij is used to transform the feature Whj. In this way, the new feature of each node is the weighted average of the features of its neighbor nodes, and the attention weight dynamically measures the importance of the neighbor nodes. By performing weighted summation, the new feature vector of node i is obtained;

[0063] Step 2 is as follows:

[0064] The feature fusion module adopts three feature fusion strategies, namely weighted summation fusion, concatenation fusion, and attention mechanism fusion. Then, the outputs of the three feature fusion strategies are obtained, and the three outputs are input into the adaptive feature fusion to obtain the final output;

[0065] In weighted summation fusion, the output scores of CNN and AGCN are classified through linear layers respectively to obtain two classification probability score matrices. Two learnable weight parameters (γ1 and γ2) are used to perform weighted summation on these two score matrices to generate the final classification probability. The weight parameters are normalized to the interval [0,1] through the sigmoid function to ensure the stability of the fusion process. In this way, directly performing weighted fusion on the classification scores can balance the contributions of CNN and AGCN to a certain extent and improve the classification accuracy.

[0066] In the concatenation fusion strategy, the features of CNN and AGCN are concatenated together in the channel dimension to form a high-dimensional feature vector, and then a fully connected layer is used to perform a non-linear transformation on the concatenated features to generate the final classification result. Through this concatenation operation, the integrity of the original features can be retained, and different information extracted by CNN and AGCN can be fully utilized. At the same time, the fully connected layer can learn the complex relationships between the concatenated features and improve the expression ability of the model.

[0067] In the attention mechanism fusion strategy, the attention mechanism is used to dynamically adjust the contribution ratio of the features of the CNN branch and the AGCN branch. Specifically, first, the features of the two branches are projected to the same feature dimension through linear layers, and then the projected features are concatenated together and input into a multi-layer perceptron MLP. The MLP outputs an attention weight whose value is in the range [0,1]. Finally, the features of the two branches are weighted and summed according to the attention weight to obtain the final fusion result. This fusion method can adaptively adjust the importance of the features of the two branches according to different input features, thereby improving the fusion effect.

[0068] Finally, through the adaptive fusion strategy, according to the differences in input features, the optimal combination ratio of the three fusion strategies is automatically learned. The specific approach is to concatenate the outputs Y1, Y2, and Y3 of the three fusion strategies in the feature dimension and then input them into an adaptive fusion module. This module outputs three adaptive weights, whose values are in the range of [0, 1] and the sum of the three weights is 1. Finally, the weighted sum of the outputs of the three fusion strategies is calculated according to the adaptive weights to obtain the final fusion result Y. This fusion method can make full use of the advantages of the three fusion strategies, dynamically adjust their contribution ratios according to different input features, and thus improve the performance and generalization ability of the model.

[0069] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and conduct training.

[0070] Step 3 is specifically implemented according to the following steps:

[0071] Set the learning rate to 0.001, Epoch to 200, the optimizer to Adam, randomly select 5% of the labeled pixels from the Salinas dataset as the training set and input it into the network model obtained in Step 2, and finally obtain a trained classification network model.

[0072] Step 4: Obtain the classification results of the hyperspectral image, and evaluate the indicators, data, and image information of the classification results.

[0073] Step 4 is specifically implemented according to the following steps:

[0074] Select two image classification evaluation indicators, namely the average accuracy AA and the Kappa coefficient Kappa. The value range of AA is between 0 and 1, and the value range of the Kappa value is between -1 and 1. The higher the value of the indicator, the better the classification effect.

[0075] Example 1

[0076] The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention has a flow chart as Figure 1 shown, and is specifically implemented according to the following steps:

[0077] Step 1: Collect hyperspectral images to construct a hyperspectral dataset, which contains spectral information of multiple continuous and narrow bands from visible light to near-infrared. Divide the constructed hyperspectral dataset into a training set, a validation set, and a test set to ensure a reasonable distribution of each category in each subset.

[0078] Step 2: Construct a dual-branch feature fusion classification network model combining CNN and AGCN;

[0079] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and conduct training;

[0080] Step 4: Obtain the classification results of the hyperspectral image ground objects, and evaluate the indicators, data, and image information of the classification results.

[0081] Example 2

[0082] The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:

[0083] Step 1: Collect hyperspectral images to construct a hyperspectral dataset, which contains spectral information of multiple continuous and narrow bands from visible light to near-infrared. Divide the constructed hyperspectral dataset into a training set, a validation set, and a test set to ensure a reasonable distribution of each category in each subset.

[0084] Step 1 is specifically implemented according to the following steps:

[0085] Take the hyperspectral dataset as the input, first perform preprocessing, that is, divide it into a training set, a test set, and a validation set. The training set is used for model training to enable the model to learn the mapping relationship between the spectral-spatial features of the hyperspectral image and the ground object categories; the validation set is used to adjust hyperparameters during the model training process to prevent overfitting of the model. By evaluating the model performance on the validation set, the optimal model parameters are selected; the test set is used to objectively evaluate the generalization ability and classification performance of the model after the model training is completed, and obtain the final classification accuracy index of the model.

[0086] Step 2: Combine Figure 2 , and construct a dual-branch feature fusion classification network model combining CNN and AGCN;

[0087] The dual-branch feature fusion classification network model combining CNN and AGCN in Step 2 includes two branches of CNN and AGCN and a feature fusion module. The CNN branch includes three parts, namely data preprocessing, multi-scale residual convolution block, and spatial attention mechanism. The AGCN branch also includes three parts, namely data preprocessing, adaptive graph convolution, and graph attention network.

[0088] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and conduct training;

[0089] Step 4: Obtain the classification results of the hyperspectral image ground objects, and evaluate the indicators, data, and image information of the classification results.

[0090] Example 3

[0091] A hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion, the flow chart is as Figure 1 shown, and it is specifically implemented according to the following steps:

[0092] Step 1: Collect hyperspectral images to construct a hyperspectral dataset, which contains spectral information of multiple continuous and narrow bands from visible light to near infrared. Divide the constructed hyperspectral dataset into a training set, a validation set, and a test set to ensure a reasonable distribution of each category in each subset.

[0093] Step 1 is specifically implemented according to the following steps:

[0094] Take the hyperspectral dataset as input and first perform preprocessing, that is, divide it into a training set, a test set, and a validation set. The training set is used for model training to let the model learn the mapping relationship between the spectral-spatial features of hyperspectral images and the ground object categories; the validation set is used to adjust hyperparameters during model training to prevent overfitting. By evaluating the model performance on the validation set, the optimal model parameters are selected; the test set is used to objectively evaluate the generalization ability and classification performance of the model after the model training is completed, and the final classification accuracy index of the model is obtained.

[0095] Step 2: Combine Figure 2 , and construct a dual-branch feature fusion classification network model that combines CNN and AGCN;

[0096] The dual-branch feature fusion classification network model that combines CNN and AGCN in Step 2 includes two branches, CNN and AGCN, and a feature fusion module. The CNN branch includes three parts, namely data preprocessing, multi-scale residual convolution block, and spatial attention mechanism. The AGCN branch also includes three parts, namely data preprocessing, adaptive graph convolution, and graph attention network.

[0097] Combine Figure 3 , Figure 4 , in Step 2, the CNN branch first preprocesses the input data, that is, performs denoising and spectral feature extraction on a pixel-by-pixel basis while retaining the local spatial structure information of the image, which can enhance the expression ability of spectral features and suppress noise at the same time. In the first layer, the input data is first normalized, the number of channels is converted using a 1x1 convolution kernel, and then passed through the LeakyReLU activation function; in the second layer, the output of the first layer is first normalized, and a 3x3 convolution kernel is used for feature transformation, and then passed through the LeakyReLU activation function;

[0098] The feature map after denoising and spectral transformation is further processed through a multi-scale convolutional layer and residual connection to enhance the feature representation ability. The specific convolutional steps are as follows: Convolution kernels of different sizes are used, including 7x7, 5x5, 3x3, and 1x1. The 7x7 convolutional kernel can capture local spatial features in a larger range and is suitable for extracting the macroscopic structure information of the image. The 5x5 and 3x3 convolutional kernels capture medium and small-range local features respectively and can extract the detailed information of the image. The 1x1 convolutional kernel is used to adjust the channel dimension, reduce the computational amount, and perform a convolution operation on the input feature map. Feature maps of multiple scales are generated using different-sized convolutional kernels, with each convolutional kernel generating one feature map. The feature maps of different scales are concatenated together in the channel dimension to form a multi-scale feature map. Normalization is added after each convolutional layer to standardize the input data, reduce internal covariance, and accelerate the convergence of the model. Then, the LeakyReLU function is used to introduce non-linearity and enhance the model's representation ability. At the same time, in each convolutional block, the input feature is added to the output of the convolutional layer to form a residual connection. In the convolutional block, the input is x, and the output after the convolutional layer and activation function is F(x). The residual connection directly passes the input to the output in the form of x + F(x).

[0099] The main function of the channel attention module CAM is to adaptively adjust the importance of each channel in the feature map, enabling the network to pay more attention to those channels containing more useful information, thereby improving the performance of the model. It dynamically adjusts the weights of each position in the multi-scale feature map output by the multi-scale CNN convolutional module and generates a spatial attention map. The specific steps of the channel attention module CAM are as follows: Max-pooling and average-pooling are performed on the feature map obtained through multi-scale convolution to obtain two feature maps. Then, convolution operations are used to process the pooled feature maps to generate two attention maps. The two attention maps are merged and normalized through the Sigmoid function to generate the final spatial attention map.

[0100] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and perform training.

[0101] Step 4: Obtain the classification results of the hyperspectral image ground objects, and evaluate the indicators, data, and image information of the classification results.

[0102] Example 4

[0103] The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:

[0104] Step 1: Collect hyperspectral images to construct a hyperspectral dataset, which contains spectral information of multiple continuous and narrow bands from visible light to near-infrared. Divide the constructed hyperspectral dataset into a training set, a validation set, and a test set to ensure a reasonable distribution of each category in each subset.

[0105] Step 1 is specifically implemented according to the following steps:

[0106] Take the hyperspectral dataset as the input and first perform preprocessing, that is, divide it into a training set, a test set, and a validation set. The training set is used for model training to enable the model to learn the mapping relationship between the spectral-spatial features of hyperspectral images and the ground object categories; the validation set is used to adjust hyperparameters during model training to prevent overfitting. By evaluating the model performance on the validation set, the optimal model parameters are selected; the test set is used to objectively evaluate the generalization ability and classification performance of the model after model training to obtain the final classification accuracy index of the model.

[0107] Step 2: Combine Figure 2 to construct a dual-branch feature fusion classification network model that combines CNN and AGCN;

[0108] The dual-branch feature fusion classification network model that combines CNN and AGCN in Step 2 includes two branches, CNN and AGCN, and a feature fusion module. The CNN branch consists of three parts, namely data preprocessing, multi-scale residual convolution blocks, and spatial attention mechanisms. The AGCN branch also consists of three parts, namely data preprocessing, adaptive graph convolution, and graph attention networks.

[0109] In Step 2, the AGCN branch is a sub-network for processing superpixel-level features, which extracts graph signal features and utilizes the relationships between superpixels. First, perform preprocessing, that is, preprocess the input data using superpixel segmentation, which is to perform superpixel segmentation of the original data with different sizes, and then use the processed superpixel data as the input of the AGCN branch. The input data of the AGCN branch is the feature matrix S obtained during superpixel segmentation, with a shape of (N, in-channels), where N is the number of superpixels and in-channels is the feature dimension of each superpixel. Calculate the Euclidean distance between all superpixel features to obtain the distance matrix, and the formula is:

[0110]

[0111] S i and S j are the i-th and j-th rows of the feature matrix S respectively, and S i,k and S j,k are the k-th elements of the i-th and j-th rows of the feature matrix S respectively;

[0112] The generated Gaussian similarity matrix is the dynamic adjacency matrix D, which is used to represent the similarity relationship between superpixels. The formula is as follows:

[0113]

[0114] where σ is a learnable parameter;

[0115] Then, fuse the static and dynamic adjacency matrices. Matrix A is a pre-defined static adjacency matrix, which represents the fixed connection relationship between superpixels. Use the fusion weights γ and (1 - γ) to fuse the static adjacency matrix A and the dynamic adjacency matrix to obtain the fused adjacency matrix. Use the node attention mechanism to weight the node features, enhance important features and suppress unimportant features, calculate the attention weights, and then map the input features to a new feature space through a linear layer. Use the fused adjacency matrix for graph convolution operations. The adaptive graph convolution module needs to go through three processes from input to graph convolution;

[0116] The fused graph attention mechanism GAT highlights the influence of important neighbor nodes by introducing the attention mechanism and dynamically assigning different weights to the neighbors of each node. Specifically, the core of the attention mechanism is to calculate the importance scores between nodes, which is achieved through the following steps. The fused graph attention mechanism GAT first performs a linear transformation on the input node feature matrix, aiming to map the node features to a new feature space. This step is achieved through a learnable weight matrix W, and the initialization of the weight matrix follows the Xavier uniform distribution. The feature matrix Wh obtained after the linear transformation is used for subsequent attention calculations. Repeat and expand each node feature vector of the linearly transformed feature matrix Wh to form all possible node pair combinations. That is, for each node i, combine its feature vector Whi with the feature vectors Whj of all other nodes j to form a new feature matrix. Then, use a learnable attention vector a to calculate the attention scores between node pairs, and then introduce non-linearity through a LeakyReLU activation function. The formula is as follows:

[0117] e ij = LeakyReLU(a T [Wh i [Wh j ) (3)

[0118] The calculated attention scores e ij need to be normalized to ensure the legality and stability of the weights. In GAT, the softmax function is used to normalize the attention scores so that their sum is 1 among all the neighbors of node i. In addition, only when the value in the adjacency matrix is greater than 0, node j is regarded as a valid neighbor of node i. After normalization, the attention weights aij is used to perform weighted summation on the transformed feature \(W_{hj}\), so that the new feature of each node is the weighted average of the features of its neighbor nodes, and the attention weight dynamically measures the importance of the neighbor nodes, thereby obtaining the new feature vector of node \(i\);

[0119] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and perform training;

[0120] Step 4: Obtain the classification results of the hyperspectral image ground objects, and evaluate the indicators, data, and image information of the classification results.

[0121] Embodiment 5

[0122] The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:

[0123] Step 1: Collect hyperspectral images to construct a hyperspectral dataset, which contains spectral information of multiple continuous and narrow bands from visible light to near infrared. Divide the constructed hyperspectral dataset into a training set, a validation set, and a test set to ensure a reasonable distribution of each category in each subset.

[0124] Step 1 is specifically implemented according to the following steps:

[0125] Take the hyperspectral dataset as the input, first perform preprocessing, that is, divide it into a training set, a test set, and a validation set. The training set is used for model training to let the model learn the mapping relationship between the spectral-spatial features of the hyperspectral image and the ground object categories; the validation set is used to adjust the hyperparameters during the model training process to prevent overfitting of the model. By evaluating the model performance on the validation set, the optimal model parameters are selected; the test set is used to objectively evaluate the generalization ability and classification performance of the model after the model training is completed, and obtain the final classification accuracy index of the model.

[0126] Step 2: Combine Figure 2 , and construct a dual-branch feature fusion classification network model combining CNN and AGCN;

[0127] The dual-branch feature fusion classification network model combining CNN and AGCN in Step 2 includes two branches of CNN and AGCN and a feature fusion module. The CNN branch includes three parts, namely data preprocessing, multi-scale residual convolution block, and spatial attention mechanism. The AGCN branch also includes three parts, namely data preprocessing, adaptive graph convolution, and graph attention network.

[0128] Step 2 is specifically as follows:

[0129] The feature fusion module adopts three feature fusion strategies, namely weighted sum fusion, concatenation fusion, and attention mechanism fusion. Then, the outputs of the three feature fusion strategies are obtained, and the three outputs are input into the adaptive feature fusion to obtain the final output;

[0130] In weighted sum fusion, the output scores of CNN and AGCN are classified through linear layers respectively to obtain two classification probability score matrices. Two learnable weight parameters (γ1 and γ2) are used to perform weighted sum on these two score matrices to generate the final classification probability. The weight parameters are normalized to the range [0,1] through the sigmoid function to ensure the stability of the fusion process. Directly performing weighted fusion on the classification scores can balance the contributions of CNN and AGCN to a certain extent and improve the classification accuracy.

[0131] In the concatenation fusion strategy, the features of CNN and AGCN are concatenated together in the channel dimension to form a high-dimensional feature vector. Then, a fully connected layer is used to perform non-linear transformation on the concatenated features to generate the final classification result. Through this concatenation operation, the integrity of the original features can be retained, and different information extracted by CNN and AGCN can be fully utilized. At the same time, the fully connected layer can learn the complex relationships between the concatenated features and improve the expression ability of the model.

[0132] In the attention mechanism fusion strategy, the attention mechanism is used to dynamically adjust the contribution ratios of the features of the CNN branch and the AGCN branch. Specifically, first, the features of the two branches are projected to the same feature dimension through linear layers respectively, and then the projected features are concatenated together and input into a multi-layer perceptron MLP. The MLP outputs an attention weight whose value is in the range [0,1]. Finally, the features of the two branches are weighted and summed according to the attention weight to obtain the final fusion result. This fusion method can adaptively adjust the importance of the features of the two branches according to different input features, thereby improving the fusion effect.

[0133] Finally, through the adaptive fusion strategy, according to different input features, the optimal combination ratio of the three fusion strategies is automatically learned. The specific method is to concatenate the outputs Y1, Y2, and Y3 of the three fusion strategies in the feature dimension and then input them into an adaptive fusion module. The module outputs three adaptive weights whose values are in the range [0,1] and the sum of the three weights is 1. Finally, the outputs of the three fusion strategies are weighted and summed according to the adaptive weights to obtain the final fusion result Y. This fusion method can make full use of the advantages of the three fusion strategies and dynamically adjust their contribution ratios according to different input features, thereby improving the performance and generalization ability of the model.

[0134] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and conduct training.

[0135] Step 3 is specifically implemented according to the following steps:

[0136] Set the learning rate to 0.001, Epoch to 200, the optimizer to Adam, randomly select 5% of the labeled pixels from the Salinas dataset as the training set and input it into the network model obtained in Step 2, and finally obtain a trained classification network model.

[0137] Step 4: Obtain the classification results of the hyperspectral image ground objects, and evaluate the indicators, data, and image information of the classification results.

[0138] Step 4 is specifically implemented according to the following steps:

[0139] Select two image classification evaluation indicators, namely the average accuracy AA and the Kappa coefficient Kappa. The value range of AA is between 0 and 1, and the value range of the Kappa value is between -1 and 1. The higher the value of the indicator, the better the classification effect.

[0140] Example 6

[0141] This experiment was conducted on a computer equipped with a Windows 10 (64-bit) system, an Intel(R) Core(TM) i7-11700, an NVIDIA GeForce RTX 3060 GPU, and 16GB of RAM.

[0142] Execute Step 1. In the simulation experiment of the present invention, the Salinas dataset is used. The spectral coverage range in the dataset is from 400 to 2500 nanometers, including 224 bands. The data mainly consists of a large image covering various crops and unplanted land, and a total of 16 different ground object categories are divided, such as spinach, lettuce, corn, etc. The specific content of the data includes the original spectral image, ground object classification labels, and metadata. Each pixel in the original image has 224 spectral values;

[0143] Execute Step 2 to construct a dual-branch feature fusion classification network model combining CNN and AGCN;

[0144] Execute Step 3, set the training parameters for training, and during the entire training process, set the learning rate to 0.001, Epoch to 200, and the optimizer to Adam;

[0145] Execute Step 4, use the trained network model to classify the image data of the test set, and evaluate the indicators, data, and image information of the classification results.

[0146] To verify the effectiveness of the method proposed in the present invention, the fusion results of the method of the present invention are compared with three methods, namely: AMGCFN, CEGCN, and GTFN. Under the same conditions, the present invention is compared with four existing classification methods in the hyperspectral field in terms of the classification results on the Salinas hyperspectral dataset.

[0147] Table 1 Comparison results of the classification accuracies of four methods on the Salinas dataset

[0148]

[0149]

[0150] Table 1 shows the comparison of the classification accuracies of the present invention with three methods, namely AMGCFN, CEGCN, and WSSG, for different types of ground objects on the Salinas dataset, and at the same time gives two comprehensive evaluation indicators, namely the average accuracy (AA) and the Kappa coefficient. It can be seen that in multiple classes, such as Class3, Class5, Class11, etc., the classification accuracy of the present invention is higher than that of the other three methods, demonstrating the advantages of the present invention in the classification of these ground objects. The AA value of the present invention is 97.49%, which is significantly higher than 87.47% of AMGCFN, 95.81% of CEGCN, and 91.1% of WSSG, indicating that the present invention is superior in the average performance of the classification accuracies of different types of ground objects and can accurately classify different types of ground objects more evenly. As Figure 5 shown, the results of the present invention and three comparison methods (AMGCFN, CEGCN, WSSG) in hyperspectral image classification are presented, and at the same time, the true ground object image is given as a reference. Compared with the results of the present invention, there are differences in the color distribution and ground object division in some regions of AMGCFN. The boundary definition in some regions is not accurate enough, and there may be a situation where a certain ground object is misjudged as another ground object, resulting in a certain deviation between the classification result and the true ground object. Similar problems also exist in the classification results of CEGCN. The division of some ground object category regions is not precise enough. For example, the color filling in some regions is inconsistent with the true ground object, indicating that there are errors in the recognition of the corresponding ground objects, affecting the overall classification accuracy. Similarly, it can be seen from the image that there are obvious differences between the classification results of WSSG and the true ground objects. The color in some regions is chaotic, and the boundaries between different ground objects are blurred, reflecting the deficiencies in the extraction and classification of different ground object features in the hyperspectral image classification by this method, resulting in unsatisfactory classification results. From the color distribution and region division, it can be seen that the matching degree between the present invention and the true ground object image is relatively high. The boundaries of different ground object category regions are relatively clear, and the correspondence between the color distribution and the true ground object is more in line with expectations, indicating that the present invention can more accurately identify different ground object categories in hyperspectral image classification, and the classification results are visually closer to the actual situation.

[0151] Based on the comprehensive analysis of the above experimental results, the algorithm metrics and classification effect of the present invention exceed those of other comparative algorithms on the Salina dataset, and good classification results are obtained.

Claims

1. A hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion, characterized in that The implementation is specifically carried out according to the following steps: Step 1: Collect hyperspectral images to construct a hyperspectral dataset, and divide the constructed hyperspectral dataset into a training set, a validation set, and a test set; Step 2: Construct a dual-branch feature fusion classification network model combining CNN and AGCN; Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and conduct training; Step 4: Obtain the ground object classification results of the hyperspectral images, and evaluate the indicators, data, and image information of the classification results.

2. The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion according to claim 1, wherein The specific implementation of Step 1 is carried out according to the following steps: Taking the hyperspectral dataset as the input, first perform preprocessing, that is, divide it into a training set, a test set, and a validation set. The training set is used for model training to enable the model to learn the mapping relationship between the spectral-spatial features of the hyperspectral image and the ground object categories; the validation set is used to adjust the hyperparameters during the model training process to prevent the model from overfitting. By evaluating the model performance on the validation set, the optimal model parameters are selected; the test set is used to objectively evaluate the generalization ability and classification performance of the model after the model training is completed, and obtain the final classification accuracy index of the model.

3. The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion according to claim 2, wherein The dual-branch feature fusion classification network model combining CNN and AGCN in Step 2 includes two branches, CNN and AGCN, and a feature fusion module. The CNN branch includes three parts, namely data preprocessing, multi-scale residual convolution block, and spatial attention mechanism. The AGCN branch also includes three parts, namely data preprocessing, adaptive graph convolution, and graph attention network.

4. The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion according to claim 3, wherein In Step 2, the CNN branch first preprocesses the input data, that is, performs denoising and spectral feature extraction at the pixel level while retaining the local spatial structure information of the image. In the first layer, after normalizing the input data first, use a 1x1 convolution kernel to convert the number of channels, and then pass through the LeakyReLU activation function; in the second layer, first normalize the output of the first layer, use a 3x3 convolution kernel for feature transformation, and then pass through the LeakyReLU activation function; The feature map after denoising and spectral transformation is then passed through a multi-scale convolutional layer and residual connections. The specific convolutional steps are as follows: convolutional kernels of different sizes are used, including 7x7, 5x5, 3x3, and 1x1 convolutional kernels. The 7x7 convolutional kernel can capture local spatial features in a larger range and is suitable for extracting the macroscopic structure information of the image. The 5x5 and 3x3 convolutional kernels capture medium and small-range local features respectively and can extract the detailed information of the image. The 1x1 convolutional kernel is used to adjust the channel dimension and reduce the computational amount. Convolutional operations are performed on the input feature map using different-sized convolutional kernels to generate feature maps of multiple scales. Each convolutional kernel generates one feature map, and the feature maps of different scales are concatenated together in the channel dimension to form a multi-scale feature map. Normalization is added after each convolutional layer to standardize the input data, reduce internal covariance, and accelerate the convergence of the model. Then, the LeakyReLU function is used to introduce non-linearity and enhance the expressive power of the model. At the same time, in each convolutional block, the input feature is added to the output of the convolutional layer to form a residual connection. In the convolutional block, the input is x, and the output after passing through the convolutional layer and activation function is F(x). The residual connection directly passes the input to the output in the form of x + F(x). The specific steps of the channel attention module CAM are to perform max pooling and average pooling on the feature map obtained through multi-scale convolution to obtain two feature maps, then use convolutional operations to process the pooled feature maps to generate two attention maps, merge the two attention maps, and normalize them through the Sigmoid function to generate the final spatial attention map.

5. The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion according to claim 4, wherein In step 2, the AGCN branch is a sub-network for processing superpixel-level features, which extracts graph signal features and utilizes the relationships between superpixels. First, preprocessing is performed, that is, the input data is preprocessed using superpixel segmentation, which is to perform superpixel segmentation of the original data with different sizes, and then the processed superpixel data is used as the input of the AGCN branch; the input data of the AGCN branch is the feature matrix S obtained during superpixel segmentation, with a shape of (N, in-channels), where N is the number of superpixels and in-channels is the feature dimension of each superpixel. The Euclidean distance between all superpixel features is calculated to obtain the distance matrix, and the formula is: S i and S j are the i-th row and the j-th row of the feature matrix S, respectively, and S i,k and S j,k are the k-th elements of the i-th row and the j-th row of the feature matrix S, respectively; The generated Gaussian similarity matrix is the dynamic adjacency matrix D, which is used to represent the similarity relationship between superpixels, and the formula is: where σ is a learnable parameter; Then, fuse the static and dynamic adjacency matrices. Matrix A is a pre-defined static adjacency matrix representing the fixed connection relationships between superpixels. Use the fusion weights γ and (1 - γ) to fuse the static adjacency matrix A and the dynamic adjacency matrix to obtain the fused adjacency matrix. Use the node attention mechanism to weight the node features, enhance important features and suppress unimportant features, calculate the attention weights, and then map the input features to a new feature space through a linear layer. Use the fused adjacency matrix for graph convolution operations. The adaptive graph convolution module needs to go through three processes from input to graph convolution; The fused graph attention mechanism GAT first performs a linear transformation on the input node feature matrix, aiming to map the node features to a new feature space. This step is achieved through a learnable weight matrix W, and the initialization of the weight matrix follows the Xavier uniform distribution. The feature matrix Wh obtained after the linear transformation is used for subsequent attention calculations. Repeat and expand each node feature vector of the feature matrix Wh after the linear transformation to form all possible node pair combinations. That is, for each node i, combine its feature vector Wh i with the feature vectors Wh j of all other nodes j to form a new feature matrix. Then, use a learnable attention vector a to calculate the attention scores between node pairs, and then introduce non-linearity through a LeakyReLU activation function. The formula is: e ij = LeakyReLU(a T [Wh i [[Wh j ) (3) The calculated attention score e ij needs to be normalized to ensure the legality and stability of the weights. In GAT, the softmax function is used to normalize the attention scores so that they sum to 1 among all the neighbors of node i. Additionally, only when the value in the adjacency matrix is greater than 0 is node j considered a valid neighbor of node i. After normalization, the attention weight a ij is used to transform the feature Wh j. In this way, the new feature of each node is the weighted average of the features of its neighbor nodes, and the attention weights dynamically measure the importance of the neighbor nodes. By performing a weighted sum, the new feature vector of node i is obtained.

6. The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion according to claim 5, wherein The specific steps of step 2 are as follows: The feature fusion module adopts three feature fusion strategies, namely weighted sum fusion, concatenation fusion, and attention mechanism fusion. Then, obtain the outputs of the three feature fusion strategies, and input the three outputs into the adaptive feature fusion to obtain the final output; In weighted sum fusion, the output scores of CNN and AGCN are classified through linear layers respectively to obtain two classification probability score matrices. Use two learnable weight parameters (γ1 and γ2) to perform weighted sum on these two score matrices to generate the final classification probability, where the weight parameters are normalized to the range [0, 1] through the sigmoid function; In the concatenation fusion strategy, the features of CNN and AGCN are concatenated together in the channel dimension to form a high-dimensional feature vector, and then a fully connected layer is used to perform non-linear transformation on the concatenated features to generate the final classification result; In the attention mechanism fusion strategy, the attention mechanism is used to dynamically adjust the contribution ratios of the features of the CNN branch and the AGCN branch. Specifically, first, project the features of the two branches to the same feature dimension through linear layers respectively, and then concatenate the projected features and input them into a multi-layer perceptron MLP. The MLP outputs an attention weight, whose value is in the range [0, 1]. Finally, weight and sum the features of the two branches according to the attention weight to obtain the final fusion result; Finally, through an adaptive fusion strategy, according to the differences in input features, the optimal combination ratio of the three fusion strategies is automatically learned. The specific approach is to concatenate the outputs Y1, Y2, and Y3 of the three fusion strategies in the feature dimension and then input them into an adaptive fusion module. This module outputs three adaptive weights, whose values are in the range of [0, 1] and the sum of the three weights is 1. Finally, the outputs of the three fusion strategies are weighted and summed according to the adaptive weights to obtain the final fusion result Y.

7. The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion according to claim 6, wherein The specific implementation of step 3 is as follows: Set the learning rate to 0.001, Epoch to 200, and the optimizer to Adam. Randomly select 5% of the labeled pixels from the Salinas dataset as the training set and input it into the network model obtained in step 2 to finally obtain a trained classification network model.

8. The hyperspectral image classification method based on adaptive graph convolution and multi-scale feature fusion according to claim 7, wherein The specific implementation of step 4 is as follows: Select two image classification evaluation metrics, namely the average accuracy AA and the Kappa coefficient. The value range of AA is between 0 and 1, and the value range of the Kappa value is between -1 and 1. The higher the value of the metric, the better the classification effect.

Citation Information

Patent Citations

  • SAR image classification algorithm combining graph convolutional network and Markov random field

    CN113486967A

  • Hyperspectral remote sensing image classification method based on hybrid convolutional neural network

    CN115909052A

  • Image classification method of fusion network based on convolutional neural network and enhanced graph attention network

    CN116152561A

  • Deep forgery detection method and system based on graph convolution and multi-scale prompt fusion

    CN118941936A

  • Feature weighted fusion hyperspectral image classification method based on GAT and CNN

    CN119399541A

Cited By

  • Equipment defect detection method and equipment based on multiple modes

    CN120726042A