Hyperspectral Image Classification Method Based on Feature Weighted Fusion of GAT and CNN
By combining the dual-branch feature weighted fusion method of GAT and CNN, Euclidian and non-Euclidian feature information of hyperspectral images is extracted and fused, and the problem of insufficient classification accuracy and generalization performance in the prior art is solved, thereby achieving higher classification accuracy and better generalization performance.
Patent Information
- Application Number
- CN202411537987.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Existing hyperspectral image classification methods are difficult to effectively utilize Euclidian and non-Euclidian feature information of images, resulting in insufficient classification accuracy and generalization performance.
The dual-branch feature weighted fusion method combined with graph attention network (GAT) and convolutional neural network (CNN) is adopted to extract superpixel-level graph structural features through GAT, and CNN extracts pixel-level spectral-space features and performs weighted fusion to achieve multi-scale fusion of features.
It improves the accuracy and generalization performance of hyperspectral image classification, and can better capture the non-Euclidean feature information and spatial structure information of the image.
Smart Images

Figure CN119399541B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hyperspectral image classification, and specifically relates to a hyperspectral image classification method based on feature weighted fusion of GAT and CNN. Background Art
[0002] Hyperspectral images have hundreds or even thousands of continuous bands in the visible and infrared bands, and can capture the characteristic spectral information of objects and surface materials in different bands. Compared with traditional color or single-band images that can only provide limited color or band information, hyperspectral images can provide richer and more detailed spectral features of ground objects, making it possible to perform more accurate analysis and identification of ground objects. Therefore, hyperspectral images have been widely used in fields such as agricultural crop monitoring, environmental pollution detection, urban planning, mineral exploration, weather forecasting, and military reconnaissance. With its rich spectral information and diverse application capabilities, hyperspectral images have become an indispensable and important part of modern remote sensing technology. The prerequisite for these applications is to accurately classify each pixel in the hyperspectral image (HSI).
[0003] In the past few decades, various machine learning-based classifiers have been developed for hyperspectral image classification. Early classification methods usually mapped high-dimensional hyperspectral information to low dimensions and processed the low-dimensional data. Therefore, how to establish a mapping function and find a separable hyperplane became the research goal. For example, logistic regression and extreme learning machines have been used to develop pixel-level HSI classifiers. However, pixelization methods usually produce quite large errors or outliers in the final classification map. To alleviate this problem, kernel tricks such as kernel support vector machines and multi-kernel learning are used to improve linear separability. But these methods usually focus on the design of the classifier and ignore the representation and learning of features. To make full use of spectral information, typical methods based on representation, such as sparse representation, low-rank representation, and collaborative representation, have been developed. Through representation learning, the inherent data structure of the spectrum can be revealed and the dependence on labeled samples can be reduced. In addition, the spatial structure of HSI has been explored, such as graph construction, superpixel segmentation, morphological segmentation, etc., to promote spectral-spatial feature learning. By explicitly modeling the spatial structure of HSI, its spatial information can be better utilized. However, limited by artificial features and empirical parameters, the above methods cannot learn robust deep feature representations from HSI.
[0004] Compared with traditional machine learning classification methods, deep learning methods can automatically learn adaptive and robust deep features from training data. Currently, many classic deep learning methods have been applied to HSI classification and achieved good results. Such as variants from one-dimensional convolutional neural network (CNN) to 3D CNN, and from single CNN to hybrid CNN. The advantage of deep learning methods in HSI image classification lies in their ability to better utilize big data and powerful computing capabilities, thereby improving classification accuracy, reducing the need for manual feature engineering, and being able to handle more complex and high-dimensional data features. Due to the high computational complexity, this hybrid CNN requires high computing power and a long training time. Previous deep learning models were designed for Euclidean data and often ignored the intrinsic correlation between adjacent land covers. In recent years, due to the ability to perform convolutional operations on graphs of arbitrary structures, graph neural networks (GNNs) have received increasing attention. By encoding HSI into a graph, the correlation between adjacent land covers can be explicitly utilized, and GNNs can better simulate the spatial context structure of HSI. GNNs can not only perform descriptive learning on non-Euclidean data but also perform end-to-end representation learning on both node feature information and structural information simultaneously. HSI data can be converted into graph data through a superpixel-based method, and then the GNN method can effectively model the spectral-spatial context information. In this way, the number of labels is implicitly expanded, and the small sample problem is alleviated to a certain extent. The superpixel-based GNN can simulate various spatial structures of land covers on the graph, but it cannot generate fine-grained individual features for each pixel. In contrast, CNNs can learn local spectral-spatial features at the pixel level, but their receptive field is usually limited to a small square window. Therefore, it may be difficult to capture the large-scale context structure of HSI. How to integrate the advantages of superpixel-level GNN and pixel-level CNN and enable data intercommunication has gradually become a key issue in the field of HSI classification. Summary of the Invention
[0005] The present invention aims to provide an innovative hyperspectral image classification method, which can combine the advantages of the Graph Attention Network (GAT) and the Convolutional Neural Network (CNN), and achieve feature fusion by extracting features of images at different scales. GAT is an advanced deep learning architecture suitable for processing graph-structured data. In the application of hyperspectral image classification, GAT can effectively learn the relative importance between nodes, thereby accurately capturing the non-Euclidean feature information in the graph. At the same time, this method utilizes the powerful pixel information capture ability of depthwise separable convolution and the processing ability of GAT for superpixel data to comprehensively extract the Euclidean and non-Euclidean feature information of hyperspectral images. Such feature fusion ensures high accuracy of the classification results and good generalization performance. The technical solution adopted by the present invention is a hyperspectral image classification method based on feature weighted fusion of GAT and CNN, which is characterized in that it is specifically implemented according to the following steps:
[0006] Step 1: Divide the hyperspectral dataset into a training set, a test set, and a validation set;
[0007] Step 2: Construct a dual-branch feature weighted fusion classification network model combining CNN and GAT;
[0008] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set parameters, and perform training to obtain training metrics;
[0009] Step 4: Evaluate the metrics, data, and image information of the classification results.
[0010] The present invention is further characterized in that,
[0011] Step 1 is specifically implemented according to the following steps:
[0012] Step 1.1: Use the classical hyperspectral dataset Indian Pines as the model input for experiments. Preprocess the dataset separately on the GAT and CNN dual branches, and divide it into a training set, a test set, and a validation set. The training set and the validation set are both randomly selected at 1% of the total data, and the remaining data is used as the test set.
[0013] Step 2 is specifically implemented according to the following steps:
[0014] Step 2.1. The dual-branch weighted fusion model combining depthwise separable convolution and GAT consists of two branches, and each branch has three parts, namely data preprocessing, multi-scale division, and weighted fusion parts. For the GAT branch, superpixel segmentation is used to preprocess the original pixel image to obtain superpixel information at different segmentation scales. The obtained information is respectively input into three GAT networks, and the result data is concatenated as the output of the GAT total branch. For the CNN branch, the spectral convolution layer is used to perform dimensionality reduction and denoising processing on the input original data. The processed data is downsampled and then input into the CNN branch, and the output result is upsampled and concatenated as the output of the CNN total branch;
[0015] Step 2.2. For the input data of the CNN branch, spectral convolution is used for preprocessing. First, the input image data is subjected to channel compression conversion, then the converted data is normalized, and then it passes through a spectral convolution layer with a convolution kernel size of 1×1. Finally, it passes through the LeakyReLU function to increase the generalization ability, and the processed data is used as the input of the CNN total branch. For the input data of the GAT branch, superpixel segmentation is used for preprocessing. The specific method is to perform superpixel segmentation of the original data with different sizes to obtain segmentation information at different scales, and then the processed data is used as the input of the GAT total branch;
[0016] Step 2.3: For the multi-scale division of different branches, perform max pooling with a pooling kernel size of 2 and a pooling stride of 2 on the input data obtained by the CNN branch through the spectral convolution layer twice to obtain three different scales of pixel-level information, and input the three types of data into three branches of the CNN network. Each CNN branch is divided into three layers, and each layer sequentially goes through three steps: the position attention module PAM, the channel attention module CAM, and the depthwise separable convolution layer. The position attention module PAM is used to capture the spatial position relationship in the image, and by weighting the features at each position, it strengthens the model's understanding and processing ability of spatial information. The position attention module PAM generates an attention map by calculating the similarity between different positions in the feature map, thereby achieving the emphasis on key positions. The channel attention module CAM adjusts the response intensity of each channel by analyzing the importance of different channels, so that the model can pay more attention to the feature channels that are more important for the final task. The depthwise separable convolution is carried out in two steps: first, use pointwise depth convolution to process each channel of the input respectively, and then use channel-wise pointwise convolution to combine the outputs of the depth convolution. The output channels of each part of the first two layers of the first branch are all 128, and the output channels of each part of the last layer are 128, 128, and 64 in sequence. The output channels of the latter two branches are halved successively with reference to the pooled data. Finally, the output results of the latter two branches are spliced with the output result of the first branch after passing through the DYS upsampling module. The DYS upsampling module selects specific sample points from the given sample set using point sampling and performs a spatial rearrangement operation on the original image, and finally organizes the spatial structure of the image data to generate a sampled image. The result after splicing is used as the output of the CNN total branch; for the GAT branch, the processed image data is segmented at superpixel scales of 100, 50, and 10 and input into three branches of the GAT respectively. Each GAT branch consists of two layers of GAT networks, and the input is the adjacency matrix A and the superpixel position matrix Q obtained through superpixel segmentation. The output channels are 128 and 64 respectively. Finally, the outputs of the three branches of the GAT are spliced as the output of the overall GAT total branch;
[0017] Step 2.4: For the fusion of superpixel data and pixel data, multiply the output data of each GAT branch by the corresponding scale superpixel position matrix Q obtained by its segmentation to convert it into output data of the same size and channels. Reshape the data of the CNN total branch into data with exactly the same size as the output of the GAT total branch for weighted fusion testing, and take the fusion ratio with the best effect as the final network fusion parameter.
[0018] Step 3 is specifically implemented according to the following steps:
[0019] Set the BatchSize to 256, the learning rate to 0.001, the Epoch to 300, and the optimizer to Adam. Randomly select 1% of the labeled pixels of each class in the Indian Pines dataset as the training set and input it into the network model obtained in the above step 2, and finally obtain a trained classification network model.
[0020] Step 4 is specifically implemented according to the following steps:
[0021] Randomly select 1% of the labeled pixels of each class in the Indian Pines dataset except the training set as the validation set, input it into the classification network model obtained by training with the training set above, and bring in the trained weight coefficients, and finally output the hyperspectral classification result map and various index data.
[0022] The beneficial effect of the present invention is that the hyperspectral image classification method based on feature weighted fusion of GAT and CNN first constructs a local feature extraction branch with GAT as the core, uses superpixel segmentation to convert traditional image data into graph structure data, and fully extracts the local features of the image; at the same time, a pixel-level feature extraction branch based on CNN is constructed, which can capture the pixel-level features of the image and effectively extract the overall information of the image. By combining the superpixel-level graph structure data with the pixel-level image data, a dual-branch weighted fusion network is constructed to realize feature extraction and information fusion from pixel points to graph structures, which not only retains the spatial structure information, but also retains the complex relationships and local structure information between graph nodes. Description of the Drawings
[0023] Figure 1 is the overall flowchart of the method of the present invention;
[0024] Figure 2 is the network framework diagram adopted by the present invention;
[0025] Figure 3 is the framework diagram of the DYS module in the network model of the present invention;
[0026] Figure 4 is the framework diagram of the PAM and CAM modules of the present invention;
[0027] Figure 5 is the comparison diagram of the classification results of this method and other methods. Detailed Embodiments
[0028] The present invention will be described in detail below in conjunction with the drawings and specific embodiments.
[0029] The hyperspectral image classification method based on feature weighted fusion of GAT and CNN of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:
[0030] Step 1: Divide the hyperspectral dataset into a training set, a test set, and a validation set;
[0031] Step 1 is specifically implemented according to the following steps:
[0032] Step 1.1: To comprehensively evaluate the performance of the model, use the classic hyperspectral dataset Indian Pines as the model input for experiments. Preprocess the dataset separately in the GAT and CNN dual branches, and divide it into a training set, a test set, and a validation set. The training set is used to train the model parameters in the subsequent steps, the test set is used to test the effectiveness of the model, and the validation set is used to evaluate the generalization ability of the model during training and optimize the model hyperparameters. Both the training set and the validation set randomly select 1% of the total data, and the remaining data is used as the test set.
[0033] Step 2: Construct a dual-branch feature weighted fusion network model combining CNN and GAT;
[0034] Step 2 is specifically implemented according to the following steps:
[0035] Step 2.1: The dual-branch weighted fusion model combining depthwise separable convolution and GAT consists of two branches in total. Each branch has 3 parts, namely data preprocessing, multi-scale division, and weighted fusion parts. The overall network framework is as Figure 2 shown. For the GAT branch, use superpixel segmentation to preprocess the original pixel image to obtain superpixel information at different segmentation scales, input the obtained information into 3 GAT networks respectively, and splice the result data as the output of the GAT total branch. For the CNN branch, use the spectral convolution layer to perform dimensionality reduction and denoising processing on the input original data, downsample the processed data and input it into the CNN branch, and splice the output result as the output of the CNN total branch;
[0036] Step 2.2: For the input data of the CNN branch, perform preprocessing using spectral convolution. First, perform channel compression conversion on the input image data, then perform normalization processing on the converted data, then pass it through a spectral convolution layer with a convolution kernel size of 1×1, and finally pass it through the LeakyReLU function to increase the generalization ability. Use the processed data as the input of the CNN total branch. For the input data of the GAT branch, use superpixel segmentation for preprocessing. The specific method is to perform superpixel segmentation of different sizes on the original data to obtain segmentation information at different scales, and then use the processed data as the input of the GAT total branch;
[0037] Step 2.3. For the multi-scale partitioning of different branches, the data obtained by the CNN branch through the spectral convolution layer is subjected to two maximum poolings with a pooling kernel size of 2 and a pooling stride of 2, resulting in pixel-level information of three different scales. The three types of data are input into three branches of the CNN network. Each CNN branch is divided into three layers, and each layer sequentially goes through three steps: the position attention module PAM, the channel attention module CAM, and the depthwise separable convolution layer. The position attention module PAM is an attention mechanism used to capture the spatial position relationships in the image. By weighting the features at each position, it enhances the model's understanding and processing ability of spatial information. The position attention module PAM generates an attention map by calculating the similarity between different positions in the feature map, thereby emphasizing the key positions. The channel attention module CAM is an attention mechanism that focuses on the channel dimension of the feature map. By analyzing the importance of different channels, it adjusts the response intensity of each channel, enabling the model to pay more attention to the feature channels that are more important for the final task. The depthwise separable convolution is an efficient convolution operation that is carried out in two steps: first, the depthwise convolution is used to process each input channel separately, and then the pointwise convolution is used to combine the outputs of the depthwise convolution. This design greatly reduces the number of model parameters and computational complexity while maintaining or improving the performance. The specific model structures of the PAM and CAM attention modules are respectively as Figure 4 shown. For the first branch, the output channels of each part in the first two layers are all 128, and the output channels of each part in the last layer are 128, 128, and 64 in sequence. For the last two branches, the output channels are halved successively according to the pooled data. Finally, the output results of the last two branches are upsampled through the DYS module and concatenated with the output result of the first branch as the output of the CNN main branch; for the GAT branch, the processed image data is segmented at superpixel scales of 100, 50, and 10 and input into three branches of the GAT respectively. Each GAT branch consists of two layers of GAT networks. The input is the adjacency matrix A and the superpixel position matrix Q obtained through superpixel segmentation, and the output channels are 128 and 64 respectively. Finally, the outputs of the three branches of the GAT are concatenated as the output of the overall GAT main branch;
[0038] Step 2.4. For the fusion of superpixel data and pixel data, the output data of each GAT branch is multiplied by its corresponding superpixel position matrix Q of the same scale to be converted into output data of the same size and channels. The data of the CNN main branch is reshaped into data with exactly the same size as the output of the GAT main branch for weighted fusion testing, and the fusion ratio with the best effect is used as the final network fusion parameter.
[0039] Step 3. Input the training set partitioned in Step 1 into the classification network model constructed in Step 2, set the parameters, and perform training to obtain training metrics;
[0040] Step 3 is specifically implemented according to the following steps:
[0041] Set the size of BatchSize to 256, the learning rate to 0.001, Epoch to 300, the optimizer to Adam, and randomly select 1% of the labeled pixels of each class in the Indian Pines dataset as the training set and input it into the network model obtained in the above Step 2, and finally obtain a trained classification network model.
[0042] Step 4: Evaluate the metrics, data, and image information of the classification results.
[0043] Step 4 is specifically implemented according to the following steps:
[0044] Randomly select 1% of the labeled pixels of each class in the Indian Pines dataset excluding the training set as the validation set, input it into the classification network model obtained by training with the training set above, and bring in the trained weight coefficients, and finally output the hyperspectral classification result map and various metric data.
[0045] Example 1
[0046] The hyperspectral image classification method based on feature weighted fusion of GAT and CNN of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:
[0047] Step 1: Divide the hyperspectral dataset into a training set, a test set, and a validation set;
[0048] Step 2: Construct a dual-branch feature weighted fusion network model combining CNN and GAT;
[0049] Step 3: Input the training set divided in Step 1 into the classification network model constructed in Step 2, set the parameters, and perform training to obtain training metrics;
[0050] Step 4: Evaluate the metrics, data, and image information of the classification results.
[0051] Example 2
[0052] The hyperspectral image classification method based on feature weighted fusion of GAT and CNN of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:
[0053] Step 1: Divide the hyperspectral dataset into a training set, a test set, and a validation set;
[0054] Step 1 is specifically implemented according to the following steps:
[0055] Step 1.1: To comprehensively evaluate the performance of the model, the classical hyperspectral dataset Indian Pines was used as the model input for experiments. The dataset was preprocessed separately in the GAT and CNN dual branches and divided into a training set, a test set, and a validation set. The training set was used to train the model parameters in the subsequent steps, the test set was used to test the effectiveness of the model, and the validation set was used to evaluate the generalization ability of the model during training and optimize the model hyperparameters. The training set and the validation set were both randomly selected at 1% of the total data, and the remaining data was used as the test set.
[0056] Step 2: Construct a dual-branch feature weighted fusion network model combining CNN and GAT;
[0057] Step 2 is specifically implemented according to the following steps:
[0058] Step 2.1: The dual-branch weighted fusion model combining depthwise separable convolution and GAT consists of two branches in total. Each branch has three parts, namely data preprocessing, multi-scale division, and weighted fusion parts. The overall network framework is as Figure 2 shown. For the GAT branch, superpixel segmentation was used to preprocess the original pixel image to obtain superpixel information at different segmentation scales. The obtained information was respectively input into 3 GAT networks, and the result data was concatenated as the output of the GAT total branch. For the CNN branch, a spectral convolution layer was used to perform dimensionality reduction and denoising processing on the input original data. The processed data was downsampled and then input into the CNN branch, and the output result was concatenated as the output of the CNN total branch;
[0059] Step 2.2: For the input data of the CNN branch, spectral convolution was used for preprocessing. First, the input image data was subjected to channel compression conversion, then the converted data was normalized, and then it was passed through a spectral convolution layer with a convolution kernel size of 1×1. Finally, it passed through the LeakyReLU function to increase the generalization ability. The processed data was used as the input of the CNN total branch. For the input data of the GAT branch, superpixel segmentation was used for preprocessing. The specific method was to perform superpixel segmentation of the original data with different sizes to obtain segmentation information at different scales, and then the processed data was used as the input of the GAT total branch;
[0060] Step 2.3. For the multi-scale partitioning of different branches, perform max pooling with a pooling kernel size of 2 and a pooling stride of 2 on the input data obtained by the CNN branch through the spectral convolution layer twice to obtain pixel-level information of three different scales, and input the three types of data into three branches of the CNN network. Each CNN branch is divided into three layers, and each layer sequentially goes through three steps: the position attention module PAM, the channel attention module CAM, and the depthwise separable convolution layer. The position attention module PAM is used to capture the spatial position relationship in the image, and by weighting the features at each position, it strengthens the model's understanding and processing ability of spatial information. The position attention module PAM generates an attention map by calculating the similarity between different positions in the feature map, thereby emphasizing the key positions. The channel attention module CAM adjusts the response intensity of each channel by analyzing the importance of different channels, enabling the model to pay more attention to the feature channels that are more important for the final task. The depthwise separable convolution is carried out in two steps: first, use pointwise depth convolution to process each input channel separately, and then use channel-wise pointwise convolution to combine the outputs of the depth convolution. The output channels of each part of the first two layers of the first branch are all 128, and the output channels of each part of the last layer are 128, 128, and 64 in sequence. The output channels of the last two branches are halved successively with reference to the pooled data. Finally, the output results of the last two branches are spliced with the output result of the first branch after passing through the DYS upsampling module. The DYS upsampling module selects specific sample points from the given sample set using point sampling, performs a spatial rearrangement operation on the original image, and finally organizes the spatial structure of the image data to generate the sampled image. The result after splicing is used as the output of the CNN main branch; for the GAT branch, the processed image data is segmented at superpixel scales of 100, 50, and 10 and input into three branches of the GAT respectively. Each GAT branch consists of two layers of GAT networks, and the input is the adjacency matrix A and the superpixel position matrix Q obtained through superpixel segmentation. The output channels are 128 and 64 respectively. Finally, the outputs of the three branches of the GAT are spliced as the output of the overall GAT main branch;
[0061] Step 2.4. For the fusion of superpixel data and pixel data, multiply the output data of each GAT branch by its corresponding scale superpixel position matrix Q to convert it into output data of the same size and channels. Reshape the data of the CNN main branch into data with exactly the same size as the output of the GAT main branch for weighted fusion testing, and take the fusion ratio with the best effect as the final network fusion parameter.
[0062] Step 3. Input the training set partitioned in Step 1 into the classification network model constructed in Step 2, set parameters, and perform training to obtain training metrics;
[0063] Step 3 is specifically implemented according to the following steps:
[0064] Set the BatchSize to 256, the learning rate to 0.001, the Epoch to 300, and the optimizer to Adam. Randomly select 1% of the labeled pixels of each class in the Indian Pines dataset as the training set and input it into the network model obtained in the above step 2, and finally obtain a trained classification network model.
[0065] Step 4: Evaluate the metrics, data, and image information of the classification results.
[0066] Step 4 is specifically implemented according to the following steps:
[0067] Randomly select 1% of the labeled pixels of each class in the Indian Pines dataset except the training set as the validation set, input it into the classification network model obtained by training with the training set above, and bring in the trained weight coefficients, and finally output the hyperspectral classification result map and various index data. Record the overall classification accuracy OA, the average classification accuracy AA, and the Kappa coefficient, and make a comparison with the current various classification methods as shown in Table 1 and conduct an analysis.
[0068] Example 3
[0069] The following further illustrates the effect of the present invention in combination with simulation experiments:
[0070] This experiment was carried out on a computer equipped with a Windows 10 (64-bit) system, an Intel(R) Core(TM) i7-12700F, an NVIDIA GeForce RTX 3060Ti GPU, and 16GB of RAM.
[0071] Execute step 1. The data uses the Indian_pines dataset. The data of different types in this dataset has relatively similar spectral curves, and the sample categories are extremely unbalanced. The image spatial resolution used is 20m×20m, covering 145×145 pixels. The wavelength range is 0.4 - 2.5μm, including 220 continuous bands. After removing 20 water absorption bands and noise bands, 200 bands are retained. Approximately half of the data (10,249 out of a total of 21,025 pixels) is labeled with 16 different categories;
[0072] Execute step 2 to construct a dual-branch feature weighted fusion model of CNN and GAT;
[0073] Execute step 3, set the training parameters for training. During the entire training process, set the BatchSize to 256, the learning rate to 0.001, the Epoch to 300, and the optimizer to Adam;
[0074] Execute Step 4, use the trained network model to classify the image data in the test set, and evaluate the indicators, data, and image information of the classification results.
[0075] To verify the effectiveness of the method proposed in the present invention, the fusion results of the method of the present invention are compared with three methods, namely: CEGCN, He3DCNN, and FDGC. The comparison chart of the classification results is shown as Figure 5 shown. Table 1 shows the objective evaluation index results compared with the other three methods, including the accuracy of each class, OA (overall accuracy), AA (average accuracy), and Kappa (kappa coefficient).
[0076] Table 1 Index values of the fusion results of four classification methods on the Indian_pines dataset
[0077] CEGCN He3DCNN FDGC Ours Class1 9.59 35.50 26.61 97.30 Class2 8138 6840 8602 86.85 Class3 41.37 50.00 62.00 70.67 Class4 14.63 60.30 45.22 96.12 Class5 60.33 87.00 88.89 97.91 Class6 98.90 90.10 99.82 98.06 Class7 29.34 80.80 77.14 100 Class8 88.61 95.40 100 100 Class9 21.28 53.80 47.37 81.82 Class10 67.21 69.10 80.85 80.29 Class11 88.17 73.30 81.99 87.86 Class12 33.59 51.90 68.97 75.71 Class13 99.59 91.90 83.19 99.48 Class14 99.85 94.40 100 99.91 Class15 50.18 51.00 55.63 73.89 Class16 25.03 70.20 67.92 95.79 OA(%) 56.82 50.30 82.46 88.19 AA(%) 75.12 73.40 81.06 90.10 Kappa×100 71.09 69.50 78 87.97
[0078] Through experimental verification, the proposed algorithm has obtained excellent performance indicators in the simulation environment.
Claims
1. A hyperspectral image classification method based on feature weighted fusion of GAT and CNN, characterized in that: Follow the steps below to implement it: Step 1: Divide the hyperspectral data set into training set, test set and validation set; Step 2: Construct a dual-branch feature weighted fusion network model combining CNN and GAT; The step 2 is specifically implemented according to the following steps: Step 2.1, the dual-branch weighted fusion model combining deep separable convolution and GAT contains two branches in total, each branch has three parts, namely data preprocessing, multi-scale division and weighted fusion. For the GAT branch, the original pixel image is preprocessed by superpixel segmentation to obtain superpixel information of different segmentation scales, and the obtained information is input into three GAT networks respectively and the result data is spliced as the output of the GAT main branch. For the CNN branch, the spectral convolution layer is used to perform dimensionality reduction and denoising on the input original data, and the processed data is downsampled and input into the CNN branch and the output results are spliced as the output of the CNN main branch. Step 2.2: For the input data of the CNN branch, spectral convolution is used for preprocessing. First, the input image data is subjected to channel compression conversion, and then the converted data is standardized, and then it is passed through a spectral convolution layer with a convolution kernel size of 1×1. Finally, the LeakyReLU function is used to increase the generalization ability, and the processed data is used as the input of the CNN main branch. For the input data of the GAT branch, superpixel segmentation is used for preprocessing. Specifically, the original data is segmented into superpixels of different sizes to obtain segmentation information of different scales, and then the processed data is used as the input of the GAT main branch. Step 2.3: For the multi-scale division of different branches, the input data obtained by the CNN branch through the spectral convolution layer is pooled twice with a kernel size of 2 and a pooling step of 2, and the pixel-level information of three different scales is obtained. The three kinds of data are input into the three branches of the CNN network. Each CNN branch is divided into three layers, and each layer passes through the position attention module PAM, the channel attention module CAM and the depth-separable convolution layer in three steps; Step 3: input the training set divided in step 1 into the classification network model constructed in step 2, set parameters, and perform training to obtain training indicators; Step 4: Evaluate the indicators, data, and image information of the classification results.
2. The hyperspectral image classification method based on feature weighted fusion of GAT and CNN according to claim 1 is characterized in that: The step 1 is specifically implemented according to the following steps: Step 1.1: Use the classic hyperspectral dataset Indian Pines as the model input for the experiment. Preprocess the dataset in GAT and CNN branches respectively and divide it into training set, test set and validation set. The training set and validation set are randomly selected from 1% of the total data, and the rest of the data is used as the test set.
3. The hyperspectral image classification method based on feature weighted fusion of GAT and CNN according to claim 2 is characterized in that: In the step 2, the position attention module PAM is used to capture the spatial position relationship in the image, and strengthen the model's understanding and processing ability of spatial information by weighting the features of each position. The position attention module PAM generates an attention map by calculating the similarity of different positions in the feature map, thereby emphasizing the key positions. The channel attention module CAM adjusts the response strength of each channel by analyzing the importance of different channels, so that the model can pay more attention to the feature channels that are more important to the final task. The depth separable convolution is divided into two steps: first, each channel of the input is processed separately using point-by-point depth convolution, and then the output of the depth convolution is merged using channel-by-channel point-by-channel convolution. The output channels of each part of the first two layers of the first branch are 128, and the output channels of each part of the last layer are 128, 128, and 64 respectively. The last two branches refer to the pooling number. The number of output channels of the data is halved in turn, and finally the output results of the last two branches are spliced with the output results of the first branch after passing through the DYS upsampling module. The DYS upsampling module uses point sampling to select specific sample points from a given sample set, and performs spatial rearrangement operations on the original image. Finally, the spatial structure of the image data is organized to generate a sampled image, and the spliced result is used as the output of the CNN main branch; for the GAT branch, the processed image data is segmented at superpixel scales of 100, 50, and 10 and input into the three branches of GAT respectively. Each GAT branch consists of two layers of GAT networks, and the adjacency matrix A and superpixel position matrix Q obtained by superpixel segmentation are input. The output channels are 128 and 64 respectively. Finally, the outputs of the three branches of GAT are spliced as the output of the overall GAT main branch; For the fusion of superpixel data and pixel data, the output data of each GAT branch is multiplied by the superpixel position matrix Q of its corresponding scale to convert it into output data of the same size and channel. The data of the CNN main branch is reshaped to be exactly the same size as the output of the GAT main branch for weighted fusion test, and the fusion ratio with the best effect is used as the final network fusion parameter.
4. The hyperspectral image classification method based on feature weighted fusion of GAT and CNN according to claim 3 is characterized in that: The step 3 is specifically implemented according to the following steps: The BatchSize is set to 256, the learning rate is set to 0.001, the Epoch is set to 300, the optimizer is Adam, and 1% of the labeled pixels of each type of data in the Indian Pines dataset are randomly selected as the training set and input into the network model obtained in step 2, and finally a trained classification network model is obtained.
5. The hyperspectral image classification method based on feature weighted fusion of GAT and CNN according to claim 4 is characterized in that: The step 4 is specifically implemented according to the following steps: 1% of the labeled pixels in each category of the Indian Pines dataset except the training set are randomly selected as the validation set, and are input into the classification network model obtained after training with the training set in the previous article, and the trained weight coefficients are brought in, and finally the hyperspectral classification result map and various indicator data are output.
Citation Information
Patent Citations
Small sample hyperspectral image classification method based on 3D deep convolutional neural network
CN115147742A
Hyperspectral ground feature classification method based on graph convolution and convolution fusion
CN116664954A