Hyperspectral object classification method based on graph convolution and convolution fusion

By constructing a hyperspectral geographic classification method that integrates graph convolution and convolution, the problem of difficulty in extracting spatial information and spectral information in the high-spectral image simultaneously in the prior art is solved, and higher classification accuracy and more stable performance are achieved.

CN116664954BActive Publication Date: 2025-05-16XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310791825.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-05-16
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

The existing hyperspectral remote sensing image classification technology is difficult to extract spatial information and spectral information in hyperspectral images at the same time, and is easily affected by noisy information in the spectral channel during feature extraction, resulting in low classification accuracy.

Method used

Using a hyperspectral land object classification method based on graph convolution and convolution fusion, the region-level feature extraction module and the pixel-level feature extraction module are constructed in parallel and connected in series with the feature fusion module to form a feature extraction network, which can simultaneously extract the dependence between regions in the hyperspectral image and the spatial and spectral features on a single pixel.

Benefits of technology

A more comprehensive feature extraction of hyperspectral remote sensing images is achieved, classification accuracy is improved, computing resources consumption is reduced, and feature extraction capabilities are improved under the conditions of small sample data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116664954B_ABST
    Figure CN116664954B_ABST
Patent Text Reader

Abstract

The present invention discloses a hyperspectral feature classification method based on graph convolution and convolution fusion, which mainly solves the problem that the existing technology relies on a large number of sample training for image classification and cannot fully capture the long-distance dependency of the image. Its implementation scheme is: obtain a hyperspectral data set and a data annotation set from a public website, and divide the training set and test set samples; construct a feature extraction network composed of a regional feature extraction module, a feature fusion conversion module, and a pixel-level feature extraction module in series; use the training set to calculate the loss of the entire network through the cross entropy loss function, and use the stochastic gradient descent method to iteratively optimize the network parameters to obtain a trained extraction network; input the test set into the trained extraction network to obtain the hyperspectral image classification result. The present invention can fully explore the deep-level information of remote sensing image features, capture the dependency between image regions, and improve the ability to extract image features, and can be used for urban development, environmental monitoring and resource exploration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to a classification method based on hyperspectral objects, which can be used for urban development, resource exploration and environmental monitoring. Background Art

[0002] Hyperspectral remote sensing image object classification plays a vital role in the field of remote sensing and has attracted increasing attention due to its wide range of practical applications.

[0003] Existing hyperspectral classification technologies are mainly divided into the following two categories:

[0004] The first category is the hyperspectral classification method based on machine learning, that is, traditional machine learning algorithms such as support vector machine (SVM) and random forest (RF) have been widely used to accurately classify objects in remote sensing images. The features they extract mainly include shallow features, such as texture features, shape features, and color features. Since the scenes involved in hyperspectral remote sensing images are complex and diverse, and accompanied by the problems of "same object, different spectrum" and "same spectrum, different objects", this type of method cannot well represent the complex content in hyperspectral remote sensing images.

[0005] The second category is the hyperspectral remote sensing image classification method based on deep learning, the most representative of which is based on convolutional neural network and graph convolutional neural network method. This type of method can extract high-level semantic features in remote sensing images, and the classification effect is relatively good. It is one of the currently widely used hyperspectral remote sensing image classification methods. Although the method based on deep convolutional neural network has achieved great success in the task of hyperspectral remote sensing image classification, it is good at capturing global information from remote sensing images, and therefore cannot thoroughly mine the local knowledge hidden in complex remote sensing images and the band information from the spectrum; at the same time, because it uses specific hyperspectral remote sensing data division, it will be disturbed by the number of data division blocks. The number of blocks for remote sensing data division is usually very small, which leads to the misuse of a lot of noisy information in feature extraction, affecting the ability to learn image features.

[0006] In view of the problems and defects in the existing hyperspectral remote sensing image classification technology, how to provide a method that can simultaneously extract spatial information and spectral information in the hyperspectral image and effectively filter out noisy information on the spectral channel during feature extraction is a difficult problem that technicians in this field need to solve.

[0007] In an article published in the IEEE Journal of the Institute of Electrical and Electronics Engineers, Mou Lichao et al. proposed a non-local graph convolutional network that improves classification accuracy by calculating the relationship between pixels in the entire image. However, this method consumes a lot of memory resources and takes a lot of time in the training phase.

[0008] Wan Sheng et al. proposed a multi-scale graph convolution method in an article published in the IEEE Journal of the Institute of Electrical and Electronics Engineers. This method first uses a simple linear iterative clustering algorithm to divide the original image into several superpixels, and uses the Euclidean distance formula to measure the correlation between pixel blocks, thereby constructing the adjacency matrix in the input graph neural network and the mapping function between pixels in the pixel block area. At the same time, multi-scale technology is used to effectively supplement the ability to extract spatial information. However, in order to prevent the loss of key information when constructing node relationships, small sample data tends to construct a relationship graph with more nodes due to its smaller space, while a relationship graph constructed with fewer nodes better represents large sample data. Therefore, the superpixels constructed by this method using a simple linear iterative clustering algorithm are easily affected by the ability to extract features under the condition of unbalanced samples.

[0009] The article published by Xu Kejie et al. in the IEEE Journal proposed a hyperspectral image classification based on graph attention convolution designed specifically for hyperspectral characteristics. It first uses the spectral residual module to extract spectral discriminant features; then uses the graph attention network to characterize the relationship between the target node and the neighboring nodes; then uses the fully connected layer to calculate multiple groups of attention coefficients, and through feature concatenation and weighted average operations, it multi-dimensionally mines the importance of different neighbor nodes to the central node, adaptively highlights important neighbor nodes, and suppresses noise nodes to obtain expressive features. However, since graph convolution is still limited to shallow feature extraction, the classification effect needs to be further improved. Summary of the invention

[0010] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a hyperspectral object classification method based on graph convolution and convolution fusion to enhance the feature extraction capability of samples, reduce computing resources and improve classification accuracy.

[0011] To achieve the above object, the technical solution of the present invention comprises the following steps:

[0012] (1) Obtain the original data set and data annotation set of hyperspectral images from the public website, select a certain proportion of non-zero labels from the data annotation set to form a training set, and use all the remaining samples to form a test sample set;

[0013] (2) Preprocessing the original data set of hyperspectral images:

[0014] 2a) Using a simple linear iterative algorithm to segment the hyperspectral image;

[0015] 2b) calculating the region matrix V, adjacency matrix A, and region and pixel mapping matrix Q of the segmented image;

[0016] 3) Build a feature extraction network:

[0017] 3a) Establish a region-level feature extraction module consisting of three graph convolution layers, one graph pooling layer, and one graph de-pooling layer to extract the dependencies between regions;

[0018] 3b) Build a pixel-level feature extraction module consisting of a multi-scale feature extraction layer and an existing attention layer to extract spectral information and spatial features on a single pixel;

[0019] 3c) establishing a fusion module for fusing the features of 3a) and 3b);

[0020] 3d) connecting the region-level feature extraction module and the pixel-level feature module in parallel, and then connecting them in series with the fusion module to form a feature extraction network, and using the cross entropy loss function as the loss function of the feature extraction network;

[0021] (4) Train the feature extraction network:

[0022] 4a) Input the original image and the region matrix V, the adjacency matrix M, and the region and pixel mapping matrix Q into the feature extraction network, and calculate the loss value between the feature extraction network output and the true label of the training set;

[0023] 4b) Using the stochastic gradient descent method, gradually reduce the value of the loss function to update the network parameters until the set maximum number of iterations is completed to obtain a trained feature extraction network;

[0024] (5) Using the trained feature extraction network, the classification results of the test set hyperspectral images are obtained:

[0025] 5a) Input the original image and the region matrix V, the adjacency matrix M, and the region-pixel mapping matrix Q into the trained feature extraction network to obtain the output vector F;

[0026] 5b) Using the argmax function on the channel dimension of the output vector F, calculate the position index of the maximum value on the channel dimension, and the position index is the category of all pixels in the entire image;

[0027] 5c) From the categories of all pixels in the entire image, find the values ​​of the corresponding coordinates of the test set, which is the classification result of the test set.

[0028] Compared with the prior art, the present invention has the following advantages:

[0029] 1) The present invention constructs a feature extraction network based on a regional feature extraction module and a pixel-level feature extraction module in parallel, which are then connected in series with a feature fusion module. Therefore, the dependency between regions in the hyperspectral image and the spatial and spectral features of a single pixel can be extracted simultaneously, and hyperspectral classification can be performed in an end-to-end manner. The complex content in the hyperspectral remote sensing image can be fully interpreted, and the classification performance can be greatly improved.

[0030] 2) In the feature extraction network, the present invention establishes a pixel-level feature extraction module composed of a multi-scale feature extraction layer and an attention layer. By embedding position information into the channel attention mechanism, the spectral features on each pixel and the local spatial features within the pixel-level area are fully extracted, thereby further enhancing the small sample feature extraction capability.

[0031] 3) In the feature extraction network of the present invention, a regional feature extraction module is established to model the relationship between regions in the image by gradually fusing shallow features and deep features, so that the network can learn more deep features.

[0032] Simulation experiments show that the classification accuracy of the present invention is better than other existing methods, and the overall classification effect is more stable. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of the implementation process of the present invention;

[0034] Figure 2 It is a schematic diagram of segmenting a hyperspectral image using a simple linear iterative algorithm in the present invention;

[0035] Figure 3 It is a schematic diagram of the structure of the pixel-level feature extraction module in the present invention; DETAILED DESCRIPTION

[0036] The examples and effects of the present invention are further described in detail below with reference to the accompanying drawings.

[0037] Reference Figure 1 ,The hyperspectral object classification method based on graph convolution and convolution fusion in this example includes five modules: selection of training set and test set, image preprocessing module, regional feature extraction module, pixel-level feature extraction module, feature conversion and fusion module. The specific implementation steps are as follows:

[0038] Step 1: Select training set and test set.

[0039] Obtain the original dataset and data annotation set of hyperspectral images from public websites;

[0040] The samples with label 0 are removed from the data annotation set to form the foreground data set, some samples are selected from the foreground data set as the training set, and all the remaining samples constitute the test sample set.

[0041] Step 2: Preprocess the original data set of hyperspectral images.

[0042] 2.1) Use a simple linear iterative algorithm to segment the hyperspectral image and obtain the segmented hyperspectral image P'

[0043] p'=Slic(p,Z)

[0044] Where P∈H×W×B is the original hyperspectral image, Z=(H×W) / a is the number of divided blocks, H, W, B represent the length, width, and height of the image respectively, and a is a hyperparameter used to control the size of superpixels. The segmented image is as follows Figure 2 As shown;

[0045] 2.2) For the segmented image P', calculate its region matrix V:

[0046]

[0047] Among them, S i represents the i-th superpixel block, Indicates that there is i The jth pixel in N z Represents the superpixel block S Z The number of pixels contained in ;

[0048] 2.3) For the segmented image P', calculate the region adjacency matrix A:

[0049]

[0050] in Represents the element in the i-th row and j-th column of the matrix A, i.e., the i-th superpixel block V i With the jth superpixel block V j , N(·) represents the neighborhood of the superpixel, and y represents the set hyperparameter.

[0051] 2.4) For the segmented image P', calculate the region and pixel mapping relationship matrix Q:

[0052]

[0053] Where Q∈(Z,H×W), Z is the number of superpixel blocks, H and W are the image length and width respectively; Q i,j Represents the element in the i-th row and j-th column of the matrix Q, i.e., the superpixel block S j With Pixel X iThe mapping relationship is determined as follows:

[0054]

[0055] Step 3: Build a feature extraction network.

[0056] 3.1 Constructing regional feature extraction module:

[0057] 3.1.1) Establish a regional feature extraction network consisting of three graph convolutional layers, one graph pooling layer, and one graph de-pooling layer;

[0058] The first graph convolution layer and the second graph convolution layer both include a graph convolution with a dimension parameter of 128;

[0059] The third graph convolution layer includes a graph convolution with a dimension parameter of M, where M represents the number of image categories;

[0060] The first graph pooling layer includes a pooling operation with a hyperparameter K. In this example, K is set but not limited to 0.5. The first graph de-pooling layer includes a de-pooling operation.

[0061] 3.1.2) Establish the connection relationship between each layer and obtain the output feature F of the regional feature extraction module:

[0062] The input features pass through the first graph convolution layer in sequence to obtain the output first feature F1; F1 passes through the first graph pooling layer, the second graph convolution layer, and the first graph de-pooling layer in sequence to obtain the second feature F2; after the output feature F1 of the first graph convolution layer is added to the output feature F2 of the first graph de-pooling layer, it passes through the third graph convolution layer to output the feature F of the final regional feature extraction module.

[0063] 3.2) Establish pixel-level feature extraction module:

[0064] 3.2.1) Establish a multi-scale feature extraction layer consisting of a point convolution layer connected in parallel with three channel-by-channel convolution layers and then added together to extract spectral and spatial features in pixels;

[0065] The convolution kernel size of this point convolution layer is 1×1, and the channel dimension parameter is set to 128;

[0066] The convolution kernel size of the first channel-by-channel convolution layer is 3×3, and the padding layer parameter is set to 1;

[0067] The convolution kernel size of the second channel-by-channel convolution layer is 5×5, and the padding layer parameter is set to 2;

[0068] The convolution kernel size of the third channel-by-channel convolution layer is 7×7, and the padding layer parameter is set to 3;

[0069] The features output by the three channel-by-channel convolutional layers are F11, F21, and F31, and the output feature after addition is F41, that is, F41 = F11 + F21 + F31;

[0070] 3.2.2) Select the existing attention layer to embed the position information into feature learning and increase the feature perception ability of spatial position information;

[0071] 3.2.3) Reference Figure 3 , three channel-by-channel convolutional layers are connected in series with the existing attention layer to form the pixel-level feature extraction module; the output feature CNNres of the pixel-level feature extraction module is expressed as follows

[0072] CNNres=attention(F41)

[0073] Among them, attention() represents the existing attention layer.

[0074] 3.3) Feature fusion and conversion module:

[0075] 3.3.1) Convert the output feature F of 3.1.2) to obtain the graph conversion feature GCNres:

[0076] GCNres=Q×F

[0077] Among them, Q represents the region and pixel mapping relationship matrix;

[0078] 3.3.2) The two output features obtained in step 3.3.1) and step 3.2.3) are fused, and the obtained output feature F5 is expressed as follows:

[0079] F5=Softmax(Concat(GCNres,CNNres))

[0080] Among them, Concat represents the concatenation operation;

[0081] 3.4) The region-level feature extraction module is connected in series with the feature fusion conversion module and the pixel-level feature extraction module to form a feature extraction network, which is used to perform feature conversion on the original image X and the region matrix V, the adjacency relationship matrix M, and the region and pixel mapping relationship matrix Q input to the region-level feature extraction module through the feature fusion module, and its output features are then fused with the output features of the pixel-level feature extraction module to output the network features.

[0082] Step 4: Train the feature extraction network.

[0083] 4.1) Input the original image and the region matrix V, the adjacency matrix M, and the region and pixel mapping matrix Q into the feature extraction network, and calculate the cross entropy loss function Loss between the output of the feature extraction network and the training set label:

[0084] Loss=-[y1logy'+(1-y1)log(1-y')]

[0085] Among them, θ is the initial parameter of the network obtained by random initialization, y1 is the true label, and y' is the output of the network.

[0086] 4.2) Update the network model parameters θ using the gradient descent method:

[0087] Assume that the initial training batch m=0, and the maximum number of iterations in the pre-training phase T=200;

[0088] Calculate the network model parameters θ after the current training update m+1 :

[0089]

[0090] Among them, α is the learning rate during the training phase, θ m is the network parameter before the current training update, and L1(θ) is the current network parameter;

[0091] 4.3) Repeat step 4.2) until the maximum number of iterations T in the pre-training phase is reached, stop training, and obtain a trained feature extraction network.

[0092] Step 5: Get the classification results of the test set.

[0093] 5.1) Input the original image, region matrix V, adjacency matrix M, and region-pixel mapping matrix Q into the feature extraction network to output feature vector F6;

[0094] 5.2) Use the argsmax function to calculate the index of the maximum value in the channel dimension of vector F6: out =argmax(F6) where F out Represents the maximum position index of the image, F out ∈H×W×1, F6∈H×W×k is the output feature vector, H, W are the image length and width respectively, and k is the total number of categories.

[0095] 5.3) Find the value of the corresponding coordinate of the test set from all the pixels in the entire image, which is the classification result of the test set.

[0096] The effect of the present invention can be further illustrated by the following simulation:

[0097] 1. Simulation conditions

[0098] The simulation environment of the present invention selected the framework of python 3.8+pytorch 1.7 and was completed on a workstation with GeForce RTX2080Ti and 11G memory.

[0099] The four datasets used in the simulation are Indian Pines dataset, PaviaU dataset, Houston2013 dataset and Salinas dataset.

[0100] The Indian Pines dataset was collected by the airborne visible / infrared imaging spectrometer AVIRIS sensor in a farmland area in northwestern Indiana. The original hyperspectral image wavelength range is between 0.4-2.5 microns. The spatial resolution is 20m. The selected area size contains 145×145 pixels. After noise removal and atmospheric correction, 200 spectral bands are selected for experiments. The dataset contains a variety of objects and materials in real scenes such as coniferous forests and farmlands, such as corn, soybeans, and pine trees. There are 16 different ground object categories, including a total of 10,249 manually labeled data samples, and the remaining 10,776 pixels are background pixels.

[0101] The paivaU dataset is a hyperspectral remote sensing image dataset obtained in the area near the University of Pavia using the reflective optical system imaging spectrometer ROSIS-3HS sensor. The spectral coverage range is between 430-860nm, and the geometric resolution of the pixel is 1.3m. After removing the bands affected by noise, 103 spectral bands were selected for the experiment. The area size contains 610×340 pixels, of which there are a total of 42,776 labeled pixels, which are divided into 9 land categories. Mainly including asphalt, grass, gravel and other objects.

[0102] The Houston2013 dataset was taken in 2012 at the University of Houston and its vicinity using the airborne sensor ITRES-CASI1500. After calibration, 144 spectral bands were selected for the experiment. Its spatial resolution is 2.5m. The dataset was published in the IEEE Geoscience and Remote Sensing Society Data Fusion Competition in 2013. The area is very complex. For the consistency of the experiment, the training set and the test set of the provided dataset were fused, and the obtained dataset contained 349×1905 pixels. There are 15029 labeled pixels in total. It contains 15 categories in total.

[0103] The Salinas dataset was taken by the AVIRIS sensor in the Salinas Valley of California, with a spatial resolution of 3.7 meters and 224 continuous bands. After removing 20 water-absorbing bands (108-112, 154-167, 224), the actual band used for training is 204. The area contains 512×340 pixels, mainly including 16 ground object categories such as corn, soybeans, cucumbers, tomatoes, etc.

[0104] 2. Simulation content

[0105] Simulation 1. Under the above simulation conditions, the present invention and seven existing hyperspectral remote sensing image classification methods CEGCN, MDGCN, S2GAT, 3DOCT-SSAN, A2S2K-ResNet, SVM, and RF-200 are used to classify the Indian Pines dataset, and their overall accuracy OA, average classification accuracy AA, and Kappa coefficient are calculated for performance comparison. The results are shown in Table 1.

[0106] Table 1 Comparison of the classification performance of the present invention and seven existing methods on the Indian Pines dataset

[0107]

[0108] Simulation 2. Under the above simulation conditions, the present invention and the existing seven hyperspectral remote sensing image classification methods CEGCN, MDGCN, S2GAT, 3DOCT-SSAN, A2S2K-ResNet, SVM, and RF-200 are used to classify the paviaU dataset, and their respective overall accuracy OA, average classification accuracy AA, and Kappa coefficient are calculated for performance comparison. The results are shown in Table 2.

[0109] Table 2 Comparison of the classification performance of the present invention and the existing seven methods on the paviaU dataset

[0110]

[0111] Simulation 3. Under the above simulation conditions, the present invention and the existing seven hyperspectral remote sensing image classification methods CEGCN, MDGCN, S2GAT, 3DOCT-SSAN, A2S2K-ResNet, SVM, and RF-200 are used to classify the Houston 2013 dataset, and their respective overall accuracy OA, average classification accuracy AA, and Kappa coefficient are calculated for performance comparison. The results are shown in Table 3.

[0112] Table 3 Performance comparison of the present invention and the existing seven methods in the classification of the Houston2013 dataset

[0113]

[0114] Simulation 4. Under the above simulation conditions, the present invention and the existing seven hyperspectral remote sensing image classification methods CEGCN, MDGCN, S2GAT, 3DOCT-SSAN, A2S2K-ResNet, SVM, and RF-200 are used to classify the Salinas dataset, and their respective overall accuracy OA, average classification accuracy AA, and Kappa coefficient are calculated for performance comparison. The results are shown in Table 4.

[0115] Table 4 Comparison of the classification performance of the present invention and the existing seven methods on the Salinas dataset

[0116]

[0117] The sources of the seven prior arts are:

[0118] CEGCN is a method for hyperspectral image classification published by Liu Qichao et al. in IEEE, namely: Liu Q, XiaoL, Yang J, et al. CNN-enhanced graph convolutional network with pixel-andsuperpixel-level feature fusion for hyperspectral image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020, 59(10): 8657-8671.

[0119] MDGCN is a method for hyperspectral image classification published in IEEE, namely: Wan S, Gong C, Zhong P, et al. Multiscale dynamic graph convolutional network for hyperspectral image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2019, 58(5): 3162-3177.

[0120] S2GAT published a method for hyperspectral image classification in IEEE, namely: WShaA, Wang B, Wu X, et al. Semisupervised classification for hyperspectral images using graphattention networks[J]. IEEE Geoscience and Remote Sensing Letters, 2020, 18(1): 157-161.

[0121] 3DOCT-SSAN is a method for hyperspectral image classification published in IEEE, namely: X. Tang, F. Meng, X. Zhang, Y.-M. Cheung, J. Ma, F. Liu, and L. Jiao, “Hyperspectral image classificationbased on 3-d octave convolution with spatial–spectral attention network,” IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 3, pp. 2430–2447, 2020.

[0122] A2S2K-ResNet is a method for hyperspectral image classification published in IEEE, namely: SK Roy, S. Manna, T. Song, and L. Bruzzone, "Attention based adaptive spectral-spatial kernel resnet for hyperspectral image classification," IEEE Transactions on Geoscience and Remote Sensing, 2020.

[0123] SVM method for hyperspectral image classification published in IEEE, namely: JAGualtieri andS.Chettri, "Support vector machines for classification ofhyperspectraldata," in Proc.IEEE Int.Geosci.Remote Sens.Symp.(IGARSS),vol.2,Jul.2000,pp.813–815.

[0124] RF-200 published a method for hyperspectral image classification in IEEE, namely: J.Ham, Y.Chen, M.M.Crawford, and J.Ghosh, "Investigation of the random forest framework for classification of hyperspectral data," IEEE Trans.Geosci.Remote Sens., vol.43, no.

[0125] 3,pp.492–501,Mar.2005.

[0126] It is obvious from Tables 1 to 4 that the overall classification accuracy and average classification accuracy of the present invention in four commonly used hyperspectral remote sensing data sets are higher than those of the seven existing methods, indicating that the present invention has better image classification performance than the existing methods.

Claims

1. A hyperspectral image classification method based on graph convolution and convolution fusion, characterized in that: The steps include: (1) Obtain the original data set and data annotation set of hyperspectral images from the public website, select the labels that are not 0 from the data annotation set to form the training set, and use all the remaining samples to form the test sample set; (2) Preprocessing the original data set of hyperspectral images: 2a) Using a simple linear iterative algorithm to segment the hyperspectral image; 2b) calculating the region matrix V, adjacency matrix A, and region and pixel mapping matrix Q of the segmented image; 3) Build a feature extraction network: 3a) Establish a region-level feature extraction module consisting of three graph convolution layers, one graph pooling layer, and one graph de-pooling layer to extract the dependencies between regions; 3b) Build a pixel-level feature extraction module consisting of a multi-scale feature extraction layer and an existing attention layer to extract spectral information and spatial features on a single pixel; 3c) establishing a fusion module for fusing the features of 3a) and 3b); 3d) connecting the region-level feature extraction module and the pixel-level feature module in parallel, and then connecting them in series with the fusion module to form a feature extraction network, and using the cross entropy loss function as the loss function of the feature extraction network; (4) Train the feature extraction network: 4a) Input the original image and the region matrix V, the adjacency matrix M, and the region and pixel mapping matrix Q into the feature extraction network, and calculate the loss value between the feature extraction network output and the true label of the training set; 4b) Using the stochastic gradient descent method, gradually reduce the value of the loss function to update the network parameters until the set maximum number of iterations is completed to obtain a trained feature extraction network; (5) Using the trained feature extraction network, the classification results of the test set hyperspectral images are obtained: 5a) Input the original image and the region matrix V, the adjacency matrix M, and the region-pixel mapping matrix Q into the trained feature extraction network to obtain the output vector F; 5b) Using the argmax function on the channel dimension of the output vector F, calculate the position index of the maximum value on the channel dimension, and the position index is the category of all pixels in the entire image; 5c) Find the value of the corresponding coordinate of the test set in the categories of all pixels in the entire image, which is the classification result of the test set.

2. The method according to claim 1, characterized in that In step 2a), the hyperspectral image is segmented using a simple linear iterative algorithm, which is performed using the following formula: p'=Slic(p,Z) Among them, P' is the segmented image, P∈H×W×B is the original hyperspectral image, Z=(H×W) / a is the number of divided blocks, H, W, B represent the length, width and height of the image respectively, and a is a hyperparameter used to control the superpixel size.

3. The method according to claim 1, characterized in that In step 2b), the region matrix V of the segmented image is calculated using the following formula: Among them, S i represents the i-th superpixel block, Indicates that there is i The jth pixel in N z Represents the superpixel block S Z The number of pixels contained in .

4. The method according to claim 1, characterized in that In step 2b), the adjacency matrix A of the segmented image is calculated using the following formula: in Represents the element in the i-th row and j-th column of the matrix A, i.e., the i-th superpixel block V i With the jth superpixel block V j , N(·) represents the neighborhood of the superpixel, and y represents the set hyperparameter.

5. The method according to claim 1, characterized in that In step 2b), the region-pixel mapping relationship matrix Q is calculated for the segmented image, and the formula is as follows: Where Q∈(Z,H×W), Z is the number of superpixel blocks, H and W are the image length and width respectively; Q i,j Represents the element in the i-th row and j-th column of the matrix Q, i.e., the superpixel block S j With Pixel X i The mapping relationship is determined as follows:

6. The method according to claim 1, characterized in that In step 3a), three graph convolution layers, one graph pooling layer, and one graph de-pooling layer are formed in the region-level feature extraction module. The structural parameters and transmission relationships are as follows: The first graph convolution layer and the second graph convolution layer both include a graph convolution with a dimension parameter of 128; The third graph convolution layer includes a graph convolution with a dimension parameter of M, where M represents the number of image categories; The first graph pooling layer includes a pooling operation; The first image de-pooling layer includes a de-pooling operation; The input features pass through the first graph convolution layer in sequence to obtain the output first feature F1, and F1 passes through the first graph pooling layer, the second graph convolution layer, and the first graph de-pooling layer in sequence to obtain the second feature F2; after the output feature F1 of the first graph convolution layer is added to the output F2 of the first graph de-pooling layer, it passes through the third graph convolution layer to output the feature F of the final regional feature extraction module.

7. The method according to claim 1, characterized in that Step 3b) Build a pixel-level feature extraction module consisting of a multi-scale feature extraction layer and an existing attention layer, and implement it as follows: 3b1) Establish a multi-scale feature extraction layer consisting of a point convolution layer connected in parallel with three channel-by-channel convolution layers and then added together to extract the spectral and spatial features of the pixels, where: The convolution kernel size of the point convolution layer is 1×1; The convolution kernel size of the first channel-by-channel convolution layer is 3×3; The convolution kernel size of the second channel-by-channel convolution layer is 5×5; The convolution kernel size of the third channel-by-channel convolution layer is 7×7; The features output by the three channel-by-channel convolutional layers are F11, F21, and F31, and the output feature after addition is F41. Among them, F41 = F11 + F21 + F31; 3b2) Select the existing attention layer to embed the position information into feature learning and increase the feature perception ability of spatial position information; 3b3) The multi-scale feature extraction layer and the attention layer are connected in series to form the pixel-level feature extraction module.

8. The method according to claim 1, characterized in that In step 4a), the loss value between the feature extraction network output and the actual label of the training set is calculated as follows: L1(θ)=-[y1 logy'+(1-y1)log(1-y')] Among them, θ is the initial parameter of the network obtained by random initialization, y1 is the true label, and y' is the output of the entire network.

9. The method according to claim 1, characterized in that: In step 4b), the stochastic gradient descent method is used to gradually reduce the value of the loss function to update the network parameters, which is implemented as follows: 4b1) Initial training batch m = 0, maximum number of iterations in the training phase T = 200, learning rate α; 4b2) According to the loss value of the feature extraction network output and the true label of the training set, calculate the feature extraction network parameter θ after the current training update m+1 : Among them, θ is the initial parameter of the network obtained by random initialization, θ m are the network parameters before the current training update; 4b3) Repeat step 4b2) until the maximum number of iterations T is reached, completing the training of the feature extraction network.

10. The method according to claim 1, characterized in that Step 5b) Use the argmax function to calculate the position index of the maximum value on the channel dimension of the output vector F. The formula is as follows: F out =argmax(F) Among them, F out Represents the maximum position index of the image, F out ∈H×W×1, F∈H×W×k is the output feature vector, H, W are the image length and width respectively, and k is the total number of categories.

Citation Information

Patent Citations

  • Hyperspectral classification method combining graph structure and convolutional neural network

    CN113920442A

  • Image classification method of fusion network based on convolutional neural network and enhanced graph attention network

    CN116152561A