Data representation learning method and system based on hybrid domain graph neural network HGCN

By using the Hybrid Graph Neural Network (HGCN), which combines spatial and spectral graph convolutional layers and introduces residual connections, the problems of over-smoothing and understability in fitting of edge-dense datasets are solved, resulting in more stable data representation learning performance.

CN120929978APending Publication Date: 2025-11-11SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510941932.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

When processing edge-dense datasets, existing technologies suffer from over-smoothing in spectral domain graph convolutional neural networks and under-stability in spatial domain graph convolutional neural networks, leading to uncertainty in model output results.

Method used

A hybrid domain graph neural network (HGCN) is adopted, which combines multi-head attention convolutional layers in the spatial domain and Laplacian matrix spectral decomposition graph convolutional layers in the spectral domain. The model is optimized through residual connections to solve the problems of local feature extraction and global information utilization.

Benefits of technology

It improves the model's performance and stability on edge-dense datasets, enhances data representation learning capabilities, better balances global low-frequency and local high-frequency information in graph-structured data, and reduces the model's sensitivity to the proportion of convolutional layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929978A_ABST
    Figure CN120929978A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses a data representation learning method and system based on a hybrid domain graph neural network HGCN. The mixed domain graph-based neural network comprises at least one multi-head attention convolution layer based on a space domain and at least two mixed domain graph neural networks based on Laplacian matrix spectral decomposition graph convolution layers based on a spectral domain; residual connection is inserted between the multi-head attention convolution layer and the image convolution layer; the data representation learning method comprises the following steps: obtaining and preprocessing topological graph structure data to be learned including a node feature matrix; inputting the node feature matrix into the trained neural network based on the mixed domain graph; and the trained neural network based on the mixed domain graph outputs a prediction node classification label matrix as a prediction classification result of all nodes of the graph structure data. According to the method, the problem of excessive smoothing during edge-intensive data set processing in the prior art is solved, and the method has the characteristic of stable result fitting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and more specifically, to a data representation learning method and system based on Hybrid Graph Neural Network (HGCN). Background Technology

[0002] Graph neural networks are an efficient and concise deep learning network framework that plays an important role in fields such as molecular docking and social networks. The purpose of this technique is to characterize irregular topological graph structure data.

[0003] The core of graph neural networks (Graph Neural Networks) is to apply learnable convolution operators to graph-structured data to extract its features, and finally obtain the feature information of the data through a finite number of iterations. According to the definition of convolution operators, graph neural networks are further divided into spectral domain graph convolutional neural networks and spatial domain graph convolutional neural networks.

[0004] Spectral domain graph convolutional neural networks (GCNs) extract features from graph-structured data by performing eigenvalue decomposition on the Laplacian matrix, transforming the convolution operation between nodes and kernels in the graph into a composite operation of Fourier transform and dot product. These networks have achieved significant results in semi-supervised node classification tasks on small to medium-sized datasets, demonstrating stable training performance and superior average performance. Classic spectral domain GCN models include GCN.

[0005] Spatial domain graph convolutional neural networks (GATs) utilize local information aggregation mechanisms to define convolution. Specifically, they extract hidden layer features of each node by using information from its neighbors and edges. This extraction process, involving a finite number of steps, achieves information transfer across the graph and extracts the feature information from the graph structure data. Convolutional kernels defined using this method can efficiently extract high-frequency information from graph structure data, accurately characterizing the differentiated features of different local subgraphs. Furthermore, GATs combine attention mechanisms and local information aggregation mechanisms, further enhancing the local representation capabilities of spatial domain-based graph convolutional neural networks.

[0006] However, when dealing with edge-dense datasets, spectral domain graph convolutional neural networks often overemphasize the global low-frequency features of the data, making it difficult to accurately characterize the fine-grained differences between local nodes, resulting in over-smoothing and an inability to effectively extract local features. Spatial domain graph convolutional neural networks are very sensitive to local data perturbations, making the model prone to underfitting, increasing the difficulty of model parameter tuning, and causing uncertainty in the output results.

[0007] In summary, the urgent technical problem to be solved in this field is how to invent a data representation learning method and system based on hybrid domain graph neural networks (HGCN) that can simultaneously solve the over-smoothing problem of spectral domain graph convolutional neural networks when processing edge-dense datasets, and the under-fitting instability problem of spatial domain graph convolutional neural networks when processing such datasets. Summary of the Invention

[0008] To address the problem of excessive smoothing in existing technologies when processing edge-dense datasets, this invention provides a data representation learning method and system based on Hybrid Graph Neural Network (HGCN), which features stable result fitting.

[0009] To achieve the above-mentioned objectives of this invention, the technical solution adopted is as follows:

[0010] A hybrid domain graph neural network includes at least one multi-head attention convolutional layer in the spatial domain and at least two graph convolutional layers in the spectral domain based on Laplacian matrix spectral decomposition; residual connections are inserted between the multi-head attention convolutional layer and the graph convolutional layers.

[0011] Preferably, the residual connection is configured such that the first intermediate feature output by the multi-head attention convolutional layer is added to the feature output by the first graph convolutional layer, and then used as the input to the next graph convolutional layer.

[0012] A training method based on a hybrid domain graph neural network includes the following specific steps:

[0013] Acquire and preprocess the graph structure data for training, including the node feature matrix;

[0014] The node feature matrix is ​​input into at least one multi-head attention convolutional layer based on the spatial domain of the hybrid domain graph neural network for processing to obtain the first intermediate feature;

[0015] The first intermediate feature output from the multi-head attention convolutional layer is added to the feature output from the first graph convolutional layer, and then used as the input to the next graph convolutional layer; the last graph convolutional layer outputs the predicted node classification label matrix.

[0016] Calculate the cross-entropy loss function between the predicted node classification label matrix and the true node classification label matrix;

[0017] Gradient updates are performed based on the cross-entropy loss function, and the node feature matrix is ​​re-input into at least one multi-head attention convolutional layer for processing until the hybrid domain graph neural network converges, resulting in a trained hybrid domain graph neural network model.

[0018] Preferably, the training graph structure data including the node feature matrix is ​​obtained, specifically: given n nodes and m categories, the training graph structure data including the node feature matrix X and the adjacency matrix A of the edge-dense dataset is obtained.

[0019] Furthermore, the node feature matrix is ​​input into at least one multi-head attention convolutional layer for processing to obtain the first intermediate feature. The specific steps include:

[0020] There are K heads, and the l-th head corresponds to the shared attention mechanism a. l : Parameter matrix to be learned For any given point i, compute the normalized attention coefficients of its any neighboring nodes j with respect to i.

[0021]

[0022] in and Let W be the node feature vectors of the i-th and j-th nodes, respectively, and let σ be a nonlinear function. l Let be the weight matrix to be learned, be the weight vector of the shared attention mechanism, || be the concatenation function, and N be the weight matrix to be learned. i Let i be the set of neighboring nodes of node i;

[0023] The features of the neighboring nodes are weighted and summed using the normalized attention coefficients, and then processed by a nonlinear function to calculate the new node feature vector of the i-th node.

[0024]

[0025] Obtain the first intermediate feature matrix

[0026] Furthermore, the first intermediate feature input from the multi-head attention convolutional layer is added to the feature output from the first graph convolutional layer, and this sum is used as the input to the next graph convolutional layer; the last graph convolutional layer outputs the predicted node classification label matrix. The specific steps are as follows:

[0027] Calculate the renormalized Laplacian matrix based on the adjacency matrix A.

[0028]

[0029] Where diag(·) represents transforming a vector into a diagonal matrix, with each diagonal element corresponding to a one-to-one vector element. N Represents an N-dimensional identity matrix;

[0030] The first convolutional layer f1(·) based on the Laplacian matrix spectral decomposition is represented as:

[0031]

[0032] Among them W (1) Let f1 be the parameter matrix to be learned in the first layer. Let X2 = f1(X1) be the new node feature vector obtained after passing through the first convolutional layer based on the Laplacian matrix spectral decomposition.

[0033] The second convolutional layer f2(X2) based on Laplacian matrix spectral decomposition is represented as:

[0034]

[0035] Among them W (2) This is the parameter matrix to be learned in the second layer;

[0036] Let Y be the predicted node classification label matrix obtained after the Nth convolutional layer based on Laplacian matrix spectral decomposition. Then Y = f N (X N ).

[0037] Furthermore, calculation When σ is selected, LeakyReLU is used to calculate the node feature vectors. When f1(·) is constructed, ReLU is used for σ; when .... n When (·), σ is selected using Softmax.

[0038] Furthermore, the cross-entropy loss function between the predicted node classification label matrix and the true node classification label matrix is ​​calculated. The specific steps are as follows:

[0039] Let the predicted node classification label matrix be Y = (y ik ), where y ik This represents the probability that node i is predicted to be of category k; obtain the classification label matrix of the real nodes. in Indicate whether node i belongs to category k; calculate the cross-entropy loss function:

[0040]

[0041] A data representation learning method based on Hybrid Graph Neural Network (HGCN) includes the following specific steps:

[0042] Acquire and preprocess the topological graph structure data to be learned, including the node feature matrix;

[0043] Input the node feature matrix into the trained hybrid domain graph neural network;

[0044] The trained hybrid domain graph neural network outputs a predicted node classification label matrix, which serves as the predicted classification result for all nodes in the graph structure data.

[0045] A data representation learning system based on Hybrid Graph Neural Network (HGCN) includes a data acquisition module, a feature fusion module, a global feature module, and a result output module.

[0046] The data acquisition module is used to acquire and preprocess the topological graph structure data to be learned, including the node feature matrix;

[0047] The feature fusion module is used to input the node feature matrix into the spatial domain-based multi-head attention convolutional layer of the hybrid domain graph neural network for processing to obtain the first intermediate feature.

[0048] The global feature module is used to add the first intermediate feature input from the multi-head attention convolutional layer to the feature output from the first graph convolutional layer, and use the result as the input to the next graph convolutional layer; the last graph convolutional layer outputs the predicted node classification label matrix.

[0049] The result output module is used to output the predicted node classification label matrix as the predicted classification result for all nodes of the graph structure data.

[0050] The beneficial effects of this invention are as follows:

[0051] This invention proposes a hybrid-domain graph neural network (HNN) that combines the advantages of spectral and spatial domain HNNs. It utilizes multi-head attention and a global Laplacian matrix to address the problems of local feature extraction and global information utilization, respectively. Furthermore, a residual connection optimization model is introduced to improve the algorithm's performance and stability on edge-dense datasets. Thus, this invention addresses the over-smoothing problem of spectral domain HNNs when processing edge-dense datasets, and the understability of spatial domain HNNs when handling such datasets. By improving the algorithm, it enhances the data representation learning ability on edge-dense datasets. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of the hybrid domain graph neural network in Example 1.

[0053] Figure 2 This is a schematic diagram of the training method based on hybrid domain graph neural networks.

[0054] Figure 3 This is a schematic diagram illustrating the evolution of the single-head attention mechanism in Example 2.

[0055] Figure 4This is a schematic diagram illustrating the evolution of the multi-head attention mechanism in Example 2.

[0056] Figure 5 This is a flowchart illustrating a data representation learning method based on hybrid domain graph neural networks. Detailed Implementation

[0057] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0058] Example 1

[0059] like Figure 1 As shown, a hybrid domain graph neural network includes a single-layer multi-head attention convolutional layer based on the spatial domain and a two-layer graph convolutional layer based on Laplacian matrix spectral decomposition based on the spectral domain; residual connections are inserted between the multi-head attention convolutional layer and the graph convolutional layer.

[0060] In one specific embodiment, the residual connection is configured such that the first intermediate feature output by the multi-head attention convolutional layer is added to the feature output by the first graph convolutional layer, and then used as the input of the second graph convolutional layer.

[0061] Example 2

[0062] like Figure 2 As shown, a training method based on a hybrid domain graph neural network includes the following specific steps:

[0063] Acquire and preprocess the graph structure data for training, including the node feature matrix;

[0064] In this embodiment, three publicly available social network datasets are used as training datasets. All three datasets are edge-dense datasets: the Amason Computers dataset and the AmasonPhoto dataset from Amazon's shopping website, and the Twitch EN dataset from the Twitch live streaming platform. The preprocessed graph structure data used for training is randomly divided into training, validation, and test sets according to a 7:1:2 ratio.

[0065] The node feature matrix is ​​input into at least one multi-head attention convolutional layer based on the spatial domain of the hybrid domain graph neural network for processing to obtain the first intermediate feature;

[0066] The first intermediate feature output from the multi-head attention convolutional layer is added to the feature output from the first graph convolutional layer, and then used as the input to the next graph convolutional layer; the last graph convolutional layer outputs the predicted node classification label matrix.

[0067] Calculate the cross-entropy loss function between the predicted node classification label matrix and the true node classification label matrix;

[0068] Backpropagation is performed based on the cross-entropy loss function to update network parameters, thereby performing gradient updates; the node feature matrix is ​​then re-input into at least one multi-head attention convolutional layer for processing until the hybrid domain graph neural network converges, resulting in a trained hybrid domain graph neural network model.

[0069] In this embodiment, irregular topological graph structure data is obtained by: setting n nodes and m categories, and using the node feature matrix X and adjacency matrix A of the edge-dense dataset as input data.

[0070] In one specific embodiment, the node feature matrix is ​​input into at least one multi-head attention convolutional layer for processing to obtain the first intermediate feature. The specific steps include:

[0071] There are K heads, and the l-th head corresponds to the shared attention mechanism. Parameter matrix to be learned For any given point i, compute the normalized attention coefficients of its any neighboring nodes j with respect to i.

[0072]

[0073] in and Let W be the node feature vectors of the i-th and j-th nodes, respectively, and let σ be a nonlinear function. l Let be the weight matrix to be learned, be the weight vector of the shared attention mechanism, || be the concatenation function, and N be the weight matrix to be learned. i Let i be the set of neighboring nodes of node i;

[0074] In this embodiment, the evolution of the single-head attention mechanism is as follows: Figure 3 As shown, the evolution of multi-head attention mechanisms is as follows: Figure 4 As shown.

[0075] The features of the neighboring nodes are weighted and summed using the normalized attention coefficients, and then processed by a nonlinear function to calculate the new node feature vector of the i-th node.

[0076]

[0077] Obtain the first intermediate feature matrix

[0078] In one specific embodiment, calculation When σ is selected, LeakyReLU is used to calculate the node feature vectors. When σ is selected, ReLU is used.

[0079] In one specific embodiment, the first intermediate feature input from the multi-head attention convolutional layer is added to the feature output from the first graph convolutional layer, and this addition is used as the input to the next graph convolutional layer; the last graph convolutional layer outputs a predicted node classification label matrix. The specific steps are as follows:

[0080] Calculate the renormalized Laplacian matrix based on the adjacency matrix A.

[0081]

[0082] Where diag(·) represents transforming a vector into a diagonal matrix, with each diagonal element corresponding to a one-to-one vector element. N Represents an N-dimensional identity matrix;

[0083] The first convolutional layer f1(·) based on the Laplacian matrix spectral decomposition is represented as:

[0084]

[0085] Among them W (1) Let f1 be the parameter matrix to be learned in the first layer. Let X2 = f1(X1) be the new node feature vector obtained after passing through the first convolutional layer based on the Laplacian matrix spectral decomposition.

[0086] The second convolutional layer f2(X2) based on Laplacian matrix spectral decomposition is represented as:

[0087]

[0088] Among them W (2) This is the parameter matrix to be learned in the second layer;

[0089] Let Y be the predicted node classification label matrix obtained after the Nth convolutional layer based on Laplacian matrix spectral decomposition. Then Y = f N (X N ).

[0090] In one specific embodiment, when constructing f1(·), σ is selected as ReLU; when constructing f n When (·), σ is selected using Softmax.

[0091] The training results of the Hybrid Graph Neural Network (HGCN) model obtained in this invention and its implementation with other methods in node classification tasks are presented below. (HGCN-2 represents a two-layer HGCN model, as shown in Table 1, where the first layer is a multi-head attention convolutional layer and the second layer is a convolutional layer based on Laplacian matrix spectral decomposition; HGCN-3 represents a three-layer HGCN model; Accuracy represents the average accuracy of 100 classifications, with a larger value indicating better performance; Variance represents the variance of the accuracy of 100 classifications, with a smaller value indicating more stable training performance.)

[0092] Table 1

[0093]

[0094] The table above leads to the following conclusions: The classification performance and training stability of the hybrid domain graph neural network model HGCN obtained by this invention are significantly better than those of the GCN model. Furthermore, while achieving comparable classification results to the GAT model, the HGCN model demonstrates superior training stability. This verifies the effectiveness of the proposed hybrid domain graph convolution method in learning representations of graph-structured data, especially edge-dense datasets. It better balances global low-frequency and local high-frequency information in graph-structured data while removing local noise.

[0095] Example 3

[0096] like Figure 5 As shown, a data representation learning method based on a hybrid domain graph neural network includes the following specific steps:

[0097] Acquire and preprocess the topological graph structure data to be learned, including the node feature matrix;

[0098] Input the node feature matrix into the trained hybrid domain graph neural network;

[0099] The trained hybrid domain graph neural network outputs a predicted node classification label matrix, which serves as the predicted classification result for all nodes in the graph structure data.

[0100] Example 4

[0101] A data representation learning system based on Hybrid Graph Neural Network (HGCN) includes a data acquisition module, a feature fusion module, a global feature module, and a result output module.

[0102] The data acquisition module is used to acquire and preprocess the topological graph structure data to be learned, including the node feature matrix;

[0103] The feature fusion module is used to input the node feature matrix into the spatial domain-based multi-head attention convolutional layer of the hybrid domain graph neural network for processing to obtain the first intermediate feature.

[0104] The global feature module is used to add the first intermediate feature input from the multi-head attention convolutional layer to the feature output from the first graph convolutional layer, and use the result as the input to the next graph convolutional layer; the last graph convolutional layer outputs the predicted node classification label matrix.

[0105] The result output module is used to output the predicted node classification label matrix as the predicted classification result for all nodes of the graph structure data.

[0106] This system has the following beneficial effects:

[0107] Unlike traditional graph convolutional neural networks in the spectral and spatial domains, this invention proposes a hybrid domain graph neural network that integrates the characteristics of both types of graph convolutional neural networks. This allows the network to focus on both global information and local feature differences in the data, enabling it to learn feature information at multiple scales and improving model performance and stability.

[0108] Unlike conventional methods that directly combine two different types of graph convolutional layers, this invention introduces residual connections, which directly map the feature information extracted by the shallow multi-head attention layer to the deep layer. This allows the deep convolutional layers to combine the feature information output by the shallow layer with the feature information output by the previous layer for representation processing. On the one hand, this avoids the problems of gradient vanishing and gradient exploding, and on the other hand, it reduces the model's sensitivity to the ratio of the two different types of convolutional layers, thereby improving the model's stability during training.

[0109] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the claims of the present invention.

Claims

1. A hybrid domain graph neural network, characterized in that: The method comprises a hybrid domain graph neural network consisting of at least one multi-head attention convolutional layer based on the spatial domain and at least two graph convolutional layers based on Laplacian matrix spectral decomposition based on the spectral domain; residual connections are inserted between the multi-head attention convolutional layer and the graph convolutional layer.

2. The Hybrid Graph Neural Network (HGCN) based on claim 1, characterized in that: The residual connection is configured such that the first intermediate feature output by the multi-head attention convolutional layer is added to the feature output by the first graph convolutional layer, and then used as the input to the next graph convolutional layer.

3. A training method based on a hybrid domain graph neural network, characterized in that: The specific steps include the following: Acquire and preprocess the graph structure data for training, including the node feature matrix; The node feature matrix is ​​input into at least one multi-head attention convolutional layer based on the spatial domain of the hybrid domain graph neural network for processing to obtain the first intermediate feature; The first intermediate feature output from the multi-head attention convolutional layer is added to the feature output from the first graph convolutional layer, and then used as the input to the next graph convolutional layer; the last graph convolutional layer outputs the predicted node classification label matrix. Calculate the cross-entropy loss function between the predicted node classification label matrix and the true node classification label matrix; Gradient updates are performed based on the cross-entropy loss function, and the node feature matrix is ​​re-input into at least one multi-head attention convolutional layer for processing until the hybrid domain graph neural network converges, resulting in a trained hybrid domain graph neural network model.

4. The training method based on a hybrid domain graph neural network according to claim 3, characterized in that: To obtain training graph structure data including node feature matrices, specifically: given n nodes and m categories, obtain training graph structure data including node feature matrices X and adjacency matrices A from an edge-dense dataset.

5. The training method based on a hybrid domain graph neural network according to claim 4, characterized in that: The node feature matrix is ​​input into at least one multi-head attention convolutional layer for processing to obtain the first intermediate feature. The specific steps include: There are K heads, and the l-th head corresponds to the shared attention mechanism a. l : Parameter matrix to be learned For any given point i, compute the normalized attention coefficients of its any neighboring nodes j with respect to i. in and Let W be the node feature vectors of the i-th and j-th nodes, respectively, and let σ be a nonlinear function. l Let be the weight matrix to be learned, be the weight vector of the shared attention mechanism, || be the concatenation function, and N be the weight matrix to be learned. i Let i be the set of neighboring nodes of node i; The features of the neighboring nodes are weighted and summed using the normalized attention coefficients, and then processed by a nonlinear function to calculate the new node feature vector of the i-th node. Obtain the first intermediate feature matrix 6. The training method based on a hybrid domain graph neural network according to claim 5, characterized in that: The first intermediate feature output from the multi-head attention convolutional layer is added to the feature output from the first graph convolutional layer, and then used as the input to the next graph convolutional layer. The specific steps are as follows: Calculate the renormalized Laplacian matrix based on the adjacency matrix A. Where diag(·) represents transforming a vector into a diagonal matrix, with each diagonal element corresponding to a one-to-one vector element. N Represents an N-dimensional identity matrix; The first convolutional layer f1(·) based on the Laplacian matrix spectral decomposition is represented as: Among them W (1) Let f1 be the parameter matrix to be learned in the first layer. Let the new node feature vector obtained after passing through the first convolutional layer based on the Laplacian matrix spectral decomposition be f2 = f1(X1). The second convolutional layer f2(X2) based on Laplacian matrix spectral decomposition is represented as: Among them W (2) This is the parameter matrix to be learned in the second layer; Let Y be the predicted node classification label matrix obtained after the Nth convolutional layer based on Laplacian matrix spectral decomposition. Then Y = f N (X N ).

7. The training method based on a hybrid domain graph neural network according to claim 6, characterized in that: calculate When σ is selected, LeakyReLU is used to calculate the node feature vectors. When σ is used, ReLU is selected; when building f1(·), ReLU is selected; when building f n When (·), σ is selected using Softmax.

8. The training method based on a hybrid domain graph neural network according to claim 3, characterized in that: The specific steps for calculating the cross-entropy loss function between the predicted node classification label matrix and the true node classification label matrix are as follows: Let the predicted node classification label matrix be Y = (y ik ), where y ik This represents the probability that node i is predicted to be of category k; obtain the classification label matrix of the real nodes. in Indicate whether node i belongs to category k; calculate the cross-entropy loss function:

9. A data representation learning method based on Hybrid Graph Neural Network (HGCN), characterized in that: The specific steps include the following: Acquire and preprocess the topological graph structure data to be learned, including the node feature matrix; Input the node feature matrix into the trained hybrid domain graph neural network; The trained hybrid domain graph neural network outputs a predicted node classification label matrix, which serves as the predicted classification result for all nodes in the graph structure data.

10. A data representation learning system based on Hybrid Graph Neural Network (HGCN), characterized in that: It includes a data acquisition module, a feature fusion module, a global feature module, and a result output module; The data acquisition module is used to acquire and preprocess the topological graph structure data to be learned, including the node feature matrix; The feature fusion module is used to input the node feature matrix into the spatial domain-based multi-head attention convolutional layer of the hybrid domain graph neural network for processing to obtain the first intermediate feature. The global feature module is used to add the first intermediate feature input from the multi-head attention convolutional layer to the feature output from the first graph convolutional layer, and use the result as the input to the next graph convolutional layer; the last graph convolutional layer outputs the predicted node classification label matrix. The result output module is used to output the predicted node classification label matrix as the predicted classification result for all nodes of the graph structure data.