Graph neural network training method based on multi-granularity implicit information aggregation

By mapping the node feature learning process into Markov decision-making process, using reinforcement learning algorithms to obtain coarse, fine-grained information and implicit information of graph neural network nodes, and integrating these information, the shortcomings of graph neural networks in homogeneous graph structure processing are solved, and node classification performance is improved.

CN120146143APending Publication Date: 2025-06-13JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182810.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing graph neural networks perform poorly when dealing with homogeneous graph structures and ignore the importance of aggregation at different granularity levels in node representations, especially the implicit relationships between long-distance nodes.

Method used

By mapping the node feature learning process into the Markov decision-making process, using reinforcement learning algorithms to determine the aggregation hop number of nodes, obtain coarse, fine-grained information and implicit information, and integrate this information through the aggregation function to generate enhanced node features.

Benefits of technology

It realizes the generation of smoother and more accurate node embedding, which significantly improves the node classification performance of graph neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146143A_ABST
    Figure CN120146143A_ABST
Patent Text Reader

Abstract

The invention provides a graph neural network training method based on multi-granularity implicit information aggregation, and the method comprises the steps: obtaining the same number of same-matching and different-matching graph data containing node features, node tags and adjacent matrixes, so as to construct a training data set; acquiring graph data from the data set, and determining aggregation hops of each node by using a reinforcement learning algorithm to obtain coarse and fine granularity information and implicit information of each node; aggregating the coarse and fine granularity information and implicit information of the node by using an aggregation function to obtain an enhanced node feature of the node; predicting a node label of the node by using a classifier, and evaluating prediction performance; and taking the evaluated prediction performance as an award, and updating parameters of the graph neural network by taking accumulated award and maximization as a target until the data set is traversed. According to the method, node features are extracted from multi-view and multi-granularity levels, and smoother and more accurate node features are generated in combination with implicit information, so that the node classification performance of the graph neural network is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of graph data mining, and particularly relates to a method for training a graph neural network based on multi-granularity implicit information aggregation. Background Art

[0002] A graph neural network (GNN) is a type of deep neural network specifically designed to process graph-structured data. Its working mechanism can be summarized as follows: First, represent the data as nodes and edges of a graph, then utilize the unique properties of the graph to capture complex dependencies between entities, and finally, through an iterative message passing and aggregation process, combine the features of a node with those of its neighbors to update the node's representation. GNNs can learn rich and informative embeddings and can be applied to graph-based machine learning tasks such as node classification, link prediction, and graph classification. Among them, for graph structures with a homophilic assumption (nodes with a connection relationship share similar attributes and labels), GNNs perform well, while for graph structures with a heterophilic assumption (nodes with a connection relationship have different features or labels), it is difficult for GNNs to effectively aggregate information from connected nodes with different attributes.

[0003] To address the above problems, researchers have proposed the following solutions: 1) Adaptive simple graph convolution, which makes it equally adaptable to homophilic and heterophilic graph structures by selecting different filters for each feature; 2) Capturing homophilic and heterophilic information through dual-core feature transformation and introducing a selection gate to select an appropriate core for a given node; 3) Solving the spatial heterophily problem in urban graphs by designing a rotation-scaling spatial aggregation module and a heterophily-sensitive spatial interaction module; 4) Solving the label relationship prediction problem in heterophilic graphs by solving the approximation problem of the global label relationship matrix of a signed graph, making the proposed model applicable to both homophilic and heterophilic environments.

[0004] Although the existing technologies have achieved good results in dealing with heterophily, they usually perform poorly in dealing with homophily. At the same time, they often ignore the importance of aggregation at different granularity levels in node representation, where coarse-grained information captures the overall position or influence of a node (such as a user) in the network, while fine-grained information focuses on its direct interaction with closely connected nodes. In addition, these methods rarely consider the implicit relationships between distant nodes, which may not be adjacent but may have common interests or features, and this implicit information is crucial for generating smooth embeddings by capturing potential connections beyond adjacency. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method for training a graph neural network based on multi-granularity implicit information aggregation. By mapping the node feature learning process into a Markov decision process, reinforcement learning of node features is realized to obtain the coarse-grained, fine-grained, and implicit information of nodes, thereby generating smoother and more accurate node embeddings and effectively improving the node classification performance of the graph neural network.

[0006] This application provides a method for training a graph neural network based on multi-granularity implicit information aggregation, including:

[0007] Obtain the same number of assortative and disassortative graph data containing node features, node labels, and adjacency matrices to construct a training dataset;

[0008] Obtain graph data from the training dataset and use a reinforcement learning algorithm to determine the aggregation hop count of each node in the graph data to obtain the coarse-grained, fine-grained, and implicit information of each node;

[0009] For each node, use an aggregation function to aggregate the coarse-grained, fine-grained, and implicit information of the node to obtain the enhanced node features of the node;

[0010] Based on the enhanced node features of the node, use a classifier to predict the node label of the node and evaluate the prediction performance;

[0011] Use the evaluated prediction performance as a reward, aiming at maximizing the cumulative reward, and update the parameters of the reinforcement learning algorithm until the training dataset is traversed.

[0012] Further, the use of the reinforcement learning algorithm to determine the aggregation hop count of each node in the graph data to obtain the coarse-grained, fine-grained, and implicit information of each node includes:

[0013] Take the node features of any node in the graph data as the state of the current time step of the Markov decision process, and use the reinforcement learning algorithm to generate the action of the current time step to obtain the aggregation hop count of the node;

[0014] For each node, determine the information aggregation range of the node based on the aggregation hop count;

[0015] Within the information aggregation range, take the node features of the direct neighbor nodes of the node as fine-grained information, take the node features of the distant neighbor nodes of the node as coarse-grained information, and take the node features of the non-neighbor nodes with semantic similarity to the node as implicit information.

[0016] Further, after obtaining the aggregation hop count of the node, the method further includes:

[0017] Sample the next node from the neighbor nodes of the aggregated hop count as the state at the next time step of the Markov decision process, and use the reinforcement learning algorithm to generate the action at the next time step;

[0018] Store the state and action at each time step to avoid repeated calculations.

[0019] Further, the aggregating function is used to aggregate the coarse-grained, fine-grained, and implicit information of the node to obtain the enhanced node features of the node, including:

[0020] Based on the aggregated hop count, determine the aggregation layer number of the node;

[0021] Use the aggregating function to sequentially integrate the fine-grained information and the coarse-grained information of the node from the inside to the outside of the aggregation layer, and integrate the implicit information of the node with the node features containing the granularity information to obtain the enhanced node features of the node;

[0022] Among them, each aggregation layer shares the parameters of the aggregating function.

[0023] Further, based on the enhanced node features of the node, use a classifier to predict the node label of the node and evaluate the prediction performance, including:

[0024] Based on the enhanced node features of the node, use a classifier to predict the node label of the node to obtain the predicted node label of the node;

[0025] Obtain the original node label of the node;

[0026] Use the cross-entropy loss function to calculate the error between the predicted node label and the original node label of the node, and evaluate the prediction performance based on the error.

[0027] The graph neural network training method based on multi-granularity implicit information aggregation provided by this application maps the node feature learning process to a Markov decision process, realizes the reinforcement learning of node features to obtain the coarse-grained, fine-grained, and implicit information of the node, thereby generating a smoother and more accurate node embedding, and effectively improving the node classification performance of the graph neural network. Brief Description of the Drawings

[0028] Figure 1 Shows the flowchart of the graph neural network training method based on multi-granularity implicit information aggregation provided by the embodiment of this application. Detailed Embodiment

[0029] To make the objectives, technical solutions, and advantages of this technical solution clearer, the following further elaborates on this technical solution in conjunction with specific implementation manners. It should be understood that these descriptions are exemplary and not intended to limit the scope of this technical solution.

[0030] Please refer to Figure 1 the flowchart of the graph neural network training method based on multi-granularity implicit information aggregation as shown in Figure 1 the figure. As shown in

[0031] S101. Obtain the same number of assortative and disassortative graph data containing node features, node labels, and adjacency matrices to construct a training dataset.

[0032] In this step, a graph structure with an assortative hypothesis (nodes with connection relationships share similar attributes and labels) is called an assortative graph, and a graph structure with a disassortative hypothesis (nodes with connection relationships have different features or labels) is called a disassortative graph; for each assortative and disassortative graph, obtain the node features, node labels, and adjacency matrix of this graph to form the graph data of this assortative and disassortative graph; and in order to enable the graph neural network to have the same performance for assortative and disassortative graphs, the embodiments of this application obtain the same number of assortative and disassortative graph data and construct a training dataset to train the graph neural network.

[0033] S102. Obtain graph data from the training dataset, and use the reinforcement learning algorithm to determine the aggregation hop count of each node in this graph data to obtain the coarse-grained, fine-grained information, and implicit information of each node.

[0034] In this step, the embodiments of this application map the node feature learning process to a Markov decision process, so as to use the reinforcement learning algorithm to implement the reinforcement learning of node features. Before learning node features, it is first necessary to clarify the mapping relationship between the graph data in the node feature learning process and each component in the Markov decision process. It is known that the key components of the Markov decision process include state, action, and reward. Define each component as follows with the graph neural network as the interaction environment: Define the t-th iteration state s t as the node feature of the current node, define the t-th iteration action a t as the aggregation hop count of the current node, and define the t-th reward r t as the average value of the performance in the first few steps of the prediction task. Then, based on the components of the Markov decision process defined above, use the reinforcement learning algorithm to obtain the coarse-grained, fine-grained information, and implicit information of each node.

[0035] In specific implementation, the coarse-grained, fine-grained information, and implicit information of each node can be obtained through the following method:

[0036] Step 1021: Use the node features of any node in the graph data as the state at the current time step of the Markov decision process, and generate the action at the current time step using a reinforcement learning algorithm to obtain the aggregated hop count of the node.

[0037] In this step, the parameters of the reinforcement learning algorithm are adjustable. It adjusts the attention to different information granularities according to the feedback of the environment, i.e., the graph neural network, so as to optimize the reinforcement learning algorithm.

[0038] Step 1022: For each node, determine the information aggregation range of the node based on the aggregated hop count.

[0039] Step 1023: Within the information aggregation range, use the node features of the direct neighbor nodes of the node as fine-grained information, the node features of the distant neighbor nodes of the node as coarse-grained information, and the node features of the non-neighbor nodes with semantic similarity to the node as implicit information.

[0040] As an example, an Actor-Critic type reinforcement learning algorithm can be used to learn and adjust the attention to different granularity information, and the twin-delayed deep deterministic policy gradient (TD3) is used to learn the implicit information between nodes, so as to improve the representation ability of node embeddings.

[0041] In addition, after obtaining the aggregated hop count of the node in the embodiment of the present application, the method further includes:

[0042] Step 201: Sample the next node from the neighbor nodes of the aggregated hop count as the state at the next time step of the Markov decision process, and generate the action at the next time step using a reinforcement learning algorithm.

[0043] Step 202: Store the state and action of each time step to avoid repeated calculations.

[0044] Based on the above steps 201-202 to store the state and action of each time step, the state and action of the next time step can be directly obtained during training, thereby improving the training efficiency of the graph neural network.

[0045] S103: For each node, use an aggregation function to aggregate the coarse-grained, fine-grained information and implicit information of the node to obtain the enhanced node features of the node.

[0046] In specific implementation, the enhanced node features of the node can be obtained in the following way:

[0047] Step 1031: Based on the aggregated hop count, determine the aggregation layer number of the node.

[0048] Step 1032: Using an aggregation function, integrate the fine-grained information and coarse-grained information of the node in sequence from the inside to the outside of the aggregation layer, and integrate the implicit information of the node with the node features containing granularity information to obtain the enhanced node features of the node.

[0049] Among them, each aggregation layer shares the parameters of the aggregation function.

[0050] The aggregation functions shown in the above steps 1031 - 1032 can be expressed as shown in the following formulas (1) - (4):

[0051]

[0052]

[0053] In the formula, represents the feature vector of the node v at the k-th layer; X u represents the direct neighbor representation vector of the node u; k = <a t > represents the aggregation hop count, which is determined by the policy function π at the time step t. Since a t is continuous, rounding is chosen here; is the normalized adjacency matrix of the first-hop neighbors of the node v; in the formula for obtaining h v , the former term represents the comprehensive information of coarse-grained and fine-grained, and the latter term represents the implicit information of the potential connection and semantic similarity with the target node; and respectively represent the floor and ceiling functions. This formula allows each node to dynamically adjust the granularity information of aggregation, enhancing the learning ability of implicit information.

[0054] S104: Based on the enhanced node features of the node, use a classifier to predict the node label of the node and evaluate the prediction performance.

[0055] In this step, the enhanced node features can be applied not only to the node classification task, but also to graph-based machine learning tasks such as link prediction and graph classification. This application does not make any limitations here.

[0056] In specific implementation, the node label of the node can be predicted and the prediction performance can be evaluated in the following way:

[0057] Step 1041: Based on the enhanced node features of the node, use a classifier to predict the node label of the node to obtain the predicted node label of the node.

[0058] Step 1042: Obtain the original node label of the node.

[0059] Step 1043: Use the cross-entropy loss function to calculate the error between the node label predicted by this node and the original node label, and evaluate the prediction performance based on the error.

[0060] S105: Take the evaluated prediction performance as the reward, and aim at maximizing the cumulative reward to update the parameters of the reinforcement learning algorithm until the training data set is traversed.

[0061] In this step, the reward is also a key component of the Markov decision process. Under the graph neural network, the reward function is defined by the following formula:

[0062]

[0063] In the formula, φ is a hyperparameter that plays a role in determining the strength of the reward signal. is the hyperparameter of the historical step length from the current node during the optimization process, and F is the evaluation index of the graph node classification task.

[0064] Based on the reward function shown in the above formula (5), update the parameters of the reinforcement learning algorithm. By traversing the training data set, learn and adjust the attention to different granularity information and the embedding ability of implicit information, obtain smoother and more accurate node features, and further improve the prediction performance of the graph neural network.

[0065] The above content is only the preferred embodiment of the present invention. For those of ordinary skill in the art, according to the idea of the technical content of the present application, many changes can be made in the specific implementation manner and application scope. As long as these changes do not depart from the concept of the present invention, they all fall within the protection scope of this patent.

Claims

1. A graph neural network training method based on multi-granularity implicit information aggregation, characterized in that: The method comprises: Obtain the same number of homogamous and heterogamous graph data including node features, node labels and adjacency matrices to construct a training data set; Obtain graph data from the training data set, and use the reinforcement learning algorithm to determine the number of aggregated hops of each node in the graph data to obtain the coarse-grained and fine-grained information and implicit information of each node; For each node, the coarse-grained and fine-grained information and implicit information of the node are aggregated using an aggregation function to obtain the enhanced node features of the node; Based on the enhanced node features of the node, a classifier is used to predict the node label of the node, and the prediction performance is evaluated; The predicted performance of the evaluation is used as a reward, and the parameters of the reinforcement learning algorithm are updated with the goal of accumulating and maximizing the reward until the training data set is traversed.

2. The method according to claim 1, characterized in that The reinforcement learning algorithm is used to determine the aggregation hop count of each node in the graph data to obtain the coarse and fine granularity information and implicit information of each node, including: The node feature of any node in the graph data is used as the state of the current time step of the Markov decision process, and the action of the current time step is generated by the reinforcement learning algorithm to obtain the aggregate hop count of the node; For each node, determining the information aggregation range of the node based on the aggregation hop count; Within the scope of information aggregation, node features of the node's direct neighbor nodes are used as fine-grained information, node features of the node's distant neighbor nodes are used as coarse-grained information, and node features of non-neighbor nodes with semantic similarity to the node are used as implicit information.

3. The method according to claim 2, characterized in that After obtaining the aggregated hop count of the node, the method further includes: Sampling a next node from the neighboring nodes of the aggregated hop count as the state of the next time step of the Markov decision process, and generating an action for the next time step using a reinforcement learning algorithm; The state and action of each time step are stored to avoid repeated calculations.

4. The method according to claim 1, characterized in that The method of aggregating the coarse and fine granular information and implicit information of the node by using an aggregation function to obtain enhanced node features of the node includes: Based on the aggregation hop count, determining the number of aggregation layers of the node; Using an aggregation function, sequentially integrating the fine-grained information and the coarse-grained information of the node from the inside to the outside according to the aggregation layer, and integrating the implicit information of the node with the node feature containing the granular information to obtain an enhanced node feature of the node; Among them, each aggregation layer shares the parameters of the aggregation function.

5. The method according to claim 1, characterized in that The method of predicting the node label of the node by using a classifier based on the enhanced node feature of the node and evaluating the prediction performance includes: Based on the enhanced node features of the node, a classifier is used to predict the node label of the node to obtain a predicted node label of the node; Get the original node label of the node; The cross entropy loss function is used to calculate the error between the predicted node label and the original node label of the node, and the prediction performance is evaluated based on the error.