Heterogeneous graph influence node identification method based on meta-path and graph attention mechanism
By using the metapathic path and graph attention mechanism in heterogeneous graphs, the heterogeneous neighborhoods and the influence of nodes are constructed, and the problem of difficulty in capturing the complex relationship differences between nodes in metapaths in the existing technology is solved, and a higher-precision influencing node recognition is achieved.
Patent Information
- Application Number
- CN202510475370.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
AI Technical Summary
When identifying influence nodes in heterogeneous graphs, it is difficult for the prior art to fully capture the differences in complex relationships between nodes in the metapath, resulting in inaccurate identification results.
Using a method based on metapath and graph attention mechanism, a heterogeneous neighborhood is constructed through random walks, a skip-gram model is used to obtain node embeddings, and a graph attention network is used to assign different weights to adjacent nodes in the sub-graph network, and the entropy value is used to calculate the influence of nodes.
It significantly improves the recognition accuracy of influential nodes in heterogeneous graphs, and can more comprehensively capture the complex relationships between nodes, and is suitable for complex scenarios such as social network analysis, recommendation systems and knowledge graphs.
Smart Images

Figure CN119988689A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and specifically relates to a method for identifying influential nodes in heterogeneous graphs based on meta-path and graph attention mechanism. Background Art
[0002] With the widespread application of information networks, heterogeneous information networks have become an important tool for analyzing complex data. In heterogeneous networks, nodes and edges are of different types, and there are various complex relationships between nodes and edges. In heterogeneous networks, the identification of influential nodes is a crucial task, which involves discovering key nodes from the network that have a significant impact on information dissemination, disease control, social influence, etc. These nodes play a core role in the structure and function of the network. Therefore, identifying influential nodes is of great significance for understanding and controlling network dynamics. How to effectively identify influential nodes in the network is crucial for many practical applications (such as social network analysis, recommendation systems, etc.).
[0003] Most of the previous methods evaluate the importance of nodes based on node degree and graph structure, but these methods often cannot fully capture the complexity and diversity of heterogeneous networks, lack theoretical performance guarantees, and the influential nodes obtained may be closely related to the selected specific network. With the deepening of research, researchers began to explore the combination of machine learning and deep learning techniques to identify influential nodes. Compared with the method of evaluating node importance based on factors such as node degree, the machine learning-based method can comprehensively consider more factors by constructing machine learning features to identify influential nodes. However, the machine learning-based method needs to assign weights to different features, and there is a problem of low accuracy in feature selection. The deep learning-based method identifies influential nodes by learning low-dimensional representations of nodes in the network. These methods can automatically capture the complex relationships between nodes. However, when combining deep learning technology to process heterogeneous graphs, it faces the problem that nodes and edges in heterogeneous graphs have different types and attributes. In recent years, models based on meta-path embedding and graph neural networks have been widely used to capture complex relationships between nodes. Meta-path is an ordered sequence of node types and edge types defined in a heterogeneous graph, which describes the compound relationship between node types. However, most existing methods only consider the importance of different meta-paths and ignore the influence differences between adjacent nodes in each meta-path, resulting in inaccurate recognition results. Summary of the invention
[0004] In order to solve the above-mentioned technical problems existing in the prior art, the present invention provides a method for identifying influential nodes in heterogeneous graphs based on meta-path and graph attention mechanism.
[0005] The technical solution of the present invention to solve the above technical problems is: a method for identifying influential nodes in heterogeneous graphs based on meta-path and graph attention mechanism, comprising the following steps: Step S1, define the required meta-path, and use the meta-path embedding method to obtain the node embedding representation; the specific steps of S1 are: Step S 11 , select the meta-path of source node and terminal node type as target type; Step S 12 , according to the defined meta-path, the random walk method is used to construct the heterogeneous neighborhoods of various types of nodes, as follows: The heterogeneous network is represented as G = (V, E), where V and E are the sets of all nodes and edges in the heterogeneous network, respectively; Given a meta-path scheme P: , where each Represents a node type, each R i Represents a specific relationship type between node types; the transition probability between nodes is as follows: ; in, Represents the probability transfer formula between nodes, Representation Node The node type of is the neighbor node set of t+1, Representation Node The number of neighbor nodes of type t+1, Representation Node Type; Step S 13 , if the node and nodes There is an edge between nodes The type of is t+1, which means that it successfully follows the meta-path P for one step of random walk, and the transition probability is ; Step S 14 , if the node and nodes There is an edge between them, but the nodes If the type is not t+1, then the meta-path P is not followed for random walk, and the transition probability is 0; Step S 15 , if the node and nodes If there is no edge between them, the node does not follow the meta-path P for random walk, and the transition probability is 0; Step S 16 , the skip-gram model is used to obtain the feature embedding of the nodes in each meta-path; the skip-gram model learns the effective feature representation of nodes in heterogeneous networks by maximizing the probability of heterogeneous context node pairs; the maximum probability is defined as: ; Among them, arg max means finding the value of the model parameter θ that maximizes the following expression, represents a collection of context types, Representation Node a context node set of type t, is a given node and parameter θ, the context node The conditional probability of Representation Node The feature embedding of Represents the context node and nodes The dot product of the embedding vector of ; Step S 17 , extract the feature embedding of target type nodes in each meta-path; Step S2, generating a subgraph network corresponding to each meta-path according to random walk; Step S3, assigning different weight values to adjacent nodes in the subgraph network based on the graph attention network; Step S4, using entropy values to calculate the influence of nodes in the subgraph network, and aggregating the influence of nodes in all subgraphs to obtain the final influence value of each node.
[0006] Furthermore, the specific steps of S2 are: Step S 21 According to the predefined meta-path P, a random walk algorithm is used to start from the starting node and walk along the node types in the meta-path in turn to generate a node sequence ; Step S 22 ,For each random walk of each meta-path, ensure that the nodes walked conform to the node type order defined in the meta-path to preserve the structural and semantic relationships between nodes; Step S 23 ,The generated random walk sequence is used to construct the subgraph network corresponding to each meta-path, and each subgraph network reflects the connection pattern between nodes under the meta-path; Step S 24 , the constructed meta-path subgraph network only contains nodes whose node type under the meta-path belongs to the target type.
[0007] Furthermore, the specific steps of S3 are: Step S 31, extract nodes and their connection relationships in the generated subgraph network, and extract feature embeddings of target type nodes in the meta-path network, which are used as the initial feature vectors input to the graph attention network; Step S 32 , perform linear transformation on the initial feature vector of the target type node, the linear transformation formula is as follows: ; Step S 33 , the attention coefficient between each node and its neighbors is calculated as follows: ; in is the weight matrix is the attention coefficient, indicating that the node For Node The importance of LeakyReLU is the activation function. , They are the nodes after linear transformation. and nodes The embedding represents the transpose of the weight vector, and || represents the concatenation operation; Step S 34 , using the masked attention mechanism to calculate the node With all neighboring nodes The attention coefficients between are as follows: ; in Represents the node attention score, i.e., node For neighbor nodes The attention weight, exp represents the exponential function; Step S 35 , use the attention weights to aggregate the feature vectors of neighboring nodes and obtain the updated node embedding representation of each node. The formula is as follows: ; in, Representation Node The updated node feature vector, Represents the activation function and are the attention weight and feature vector respectively; The multi-head mechanism is introduced to learn the key information between nodes on a larger scale and more effectively capture the complex relationship between nodes in the graph. The formula is as follows: ; in and are the attention weight and feature vector calculated by the kth attention mechanism respectively; In the result obtained by the multi-head mechanism, each node has K output features, and the K output features are aggregated by the average value so that the feature output of each node constitutes a feature matrix.
[0008] Furthermore, the specific steps of S4 are: The cosine similarity is used to calculate the similarity between each subgraph node. The formula is as follows: ; in Representation Node and nodes The cosine similarity between and Respectively represent nodes and nodes The yth component of the eigenvector, d is the dimension of the eigenvector; Map the similarity results to between 0 and 1, the formula is as follows: ; Calculate the entropy of all nodes in each subgraph using the following formula: ; in Represents the node in the mth subgraph The entropy of Representation Node The set of neighbors in the mth subgraph, Representation Node The number of neighbors in the mth subgraph; Aggregate the entropy of all nodes in each subgraph to get the influence value of each node. The formula is as follows: .
[0009] The beneficial effects of the present invention are as follows: by introducing a graph attention network, the present invention can not only evaluate the relative importance of different meta-paths, but also fully consider the different influences of node neighbors in each meta-path on the target node. Compared with traditional methods, the present invention can more comprehensively capture the complex relationships between nodes in heterogeneous networks, and significantly improve the accuracy of identifying influential nodes. By aggregating multiple meta-paths, the present invention can establish richer semantic associations between different types of nodes and edges, and is particularly suitable for identifying key nodes in complex scenarios such as social network analysis, recommendation systems, and knowledge graphs. By measuring the influence of nodes using information entropy, the present invention can ensure the objectivity and accuracy of the calculation results. The entropy calculation not only considers the direct neighbors of the node, but also integrates the influence distribution in the entire network, making the influence evaluation more robust, especially in complex heterogeneous networks, the identification results are more reliable. The present invention has good scalability and generalization capabilities, and provides a new and efficient solution for the identification of influential nodes in heterogeneous graphs. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 Flowchart of the present invention; Figure 2 Schematic diagram of node type selection strategy in heterogeneous networks in the present invention. DETAILED DESCRIPTION
[0011] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0012] like Figure 1 As shown, Figure 1 The present invention provides a method for identifying influential nodes in heterogeneous graphs based on meta-path and graph attention mechanism, which specifically includes the following steps: Step S1, a method for obtaining node embedding representation using a meta-path selection strategy and a meta-vector method. First, a meta-path whose source and target nodes are of the target type is selected. Based on the defined meta-path, a random walk method is used to construct heterogeneous neighborhoods of nodes of each type, as follows: The heterogeneous network is represented as G = (V, E), where V and E are the sets of all nodes and edges in the heterogeneous network, respectively; Given a meta-path scheme P: , where each Indicates a node type, each R i Represents a specific type of relationship between node types. The transition probability between nodes is as follows: ; in, Representation Node The node type is the neighbor node set of t+1, that is, the next node selected must be of type t+1, Representation Node The number of neighbor nodes of type t+1, Representation Node This shows that the meta-path based random walk is a biased walk under the predefined meta-path P.
[0013] If the node and nodes There is an edge between nodes The type is , it means that the meta-path P is successfully followed for one step of random walk, and the transition probability is If the node and nodes There is an edge between them, but the nodes The type is not , then there is no random walk following metapath P, and the transition probability is 0. and nodes If there is no edge between them, the node does not follow the meta-path P for random walk, and the transition probability is 0.
[0014] Then the skip-gram model is used to obtain the feature embeddings of the nodes in each meta-path. Figure 2 The upper part shows a schematic diagram of implementing node embedding according to the present invention.
[0015] The skip-gram model learns effective feature representations of nodes in heterogeneous networks by maximizing the probability of heterogeneous context node pairs. The maximum probability is defined as: ; Among them, arg max means finding the value of the model parameter θ that maximizes the following expression, represents a collection of context types, Representation Node a context node set of type t, is a given node and parameter θ, the context node The conditional probability of . Representation Node The feature embedding of Represents the context node and nodes The dot product of the embedding vector of .
[0016] Finally, feature embeddings of target type nodes are extracted in each meta-path.
[0017] Step S2: Generate a subgraph network corresponding to each meta-path by random walk. Specifically, based on the pre-defined meta-path P, a random walk algorithm is used to start from the starting node and walk along the node types in the meta-path in turn to generate a node sequence .
[0018] For each random walk of each meta-path, ensure that the nodes walked conform to the node type order defined in the meta-path to preserve the structure and semantic relationship between nodes. The generated random walk sequence is used to construct the subgraph network corresponding to each meta-path, and each subgraph network reflects the connection pattern between nodes under the meta-path.
[0019] The constructed meta-path subgraph network only contains nodes whose node type under the meta-path belongs to the target type, that is, if there are common neighbors between the target type nodes in the meta-path network, there are connected edges in the corresponding meta-path subgraph network.
[0020] Step S3 assigns different weight values to adjacent nodes in the subgraph network based on the graph attention network (GAT) to reflect the different influences of nodes on their neighboring nodes: Extract nodes and their connection relationships in the generated subgraph network, and extract feature embeddings of target type nodes in the meta-path network, and use them as the initial feature vectors input to the graph attention network. Perform a linear transformation on the initial feature vectors of target type nodes to obtain a feature representation of uniform dimension. The linear transformation formula is as follows:
[0021] in Representation Node The initial eigenvector of , W represents the weight matrix, Indicated by The new feature vector mapped.
[0022] Defining an attention mechanism , which is a single-layer feedforward neural network composed of weight vector Initialize and use the LeakyReLU nonlinear function to calculate the attention coefficient of the node; calculate the attention coefficient, that is, the weight between each node and its neighbors, and the attention coefficient is calculated as follows:
[0023] in is the attention coefficient, indicating that the node For Node The importance of LeakyReLU is the activation function. , They are the nodes after linear transformation. and nodes The embedding represents the transpose of the weight vector, and || represents the concatenation operation.
[0024] Use masked attention mechanism to calculate only nodes With all neighboring nodes The attention coefficient between them is used to reduce the amount of calculation. The specific formula is as follows:
[0025] in Represents the node attention score, i.e., node For neighbor nodes The attention weight, exp represents the exponential function; The feature vectors of neighboring nodes are aggregated using the attention weights to obtain the updated node embedding representation of each node. The formula is as follows: ; in, Representation Node The updated node feature vector, σ represents the activation function.
[0026] The multi-head mechanism is introduced to learn the key information between nodes on a larger scale and more effectively capture the complex relationship between nodes in the graph. The formula is as follows:
[0027] in and are the attention weight and feature vector calculated by the kth attention mechanism, respectively.
[0028] In the result obtained by the multi-head mechanism, each node has K output features, and the K output features are aggregated by the average value so that the feature output of each node constitutes a feature matrix, that is, each row corresponds to the final feature representation of a node, and each subgraph corresponds to a feature matrix.
[0029] Step S4, using entropy to calculate the influence of nodes in the subgraph network, and aggregate the influence of nodes in all subgraphs to obtain the final influence value of each node. The specific steps are: First, the cosine similarity is used to calculate the similarity between each subgraph node. The formula is as follows: ; in Representation Node and nodes The cosine similarity between and Respectively represent nodes and nodes The yth component of the eigenvector, d is the dimension of the eigenvector; The similarity result is mapped to between 0 and 1, so that the final node influence value is non-negative. The formula is as follows: ; Calculate the entropy of all nodes in each subgraph using the following formula: ; in Represents the node in the mth subgraph The entropy of Representation Node The set of neighbors in the mth subgraph, Representation Node The number of neighbors in the mth subgraph; Finally, by aggregating the entropy of all nodes in each subgraph, the influence value of each node is finally obtained. The formula is as follows: .
[0030] The above shows and describes the basic principles and specific implementation process of the present invention. The present invention can evaluate the relative importance of different meta-paths, and also fully considers the different influences of node neighbors on the target node in each meta-path. It can comprehensively capture the complex relationships between nodes in heterogeneous networks. The present invention can establish richer semantic associations between different types of nodes and edges, and is suitable for key node identification in complex scenarios such as social network analysis, recommendation systems, and knowledge graphs. Based on the influence measurement and multi-path aggregation mechanism of entropy values, by aggregating the subgraph networks corresponding to different meta-paths, the information of multiple paths is integrated, making the influence evaluation more comprehensive and accurate, and overcoming the deviation that may be caused by a single meta-path. This makes the present invention more robust in complex network environments and can cope with diverse practical application needs.
Claims
1. A method for identifying influential nodes in heterogeneous graphs based on meta-path and graph attention mechanism, characterized in that: The following steps are involved: Step S1, define the required meta-path, and use the meta-path embedding method to obtain the node embedding representation; the specific steps of S1 are: Step S 11 , select the meta-path whose source node and terminal node types are the target type; Step S 12 , according to the defined meta-path, the random walk method is used to construct the heterogeneous neighborhoods of various types of nodes; the details are as follows: The heterogeneous network is represented as G = (V, E), where V and E are the sets of all nodes and edges in the heterogeneous network, respectively; Given a meta-path scheme P: , where each Represents a node type, each R i Represents a specific relationship type between node types; the transition probability between nodes is as follows: ; in, Represents the probability transfer formula between nodes, Representation Node The node type of is the neighbor node set of t+1, Representation Node The number of neighbor nodes of type t+1, Representation Node Type; Step S 13 , if the node and nodes There is an edge between nodes The type of is t+1, which means that it successfully follows the meta-path P for one step of random walk, and the transition probability is ; Step S 14 , if the node and nodes There is an edge between them, but the nodes If the type is not t+1, then the meta-path P is not followed for random walk, and the transition probability is 0; Step S 15 , if the node and nodes If there is no edge between them, the node does not follow the meta-path P for random walk, and the transition probability is 0; Step S 16 , the skip-gram model is used to obtain the feature embedding of the nodes in each meta-path; the skip-gram model learns the effective feature representation of nodes in heterogeneous networks by maximizing the probability of heterogeneous context node pairs; the maximum probability is defined as: ; Among them, arg max means finding the value of the model parameter θ that maximizes the following expression, represents a collection of context types, Representation Node a context node set of type t, is a given node and parameter θ, the context node The conditional probability of Representation Node The feature embedding of Represents the context node and nodes The dot product of the embedding vector of ; Step S 17 , extract the feature embedding of target type nodes in each meta-path; Step S2, generating a subgraph network corresponding to each meta-path according to random walk; Step S3, assigning different weight values to adjacent nodes in the subgraph network based on the graph attention network; Step S4, using entropy values to calculate the influence of nodes in the subgraph network, and aggregating the influence of nodes in all subgraphs to obtain the final influence value of each node.
2. The method for identifying influential nodes in heterogeneous graphs based on meta-path and graph attention mechanism according to claim 1 is characterized in that: The specific steps of S2 are: Step S 21 According to the predefined meta-path P, a random walk algorithm is used to start from the starting node and walk along the node types in the meta-path in turn to generate a node sequence ; Step S 22 ,For each random walk of each meta-path, ensure that the nodes walked conform to the node type order defined in the meta-path to preserve the structural and semantic relationships between nodes; Step S 23 ,The generated random walk sequence is used to construct the subgraph network corresponding to each meta-path, and each subgraph network reflects the connection pattern between nodes under the meta-path; Step S 24 , the constructed meta-path subgraph network only contains nodes whose node type under the meta-path belongs to the target type.
3. The method for identifying influential nodes in heterogeneous graphs based on meta-path and graph attention mechanism according to claim 1 is characterized in that: The specific steps of S3 are: Step S 31 , extract nodes and their connection relationships in the generated subgraph network, and extract feature embeddings of target type nodes in the meta-path network, which are used as the initial feature vectors input to the graph attention network; Step S 32 , perform linear transformation on the initial feature vector of the target type node, the linear transformation formula is as follows: ; Step S 33 , the attention coefficient between each node and its neighbors is calculated as follows: ; in, is the weight matrix, is the attention coefficient, indicating that the node For Node The importance of LeakyReLU is the activation function. , They are the nodes after linear transformation. and nodes The embedding represents the transpose of the weight vector, and || represents the concatenation operation; Step S 34 , using the masked attention mechanism to calculate the node With all neighboring nodes The attention coefficients between are as follows: ; in Represents the node attention score, i.e., node For neighbor nodes The attention weight, exp represents the exponential function; Step S 35 , use the attention weights to aggregate the feature vectors of neighboring nodes and obtain the updated node embedding representation of each node. The formula is as follows: ; in, Representation Node The updated node feature vector, represents the activation function, and are the attention weight and feature vector respectively; The multi-head mechanism is introduced to learn the key information between nodes on a larger scale and more effectively capture the complex relationship between nodes in the graph. The formula is as follows: ; in and are the attention weight and feature vector calculated by the kth attention mechanism respectively; In the result obtained by the multi-head mechanism, each node has K output features, and the K output features are aggregated by the average value so that the feature output of each node constitutes a feature matrix.
4. The method for identifying influential nodes in heterogeneous graphs based on meta-path and graph attention mechanism according to claim 1 is characterized in that: The specific steps of S4 are: The cosine similarity is used to calculate the similarity between each subgraph node. The formula is as follows: ; in Representation Node and nodes The cosine similarity between and Respectively represent nodes and nodes The yth component of the eigenvector, d is the dimension of the eigenvector; Map the similarity results to between 0 and 1, the formula is as follows: ; Calculate the entropy of all nodes in each subgraph using the following formula: ; in Represents the node in the mth subgraph The entropy of Representation Node The set of neighbors in the mth subgraph, Representation Node The number of neighbors in the mth subgraph; Aggregate the entropy of all nodes in each subgraph to get the influence value of each node. The formula is as follows: 。
Citation Information
Cited By
Heterogeneous network influence maximization method and system based on meta-path semantic perception
CN120543314A
Heterogeneous network influence maximization method and system based on meta-path semantic perception
CN120543314B
Heterogeneous cluster key node identification method and related equipment
CN121940295A