A knowledge graph node recommendation method based on the fusion of graph embedding and knowledge representation learning

By integrating the methods of graph embedding and knowledge representation learning in the knowledge graph, node vectors are generated and similar node recommendations are performed, the problem of traditional recommendation algorithms being poor in the knowledge graph is solved, and efficient and accurate similar node recommendations are achieved.

CN119807542BActive Publication Date: 2025-05-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510309119.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-13
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

Traditional recommendation algorithms perform poorly in the knowledge graph, especially ignoring the semantic information of nodes, resulting in the inability to provide effective recommendation services in cold startup situations.

Method used

Using a method based on graph embedding and knowledge representation learning fusion, the nodes are converted into vector representations through graph embedding encoder and knowledge representation learning encoder, fusion structure information and semantic information are integrated, and trained through a neural network to generate the final node vector for recommendations of similar nodes.

Benefits of technology

It realizes adaptive recommendations for different types of nodes in the knowledge graph without relying on user interaction data, avoiding cold start problems, and improving the accuracy and efficiency of recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807542B_ABST
    Figure CN119807542B_ABST
Patent Text Reader

Abstract

The present invention provides a knowledge graph node recommendation method based on the fusion of graph embedding and knowledge representation learning, belonging to the field of knowledge graph and its application. By combining the advantages of graph embedding and knowledge representation learning, the method can effectively fuse the structural information and semantic information of the node, thereby improving the accuracy and effect of the recommendation system. The present invention does not rely on the user's interactive data, but makes adaptive recommendations based on the structural relationship and semantic features of the nodes in the knowledge graph, thereby avoiding the cold start problem and fully mining the potential relationship and semantic information between nodes. The method can flexibly handle different types of nodes, adapt to a variety of complex application scenarios, and provide efficient and accurate similar node recommendations in large-scale knowledge graphs. Through the combination of graph embedding and knowledge representation learning, the present invention significantly improves the accuracy of the recommendation results, and shows obvious advantages in the lack of node semantic information and the use of structured knowledge in knowledge graph applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of knowledge graphs and their applications, and specifically relates to a knowledge graph node recommendation method based on the fusion of graph embedding and knowledge representation learning. Background Art

[0002] The core goal of the recommendation system is to help users filter out valuable data that matches their personal interests from massive amounts of information. This system collects records of interactions between users and data, learns users’ preferences for data, and provides recommendation services in a personalized manner, thereby significantly improving the efficiency and quality of information acquisition.

[0003] Traditional recommendation algorithms mainly include three types: content-based recommendation, collaborative filtering-based recommendation, and hybrid recommendation algorithms. These algorithms combine recommendation techniques with different data structures or apply them to different scenarios, which will produce different results. Although various recommendation algorithms have achieved good results in specific scenarios, they still face some challenges:

[0004] Over-reliance on interaction records between users and data: When new users or new data join the system, due to the lack of interaction records, the recommendation algorithm finds it difficult to learn the user's preferences for data, resulting in the system being unable to provide effective recommendation services. This is the so-called cold start problem.

[0005] Traditional node recommendation algorithms are not adaptable enough to knowledge graphs: Most traditional node recommendation algorithms are designed for ordinary graph structures, but with the development and application of knowledge graph technology, this data structure has been widely used in multiple industries. Knowledge graphs represent entities and concepts and their relationships through nodes and edges, effectively organizing and utilizing a large amount of knowledge. However, traditional node recommendation algorithms do not perform well on knowledge graphs, and usually ignore the semantic information of nodes in knowledge graphs. Therefore, it is necessary to develop an adaptive recommendation method that specifically targets different types of nodes in knowledge graphs so that similar nodes can be recommended without relying on user data interaction. Summary of the invention

[0006] In response to the above-mentioned technical problems and defects in the prior art, the present invention proposes a knowledge graph node recommendation method based on the fusion of graph embedding and knowledge representation learning. This method effectively combines the semantic information and structural information of nodes in the knowledge graph, and performs adaptive recommendation for different types of nodes to address the limitations of traditional recommendation algorithms in the application of knowledge graphs, and can effectively solve the above-mentioned problems.

[0007] The technical solution adopted by the present invention is as follows:

[0008] A knowledge graph node recommendation method based on the fusion of graph embedding and knowledge representation learning, comprising the following steps:

[0009] S1: Extract nodes and edges from the knowledge graph. Each node represents an entity or concept, and each edge represents the relationship between nodes, which serves as the basis for node recommendation.

[0010] S2: Use the graph embedding encoder to convert the nodes into a set of graph embedding vectors, reflecting the structural relationship and adjacency information of the nodes in the knowledge graph;

[0011] S3: Use the knowledge representation learning encoder to represent the node as a set of knowledge representation vectors. The node vector contains the node's semantic information, attribute features, and mapping relationships with other nodes.

[0012] S4: The graph embedding vector and the knowledge representation vector are fused through the encoder fuser, and the fused node representation vector is input into the neural network of the node classification model to train and predict the node label of each node;

[0013] S5: Calculate the error between the predicted label and the actual label through the loss function, and update the parameters of the node classification model to optimize the model performance;

[0014] S6: After the model training is completed, for each node, its graph embedding vector and knowledge representation vector are input into the node classification model, and the output of the middle hidden layer of the neural network is obtained as the final node vector of the node;

[0015] S7: Perform embedding recall based on the similarity of the final node vector, screen out the N candidate nodes with the highest similarity, generate a recommendation result list, and recommend the list to the user to achieve node recommendation of the knowledge graph;

[0016] S71: recall the embeddings based on the similarity of the node vectors and recommend the list to the user;

[0017] S72: For each target node, sort the nodes according to the similarity from high to low, and select the N nodes with the highest similarity as candidate nodes;

[0018] S73: Sort the screened N candidate nodes by similarity to generate a recommendation result list.

[0019] The graph embedding encoder is specifically a node2vec encoder, which converts nodes into low-dimensional graph embedding vectors.

[0020] The knowledge representation learning encoder is specifically a TransE model.

[0021] The encoding fuser specifically fuses the graph embedding vector and the knowledge representation vector through a vector connection operation; the node classification model is specifically a three-layer feedforward neural network, including an input layer, a hidden layer, and an output layer, which is specifically expressed as follows:

[0022]

[0023]

[0024] in, represents the output of the hidden layer, represents the activation function, and is the weight and bias of the neural network, i is the index of the network layer, for example, when i=1, and Acting on the input layer to the hidden layer, when i=2, and Acts on the hidden layer to the output layer; the softmax function is used to convert the value of the output layer into a probability distribution. is the predicted node label probability distribution.

[0025] The loss function is specifically a cross entropy loss function.

[0026] The similarity calculation method described in step S7 includes but is not limited to cosine similarity, Euclidean distance, Manhattan distance, etc., which are used to evaluate the similarity between node vectors, and generate a recommendation result list by selecting the N candidate nodes with the highest similarity, and recommend the list to the user to achieve efficient and accurate similar node recommendation.

[0027] The present invention provides a knowledge graph node recommendation method based on the fusion of graph embedding and knowledge representation learning. By combining the advantages of graph embedding and knowledge representation learning, the method can effectively fuse the structural information and semantic information of the node, thereby improving the accuracy and effect of the recommendation system. Unlike traditional recommendation methods, the present invention does not rely on user interaction data, but makes adaptive recommendations based on the structural relationship and semantic features of the nodes in the knowledge graph, thereby avoiding the cold start problem and fully mining the potential relationship and semantic information between nodes. The method can flexibly handle different types of nodes, adapt to a variety of complex application scenarios, and provide efficient and accurate similar node recommendations in large-scale knowledge graphs. Through the combination of graph embedding and knowledge representation learning, the present invention significantly improves the accuracy of recommendation results, and shows obvious advantages in the lack of node semantic information and the use of structured knowledge in knowledge graph applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1It is a training flow chart of a knowledge graph node recommendation method based on the fusion of graph embedding and knowledge representation learning in the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and beneficial effects of the present invention more clearly understood, the specific implementation methods of the present invention will be further described in detail below in conjunction with specific embodiments.

[0030] Reference Figure 1 The present invention provides a knowledge graph node recommendation method based on graph embedding and knowledge representation learning fusion, comprising the following steps:

[0031] S1: Extract nodes and edges from the knowledge graph. Each node represents an entity or concept, and each edge represents the relationship between nodes, which serves as the basis for node recommendation.

[0032] Specifically, we first connected to the Neo4j graph database of high-temperature alloy materials and used the Cypher query language to extract and query all triple data in the knowledge graph. We extracted and constructed a knowledge graph of high-temperature alloy materials from 205 high-temperature alloy grade data, including 69 ontology network triples and their related structural features, including physical and chemical parameters, shear modulus, microstructure, preparation process, mismatch, performance parameters, elemental composition, creep temperature and other entity features. These nodes constitute the basic components of the knowledge graph, containing rich semantic information and structural relationships.

[0033] S2: Use the graph embedding encoder to convert the extracted nodes into a set of graph embedding vectors, reflecting the structural relationship and adjacency information of the nodes in the knowledge graph.

[0034] The graph embedding encoder is specifically a node2vec encoder, which converts the nodes in the high-temperature alloy knowledge graph into low-dimensional graph embedding vectors; specifically, assuming that the knowledge graph is an undirected graph ,in is a node set, is the edge set. For each node , generating node sequences through random walks ,in is the walk length. Taking GH4169 alloy as an example, the following node sequence may be generated through random walk: GH4169 -> Ni content -> 52.5% -> heat treatment process -> solution treatment. The node2vec encoder captures the local and global structural information of the node through the random walk method and generates a low-dimensional vector representation of each node; the node2vec encoder parameter settings include embedding dimension 128, random walk length 30, random walk number 10, backoff probability p is 0.5, and exploration probability q is 2;

[0035] Then use the Skip-gram model to maximize the given central node Predicting context nodes The log-likelihood of the conditional probability of :

[0036]

[0037] in, represents a central node in the graph, Representation Node Belongs to the node collection , Representation Node Is a node The context collection The conditional probability of the nodes in is defined as:

[0038]

[0039] Where exp represents the exponential function, represents all the nodes in the graph, Represents the inner product of the embedding vectors of node u and node v. Through the above method, the node2vec encoder generates a graph embedding vector for each node , this vector can reflect the structural relationship and adjacency information of the node in the knowledge graph.

[0040] S3: Use the knowledge representation learning encoder to represent the node as a set of knowledge representation vectors. The node vector contains the node's semantic information, attribute features, and mapping relationships with other nodes.

[0041] The knowledge representation learning encoder is specifically a TransE model; for each triple in the knowledge graph , the head entity and tail entity And the relationship Map them into low-dimensional vector space respectively, and get vector Taking the preparation process of high-temperature alloy as an example, typical triples include: (GH4169, containing elements, Ni), (solution treatment, previous process, thermal deformation), etc. The TransE model maps entities and relations to the same vector space, and assumes that the relation r converts the head entity h to the tail entity t, that is:

[0042]

[0043] TransE model parameter settings: vector dimension 128, learning rate 0.01, marginal loss 1.0, negative sampling rate 1;

[0044] Define the scoring function To measure the rationality of the triple:

[0045]

[0046] The optimization goal is to minimize the scoring function of reasonable triplets and maximize the scoring function of unreasonable triplets:

[0047]

[0048] in, is the marginal loss, positive sample is the triple existing in the knowledge graph, negative sample is a randomly generated unreasonable triple, represents the head entity in the negative sample, represents the relationship in the negative sample, Represents the tail entity in the negative sample. Through the TransE model, the knowledge representation vector of each node is generated , which contains the node’s semantic information, attribute features, and mapping relationships with other nodes.

[0049] S4: The graph embedding vector and the domain knowledge vector are fused through the encoder fuser, and the fused node representation vector is input into the neural network of the node classification model to train and predict the node label of each node.

[0050] Embedding graph into vector via Encoding Fusor and knowledge representation vector Fusion, get the fused node representation vector in, Represents a vector concatenation operation.

[0051] The node classification model can use a multi-layer feedforward neural network. This embodiment uses a three-layer feedforward neural network for node classification. The specific structure is as follows:

[0052] Input layer: The input dimension is 256 (128-dimensional graph embedding vector + 256-dimensional knowledge representation vector).

[0053] Hidden layer: 128 neurons, activation function is ReLU.

[0054] Output layer: Set according to specific tasks, such as the number of node label categories.

[0055]

[0056]

[0057] in, represents the output of the hidden layer, represents the activation function, and is the weight and bias of the neural network, i is the index of the network layer, for example, when i=1, and Acting on the input layer to the hidden layer, when i=2, and The softmax function from the hidden layer to the output layer is used to convert the value of the output layer into a probability distribution. is the predicted node label probability distribution.

[0058] S5: Calculate the error between the predicted label and the actual label through the loss function, and update the parameters of the node classification model to optimize the model performance.

[0059] Calculate the predicted label through the loss function With the actual label The error between them is used to update the parameters of the node classification model to optimize the model performance. The cross entropy loss function is defined as:

[0060]

[0061] in, is the number of categories, is the one-hot encoding of the actual label, To predict the probability, use a gradient descent algorithm (such as the Adam optimizer) to optimize the loss function Minimize and update the parameters of the neural network and ;

[0062] S6: After model training is completed, for each node, input its graph embedding vector and knowledge representation vector To the node classification model, get the output of the middle hidden layer of the neural network as the final node vector of the node.

[0063] After the model training is completed, for each high-temperature alloy knowledge graph node, its graph embedding vector is input and knowledge representation vector To the node classification model, obtain the output of the middle hidden layer of the neural network as the final node vector of the node. Specifically, select the output of the hidden layer As the final node vector ;

[0064] S7: Perform embedding recall based on the similarity of node vectors, filter out the N candidate nodes with the highest similarity, generate a recommendation result list, and recommend the list to the user to achieve node recommendation of the knowledge graph.

[0065] S71: Based on node vector The similarity of the target node is used to perform Embedding recall and recommend the list to the user. The cosine similarity is used to calculate the similarity between the target node and all candidate nodes:

[0066]

[0067] in, Represent the vector representation of two different nodes respectively, Represents the L2 norm, used for normalization.

[0068] S72: For each target node, sort the nodes from high to low according to the similarity, and select the N nodes with the highest similarity as candidate nodes.

[0069] S73: Sort the screened N candidate nodes by similarity to generate a recommendation result list.

[0070] For example, for GH4169 alloy: calculate the similarity with other alloys, select the N alloys with the highest similarity (such as N=5), and generate a recommendation list, which contains detailed information of similar alloys.

Claims

1. A knowledge graph node recommendation method based on the fusion of graph embedding and knowledge representation learning, characterized in that: The following steps are involved: S1: Extract nodes and edges from the knowledge graph, where each node represents an entity or concept, and each edge represents the relationship between nodes, which serves as the basis for node recommendation; specifically connect to the high-temperature alloy material graph database, extract and query all triple data in the knowledge graph, and construct a high-temperature alloy material knowledge graph, including the ontology network triples and their structural features, including physicochemical parameters, shear modulus, microstructure, preparation process, mismatch, performance parameters, elemental composition, and creep temperature entity features; S2: Use a graph embedding encoder to convert the node into a set of graph embedding vectors, reflecting the structural relationship and adjacency information of the node in the knowledge graph; the graph embedding encoder is specifically a node2vec encoder, which converts the node into a low-dimensional graph embedding vector; S3: using a knowledge representation learning encoder to represent a node as a set of knowledge representation vectors, where the node vector contains the node's semantic information, attribute features, and mapping relationships with other nodes; the knowledge representation learning encoder is specifically a TransE model; S4: The graph embedding vector and the knowledge representation vector are fused through the encoder fuser, and the fused node representation vector is input into the neural network of the node classification model to train and predict the node label of each node; S5: Calculate the error between the predicted label and the actual label through the loss function, and update the parameters of the node classification model to optimize the model performance; S6: After the model training is completed, for each high-temperature alloy knowledge graph node, its graph embedding vector and knowledge representation vector are input into the node classification model, and the output of the middle hidden layer of the node classification model is obtained as the final node vector of the node; the node classification model is specifically a multi-layer feedforward neural network; S7: Perform embedding recall based on the similarity of the final node vector, filter out the N candidate nodes with the highest similarity, generate a recommendation result list, and recommend the list to the user to achieve node recommendation of the knowledge graph.

2. According to claim 1, a knowledge graph node recommendation method based on graph embedding and knowledge representation learning fusion is characterized in that: The encoding fuser specifically fuses the graph embedding vector and the knowledge representation vector through a vector connection operation; the multi-layer feedforward neural network includes an input layer, a hidden layer, and an output layer, which are specifically expressed as follows: h (1) =σ(W (1) v+b (1) ); Among them, h (1) represents the output of the hidden layer, σ represents the activation function, W (i) and b (i) are the weights and biases of the neural network, i is the index of the network layer; the softmax function is used to convert the value of the output layer into a probability distribution, is the predicted node label probability distribution.

3. According to claim 2, a knowledge graph node recommendation method based on graph embedding and knowledge representation learning fusion is characterized in that: The loss function is specifically a cross entropy loss function.

4. According to claim 3, a knowledge graph node recommendation method based on graph embedding and knowledge representation learning fusion is characterized in that: The S7 is specifically as follows: S71: Perform Embedding recall based on the similarity of the final node vector and recommend the list to the user; S72: For each target node, sort them from high to low according to the similarity, and select the N nodes with the highest similarity as candidate nodes; S73: Sort the screened N candidate nodes by similarity to generate a recommendation result list.

5. According to claim 4, a knowledge graph node recommendation method based on graph embedding and knowledge representation learning fusion is characterized in that: In S71, cosine similarity is used to calculate the similarity between the target node and all candidate nodes: Among them, z i , z j They represent the vector representations of two different nodes respectively, and ||·|| represents the L2 norm.

Citation Information

Patent Citations

  • Multi-task recommendation algorithm fusing user behaviors and knowledge graph

    CN117370674A

  • Personalized article recommendation method based on knowledge graph and attention mechanism

    CN118820594A