A Few-Shot Knowledge Graph Completion Method Based on Reinforcement Learning
By introducing reinforcement learning methods into knowledge graph completion, building action space and conducting path exploration, the problems of interpretability and data sparseness in knowledge graph completion are solved, and a more efficient knowledge graph completion and reasoning process is achieved.
Patent Information
- Application Number
- CN202311112645.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-08-31
AI Technical Summary
The existing knowledge graph completion technology has interpretability problems and data sparsity problems, making it difficult to provide reliable data support in special application fields such as smart medical care and military.
A small sample knowledge graph completion method based on reinforcement learning is adopted, and the action space is built and path exploration is carried out through the combination of decision network and meta-learning network, which improves the interpretability and data density of the completion process.
It effectively solves the problems of insufficient number of entities and insufficient prediction information in the inference process caused by knowledge graph sparsity, improves the interpretability of the completion method, and enhances users' trust in the model.
Smart Images

Figure CN117150041B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to knowledge graph reasoning technology, and particularly to a small-sample knowledge graph completion method based on reinforcement learning. Background Art
[0002] Data in a knowledge graph usually exists in the form of triples, specifically represented as (head entity, relation, tail entity). A large number of triple data constitute a huge network graph. Any thing and concept in nature can be represented as an entity node in the knowledge graph, and any connection can be represented as an edge in the knowledge graph. For example, if there is a triple (Socrates, born in, Athens) in the knowledge graph, the semantics of this triple represents "Socrates was born in Athens", where "Socrates" and "Athens" are two entity nodes, and "born in" is a relation, and they together constitute an edge in the graph structure.
[0003] Currently, knowledge graphs have become one of the main data sources in many scientific research and application fields. For example, many existing information retrieval, data mining, data analysis and other tasks in the field of natural language processing (NLP) require the data support of knowledge graphs. Related applications of knowledge graphs have also spread to all walks of life, such as social networks, intelligent conversations, personalized intelligent recommendations, intelligent Q&A, intelligent search, etc., and are further applied in vertical fields such as finance, social security, medical services, and public opinion response.
[0004] However, in practical applications, the knowledge graph data obtained automatically by models or manually is usually redundant and missing, and there may even be abnormal and incorrect knowledge, which will greatly reduce the usability of knowledge graph data in tasks, and even the quality of the data will directly affect the development and progress of related technologies. Some representation learning methods have been applied to knowledge graph completion. The representation learning method maps complex physical objects and corresponding relationship information into a low-dimensional continuous vector space, and measures the similarity between physical objects and relationships based on vectors, thereby achieving effective knowledge graph completion. Existing deep learning networks contain hundreds of thousands or even tens of millions of neurons, and these models can achieve high effects in terms of speed, stability, accuracy, etc. in the knowledge graph completion task.
[0005] However, since some special application fields such as intelligent healthcare, military, finance, and traffic safety are very sensitive to risk perception, any decision made by the inference model may be related to the life and property safety of users. If users do not have an intuitive understanding of the completion process and results, they cannot establish trust between users and machine models.
[0006] Secondly, since the sample data in the small-sample knowledge graph is sparse, that is, there is a lack of corresponding triple information for entities and relationships in the knowledge graph, or there is also a lack of sufficient path information for inference between the source entity and the target entity during the inference process. Due to the existence of this sparse relationship, the number of available entity pairs for knowledge graph inference is insufficient, resulting in an inability to obtain sufficient prediction information during the inference process, and the path will become unclear. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to propose a small-sample knowledge graph completion method based on reinforcement learning to solve the interpretability problem and the knowledge graph data sparsity problem existing in the existing knowledge graph completion technology, so as to better provide data support for its downstream tasks.
[0008] The technical solution adopted by the present invention to solve the above technical problems is:
[0009] A small-sample knowledge graph completion method based on reinforcement learning, comprising the following steps:
[0010] S1. Input the triple to be completed (h0, r,?), where g0 represents the head entity of the triple to be completed, r represents the target relationship, and? represents the tail entity to be completed; use the head entity h0 and the target relationship r as the initial input of the decision network at time step t0.
[0011] S2. Use the input head entity as the agent of the decision network at the current time step t; based on the input head entity and the target relationship, update the state S of the decision network at the current time step t t :
[0012] S t =(h t , r, E t )
[0013] where h t represents the head entity input at the current time step t, and E t represents the historical path representation of the agent, and its initial value is empty;
[0014] Based on the input head entity and the target relationship, construct the action space A of the decision network at the current time step t t The action space A t includes the following two parts:
[0015] The existing action space constructed based on the knowledge graph to be completed where represents the tail entity of the kth triple with the head entity h and the relationship r in the triple set t of the knowledge graph to be completed;
[0016] Using the meta-learning network of the knowledge graph to be completed, extract the meta-representation of the target relation r, and based on the head entity h t and the meta-representation of the target relation r, infer the additional action space where represents the tail entity of the l-th triple inferred by the meta-learning network, and r l represents the relation in the triple corresponding to in the knowledge graph to be completed;
[0017] S3. The policy function based on the decision network determines the action a t at the current time step t from the action space A t , and based on the tail entity e t obtained by taking this action a t , construct the inference result (h0, r t ′, e t ) at the current time step, where r t ′ represents the Hadamard product of the entity relations passed by the historical path and is expressed as represents the Hadamard product, and r t represents the relation in the triple corresponding to the tail entity e t in the knowledge graph to be completed;
[0018] S4. Based on the inference result (h0, r t ′, e t ) at the current time step, judge the similarity between r t ′ and r, and between e t and (h0 + r). If the preset similarity threshold is satisfied, the completion is completed; otherwise, based on the tail entity e t of the inference result at the current time step, construct the head entity h t+1 of the decision network at the next time step; update the historical path representation E t of the agent based on the action a t at the current time step t; then, return to step S2;
[0019] Before performing few-shot knowledge graph completion, the decision network and the meta-learning network are trained as follows:
[0020] B1. Use the knowledge graph to be completed to construct a support set and a query set, and train the meta-learning network based on the support set and the query set to obtain the meta-learning network of the knowledge graph to be completed that has completed training;
[0021] B2. Embed the completed meta - learning network into the policy network as the inference module for the additional action space in the policy network. Then, use the triples in the knowledge graph to be completed as training triples to train the policy network.
[0022] Further, the policy network is an LSTM network, and the policy function is:
[0023]
[0024] where σ represents the softmax operator; W1 and W2 are the parameters of the policy network respectively; denotes the product of tensors; π θ (a t ∣S t ) represents the probability distribution of each element in the action space A t at the current time step being selected, and determine the element with the maximum probability in the action space A t as the action a t at the current time step t.
[0025] Further, using the triples in the knowledge graph to be completed as training triples to train the policy network specifically includes:
[0026] B21. Randomly select a triple from the knowledge graph to be completed as a training triple, and specify the head entity and relationship of the training triple as the head entity h0 and the target relationship r, and specify its tail entity as the true - value tail entity e0;
[0027] B22. Obtain the inference result (h0, r t ′, e t ) at the current time step according to steps S2 - S3;
[0028] B23. Based on the triple (h0, r t ′, e t ), judge whether it is (h0, r, e0). If so, complete the completion, obtain the reward R b = 1, calculate the reward value R(a t ) at the current time step, and then execute step B24;
[0029] Otherwise, judge the similarity between (h0, r t ′, e t ) and (h0, r, e0). If the preset similarity threshold is satisfied, complete the completion, obtain the reward R b = 1, calculate the reward value R(a t ) at the current time step, and then execute step B24;
[0030] Otherwise, obtain the reward Rb = 0, calculate the reward value R(a t ) at the current time step; based on the tail entity e t of the inference result at the current time step, construct the head entity h t+1 of the decision network at the next time step; based on the action a t at the current time step t, update the historical path representation E t of the agent; then, return to step B22;
[0031] The reward value R(a t ) is calculated according to the following reward function:
[0032] R(α t ) = R b + (1 - R b )f(h o , r t ′, e t )
[0033]
[0034] where, represents the cosine similarity function;
[0035] B24. Calculate the training loss according to the following loss function, and update the parameters of the policy network according to the reverse gradient:
[0036]
[0037] where, T represents the total number of steps required to reach the final inference result, π θ (a t |S t ) represents the probability of taking action a t in state S t , and R(a t ) represents the reward value that can be obtained by taking this action;
[0038] The optimization process of the policy network parameters θ C is as follows:
[0039]
[0040] where, β is the preset learning rate parameter, represents the gradient of the loss;
[0041] B25. Judge whether the training converges or reaches the preset number of training rounds. If so, complete the training; otherwise, return to step B21.
[0042] Further, according to the following formula, based on r t ′ and r, and et The similarity with (h0+r) is judged to determine whether it meets the preset similarity threshold λ:
[0043]
[0044] Among them, represents the cosine similarity function.
[0045] Furthermore, using the meta-learning network of the knowledge graph to be completed, the meta-representation of the target relationship r is extracted, and based on the head entity h t and the meta-representation of the target relationship r, the additional action space is inferred including:
[0046] A1. Based on the relationship r, extract the support set from the knowledge graph to be completed, and the triples in the support set are represented as where i is the serial number of the triples in the support set;
[0047] A2. Construct the neighbor set of each head entity h i and the neighbor set of each tail entity and each tail entity of the neighbor set Initialize the feature vectors of each head entity h i , each tail entity and its neighbor set and neighbor set of each entity in it through the pre-trained model;
[0048] A3. Update the feature vectors of each head entity h i and tail entity by aggregating neighbor information;
[0049] A4. Based on the feature vectors of each updated head entity h i and tail entity construct the embedding vectors of the head and tail entity pairs of each triple in the support set
[0050] A5. Extract the meta-representation of the target relationship r based on the embedding vectors of the head and tail entity pairs of each triple in the support set ;
[0051] A6. Use the head entity h t and the meta-representation of the target relationship r to construct a query set based on the knowledge graph to be completed, screen the triples that meet the preset screening conditions in the query set as the expansion results, and construct an additional action space based on the expansion results
[0052] Furthermore, in step A3, by aggregating neighbor information, for each head entity h iTail entity The feature vector of is updated as:
[0053]
[0054]
[0055] where σ represents the sigmoid activation function, e n represents the initialized feature vector of the nth neighbor entity of the current computing entity e obtained by the pre-trained model, r n represents the initialized feature vector of the relationship between the current computing entity e and its neighbor entity e n obtained by the pre-trained model, α n is the attention weight of the nth neighbor entity, represents the concatenation operator, u rt , and b rt are learnable parameters, and T represents matrix transpose.
[0056] Furthermore, in step A4, based on the feature vectors of each head entity h i and tail entity , construct the embedding vectors of the head and tail entity pairs of each triple in the support set as:
[0057]
[0058] where, represents the concatenation operator, f θ (h i ) and represent the feature vectors of the updated head entity and tail entity of the ith triple in the support set, respectively.
[0059] Furthermore, in step A5, based on the embedding vectors of the head and tail entity pairs of each triple in the support set extract the meta-representation of the target relationship r, including:
[0060] A51. Input the embedding vectors of each triple head and tail entity pair into the recurrent neural network for encoding in sequence, and the number of layers of the recurrent neural network is the same as the number of triples in the support set;
[0061] A52. According to the following formula, update the hidden features output by the recurrent neural network by adding residual connections:
[0062]
[0063] where m iDenote the hidden feature of the $i$-th entity pair output by the recurrent neural network, $m$ i ' denotes the updated hidden feature of the $i$-th entity pair;
[0064] A53. Aggregate the hidden features of each entity pair based on attention according to the following formula to obtain the meta-representation $f$ of the target relationship $r$ ∈ (R r ) :
[0065]
[0066]
[0067] where $u$ R , and $b$ R are learnable parameters, and $T$ represents matrix transpose.
[0068] Furthermore, in step A6, use the head entity $h$ t and the meta-representation of the target relationship $r$ to construct a query set based on the knowledge graph to be completed, filter the triples in the query set that meet the preset filtering conditions as the expansion results, and construct an additional action space based on the expansion results including:
[0069] A61. Use the head entity $h$ t and the meta-representation of the target relationship $r$ to extract entity data based on the knowledge graph to be completed to construct a query set
[0070] A62. Use the meta-representation of the target relationship $r$ to obtain the similarity scores between the head and tail entity pairs of each triple in the query set and the relationship $r$ The similarity score function is expressed as:
[0071]
[0072] where $P$ r represents the normal vector of the hyperplane of the relationship $r$ in the knowledge graph, represents the $j$-th tail entity in the support set obtained based on the relationship $r$, and $f$ ∈ (R r ) is the meta-representation of the target relationship $r$;
[0073] A63. Based on the similarity scores, filter the triples in the query set that meet the preset filtering conditions as the expansion results, and construct an additional action space based on the expansion results
[0074] Furthermore, in step B1, use the knowledge graph to be completed to construct a support set and a query set, and train the meta-learning network based on the support set and the query set, including:
[0075] B11. Randomly specify a relationship as the training relationship, and extract the triples with the relationship as the training relationship from the knowledge graph to be completed as the training set;
[0076] B12. Divide the training set into two parts to form the support set S r and the positive query set By replacing the tail entity of the triples in the positive support set way to construct a negative query set The positive query set The triples in are expressed as The negative query set The triples in are expressed as
[0077] B13. Based on the support set S r , according to steps A2 - A5, obtain the meta - representation of the training relationship r;
[0078] B14. Based on the meta - representation of the training relationship r, calculate the matching scores of each triple in the positive query set and the negative query set respectively, and calculate the loss according to the following loss function:
[0079]
[0080]
[0081]
[0082] where, represents the matching score of the i - th triple in the positive query set , represents the matching score of the i - th triple in the negative query set , ξ is the safety margin distance, and γ is the discount factor;
[0083] Then, through the calculated loss l(S r ), update the hyperplane normal vector P r of the training relationship in the knowledge graph:
[0084]
[0085] where, l p represents the learning rate when the current model updates the hyperplane normal vector P r ;
[0086] B15. Determine whether the training converges or reaches the preset number of training rounds. If so, complete the training; otherwise, return to step B11.
[0087] The beneficial effects of the present invention are as follows:
[0088] (1) Capture knowledge from the knowledge graph through a small-sample method, match entity pairs associated with each relationship, thereby expanding the data in the knowledge graph during the reasoning process, forming a potential action space, and providing data support for the reasoning of the meta-learning network of the knowledge graph. It solves the problems such as insufficient number of available entity pairs for knowledge graph reasoning caused by the sparsity of the knowledge graph and inability to obtain sufficient prediction information during the reasoning process.
[0089] (2) Use reinforcement learning as the basic model for knowledge graph completion. First, determine the state of the agent at the current time step; secondly, obtain the actions that the agent can take. The actions consist of two parts, one is the original action space, and the other is the additional action space obtained based on meta-learning. Therefore, through the decision-making network based on reinforcement learning, path exploration is carried out in the knowledge graph, converting the completion process into a sequential decision-making problem and retaining the path information experienced by the agent, thereby improving the interpretability of the completion method and building a trust bridge between users and the model. Description of the Drawings
[0090] Figure 1 It is a schematic flowchart of the small-sample knowledge graph completion method based on reinforcement learning in the embodiment of the present invention;
[0091] Figure 2 It is a flowchart of constructing an additional action space in the embodiment of the present invention. Detailed Embodiments
[0092] The present invention aims to provide a small-sample knowledge graph completion method based on reinforcement learning to solve the interpretability problem and the knowledge graph data sparsity problem existing in the existing knowledge graph completion technology, so as to better provide data support for its downstream tasks. In the present invention, first, a support set and a query set are constructed using the knowledge graph to be completed, and the meta-learning network is trained based on the support set and the query set to obtain the meta-learning network of the knowledge graph to be completed after training; then the meta-learning network after training is embedded into the policy network as an inference module for the additional action space in the policy network, and then the triples in the knowledge graph to be completed are used as training triples to train the policy network. After obtaining the policy network integrated with the meta-learning network after training, it can be used to perform the knowledge graph completion task.
[0093] When performing the knowledge graph completion task, at the current time step, an action space is constructed based on the input head entity as an agent and the agent and the target relationship as the input of the decision network. Among them, the input head entity is constructed according to the tail entity inferred in the previous time step. The action space includes the existing action space in the knowledge graph that has the same head entity and relationship as it, and also includes an additional action space that does not exist in the original knowledge graph obtained by using the meta-learning network to extract features of the target relationship and reasoning based on the head entity and the target relationship features. Then, based on the policy function of the decision network, an inference result is constructed according to the tail entity corresponding to the action in the action space. Then, by calculating the similarity between the inference result and the query triple, it is judged whether the completed tail entity meets the requirements. If it does not meet the requirements, a new head entity is constructed again according to the inferred tail entity, and the new head entity and the target relationship are input into the decision network for continuous iteration.
[0094] Since the above solution of the present invention captures knowledge from the knowledge graph, matches entity pairs associated with each relationship, thereby expanding the data in the knowledge graph during the reasoning process, forming a potential action space, and providing data support for the reasoning of the meta-learning network of the knowledge graph, thus solving the problems of insufficient available entity pairs for knowledge graph reasoning and inability to obtain sufficient prediction information during the reasoning process due to the sparsity of the knowledge graph. Moreover, during the path exploration process of the agent, the path information experienced by the agent is retained and continuously updated as time steps progress, thereby improving the interpretability of the completion method.
[0095] Embodiment:
[0096] For a small sample knowledge graph where ε and respectively represent the sets of entities and relationships in the knowledge graph, is the set of triples in the knowledge graph.
[0097] This embodiment aims at a given small sample knowledge graph data and a source query triple (h0, r,?), where h0 is the head entity and r is the target relationship to be queried. By using the reinforcement learning strategy, the graph structure formed by the knowledge graph data is explored and the tail entity is predicted to complete the knowledge graph. The flow of the small sample knowledge graph completion method based on reinforcement learning provided by this embodiment is as Figure 1 shown, and it includes the following implementation steps:
[0098] S1. Input the query triple (h t , r,?) at the current time step t
[0099] In this step, if the current time step is the t0 moment, that is, querying the initial moment, the input query triple is (h0, r,?), if the current time step is not the t0 moment, the head entity h in the input query triple t is obtained by predicting the tail entity e at the previous time step t-1 and constructing it
[0100] S2. Update the state S of the decision network at the current time step t based on the head entity and the target relationship input at the current time step t t , and construct the action space A of the decision network at the current time step t t
[0101] In this step, use the head entity h input at the current time step t t as the agent of the decision network at the current time step t; update the state S of the decision network at the current time step t based on the input head entity and the target relationship t :
[0102] S t =(h t , r, E t )
[0103] where h t represents the head entity input at the current time step t, and E t represents the historical path representation of the agent. The action a taken at each moment t is the path passed at time t. For example: when t = 0, E t is empty; when t = 1, the action a1 taken will be saved to E t ; when t = 2, the action a2 taken will be saved to E t ... and so on, so as to retain the historical path of the agent to provide interpretability for the completion process
[0104] The action space A of the decision network at the current time step t t consists of two parts, expressed as: where is the existing action space constructed based on the knowledge graph to be completed where represents the tail entity of the k-th triple in the triple set T of the knowledge graph to be completed, where the head entity is h t and the relationship is r
[0105] is the additional action space that does not exist in the original knowledge graph obtained by reasoning based on the head entity and the target relationship features at the current time step where Denote the tail entity of the l-th triple obtained by the meta-learning network inference, r l Denote in the knowledge graph to be completed The relationship in the corresponding triple
[0106] Infer the above additional action space The process is as follows Figure 2 Shown as follows, specifically including
[0107] A1. Extract the support set from the knowledge graph to be completed based on the target relationship of the query
[0108] In this step, based on the target relationship r of the query, extract the support set from the knowledge graph to be completed. The triples in the support set are expressed as where i is the serial number of the triple in the support set
[0109] A2. Initialize the feature vectors of the head, tail entities and neighbor entities in the support set
[0110] In this step, construct the neighbor set of each head entity h i and the neighbor set of each tail entity and each tail entity and the neighbor set of each tail entity Through a pre-trained model such as the TransE model, initialize each head entity h i 、each tail entity and its neighbor set and the neighbor set The feature vectors of each entity in
[0111] A3. Aggregate neighbor information and update the feature vectors of the head and tail entities
[0112] In this step, by aggregating neighbor information, update the feature vectors of each head entity h i and the tail entity The calculation formula is
[0113]
[0114]
[0115] where σ represents the sigmoid activation function, e n represents the initialized feature vector of the n-th neighbor entity of the current calculated entity e obtained by the pre-trained model, r n represents the initialized feature vector of the relationship between the current calculated entity e and its neighbor entity e n obtained by the pre-trained model, α n is the attention weight of the n-th neighbor entity, represents the connection operator, u rt 、 and b rt are learnable parameters, and T represents matrix transpose.
[0116] A4. Based on the feature vectors of the updated head and tail entities, construct the embedding vectors of the head and tail entity pairs in each triple in the support set
[0117] In this step, based on each updated head entity h i and tail entity feature vectors, construct the embedding vectors of the head and tail entity pairs in each triple in the support set It is expressed as:
[0118]
[0119] Among them, represents the concatenation operator, and f θ (h i ) and respectively represent the feature vectors of the updated head entity and tail entity of the i-th triple in the support set.
[0120] A5. Extract the meta-representation of the target relationship based on the embedding vectors of the head and tail entity pairs in each triple in the support set
[0121] In this step, based on the embedding vectors of the head and tail entity pairs in each triple in the support set extract the meta-representation of the target relationship r, which specifically includes:
[0122] A51. Input the embedding vectors of each head and tail entity pair into the recurrent neural network for encoding in sequence. The number of layers of the recurrent neural network is the same as the number of triples in the support set; the recurrent neural network, that is, the RNN network or the one obtained by improving on the basis of RNN. In this embodiment, the recurrent neural network uses the RNN network.
[0123] A52. According to the following formula, update the hidden features output by the recurrent neural network by adding residual connections:
[0124]
[0125] Among them, m i represents the hidden feature of the i-th entity pair output by the recurrent neural network, and m i ′ represents the updated hidden feature of the i-th entity pair;
[0126] A53. According to the following formula, aggregate the hidden features of each entity pair based on attention to obtain the meta-representation f ∈ (R r ):
[0127]
[0128]
[0129] Among them, u R , and b R are learnable parameters, and T represents matrix transpose.
[0130] A6. Based on the meta-representation of the head entity and the target relationship, construct a query set based on the knowledge graph to be completed, filter out the triples that meet the conditions, and construct an additional action space
[0131] In this step, use the meta-representation of the head entity h t and the target relationship r to construct a query set based on the knowledge graph to be completed, filter out the triples in the query set that meet the preset filtering conditions as the expansion results, and construct an additional action space based on the expansion results Specifically, it includes the following steps:
[0132] A61. Use the meta-representation of the head entity h t and the target relationship r to extract entity data based on the knowledge graph to be completed to construct a query set For the above-mentioned extraction of entity data, the range of entities to be extracted can be determined according to the entity type corresponding to the relationship type. For example, if the relationship belongs to the interpersonal relationship, the entity type is the entity type representing people such as name and position. In this embodiment, all the tail entity data of the knowledge graph to be completed is extracted to construct a query set.
[0133] A62. Use the meta-representation of the target relationship r to obtain the similarity scores between the head and tail entity pairs of each triple in the query set and the relationship r The similarity score function is expressed as:
[0134]
[0135] Among them, P r represents the normal vector of the hyperplane of the relationship r in the knowledge graph to be completed, represents the j-th tail entity in the support set obtained based on the relationship r, and f ∈ (R r ) is the meta-representation of the target relationship r;
[0136] A63. Based on the similarity scores, filter out the triples in the query set that meet the preset filtering conditions as the expansion results, and construct an additional action space based on the expansion results Among them, represents the tail entity of the l-th triple obtained by the inference of the meta-learning network, and r l represents in the knowledge graph to be completed The relationship in the corresponding triple.
[0137] Through the construction of the above-mentioned additional action space in this embodiment, on the basis of the existing action space in the knowledge graph, the action space information can be further enriched, so that more prediction information can be obtained to solve the sparsity problem in the reasoning process and improve the accuracy of reasoning.
[0138] S3. The policy function based on the decision network constructs the reasoning result according to the tail entity corresponding to the action in the action space
[0139] In this step, the policy function based on the decision network determines the action a at the current time step t from the action space A t and constructs the reasoning result (h0, r t , based on the tail entity e obtained by taking this action a t ′, e t t t t t ) at the current time step, where r t ′ represents the Hadamard product of the entity relationships passed by the historical path and is expressed as represents the Hadamard product, and r θ represents the relationship in the triple corresponding to the tail entity e in the knowledge graph to be completed t .
[0140] As an exemplary choice, in this embodiment, the policy network can adopt an LSTM network, and the policy function is:
[0141]
[0142] where σ represents the softmax operator; W1 and W2 are respectively the parameters of the policy network; represents the product of tensors; π t (a t ∣S t ) represents the probability distribution of each element in the action space A at the current time step being selected, and the element with the largest probability in the action space A t is determined as the action a at the current time step t t . t .
[0143] S4. Determine whether the completed tail entity meets the requirements by calculating the similarity between the reasoning result and the source query triple
[0144] In this step, based on the reasoning result (h0, r t ′, e t), determine the similarity between this inference result and the source query triple (h0, r,?). When determining the similarity, use the method of judging the similarity between r t ' and r, and between e t and (h0 + r). It is expressed by the formula:
[0145]
[0146] where represents the cosine similarity function, and λ is the similarity threshold.
[0147] That is, by judging whether the above formula holds, it is determined whether the preset similarity threshold is satisfied.
[0148] If the preset similarity threshold is satisfied, the completion result is obtained, and the tail entity e in the inference result (h0, r t ', e t ) t is the tail entity for completing the source query triple (h0, r,?).
[0149] If the similarity threshold cannot be satisfied, further search is required. That is, based on the tail entity e of the inference result at the current time step t , construct the head entity h at the next time step of the decision network t+1 ; update the historical path representation E of the agent based on the action a at the current time step t t ; then, return to step S1. t
[0150] In this embodiment, a training method for the above decision network and meta-learning network is also provided, specifically including:
[0151] B1. Use the knowledge graph to be completed to construct a support set and a query set, and train the meta-learning network based on the support set and the query set to obtain the meta-learning network of the knowledge graph to be completed that has completed training;
[0152] Specifically, it includes the following steps:
[0153] B11. Randomly specify a relationship as the training relationship, and extract the triples with the relationship as the training relationship from the knowledge graph to be completed as the training set;
[0154] B12. Divide the training set into two parts to form the support set S r and the positive query set By replacing the tail entity of the triples in the positive support set , construct the negative query set The triples in the positive query set are expressed as The triples in the negative query set The middle triple is represented as
[0155] B13. Based on the support set S r , obtain the meta-representation of the training relation r according to steps A2 to A5;
[0156] B14. Based on the meta-representation of the training relation r, calculate the matching scores of each triple in the positive query set and the negative query set respectively, and calculate the loss according to the following loss function:
[0157]
[0158]
[0159]
[0160] where represents the matching score of the i-th triple in the positive query set , represents the matching score of the i-th triple in the negative query set , ξ is the safety margin distance, and γ is the discount factor;
[0161] Then, through the calculated loss L(S r ), update the hyperplane normal vector P r of the training relation in the knowledge graph:
[0162]
[0163] where l p represents the learning rate when the current model updates the hyperplane normal vector P r on the support set data;
[0164] B15. Determine whether the training converges or reaches the preset number of training rounds. If so, complete the training; otherwise, return to step B11.
[0165] B2. Embed the trained meta-learning network into the policy network as an inference module for the additional action space in the policy network. Then, use the triples in the knowledge graph to be completed as training triples to train the policy network, specifically including:
[0166] B21. Randomly select a triple from the knowledge graph to be completed as a training triple, and specify the head entity and relation of the training triple as the head entity h0 and the target relation r, and specify its tail entity as the true tail entity e0;
[0167] B22. According to steps S2 - S3, obtain the inference result (h0, r t ′, e t ) at the current time step;
[0168] B23. Based on the triple (h0, r t ′, e t ), determine whether it is (h0, r, e0). If so, complete the complementation, obtain the reward R b = 1, calculate the reward value R(a t ) at the current time step, and then execute step B24;
[0169] Otherwise, judge the similarity between (h0, r t ′, e t ) and (h0, r, e0). If the preset similarity threshold is satisfied, complete the complementation, obtain the reward R b = 1, calculate the reward value R(a t ) at the current time step, and then execute step B24;
[0170] Otherwise, obtain the reward R b = 0, calculate the reward value R(a t ) at the current time step; based on the tail entity e t of the inference result at the current time step, construct the head entity h t+1 of the decision network at the next time step; based on the action a t at the current time step t, update the historical path representation E t of the agent; then, return to step B22;
[0171] The reward value R(a t ) is calculated according to the following reward function:
[0172] R(a t ) = R b + (1 - R b )f(h o , r t ′, e t )
[0173]
[0174] Among them, represents the cosine similarity function;
[0175] B24. Calculate the training loss according to the following loss function, and update the parameters of the policy network according to the reverse gradient:
[0176]
[0177] where T represents the total number of steps required to reach the final inference result, π θ (a t ∣S t ) represents the probability of taking action a t in state S t , and R(a t ) represents the reward value that can be obtained by taking this action;
[0178] The optimization process of the policy network parameters θ C is as follows:
[0179]
[0180] where β is a preset learning rate parameter, represents the gradient of the loss;
[0181] B25. Judge whether the training converges or reaches the preset number of training epochs. If so, complete the training; otherwise, return to step B21.
[0182] Finally, it should be noted that the above embodiments are only preferred embodiments and are not intended to limit the present invention. It should be pointed out that for those of ordinary skill in the art in the technical field, without departing from the spirit and scope of the present invention as protected by the claims, several modifications, equivalent replacements, improvements, etc. should all be included within the protection scope of the present invention.
Claims
1. A few-shot knowledge graph completion method based on reinforcement learning, characterized in that The method includes: S1. Input the triple to be completed (h0, r,?), where h0 represents the head entity of the triple to be completed, r represents the target relation, and? represents the tail entity to be completed; use the head entity h0 and the target relation r as the initial input at time t0 of the decision-making network. S2. Use the input head entity as the agent of the decision-making network at the current time step t; update the state S of the decision-making network at the current time step t based on the input head entity and the target relationship. t : S t = (h t , r, E t ) Among them, h t represents the head entity input at the current time step t, and E t represents the historical path representation of the agent, and its initial value is empty; Construct the action space \(A\) of the decision-making network at the current time step \(t\) based on the input head entity and target relationship t , where the action space \(A\) t includes the following two parts: The existing action space constructed based on the knowledge graph to be completed Among them, represents the tail entity of the k-th triple with the head entity h in the triple set of the knowledge graph to be completed t and the relationship r; Using the meta-learning network of the knowledge graph to be completed, extract the meta-representation of the target relation r, and based on the head entity h t and the meta-representation of the target relation r, infer the additional action space where represents the tail entity of the l-th triple inferred by the meta-learning network, and r l represents in the knowledge graph to be completed the relation in the corresponding triple; S3. The policy function based on the decision network determines the action a at the current time step t from the action space A t and constructs the inference result (h0, r t ′, e t ) at the current time step based on the tail entity e obtained by taking the action a t , where r′ t represents the Hadamard product of the entity relationships passed by the historical path and is expressed as t denotes the Hadamard product, and r t represents the relationship in the triple corresponding to the tail entity e in the knowledge graph to be completed t t ; S4. Based on the inference result (h0, r t ′, e t ) at the current time step, judge the similarity between r t ′ and r, and between e t and (h0 + r). If the preset similarity threshold is satisfied, then the completion is finished; otherwise, based on the tail entity e t of the inference result at the current time step, construct the head entity h t+1 of the decision network at the next time step; update the historical path representation E t of the agent based on the action a t at the current time step t; then, return to step S2; Before performing few-shot knowledge graph completion, the decision-making network and the meta-learning network are trained as follows: B1. Use the knowledge graph to be completed to construct a support set and a query set, and train the meta-learning network based on the support set and the query set to obtain the meta-learning network of the knowledge graph to be completed that has completed training. B2. Embed the meta-learning network that has completed training into the policy network as an inference module in the additional action space of the policy network. Then, use the triples in the knowledge graph to be completed as training triples to train the policy network.
2. The small-sample knowledge graph completion method based on reinforcement learning according to claim 1, characterized in that The policy network is an LSTM network, and the policy function is: Among them, σ represents the softmax operator; W1 and W2 are the parameters of the policy network respectively; represents the product of tensors; π θ (a t |S t ) represents the probability distribution of each element in the action space A t at the current time step, and determines the element with the highest probability in the action space A t as the action a at the current time step t t .
3. A few-shot knowledge graph completion method based on reinforcement learning according to claim 1 or 2, characterized in that Using the triples in the knowledge graph to be completed as training triples to train the policy network includes: B21. Randomly select a triple from the knowledge graph to be completed as a training triple, and specify the head entity and relation of the training triple as the head entity h0 and the target relation r, and specify its tail entity as the true value e0 of the tail entity. B22. According to steps S2 to S3, obtain the inference result (h0, r t ′, e t ) at the current time step; B23. Based on the triple (h0, r t ′, e t ), determine whether it is (h0, r, e0). If so, complete the complementation and obtain the reward R b = 1, calculate the reward value R(a t ), and then execute step B24; Otherwise, judge (h0, r t ′, e t ) and the similarity between (h0, r, e0). If the preset similarity threshold is satisfied, the completion is completed and the reward R b = 1, calculate the reward value R(a t ) at the current time step, and then execute step B24; Otherwise, obtain the reward R b = 0, calculate the reward value R(a t ); based on the tail entity e t of the inference result at the current time step, construct the head entity h t+1 of the decision network at the next time step; based on the action a t at the current time step t, update the historical path representation E t of the agent; then, return to step B22; The reward value R(a t ) is calculated according to the following reward function: R(a t ) = R b + (1 - R b ) f(h o , r t ', e t ) Among them, represents the cosine similarity function; B24. Calculate the training loss according to the following loss function, and update the parameters of the policy network according to the reverse gradient. Among them, T represents the total number of steps required to reach the final inference result, and π θ (a t |S t ) represents the probability of taking action a t in state S t , and R(a t ) represents the reward value that can be obtained by taking this action; Policy network parameters θ C The optimization process is as follows: where β is a preset learning rate parameter, represents the gradient of the loss; B25. Determine whether the training converges or reaches the preset number of training rounds. If so, the training is completed; otherwise, return to step B21.
4. A few-shot knowledge graph completion method based on reinforcement learning according to claim 1 or 2, characterized in that Based on r according to the following formula t ' and r, as well as e t and (h0 + r), determine whether it meets the preset similarity threshold λ: Among them, represents the cosine similarity function.
5. A small-sample knowledge graph completion method based on reinforcement learning according to claim 1 or 2, characterized in that Using the meta-learning network of the knowledge graph to be completed, extract the meta-representation of the target relationship r, and based on the head entity h t and the meta-representation of the target relationship r, infer the additional action space including: A1. Extract a support set from the knowledge graph to be completed based on the relationship r, and the triples in the support set are represented as where i is the serial number of the triples in the support set. A2. Construct the neighbor sets of each head entity h i and the neighbor sets of each tail entity through the pre-trained model, initialize the feature vectors of each head entity h and the neighbor sets of each tail entity Initialize the feature vectors of each head entity h i , each tail entity and its neighbor set and the neighbor set for each entity in them; A3. Update the feature vectors of each head entity h i and tail entity by aggregating neighbor information; A4. Based on the feature vectors of each updated head entity h i and the tail entity construct the embedding vectors of the head and tail entity pairs of each triple in the support set A5. Based on the embedding vectors of the head and tail entity pairs in the support set Extract the meta-representation of the target relation r; A6. Utilize the meta-representation of the head entity h t and the target relation r to construct a query set based on the knowledge graph to be completed, screen out the triples that meet the preset screening conditions in the query set as the expansion results, and construct an additional action space based on the expansion results 6. The small-sample knowledge graph completion method based on reinforcement learning according to claim 5, wherein Step A3, update the feature vectors of each head entity h i and tail entity by aggregating neighbor information, as follows: Among them, σ represents the sigmoid activation function, e n represents the initialized feature vector obtained by the pre-trained model for the nth neighbor entity of the current computing entity e, r n represents the initialized feature vector obtained by the pre-trained model for the relationship between the current computing entity e and its neighbor entity e n The relationship between them, α n is the attention weight of the nth neighbor entity, represents the concatenation operator, u rt 、 and b rt are learnable parameters, and T represents matrix transpose.
7. The few-shot knowledge graph completion method based on reinforcement learning according to claim 5, characterized in that Step A4, based on the updated feature vectors of each head entity h i and tail entity to construct the embedding vectors of the head and tail entity pairs of each triple in the support set as follows: Among them, represents a connection operator, and f θ (h i ) and respectively represent the feature vectors of the updated head entity and tail entity of the i-th triple in the support set.
8. The small-sample knowledge graph completion method based on reinforcement learning according to claim 5, characterized in that Step A5, based on the embedding vectors of the head and tail entity pairs of each triple in the support set Extract the meta-representation of the target relation r, including: A51. Input the embedding vectors of the head and tail entity pairs of each triple into a recurrent neural network for encoding in sequence. The number of layers of the recurrent neural network is the same as the number of triples in the support set; A52. Update the hidden features output by the recurrent neural network by adding a residual connection according to the following formula: where m i represents the hidden feature of the i-th entity pair output by the recurrent neural network, and m' i represents the updated hidden feature of the i-th entity pair; A53. Aggregate the hidden features of each entity pair based on attention according to the following formula to obtain the meta-representation \(f\) of the target relationship \(r\). ∈ (R r ): where, u R , and b R are learnable parameters, and T represents matrix transpose.
9. The small-sample knowledge graph completion method based on reinforcement learning according to claim 5, characterized in that Step A6, using the meta-representation of the head entity h t and the target relation r, construct a query set based on the knowledge graph to be completed, filter the triples in the query set that meet the preset filtering conditions as the expansion results, and construct an additional action space based on the expansion results including: A61. Utilize the meta-representation of the head entity h t and the target relation r to extract entity data based on the knowledge graph to be completed and construct a query set A62. Obtain the similarity score between the head and tail entity pairs of each triple in the query set and the relationship r by using the meta-representation of the target relationship r The similarity score function is expressed as: Among them, P r represents the normal vector of the hyperplane of the relationship r in the knowledge graph to be completed, represents the j-th tail entity in the support set obtained based on the relationship r, f ∈ (R r ) is the meta-representation of the target relationship r; A63. Based on the similarity score, filter out the triples in the query set that meet the preset filtering conditions as the expansion results, and construct an additional action space based on the expansion results 10. A small-sample knowledge graph completion method based on reinforcement learning according to claim 9, characterized in that, In step B1, using the knowledge graph to be completed to construct a support set and a query set, and training the meta-learning network based on the support set and the query set includes: B11. Randomly specify a relation as the training relation, and extract the triples with the relation as the training relation from the knowledge graph to be completed as the training set. B12. Divide the training set into two parts to form the support set S r and the positive query set By replacing the tail entity of the triples in the positive support set to construct the negative query set The triples in the positive query set are represented as The triples in the negative query set are represented as B13. Based on the support set S r , obtain the meta-representation of the training relationship r according to steps A2 to A5; B14. Calculate the matching scores of each triple in the positive query set and the negative query set respectively based on the meta-representation of the training relationship r, and calculate the loss according to the following loss function: Among them, represents the matching score of the i-th triple in the positive query set, represents the matching score of the i-th triple in the negative query set, ξ is the safety margin distance, and γ is the discount factor; Then, based on the calculated loss L(S r ), update the hyperplane normal vector P r of the training relationships in the knowledge graph: Among them, l p represents the learning rate when the current model updates the hyperplane normal vector P r on the support set data; B15. Determine whether the training converges or reaches the preset number of training rounds. If so, the training is completed; otherwise, return to step B11.
Citation Information
Patent Citations
Few-sample knowledge graph completion method based on meta-learning
CN113239131A
Knowledge graph reasoning method based on logic rules and reinforcement learning
CN115660086A