A Knowledge Graph Conversational Question Answering Method Based on Reinforcement Learning for Problem Constraint and Focus Entity Transfer

The integration of reinforcement learning and BERT for knowledge graph-based conversational systems addresses entity focus shifts and semantic gaps, enhancing answer accuracy by dynamically maintaining context entity sets and utilizing 1-hop and 2-hop path constraints.

CN116414956BActive Publication Date: 2025-07-15DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310096306.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2025-07-15
Estimated Expiration
2043-02-10

AI Technical Summary

Technical Problem

The existing knowledge graph question and answer system is difficult to deal with subsequent incomplete problems, lacks the combination of various constraint information in the problem and dialogue focus transfer, and the existing conversational question and answer system based on reinforcement learning relies on feedback from non-expert users, resulting in strong training dependence.

Method used

Using reinforcement learning-based method, dynamically maintains the context entity set, uses BERT pre-trained model to obtain dialogue semantic information, combines problem rounds, entity rounds and various constraint information, and realizes training on the 1-hop path and 2-hop constraint path to solve the problem of focus entity transfer.

Benefits of technology

Improve the accuracy of the answer, accurately capture the dialogue context information, and comprehensively utilize the backbone information of the 1-hop path and the constraint information of the 2-hop path, improving the accuracy of the question-and-answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116414956B_ABST
    Figure CN116414956B_ABST
Patent Text Reader

Abstract

A knowledge graph conversational question-answering method based on reinforcement learning for problem constraint and focus entity transfer, belonging to the technical field of question-answering systems. It captures all possible focus entities involved during the conversation by dynamically maintaining a context entity set to solve the problem of focus entity transfer during the conversation; it uses the BERT pre-trained model to obtain the semantic information of the conversation context to solve the problem of semantic loss of subsequent questions during the conversation; the model combines the question round, entity round and various constraint information, and uses the reinforcement learning method to realize the training of the model on the 1-hop path and 2-hop constraint path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of question-and-answer systems, and adopts a knowledge graph conversational question-and-answer method combining reinforcement learning and deep learning. Background Art

[0002] As an important way for users to obtain information, the question-and-answer system is constantly developing with the progress of technology. Since the knowledge graph contains factual information in the form of triples that can be directly used for question and answer, question answering based on the knowledge graph has been deeply studied and continuous new progress has been made in solving complex problems. However, with the gradual increase in users' information needs, users often hope to obtain more information on a topic and thus ask subsequent questions on a topic. In such subsequent questions, problems such as entity omission, predicate omission, and ungrammaticality may occur. Therefore, one of the challenges in solving conversational question answering is to model the conversation history and at the same time use an appropriate method to capture the transfer of the focus entity during the conversation process to better answer the current question.

[0003] Existing knowledge graph question-and-answer systems all require the form of the question to be complete, and they basically cannot handle subsequent incomplete questions in conversational question answering. On this basis, relevant researchers have started the research on conversational question-and-answer systems. In existing knowledge graph-based conversational question-and-answer systems, the unsupervised method maintains the entire conversation history information by expanding the context subgraph to answer subsequent questions. This method relies on the similarity evaluation index, and the unsupervised method has certain limitations in improving accuracy. The end-to-end model based on semantic analysis uses a grammar-guided decoder to control the generation of the action sequence and uses a conversation memory management component to utilize the conversation history context, which can effectively handle context reference and ellipsis phenomena in multi-turn question answering, but this method focuses on rule-based methods. Reinforcement learning has achieved good results in knowledge graph reasoning. The conversational question-and-answer system based on reinforcement learning needs to define a suitable action space and reward mechanism. For example, a conversational question-and-answer method for reinforcement learning from the reconstruction information of user questions makes full use of the feedback information of simulating whether the user's answer is correct to train the model, which depends on the explicit feedback of non-expert users and thus generates additional dependencies.

[0004] At the same time, in the current knowledge graph-based conversational question-and-answer systems, there is still a lack of combination of various constraint information in the question and the transfer of the conversational focus. Most of them only perform reasoning on the 1-hop path. Summary of the Invention

[0005] Based on the above, the present invention proposes a knowledge graph conversational question answering method based on reinforcement learning for problem constraint and focus entity transfer. By dynamically maintaining a context entity set to capture all possible focus entities involved during the conversation, the problem of focus entity transfer during the conversation is solved. The BERT pre-trained model is used to obtain the semantic information of the conversation context, solving the problem of semantic loss of subsequent questions during the conversation. The model combines question rounds, entity rounds, and various constraint information, and uses reinforcement learning to train the model on 1-hop paths and 2-hop constraint paths.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A knowledge graph conversational question answering method based on reinforcement learning for problem constraint and focus entity transfer, the steps are as follows:

[0008] Step 1: Data preparation and preprocessing

[0009] (1-1) Install the Wikidata knowledge graph dump: The Wikidata knowledge graph includes triple information <s, p, o> and qualifiers for supplementing and enhancing the triple information. The process of realizing question answering is a process of reasoning on triples and qualifiers on the knowledge graph;

[0010] The knowledge graph is a relational graph containing a set of known facts, which are stored in the form of triples <s, p, o>. Among them, the subject s and the object o are knowledge graph entities, and the predicate p is the relationship between them, that is, the knowledge graph path; in Wikidata, in addition to triple information, there are also qualifiers for supplementing and enhancing the triple information. The process of realizing question answering is a process of reasoning on triples and qualifiers on the knowledge graph; Load the dump data of Wikidata into the Neo4j graph database for subsequent access;

[0011] (1-2) Obtain the conversational question answering dataset: Adopt the publicly available dataset ConvRef;

[0012] The conversational question answering dataset uses the publicly available dataset ConvRef, which contains approximately 11,000 natural conversations. Each conversation contains at most 5 consecutive questions, and each question provides at most 5 rephrased questions. The rephrased questions are different expressions of the current question. When the current question cannot be answered correctly, the rephrased questions can be tried for answering; ConvRef can be evaluated through Wikidata; For each conversation, only the question-answer pair information in the publicly available dataset is used for model training;

[0013] (1-3) Detect the entities included in the question: Use the entity detection tool TagMe to detect all the entities that may be included in all the questions in the conversation and answer of ConvRef;

[0014] Step 2: Initialize the context entity set

[0015] The context entity set is used to maintain the relevant information of all the knowledge graph entities involved in the conversation and answer process; The initialization of the context entity set is initialized according to the question entity detection result in (1-3) of Step 1. The entities detected in the first question of the conversation are put into the context entity set, and the round of the entity is recorded as 0. The entity round is the question round where the entity first appears;

[0016] Step 3: Implement the construction of the conversational question and answer model

[0017] (3-1) Extraction of semantic information: First, use the BERT pre-trained model to obtain the embedding vector q of the current question t and the embedding vector q of the historical conversation questions t-1 . If the current question is the first-round question, the historical conversation question is set to an empty string;

[0018] Then concatenate q t and q t-1 , input them into the first fully connected layer, activate through ReLU and then input into the second fully connected layer, and finally output the context question vector representation Conq:

[0019] Conq = W2 × (ReLU(W1 × [q t-1 ; q t ))

[0020] Among them, W1 and W2 represent two fully connected layers, and ReLU represents the activation function;

[0021] Then obtain all the knowledge graph 1-hop paths starting from the entities in the context entity set from the knowledge graph, and use the BERT pre-trained model to obtain the embedding vector representations of all the paths to form the reinforcement learning action space As;

[0022] (3-2) Integrate the semantic information and various information included in the question to obtain the path probability of the knowledge graph to achieve reasoning: First, take the inner product of the context question vector representation Conq and the action space As to obtain the similarity evaluation vector of the question and the path; Input the similarity evaluation vector, time constraint, superlative constraint, current question round, and entity discovery round into a fully connected network for these 5-dimensional vectors:

[0023] Z = [As × Conq, time, max, qturn, eturn(e i )]

[0024] Among them, Conq is the context problem vector representation, As is the path action space, time and max represent whether there are time constraints and superlative constraints in the problem, qturn represents the problem round where the current problem is located, and eturn(e i ) represents the entity round of the entity e i at the starting point of the current path;

[0025] Finally, the selection probability of each path is obtained through the softmax function, and the entity obtained from the knowledge graph 1-hop path with the largest selection probability is the candidate answer entity. The definition of the 1-hop path probability distribution P(As) is as follows:

[0026] P(As) = σ(W3 × Z)

[0027] Among them, W3 is a fully connected network, and σ represents the softmax operation. According to the output probability distribution of the 1-hop path, the path with the largest probability can be selected to obtain the candidate answer entity;

[0028] When the number of candidate answers obtained from the 1-hop path is greater than 1, the training of the 2-hop constrained path will be triggered; the difference between the input of the 2-hop path and the 1-hop path lies in the path information, and the path starting from the candidate answer will be input; the definition of the output probability distribution P(A s2 ) is as follows:

[0029] P(A s2 ) = σ(W3 × Z2)

[0030] Z2 = [A s2 ×Conq,time,max,qturn,eturn(e j )]

[0031] Among them, A s2 represents the action space formed by the 2-hop path; e j is the starting entity of the 1-hop path; in order to enable the model to have the ability to simultaneously recognize the main body and constraints in the question sentence, the above selection of the 1-hop and 2-hop paths is realized through the same network;

[0032] Step 4: Model training based on reinforcement learning

[0033] To train the model, all dialogue questions and their golden answers in the public training set are used;

[0034] (4-1) For the 1-hop path, the accuracy score of the entity obtained from each 1-hop path relative to the golden answer is used as the reward for training the 1-hop path; to update the parameters θ of the policy function, the objective function is defined as follows:

[0035] J(θ) = E (q,ans)∈con E a∽π(θ) [Pre(q, ans, a)]

[0036] Among them, (q, ans) is a round of question and answer in the current dialogue con, ans is the golden answer, a is the action selected by the probability of the policy network, and Pre() is the accuracy of predicting the answer; Since the answers "Yes" and "No" to existential questions (i.e., judgment questions) cannot be directly obtained in the knowledge graph, only existential questions with the answer "Yes" and at least 2 entities can be recognized are trained, and one of the entities is regarded as the answer entity;

[0037] (4-2) For those paths that can obtain answers and whose accuracy is not 1, the 2-hop path will be expanded again, and the 2-hop paths will be sorted according to the manually defined constraint similarity rules, and then rewards will be given according to the sorting results:

[0038] The specific objective function is defined as follows:

[0039]

[0040] Among them, Score() is the reward obtained by the constraint entity corresponding to the action path a2 by applying the similarity rule;

[0041] Using the likelihood method, the gradients of the 1-hop path policy and the 2-hop constraint path policy are as follows:

[0042]

[0043]

[0044] After completing the training of the current problem, it is necessary to put the best answer of the current problem training and the entities detected in the next round of questions into the context entity set to expand the context entity set;

[0045] Step Five: Obtain answers from the knowledge graph through the trained model

[0046] Apply the trained model to the test set for testing to obtain answers; The process of obtaining answers is mainly divided into 2 steps: 1) Input the question information and the knowledge graph path information into the model to obtain the path probability, select the path according to the path probability, and then obtain the candidate answers according to the selected path; 2) If the number of candidate answers is greater than 1, continue to obtain the 2-hop path probability through the model, obtain the constraint entities according to the path probability, and sort the candidate answers according to the constraint entities.

[0047] Furthermore, the rules for sorting the candidate answers described in Step Five are as follows:

[0048] If there is a superlative constraint in the question, then for all candidate answers, sort the constraint entities corresponding to their 2-hop paths by time and return the sorting result to the user;

[0049] If a time constraint is identified in the question, then for all candidate answers, sort them according to the proximity of the constraint entities corresponding to their 2-hop paths to this time constraint, and then return the sorting result to the user;

[0050] If there is no superlative constraint and time constraint in the question, only consider entity constraints; for all candidate answers, if the constraint entity obtained by a certain 2-hop path can be identified in the question, rank it in the front; otherwise, use the intersection over union as a measure of similarity to calculate the similarity between the constraint entity and the question and perform sorting, and finally return the sorting result to the user.

[0051] The present invention has the following beneficial effects:

[0052] (1) The present invention uses a reinforcement learning policy network to select paths by capturing entity rounds, question rounds, and various constraint conditions in the question sentence, making the accuracy of the answer better.

[0053] (2) The present invention uses a context entity set and a BERT pre-trained model to maintain the context history from two aspects of entity transfer and semantic inheritance, making the capture of dialogue context information more accurate.

[0054] (3) The present invention applies corresponding reward mechanisms on the 1-hop path and 2-hop path of the knowledge graph to train the parameters of the policy network model, making it integrate the backbone information of 1-hop and the constraint information of 2-hop, and making the result more accurate. Brief Description of the Drawings

[0055] Figure 1 is the basic process diagram of the dialogue-based question answering based on reinforcement learning of the present invention;

[0056] Figure 2 is the structure of the policy network of the present invention, showing the selection process of the 1-hop path and the 2-hop constraint path;

[0057] Figure 3 is the schematic diagram of the dialogue-based question answering process of the present invention. Detailed Embodiments

[0058] In order to describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation schemes.

[0059] First, relevant definitions are given for some preliminary knowledge and concepts involved.

[0060] Definition 1. q t : question qt is the sequence of words proposed by the user in the tth round of dialogue question answering, which implies the intention of the user's query. t ={w1,w2,...,w n}, n represents the number of words in the current sentence.

[0061] Definition 2.q t-1 : Historical issues t-1 Refers to the question q in the conversation. t In the process of dialogue question and answer, due to semantic connection, the previous round of questions often implies important information for answering the current round of questions, especially predicate information.

[0062] Definition 3. Predict the answer: for question q t The predicted answer must be a fact limited to the knowledge graph, including entities, text strings, quantities, or "Yes" and "No" indicating judgments in the knowledge graph.

[0063] Definition 4. Question turn: The question turn qturn indicates when the question is asked. When qturn = 0, the question is the first question, and the user asks a question based on a certain topic at this time. Questions with qturn>0 will be followed by subsequent questions around this topic or related topics.

[0064] Definition 5. Entity turn: Entity turn records when an entity is raised in a conversation, which plays a crucial role in helping the model capture topic shifts. Intuitively, entities that appear in the first round of questions and entities and their answer entities that appear in the previous round of questions are more likely to become the focus entities of the current question.

[0065] Combined with the above definitions, we describe the final problem as: Based on the current round of questions q t and historical issues t-1 , modeling the contextual semantics of the entire conversational question and answer, and maintaining all possible subject entities of the current question through the context entity set, combining the question turn qturn, entity turn eturn and various constraint information, inferring the question q from the knowledge graph t The best answer.

[0066] This paper proposes a knowledge graph conversational question answering method based on question constraints and focus entity transfer based on reinforcement learning. Figure 2For the policy network of reinforcement learning, the semantic embedding representations of the current question and historical questions are obtained through the BERT pre-trained model respectively, and then they are passed through two fully connected networks and an activation layer to obtain a context embedding representation. Similarly, the semantic embedding representations of all knowledge graph candidate paths are obtained through the BERT pre-trained model, and then the context embedding is used to perform an inner product with the knowledge graph path semantic embedding to obtain a similarity rating score, which is used as one dimension of the input of a 5*1 fully connected layer. The other four dimensions are the entity round, question round, time constraint in the question, and superlative constraint respectively.

[0067] We take the publicly available dataset ConvRef dataset as an example for illustration. The ConvRef dataset is a dataset containing question reconstruction information created through the conversations in another dataset ConvQuestions. It sets up to 4 reconstructed questions for each question in ConvQuestions. The reconstructed questions are different expressions of the original question. When the original question cannot be answered, the question reconstruction can be triggered to answer again. This dataset contains approximately 11,200 conversations and 20,500 reconstructed questions. Each conversation has 5 rounds of Q&A, and the content covers 5 fields including books, movies, football, music, and TV dramas. It can be evaluated through the knowledge graph Wikidata. The ConvRef dataset website is https: / / conquer.mpi-inf.mpg.de / .

[0068] Specifically in implementation, it includes the following steps:

[0069] Step 1: Data preparation and preprocessing;

[0070] (1-1) The process of question and answer is the process of reasoning about triples and qualifiers on the knowledge graph. Therefore, first, the dump data of the knowledge graph Wikidata needs to be loaded into the Neo4j graph database for subsequent access. The present invention uses the Wikidata dump data on April 26, 2020 modified by Kaiser et al., which contains approximately 200 million triples, and redundant labels such as urls, external ids, and language tags have been deleted. Then this Wikidata dump data is loaded into the Neo4j graph database. The website of this dump data is https: / / github.com / PhilippChr / wikidata-core-for-QA.

[0071] (1-2) The dialogue question and answer dataset uses the publicly available dataset ConvRef. For each round of conversation, only the question-answer pair information is used for model training;

[0072] (1-3) Detect the entities included in the question: Use the entity detection tool TagMe to detect the entities that may be included in all questions in the dialogue Q&A, and implement it by directly calling the official API of TagMe through the Python language.

[0073] Step 2: Initialize the context entity set

[0074] The context entity set is used to maintain the relevant information of all knowledge graph entities involved in the dialogue Q&A process. Initialize the context entity set of the current dialogue according to the question entity detection result of TagMe, put the entities detected in the first question of the dialogue into the context entity set, and record the round of the entity as 0. Subsequently, during the process of the dialogue, new discovered relevant entities will be continuously added to the context entity set, and the round of the newly added entity will be recorded as the question round when the entity appears, so as to realize the expansion of the context entity set.

[0075] Step 3: Implement the construction of the conversational Q&A model

[0076] (3-1) Extraction of semantic information: First, obtain the embedding vector q of the current question through the BERT pre-trained model t and the embedding vector q of the historical questions in the dialogue t-1 . If the current question is the first-round question, set the historical questions in the dialogue as an empty string. Then concatenate q t and q t-1 , input the concatenated result into the first fully connected layer, activate it through ReLU, and then input it into the second fully connected layer. Finally, output the context question vector representation Conq:

[0077] Conq = W2 × (ReLU(W1 × [q t-1 ; q t ))

[0078] Among them, W1 and W2 represent two fully connected layers, ReLU represents the activation function, and the context question vector Conq synthesizes the historical information of the dialogue, which helps to complete the semantic missing of the current question. Then obtain all 1-hop paths of the knowledge graph starting from the entities in the context entity set from Wikidata saved in the Neo4j graph database, and use the BERT pre-trained model to obtain the embedding vector representations of all paths, and concatenate them to form the reinforcement learning action space matrix As;

[0079] (3-2) Integrate the semantic information and various constraint information included in the question to obtain the path probability of the knowledge graph for reasoning: First, take the inner product of the context question vector representation Conq and the action space As to obtain the similarity evaluation vector of the question and the path. Input the similarity evaluation vector, time constraint, superlative constraint, current question round, and starting entity round into a fully connected network:

[0080] Z = [As × Conq, time, max, qturn, eturn(e i )]

[0081] Where Conq is the context question vector representation, As is the path action space, time and max indicate whether there are time constraints and superlative constraints in the question, represented by 0 or 1, qturn represents the question round where the current question is located, and eturn(e i ) represents the entity round of the entity e from which the current path starts i ;

[0082] Finally, the selection probability of each path is obtained through the softmax function, and the entity obtained from the knowledge graph 1-hop path with the largest selection probability is the candidate answer entity. The definition of the 1-hop path probability distribution P(As) is as follows:

[0083] P(As) = σ(W3 × Z)

[0084] Where W3 is a fully connected network, and σ represents the softmax operation. According to the output probability distribution of the 1-hop path, the path with the largest probability can be selected to obtain the candidate answer entity

[0085] When there is more than one candidate answer obtained from the 1-hop path, the training of the 2-hop constrained path will be triggered. The input of the 2-hop path is similar to that of the 1-hop path. The difference is that for the path information, the path starting from the candidate answer will be input. The output probability distribution P(A s2 ) is defined as follows:

[0086] P(A s2 ) = σ(W3 × Z2)

[0087] Z2 = [A s2 × Conq, time, max, qturn, eturn(e j )]

[0088] Where A s2 represents the action space composed of the 2-hop path. ej is the starting entity of the 1-hop path; in order to enable the model to have the ability to identify both the main body and constraints in the question sentence, the above selection of the 1-hop and 2-hop paths is implemented through the same network; the overall model structure is as Figure 2 shown

[0089] Step 4: Model training based on reinforcement learning

[0090] To train the model, all the dialogue questions and their golden answers in the publicly available training set ConvRef are used

[0091] (4-1) For the 1-hop paths, the precision scores of the entities obtained from each 1-hop path relative to the gold answer are used as the rewards for training the 1-hop paths. To update the parameters θ of the policy function, the objective function is defined as follows:

[0092] J(θ) = E (q,ans)∈con E a∽π(θ) [Pre(q, ans, a)]

[0093] Among them, (q, ans) is a round of question and answer in the current conversation con, ans is the gold answer, a is the action selected by the probability of the policy network, and Pre() is the precision of the predicted answer. Since the answers "Yes" and "No" to existence questions (i.e., judgment questions) cannot be directly obtained in the knowledge graph, only existence questions with the answer "Yes" and at least 2 entities that can be recognized are trained, and one of the entities is regarded as the answer entity.

[0094] (4-2) For the paths that can obtain answers and whose precision is not 1, the 2-hop paths will be expanded again and sorted according to the constraint similarity rules defined below, and then rewards will be given according to the sorting results:

[0095] 1 If the highest-level constraint (such as first, etc.) is recognized in the question, then for all 2-hop paths, they are sorted according to their corresponding constraint entities in time. The paths with non-time-type constraint entities will be ranked behind the paths with time constraint entities, and then rewards will be given according to the sorting results of the paths;

[0096] 2 If the time constraint (such as 2013, etc.) is recognized in the question, the 2-hop paths will be sorted and rewarded in a similar way to the highest-level constraint, and the constraint entities closer to the time constraint will be ranked in the front;

[0097] 3 If there is no highest-level constraint and time constraint in the question, only entity constraints are considered. If the constraint entity obtained by a 2-hop path can be recognized in the question, a reward of 1 will be given to this 2-hop path. Otherwise, the intersection over union is used as a measure of similarity to calculate the similarity between the constraint entity and the question and sort them, and then rewards will be given according to the sorting results of the paths;

[0098] It should be noted that since Wikidata has qualifier information in addition to triple information, the above rules will consider qualifiers at the same time. At the same time, there may be cases where the similarity results are the same. The present invention will sort them again according to the popularity of the corresponding answer entities, and rank the entities with higher popularity in the front (the popularity is the number of neighbor entities. From common sense, the higher the popularity, the more likely it is to be asked).

[0099] The specific objective function is defined as follows:

[0100]

[0101] Among them, Score() is the reward obtained by the constraint entity corresponding to the action path a2 through applying the similarity rule;

[0102] Using the likelihood method, the gradients of the 1-hop path strategy and the 2-hop constrained path strategy are as follows:

[0103]

[0104]

[0105] Step Five: Obtain the answer from the knowledge graph through the trained model

[0106] After the model training is completed, applying the strategy on the test set can obtain the answer. The process of obtaining the answer is mainly divided into two steps: 1) Input the question information and the knowledge graph path information into the model to obtain the path probability, and select the action according to the probability to obtain the candidate answers. 2) If the number of candidate answers is greater than 1, continue to obtain the 2-hop path probability through the model, obtain the constraint entity according to the probability, and sort the candidate answers according to the similarity rule defined as follows:

[0107] 1 If there is a superlative constraint (such as first, etc.) in the question, then for all candidate answers, sort them by time according to the constraint entity corresponding to their 2-hop path and return the sorting result to the user;

[0108] 2 If a time constraint (such as 2013, etc.) is recognized in the question, then for all candidate answers, sort them according to the proximity of the constraint entity corresponding to their 2-hop path to this time constraint, and then return the sorting result to the user;

[0109] 3 If there is no superlative constraint and time constraint in the question, only consider the entity constraint. For all candidate answers, if the constraint entity obtained by a 2-hop path can be recognized in the question, rank it in the front. Otherwise, use the intersection over union as the measure of similarity to calculate the similarity between the constraint entity and the question and sort them. Finally, return the sorting result to the user;

[0110] Since the ConvRef dataset contains question reconstruction information, so as done in other experiments, when the accuracy of the obtained answer is not equal to 1, the reconstructed question can be used to try to answer again.

[0111] Step 6: Model and Test Details: The reinforcement learning code of the present invention is written in the Python language and uses the deep learning library Pytorch. The model is trained for 30 epochs on the training set. When updating the parameters, the Adam optimizer is used, and the warm-up learning rate is adopted. Specifically, the learning rate will rapidly increase from 0 to the maximum value of 10 at the beginning of training -5 and will gradually decrease during subsequent training. Using the warm-up learning rate can make the model converge faster and achieve better results. Finally, the code runs on the Ubuntu 20.04 system, configured with an Intel(R) Xeon(R) Silver 4210R CPU @ 2.40 GHz, 32G RAM, and an NVIDIA GeForce RTX 3090 GPU. In addition, the Neo4j database we use to store Wikidata runs on another computer with 64G RAM to meet the requirements of data storage and access.

[0112] According to the operation process of the above steps, the knowledge graph conversational question-answering method based on reinforcement learning proposed by the present invention can be realized. To verify the technical effect of the present invention in the question-answering system, the present invention conducts experiments using the publicly available dataset ConvRef and evaluates through the knowledge graph Wikidata. For the experiment on the ConvRef dataset, only the ConvRef training set with the question reconstruction information removed is used, and then the test is conducted on the ConvRef test set. This is because ConvRef is obtained by supplementing the question reconstruction information of another dataset ConvQuestions, and the present invention does not need to use the reconstruction information in the ConvRef training set.

[0113] The evaluation metrics of ConvRef are precision (P@1), mean reciprocal rank (MMR), hit rate (H@5), and RefTriggers (number of reconstruction triggers). To illustrate the effect, the present invention is compared with related methods using the same dataset, and the comparison results are shown in Table 1.

[0114] The present invention improves by 4.1% under the P@1 metric of the ConvRef dataset and uses 2525 fewer reconstructed questions. From the comparison with the optimal method PRALINE, it can be seen that the present invention improves by 5.9% compared with the PRALINE result under the P@1 metric. For the P@1 metric, it mainly evaluates the first returned answer, indicating that the answer ranking method based on the 2-hop constrained path of the present invention greatly improves the result.

[0115] Table 1: Results on the ConvRef dataset

[0116]

Claims

1. A knowledge graph conversational question-answering method based on reinforcement learning for problem constraint and focus entity transfer, characterized in that By dynamically maintaining a set of context entities to capture all possible focus entities involved during the conversation, the problem of focus entity transfer during the conversation is solved; the BERT pre-trained model is used to obtain the semantic information of the conversation context to solve the problem of semantic missingness in subsequent questions during the conversation; the model combines the question round, entity round, and various constraint information, and uses reinforcement learning to train the model on 1-hop paths and 2-hop constraint paths; the specific steps are as follows: Step 1: Data preparation and preprocessing Step 2: Initialize the context entity set Step 3: Implement the construction of the conversational question-answering model (3-1) Extraction of semantic information: First, obtain the embedding vector q of the current question and the embedding vector q of the historical dialogue questions respectively through the BERT pre-trained model. If the current question is the first-round question, set the historical dialogue question as an empty string. t and the embedding vector q of the historical dialogue questions t-1 , if the current question is the first-round question, set the historical dialogue question as an empty string; Then concatenate q t and q t-1 and input the result into the fully connected layer 1, then activate it through ReLU and input it into the fully connected layer 2, and finally output the context question vector representation Conq: Conq = W2 × (ReLU(W1 × [q t-1 ; q t )) Among them, W1 and W2 represent two fully connected layers, and ReLU represents the activation function; Then, all knowledge graph 1-hop paths starting from the entities in the context entity set are obtained from the knowledge graph, and the BERT pre-trained model is used to obtain the embedding vector representations of all paths, constituting the reinforcement learning action space As; (3-2) Integrate the semantic information and various information contained in the question to obtain the path probability of the knowledge graph for reasoning: First, take the inner product of the context question vector representation Conq and the action space As to obtain the similarity evaluation vector of the question and the path; input the similarity evaluation vector, time constraint, superlative constraint, current question round, and entity discovery round into a fully connected network for these 5-dimensional vectors: Z = [As×Conq, time, max, qturn, eturn(e i )] Among them, Conq is the context problem vector representation, As is the path action space, time and max indicate whether there are time constraints and superlative constraints in the problem, qturn represents the problem round where the current problem is located, and eturn(e i ) represents the entity round of the current path starting entity e i ; Finally, the selection probability of each path is obtained through the softmax function, and the entity obtained from the knowledge graph 1-hop path with the largest selection probability is the candidate answer entity. The definition of the 1-hop path probability distribution P(As) is as follows: P(As) = σ(W3 × Z) Among them, W3 is the fully connected network, and σ represents the softmax operation. According to the output probability distribution of the 1-hop path, the path with the largest probability can be taken to obtain the candidate answer entity; When the number of candidate answers obtained by the 1-hop path is greater than 1, the training of the 2-hop constrained path will be triggered; the difference between the input of the 2-hop path and the 1-hop path lies in that for the path information, the path starting from the candidate answer will be input; the output probability distribution P(A s2 ) is defined as follows: P(A s2 ) = σ(W3 × Z2) Z2 = [A s2 ×Conq,time,max,qturn,eturn(e j )] Among them, A s2 represents the action space composed of 2-hop paths; ej is the starting entity of the 1-hop path; Step 4: Model training based on reinforcement learning Step 5: Obtain answers from the knowledge graph through the trained model.

2. The knowledge graph conversational question answering method based on reinforcement learning for problem constraint and focus entity transfer according to claim 1, characterized in that, The specific content of the said Step 1 is as follows: (1-1) Install the Wikidata knowledge graph dump: The Wikidata knowledge graph includes triple information <s, p, o> and qualifiers that supplement and enhance the triple information. The process of realizing question answering is a process of reasoning about triples and qualifiers on the knowledge graph; (1-2) Obtain the conversational question-answering dataset: Adopt the publicly available dataset ConvRef; (1-3) Detect the entities contained in the question: Use the entity detection tool TagMe to detect all entities that may be contained in the questions in the conversational questions and answers of ConvRef.

3. A knowledge graph conversational question answering method based on reinforcement learning for problem constraint and focus entity transfer according to claim 1, characterized in that, The specific content of the said Step 2 is as follows: The context entity set is used to maintain the relevant information of all knowledge graph entities involved in the process of conversational question answering; the initialization of the context entity set is carried out according to the question entity detection result in (1-3) of Step 1. The entities detected in the first question of the conversation are put into the context entity set, and the entity round is recorded as 0. The entity round is the question round when the entity first appears.

4. A knowledge graph conversational question answering method based on reinforcement learning for problem constraint and focus entity transfer according to claim 1, characterized in that, The specific content of the said Step 4 is as follows: (4-1) For the 1-hop paths, the precision scores of the entities obtained from each 1-hop path relative to the golden answer are used as the rewards for training the 1-hop paths; to update the parameters θ of the policy function, the objective function is defined as follows: J(θ) = E (q,ans)∈con E a∽π(θ) [Pre(q, ans, a)] where (q, ans) is a question-answer pair in the current conversation con, ans is the golden answer, a is the action selected with probability by the policy network, and Pre() is the precision of the predicted answer; since the answers "Yes" and "No" to existence questions (i.e., judgment questions) cannot be directly obtained from the knowledge graph, only existence questions with the answer "Yes" and at least 2 entities identified are trained, and one of the entities is regarded as the answer entity. (4-2) For those paths that can obtain answers and whose precision is not 1, the 2-hop paths will be further expanded, and the 2-hop paths will be sorted according to the manually defined constraint similarity rules, and then rewards will be given according to the sorting results: The specific objective function is defined as follows: where Score() is the reward obtained by the constraint entity corresponding to the action path a2 by applying the similarity rules. Using the likelihood method, the gradients of the 1-hop path policy and the 2-hop constraint path policy are as follows: After completing the training of the current question, the best answer to the current question training and the entities detected in the next round need to be put into the context entity set to expand the context entity set.

5. A knowledge graph conversational question answering method based on reinforcement learning for problem constraint and focus entity transfer according to claim 1, characterized in that, The specific steps of Step Five are as follows: Apply the trained model to the test set for testing to obtain answers; the process of obtaining answers is mainly divided into two steps: 1) Input the question information and the knowledge graph path information into the model to obtain the path probabilities, select paths according to the path probabilities, and then obtain candidate answers according to the selected paths; 2) If the number of candidate answers is greater than 1, continue to obtain the 2-hop path probabilities through the model, obtain the constraint entities according to the path probabilities, and sort the candidate answers according to the constraint entities.

6. The method for knowledge graph conversational question answering based on reinforcement learning for problem constraint and focus entity transfer according to claim 5, characterized in that The rules for sorting candidate answers described in Step Five are as follows: A If there is a superlative constraint in the question, for all candidate answers, sort them according to the constraint entities corresponding to their 2-hop paths in terms of time and return the sorting results to the user; B If a time constraint is identified in the question, for all candidate answers, sort them according to the proximity of the constraint entities corresponding to their 2-hop paths to the time constraint, and then return the sorting results to the user; C If there is no superlative constraint and time constraint in the question, only entity constraints are considered; for all candidate answers, if the constraint entity obtained by a 2-hop path can be identified in the question, it will be ranked in the front; otherwise, use the intersection over union as a measure of similarity to calculate the similarity between the constraint entity and the question and sort them, and finally return the sorting results to the user.

Citation Information

Patent Citations

  • Knowledge graph multi-hop question and answer method based on reinforcement learning path reasoning

    CN115640410A

  • Knowledge graph question-answer method and apparatus based on deep learning technology, and device

    WO2021139283A1