Question and answer agent training method and device, equipment and storage medium
By training the question generator to generate Q&A samples on the timing knowledge graph, the problem of poor training effect of Q&A model in the prior art is solved, and more efficient Q&A agent training and output quality are achieved.
Patent Information
- Application Number
- CN202510114674.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
The question-and-answer samples constructed based on manual templates in the prior art leads to poor training effects of the question-and-answer model and is unable to effectively cover the complex semantic space.
By training the question generator based on the timing knowledge graph, high-quality Q&A pair samples are generated and the Q&A agent is trained using these samples.
It improves the data diversity of the Q&A dataset, can cover complex semantic space, and optimizes the training effect and output quality of Q&A agents.
Smart Images

Figure CN120069065A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a training method, device, equipment and storage medium for a question-and-answer intelligent agent. Background Technique
[0002] The Temporal Knowledge Graph Question and Answer (TKGQA) task is an important branch of the knowledge graph question and answer task, aiming to find an entity or timestamp from a temporal knowledge graph (TKG) to answer temporal reasoning questions.
[0003] In the related art, in order to improve the accuracy of question and answer based on the temporal knowledge graph, a variety of temporal reasoning question templates are constructed, and entity aliases in the database are used to fill these temporal reasoning question templates, so as to obtain tens of thousands of question-and-answer pair samples, and then these question-and-answer pair samples are used to train a question-and-answer model.
[0004] Obviously, the semantic space that can be covered by the question-and-answer pair samples constructed based on artificial templates is limited, resulting in poor training effects of the question-and-answer model. Summary of the Invention
[0005] The embodiments of the present application provide a training method, device, equipment and storage medium for a question-and-answer intelligent agent, and the technical solutions are as follows:
[0006] On the one hand, the embodiments of the present application provide a training method for a question-and-answer intelligent agent, and the method includes:
[0007] Based on a temporal knowledge graph and a first question-and-answer data set corresponding to the temporal knowledge graph, train a question generator, where the temporal knowledge graph includes at least two temporal fact groups, and each temporal fact group includes fact information with temporal characteristics;
[0008] Input the temporal knowledge graph and question generation prompt words into the question generator, and output a second question-and-answer data set through the question generator. The question generation prompt words are used to indicate the generation of question-and-answer pair samples for at least one temporal fact group, and the data volume of the second question-and-answer data set is higher than that of the first question-and-answer data set;
[0009] Based on the second question-and-answer data set, train a question-and-answer intelligent agent, where the question-and-answer intelligent agent is used to output an answer corresponding to the temporal question based on the input temporal question, and the temporal question refers to a question with temporal reasoning logic.
[0010] On the other hand, the embodiments of the present application provide a training device for a question-and-answer intelligent agent, and the device includes:
[0011] A first training module, configured to train a question generator based on a temporal knowledge graph and a first question-and-answer data set corresponding to the temporal knowledge graph, where the temporal knowledge graph includes at least two temporal fact groups, and each temporal fact group includes factual information with temporal characteristics;
[0012] A first output module, configured to input the temporal knowledge graph and question generation prompt words into the question generator, and output a second question-and-answer data set through the question generator. The question generation prompt words are used to indicate the generation of question-and-answer pair samples for at least one temporal fact group, and the data volume of the second question-and-answer data set is higher than that of the first question-and-answer data set;
[0013] A second training module, configured to train a question-and-answer agent based on the second question-and-answer data set. The question-and-answer agent is configured to output an answer corresponding to the temporal question based on the input temporal question, where the temporal question refers to a question with temporal reasoning logic.
[0014] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the training method of the question-and-answer agent as described in the above aspect.
[0015] On the other hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the training method of the question-and-answer agent as described in the above aspect.
[0016] On the other hand, an embodiment of the present application provides a computer program product, which includes at least one instruction, and the at least one instruction is stored in a computer-readable storage medium. A processor of a computer device reads the at least one instruction from the computer-readable storage medium, and the processor executes the at least one instruction, so that the computer device executes the training method of the question-and-answer agent as described in the above aspect.
[0017] In the embodiment of the present application, by first training a question generator based on a temporal knowledge graph and a first question-and-answer data set corresponding to the temporal knowledge graph, the question generator can generate high-quality question-and-answer pair samples. Then, inputting the temporal knowledge graph and question generation prompt words into the question generator to obtain a second question-and-answer data set output by the question generator can improve the data diversity of the question-and-answer data set, so that the question-and-answer data set can cover a complex semantic space. Furthermore, using the second question-and-answer data set to train the question-and-answer agent can optimize the training effect of the question-and-answer agent and improve the output quality of the question-and-answer agent. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0019] Figure 1 Shows a block diagram of the structure of a computer system provided by an exemplary embodiment of the present application;
[0020] Figure 2 Shows a flowchart of a method for training a question-and-answer intelligent agent provided by an exemplary embodiment of the present application;
[0021] Figure 3 Shows a flowchart of a process for constructing a temporal knowledge graph provided by an exemplary embodiment of the present application;
[0022] Figure 4 Shows a schematic diagram of determining a factual association relationship provided by an exemplary embodiment of the present application;
[0023] Figure 5 Shows a flowchart of a process for enriching a data set by question rewriting provided by an exemplary embodiment of the present application;
[0024] Figure 6 Shows a flowchart of a process for screening a data set through quality assessment provided by an exemplary embodiment of the present application;
[0025] Figure 7 Shows a flowchart of a process for training a question-and-answer intelligent agent provided by an exemplary embodiment of the present application;
[0026] Figure 8 Shows a schematic diagram of the training process of a question-and-answer intelligent agent provided by an exemplary embodiment of the present application;
[0027] Figure 9 Shows a comparison diagram of dimensionality reduction representation of a data set provided by an exemplary embodiment of the present application;
[0028] Figure 10 Shows a block diagram of the structure of a training device for a question-and-answer intelligent agent provided by an exemplary embodiment of the present application;
[0029] Figure 11 Shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of this application more clear, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0031] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0032] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0033] It should be understood that although the terms first, second, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0034] First, a brief introduction to the nouns involved in the embodiments of this application:
[0035] Temporal Knowledge Graph: An extended knowledge graph that not only contains static entities and relationships but also introduces the time dimension and can represent the changes of entities and relationships over time. That is, the fact groups in the temporal knowledge graph consist of entities, relationships, and timestamps.
[0036] Fact group: Includes factual information with temporal characteristics. That is, the fact group is used to characterize the relationships between entities at a specific time point or time period. The fact group includes entities, relationships, and timestamps. In the embodiments of this application, the fact group can be in the form of a quadruple, expressed as (subject, relationship, object, time point); or in the form of a quintuple, expressed as (subject, relationship, object, start time, end time).
[0037] Please refer to Figure 1, which shows a structural block diagram of a computer system provided by an exemplary embodiment of the present application. The computer system may include a terminal 110, a server 120, and a terminal 130. Among them, data communication is carried out between the terminal 110 and the server 120, and between the server 120 and the terminal 130 through a communication network. Optionally, the communication network may be a wired network or a wireless network, and the communication network may be at least one of a local area network, a metropolitan area network, and a wide area network.
[0038] The terminal 110 is an electronic device installed with a model training client. Among them, the terminal 110 may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart TV, a wearable device, a vehicle-mounted terminal, etc. Figure 1 Only taking the terminal 110 as a desktop computer as an example for illustration, but not limited thereto.
[0039] The server 120 includes at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. In an embodiment of the present application, the server 120 may be a background server deployed with a question generator and a question-and-answer intelligent agent.
[0040] The terminal 130 is an electronic device installed with an application program having a question-and-answer function. The question-and-answer function may be a function of a native application in the terminal 130, or a function of a third-party application; the terminal 130 may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart TV, a wearable device, a vehicle-mounted terminal, etc. Figure 1 Only taking the terminal 130 as a desktop computer as an example for illustration, but not limited thereto.
[0041] In some embodiments, there is data interaction between the server 120 and the terminal 110. Schematically, as Figure 1 shown, the server 120 trains a question generator based on a temporal knowledge graph and a first question-and-answer data set corresponding to the temporal knowledge graph. After completing the training of the question generator, the server 120 further inputs the temporal knowledge graph and a question generation prompt word sent by the terminal 110 into the question generator, and outputs a second question-and-answer data set through the question generator, so as to train a question-and-answer intelligent agent based on the second question-and-answer data set.
[0042] In some embodiments, there is data interaction between the server 120 and the terminal 130. Schematically, as Figure 1 shown, after the training of the question-and-answer intelligent agent is completed, the server 120 can send an inference interface corresponding to the question-and-answer intelligent agent to the terminal 130, so that based on the temporal question input by the user in the terminal 130, the server 120 returns an answer corresponding to the temporal question to the terminal 130.
[0043] Based on the above introduction, the training method of the question-and-answer intelligent agent provided in this application will be described. This method can be executed by a server or a terminal, or jointly executed by a server and a terminal.
[0044] Please refer to Figure 2 , which shows a flowchart of the training method of the question-and-answer intelligent agent provided by an exemplary embodiment of this application. In this embodiment, it is described by taking this method being used in a computer device (including a terminal and / or a server) as an example. This method includes the following steps:
[0045] Step 210, based on the temporal knowledge graph and the first question-and-answer data set corresponding to the temporal knowledge graph, train a question generator. The temporal knowledge graph includes at least two temporal fact groups, and each temporal fact group includes factual information with temporal characteristics.
[0046] Optionally, the temporal knowledge graph is an extended knowledge graph. It not only contains static entities and relationships but also introduces a time dimension, capable of representing the changes of entities and relationships over time. Compared with the knowledge graph, the temporal knowledge graph can capture and represent dynamic information more comprehensively and is applicable to scenarios where time factors need to be considered.
[0047] Optionally, the temporal knowledge graph includes at least two temporal fact groups, and each fact group includes factual information with temporal characteristics. That is, the fact group is used to characterize the relationships between entities at a specific time point or time period. The fact group includes entities, relationships, and timestamps. In the embodiments of this application, the fact group can be in the form of a quadruple, expressed as (subject, relationship, object, time point); or in the form of a quintuple, expressed as (subject, relationship, object, start time, end time).
[0048] Optionally, the temporal knowledge graph can include multiple temporal fact subgraphs, and the fact subgraph is composed of at least two temporal fact groups with factual association relationships. Among them, the at least two temporal fact groups with factual association relationships can include the same subject, or the same object, or have a time logical relationship.
[0049] Optionally, the question generator is used to generate question-and-answer pair samples required for training the question-and-answer intelligent agent. The question-and-answer pair samples include question samples and answer samples. The network structure of the question generator can be the network structure of a large language model (LLM). The input of the question generator is the temporal knowledge graph and question generation prompt words, and the output is question-and-answer pairs.
[0050] In some embodiments, before using the question generator to generate Q&A pair samples, in order to improve the data generation quality of the question generator, the question generator also needs to be trained first. Optionally, the computer device can train the question generator according to the temporal knowledge graph and the first Q&A dataset corresponding to the temporal knowledge graph.
[0051] Among them, the first Q&A dataset contains high-quality Q&A pair samples, but the sample size of the contained Q&A pair samples is small. In a possible implementation manner, the question generator can be the second large language model, and the Q&A pair samples in the first Q&A dataset are generated by the third large language model based on the input temporal knowledge graph and question generation prompt words. For example, the first Q&A dataset can be generated by GPT (Generative Pre-Trained Transformer) based on the input temporal knowledge graph and question generation prompt words.
[0052] In a possible implementation manner, the computer device inputs the temporal knowledge graph, the question generation prompt words, and the first Q&A dataset into the question generator. The question generator learns the first Q&A dataset and outputs predicted Q&A pairs. Then, according to the Q&A pair generation loss between the predicted Q&A pairs and the true values of the Q&A pairs in the first Q&A dataset, the question generator is trained. Through multiple backpropagation updates, the trained question generator can be obtained.
[0053] Step 220: Input the temporal knowledge graph and the question generation prompt words into the question generator, and output the second Q&A dataset through the question generator. The question generation prompt words are used to indicate the generation of Q&A pair samples for at least one temporal fact group, and the data volume of the second Q&A dataset is higher than that of the first Q&A dataset.
[0054] After the question generator is trained, the computer device can generate a large number of Q&A pair samples through the question generator to enrich the Q&A dataset. In a possible implementation manner, the computer device inputs the temporal knowledge graph and the question generation prompt words into the question generator, and the Q&A pair samples output by the question generator can be obtained. Thus, a large number of Q&A pair samples constitute the second Q&A dataset.
[0055] Optionally, the question generation prompt is used to indicate the generation of Q&A pair samples for at least one group of temporal facts. Optionally, the question generation prompt may include a question generation prompt statement, an explanation statement for the group of temporal facts, and may also include requirements for the output format of the Q&A pairs. Exemplarily, the question generation prompt may be "Select at least one fact from the given temporal knowledge graph to generate questions and answers", the explanation statement for the group of temporal facts may be "Each group of facts is presented in the form of (subject, relationship, object, start time, end time), indicating that 'the relationship between the subject and the object between the start time and the end time is the object'", and the requirements for the output format of the Q&A pairs may be "Answer in English, and the questions and answers are output in the format of a JSON list: {"question": str, "answer": str}", etc.
[0056] Optionally, the data volume of the second Q&A dataset is higher than that of the first Q&A dataset. Moreover, since the second Q&A dataset is generated by a large language model, compared with the Q&A dataset generated by artificial templates, the second Q&A dataset can cover a larger and more complex semantic space and has higher data quality.
[0057] Step 230: Train a Q&A agent based on the second Q&A dataset. The Q&A agent is used to output an answer corresponding to the input temporal question. A temporal question refers to a question with temporal reasoning logic.
[0058] Optionally, the Q&A agent may include a large language model. The large language model processes the question input by the user and understands the question intention, so as to output the answer corresponding to the question.
[0059] In the embodiments of the present application, the Q&A agent is used to output an answer corresponding to the input temporal question, where the temporal question refers to a question with temporal reasoning logic. For example, the temporal question is "In which year did Xiaoming enter No. 1 Middle School?".
[0060] In some embodiments, after obtaining the second Q&A dataset, the computer device can use the second Q&A dataset to train the Q&A agent. Optionally, the computer device inputs the question samples included in the Q&A pair samples into the Q&A agent to obtain the answer prediction result output by the Q&A agent, and then trains the Q&A agent based on the answer prediction result and the answer samples. Through multiple rounds of training iterations, the training of the Q&A agent can be completed, and the trained Q&A agent can be applied to the Q&A scenario.
[0061] In summary, in the embodiments of the present application, by first training a question generator based on a temporal knowledge graph and a first question-and-answer data set corresponding to the temporal knowledge graph, enabling the question generator to generate high-quality question-and-answer pair samples, and then inputting the temporal knowledge graph and question generation prompt words into the question generator to obtain a second question-and-answer data set output by the question generator, the data diversity of the question-and-answer data set can be improved, so that the question-and-answer data set can cover complex semantic spaces. Furthermore, using the second question-and-answer data set to train a question-and-answer agent can optimize the training effect of the question-and-answer agent and improve the output quality of the question-and-answer agent.
[0062] In some embodiments, in order to improve the training efficiency and question generation quality of the question generator, before obtaining the question-and-answer pair data set, it is first necessary to construct a temporal knowledge graph to lay a foundation for subsequent question generation. Optionally, the process of constructing the temporal knowledge graph may include the following steps:
[0063] Step 301, obtain at least two temporal fact groups.
[0064] Optionally, the computer device may obtain at least two temporal fact groups from a large knowledge base.
[0065] In a possible implementation manner, the computer device may extract fact information with temporal characteristics from the knowledge base according to keywords characterizing time features, and thus represent the fact information in the form of a temporal fact group.
[0066] Step 302, determine the fact association relationships between at least two temporal fact groups, where the fact association relationships include at least one of a subject association relationship, an object association relationship, and a time association relationship.
[0067] To ensure the temporality and relevance of the data, the computer device also needs to determine the fact association relationships between at least two temporal fact groups. Among them, the fact association relationships between at least two temporal fact groups may include at least one of a subject association relationship, an object association relationship, and a time association relationship.
[0068] Optionally, when at least two temporal fact groups correspond to the same fact subject, it can be determined that at least two temporal fact groups have a subject association relationship. Schematically, as Figure 4 shown, the temporal fact group 1 is (Xiaoming, studying at, the First Primary School, 2001, 2007), and the temporal fact group 2 is (Xiaoming, obtaining, the title of excellent student, 2004). Obviously, the temporal fact group 1 and the temporal fact group 2 contain the same fact subject, so it can be determined that there is a subject association relationship between the temporal fact group 1 and the temporal fact group 2.
[0069] Optionally, when at least two time-series fact groups correspond to the same fact object, it can be determined that at least two time-series fact groups have an object association relationship. Schematically, as Figure 4 shown, the time-series fact group 1 is (Xiaoming, studying at, the First Primary School, 2001, 2007), and the time-series fact group 2 is (Xiaogang, working at, the First Primary School, 2002, 2010). Obviously, the time-series fact group 1 and the time-series fact group 2 contain the same fact object. Therefore, it can be determined that there is an object association relationship between the time-series fact group 1 and the time-series fact group 2.
[0070] Optionally, when there is a time logic relationship between at least two time-series fact groups, it can be determined that at least two time-series fact groups have a time association relationship. Schematically, as Figure 4 shown, the time-series fact group 1 is (Xiaoming, studying at, the First Primary School, 2001, 2007), and the time-series fact group 2 is (Event A, occurring at, null, 2003). Obviously, the time period included in the time-series fact group 1 and the time point included in the time-series fact group 2 have an intersection. Therefore, it can be determined that there is a time association relationship between the time-series fact group 1 and the time-series fact group 2.
[0071] Step 303: Construct a time-series knowledge graph based on at least two time-series fact groups and the fact association relationships between at least two time-series fact groups.
[0072] Furthermore, after determining the fact association relationships corresponding to each time-series fact group, the computer device can construct a time-series knowledge graph according to at least two time-series fact groups and the fact association relationships between at least two time-series fact groups.
[0073] In a possible implementation manner, the computer device can first construct a time-series fact sub-graph according to at least two time-series fact groups with fact association relationships, and then construct a time-series knowledge graph according to the obtained multiple time-series fact sub-graphs.
[0074] In the above embodiments, after obtaining the time-series fact groups, by first determining the fact association relationships between at least two time-series fact groups and then constructing a time-series knowledge graph, the time-series and relevance of each time-series fact group are characterized by the time-series knowledge graph, which is beneficial to improving the understanding efficiency of the fact information by the question generator and thus improving the question generation quality.
[0075] In some embodiments, in order to achieve the diversity of question generation, in addition to generating questions based on as much fact information as possible, the correspondence between questions and fact information can also be considered. For example, one question corresponds to one time-series fact group, that is, a simple single-hop question is generated; or one question corresponds to multiple time-series fact groups, that is, a complex multi-hop question is generated.
[0076] Optionally, by setting different question generation prompt words, the question generator can be controlled to generate questions with different corresponding relationships.
[0077] In a possible implementation, the computer device can input the temporal knowledge graph and the first question generation prompt word into the question generator, and output a second Q&A dataset of the first type through the question generator.
[0078] Among them, the first question generation prompt word is used to indicate generating Q&A pair samples for a single temporal fact group, and the second Q&A dataset of the first type contains simple single-hop Q&A pair samples. For example, the first question generation prompt word is "According to the given temporal knowledge graph, select a fact from it to generate questions and answers".
[0079] In another possible implementation, the computer device can input the temporal knowledge graph and the second question generation prompt word into the question generator, and output a second Q&A dataset of the second type through the question generator.
[0080] Among them, the second question generation prompt word is used to indicate generating Q&A pair samples for at least two temporal fact groups, and the at least two temporal fact groups have fact association relationships. The second Q&A dataset of the second type contains complex multi-hop Q&A pair samples. For example, the second question generation prompt word is "According to the given temporal knowledge graph, select at least two facts from it to generate questions and answers".
[0081] Furthermore, the computer device can obtain the second Q&A dataset based on the second Q&A dataset of the first type and the second Q&A dataset of the second type, so that the second Q&A dataset includes both simple single-hop Q&A pair samples and complex multi-hop Q&A pair samples.
[0082] In the above embodiments, by setting different question generation prompt words, the question generator can generate Q&A pair samples for different numbers of temporal fact groups, that is, obtain Q&A pair samples with different degrees of complexity, thereby improving the diversity of Q&A pair samples and being beneficial to improving the training effect of subsequent Q&A agents.
[0083] In some embodiments, after outputting the second Q&A dataset through the question generator, in order to further increase the diversity of Q&A pair samples, the question samples can also be rewritten to obtain question samples with similar semantics under different expressions, thereby enriching the second Q&A dataset.
[0084] Optionally, the process of enriching the second Q&A dataset by question rewriting can include the following steps:
[0085] Step 501: Input the question samples included in the Q&A pair samples in the second Q&A dataset and the question rewriting prompt words into the question rewriting network, and obtain the question rewriting samples output by the question rewriting network.
[0086] In a possible implementation manner, after obtaining the second Q&A dataset, the computer device can obtain the question samples included in the Q&A pair samples from the second Q&A dataset, and thus input the question samples and the question rewriting prompt words into the question rewriting network to obtain the question rewriting samples output by the question rewriting network.
[0087] Optionally, the question rewriting samples are texts that are semantically similar to the question samples but have different expressions. For example, the question sample can be "When does the meeting start?", and the question rewriting sample can be "What is the start time of the meeting?".
[0088] Optionally, the question rewriting prompt words are used to instruct the question rewriting network to rewrite the question samples. Optionally, the question rewriting prompt words can include question rewriting methods, and the question rewriting methods can be adjusting the questioning method, adjusting the word order structure, synonym replacement, voice conversion, adjusting the logical relationship, etc.
[0089] Optionally, the network structure of the question rewriting network can adopt the network structure of a large language model. The input of the question rewriting network is the question samples and the question rewriting prompt words, and the output is the question rewriting samples.
[0090] Optionally, before using the question rewriting network, the computer device can train the question rewriting network through text samples and text rewriting true values. In a possible implementation manner, the computer device inputs the text samples and the text rewriting prompt words into the question rewriting network to obtain the text rewriting results output by the question rewriting network, and thus trains the question rewriting network according to the mean square error loss between the text rewriting results and the text rewriting true values.
[0091] In another possible implementation manner, the computer device can also call the GPT interface and use the language generation ability of GPT to rewrite the question samples to obtain the question rewriting samples.
[0092] Step 502: Add the question rewriting samples to the second Q&A dataset to obtain the updated second Q&A dataset.
[0093] In a possible implementation manner, after obtaining the question rewriting samples corresponding to the question samples, the computer device can add the question rewriting samples to the second Q&A dataset to obtain the updated second Q&A dataset. Among them, the question rewriting samples and the question samples in the updated second Q&A dataset correspond to the same answer sample.
[0094] Optionally, considering the variability of question rewriting, to ensure the data quality of the second Q&A dataset, before adding the question rewriting samples to the second Q&A dataset, it is also necessary to first determine the rewriting quality of the question rewriting samples. Only when the rewriting quality of the question rewriting samples meets the quality requirements, the computer device will add the question rewriting samples to the second Q&A dataset, thereby obtaining an updated second Q&A dataset.
[0095] Optionally, the computer device can determine the rewriting quality of the question rewriting samples from multiple perspectives such as semantic similarity, syntactic correctness, question fluency, style consistency, and application scenario relevance. Optionally, the computer device can use a rewriting quality evaluation model. By inputting the question sample and the question rewriting sample into the rewriting quality evaluation model, a rewriting quality score output by the rewriting quality evaluation model can be obtained.
[0096] In a possible implementation manner, the computer device can determine the rewriting quality of the question rewriting samples from two perspectives: text similarity and semantic similarity. Optionally, the computer device first determines the text similarity and semantic similarity between the question sample and the question rewriting sample, and then determines the corresponding rewriting quality score of the question rewriting sample based on the text similarity and semantic similarity. Furthermore, by comparing the rewriting quality score with the quality score threshold, when the rewriting quality score corresponding to the question rewriting sample is higher than the quality score threshold, the computer device will add the question rewriting sample to the second Q&A dataset to obtain an updated second Q&A dataset.
[0097] Among them, the rewriting quality score is negatively correlated with the text similarity, and the rewriting quality score is positively correlated with the semantic similarity. Optionally, the text similarity can be calculated and determined using the edit distance (Levenshtein distance), and the semantic similarity can be measured by the cosine similarity between the question sample vector and the question rewriting sample vector.
[0098] Optionally, the formula for the rewriting quality score can be expressed as where P orig represents the question sample, P rew represents the question rewriting sample, δ(P orig , P rew ) is the text similarity calculated using the Levenshtein distance, V(P orig ) is the vector representation corresponding to the question sample, V(P rew ) is the vector representation corresponding to the question rewriting sample, and cos(V(P orig ), V(P rew )) is the cosine similarity between the two vectors, which is used to measure the semantic similarity between the question sample and the question rewriting sample.
[0099] Optionally, the quality scoring threshold can be a preset fixed value or an adjustable value determined based on the rewritten quality scores corresponding to the rewritten samples of each question. The embodiments of the present application do not limit this.
[0100] Step 503: Train the Q&A agent based on the updated second Q&A dataset.
[0101] In a possible implementation, after obtaining the updated second Q&A dataset, the computer device can use the Q&A pair samples in the updated second Q&A dataset to train the Q&A agent.
[0102] For the training process of the Q&A agent, reference can be made to the following embodiments, which will not be elaborated here.
[0103] In the above embodiments, by obtaining the question rewritten samples corresponding to the question samples through question rewriting, the sample size in the second Q&A dataset can be further increased, and the diversity of the Q&A pair samples can be improved. And by determining the rewritten quality of the question rewritten samples and only adding the question rewritten samples that meet the quality requirements to the second Q&A dataset, the sample quality of the second Q&A dataset can be guaranteed without affecting the training effect of the subsequent Q&A agent.
[0104] In some embodiments, to optimize the training effect of the Q&A agent and enable the Q&A agent to output high-quality, accurate, and easy-to-understand answers, before training the Q&A agent using the second Q&A dataset, the quality of each Q&A pair sample in the second Q&A dataset can also be evaluated first to filter out the Q&A pair samples with poor quality.
[0105] Optionally, the process of filtering the second Q&A dataset through quality evaluation may include the following steps:
[0106] Step 601: Input each Q&A pair sample in the second Q&A dataset into the quality evaluation network to obtain the Q&A pair evaluation values corresponding to each Q&A pair sample output by the quality evaluation network. The i-th Q&A pair evaluation value is used to characterize the Q&A pair quality of the i-th Q&A pair sample.
[0107] In a possible implementation, the computer device can input each Q&A pair sample in the second Q&A dataset into the quality evaluation network, and the quality evaluation network outputs the Q&A pair evaluation values corresponding to each Q&A pair sample.
[0108] Optionally, the network structure of the quality evaluation network can adopt the network structure of a large language model. The input of the quality evaluation model is the Q&A pair sample and the quality evaluation prompt, and the output is the Q&A pair evaluation value corresponding to the Q&A pair sample.
[0109] Optionally, before using the quality assessment network, the computer device can train the quality assessment network by using text samples and the text assessment truth values manually labeled for the text samples. In a possible implementation, the computer device inputs the text samples and quality assessment prompt words into the quality assessment network to obtain the text assessment values output by the quality assessment network, and then trains the quality assessment network according to the cross-entropy loss between the text assessment truth values and the text assessment values.
[0110] Optionally, the i-th Q&A pair assessment value is used to characterize the Q&A pair quality of the i-th Q&A pair sample. Optionally, in order to measure the sample quality of the Q&A pair sample from multiple quality assessment dimensions, the i-th Q&A pair assessment value may include at least one dimension assessment value, and the quality assessment dimension corresponding to the dimension assessment value may include one of question logic, answer accuracy, language fluency, time sequence clarity, and reasoning complexity.
[0111] Among them, question logic is used to measure whether the expression of the question sample is clear and reasonable, and whether it can clearly convey the content and intention of the question. Answer accuracy is used to measure whether the answer sample accurately answers the question and whether it provides accurate information. Language fluency is used to measure whether the expressions of the question sample and the answer sample are natural and smooth, and whether they conform to the expression habits of natural language. Time sequence clarity is used to measure whether the time sequence involved in the question sample and the answer sample is clear and reasonable, and whether it can clearly express the sequence of events. Reasoning complexity is used to measure whether the difficulty of the question sample is appropriate and whether it can appropriately exercise the reasoning ability of the Q&A agent.
[0112] In a possible implementation, in order to obtain the dimension assessment values of the Q&A pair samples under different quality assessment dimensions, the computer device can input each Q&A pair sample in the second Q&A dataset and the dimension assessment prompt words corresponding to each quality assessment dimension into the quality assessment network, so as to obtain the dimension assessment values of each Q&A pair sample under each quality assessment dimension output by the quality assessment network.
[0113] Among them, the dimension assessment prompt word corresponding to question logic can be "Please evaluate this Q&A pair sample from the perspective of question logic"; the dimension assessment prompt word corresponding to answer accuracy can be "Please evaluate this Q&A pair sample from the perspective of answer accuracy"; the dimension assessment prompt word corresponding to language fluency can be "Please evaluate this Q&A pair sample from the perspective of language fluency"; the dimension assessment prompt word corresponding to time sequence clarity can be "Please evaluate this Q&A pair sample from the perspective of time sequence clarity"; the dimension assessment prompt word corresponding to reasoning complexity can be "Please evaluate this Q&A pair sample from the perspective of reasoning complexity".
[0114] Optionally, the dimension evaluation values under each quality evaluation dimension can be any value between 0 and 1. For example, the evaluation value of the i-th question-answer pair can be expressed as [0.3, 0.7, 0.4, 0.9, 0.8].
[0115] Optionally, in order to obtain the dimension evaluation values of the question-answer pair samples under the five quality evaluation dimensions, a fully connected layer and an activation function can also be added to the last output layer of the large language model to construct a quality evaluation network. The fully connected layer contains 5 output heads, which are respectively used to output the dimension evaluation values under the 5 quality evaluation dimensions; the activation function is used to normalize the probability distribution output by the fully connected layer, so as to obtain the dimension evaluation values under each quality evaluation dimension between 0 and 1.
[0116] Step 602, filter the second question-answer data set based on the question-answer pair evaluation values and evaluation thresholds corresponding to each question-answer pair sample to obtain the filtered second question-answer data set.
[0117] In a possible implementation manner, after determining the question-answer pair evaluation values corresponding to each question-answer pair sample, the computer device can compare the question-answer pair evaluation values with the evaluation threshold, and filter the second question-answer data set according to the comparison result, so as to obtain the filtered second question-answer data set.
[0118] Among them, the evaluation threshold can be a preset fixed value, or an adjustable value determined based on the question-answer pair evaluation values corresponding to each question-answer pair sample. The embodiments of the present application do not limit this.
[0119] Optionally, when the i-th question-answer pair evaluation value corresponding to the i-th question-answer pair sample is lower than the evaluation threshold, the computer device can filter out the i-th question-answer pair sample from the second question-answer data set, so as to obtain the filtered second question-answer data set.
[0120] Optionally, when the i-th question-answer pair evaluation value includes the dimension evaluation values under multiple quality evaluation dimensions, the computer device can compare each dimension evaluation value in the i-th question-answer pair evaluation value with the evaluation threshold, and when there is a dimension evaluation value in the i-th question-answer pair evaluation value that is lower than the evaluation threshold, filter out the i-th question-answer pair sample corresponding to the i-th question-answer pair evaluation value from the second question-answer data set, so as to obtain the filtered second question-answer data set.
[0121] Among them, different quality evaluation dimensions correspond to their respective evaluation thresholds. Optionally, the evaluation thresholds corresponding to different quality evaluation dimensions can be the same or different. The embodiments of the present application do not limit this. For example, the evaluation thresholds corresponding to different quality evaluation dimensions are all 0.6; for another example, the evaluation thresholds for question logic and answer accuracy are 0.8, and the evaluation thresholds for language fluency, time sequence clarity, and reasoning complexity are 0.6.
[0122] Step 603: Train a Q&A agent based on the filtered second Q&A dataset.
[0123] In a possible implementation, after obtaining the filtered second Q&A dataset, the computer device can use the Q&A pair samples in the filtered second Q&A dataset to train the Q&A agent.
[0124] For the training process of the Q&A agent, reference can be made to the following embodiments, which will not be elaborated here.
[0125] In the above embodiments, by evaluating the quality of each Q&A pair sample through a quality evaluation network, the Q&A pair evaluation value corresponding to each Q&A pair sample is obtained, and the Q&A pair samples with Q&A pair evaluation values lower than the evaluation threshold are filtered out from the second Q&A dataset, which can improve the data quality of the second Q&A dataset and optimize the training effect of the subsequent Q&A agent.
[0126] Moreover, by evaluating the quality of the Q&A pair samples from multiple quality evaluation dimensions, a more comprehensive Q&A pair evaluation value for each Q&A pair sample can be obtained, thereby improving the comprehensiveness and accuracy of the quality evaluation.
[0127] In some embodiments, in order to improve the answer prediction efficiency and accuracy of the Q&A agent, answer prediction can be achieved through the cooperation of multiple agents. At the same time, during the training process, the result scoring of the answer prediction results can be increased, and only the answer prediction results with higher quality are used for Q&A loss calculation, thereby improving the training effect of the Q&A agent.
[0128] Optionally, the training process of the Q&A agent may include the following steps:
[0129] Step 231: Input the question samples included in the Q&A pair samples in the second Q&A dataset into the Q&A agent to obtain the answer prediction results corresponding to the question samples output by the Q&A agent.
[0130] In a possible implementation, after obtaining the second Q&A dataset, the computer device can input the question samples included in the Q&A pair samples in the second Q&A dataset into the Q&A agent, so that the Q&A agent predicts the answer to the question sample and outputs the answer prediction result corresponding to the question sample.
[0131] Optionally, the Q&A agent may include a first large language model and a classification network, where the output end of the large language model is connected to the input end of the classification network, that is, the output data of the large language model is the input data of the classification network.
[0132] Among them, the classification heads in the classification network correspond to different candidate answers, and the candidate answers are entities or times in the temporal knowledge graph. That is, the classification heads in the classification network correspond to entities or times in the temporal knowledge graph, and the entity can be a subject or an object.
[0133] In a possible implementation manner, the computer device inputs the question sample included in the question-and-answer pair sample in the second question-and-answer dataset into the first large language model in the question-and-answer agent, so as to obtain the hidden state vector output by the first large language model, and then inputs the hidden state vector into the classification network, and then the answer prediction result corresponding to the question sample output by the classification network can be obtained.
[0134] Optionally, the classification network may include a Multilayer Perceptron (MLP) and an activation function. Among them, the weight matrix of the multilayer perceptron is used to classify the answers. The input dimension of the weight matrix is the same as the vector dimension of the hidden state vector output by the first large language model, and the output dimension of the weight matrix is the same as the number of candidate answers.
[0135] In a possible implementation manner, the computer device inputs the hidden state vector into the multilayer perceptron, outputs the prediction scores corresponding to each candidate answer through the multilayer perceptron, and then after normalization processing by the activation function, the probability distribution corresponding to each candidate answer can be obtained, and then the classification network outputs the candidate answer with the highest probability value as the answer prediction result.
[0136] Optionally, the output process of the answer prediction result can be expressed as Y pred = argmax c∈C (softmax(W × h ∈d + b)) c , where Y pred represents the probability value corresponding to the answer prediction result, C is the set of candidate answers, c is the candidate answer, softmax represents the activation function, W represents the weight matrix of the MLP layer, its dimension is |c| × d, |C| is the number of candidate answers, d is the vector dimension of the hidden state vector, and b is the bias vector.
[0137] Optionally, in order to improve the question understanding ability and output efficiency of the question-and-answer agent, the answer prediction of the temporal question can be realized by means of multi-agent collaboration. Optionally, before applying the question-and-answer agent to output the answer prediction result, the question understanding agent can also be used to understand and analyze the question sample, and the label agent can be used to classify the question sample label, so as to jointly use the question analysis result, the question type label and the question sample as the input of the question-and-answer agent.
[0138] Among them, the problem understanding agent is responsible for understanding and analyzing the semantics hidden in the problem samples, such as parsing the logical chain in complex problems; the tagging agent is responsible for characterizing the portrait tags of the problem samples through context learning. For example, the tagging agent will give tags such as "the answer is an entity", "before / after", etc.
[0139] In a possible implementation, the computer device inputs the problem samples included in the Q&A pair samples in the second Q&A dataset into the problem understanding agent. Through the problem understanding agent's understanding and analysis of the problem samples, the problem analysis results corresponding to the problem samples output by the problem understanding agent can be obtained. At the same time, the computer device inputs the problem samples included in the Q&A pair samples in the second Q&A dataset into the tagging agent. Through the tagging agent's classification of the problem samples by tags, the problem type tags corresponding to the problem samples output by the tagging agent can be obtained.
[0140] Optionally, the network structures of both the problem understanding agent and the tagging agent can adopt the network structure of the large language model. In a possible implementation, the computer device can train the problem understanding agent through text samples and the corresponding text analysis ground truth. And train the tagging agent through text samples and the corresponding type tag ground truth.
[0141] In a possible implementation, the computer device can also call the GPT interface. By inputting the problem samples and problem understanding prompt words into the GPT network, the problem analysis results corresponding to the problem samples output by the GPT network can be obtained. Similarly, by inputting the problem samples and tag classification prompt words into the GPT network, the problem type tags corresponding to the problem samples output by the GPT network can be obtained.
[0142] Furthermore, after obtaining the problem analysis results and problem type tags corresponding to the problem samples, the computer device can input the problem samples included in the Q&A pair samples in the second Q&A dataset, the problem analysis results corresponding to the problem samples, and the problem type tags corresponding to the problem samples into the Q&A agent, so as to obtain the answer prediction results corresponding to the problem samples output by the Q&A agent.
[0143] Step 232, based on the answer prediction result corresponding to the problem sample and the answer sample, determine the Q&A loss through the mean squared error loss function.
[0144] In a possible implementation, after obtaining the answer prediction result corresponding to the problem sample, the computer device can calculate the Q&A loss between the answer prediction result and the answer sample through the mean squared error (MSE) loss function, so as to train the Q&A agent through this Q&A loss.
[0145] Optionally, to improve the accuracy of the Q&A loss, the computer device can also perform a reverse check on the answer prediction result, score the answer prediction result through a result scoring network, and then screen the answer prediction results used for loss calculation according to the result scores corresponding to the answer prediction results.
[0146] In a possible implementation, the computer device inputs the answer prediction result corresponding to the question sample into the result scoring network, scores the answer prediction result through the result scoring network, and obtains the result score and the scoring reason corresponding to the answer prediction result output by the result scoring network. Thus, when the result score is not less than the score threshold, the computer device determines the Q&A loss based on the answer prediction result corresponding to the question sample and the answer sample through the mean squared error loss function.
[0147] When the result score is less than the score threshold, the computer device does not use this answer prediction result for loss calculation. Instead, it inputs the question sample and the scoring reason corresponding to the answer prediction result into the Q&A agent, and the Q&A agent re-predicts the answer for the question sample, thereby obtaining the answer prediction result corresponding to the question sample re-output by the Q&A agent.
[0148] Exemplarily, the result score corresponding to the answer prediction result can be any score between 1 and 10, and the score threshold can be set to 5. Thus, for the answer prediction results with a result score lower than 5, they are not used for loss calculation.
[0149] Optionally, to improve the scoring accuracy of the result scoring network, the computer device can train the result scoring network based on positive and negative answer samples and the true result scores corresponding to each positive and negative answer sample.
[0150] Optionally, the network structure of the result scoring network can also adopt the network structure of a large language model. In a possible implementation, the computer device can input the answer prediction result and the result scoring prompt word into the result scoring network, thereby obtaining the result score and the scoring reason output by the result scoring network.
[0151] Step 233: Train the Q&A agent based on the Q&A loss.
[0152] In a possible implementation, the computer device trains the Q&A agent based on the Q&A loss. Thus, after multiple rounds of training iterations, the trained Q&A agent can be obtained and applied to the Q&A scenario.
[0153] In the above embodiments, during the process of training the Q&A agent, through multi-agent collaborative reasoning, first, the problem understanding agent outputs the problem analysis result corresponding to the problem sample, and the label agent outputs the problem type label corresponding to the problem sample. Then, the problem sample, the problem analysis result, and the problem type label are jointly input into the Q&A agent. The Q&A agent combines the problem analysis result and the problem type label to predict the answer for the problem sample, which can enhance the understanding and reasoning ability of the Q&A agent and optimize the training effect of the Q&A agent.
[0154] In addition, before calculating the loss, by scoring the answer prediction result, for the answer prediction result with a result score lower than the score threshold, it is not used for calculating the Q&A loss. Instead, the problem sample and the scoring reason are input into the Q&A agent, and the Q&A agent re-predicts the answer, which can improve the accuracy of loss calculation and further optimize the training effect of the Q&A agent.
[0155] Please refer to Figure 8 , which shows a schematic diagram of the training process of the Q&A agent provided by another exemplary embodiment of the present application. Next, the training method of the Q&A agent proposed in the present application will be described in combination with Figure 8 and the above embodiments.
[0156] Step 1, obtain at least two temporal fact groups.
[0157] In a possible implementation manner, the computer device can extract the fact information with temporal characteristics from the knowledge base according to the keywords representing the time characteristics, and then represent the fact information in the form of temporal fact groups. Optionally, the temporal fact group can be in the form of a quadruple, expressed as (subject, relationship, object, time point); or it can be in the form of a quintuple, expressed as (subject, relationship, object, start time, end time).
[0158] Step 2, determine the fact association relationship between at least two temporal fact groups, and the fact association relationship includes at least one of a subject association relationship, an object association relationship, and a time association relationship.
[0159] To improve the temporality and relevance of the data, the computer device can determine the fact association relationship between at least two temporal fact groups. Optionally, when at least two temporal fact groups correspond to the same fact subject, it can be determined that at least two temporal fact groups have a subject association relationship; when at least two temporal fact groups correspond to the same fact object, it can be determined that at least two temporal fact groups have an object association relationship; when there is a time logical relationship between at least two temporal fact groups, it can be determined that at least two temporal fact groups have a time association relationship.
[0160] Step 3: Construct a temporal knowledge graph based on at least two groups of temporal facts and the factual association relationships between at least two groups of temporal facts.
[0161] Further, after determining the factual association relationships corresponding to each group of temporal facts, the computer device can construct a temporal knowledge graph based on at least two groups of temporal facts and the factual association relationships between at least two groups of temporal facts.
[0162] Step 4: Input the temporal knowledge graph and the question generation prompt words into a large language model, and output the first Q&A dataset 801 through the large language model.
[0163] Before using the question generator to generate Q&A pair samples, in order to enable the question generator to learn the question generation ability of the large language model and improve the sample quality, in a possible implementation, the computer device first inputs the temporal knowledge graph and the question generation prompt words into the large language model (such as GPT), so as to obtain the Q&A pair samples output by the large language model, which constitute the first Q&A dataset. The first Q&A dataset contains high-quality but small-number Q&A pair samples.
[0164] Among them, the question generation prompt words can be used to indicate the generation of Q&A pair samples for a single group of temporal facts, that is, simple Q&A pairs; or they can be used to indicate the generation of Q&A pair samples for at least two groups of temporal facts, that is, complex Q&A pairs.
[0165] Step 5: Train the question generator 802 based on the temporal knowledge graph and the first Q&A dataset 801.
[0166] After obtaining the first Q&A dataset, the computer device can train the question generator based on the temporal knowledge graph and the first Q&A dataset, so that the question generator learns the question generation ability of the large language model.
[0167] In a possible implementation, the computer device inputs the temporal knowledge graph, the question generation prompt words, and the first Q&A dataset into the question generator. The question generator learns the first Q&A dataset and outputs predicted Q&A pairs. Then, according to the Q&A pair generation loss between the predicted Q&A pairs and the true values of the Q&A pairs in the first Q&A dataset, the question generator is trained. Through multiple backpropagation updates, the trained question generator can be obtained.
[0168] Step 6: Input the temporal knowledge graph and the question generation prompt words into the question generator 802, and output the second Q&A dataset 803 through the question generator 802. Among them, the data volume of the second Q&A dataset 803 is higher than that of the first Q&A dataset 801.
[0169] After the question generator is trained, the computer device can input the temporal knowledge graph and question generation prompt words into the question generator, and output the second Q&A dataset through the question generator.
[0170] Optionally, the data volume of the second Q&A dataset is higher than that of the first Q&A dataset. Moreover, since the second Q&A dataset is generated by the model, compared with the Q&A dataset generated by manual templates, the second Q&A dataset can cover a larger and more complex semantic space and has higher data quality.
[0171] Step 7: Input the question samples included in the Q&A pair samples in the second Q&A dataset 803 into the question rewriting network 804 to obtain the question rewriting samples output by the question rewriting network 804.
[0172] To further increase the diversity of the Q&A pair samples, after obtaining the second Q&A dataset, the computer device can obtain the question samples included in the Q&A pair samples from the second Q&A dataset, and then input the question samples and question rewriting prompt words into the question rewriting network to obtain the question rewriting samples output by the question rewriting network.
[0173] Optionally, the question rewriting samples are texts that are semantically similar to the question samples but have different expressions. Optionally, the question rewriting prompt words are used to instruct the question rewriting network to rewrite the question samples. Optionally, the question rewriting prompt words can include question rewriting methods, and the question rewriting methods can be adjusting the questioning method, adjusting the word order structure, synonym replacement, voice conversion, adjusting the logical relationship, etc.
[0174] Step 8: Add the question rewriting samples to the second Q&A dataset 803 to obtain the updated second Q&A dataset 803.
[0175] After obtaining the question rewriting samples corresponding to the question samples, the computer device can add the question rewriting samples to the second Q&A dataset to obtain the updated second Q&A dataset. Among them, the question rewriting samples and the question samples in the updated second Q&A dataset correspond to the same answer sample.
[0176] In addition, considering the variability of question rewriting, to ensure the data quality of the second Q&A dataset, before adding the question rewriting samples to the second Q&A dataset, the computer device also needs to first determine the rewriting quality of the question rewriting samples. Thus, only when the rewriting quality of the question rewriting samples meets the quality requirements, the computer device adds the question rewriting samples to the second Q&A dataset to obtain the updated second Q&A dataset.
[0177] Regarding the method for determining the rewriting quality of rewritten samples of questions, in one possible implementation, the computer device can first determine the text similarity and semantic similarity between the question samples and the rewritten question samples, and then determine the rewriting quality score corresponding to the rewritten question samples based on the text similarity and semantic similarity. Furthermore, by comparing the rewriting quality score with the quality score threshold, if the rewriting quality score corresponding to the rewritten question sample is higher than the quality score threshold, the rewritten question sample is added to the second Q&A dataset to obtain the updated second Q&A dataset.
[0178] Step 9: Input the Q&A pair samples in the second Q&A dataset 803 into the quality evaluation network 805 to obtain the Q&A pair evaluation values corresponding to the Q&A pair samples output by the quality evaluation network 805.
[0179] To further improve the data quality of the second Q&A dataset, after question rewriting, the computer device can also perform quality evaluation on each Q&A pair sample in the second Q&A dataset to filter out the Q&A pair samples with low quality.
[0180] In one possible implementation, the computer device can input each Q&A pair sample in the second Q&A dataset into the quality evaluation network, and the quality evaluation network outputs the Q&A pair evaluation values corresponding to each Q&A pair sample.
[0181] Optionally, the i-th Q&A pair evaluation value is used to represent the Q&A pair quality of the i-th Q&A pair sample. Optionally, to measure the sample quality of the Q&A pair sample from multiple quality evaluation dimensions, the i-th Q&A pair evaluation value can include at least one dimension evaluation value, and the quality evaluation dimension corresponding to the dimension evaluation value can include one of question logic, answer accuracy, language fluency, time sequence clarity, and reasoning complexity.
[0182] Step 10: Filter the second Q&A dataset 803 based on the Q&A pair evaluation value corresponding to the Q&A pair sample and the evaluation threshold to obtain the filtered second Q&A dataset 803.
[0183] In one possible implementation, after determining the Q&A pair evaluation values corresponding to each Q&A pair sample, the computer device can compare the Q&A pair evaluation values with the evaluation threshold, and filter the second Q&A dataset according to the comparison result to obtain the filtered second Q&A dataset.
[0184] Optionally, when the evaluation value of the i-th question-answer pair includes dimension evaluation values under multiple quality evaluation dimensions, the computer device can compare each dimension evaluation value in the evaluation value of the i-th question-answer pair with the evaluation threshold, and filter out the i-th question-answer pair sample corresponding to the evaluation value of the i-th question-answer pair from the second question-answer data set when there is a dimension evaluation value in the evaluation value of the i-th question-answer pair that is lower than the evaluation threshold, so as to obtain a filtered second question-answer data set.
[0185] Step 11: Input the question sample included in the question-answer pair sample in the second question-answer data set 803 into the question understanding agent 806, and obtain the question analysis result corresponding to the question sample output by the question understanding agent 806.
[0186] After question rewriting update and quality evaluation filtering, a second question-answer data set for training the question-answer agent can be obtained. In order to improve the understanding efficiency of the question-answer agent for question samples, in a possible implementation manner, a multi-agent collaboration method can be adopted. Before inputting the question sample into the question-answer agent, the question sample is first understood and analyzed and labeled by the question understanding agent and the label agent.
[0187] Optionally, the computer device inputs the question sample included in the question-answer pair sample in the second question-answer data set into the question understanding agent, and through the understanding and analysis of the question sample by the question understanding agent, the question analysis result corresponding to the question sample output by the question understanding agent can be obtained.
[0188] Step 12: Input the question sample included in the question-answer pair sample in the second question-answer data set 803 into the label agent 807, and obtain the question type label corresponding to the question sample output by the label agent 807.
[0189] At the same time, the computer device inputs the question sample included in the question-answer pair sample in the second question-answer data set into the label agent, and through the label classification of the question sample by the label agent, the question type label corresponding to the question sample output by the label agent can be obtained.
[0190] Step 13: Input the question sample included in the question-answer pair sample in the second question-answer data set 803, the question analysis result corresponding to the question sample, and the question type label corresponding to the question sample into the question-answer agent 808, and obtain the answer prediction result 809 corresponding to the question sample output by the question-answer agent 808.
[0191] After obtaining the question analysis result and the question type label corresponding to the question sample, the computer device can input the question sample included in the question-answer pair sample in the second question-answer data set, the question analysis result corresponding to the question sample, and the question type label corresponding to the question sample into the question-answer agent, so as to obtain the answer prediction result corresponding to the question sample output by the question-answer agent.
[0192] In a possible implementation, the computer device inputs the question samples included in the question-and-answer pair samples in the second question-and-answer dataset, the question analysis results corresponding to the question samples, and the question type tags corresponding to the question samples into the first large language model in the question-and-answer agent, so as to obtain the hidden state vector output by the first large language model, and then inputs the hidden state vector into the classification network, and the answer prediction result corresponding to the question sample output by the classification network can be obtained.
[0193] Step 14: Input the answer prediction result 809 corresponding to the question sample into the result scoring network to obtain the result score and the scoring reason corresponding to the answer prediction result 809 output by the result scoring network.
[0194] To improve the accuracy of the question-and-answer loss, the computer device can also perform a reverse check on the answer prediction result, score the answer prediction result through the result scoring network, and thus screen the answer prediction results used for loss calculation according to the result scores corresponding to the answer prediction results.
[0195] In a possible implementation, the computer device inputs the answer prediction result corresponding to the question sample into the result scoring network, scores the answer prediction result through the result scoring network, and obtains the result score and the scoring reason corresponding to the answer prediction result output by the result scoring network.
[0196] Step 15: When the result score is not less than the score threshold, based on the answer prediction result 809 corresponding to the question sample and the answer sample, determine the question-and-answer loss through the mean square error loss function.
[0197] Furthermore, when the result score is not less than the score threshold, the computer device determines the question-and-answer loss based on the answer prediction result corresponding to the question sample and the answer sample through the mean square error loss function.
[0198] Step 16: When the result score is less than the score threshold, input the question sample and the scoring reason corresponding to the answer prediction result into the question-and-answer agent 808 to obtain the answer prediction result 809 corresponding to the question sample re-output by the question-and-answer agent 808.
[0199] When the result score is less than the score threshold, the computer device does not use this answer prediction result for loss calculation, but inputs the question sample and the scoring reason corresponding to the answer prediction result into the question-and-answer agent, and re-predicts the answer to the question sample through the question-and-answer agent, so as to obtain the answer prediction result corresponding to the question sample re-output by the question-and-answer agent.
[0200] Step 17: Train the question-and-answer agent 808 based on the question-and-answer loss.
[0201] In a possible implementation, the computer device trains the question-and-answer agent based on the question-and-answer loss. After multiple rounds of training iterations, the trained question-and-answer agent can be obtained and applied to the question-and-answer scenario.
[0202] Combining the above embodiments, the beneficial effects of the training method for the question-and-answer agent provided by the embodiments of the present application at least include the following points:
[0203] In the embodiments of the present application, the question generator learns the question generation ability of the large language model, so as to use the question generator to generate the second question-and-answer data set. While reducing the resource consumption of model interface calls, the diversity of data samples is improved, so that the second question-and-answer data set can cover the complex semantic space.
[0204] Please refer to Figure 9 , which shows a comparison diagram of the dimensionality reduction representation of the data set provided by an exemplary embodiment of the present application. After performing embedding vector representation and dimensionality reduction representation on the existing data set and the second question-and-answer data set in the present application, the distribution effects of the dimensionality reduction representations corresponding to the two data sets in the semantic space can be obtained. Obviously, the second question-and-answer data set in the present application can cover more semantic spaces and is more evenly distributed, while the existing data set has an obvious separation trend.
[0205] Moreover, when constructing the second question-and-answer data set in the present application, by learning the temporal fact groups and the fact association relationships between the temporal fact groups in the temporal knowledge graph through the question generator, the semantic information in the second question-and-answer data set can be further enriched. As shown in Table 1, it shows the quantitative comparison data of the semantic richness between the second question-and-answer data set and the existing data set in the present application. Obviously, the second question-and-answer data set in the present application has more temporal relationship numbers and average vocabulary.
[0206] Table 1
[0207] Dataset Temporal relationship number Average vocabulary size CronQuestions 5 0.144 MultiQA 55 0.031 The dataset in this application 677 1.915
[0208] At the same time, in the present application, through multi-agent collaborative reasoning and scoring verification of the answer prediction results, the training effect of the question-and-answer agent is optimized, and the accuracy of answer prediction of the question-and-answer agent is improved. As shown in Table 2, it shows the comparison data of the result hit rate between the question-and-answer agent in the present application and the question-and-answer model in the prior art. Obviously, the proportion of the answer prediction results output by the question-and-answer agent in the present application that are the candidate answers with the highest probability reaches 50.8%, and the proportion of the answer prediction results output by the question-and-answer agent in the present application that belong to the top 10 candidate answers in terms of probability reaches 71.3%, showing a significant improvement compared with the question-and-answer model in the prior art.
[0209] Table 2
[0210]
[0211] In addition, the question generator trained in this application can quickly construct question-and-answer pairs based on a given knowledge graph, so that after simple fine-tuning, it can be applied to data production tasks in downstream scenarios such as intelligent customer service Q&A and personalized recommendation, thereby providing high-quality and reliable sample data. For example, by inputting content recommendation prompts and the knowledge graph into the question generator, recommended content question-and-answer pairs output by the question generator can be obtained.
[0212] Please refer to Figure 10 , which shows a structural block diagram of a training device for a Q&A intelligent agent provided by an exemplary embodiment of this application. The device includes:
[0213] A first training module 1001, configured to train a question generator based on a temporal knowledge graph and a first question-and-answer data set corresponding to the temporal knowledge graph. The temporal knowledge graph includes at least two temporal fact groups, and each temporal fact group includes fact information with temporal characteristics;
[0214] A first output module 1002, configured to input the temporal knowledge graph and question generation prompts into the question generator, and output a second question-and-answer data set through the question generator. The question generation prompts are used to indicate the generation of question-and-answer pair samples for at least one temporal fact group, and the data volume of the second question-and-answer data set is higher than that of the first question-and-answer data set;
[0215] A second training module 1003, configured to train a Q&A intelligent agent based on the second question-and-answer data set. The Q&A intelligent agent is used to output an answer corresponding to the temporal question based on the input temporal question, and the temporal question refers to a question with a temporal reasoning logic.
[0216] Optionally, the second training module 1003 includes:
[0217] A question rewriting unit, configured to input a question sample included in the question-and-answer pair sample in the second question-and-answer data set and question rewriting prompts into a question rewriting network, and obtain a question rewriting sample output by the question rewriting network;
[0218] A data update unit, configured to add the question rewriting sample to the second question-and-answer data set to obtain an updated second question-and-answer data set;
[0219] A first training unit, configured to train the Q&A intelligent agent based on the updated second question-and-answer data set.
[0220] Optionally, the data update unit is configured to:
[0221] Determine the rewriting quality of the question rewriting sample;
[0222] When the rewriting quality of the problem rewriting sample meets the quality requirements, add the problem rewriting sample to the second Q&A dataset to obtain the updated second Q&A dataset.
[0223] Optionally, the data updating unit is configured to:
[0224] Determine the text similarity and semantic similarity between the problem sample and the problem rewriting sample;
[0225] Based on the text similarity and the semantic similarity, determine the rewriting quality score corresponding to the problem rewriting sample, where the rewriting quality score is negatively correlated with the text similarity and positively correlated with the semantic similarity;
[0226] The data updating unit is further configured to:
[0227] When the rewriting quality score corresponding to the problem rewriting sample is higher than the quality score threshold, add the problem rewriting sample to the second Q&A dataset to obtain the updated second Q&A dataset.
[0228] Optionally, the second training module 1003 includes:
[0229] A quality evaluation unit, configured to input each Q&A pair sample in the second Q&A dataset into a quality evaluation network to obtain the Q&A pair evaluation value corresponding to each Q&A pair sample output by the quality evaluation network, where the i-th Q&A pair evaluation value is used to characterize the Q&A pair quality of the i-th Q&A pair sample;
[0230] A data filtering unit, configured to filter the second Q&A dataset based on the Q&A pair evaluation value corresponding to each Q&A pair sample and an evaluation threshold to obtain a filtered second Q&A dataset;
[0231] A second training unit, configured to train the Q&A agent based on the filtered second Q&A dataset.
[0232] Optionally, the i-th Q&A pair evaluation value includes at least one dimension evaluation value, and the quality evaluation dimension corresponding to the dimension evaluation value includes one of problem logic, answer accuracy, language fluency, time sequence clarity, and reasoning complexity;
[0233] The data filtering unit is configured to:
[0234] When there is a dimension evaluation value in the i-th Q&A pair evaluation value that is lower than the evaluation threshold, filter out the i-th Q&A pair sample corresponding to the i-th Q&A pair evaluation value from the second Q&A dataset to obtain the filtered second Q&A dataset.
[0235] Optionally, the quality assessment unit is configured to:
[0236] Input each Q&A pair sample in the second Q&A dataset and the dimension evaluation prompt words corresponding to each quality assessment dimension into the quality assessment network, and obtain the dimension evaluation values of each Q&A pair sample output by the quality assessment network under each quality assessment dimension.
[0237] Optionally, the device further includes:
[0238] A data acquisition module, configured to acquire the at least two temporal fact groups;
[0239] A relationship determination module, configured to determine the fact association relationship between the at least two temporal fact groups, where the fact association relationship includes at least one of a subject association relationship, an object association relationship, and a time association relationship;
[0240] A knowledge graph generation module, configured to construct the temporal knowledge graph based on the at least two temporal fact groups and the fact association relationship between the at least two temporal fact groups.
[0241] Optionally, the relationship determination module is configured to:
[0242] When the at least two temporal fact groups correspond to the same fact subject, determine that the at least two temporal fact groups have the subject association relationship;
[0243] When the at least two temporal fact groups correspond to the same fact object, determine that the at least two temporal fact groups have the object association relationship;
[0244] When there is a time logic relationship between the at least two temporal fact groups, determine that the at least two temporal fact groups have the time association relationship.
[0245] Optionally, the first output module 1002 is configured to:
[0246] Input the temporal knowledge graph and the first question generation prompt words into the question generator, and output a second Q&A dataset of the first type through the question generator, where the first question generation prompt words are used to indicate generating the Q&A pair samples for a single temporal fact group;
[0247] Input the temporal knowledge graph and the second question generation prompt words into the question generator, and output a second Q&A dataset of the second type through the question generator, where the second question generation prompt words are used to indicate generating the Q&A pair samples for at least two temporal fact groups that have the fact association relationship;
[0248] Based on the second Q&A dataset of the first type and the second Q&A dataset of the second type, the second Q&A dataset is obtained.
[0249] Optionally, the second training module 1003 includes:
[0250] A result output unit, configured to input the question sample included in the Q&A pair sample in the second Q&A dataset into the Q&A agent, and obtain the answer prediction result corresponding to the question sample output by the Q&A agent;
[0251] A loss determination unit, configured to determine the Q&A loss through a mean squared error loss function based on the answer prediction result corresponding to the question sample and the answer sample;
[0252] A third training unit, configured to train the Q&A agent based on the Q&A loss.
[0253] Optionally, the Q&A agent includes a first large language model and a classification network, and the output data of the first large language model is the input data of the classification network;
[0254] The result output unit is configured to:
[0255] Input the question sample included in the Q&A pair sample in the second Q&A dataset into the first large language model in the Q&A agent, and obtain the hidden state vector output by the first large language model;
[0256] Input the hidden state vector into the classification network, and obtain the answer prediction result corresponding to the question sample output by the classification network, where the classification head in the classification network corresponds to the entity or time in the temporal knowledge graph.
[0257] Optionally, the device further includes:
[0258] A second output module, configured to input the question sample included in the Q&A pair sample in the second Q&A dataset into a question understanding agent, and obtain the question analysis result corresponding to the question sample output by the question understanding agent;
[0259] A third output module, configured to input the question sample included in the Q&A pair sample in the second Q&A dataset into a label agent, and obtain the question type label corresponding to the question sample output by the label agent;
[0260] The result output unit is configured to:
[0261] Input the question sample, the question analysis result corresponding to the question sample, and the question type label corresponding to the question sample included in the second Q&A dataset into the Q&A agent to obtain the answer prediction result corresponding to the question sample output by the Q&A agent.
[0262] Optionally, the loss determination unit is configured to:
[0263] Input the answer prediction result corresponding to the question sample into the result scoring network to obtain the result score and the scoring reason corresponding to the answer prediction result output by the result scoring network;
[0264] When the result score is not less than the score threshold, determine the Q&A loss based on the answer prediction result corresponding to the question sample and the answer sample through the mean square error loss function;
[0265] The device further includes:
[0266] A fourth output module, configured to input the question sample and the scoring reason corresponding to the answer prediction result into the Q&A agent when the result score is less than the score threshold, to obtain the answer prediction result corresponding to the question sample re-output by the Q&A agent.
[0267] Optionally, the question generator is the second large language model; the Q&A pair samples in the first Q&A dataset are generated by the third large language model based on the input temporal knowledge graph and question generation prompt words.
[0268] In summary, in the embodiments of the present application, by first training a question generator based on a temporal knowledge graph and a first Q&A dataset corresponding to the temporal knowledge graph, enabling the question generator to generate high-quality Q&A pair samples, and then inputting the temporal knowledge graph and question generation prompt words into the question generator to obtain the second Q&A dataset output by the question generator, the data diversity of the Q&A dataset can be improved, so that the Q&A dataset can cover complex semantic spaces. Furthermore, using the second Q&A dataset to train the Q&A agent can optimize the training effect of the Q&A agent and improve the output quality of the Q&A agent.
[0269] It should be noted that: for the device provided in the above embodiments, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0270] Please refer toFigure 11 , which shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Specifically: The computer device 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including a random access memory 1102 and a read-only memory 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The computer device 1100 may further include a basic input / output system (Input / Output, I / O system) 1106 for facilitating the transfer of information between various components within the computer, and a mass storage device 1107 for storing an operating system 1113, application programs 1114, and other program modules 1115.
[0271] In some embodiments, the basic input / output system 1106 includes a display 1108 for displaying information and input devices 1109 such as a mouse, keyboard, etc. for user input of information. Among them, both the display 1108 and the input devices 1109 are connected to the central processing unit 1101 through an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may further include an input / output controller 1110 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1110 also provides output to a display screen, printer, or other types of output devices.
[0272] The mass storage device 1107 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer-readable medium provide non-volatile storage for the computer device 1100. That is to say, the mass storage device 1107 may include a computer-readable medium (not shown) such as a hard disk or a drive.
[0273] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several types. The above-mentioned system memory 1104 and mass storage device 1107 can be collectively referred to as memory.
[0274] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 1101. The one or more programs contain instructions for implementing the above method. The central processing unit 1101 executes the one or more programs to implement the training method of the question-and-answer intelligent agent provided by each of the above method embodiments.
[0275] According to various embodiments of the present application, the computer device 1100 can also run on a remote computer on the network through a network such as the Internet. That is, the computer device 1100 can be connected to the network 1111 through the network interface unit 1112 connected to the system bus 1105. Or rather, the network interface unit 1112 can also be used to connect to other types of networks or remote computer systems (not shown).
[0276] The embodiments of the present application also provide a computer-readable storage medium. At least one instruction is stored in the readable storage medium, and the at least one instruction is loaded and executed by a processor to implement the training method of the question-and-answer intelligent agent described in the above embodiments.
[0277] Optionally, the computer-readable storage medium may include: ROM, RAM, solid state drives (SSD), or optical discs, etc. Among them, RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).
[0278] An embodiment of the present application provides a computer program product, which includes at least one instruction stored in a computer-readable storage medium. A processor of a computer device reads the at least one instruction from the computer-readable storage medium, and the processor executes the at least one instruction, so that the computer device executes the training method of the question-and-answer intelligent agent described in the above embodiment.
[0279] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiment can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.
[0280] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for training a question-answering agent, characterized in that: The method comprises: Training a question generator based on a time series knowledge graph and a first question-answering dataset corresponding to the time series knowledge graph, wherein the time series knowledge graph includes at least two time series fact groups, each time series fact group includes fact information with time series features; Inputting the temporal knowledge graph and the question generation prompt words into the question generator, and outputting a second question-answering dataset through the question generator, wherein the question generation prompt words are used to indicate generating a question-answering pair sample for at least one temporal fact group, and the data volume of the second question-answering dataset is greater than the data volume of the first question-answering dataset; Based on the second question-answering data set, a question-answering agent is trained, and the question-answering agent is used to output answers corresponding to the time series questions based on the input time series questions, and the time series questions refer to questions with time series reasoning logic.
2. The method according to claim 1, characterized in that The step of training a question-answering agent based on the second question-answering dataset comprises: Inputting the question samples and question rewriting prompt words contained in the question-answer pair samples in the second question-answer data set into a question rewriting network to obtain question rewriting samples output by the question rewriting network; Adding the question rewriting sample to the second question and answer dataset to obtain an updated second question and answer dataset; The question-answering agent is trained based on the updated second question-answering dataset.
3. The method according to claim 2, characterized in that The step of adding the question rewriting sample to the second question and answer dataset to obtain an updated second question and answer dataset includes: determining the rewriting quality of the rewriting sample of the problem; When the rewriting quality of the question rewriting sample meets the quality requirement, the question rewriting sample is added to the second question and answer dataset to obtain the updated second question and answer dataset.
4. The method according to claim 3, characterized in that: The determining the rewriting quality of the problem rewriting sample includes: Determining textual similarity and semantic similarity between the question sample and the question rewritten sample; Based on the text similarity and the semantic similarity, determining a rewriting quality score corresponding to the problem rewriting sample, wherein the rewriting quality score is negatively correlated with the text similarity, and the rewriting quality score is positively correlated with the semantic similarity; When the rewriting quality of the question rewriting sample meets the quality requirement, adding the question rewriting sample to the second question and answer dataset to obtain the updated second question and answer dataset includes: When the rewriting quality score corresponding to the question rewriting sample is higher than the quality score threshold, the question rewriting sample is added to the second question and answer dataset to obtain the updated second question and answer dataset.
5. The method according to any one of claims 1 to 4, characterized in that: The step of training a question-answering agent based on the second question-answering dataset comprises: Input each question-answer pair sample in the second question-answer data set into a quality assessment network, and obtain a question-answer pair evaluation value corresponding to each question-answer pair sample output by the quality assessment network, wherein the i-th question-answer pair evaluation value is used to characterize the question-answer pair quality of the i-th question-answer pair sample; Based on the question-answer pair evaluation value and the evaluation threshold corresponding to each question-answer pair sample, filtering the second question-answer data set to obtain a filtered second question-answer data set; The question-answering agent is trained based on the filtered second question-answering dataset.
6. The method according to claim 5, characterized in that The evaluation value of the i-th question-answer pair includes at least one dimension evaluation value, and the quality evaluation dimension corresponding to the dimension evaluation value includes one of question logic, answer accuracy, language fluency, temporal clarity, and reasoning complexity; The filtering the second question and answer data set based on the question and answer pair evaluation value and the evaluation threshold corresponding to each question and answer pair sample to obtain the filtered second question and answer data set includes: In the case that there is a dimension evaluation value lower than the evaluation threshold in the evaluation value of the i-th question and answer pair, the i-th question and answer pair sample corresponding to the evaluation value of the i-th question and answer pair is filtered out from the second question and answer data set to obtain the filtered second question and answer data set.
7. The method according to claim 6, characterized in that The step of inputting each question-answer pair sample in the second question-answer data set into a quality assessment network to obtain a question-answer pair assessment value corresponding to each question-answer pair sample output by the quality assessment network includes: Each question-answer pair sample in the second question-answer data set and the dimension assessment prompt words corresponding to each quality assessment dimension are input into the quality assessment network to obtain the dimension assessment value of each question-answer pair sample output by the quality assessment network under each quality assessment dimension.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Acquire the at least two time series fact groups; Determine a fact association relationship between the at least two time series fact groups, wherein the fact association relationship includes at least one of a subject association relationship, an object association relationship, and a time association relationship; The temporal knowledge graph is constructed based on the at least two temporal fact groups and the fact association relationship between the at least two temporal fact groups.
9. The method according to claim 8, characterized in that The determining of the fact association relationship between the at least two time series fact groups includes: In a case where the at least two time series fact groups correspond to the same fact subject, determining that the at least two time series fact groups have the subject association relationship; In a case where the at least two time series fact groups correspond to the same fact object, determining that the at least two time series fact groups have the object association relationship; In the case that there is a temporal logical relationship between the at least two temporal fact groups, it is determined that the at least two temporal fact groups have the temporal association relationship.
10. The method according to claim 8, characterized in that The step of inputting the time series knowledge graph and the question generation prompt words into the question generator, and outputting a second question-answering data set through the question generator, comprises: Inputting the temporal knowledge graph and a first question generation prompt word into the question generator, and outputting a second question-answer data set of the first type through the question generator, wherein the first question generation prompt word is used to indicate generation of the question-answer pair sample for a single temporal fact group; Inputting the temporal knowledge graph and a second question generation prompt word into the question generator, and outputting a second question-answer data set of a second type through the question generator, wherein the second question generation prompt word is used to indicate generating the question-answer pair samples for at least two temporal fact groups, and the at least two temporal fact groups have the fact association relationship; Based on the second question and answer dataset of the first type and the second question and answer dataset of the second type, the second question and answer dataset is obtained.
11. The method according to any one of claims 1 to 10, characterized in that: The step of training a question-answering agent based on the second question-answering dataset comprises: Inputting the question sample included in the question-answer pair sample in the second question-answer data set into the question-answering agent, and obtaining the answer prediction result corresponding to the question sample output by the question-answering agent; Based on the answer prediction result corresponding to the question sample and the answer sample, determining the question-answering loss by a mean square error loss function; Based on the question-answering loss, the question-answering agent is trained.
12. The method according to claim 11, characterized in that The question-answering agent includes a first large language model and a classification network, and output data of the first large language model is input data of the classification network; The step of inputting the question sample included in the question-answer pair sample in the second question-answer data set into the question-answering agent to obtain the answer prediction result corresponding to the question sample output by the question-answering agent includes: Inputting the question samples included in the question-answer pair samples in the second question-answer data set into the first large language model in the question-answering agent to obtain a hidden state vector output by the first large language model; The hidden state vector is input into the classification network to obtain the answer prediction result corresponding to the question sample output by the classification network, and the classification head in the classification network corresponds to the entity or time in the time series knowledge graph.
13. The method according to claim 11, characterized in that The method further comprises: Inputting the question samples included in the question-answer pair samples in the second question-answer data set into the question understanding agent, and obtaining the question analysis results corresponding to the question samples output by the question understanding agent; Inputting the question sample contained in the question-answer pair sample in the second question-answer data set into the label agent, and obtaining the question type label corresponding to the question sample output by the label agent; The step of inputting the question sample included in the question-answer pair sample in the second question-answer data set into the question-answering agent to obtain the answer prediction result corresponding to the question sample output by the question-answering agent includes: The question samples contained in the question and answer pair samples in the second question and answer data set, the question analysis results corresponding to the question samples, and the question type labels corresponding to the question samples are input into the question and answer agent to obtain the answer prediction results corresponding to the question samples output by the question and answer agent.
14. The method according to claim 11, characterized in that The question-answering loss is determined by a mean square error loss function based on the answer prediction result corresponding to the question sample and the answer sample, including: Inputting the answer prediction result corresponding to the question sample into the result scoring network, and obtaining the result score and scoring reason corresponding to the answer prediction result output by the result scoring network; When the result score is not less than the score threshold, determining the question-answering loss by a mean square error loss function based on the answer prediction result corresponding to the question sample and the answer sample; The method further comprises: When the result score is less than the score threshold, the question sample and the scoring reasons corresponding to the answer prediction result are input into the question-answering agent to obtain the answer prediction result corresponding to the question sample re-output by the question-answering agent.
15. The method according to any one of claims 1 to 14, characterized in that: The question generator is the second largest language model; the question-answer pair samples in the first question-answer data set are generated by the third largest language model based on the input temporal knowledge graph and question generation prompt words.
16. A training device for a question-answering agent, characterized in that: The device comprises: A first training module, configured to train a question generator based on a time series knowledge graph and a first question-answering dataset corresponding to the time series knowledge graph, wherein the time series knowledge graph includes at least two time series fact groups, each time series fact group includes fact information having time series features; A first output module, configured to input the temporal knowledge graph and question generation prompt words into the question generator, and output a second question-answering dataset through the question generator, wherein the question generation prompt words are used to indicate generation of question-answering pair samples for at least one temporal fact group, and the data volume of the second question-answering dataset is greater than the data volume of the first question-answering dataset; The second training module is used to train a question-answering agent based on the second question-answering data set, and the question-answering agent is used to output answers corresponding to the time series questions based on the input time series questions, and the time series questions refer to questions with time series reasoning logic.
17. A computer device, characterized in that: The computer device includes a processor and a memory; the memory stores at least one instruction, and the at least one instruction is used to be executed by the processor to implement the training method of the question-answering agent as described in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the training method of a question-answering agent as described in any one of claims 1 to 15.
19. A computer program product, characterized in that The computer program product includes at least one instruction, which is stored in a computer-readable storage medium; the processor of the computer device reads the at least one instruction from the computer-readable storage medium, and the processor executes the at least one instruction, so that the computer device implements the training method of the question-answering agent as described in any one of claims 1 to 15.
Citation Information
Cited By
Training method and device of vertical domain question and answer model, equipment and storage medium
CN121117615A