Question and answer model training method and device, equipment, medium and product
By simulating the multi-round dialogue process, generating training data and training the Q&A model, the existing Q&A model lacks the questioning ability in multiple rounds of dialogue is solved, and more efficient Q&A performance is achieved.
Patent Information
- Application Number
- CN202510200027.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing question-and-answer model lacks the ability to follow up in multiple rounds of conversations, especially when the user's question description is not clear.
By obtaining the first dialogue record of the target knowledge field, N rounds of sample generation process are performed, training data is generated, and the question-and-answer model is trained using this data. This method simulates the process of ‘study’ and ‘re-question’ in multiple rounds of dialogues, and constructs diversified training data to improve the model’s questioning ability.
It significantly improves the question-and-answer model's ability to follow up in multiple rounds of conversations, enhances the model's Q&A performance in the target knowledge field, and can handle user's problems more effectively, especially when the problem description is unclear.
Smart Images

Figure CN120124744A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for training a question-and-answer model. Background Art
[0002] With the continuous development of computer technologies, language models that can be applied to different scenarios have emerged. For example, a language model can be applied to a question-and-answer scenario to answer users' questions through conversations. A language model applied to a question-and-answer scenario can also be referred to as a question-and-answer model.
[0003] Considering that the question-and-answer scenario is usually related to a specific knowledge field, in the related art, the pre-trained question-and-answer model is usually fine-tuned using the conversation records of the specific knowledge field to enhance the reasoning ability of the question-and-answer model in the specific knowledge field. However, the question-and-answer model after fine-tuning by the above method performs poorly in multi-round conversations, especially when the user's question description is unclear, it is difficult to conduct effective follow-up questions. Summary of the Invention
[0004] The present application provides a method for training a question-and-answer model. This method can improve the follow-up question ability of the question-and-answer model in multi-round conversations and enhance the question-and-answer performance of the question-and-answer model in the target knowledge field. The present application also provides an apparatus, an electronic device, a computer-readable storage medium, and a computer program product corresponding to the above method.
[0005] In a first aspect, the present application provides a method for training a question-and-answer model, the method comprising:
[0006] Obtaining a first conversation record of a target knowledge field; wherein, the first conversation record includes multi-round conversations between a user role and a customer service role for a target question in the target knowledge field, and the target reply of the customer service role for the target question in the first conversation record includes target knowledge information of the target knowledge field;
[0007] Performing N rounds of sample generation processes based on the first conversation record to obtain training data;
[0008] Training a target question-and-answer model using the training data;
[0009] Wherein, the nth round of sample generation process includes: using the target question-and-answer model to answer the nth question based on the knowledge information of the target knowledge field and the conversation context corresponding to the nth round of sample generation process to obtain the nth reply; 1≤n≤N;
[0010] When n is 1, the nth question is the first question sent by the user role in the first conversation record; when n is not 1, the nth question is a question generated by the user role model based on the (n - 1)th follow-up question, and the (n - 1)th follow-up question is generated by the customer service role model during the (n - 1)th round of sample generation;
[0011] When n is 1, the conversation context corresponding to the nth round of sample generation is the nth question; when n is not 1, the conversation context corresponding to the nth round of sample generation includes: a conversation sequence composed of n - 1 questions and n - 1 follow-up questions in the previous n - 1 rounds of sample generation and the nth question;
[0012] The condition for ending the execution of the sample generation process is that the Nth reply includes at least part of the target knowledge information.
[0013] In a second aspect, the present application provides a training device for a question-and-answer model, and the device includes:
[0014] An acquisition module, configured to acquire a first conversation record of a target knowledge domain; wherein, the first conversation record includes multiple rounds of conversations between a user role and a customer service role for answering a target question in the target knowledge domain, and the target reply of the customer service role to the target question in the first conversation record includes the target knowledge information of the target knowledge domain;
[0015] A generation module, configured to execute N rounds of sample generation processes based on the first conversation record to obtain training data;
[0016] A training module, configured to use the training data to train a target question-and-answer model;
[0017] Wherein, the nth round of sample generation process includes: using the target question-and-answer model, based on the knowledge information of the target knowledge domain and the conversation context corresponding to the nth round of sample generation process, answering the nth question to obtain the nth reply; 1 ≤ n ≤ N;
[0018] When n is 1, the nth question is the first question sent by the user role in the first conversation record; when n is not 1, the nth question is a question generated by the user role model based on the (n - 1)th follow-up question, and the (n - 1)th follow-up question is generated by the customer service role model during the (n - 1)th round of sample generation;
[0019] When n is 1, the dialogue context corresponding to the nth round of sample generation process is the nth question; when n is not 1, the dialogue context corresponding to the nth round of sample generation process includes: a dialogue sequence composed of n - 1 questions and n - 1 follow-up questions in the previous n - 1 rounds of sample generation processes and the nth question.
[0020] The condition for ending the execution of the sample generation process is that the Nth reply includes at least part of the target knowledge information.
[0021] In a third aspect, the present application provides an electronic device, which includes a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory, so that the electronic device executes the training method of the question-and-answer model in the first aspect or any implementation manner of the first aspect.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions direct an electronic device to execute the training method of the question-and-answer model described in the first aspect or any implementation manner of the first aspect above.
[0023] In a fifth aspect, the present application provides a computer program product containing instructions, which, when running on an electronic device, causes the electronic device to execute the training method of the question-and-answer model described in the first aspect or any implementation manner of the first aspect above.
[0024] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.
[0025] From the above technical solutions, it can be seen that the present application has the following advantages:
[0026] The present application provides a training method for a question-and-answer model. The method first obtains a first dialogue record in a target knowledge field, where the first dialogue record includes multiple rounds of conversations between a user role and a customer service role for a target question in the target knowledge field, and the target reply of the customer service role for the target question in the first dialogue record includes target knowledge information in the target knowledge field. Then, based on the first dialogue record, N rounds of sample generation processes are executed to obtain training data, and the training data is used to train the target question-and-answer model.
[0027] Among them, the n-th round of sample generation process includes: using the target Q&A model, based on the knowledge information of the target knowledge domain and the conversation context corresponding to the n-th round of sample generation process, answering the n-th question to obtain the n-th reply, where 1 ≤ n ≤ N. When n = 1, the n-th question is the first question sent by the user role in the first conversation record; when n ≠ 1, the n-th question is the question generated by the user role model based on the (n - 1)-th follow-up question, and the (n - 1)-th follow-up question is generated by the customer service role model in the (n - 1)-th round of sample generation process. When n = 1, the conversation context corresponding to the n-th round of sample generation process is the n-th question; when n ≠ 1, the conversation context corresponding to the n-th round of sample generation process includes: a conversation sequence composed of the n - 1 questions and n - 1 follow-up questions generated in the previous n - 1 rounds of sample generation process and the n-th question. The condition for ending the execution of the sample generation process is: the N-th reply includes at least part of the target knowledge information.
[0028] In this method, for the Q&A scenario of multi-round conversations in the target knowledge domain, by obtaining the first conversation record generated during the actual Q&A process, based on the first conversation record, using the model playing the user role (i.e., the question proposer) and the model playing the customer service role (i.e., the question answerer), simulating the processes of "follow-up question" and "re-question" in multi-round conversations, and constructing diverse training data. In this way, using this training data for model fine-tuning can improve the follow-up question ability of the Q&A model in multi-round conversations and enhance the Q&A performance of the Q&A model in the target knowledge domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the technical method of the embodiments of the present application, the drawings required in the embodiments will be briefly introduced below.
[0030] Figure 1 It is a schematic flowchart of a method for training a Q&A model provided by an embodiment of the present application;
[0031] Figure 2 It is a schematic diagram of a sample directed graph provided by an embodiment of the present application;
[0032] Figure 3 It is a schematic structural diagram of a training device for a Q&A model provided by an embodiment of the present application;
[0033] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The terms "first" and "second" in the embodiments of the present application are only for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0035] First, some technical terms and application scenarios involved in the embodiments of the present application are introduced.
[0036] A language model can be understood as a natural language processing model based on deep learning technology. A language model usually has the ability to understand, process, and generate natural language. The language model can be applied to different scenarios. For example, the language model can be applied to scenarios such as question answering, text classification, and content generation.
[0037] When applying the language model to the question answering scenario, the language model can also be called a question answering model. The question answering model can, in the form of a virtual assistant, intelligent customer service, etc., answer the user's questions through dialogue. Usually, the question answering scenario can be related to a specific knowledge domain. For example, the specific knowledge domain can be the information technology (IT) domain. The user can send questions related to the IT domain, and the question answering model answers the questions and returns knowledge information related to the IT domain to the user.
[0038] In the related art, the inference ability of the language model in a specific knowledge domain is improved by fine-tuning the language model. Specifically, the training process of the language model is usually divided into a pre-training stage and a post-training stage. In the pre-training stage, the language model is trained by self-supervised learning on a data set. In the post-training stage, the model is fine-tuned by supervised fine tuning (SFT) and reinforcement learning from human feedback (RLHF) to improve the performance of the language model in a specific knowledge domain and optimize the output quality of the language model.
[0039] In the question answering scenario, in the post-training stage, the pre-trained question answering model is usually fine-tuned using the conversation records in a specific knowledge domain to improve the question answering ability of the question answering model in the specific knowledge domain. However, the performance of the question answering model fine-tuned by the above method is not good in multi-round conversations, especially when the information provided by the user is incomplete and the question description is unclear, it is difficult to conduct effective follow-up questions.
[0040] In view of this, the present application provides a method for training a question-and-answer model. The method first obtains a first conversation record in a target knowledge domain. The first conversation record includes multiple rounds of conversations between a user role and a customer service role regarding a target question in the target knowledge domain. The target response of the customer service role to the target question in the first conversation record includes target knowledge information in the target knowledge domain. Then, based on the first conversation record, N rounds of sample generation processes are executed to obtain training data, and the training data is used to train the target question-and-answer model.
[0041] Among them, the nth round of sample generation process includes: using the target question-and-answer model, based on the knowledge information in the target knowledge domain and the conversation context corresponding to the nth round of sample generation process, answering the nth question to obtain the nth response, where 1 ≤ n ≤ N. When n = 1, the nth question is the first question sent by the user role in the first conversation record; when n ≠ 1, the nth question is a question generated by the user role model based on the (n - 1)th follow-up question, and the (n - 1)th follow-up question is generated by the customer service role model in the (n - 1)th round of sample generation process. When n = 1, the conversation context corresponding to the nth round of sample generation process is the nth question; when n ≠ 1, the conversation context corresponding to the nth round of sample generation process includes: a conversation sequence composed of the n - 1 questions and n - 1 follow-up questions generated in the previous n - 1 rounds of sample generation processes and the nth question. The condition for ending the execution of the sample generation process is that the Nth response includes at least part of the target knowledge information.
[0042] In this method, for the question-and-answer scenario of multiple rounds of conversations in the target knowledge domain, by obtaining the first conversation record generated during the actual question-and-answer process, and based on the first conversation record, using the model playing the user role (i.e., the question proposer) and the model playing the customer service role (i.e., the question answerer), simulating the processes of "follow-up questions" and "re-asking questions" in multiple rounds of conversations, diverse training data is constructed. In this way, using this training data for model fine-tuning can improve the follow-up question ability of the question-and-answer model in multiple rounds of conversations and enhance the question-and-answer performance of the question-and-answer model in the target knowledge domain.
[0043] To facilitate understanding of the technical solution provided by the embodiments of the present application, the following will be described in conjunction with the accompanying drawings. Refer to Figure 1 the flowchart of a method for training a question-and-answer model shown in the figure. The method specifically includes:
[0044] S101: Obtain a first conversation record in the target knowledge domain.
[0045] Among them, the target knowledge domain can be understood as any knowledge domain related to professional technology. In actual business scenarios, the target knowledge domain can be related to specific businesses. For example, in the product sales scenario, the target knowledge domain can be related to the product sales business. Another example is that in the IT operation and maintenance scenario, the target knowledge domain can be related to the IT operation and maintenance business.
[0046] In the embodiments of the present application, since it is necessary to enhance the reasoning ability of the question-and-answer model in the target knowledge domain and multi-round conversations, therefore, the first conversation record may include multi-round conversations between the user role and the customer service role for answering questions about the target question in the target knowledge domain. The target reply of the customer service role to the target question in the first conversation record includes the target knowledge information of the target knowledge domain.
[0047] In other words, the user role can be understood as the questioner, who asks questions about the target question in the target knowledge domain, and the customer service role can be understood as the answerer, who answers the target question in the target knowledge domain by recommending the target knowledge information of the target knowledge domain to the user role. That is, in the first conversation record, multi-round conversations are carried out between the user role and the customer service role. The user role asks the target question, and the customer service role answers the target question, and the target knowledge information used to answer the target question is used as the final target reply.
[0048] For example, the user role can be a user, the customer service role can be a customer service staff, and the target knowledge domain can be the IT operation and maintenance domain. The first conversation record can be as follows:
[0049] "User: It doesn't work.
[0050] Customer service staff: What doesn't work?
[0051] User: Word doesn't work.
[0052] Customer service staff: Your license has expired. You can reapply from [Apply for license](http: / / how-to-apply).
[0053] In the above first conversation record as an example, the target question is "Word doesn't work", the target reply is "Your license has expired. You can reapply from [Apply for license](http: / / how-to-apply)", and the target knowledge information is "[Apply for license](http: / / how-to-apply)".
[0054] The embodiments of the present application do not limit the manner of obtaining the first conversation record. In some embodiments, the historical conversation record in the Q&A system of the target knowledge domain may be used as the first conversation record, where the Q&A system of the target knowledge domain can be used for interaction between the user role and the customer service role regarding relevant questions in the target knowledge domain. In other embodiments, the conversation record may also be presented in the form of a work order, that is, completing a question answering is equivalent to completing a work order. In this case, the historical work orders related to the target knowledge domain in the work order system may be used as the first conversation record.
[0055] Since training data needs to be generated based on the first conversation record subsequently, in order to ensure the quality of model fine-tuning, in the embodiments of the present application, there are requirements for the conversation quality of the first conversation record. In some possible implementation manners, the first conversation record may be a conversation record that the user role expresses satisfaction with. For example, after the user role and the customer service role complete multiple rounds of conversations, the user role can give feedback on the conversation record (such as satisfied or dissatisfied), and determine the conversation record that the user role expresses satisfaction with as the first conversation record to ensure that the quality of the first conversation record meets the requirements of model fine-tuning. In other possible implementation manners, the first conversation record may also be a conversation record that passes the quality inspection. For example, the quality inspection personnel check the quality of the conversation record (such as passing the quality inspection or failing to pass the quality inspection), and determine the conversation record that passes the quality inspection as the first conversation record to ensure that the quality of the first conversation record meets the requirements of model fine-tuning.
[0056] S102: Based on the first conversation record, perform N rounds of sample generation processes to obtain training data.
[0057] In the embodiments of the present application, the sample generation process can be understood as a process of constructing a new multi-round conversation based on the first conversation record. Among them, one round of sample generation process may include three steps: "determining the question", "determining the reply", and "determining the follow-up question".
[0058] Specifically, taking the nth round of sample generation process as an example for illustration, 1 ≤ n ≤ N, that is, the nth round of sample generation process can be understood as any round of sample generation process in the N rounds of sample generation processes. The nth round of sample generation process may include: using the target Q&A model, based on the knowledge information of the target knowledge domain and the conversation context corresponding to the nth round of sample generation process, answering the nth question to obtain the nth reply.
[0059] The target question-answering model can be understood as a pre-trained language model that requires model fine-tuning. That is, in the embodiments of the present application, model fine-tuning is performed on the target question-answering model, and the fine-tuned target question-answering model is used for question answering. For example, the fine-tuned target question-answering model can act as a virtual assistant, intelligent customer service, etc., and answer questions in the target knowledge domain to achieve intelligent question answering.
[0060] In each round of sample generation process, the target question-answering model executes the step of "determining the reply". In the embodiments of the present application, the retrieval-augmented generation (RAG) technology is used to assist the target question-answering model in generating replies. Among them, the RAG technology is divided into two stages: the information retrieval stage and the reply generation stage. In the information retrieval stage, by calling the recall model, knowledge information in the knowledge base of the target knowledge domain is retrieved. In the reply generation stage, the target question-answering model answers the nth question with the knowledge information obtained in the information retrieval stage to generate the nth reply.
[0061] Among them, when n is 1, the nth question is the first question sent by the user role in the first conversation record. That is, in the first round of sample generation process, the first question is the first question sent by the user role in the first conversation record. Starting from the first question sent by the user role, a multi-round conversation is constructed.
[0062] It should be noted that the first question sent by the user role is not necessarily the same as the target question. In some embodiments, in the first conversation record, the first question sent by the user role is the target question. At this time, the first question is the target question. In other embodiments, in the first conversation record, the first question sent by the user role does not clearly describe the question, and the target question is only proposed in the subsequent questions of the user role. At this time, the first question is not the target question. For example, in the previous example in the IT operation and maintenance field, the first question sent by the user role is "not working", and the target question is "word is not working", and the two are different.
[0063] When n is not 1, the nth question is the question generated by the user role model based on the (n - 1)th follow-up question, and the (n - 1)th follow-up question is generated by the customer service role model in the (n - 1)th round of sample generation process.
[0064] That is, starting from the second round of sample generation process, the user role model executes the step of "determining the question". In each round of sample generation process including the first round of sample generation process, the customer service role model executes the step of "determining the follow-up question".
[0065] Among them, the user role model can be understood as a language model representing the user role, and the customer service role model can be understood as a language model representing the customer service role. The user role model and the customer service role model can be self-developed language models.
[0066] That is to say, the user role model can simulate and play the role of the user, and from the perspective of the user, perform the step of "determining the problem". The customer service role can simulate and play the role of the customer service staff, and from the perspective of the customer service staff, perform the step of "determining the follow-up question".
[0067] In the embodiment of the present application, the condition for ending the execution of the sample generation process is that the Nth reply includes at least part of the target knowledge information. In other words, the sample generation process is executed N rounds in total. When the condition for ending the execution of the sample generation process is met in the Nth round of the sample generation process, that is, when the Nth reply includes at least part of the target knowledge information, the sample generation process ends and the (N + 1)th round of the sample generation process is no longer executed.
[0068] For example, when the target knowledge information only includes one knowledge article, the condition for ending the execution of the sample generation process is that the Nth reply includes that knowledge article of the target knowledge information. Another example is that when the target knowledge information includes multiple knowledge articles, the condition for ending the execution of the sample generation process is that the Nth reply includes the number of knowledge articles exceeding the proportion threshold in the target knowledge information.
[0069] That is to say, if in the Nth round of the sample generation process, the Nth reply generated by the target question and answer model includes at least part of the target knowledge information, it indicates that the Nth reply is the target reply in the first conversation record, the target question and answer model has generated the correct reply to the target question, and based on the first conversation record, a new conversation record starting from "the first question sent by the user role in the first conversation record" and ending with the "Nth reply" equivalent to the target reply is constructed, automatically constructing a high-quality multi-round conversation.
[0070] Taking the example that there are 2 rounds of sample generation processes in total, in the first round of the sample generation process, the first question is the first question sent by the user role in the first conversation record. The target question and answer model answers the first question based on the knowledge information in the target knowledge field and the conversation context corresponding to the first round of the sample generation process, obtaining the first reply. The first reply does not meet the condition for ending the execution of the sample generation process, and the customer service role model asks a follow-up question about the first question, obtaining the first follow-up question.
[0071] In the process of generating the second-round samples, based on the first follow-up question, the user role model proposes the second question. The target Q&A model answers the second question based on the knowledge information in the target knowledge domain and the dialogue context corresponding to the second-round sample generation process, obtaining the second reply. The second reply meets the condition for ending the sample generation process. However, in order to construct training data with the same path prefix, the customer service role model follows up on the second question, obtaining the second follow-up question.
[0072] In the embodiments of the present application, the target Q&A model answers the nth question in combination with the dialogue context corresponding to the nth-round sample generation process. Among them, the dialogue context corresponding to the nth-round sample generation process can be understood as the generated dialogue sequence including questions and follow-up questions. That is to say, the target Q&A model answers the nth question through the generated dialogue sequence, obtaining the nth reply.
[0073] Specifically, when n is 1, the dialogue context corresponding to the nth-round sample generation process is the nth question. Since in the first-round sample generation process, no other dialogue content has been generated, therefore, the dialogue context corresponding to the first-round sample generation process is the first question itself.
[0074] When n is not 1, the dialogue context corresponding to the nth-round sample generation process includes: a dialogue sequence composed of n - 1 questions and n - 1 follow-up questions in the previous n - 1 rounds of sample generation processes and the nth question.
[0075] Since the n - 1 replies generated in the previous n - 1 rounds of sample generation do not meet the condition for ending the sample generation process, that is, the n - 1 replies generated in the previous n - 1 rounds of sample generation are not the correct replies to the target question. In the previous n - 1 rounds of sample generation, the target Q&A model cannot give a correct reply based on the existing dialogue sequence and should conduct a follow-up question. Therefore, the previous n - 1 replies are not included in the dialogue context corresponding to the nth-round sample generation process.
[0076] The following explains the processes of the target Q&A model generating replies, the user role model generating questions, and the customer service role model generating follow-up questions.
[0077] The target Q&A model can generate replies based on the way of prompt learning. Among them, the prompt can be used to guide the language model to perform specific outputs in generative tasks (such as text generation tasks, Q&A tasks, dialogue tasks). By configuring the prompt, it helps the language model understand the background and requirements of the task, enabling the language model to handle different types of natural language processing tasks without retraining the language model, and increasing the scalability and flexibility of the language model.
[0078] Specifically, generate a first prompt, send the first prompt to the target Q&A model, and receive the nth response returned by the target Q&A model.
[0079] Among them, the first prompt includes: knowledge information in the target knowledge domain, the dialogue context corresponding to the nth round of sample generation process, and prompt information for indicating an answer to the nth question. Since the first prompt includes the above information, the target Q&A model can, based on the prompting ability of the first prompt, combine the knowledge information in the target knowledge domain and the dialogue context corresponding to the nth round of sample generation process, and in the nth round of sample generation process, reply to the nth question to obtain the nth response.
[0080] Similarly, when n is not 1, the user role model can generate questions based on prompt learning. Specifically, generate a third prompt, send the third prompt to the user role model, and receive the nth question returned by the user role model.
[0081] Among them, the third prompt includes: prompt information for indicating to simulate the user role, a dialogue sequence consisting of n - 1 questions and n - 1 follow-up questions generated during the previous n - 1 rounds of sample generation process, the first dialogue record, and prompt information for indicating to generate questions.
[0082] Since the user role model needs to re-ask questions based on the existing dialogue sequence, therefore, the user role model needs to know the complete and correct first dialogue record to ensure that the generated nth question fits the first dialogue record and avoid generating questions that do not match the first dialogue record.
[0083] Since the third prompt includes the above information, the user role model can, based on the prompting ability of the third prompt, simulate the user's role, combine the existing dialogue sequence, and in the nth round of sample generation process, generate the nth question that fits the first dialogue record.
[0084] Similarly, when n is not 1, the customer service role model can generate follow-up questions based on prompt learning. Specifically, generate a second prompt, send the second prompt to the customer service role model, and receive the (n - 1)th follow-up question returned by the customer service role model.
[0085] Among them, the second prompt includes: prompt information for indicating to simulate the customer service role, a dialogue sequence consisting of n - 1 questions and n - 2 follow-up questions generated during the previous n - 1 rounds of sample generation process, and prompt information for indicating to generate follow-up questions.
[0086] Since the customer service role model has not generated the (n - 1)-th follow-up question during the sample generation process of the (n - 1)-th round, at this time, the existing dialogue sequence only includes (n - 1) questions and (n - 2) follow-up questions. Since the second prompt includes the above information, the customer service role model can, based on the prompting ability of the second prompt, simulate the role of the second object, and in combination with the existing dialogue sequence, during the sample generation process of the (n - 1)-th round, ask a follow-up question about the unclear (n - 1)-th question to obtain the (n - 1)-th follow-up question.
[0087] Taking the example in the IT operation and maintenance field mentioned above, during the sample generation process of the first round, the first question is "It doesn't work". The target Q&A model replies to the first question. Since the description of the first question is unclear, the first reply generated by the target Q&A model does not include the target knowledge information, and the sample generation process continues. The customer service role model asks a follow-up question about the first question. At this time, the second prompt can be as follows:
[0088] "You are responsible for playing the role of a customer service staff. According to the following dialogue record, ask reasonable follow-up questions about the user's question in order to help the user locate and solve the problem as soon as possible.
[0089] Dialogue record:
[0090] User: It doesn't work"
[0091] Send this second prompt to the customer service role model, and the customer service role model outputs the first follow-up question "What doesn't work for you?", completing the sample generation process of the first round.
[0092] During the sample generation process of the second round, the user role model re-describes the problem based on the first follow-up question. At this time, the third prompt can be as follows:
[0093] "The first dialogue record is:
[0094] User: It doesn't work
[0095] Customer service staff: What doesn't work?
[0096] User: Word doesn't work
[0097] Customer service staff: Your license has expired. You can re-apply from [Apply for license](http: / / how-to-apply)
[0098] You are responsible for playing the role of a user. According to the first dialogue record, for the follow-up question of the customer service staff in the following dialogue record, faithfully based on the above first dialogue record, give an answer that conforms to the above first dialogue record and the follow-up question of the customer service staff.
[0099] Dialogue record:
[0100] User: It doesn't work
[0101] Customer service staff: May I ask what doesn't work?
[0102] Send the third prompt word to the user role model, and the user role model outputs the second question "Word doesn't work". The target Q&A model replies to the second question. Since the second question clearly describes the target problem, the target Q&A model can generate the target knowledge information (i.e., "[Apply for permission](http: / / how-to-apply)"), meeting the condition to end the sample generation process. At this time, in order to construct training data with the same path prefix, the customer service role model can also follow up on the second question to get the second follow-up question.
[0103] After completing the N-round sample generation process, the Nth reply meets the condition to end the sample generation process, and the process of constructing a multi-round dialogue ends. In the embodiment of the present application, the constructed multi-round dialogue is converted into a sample directed graph, and the association relationship between "questions", "replies" and "follow-up questions" in the constructed multi-round dialogue is represented by the sample directed graph.
[0104] Specifically, according to the N questions, N follow-up questions and N replies in the N-round sample generation process, a sample directed graph is constructed, and based on the sample directed graph, positive samples and negative samples are determined.
[0105] Among them, the sample directed graph consists of a first vertex representing a question, a second vertex representing a reply, a third vertex representing a follow-up question, and multiple directed edges. In other words, "questions", "replies" and "follow-up questions" in the multi-round dialogue are represented as different types of vertices in the sample directed graph.
[0106] The directed edges in the sample directed graph include: an edge from the first vertex representing the mth question to the second vertex representing the mth reply, an edge from the first vertex representing the mth question to the third vertex representing the mth follow-up question, and an edge from the third vertex representing the mth follow-up question to the first vertex representing the (m + 1)th question, where 1 ≤ m ≤ N.
[0107] In the N-round sample generation process, the mth reply is the answer generated by the target Q&A model for the mth question. Therefore, there is an association relationship between the first vertex representing the mth question and the second vertex representing the mth reply. In the N-round sample generation process, the mth follow-up question is the follow-up question generated by the customer service role model for the mth question. Therefore, there is an association relationship between the first vertex representing the mth question and the third vertex representing the mth follow-up question. In N
[0108] During the generation process of round samples, the (m + 1)-th question is a re-question generated for the m-th follow-up question by the user role model. Therefore, there is an association relationship between the third vertex representing the m-th follow-up question and the first vertex representing the (m + 1)-th question.
[0109] See Figure 2 the schematic diagram of a sample directed graph as shown. In the sample directed graph, Qi represents the first vertex, Ai represents the second vertex, ai represents the third vertex, and i represents the i-th round of sample generation process. From Figure 2 it can be seen that there are edges from Q1 to A1, from Q1 to a1, from a1 to Q2, etc. in the sample directed graph.
[0110] Since there are multiple first vertices representing multiple responses generated by the target Q&A model in the sample directed graph, and there are incorrect responses (i.e., the first N - 1 responses) and correct responses (i.e., the N-th response) among the multiple responses, therefore, from the sample directed graph, positive samples representing the correct conversation for the target question and negative samples representing the incorrect conversation for the target question can be determined.
[0111] In some possible implementation manners, from the sample directed graph, determine the first path from the first vertex representing the 1st question to the second vertex representing the N-th response, and then, determine the first path and the sub-path from the first vertex representing the 1st question to any third vertex representing a follow-up question in the first path as positive samples.
[0112] That is to say, the first path can be understood as the path in the sample directed graph from the first vertex representing the first question to the second vertex representing the last response. For example, in Figure 2 it, the first path is "Q1 - a1 - Q2 - a2 - Q3 - …… - QN - AN".
[0113] It can be found that except for the second vertex representing the N-th response in the first path, the remaining vertices are all the first vertices representing questions and the third vertices representing follow-up questions. The N-th response satisfies the condition for ending the sample generation process, that is, only when the target Q&A model generates the N-th response, does it indicate that the target Q&A model has made a correct response to the target question. That is to say, the goal of model fine-tuning for the target Q&A model is: before the N-th question appears in the multi-round conversation, the target Q&A model should conduct follow-up questions based on the question instead of directly making a response. Therefore, the first path represents the correct multi-round conversation and belongs to positive samples.
[0114] Further, since the first path represents a correct multi-turn conversation, the sub-path from the first vertex to any third vertex representing a follow-up question in the first path can represent the follow-up questions that the target Q&A model should generate for the existing conversation sequence. Therefore, they also represent correct conversations and belong to positive samples. For example, in Figure 2 the sub-paths from the first vertex to any third vertex representing a follow-up question in the first path can be "Q1-a1", "Q1-a1-Q2-a2", etc.
[0115] For negative samples, in the sample directed graph, the paths from the first vertex representing the first question to the second vertex representing the k-th reply, and the paths from the first vertex representing the first question to the third vertex representing the N-th follow-up question are determined as negative samples. Where 1 ≤ k < N.
[0116] That is to say, the k-th reply is any reply among the first N - 1 replies. The goal of fine-tuning the target Q&A model is: for the first N - 1 questions, the target Q&A model should ask follow-up questions based on the questions instead of directly giving replies. Therefore, the path from the first vertex representing the first question to the second vertex representing the k-th reply represents an incorrect multi-turn conversation and belongs to negative samples.
[0117] The path from the first vertex representing the first question to the third vertex representing the N-th follow-up question is the path from the first vertex representing the first question to the last vertex representing the third follow-up question in the sample directed graph. Since the N-th reply is the correct reply, the goal of fine-tuning the target Q&A model is: for the N-th question, the target Q&A model should give a reply instead of asking a follow-up question. Therefore, the path from the first vertex representing the first question to the third vertex representing the N-th follow-up question represents an incorrect multi-turn conversation and belongs to negative samples.
[0118] In this way, from the sample directed graph, the paths belonging to positive samples and the paths belonging to negative samples are screened out. For the constructed multi-turn conversations, both positive and negative samples are determined simultaneously, improving the efficiency of model fine-tuning. Moreover, there are positive and negative samples with the same path prefix. For example, the positive sample is "Q1-a1-Q2-a2" and the negative sample is "Q1-a1-Q2-A2", which have the same path prefix "Q1-a1-Q2", facilitating targeted adjustment in the subsequent model fine-tuning process.
[0119] S103: Train the target Q&A model using the training data.
[0120] In the post-training stage of the target question-and-answer model, the training data can be divided into the forms of "input" and "output". Among them, "input" and "output" can also be understood as question-and-answer pairs. The "input" part can be understood as the prompt words input into the target question-and-answer model, and the "output" part can be understood as the content that is expected or not expected to be generated by the target question-and-answer model.
[0121] Specifically, for positive samples, the "input" part can be called the positive sample input, and the "output" part can be called the positive sample output. That is, when the positive sample input is sent to the target question-and-answer model, it is expected that the target question-and-answer model will return the positive sample output. For negative samples, the "input" part can be called the negative sample input, and the "output" part can be called the negative sample output. That is, when the negative sample input is sent to the target question-and-answer model, it is not expected that the target question-and-answer model will return the negative sample output.
[0122] In the embodiments of the present application, for each positive sample, the following steps are performed: According to the path composed of the vertices except the last vertex in the positive sample and the knowledge information in the target knowledge field, determine the positive sample input, and, determine the last vertex in the positive sample as the positive sample output, and construct a positive sample question-and-answer pair with the positive sample input and the positive sample output.
[0123] For each negative sample, the following steps are performed: According to the path composed of the vertices except the last vertex in the negative sample and the knowledge information in the target knowledge field, determine the negative sample input, and, determine the last vertex in the negative sample as the negative sample output, and construct a negative sample question-and-answer pair with the negative sample input and the negative sample output.
[0124] The last vertices in both positive samples and negative samples are related to the target question-and-answer model. In positive samples, the last vertex represents the correct reply that the target question-and-answer model should generate for the target question, or represents the follow-up question that the target question-and-answer model should generate for other questions with unclear descriptions. In negative samples, the last vertex represents the reply that the target question-and-answer model should not generate for other questions with unclear descriptions, or represents the follow-up question that the target question-and-answer model should not generate for the target question. Therefore, according to the content of the vertices except the last vertex in the positive sample, determine the positive sample input, determine the content of the last vertex in the positive sample as the positive sample output, according to the content of the vertices except the last vertex in the negative sample, determine the negative sample input, and determine the content of the last vertex in the negative sample as the negative sample output.
[0125] It can be understood that since the positive sample input and the negative sample input are equivalent to the prompt words input into the target question-and-answer model, therefore, the relevant content in the positive samples and negative samples needs to be filled into the prompt word template, and then the positive sample input and the negative sample input are formed.
[0126] Taking the positive sample as "Q1-a1-Q2-a2" and the negative sample as "Q1-a1-Q2-A2" as an example, the positive sample input and the negative sample input can be as follows:
[0127] "As a Q&A expert, your job is to answer or follow up on the user's questions based on the following knowledge information, and finally return the information that can answer the questions to the user.
[0128] Knowledge information: {Knowledge information in the target knowledge domain}
[0129] Conversation record:
[0130] User: {Q1}
[0131] Intelligent customer service: {a1}
[0132] User: {Q2}"
[0133] Since the positive sample and the negative sample have the same path prefix, the positive sample input and the negative sample input are the same. The positive sample output is "a2", and the negative sample output is "A2".
[0134] After constructing the positive sample Q&A pairs and the negative sample Q&A pairs, the positive sample Q&A pairs and the negative sample Q&A pairs are used to train the target Q&A model. In some embodiments, the positive sample Q&A pairs are used to perform SFT on the target Q&A model to improve the multi-turn conversation ability of the target Q&A model. In other embodiments, the positive sample Q&A pairs and the negative sample Q&A pairs are used to perform RLHF on the target Q&A model. For example, the RLHF is performed using the optimization method based on Kahneman and Tversky (kahneman-tversky optimization, KTO), the RLHF is performed using the direct preference optimization (direct preference optimization, DPO) method, etc.
[0135] Furthermore, the embodiments of the present application also support data augmentation of the training data in different ways to strengthen the effect of model fine-tuning and improve the robustness of the target Q&A model after model fine-tuning.
[0136] In some possible implementation manners, the content corresponding to the first vertex and the second vertex in the positive sample is replaced with semantically similar content to generate a new positive sample, and the content corresponding to the first vertex and the second vertex in the negative sample is replaced with semantically similar content to generate a new negative sample.
[0137] That is to say, the content representing the question and the content representing the reply in the training data are replaced with semantically similar expressions to enrich the number of positive samples and negative samples. And by adjusting the expression ways of the positive samples and the negative samples, the Q&A ability of the target Q&A model under different similar expressions is enhanced.
[0138] In some other possible implementation manners, the knowledge information of the target knowledge domain may include multiple knowledge articles. Perform at least one of the following modification operations: randomly add at least one knowledge article to the knowledge information of the target knowledge domain, adjust the arrangement order of the multiple knowledge articles in the knowledge information of the target knowledge domain, replace the content representing the titles of the multiple knowledge articles in the knowledge information of the target knowledge domain with semantically similar content, and replace the content representing the article content of the multiple knowledge articles in the knowledge information of the target knowledge domain with semantically similar content. Then, use the knowledge information of the target knowledge domain after performing the modification operations to determine new positive sample inputs and new negative sample inputs, construct new positive sample question-and-answer pairs with the new positive sample inputs and positive sample outputs, and construct new negative sample question-and-answer pairs with the new negative sample inputs and negative sample outputs.
[0139] That is to say, adjust the content related to the "knowledge information of the target knowledge domain" in the positive sample inputs and negative sample inputs, enrich the number of positive samples and negative samples by adding new knowledge articles, shuffling the order of multiple knowledge articles, or replacing the titles and content of knowledge articles. And, enhance the ability of the target question-and-answer model to answer questions by combining different knowledge information by adjusting the content related to the "knowledge information of the target knowledge domain".
[0140] In this method, for the question-and-answer scenario of multi-turn conversations in the target knowledge domain, by obtaining the first conversation record generated during the actual question-and-answer process, on the basis of the first conversation record, use the model playing the role of the user (i.e., the question asker) and the model playing the role of the customer service (i.e., the question answerer) to simulate the processes of "asking for clarification" and "re-asking" in multi-turn conversations, and construct diverse training data. In this way, use this training data to fine-tune the model, improve the ability of the question-and-answer model to ask for clarification in multi-turn conversations, and enhance the question-and-answer performance of the question-and-answer model in the target knowledge domain.
[0141] As described above in conjunction with Figure 1 and Figure 2 the training method of the question-and-answer model provided by the embodiments of the present application has been introduced in detail. Next, the devices and equipment provided by the embodiments of the present application will be introduced with reference to the accompanying drawings.
[0142] See Figure 3 the structural schematic diagram of the training device of the question-and-answer model shown in the figure. The device 30 includes:
[0143] An acquisition module 301, configured to acquire a first conversation record of a target knowledge domain; wherein, the first conversation record includes multiple rounds of conversations between a user role and a customer service role for a target question in the target knowledge domain, and a target reply of the customer service role for the target question in the first conversation record includes target knowledge information of the target knowledge domain;
[0144] A generation module 302, configured to perform N rounds of sample generation processes based on the first conversation record to obtain training data;
[0145] A training module 303, configured to use the training data to train a target Q&A model;
[0146] Wherein, the nth round of sample generation process includes: using the target Q&A model, based on the knowledge information of the target knowledge domain and the conversation context corresponding to the nth round of sample generation process, answering the nth question to obtain the nth reply; 1 ≤ n ≤ N;
[0147] When n is 1, the nth question is the first question sent by the user role in the first conversation record; when n is not 1, the nth question is a question generated by a user role model based on the (n - 1)th follow-up question, and the (n - 1)th follow-up question is generated by a customer service role model in the (n - 1)th round of sample generation process;
[0148] When n is 1, the conversation context corresponding to the nth round of sample generation process is the nth question; when n is not 1, the conversation context corresponding to the nth round of sample generation process includes: a conversation sequence composed of n - 1 questions, n - 1 follow-up questions, and the nth question in the previous n - 1 rounds of sample generation processes;
[0149] The condition for ending the execution of the sample generation process is: the Nth reply includes at least part of the target knowledge information.
[0150] In some possible implementation manners, the generation module 302 is specifically configured to:
[0151] Construct a sample directed graph according to N questions, N follow-up questions, and N replies in N rounds of sample generation processes; wherein, the sample directed graph is composed of a first vertex representing a question, a second vertex representing a reply, a third vertex representing a follow-up question, and multiple directed edges, and the directed edges include: an edge pointing from the first vertex representing the mth question to the second vertex representing the mth reply, an edge pointing from the first vertex representing the mth question to the third vertex representing the mth follow-up question, and an edge pointing from the third vertex representing the mth follow-up question to the first vertex representing the (m + 1)th question, 1 ≤ m ≤ N;
[0152] Determine positive samples and negative samples according to the sample directed graph; wherein, the positive samples represent correct conversations for the target problem, and the negative samples represent incorrect conversations for the target problem.
[0153] In some possible implementation manners, the generating module 302 is specifically configured to:
[0154] Determine a first path from a first vertex representing the first question to a second vertex representing the Nth reply in the sample directed graph;
[0155] Determine the first path and sub-paths from the first vertex representing the first question to any third vertex representing a follow-up question in the first path as positive samples;
[0156] Determine, from the sample directed graph, paths from the first vertex representing the first question to a second vertex representing the kth reply and paths from the first vertex representing the first question to a third vertex representing the Nth follow-up question as negative samples; where 1 ≤ k < N.
[0157] In some possible implementation manners, the generating module 302 is further configured to:
[0158] Replace the content corresponding to the first vertex and the second vertex in the positive samples with semantically similar content to generate new positive samples; and,
[0159] Replace the content corresponding to the first vertex and the second vertex in the negative samples with semantically similar content to generate new negative samples.
[0160] In some possible implementation manners, the training module 303 is specifically configured to:
[0161] For each positive sample, perform the following steps: Determine a positive sample input according to the path formed by vertices except the last vertex in the positive sample and the knowledge information in the target knowledge domain, and determine the last vertex in the positive sample as the positive sample output; Construct a positive sample question-answer pair with the positive sample input and the positive sample output;
[0162] For each negative sample, perform the following steps: Determine a negative sample input according to the path formed by vertices except the last vertex in the negative sample and the knowledge information in the target knowledge domain, and determine the last vertex in the negative sample as the negative sample output; Construct a negative sample question-answer pair with the negative sample input and the negative sample output;
[0163] Train the target question-answer model using the positive sample question-answer pairs and the negative sample question-answer pairs.
[0164] In some possible implementations, the knowledge information of the target knowledge domain includes multiple knowledge articles; the generating module 302 is further configured to:
[0165] Perform at least one of the following modification operations: randomly add at least one knowledge article to the knowledge information of the target knowledge domain, adjust the arrangement order of the multiple knowledge articles in the knowledge information of the target knowledge domain, replace the content representing the titles of the multiple knowledge articles in the knowledge information of the target knowledge domain with semantically similar content, and replace the content representing the article content of the multiple knowledge articles in the knowledge information of the target knowledge domain with semantically similar content;
[0166] Use the knowledge information of the target knowledge domain after performing the modification operation to determine a new positive sample input and a new negative sample input;
[0167] Construct a new positive sample question-and-answer pair with the new positive sample input and the positive sample output, and construct a new negative sample question-and-answer pair with the new negative sample input and the negative sample output.
[0168] In some possible implementations, the generating module 302 is specifically configured to:
[0169] Generate a first prompt; wherein, the first prompt includes: the knowledge information of the target knowledge domain, the dialogue context corresponding to the nth round of sample generation process, and the prompt information for indicating an answer to the nth question;
[0170] Send the first prompt to the target question-and-answer model and receive the nth reply returned by the target question-and-answer model.
[0171] In some possible implementations, the generating module 302 is specifically configured to:
[0172] Generate a second prompt; wherein, the second prompt includes: the prompt information for indicating to simulate the customer service role, the dialogue sequence composed of n - 1 questions and n - 2 follow-up questions generated during the previous n - 1 rounds of sample generation process, and the prompt information for indicating to generate a follow-up question;
[0173] Send the second prompt to the customer service role model and receive the (n - 1)th follow-up question returned by the customer service role model.
[0174] In some possible implementations, the generating module 302 is specifically configured to:
[0175] Generate a third prompt; wherein, the third prompt includes: a prompt message for indicating an analog user role, a dialogue sequence composed of n-1 questions and n-1 follow-up questions generated during the generation process of the previous n-1 rounds of samples, the first dialogue record, and a prompt message for indicating the generation of questions;
[0176] Send the third prompt to the user role model and receive the nth question returned by the user role model.
[0177] The training device 30 of the question-and-answer model according to the embodiment of the present application can correspondingly execute the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module / unit of the training device 30 of the question-and-answer model are respectively for realizing Figure 1 The corresponding processes of the respective methods in the illustrated embodiments, and for the sake of brevity, will not be described herein again.
[0178] The embodiment of the present application also provides an electronic device. This electronic device is specifically used to implement the functions of the training device 30 of the question-and-answer model as shown in Figure 3 the illustrated embodiments.
[0179] Figure 4 A structural schematic diagram of an electronic device 400 is provided, as shown in Figure 4 The electronic device 400 includes a bus 401, a processor 402, a communication interface 403, and a memory 404. The processor 402, the memory 404, and the communication interface 403 communicate with each other through the bus 401.
[0180] The bus 401 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0181] The processor 402 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0182] The communication interface 403 is used for external communication. For example, the communication interface 403 can be used for communication with a terminal.
[0183] The memory 404 may include volatile memory, such as random access memory (RAM). The memory 404 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0184] Executable code is stored in the memory 404, and the processor 402 executes the executable code to perform the training method of the aforementioned question-and-answer model.
[0185] Specifically, in the case of implementing Figure 3 the illustrated embodiment, and Figure 3 when each module or unit of the training device 30 of the question-and-answer model described in the embodiment is implemented by software, the software or program code required to execute Figure 3 the functions of each module / unit in may be partially or entirely stored in the memory 404. The processor 402 executes the program code corresponding to each unit stored in the memory 404 to perform the training method of the aforementioned question-and-answer model.
[0186] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center that includes one or more available media. The available media may be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid state drives), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the training method of the question-and-answer model applied to the training device 30 of the present application.
[0187] The embodiment of the present application also provides a computer program product that includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the processes or functions described in the embodiment of the present application are fully or partially generated.
[0188] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center by wire (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).
[0189] When the computer program product is executed by a computer, the computer executes any one of the methods for training the foregoing question-and-answer model. The computer program product may be a software installation package. In the case where any one of the methods for training the foregoing question-and-answer model is required, the computer program product can be downloaded and executed on the computer.
[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0191] The units involved in the embodiments described in the present application can be implemented in software or in hardware. Among them, the name of the unit / module does not constitute a limitation on the unit itself in some cases.
[0192] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0193] In the context of the embodiments of the present application, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0194] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or apparatuses disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0195] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist simultaneously. Among them, A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c may be single or multiple.
[0196] It should also be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0197] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in a software module executed by a processor, or in a combination thereof. The software module may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the art.
[0198] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A training method for a question-answering model, characterized in that: The method comprises: Acquire a first dialogue record of a target knowledge domain; wherein the first dialogue record includes multiple rounds of dialogues between a user role and a customer service role regarding a target question in the target knowledge domain, and the target reply of the customer service role to the target question in the first dialogue record includes target knowledge information of the target knowledge domain; Based on the first conversation record, perform N rounds of sample generation process to obtain training data; Using the training data, training a target question-answering model; The nth round sample generation process includes: using the target question-answering model, based on the knowledge information of the target knowledge domain and the conversation context corresponding to the nth round sample generation process, answering the nth question to obtain the nth reply; 1≤n≤N; When n is 1, the nth question is the first question sent by the user role in the first conversation record; when n is not 1, the nth question is a question generated by the user role model based on the n-1th follow-up question, and the n-1th follow-up question is generated by the customer service role model in the n-1th round of sample generation; When n is 1, the conversation context corresponding to the n-th round of sample generation process is the n-th question; when n is not 1, the conversation context corresponding to the n-th round of sample generation process includes: a conversation sequence consisting of n-1 questions and n-1 follow-up questions in the first n-1 rounds of sample generation process and the n-th question; The condition for ending the sample generation process is that the Nth reply includes at least part of the target knowledge information.
2. The method according to claim 1, characterized in that The performing N rounds of sample generation process based on the first conversation record to obtain training data includes: According to N questions, N follow-up questions and N replies in the N rounds of sample generation process, a sample directed graph is constructed; wherein the sample directed graph is composed of a first vertex representing the question, a second vertex representing the reply, a third vertex representing the follow-up question and a plurality of directed edges, wherein the directed edges include: an edge pointing from the first vertex representing the mth question to the second vertex representing the mth reply, an edge pointing from the first vertex representing the mth question to the third vertex representing the mth follow-up question and an edge pointing from the third vertex representing the mth follow-up question to the first vertex representing the m+1th question, 1≤m≤N; According to the sample directed graph, positive samples and negative samples are determined; wherein the positive samples represent correct dialogues for the target question, and the negative samples represent incorrect dialogues for the target question.
3. The method according to claim 2, characterized in that Determining positive samples and negative samples according to the sample directed graph includes: Determine, from the sample directed graph, a first path from a first vertex representing the first question to a second vertex representing the Nth answer; Determine the first path and a subpath in the first path from the first vertex representing the first question to any third vertex representing a follow-up question as positive samples; From the sample directed graph, the path from the first vertex representing the first question to the second vertex representing the kth answer, and the path from the first vertex representing the first question to the third vertex representing the Nth follow-up question are determined as negative samples; wherein 1≤k<N.
4. The method according to claim 3, characterized in that The method further comprises: Replacing the content corresponding to the first vertex and the second vertex in the positive sample with semantically similar content to generate a new positive sample; and The content corresponding to the first vertex and the second vertex in the negative sample is replaced with semantically similar content to generate a new negative sample.
5. The method according to claim 3, characterized in that: The step of training the target question answering model using the training data includes: For each positive sample, the following steps are performed: according to the path composed of vertices other than the last vertex in the positive sample and the knowledge information of the target knowledge domain, the positive sample input is determined, and the last vertex in the positive sample is determined as the positive sample output; the positive sample input and the positive sample output are used to construct a positive sample question-answer pair; For each negative sample, the following steps are performed: according to the path composed of vertices other than the last vertex in the negative sample and the knowledge information of the target knowledge domain, the negative sample input is determined, and the last vertex in the negative sample is determined as the negative sample output; the negative sample input and the negative sample output are used to construct a negative sample question-answer pair; The target question-answering model is trained using the positive sample question-answering pairs and the negative sample question-answering pairs.
6. The method according to claim 5, characterized in that The knowledge information in the target knowledge domain includes a plurality of knowledge articles; Before training the target question-answer model using the positive sample question-answer pairs and the negative sample question-answer pairs, the method further includes: Perform at least one of the following modification operations: randomly adding at least one knowledge article in the knowledge information of the target knowledge field, adjusting the arrangement order of the plurality of knowledge articles in the knowledge information of the target knowledge field, replacing the content representing the titles of the plurality of knowledge articles in the knowledge information of the target knowledge field with semantically similar content, and replacing the content representing the article content of the plurality of knowledge articles in the knowledge information of the target knowledge field with semantically similar content; Determine new positive sample input and new negative sample input by using the knowledge information of the target knowledge domain after performing the modification operation; The new positive sample is input and the positive sample is output to construct a new positive sample question-answer pair, and the new negative sample is input and the negative sample is output to construct a new negative sample question-answer pair.
7. The method according to any one of claims 1 to 6, characterized in that: The using the target question-answering model to answer the nth question based on the knowledge information of the target knowledge domain and the conversation context corresponding to the nth round of sample generation process to obtain the nth reply includes: Generate a first prompt word; wherein the first prompt word includes: knowledge information of the target knowledge domain, a conversation context corresponding to the n-th round sample generation process, and prompt information for indicating an answer to the n-th question; The first prompt word is sent to the target question-answering model, and an nth response returned by the target question-answering model is received.
8. The method according to any one of claims 1 to 6, characterized in that: If n is not 1, the n-1th question is generated in the following manner: Generate a second prompt word; wherein the second prompt word includes: prompt information for indicating a simulated customer service role, a dialogue sequence consisting of n-1 questions and n-2 follow-up questions generated in the first n-1 rounds of sample generation, and prompt information for indicating the generation of follow-up questions; The second prompt word is sent to the customer service role model, and the n-1th follow-up question returned by the customer service role model is received.
9. The method according to any one of claims 1 to 6, characterized in that: The n is not 1, and the nth question is generated in the following manner: Generate a third prompt word; wherein the third prompt word includes: prompt information for indicating the simulated user role, a dialogue sequence consisting of n-1 questions and n-1 follow-up questions generated in the first n-1 rounds of sample generation, the first dialogue record, and prompt information for indicating the generated question; The third prompt word is sent to the user role model, and the nth question returned by the user role model is received.
10. A training device for a question-answering model, characterized in that: The device comprises: An acquisition module is configured to acquire a first dialogue record in a target knowledge domain; wherein the first dialogue record includes multiple rounds of dialogues between a user role and a customer service role regarding a target question in the target knowledge domain, and a target reply of the customer service role to the target question in the first dialogue record includes target knowledge information in the target knowledge domain; A generation module, configured to perform N rounds of sample generation process based on the first conversation record to obtain training data; A training module, used to train a target question-answering model using the training data; The nth round sample generation process includes: using the target question-answering model, based on the knowledge information of the target knowledge domain and the conversation context corresponding to the nth round sample generation process, answering the nth question to obtain the nth reply; 1≤n≤N; When n is 1, the nth question is the first question sent by the user role in the first conversation record; when n is not 1, the nth question is a question generated by the user role model based on the n-1th follow-up question, and the n-1th follow-up question is generated by the customer service role model in the n-1th round of sample generation; When n is 1, the conversation context corresponding to the n-th round of sample generation process is the n-th question; when n is not 1, the conversation context corresponding to the n-th round of sample generation process includes: a conversation sequence consisting of n-1 questions and n-1 follow-up questions in the first n-1 rounds of sample generation process and the n-th question; The condition for ending the sample generation process is that the Nth reply includes at least part of the target knowledge information.
11. An electronic device, characterized in that: The electronic device comprises a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The method comprises instructions, wherein the instructions instruct an electronic device to execute the method as claimed in any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product comprises computer readable instructions for implementing the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Training method and device of dialogue generation model and dialogue generation method and device
CN114547272A
Interaction method and device, computer equipment and storage medium
CN117453871A
Text generation method and apparatus, device, and non-volatile readable storage medium
WO2024051115A1
Cited By
Training data generation method, electronic equipment, storage medium and product
CN120994799A