Training of Question Answering Model, Question Answering Method, and Device

By obtaining sample questions and answer steps, generating answers using a large language model, and pre-training the model with pre-training data, the existing large language model has solved the problem of low accuracy in answering questions, and achieved higher answer accuracy and model performance.

CN117932015BActive Publication Date: 2025-06-20BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311763895.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20
Estimated Expiration
2043-12-20

AI Technical Summary

Technical Problem

The existing large language models have low answer accuracy when deducing and answering questions.

Method used

By obtaining sample questions and answer steps, using a large language model to generate answers, and pre-training the step planning model and large language model are combined with pre-training data to improve the accuracy of the answer.

Benefits of technology

It improves the accuracy and performance of large language models when answering complex questions, and enhances the overall effect of the question-solving model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117932015B_ABST
    Figure CN117932015B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method and a question answering method, apparatus, electronic device, and readable storage medium for a question answering model. The training method of the question answering model includes: obtaining a first sample question; inputting the first sample question and an answer step extraction template into a large language model to obtain a first sample answer step; inputting the first sample question, the first sample answer step, and an answer extraction template into the large language model to obtain a first sample answer; pre-training a step planning model according to the first sample question and the first sample answer step; pre-training the large language model according to the first sample question, the first sample answer step, and the first sample answer; and obtaining a question answering model according to the pre-trained step planning model and the large language model. The question answering method includes: obtaining a question to be answered; inputting the question to be answered into the step planning model to obtain an answer step; and inputting the question to be answered and the answer step into the large language model to obtain an answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular to artificial intelligence technology fields such as large models, natural language processing, and deep learning. A training method and a question answering method, apparatus, electronic device, and readable storage medium for a question answering model are provided. Background Art

[0002] A large language model (LLM) refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of natural language text, etc. The large language model has a certain reasoning ability, enabling the large language model to answer the input questions and thus obtain answers to the questions. However, in the prior art, when the large language model answers questions through reasoning, there is a problem that the accuracy of the obtained answers is relatively low. Summary of the Invention

[0003] According to a first aspect of the present disclosure, a training method for a question answering model is provided, including: obtaining a first sample question; inputting the first sample question and a solution step extraction template into a large language model to obtain a first sample solution step output by the large language model; inputting the first sample question, the first sample solution step, and an answer extraction template into the large language model to obtain a first sample answer output by the large language model; pre-training a step planning model according to the first sample question and the first sample solution step; pre-training the large language model according to the first sample question, the first sample solution step, and the first sample answer; and obtaining a question answering model according to the pre-trained step planning model and large language model.

[0004] According to a second aspect of the present disclosure, a question answering method is provided, including: obtaining a question to be answered; inputting the question to be answered into a step planning model in the question answering model to obtain a solution step output by the step planning model; and inputting the question to be answered and the solution step into a large language model in the question answering model to obtain an answer output by the large language model.

[0005] According to a third aspect of the present disclosure, there is provided a training device for a question answering model, including: a first acquisition unit configured to acquire a first sample question; a first processing unit configured to input the first sample question and an answer step extraction template into a large language model to obtain a first sample answer step output by the large language model; a second processing unit configured to input the first sample question, the first sample answer step and an answer extraction template into the large language model to obtain a first sample answer output by the large language model; a first pre-training unit configured to pre-train a step planning model according to the first sample question and the first sample answer step; a second pre-training unit configured to pre-train the large language model according to the first sample question, the first sample answer step and the first sample answer; and a construction unit configured to obtain a question answering model according to the pre-trained step planning model and the large language model.

[0006] According to a fourth aspect of the present disclosure, there is provided a question answering device, including: a second acquisition unit configured to acquire a question to be answered; a first answering unit configured to input the question to be answered into the step planning model in the question answering model to obtain an answer step output by the step planning model; and a second answering unit configured to input the question to be answered and the answer step into the large language model in the question answering model to obtain an answer output by the large language model.

[0007] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.

[0008] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.

[0009] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program, which when executed by a processor implements the method as described above.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0013] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;

[0016] Figure 5 is a schematic diagram according to the fifth embodiment of the present disclosure;

[0017] Figure 6 is a block diagram of an electronic device for implementing the training method or the question answering method of the question answering model of the embodiments of the present disclosure. Detailed implementation manners

[0018] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and mechanisms are omitted below for clarity and conciseness.

[0019] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. As Figure 1 shown, the training method of the question answering model in this embodiment specifically includes the following steps:

[0020] S101. Obtain a first sample question;

[0021] S102. Input the first sample question and the answer step extraction template into a large language model to obtain the first sample answer steps output by the large language model;

[0022] S103. Input the first sample question, the first sample answer steps, and the answer extraction template into the large language model to obtain the first sample answer output by the large language model;

[0023] S104. Pre-train a step planning model according to the first sample question and the first sample answer steps;

[0024] S105. Pre-train the large language model according to the first sample question, the first sample answer steps, and the first sample answer;

[0025] S106. Obtain a question answering model according to the step planning model and the large language model obtained through pre-training.

[0026] For the training method of the question answering model in this embodiment, on the one hand, according to the first sample question and the preset answer step extraction template and answer extraction template, use the large language model to respectively obtain the first sample answer steps and the first sample answers, achieving the purpose of obtaining pre-training data according to the large language model and being able to reduce the acquisition cost of pre-training data; on the other hand, pre-train the large language model by combining the first sample answer steps corresponding to the first sample question in the pre-training data, so that the large language model can generate an answer corresponding to the question according to the connection between the question and the answer steps, thereby improving the pre-training effect of the large language model and enabling the large language model in the question answering model to generate more accurate answers.

[0027] In this embodiment, when executing S101, the questions mined from the Internet can be used as the first sample questions; the number of the obtained first sample questions in this embodiment is not limited, and can be one or multiple.

[0028] In this embodiment, when executing S101 to obtain the first sample questions, the answers corresponding to the first sample questions can also be further obtained.

[0029] After this embodiment executes S101 to obtain the first sample questions, execute S102 to input the first sample questions and the answer step extraction template into the large language model to obtain the first sample answer steps output by the large language model; the first sample answer steps obtained in this embodiment are used to represent the answer logic when answering the first sample questions.

[0030] In this embodiment, the answer step extraction template is preset and is used as a prompt to be input into the large language model, so that the large language model combines the answer step extraction template and outputs the corresponding first sample answer steps according to the input first sample questions.

[0031] For example, the answer step extraction template in this embodiment can be: You are a reasoning master. Given a question, you can extract the answer steps corresponding to this question. Please do not give too many details and focus on the core logic. The following are some examples of this task: [Question]: Children play games because they are bored. Who does the "they" in the sentence refer to? A. Children, B. Games; [Answer steps]: 1. Analyze the sentence and find the possible names corresponding to the pronouns; 2. Consider the syntactic structure and semantic logic respectively to judge the possible referents; 3. Combine the syntactic structure and semantic logic to give the final answer.

[0032] When executing S102 in this embodiment, the first sample question and the above-mentioned answer step extraction template are input into the large language model together, so that the large language model combines the answer step extraction template to output the first sample answer steps corresponding to the first sample question.

[0033] It can be understood that if this embodiment executes S101 to obtain multiple first sample questions, when executing S102, this embodiment will respectively obtain the first sample answer steps corresponding to each first sample question according to the answer step extraction template.

[0034] For example, if the first sample question is "A forest ranger needs to move a python, a squirrel, and pinecones from the east to the west of the forest, and can only carry one at a time; when the forest ranger is not present, the python will eat the squirrel, and the squirrel will eat the pinecones; how to ensure their safety?", when executing S102 in this embodiment, this first sample question and the above-mentioned answer step extraction template are input into the large language model together, and the output result of the large language model can be "1. Analyze the mutually exclusive relationships between each object and determine the combinations of objects that cannot coexist; 2. Select a movement strategy that neither violates the mutually exclusive relationships nor violates the mutually exclusive relationships and can achieve the goal, and list the specific operation steps", and this output result is used as the first sample answer steps corresponding to the first sample question.

[0035] After this embodiment executes S102 to obtain the first sample answer steps corresponding to the first sample question, it executes S103 to input the first sample question, the first sample answer steps, and the answer extraction template into the large language model to obtain the first sample answer output by the large language model; the first sample answer in this embodiment includes the answer corresponding to each step of the answer steps.

[0036] In this embodiment, the answer extraction template is preset, and it is used as a prompt to be input into the large language model, so that the large language model combines this answer extraction template and outputs the first sample answer corresponding to the first sample question according to the input first sample question and the first sample answer steps.

[0037] For example, the answer extraction template in this embodiment can be: You are a master of reasoning. Given the solution steps to a problem, you can answer the question according to the logic of the solution steps. Specifically, the requirements are as follows: 1. Please organize the logic of the answer according to the core logic of the given solution steps; 2. Do not include the titles of the relevant solution steps. Please organize the answer in a reasonable structured way, and note that key logic should not be omitted when answering, and please answer in detail; 3. If there are serious errors in the solution steps, please answer according to your own understanding and your own solution steps. The following are some examples of this task:

Question

Solution steps

Answer

[0038] "They are bored" is a subordinate clause used to explain the reason why "Children play games". "They" is more likely to replace the subject in the main clause, that is, "children"; 2. The second thing to consider is the semantic logic; generally speaking, games themselves do not get bored, but people do. "They" should refer to "children"; Combining the syntactic structure and semantic logic, it can be determined that "they" in this sentence refers to "children", so the answer is A. Children.

[0039] When this embodiment executes S103, the first sample question, the first sample solution steps, and the above answer extraction template are input into the large language model together, so that the large language model combines the answer extraction template to output the first sample answer corresponding to the first sample solution steps.

[0040] It can be understood that if this embodiment executes S101 to obtain multiple first sample questions, when this embodiment executes S103, it will obtain the first sample answers corresponding to each first sample question according to the answer extraction template.

[0041] For example, if the first sample problem is "A forest ranger needs to move a python, a squirrel, and a pinecone from the east to the west of the forest, and can only carry one at a time; when the ranger is away, the python will eat the squirrel, and the squirrel will eat the pinecone; how to ensure their safety?", and if the first sample solution steps are "1. Analyze the mutual exclusion relationship between each object and determine the combinations of objects that cannot coexist; 2. Select a movement strategy that neither violates the mutual exclusion relationship nor achieves the goal, and list the specific operation steps", in this embodiment, when executing S103, the first sample problem, the first sample solution steps, and the above answer extraction template are input into the large language model, and the output result of the large language model can be "First, analyze the mutual exclusion relationship between the python, the squirrel, and the pinecone: the python and the squirrel cannot be left alone together, and the squirrel and the pinecone cannot be left alone together"; therefore, to ensure that all things can be safely transported to the west of the forest: 1. First, the ranger takes the squirrel to the west of the forest; 2. Then, the ranger returns and takes the python over; 3. Since the python and the squirrel cannot be left alone together, the ranger brings the squirrel back; next, the ranger takes the pinecone to the west of the forest; finally, the ranger returns and takes the squirrel to the west of the forest".

[0042] After obtaining the first sample answer when executing S103 in this embodiment, the triple composed of the first sample problem, the first sample solution steps corresponding to the first sample problem, and the first sample answer can be directly used as pre-training data, and then this pre-training data is used to pre-train the step planning model and the large language model.

[0043] However, the triples obtained through the large language model cannot guarantee their accuracy. If the accuracy of the triples used as pre-training data is relatively low, it may reduce the pre-training effect of the step planning model and the large language model, and further lead to relatively low accuracy of the solution steps and answers output by the finally obtained problem-solving model.

[0044] Therefore, after obtaining the first sample answer when executing S103 in this embodiment, it may also include the following content: input the first sample problem, the first sample solution steps, the first sample answer, and the data evaluation template into the large language model to obtain the data evaluation result output by the large language model. The data evaluation result obtained in this embodiment can be one of "evaluation passed" and "evaluation not passed", and can also be an evaluation score; when it is determined that the obtained data evaluation result meets the preset requirements (determine that the data evaluation result is "evaluation passed" or the evaluation score exceeds the preset score threshold), the first sample problem, the first sample solution steps, and the first sample answer are used as pre-training data.

[0045] That is to say, in this embodiment, a preset data evaluation template can be used to evaluate the data of the constructed triples using a large language model, so that only the triples that pass the data evaluation are used as pre-training data, improving the accuracy of the obtained pre-training data, and further improving the accuracy of the pre-training of the step planning model and the large language model using the obtained pre-training data.

[0046] In this embodiment, the data evaluation template is preset and is used as a prompt to be input into the large language model, so that the large language model combines this data evaluation template and outputs a data evaluation result according to the input first sample question, first sample answer steps, and first sample answer.

[0047] For example, the data evaluation template in this embodiment can be: You are a reasoning master who can obtain the data evaluation result corresponding to the given question, answer steps, and answer on the premise of the given question, answer steps, and answer. The data evaluation result can be whether it passes the evaluation or an evaluation score. The following are some examples of this task:

Question

[0048]

Answer steps

Answer

[0049] 2. The second thing to consider is the semantic logic; generally speaking, games themselves are not bored, but people are bored, and "they" should refer to "children"; Combining the syntactic structure and semantic logic, it can be determined that "they" in this sentence refers to "children", so the answer is A. Children;

Data evaluation result

[0050] If in this embodiment, it is determined that the obtained data evaluation result does not meet the preset requirements, then the first sample question, first sample answer steps, and first sample answer for this evaluation will be discarded.

[0051] In this embodiment, when performing S103 and determining that the obtained data evaluation result meets the preset requirements, and using the first sample question, the first sample solution steps, and the first sample answer as pre-training data, the following content may also be included: input the first sample question into the data generation model to obtain the candidate solution steps and / or candidate answers output by the data generation model. The data generation model in this embodiment is pre-trained in advance and can output the solution steps and / or answers corresponding to the input question according to the input question; when it is determined that the candidate solution steps are similar to the first sample solution steps (the similarity between the two is greater than or equal to the preset similarity threshold) and / or the candidate answers are similar to the first sample answers (the similarity between the two is greater than or equal to the preset similarity threshold), the first sample question, the first sample solution steps, and the first sample answer are used as pre-training data.

[0052] That is to say, this embodiment can also use the pre-trained data generation model to perform final cleaning on the triples after initially cleaning the obtained triples using the large language model, further improving the accuracy of the obtained pre-training data.

[0053] When this embodiment performs S103, if it is determined that the candidate solution steps are not similar to the first sample solution steps and / or the candidate answers are not similar to the first sample answers, the first sample question, the first sample solution steps, and the first sample answer used this time can be discarded; alternatively, the method of manual annotation can be used to obtain the solution steps and answers corresponding to the first sample question.

[0054] In addition, when this embodiment performs S103, general rules can also be set to discard the triples that do not meet the general rules; for example, the general rules can include a preset number of steps, and when it is determined that the number of steps in the first sample solution steps is less than the preset number of steps, it is determined that the triples do not meet the general rules and are discarded.

[0055] After this embodiment obtains the first sample answer when performing S103, S104 and S105 can be respectively executed to pre-train the step planning model and the large language model; this embodiment does not limit the execution order of S104 and S105.

[0056] When this embodiment pre-trains the step planning model according to the first sample question and the first sample solution steps in S104, the first sample question can be first input into the step planning model to obtain the first predicted solution steps output by the step planning model; obtain the first loss function value according to the first sample solution steps and the first predicted solution steps, and adjust the parameters of the step planning model according to the obtained first loss function value to obtain the pre-trained step planning model.

[0057] In this embodiment, when performing pre-training on the large language model according to the first sample question, the first sample solution steps, and the first sample answer in S105, the first sample question and the first sample solution steps can be first input into the large language model to obtain the first predicted answer output by the large language model; the second loss function value is obtained according to the first sample answer and the first predicted answer, and the parameters of the large language model are adjusted according to the obtained second loss function value to obtain the large language model after pre-training.

[0058] After this embodiment performs pre-training on the step planning model in S104 and pre-training on the large language model in S105, it performs S106 to obtain a question answering model according to the pre-trained step planning model and the large language model.

[0059] When this embodiment performs S106, the pre-trained step planning model and the large language model can be directly obtained to construct a question answering model, so that after the constructed question answering model obtains the input question, the obtained input question is first input into the step planning model to obtain the solution steps output by the step planning model, and then the obtained input question and the solution steps are input into the large language model to obtain the answer output by the large language model.

[0060] When this embodiment performs S106 to obtain a question answering model according to the pre-trained step planning model and the large language model, it may further include the following: obtaining a second sample question and determining the question type of the second sample question; obtaining the solution steps corresponding to the determined question type as the second sample solution steps of the second sample question. In this embodiment, the solution steps corresponding to different question types are obtained by means of pre-annotation; performing supervised fine-tuning (SFT, Supervised Fine-Tuning) on the pre-trained step planning model according to the second sample question and the second sample solution steps; obtaining a question answering model according to the pre-trained large language model and the step planning model obtained by supervised fine-tuning.

[0061] That is to say, this embodiment can also improve the training effect of the step planning model by performing supervised fine-tuning on the step planning model, and then obtain a question answering model according to the step planning model obtained by supervised fine-tuning; and this embodiment obtains the solution steps according to the question type of the sample question, so that the question answering model can output solution steps with similar quantity or logic when facing questions of the same type.

[0062] When this embodiment executes S106 to obtain a question answering model according to the step planning model and the large language model obtained by pre-training, the following content may also be included: obtaining a second sample question and determining the question type of the second sample question; obtaining the answer steps corresponding to the determined question type as the second sample answer steps of the second sample question; determining the answer step type of the second sample answer steps, and obtaining the answer corresponding to the determined answer step type as the second sample answer of the second sample question; performing supervised fine-tuning (SFT, Supervised Fine-Tuning) on the large language model obtained by pre-training according to the second sample question, the second sample answer steps, and the second sample answer; and obtaining a question answering model according to the step planning model obtained by pre-training and the large language model obtained by supervised fine-tuning.

[0063] That is to say, this embodiment can also improve the training effect of the large language model by performing supervised fine-tuning on the large language model, and then obtain a question answering model according to the large language model obtained by supervised fine-tuning; and this embodiment obtains an answer according to the answer type of the sample answer steps, so that the question answering model can output answers that are numerically or logically similar when dealing with the same type of answer steps.

[0064] It can be understood that when this embodiment executes S106, it can also perform supervised fine-tuning on the step planning model and the large language model obtained by pre-training at the same time, and then obtain a question answering model according to the step planning model and the large language model obtained by supervised fine-tuning.

[0065] An example of the training process of the question answering model of this embodiment is as follows: obtaining a first sample question. This embodiment can obtain the first sample question in text format, or after performing speech recognition on the obtained question in speech format, use the recognition result as the first sample question; input the obtained first sample question and the text format answer step extraction template into the large language model to obtain the first sample answer steps in text format output by the large language model; input the first sample question, the first sample answer steps, and the text format answer extraction template into the large language model to obtain the first sample answer in text format output by the large language model; pre-train the step planning model according to the first sample question in text format and the first sample answer steps, and pre-train the large language model according to the first sample question in text format, the first sample answer steps, and the first sample answer; and obtain a question answering model according to the step planning model and the large language model obtained after pre-training.

[0066] Figure 2 It is a schematic diagram according to the second embodiment of the present disclosure. Figure 2The structural diagram of the question answering model of this embodiment is shown: The question answering model in this embodiment includes a step planning model and a large language model. The step planning model is used to output answer steps according to the input question, and the large language model is used to output an answer corresponding to the input question according to the input question and the answer steps output by the step planning model; The output answer may include multiple sub-answers, and each sub-answer corresponds to an answer step.

[0067] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure. As Figure 3 shown, the question answering method of this embodiment specifically includes the following steps:

[0068] S301. Obtain the question to be answered;

[0069] S302. Input the question to be answered into the step planning model in the question answering model, and obtain the answer steps output by the step planning model;

[0070] S303. Input the question to be answered and the answer steps into the large language model in the question answering model, and obtain the answer output by the large language model.

[0071] That is to say, this embodiment combines the pre-trained question answering model, and adopts the method of first obtaining the answer steps of the question to be answered and then obtaining the answer according to the obtained answer steps, which can improve the accuracy of the answer corresponding to the question to be answered and enhance the performance of the large language model in the question answering model when answering complex reasoning questions.

[0072] An example is given to illustrate the answering process of the question in this embodiment: Obtain the question to be answered. In this embodiment, the question to be answered in text format input at the input end can be obtained, or after speech recognition of the question in voice format input at the input end, the recognition result can be used as the question to be answered; Input the question to be answered into the step planning model in the question answering model, and obtain the answer steps in text format output by the step planning model; Input the question to be answered in text format and the answer steps into the large language model in the question answering model, and obtain the answer in text format output by the large language model as the answer result corresponding to the question to be answered, and return the obtained answer to the input end to be displayed to the user.

[0073] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure. As Figure 4 shown, the training device 400 of the question answering model of this embodiment includes:

[0074] The first obtaining unit 401 is used to obtain the first sample question;

[0075] The first processing unit 402 is configured to input the first sample question and the answer step extraction template into the large language model to obtain the first sample answer steps output by the large language model;

[0076] The second processing unit 403 is configured to input the first sample question, the first sample answer steps, and the answer extraction template into the large language model to obtain the first sample answer output by the large language model;

[0077] The first pre-training unit 404 is configured to pre-train the step planning model according to the first sample question and the first sample answer steps;

[0078] The second pre-training unit 405 is configured to pre-train the large language model according to the first sample question, the first sample answer steps, and the first sample answer;

[0079] The construction unit 406 is configured to obtain the question answering model according to the pre-trained step planning model and the large language model.

[0080] The first acquisition unit 401 may use the questions mined from the Internet as the first sample questions; in this embodiment, the number of the acquired first sample questions is not limited, and it may be one or multiple.

[0081] When the first acquisition unit 401 acquires the first sample questions, it may further acquire the answers corresponding to the first sample questions.

[0082] In this embodiment, after the first acquisition unit 401 acquires the first sample questions, the first processing unit 402 inputs the first sample questions and the answer step extraction template into the large language model to obtain the first sample answer steps output by the large language model; the first sample answer steps obtained by the first processing unit 402 are used to represent the answering logic when answering the first sample questions.

[0083] In this embodiment, the answer step extraction template is pre-set and is used as a prompt to be input into the large language model, so that the large language model combines the answer step extraction template and outputs the corresponding first sample answer steps according to the input first sample questions.

[0084] The first processing unit 402 inputs the first sample questions and the above answer step extraction template into the large language model together, so that the large language model combines the answer step extraction template to output the first sample answer steps corresponding to the first sample questions.

[0085] It can be understood that if multiple first sample questions are acquired, the first processing unit 402 will respectively acquire the first sample answer steps corresponding to each first sample question according to the answer step extraction template.

[0086] After the step of obtaining the first sample answer corresponding to the first sample question by the first processing unit 402 in this embodiment, the second processing unit 403 inputs the first sample question, the first sample answer step, and the answer extraction template into the large language model to obtain the first sample answer output by the large language model; the first sample answer obtained by the second processing unit 403 includes answers corresponding to each answer step.

[0087] In this embodiment, the answer extraction template is preset and is used as a prompt to be input into the large language model, so that the large language model combines this answer extraction template and outputs the first sample answer corresponding to the first sample question according to the input first sample question and the first sample answer step.

[0088] The second processing unit 403 inputs the first sample question, the first sample answer step, and the above answer extraction template into the large language model together, so that the large language model combines the answer extraction template to output the first sample answer corresponding to the first sample answer step.

[0089] It can be understood that if multiple first sample questions are obtained, the second processing unit 403 will respectively obtain the first sample answers corresponding to each first sample question according to the answer extraction template.

[0090] After the second processing unit 403 obtains the first sample answer, it can directly use the triple composed of the first sample question, the first sample answer step corresponding to the first sample question, and the first sample answer as pre-training data, and then use this pre-training data to pre-train the step planning model and the large language model.

[0091] However, the triples obtained through the large language model cannot guarantee their accuracy. If the accuracy of the triples used as pre-training data is relatively low, it may reduce the pre-training effect of the step planning model and the large language model, and further lead to relatively low accuracy of the answer steps and answers output by the finally obtained question answering model.

[0092] Therefore, after the second processing unit 403 obtains the first sample answer, it may further include the following: inputting the first sample question, the first sample answer step, the first sample answer, and the data evaluation template into the large language model to obtain the data evaluation result output by the large language model; when it is determined that the obtained data evaluation result meets the preset requirements, using the first sample question, the first sample answer step, and the first sample answer as pre-training data.

[0093] That is to say, the second processing unit 403 can use a large language model to evaluate the formed triples with the help of a preset data evaluation template, so as to use only the triples that pass the data evaluation as pre-training data, improving the accuracy of the obtained pre-training data, and further improving the accuracy of the pre-training of the step planning model and the large language model using the obtained pre-training data.

[0094] In this embodiment, the data evaluation template is preset and is used as a prompt to be input into the large language model, so that the large language model combines the data evaluation template and outputs a data evaluation result according to the input first sample question, first sample answer steps, and first sample answer.

[0095] If in this embodiment it is determined that the obtained data evaluation result does not meet the preset requirements, then the first sample question, first sample answer steps, and first sample answer that are being evaluated this time are discarded.

[0096] When the second processing unit 403 determines that the obtained data evaluation result meets the preset requirements and uses the first sample question, first sample answer steps, and first sample answer as pre-training data, it may further include the following: input the first sample question into the data generation model to obtain candidate answer steps and / or candidate answers output by the data generation model; when it is determined that the candidate answer steps are similar to the first sample answer steps and / or the candidate answers are similar to the first sample answers, use the first sample question, first sample answer steps, and first sample answer as pre-training data.

[0097] That is to say, the second processing unit 403 can also use the pre-trained data generation model to perform final cleaning on the triples after initially cleaning the obtained triples using the large language model, further improving the accuracy of the obtained pre-training data.

[0098] If it is determined that the candidate answer steps are not similar to the first sample answer steps and / or the candidate answers are not similar to the first sample answers, the second processing unit 403 can discard the first sample question, first sample answer steps, and first sample answer used this time; it can also use the method of manual annotation to obtain the answer steps and answers corresponding to the first sample question.

[0099] In addition, the second processing unit 403 can also set general rules to discard the triples that do not meet the general rules.

[0100] After the second processing unit 403 obtains the first sample answer in this embodiment, the first pre-training unit 404 and the second pre-training unit 405 can respectively perform pre-training on the step planning model and the large language model.

[0101] When the first pre-training unit 404 pre-trains the step planning model according to the first sample question and the first sample answer steps, it can first input the first sample question into the step planning model to obtain the first predicted answer steps output by the step planning model; obtain the first loss function value according to the first sample answer steps and the first predicted answer steps, and adjust the parameters of the step planning model according to the obtained first loss function value to obtain the step planning model after pre-training.

[0102] When the second pre-training unit 405 pre-trains the large language model according to the first sample question, the first sample answer steps and the first sample answer, it can first input the first sample question and the first sample answer steps into the large language model to obtain the first predicted answer output by the large language model; obtain the second loss function value according to the first sample answer and the first predicted answer, and adjust the parameters of the large language model according to the obtained second loss function value to obtain the large language model after pre-training.

[0103] In this embodiment, after the first pre-training unit 404 and the second pre-training unit 405 complete pre-training, the construction unit 406 obtains a question answering model according to the step planning model and the large language model obtained by pre-training.

[0104] The construction unit 406 can directly obtain the step planning model and the large language model after pre-training to construct a question answering model, so that the constructed question answering model, after obtaining the input question, first inputs the obtained input question into the step planning model to obtain the answer steps output by the step planning model, and then inputs the obtained input question and the answer steps into the large language model to obtain the answer output by the large language model.

[0105] When the construction unit 406 obtains a question answering model according to the step planning model and the large language model obtained by pre-training, it may further include the following: obtaining a second sample question and determining the question type of the second sample question; obtaining the answer steps corresponding to the determined question type as the second sample answer steps of the second sample question; performing supervised fine-tuning (SFT, Supervised Fine-Tuning) on the step planning model obtained by pre-training according to the second sample question and the second sample answer steps; obtaining a question answering model according to the large language model obtained by pre-training and the step planning model obtained by supervised fine-tuning.

[0106] That is to say, the construction unit 406 can also improve the training effect of the step planning model by performing supervised fine-tuning on the step planning model, and then obtain the problem-solving model according to the step planning model obtained by the supervised fine-tuning; and in this embodiment, the solution steps are obtained according to the problem type of the sample problem, so that when the problem-solving model is for the same type of problem, it can output solution steps that are numerically or logically similar.

[0107] When the construction unit 406 obtains the problem-solving model according to the step planning model and the large language model obtained by pre-training, it may further include the following: obtaining a second sample problem and determining the problem type of the second sample problem; obtaining the solution steps corresponding to the determined problem type as the second sample solution steps of the second sample problem; determining the solution step type of the second sample solution steps, and obtaining the answer corresponding to the determined solution step type as the second sample answer of the second sample problem; performing supervised fine-tuning (SFT, Supervised Fine-Tuning) on the large language model obtained by pre-training according to the second sample problem, the second sample solution steps, and the second sample answer; obtaining the problem-solving model according to the step planning model obtained by pre-training and the large language model obtained by supervised fine-tuning.

[0108] That is to say, the construction unit 406 can also improve the training effect of the large language by performing supervised fine-tuning on the large language model, and then obtain the problem-solving model according to the large language obtained by the supervised fine-tuning; and in this embodiment, the answer is obtained according to the solution type of the sample solution steps, so that when the problem-solving model is for the same type of solution steps, it can output answers that are numerically or logically similar.

[0109] It can be understood that the construction unit 406 can also perform supervised fine-tuning on the step planning model and the large language model obtained by pre-training at the same time, and then obtain the problem-solving model according to the step planning model and the large language model obtained by the supervised fine-tuning.

[0110] Figure 5 It is a schematic diagram according to the fifth embodiment of the present disclosure. As Figure 5 shown, the problem-solving device 500 of this embodiment includes:

[0111] A second acquisition unit 501, configured to acquire a problem to be solved;

[0112] A first solution unit 502, configured to input the problem to be solved into the step planning model in the problem-solving model, and acquire the solution steps output by the step planning model;

[0113] A second solution unit 503, configured to input the problem to be solved and the solution steps into the large language model in the problem-solving model, and acquire the answer output by the large language model.

[0114] In the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0115] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0116] As Figure 6 shown, it is a block diagram of an electronic device for a training method or a question answering method of a question answering model according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0117] As Figure 6 shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0118] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the training method or the question answering method of the question answering model. For example, in some embodiments, the training method or the question answering method of the question answering model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 608.

[0120] In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the training method or the question answering method of the question answering model described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the training method or the question answering method of the question answering model by any other suitable means (e.g., by means of firmware).

[0121] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable apparatus for vehicle positioning or training of a positioning model, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on a remote machine or server.

[0123] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can include or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for presenting information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0125] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0126] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (“Virtual Private Server”, or simply “VPS”). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0127] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0128] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A training method for a question answering model, comprising: Obtain the first sample question; Input the first sample question and the answer step extraction template into the large language model to obtain the first sample answer steps output by the large language model. The first sample answer steps are used for the large language model to obtain the first sample answer corresponding to the first sample question and for pre-training the step planning model and the large language model; Input the first sample question, the first sample answer steps and the answer extraction template into the large language model to obtain the first sample answer output by the large language model; Pre-train the step planning model according to the first sample question and the first sample answer steps; Pre-train the large language model according to the first sample question, the first sample answer steps and the first sample answer; Obtain a question answering model according to the pre-trained step planning model and large language model; The method further includes: Input the first sample question, the first sample answer steps, the first sample answer and the data evaluation template into the large language model to obtain the data evaluation result output by the large language model; When it is determined that the data evaluation result meets the preset requirements, use the first sample question, the first sample answer steps and the first sample answer as pre-training data.

2. The method according to claim 1, wherein, The using the first sample question, the first sample answer steps and the first sample answer as pre-training data includes: Input the first sample question into the data generation model to obtain the candidate answer steps and / or candidate answers output by the data generation model; When it is determined that the candidate answer steps are similar to the first sample answer steps and / or the candidate answers are similar to the first sample answers, use the first sample question, the first sample answer steps and the first sample answer as the pre-training data.

3. The method according to claim 1, wherein, The obtaining a question answering model according to the pre-trained step planning model and large language model includes: Obtain a second sample question and determine the question type of the second sample question; Obtain the answer steps corresponding to the question type as the second sample answer steps of the second sample question; Perform supervised fine-tuning on the pre-trained step planning model according to the second sample question and the second sample answer steps; Obtain the question answering model according to the pre-trained large language model and the step planning model obtained by supervised fine-tuning.

4. The method according to claim 1, wherein, The obtaining a question answering model according to the pre-trained step planning model and large language model includes: Obtain a second sample question and determine the question type of the second sample question; Obtain the answer steps corresponding to the question type as the second sample answer steps of the second sample question; Determine the answer step type of the second sample answer steps, and obtain the answer corresponding to the answer step type as the second sample answer of the second sample question; Perform supervised fine-tuning on the pre-trained large language model according to the second sample question, the second sample answer steps and the second sample answer; Obtain the problem-solving model according to the step planning model obtained by pre-training and the large language model obtained by supervised fine-tuning.

5. The method according to claim 1, wherein, The obtaining of the problem-solving model according to the step planning model obtained by pre-training and the large language model includes: Obtain a second sample problem and determine the problem type of the second sample problem; Obtain the solution steps corresponding to the problem type as the second sample solution steps of the second sample problem; Determine the solution step type of the second sample solution steps, and obtain the answer corresponding to the solution step type as the second sample answer of the second sample problem; Perform supervised fine-tuning on the step planning model obtained by pre-training according to the second sample problem and the second sample solution steps; Perform supervised fine-tuning on the large language model obtained by pre-training according to the second sample problem, the second sample solution steps and the second sample answer; Obtain the problem-solving model according to the step planning model and the large language model obtained by supervised fine-tuning.

6. The method according to claim 1, wherein, The pre-training of the step planning model according to the first sample problem and the first sample solution steps includes: Input the first sample problem into the step planning model to obtain the first predicted solution steps output by the step planning model; Obtain the first loss function value according to the first sample solution steps and the first predicted solution steps; Adjust the parameters of the step planning model according to the first loss function value to obtain the step planning model after pre-training.

7. The method according to claim 1, wherein The pre-training of the large language model according to the first sample problem, the first sample solution steps and the first sample answer includes: Input the first sample problem and the first sample solution steps into the large language model to obtain the first predicted answer output by the large language model; Obtain the second loss function value according to the first sample answer and the first predicted answer; Adjust the parameters of the large language model according to the second loss function value to obtain the large language model after pre-training.

8. A problem-solving method, comprising: Obtain the problem to be solved; Input the problem to be solved into the step planning model in the problem-solving model to obtain the solution steps output by the step planning model; Input the problem to be solved and the solution steps into the large language model in the problem-solving model to obtain the answer output by the large language model; The problem-solving model is trained according to the method of any one of claims 1 to 7.

9. A training device for a problem-solving model, comprising: A first obtaining unit for obtaining a first sample problem; A first processing unit for inputting the first sample problem and the solution step extraction template into the large language model to obtain the first sample solution steps output by the large language model, where the first sample solution steps are used for the large language model to obtain the first sample answer corresponding to the first sample problem and for pre-training the step planning model and the large language model; A second processing unit for inputting the first sample problem, the first sample solution steps and the answer extraction template into the large language model to obtain the first sample answer output by the large language model; The first pre-training unit is used to pre-train the step planning model according to the first sample question and the first sample solution steps; The second pre-training unit is used to pre-train the large language model according to the first sample question, the first sample solution steps and the first sample answer; The construction unit is used to obtain a question answering model according to the pre-trained step planning model and the large language model; The second processing unit is further configured to execute: Input the first sample question, the first sample solution steps, the first sample answer and the data evaluation template into the large language model to obtain the data evaluation result output by the large language model; When it is determined that the data evaluation result meets the preset requirements, use the first sample question, the first sample solution steps and the first sample answer as pre-training data.

10. The device according to claim 9, wherein When the second processing unit uses the first sample question, the first sample solution steps and the first sample answer as pre-training data, it specifically executes: Input the first sample question into the data generation model to obtain candidate solution steps and / or candidate answers output by the data generation model; When it is determined that the candidate solution steps are similar to the first sample solution steps and / or the candidate answers are similar to the first sample answers, use the first sample question, the first sample solution steps and the first sample answer as the pre-training data.

11. The device according to claim 9, wherein When the construction unit obtains a question answering model according to the pre-trained step planning model and the large language model, it specifically executes: Obtain a second sample question and determine the question type of the second sample question; Obtain the solution steps corresponding to the question type as the second sample solution steps of the second sample question; Perform supervised fine-tuning on the pre-trained step planning model according to the second sample question and the second sample solution steps; Obtain the question answering model according to the pre-trained large language model and the step planning model obtained by supervised fine-tuning.

12. The device according to claim 9, wherein When the construction unit obtains a question answering model according to the pre-trained step planning model and the large language model, it specifically executes: Obtain a second sample question and determine the question type of the second sample question; Obtain the solution steps corresponding to the question type as the second sample solution steps of the second sample question; Determine the solution step type of the second sample solution steps, and obtain the answer corresponding to the solution step type as the second sample answer of the second sample question; Perform supervised fine-tuning on the pre-trained large language model according to the second sample question, the second sample solution steps and the second sample answer; Obtain the question answering model according to the pre-trained step planning model and the large language model obtained by supervised fine-tuning.

13. The device according to claim 9, wherein When the construction unit obtains a question answering model according to the pre-trained step planning model and the large language model, it specifically executes: Obtain a second sample question and determine the question type of the second sample question; Obtain the solution steps corresponding to the problem type as the second sample solution steps of the second sample problem; Determine the solution step type of the second sample solution steps, and obtain the answer corresponding to the solution step type as the second sample answer of the second sample problem; Perform supervised fine-tuning on the pre-trained step planning model according to the second sample problem and the second sample solution steps; Perform supervised fine-tuning on the pre-trained large language model according to the second sample problem, the second sample solution steps and the second sample answer; Obtain the problem-solving model according to the step planning model and the large language model obtained by supervised fine-tuning.

14. The apparatus according to claim 9, wherein, When the first pre-training unit pre-trains the step planning model according to the first sample problem and the first sample solution steps, it specifically executes: Input the first sample problem into the step planning model to obtain the first predicted solution steps output by the step planning model; Obtain the first loss function value according to the first sample solution steps and the first predicted solution steps; Adjust the parameters of the step planning model according to the first loss function value to obtain the step planning model after pre-training.

15. The apparatus according to claim 9, wherein, When the second pre-training unit pre-trains the large language model according to the first sample problem, the first sample solution steps and the first sample answer, it specifically executes: Input the first sample problem and the first sample solution steps into the large language model to obtain the first predicted answer output by the large language model; Obtain the second loss function value according to the first sample answer and the first predicted answer; Adjust the parameters of the large language model according to the second loss function value to obtain the large language model after pre-training.

16. A problem-solving apparatus, comprising: A second acquisition unit, configured to acquire the problem to be solved; A first solution unit, configured to input the problem to be solved into the step planning model in the problem-solving model to obtain the solution steps output by the step planning model; A second solution unit, configured to input the problem to be solved and the solution steps into the large language model in the problem-solving model to obtain the answer output by the large language model; The problem-solving model is trained according to the device of any one of claims 9 to 15.

17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method of any one of claims 1 to 8.

19. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Generation method, model training method, equipment and storage medium

    CN116821308A

  • Construction method and application of question and answer interaction model with cognitive reasoning ability

    CN116991996A