Question answering method, apparatus and device, and program product

By introducing the correction steps of the comment model in the initial replies output by the large language model, the problem that the answer accuracy of the large language model is difficult to guarantee on the problem of high logic is solved, and higher answer accuracy and accuracy of the answer steps are achieved.

CN119990313APending Publication Date: 2025-05-13IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510056782.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When answering questions with many steps and high logic, existing large language models face problems that are difficult to guarantee the correctness of the answer.

Method used

By entering the initial replies output from the question and model into the comment model, evaluation information for each answer step is obtained, and the initial replies are corrected based on this until the error-free answering step is reached.

Benefits of technology

It significantly improves the accuracy of the answers to the questions, ensures the accuracy of the answering steps, and improves the model's performance on highly logical questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990313A_ABST
    Figure CN119990313A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a question answering method and device, equipment and a program product. The method comprises the steps that a target question is input into an answering model, an initial reply output by the answering model is obtained, and the initial reply comprises at least one answering step; the target question and the initial reply are input into a comment model, comments output by the comment model are obtained, and the comments comprise evaluation information of all answering steps in the initial reply; the initial reply is corrected based on the comment, a target reply corresponding to the target question is obtained, the answer model is obtained after a large language model is trained based on at least one first question and answer pair, and a first answer in the first question and answer pair comprises an answer obtained after a first comment output based on the comment model is corrected. According to the method, the answer output by the model can be corrected on the aspect of answering steps through comments, and the question answer with higher correctness is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device, equipment and program product for answering a question. Background Art

[0002] In recent years, large language models have received increasing attention. Existing research shows that large language models have certain potential in solving problems such as mathematical problems. However, for problems such as mathematical problems that have many steps and strong logic, large language models still face great challenges in solving such problems. At present, there are three common methods for solving problems with strong logic based on large language models. The first method directly corrects the defective large language model during the training process, but this method directly uses feedback information to fine-tune the model parameters, which is costly and time-consuming, and cannot fine-tune closed-source models. The second method is to guide the model output through automatic feedback during the generation stage, but the accuracy of the answer output by this method depends on the high quality of the automatic feedback. The third method is to make corrections after the large language model is generated, without updating the model parameters, but this method cannot guarantee the accuracy of the output results. Summary of the invention

[0003] Based on the above-mentioned defects and shortcomings of the prior art, the present application proposes a question answering method, device, equipment and program product, which can correct the answers output by the model at the level of the answering steps through comments, and obtain more accurate answers to the questions.

[0004] According to a first aspect of the present application, a method for answering a question is provided, comprising: inputting a target question into a question answering model to obtain an initial answer output by the question answering model, wherein the initial answer includes at least one question answering step; inputting the target question and the initial answer into a comment model to obtain a comment output by the comment model, wherein the comment includes evaluation information for each of the question answering steps in the initial answer; correcting the initial answer based on the comment to obtain a target answer corresponding to the target question, wherein the question answering model is obtained by training a large language model based on at least one first question-answer pair, and a first answer in the first question-answer pair includes an answer obtained by correcting the first comment output by the comment model.

[0005] According to the question answering method provided in the first aspect of the present application, the first answer is obtained as follows: the first question is input into the large language model to obtain a first initial reply output by the large language model; the first question and the first initial reply are input into the comment model to obtain the first comment output by the comment model; it is determined whether the first comment indicates that there are no incorrect answering steps in the first initial reply, and if not, the first initial reply is corrected based on the first comment to obtain a first corrected reply, the first corrected reply and the first question are re-input into the comment model to obtain the first comment output by the comment model again, until the first comment indicates that there are no incorrect answering steps in the first corrected reply, and the first corrected reply is used as the first answer; if so, the first initial reply is used as the first answer.

[0006] According to the question answering method provided in the first aspect of the present application, the comment model is obtained after training the initial comment model based on the second question, the second answer and the second comment; after correcting the initial reply based on the comment to obtain the target reply corresponding to the target question, it also includes: updating the target question to the second question, updating the target reply to the second answer, and updating the comment to the second comment; based on the updated second question, the updated second answer and the updated second comment, the comment model is trained again to obtain an optimized comment model.

[0007] According to the question answering method provided in the first aspect of the present application, after obtaining the optimized comment model, it also includes: based on the optimized comment model, re-obtaining the first optimized question and answer pair, wherein the first optimized answer in the first optimized question and answer pair is obtained after correcting the first optimized comment output by the optimized comment model; based on the first optimized question and answer pair, the answering model is trained again to obtain the optimized answering model.

[0008] According to the question answering method provided in the first aspect of the present application, the training process of the comment model includes: correcting the second answer based on the second comment to obtain a second corrected answer; obtaining a first correctness corresponding to the second answer, and obtaining a second correctness corresponding to the second corrected answer; determining a reward in the initial comment model training process based on the difference between the first correctness and the second correctness; with the goal of maximizing the reward, training the initial comment model based on the second question, the second answer and the second comment to obtain the comment model.

[0009] According to the question answering method provided in the first aspect of the present application, the obtaining of the first correctness corresponding to the second answer and the obtaining of the second correctness corresponding to the second revised answer include: inputting the second question and the second answer into a correction model to obtain the first correctness output by the correction model; inputting the second question and the second revised answer into the correction model to obtain the second correctness output by the correction model; wherein the correction model is obtained by training an initial correction model based on a third question, a standard answer, a third answer and a third correctness, and the third answer includes an answer obtained based on the large language model or the question answering model.

[0010] According to the question answering method provided in the first aspect of the present application, the initial answer is corrected based on the comments to obtain the target answer corresponding to the target question, including: inputting the comments into the question answering model, so that the question answering model corrects the initial answer based on the comments to obtain the target answer output by the question answering model.

[0011] According to a second aspect of the present application, a question answering device is provided, comprising: an initial question answering module, used to input a target question into a question answering model, and obtain an initial answer output by the question answering model; a comment output module, used to input the target question and the initial answer into a comment model, and obtain a comment output by the comment model, wherein the comment includes evaluation information for each question answering step in the initial answer; an answer correction module, used to correct the initial answer based on the comment, and obtain a target answer corresponding to the target question, wherein the question answering model is obtained after training a large language model based on at least one first question-answer pair, and the first answer in each of the first question-answer pairs includes an answer obtained by correcting the first comment output by the comment model.

[0012] According to a third aspect of the present application, an electronic device is provided, comprising: a memory and a processor; the memory is connected to the processor and is used to store a program; the processor is used to implement the question-solving method as described in the first aspect by running the program in the memory.

[0013] According to a fourth aspect of the present application, a computer program product is provided, comprising computer program instructions; when the computer program instructions are executed by a processor, the processor is caused to execute the problem-solving method as described in the first aspect.

[0014] In the present application, the target question is input into the answering model to obtain the initial answer output by the answering model, wherein the initial answer includes at least one answering step; the target question and the initial answer are input into the comment model to obtain the comment output by the comment model, wherein the comment includes the evaluation information of each answering step in the initial answer; the initial answer is corrected based on the comment to obtain the target answer corresponding to the target question, wherein the answering model is obtained after the large language model is trained based on at least one first question-answer pair, and the first answer in the first question-answer pair includes the answer obtained after correction based on the first comment output by the comment model. In the above process, the first answer in the first question-answer pair used in the training of the answering model includes the answer obtained after correction by the first comment output by the comment model, which improves the correctness of the first answer, thereby improving the performance of the answering model obtained based on the first question-answer pair training, and improving the correctness of the initial answer at the level of the answering model. Furthermore, since the answering model is obtained based on the training of the large language model, the initial answer initially output by the answering model has a certain error rate. This application corrects the results output by the answering model at the level of answering steps based on the comments output by the comment model, and further improves the accuracy of the target answer based on the initial answer. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0016] Figure 1 A flowchart of a method for answering a question provided in an embodiment of the present application;

[0017] Figure 2 A schematic diagram of a process for obtaining a first answer provided in an embodiment of the present application;

[0018] Figure 3 A schematic diagram of an iterative optimization process of a comment model and an answer model provided in an embodiment of the present application;

[0019] Figure 4 A schematic diagram of the training principle of a correction model provided in an embodiment of the present application;

[0020] Figure 5 A schematic diagram of a review model training principle provided in an embodiment of the present application;

[0021] Figure 6 A schematic diagram of the reinforcement learning principle of a review model provided in an embodiment of the present application;

[0022] Figure 7 A schematic diagram of an iterative optimization principle provided in an embodiment of the present application;

[0023] Figure 8 A block diagram of a question answering device provided in an embodiment of the present application;

[0024] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0026] Application Overview

[0027] A large language model (LLM) refers to a deep learning model with a large scale and a large number of parameters, which usually performs well in natural language processing tasks. By providing natural language feedback to the large language model, it can be effectively guided to correct errors and improve the performance of the large language model. For the answer output by the large language model for a certain question, the errors in the output of the large language model can be pointed out based on critical evaluation of manual feedback, including misunderstanding of the question, application errors of theorem, calculation errors, etc. Based on manual feedback, the large language model can modify its output to get a more accurate answer. This method is in line with the natural communication habits of humans and has strong interpretability, so it has received more and more attention. However, the cost and time of obtaining human feedback online are high, and it cannot be applied on a large scale.

[0028] At present, among the existing methods of correcting the output results of large language models through feedback, the first type directly corrects the defective large language models during the training process, such as reinforcement learning based on human feedback (RLHF), which is not only costly and time-consuming, but also unable to fine-tune closed-source models such as GPT-4. The second type guides the model output through automatic feedback in the generation stage. For example, in the process of answering proof questions with a large language model, the feedback provided by the intermediate reasoning steps is used to correct errors, thereby improving the answering effect. However, it will rely heavily on the quality of automatic feedback. If the feedback itself is not accurate or biased, it may lead to error accumulation and even produce erroneous output. The effect is difficult to guarantee and the interpretability is poor. The third type is to make corrections after the large language model is generated. There is no need to update the model parameters. However, for questions with many steps and strong logic, if one step of the answering step goes wrong, the entire answer may be wrong. At present, feedback is received after the model is generated, and it is difficult to locate where the error occurred. Moreover, the reward function of reinforcement learning of existing models often comes from human feedback and has low accuracy. Furthermore, the model parameters of existing models are often only corrected once, and the effect of improving the accuracy of answers to questions is limited.

[0029] Exemplary Methods

[0030] In response to at least one problem existing in the prior art, the present application provides a method for answering a question, which automatically completes the answer to the question through a software algorithm, obtains a target answer to the target question, and improves the correctness of the target answer while realizing the automation of the question answering. The software algorithm for implementing the method for answering the question can be run on any device with data processing capabilities, such as a local computer, a cloud server, a smart mobile device, etc. Optionally, the device running the software algorithm corresponding to the method for answering the question has a human-computer interaction function, so as to facilitate manual input or acquisition of information.

[0031] In one embodiment, if Figure 1 As shown, the process steps implemented by the question-solving method include:

[0032] Step 101, input the target question into the answering model to obtain an initial answer output by the answering model, wherein the initial answer includes at least one answering step.

[0033] In this embodiment, the target question refers to the question to be answered, and the question types of the target question include but are not limited to math questions, physics questions and other types. The standard answer corresponding to the target question includes multiple answering steps and strong logic. Taking the target question as a math question as an example, solving a math question usually requires at least one answering step, and the more complex the math question, the more answering steps.

[0034] In this embodiment, the answering model is a model obtained by further training on the basis of the large language model, and can better complete the answer to the target question than the untrained large language model. After the target question is input into the answering model, the answering model can directly output an initial response. The initial response includes at least one answering step. Optionally, the initial response may include all the answering steps to answer the target question, or may only include some of the answering steps to answer the target question; optionally, the initial response includes one or more answering steps starting from the first answering step and continuing. However, since the initial response is the answer directly output by the answering model based on the target question, there may be wrong answering steps in the initial response.

[0035] Step 102, input the target question and the initial answer into the comment model to obtain the comment output by the comment model, wherein the comment includes evaluation information of each answering step in the initial answer.

[0036] In this embodiment, in order to improve the accuracy of the answer, the target question and the initial answer are input into the comment model, and each answering step in the initial answer is evaluated at the level of the answering step. The comment model is a model trained on the basis of the deep learning model, which can evaluate each answering step in the initial answer and output comments based on the answering steps. Optionally, the evaluation information of each answering step in the comment includes information evaluating the correctness of any one or several answering steps, and also includes prompt information for any error in the answering step, and can also include description information on how to correct the wrong answering step. Specifically, the prompt information of the wrong answering step can prompt various types of errors such as formula logic errors, calculation errors, formula reference errors, etc. that appear in the wrong answering step. The corrected description information includes information on the correction of one or more types of errors such as formula logic errors, calculation errors, formula reference errors, etc. that exist in any wrong answering step.

[0037] Step 103, modify the initial response based on the comments to obtain a target response corresponding to the target question, wherein the question answering model is obtained by training a large language model based on at least one pair of first question-answer pairs, and the first answer in each pair of first question-answer pairs includes an answer modified based on the first comment output by the comment model.

[0038] In this embodiment, after obtaining the comments on the initial answer, if the comments indicate that each answering step of the initial answer is correct, then the initial answer can be directly used as the final target answer for the target question. If the comments indicate that there is an error in at least one answering step in the initial answer, the initial answer is corrected based on the comments to obtain the final target answer. Optionally, the process of correcting the initial answer based on the comments can be implemented through a pre-set correction processing logic. For example, if a formula reference error occurs in the initial answer, it is directly replaced based on the correct formula provided by the correction information in the comments. Optionally, the process of correcting the initial answer based on the comments can also be automatically completed through a large language model, thereby improving the degree of automation of the correction process. The large language model can be a large language model independent of the answering model, or it can be a question answering model.

[0039] In this embodiment, when the large language model is trained to obtain the answering model, the training samples used include at least one pair of first question-answer pairs, and each first question-answer pair includes a first question and a first answer corresponding to the first question. During the training process, the higher the correctness of the first answer in the first question-answer pair, the better the training effect of the answering model. Therefore, the standard answer to the first question can be used as the first answer. However, the training sample requires a large amount of data, and manually configuring the standard answers to many first questions will waste a lot of manpower and material resources. Therefore, a large number of first questions can be processed separately by the large language model. On the basis of the initial answer output by the large language model, the comment model is used to correct the initial answer output by the large language model for the comment of the initial answer output by the large language model, and the corrected target answer is used as the first answer to form the first question-answer pair. Through the large language model and the comment model, the first question-answer pair with higher correctness can be generated quickly and efficiently, thereby providing a large number of reliable training samples for the training of the answering model, thereby improving the training efficiency and training effect of the answering model.

[0040] In one embodiment, the initial answer is modified based on the comments to obtain a target answer corresponding to the target question, including: inputting the comments into the answering model so that the answering model modifies the initial answer based on the comments to obtain the target answer output by the answering model.

[0041] In this embodiment, the automatic correction of the initial reply is completed through the question answering model. Since the question answering model itself is a trained large language model, it can correct the initial reply based on the comments based on the excellent ability of the large language model to process natural language tasks, thereby improving the intelligence of the initial reply correction process. In addition, since the question answering model has already learned the initial reply, after the question answering model obtains the comments of the initial reply, the correction process is automatically triggered to output the target reply, avoiding problems such as excessive computing volume or excessive computing pressure caused by the use of multiple models. And from the user's perspective, it is more convenient to know the final target reply, or to know the initial reply and comments according to needs, thereby improving the user experience.

[0042] In one embodiment, the first answer is obtained as follows: the first question is input into the large language model to obtain a first initial reply output by the large language model; the first question and the first initial reply are input into the comment model to obtain a first comment output by the comment model; it is determined whether the first comment indicates that there are no incorrect answering steps in the first initial reply, if not, the first initial reply is revised based on the first comment to obtain a first revised reply, the first revised reply and the first question are re-input into the comment model to obtain the first comment output by the comment model again, until the first comment indicates that there are no incorrect answering steps in the first revised reply, and the first revised reply is used as the first answer; if so, the first initial reply is used as the first answer.

[0043] In this embodiment, in order to obtain a higher quality first question-answer pair, the comment model can be repeatedly called. Specifically, Figure 2 As shown, the specific process of obtaining the first answer corresponding to the first question is as follows:

[0044] Step 201, inputting a first question into a large language model to obtain a first initial answer output by the large language model;

[0045] Step 202, inputting the first question and the first initial answer into the comment model to obtain a first comment output by the comment model;

[0046] Step 203, determining whether the first comment indicates that there is no wrong answer step in the first initial answer, if so, executing step 204, if not, executing step 205;

[0047] Step 204, using the first initial reply as a first answer, forming a first question-answer pair including a first question and the first initial reply;

[0048] Step 205, revising the first initial reply based on the first comment to obtain a first revised reply;

[0049] Step 206, inputting the first revised answer and the first question into the comment model, and obtaining the first comment output by the comment model again;

[0050] Step 207, determining whether the first comment obtained again indicates that there is no wrong answer step in the first revised answer, if so, executing step 208, if not, executing step 209;

[0051] Step 208, using the first revised answer as a first answer, forming a first question-answer pair including the first question and the first revised answer;

[0052] Step 209 , revise the first revised answer again based on the first comment obtained again, obtain the revised first revised answer, and execute step 206 .

[0053] In this embodiment, if there are multiple first questions, each first question goes through the above process to obtain the first answer corresponding to each first question, thereby forming multiple pairs of high-quality first question-answer pairs. The more first question-answer pairs there are, the higher the quality, the more conducive it is to training a higher-quality answer model, and the higher the accuracy of the target answer obtained by using the answer model to answer the target question. In addition, the process of obtaining many high-quality first question-answer pairs does not require manual participation in evaluation and correction processes, and is completed by automated programs, which improves the processing efficiency of the process of obtaining the first question-answer pairs.

[0054] In one embodiment, the comment model is obtained by training the initial comment model based on the second question, the second answer and the second comment. After the initial answer is modified based on the comment to obtain the target answer corresponding to the target question, the method further includes: updating the target question to the second question, updating the target answer to the second answer, and updating the comment to the second comment; based on the updated second question, the updated second answer and the updated second comment, the comment model is trained again to obtain an optimized comment model.

[0055] In this embodiment, the comment model is obtained by training the initial comment model based on the second question, the second answer and the second comment. The specific process of obtaining the comment model includes:

[0056] Given the second question q 2 and large language models The output includes the second answer step Second answer Comment Model LM critique A second comment can be generated Second comment Used to point out First, obtain the wrong answer based on the large language model. Output answering steps:

[0057]

[0058] in, Indicates the second question in the i-th Download large language model Output the i-th second answer, i is a positive integer, N is the total number of second questions, Indicates the second question in the i-th Download large language model The second step of answering the question is output. is intercepted by the function ρ A series of answering steps starting from the first answering step.

[0059] Then, each second question is obtained by manual writing. The corresponding second comment Thus, we get the review model LM critique Training data

[0060] Based on training data Initial review model Train until the initial review model When convergence or reaching the training number threshold, the final review model LM is obtained. critique Optionally, the initial review model It can be any deep learning model determined according to needs, for example, a large language model, a neural network model specifically built according to needs, etc. Optionally, the set of second questions can be completely identical, partially identical, or completely different from the set of first questions mentioned above. In other words, the second question used to train the comment model can also be used as the first question for training the answer model; the first question used in the answer model can also be used as the second question for training the comment model, thereby making full use of the question data and improving the utilization rate of the data.

[0061] In this embodiment, the review model LM is obtained critique After that, the answering model can be based on the comment model LM critique The first comment outputted modifies the first initial answer outputted by the answering model and improves the problem-solving ability.

[0062] In this embodiment, after the trained comment model and the trained answer model are put into use, a higher quality target answer will be obtained based on the target question. Therefore, in the process of using the comment model and the answer model to answer questions, the target question, the target answer and the comments for the target answer can be continuously collected, the target question is updated to the second question, and the target answer is updated to the second answer, and the comment is updated to the second comment, so as to further expand the training data of the comment model and improve the quantity and quality of the training data corresponding to the comment model. After a certain length of time or a certain amount of collection, the comment model is trained again based on the updated second question, the updated second answer and the updated second comment, so as to further improve the performance of the comment model, make the comments output by the optimized comment model more accurate, and thus further improve the accuracy of the target answer.

[0063] In one embodiment, after obtaining the optimized comment model, it also includes: based on the optimized comment model, re-obtaining the first optimized question and answer pair, wherein the first optimized answer in the first optimized question and answer pair is obtained after correcting the first optimized comment output by the optimized comment model; based on the first optimized question and answer pair, training the answer model again to obtain the optimized answer model.

[0064] In this implementation, after the comment model is optimized using the updated second question, the updated second answer, and the updated second comment, the first answer in the original first question-answer pair can also be optimized based on the optimized comment model to obtain a first optimized question-answer pair, thereby further improving the data quality of the training sample corresponding to the answer model. Optionally, the first question in the first optimized question-answer pair includes the first question in the first question-answer pair, and can also include the target question collected before the answer model is optimized, thereby further expanding the first optimized question-answer pair, improving the training effect when the answer model is trained again, and making the optimized answer model have better performance.

[0065] In this embodiment, iterative optimization of the comment model and the answer model is achieved. Figure 3 As shown in FIG. 1 , the process of iterative optimization of the comment model and the question answering model includes:

[0066] Step 301, based on the second question, the second answer and the second comment, training a comment model or optimizing the comment model again;

[0067] Step 302: The large language model outputs a first initial answer corresponding to the first question;

[0068] Step 303, the comment model evaluates the first initial reply or the first revised reply and outputs a first comment;

[0069] Step 304, determining whether the first comment indicates that there are wrong answer steps in the first initial answer or the first revised answer, if so, executing step 305, if not, executing step 306;

[0070] Step 305, modifying the first initial reply based on the first comment to obtain a first modified reply, and executing step 303;

[0071] Step 306, taking the first initial reply or the first revised reply as the first answer, and obtaining a high-quality first question-answer pair;

[0072] If the optimized review model is used during the implementation of steps 301 to 306, a first optimized question-answer pair with higher quality is obtained;

[0073] Step 307, training a large language model based on the first question-answer pair to obtain a question-answering model;

[0074] If the first optimized question-answer pair is used to retrain the answer model, an optimized answer model is obtained;

[0075] Step 308, collecting target questions, target answers and comments during the use of the answering model;

[0076] Step 309 , after reaching the collection time or the collection quantity, the target question is updated to the second question, the target answer is updated to the second answer, and the comment is updated to the second comment, and step 301 is executed.

[0077] In this embodiment, the performance of the comment model and the question answering model is continuously improved through iterative optimization of the comment model and the question answering model, thereby continuously improving the accuracy of the target answer.

[0078] In one embodiment, the training process of the comment model includes: correcting the second answer based on the second comment to obtain a second corrected answer; obtaining a first correctness corresponding to the second answer, and obtaining a second correctness corresponding to the second corrected answer; determining a reward in the initial comment model training process based on the difference between the first correctness and the second correctness; with the goal of maximizing the reward, training the initial comment model based on the second question, the second answer and the second comment to obtain a comment model.

[0079] In this embodiment, in order to encourage the review model to generate better reviews, a reinforcement learning method based on proximal policy optimization (PPO) is used to optimize the review model. The specific process includes:

[0080] First, we need to clarify what a good review is. Given the second question q 2 、Second step of answering questions We can infer the second question q2 The corresponding complete second answer

[0081] Then, the second answer can be calculated The first accuracy, the process of calculating the accuracy, can be achieved through deep learning models or vector calculations.

[0082] Afterwards, according to the second comment You will get a second revised answer that includes the revised answer steps And calculate the second revised answer The second accuracy.

[0083] The difference between the first correctness and the second correctness is used as the advantage function of the review model. The difference between the first correctness and the second correctness is positively correlated with the reward of the review model. With the goal of maximizing the reward of the review model, the initial review model is trained to obtain the review model.

[0084] In one embodiment, obtaining a first accuracy corresponding to the second answer and obtaining a second accuracy corresponding to the second revised answer include: inputting the second question and the second answer into a correction model to obtain the first accuracy output by the correction model; inputting the second question and the second revised answer into the correction model to obtain the second accuracy output by the correction model; wherein the correction model is obtained by training an initial correction model based on a third question, a standard answer, a third answer and a third accuracy, and the third answer includes an answer obtained based on a large language model or a question-answering model.

[0085] In this embodiment, the correction model is used to obtain the correctness of the answer. The process of obtaining the correction model is as follows:

[0086] Given the third question q 3 , standard answer y, third answer Correction Model LM grade The prediction output is a value between 0 and 1 Indicates the third answer The correctness relative to the standard answer y. Therefore, during the training phase, a batch of questions containing the third question q is obtained. 3 And the data set corresponding to the standard answer y in, represents the third question of the ith i represents the i-th standard answer, i is a positive integer, and M is the total number of the third question.

[0087] Optionally, sample a large language model Based on the third topic Output of the third answer And manually review its correctness i , v i Indicates the third answer The accuracy of the training data of the correction model is obtained

[0088] Training data based on the revised model Initial correction model Train until the initial correction model When convergence or reaching the training number threshold, the final correction model LM is obtained. grade .

[0089] Optionally, an initial correction model It can be any deep learning model determined according to needs, for example, a large language model, a neural network model specifically built according to needs, etc. Optionally, the set of the third questions can be completely identical, partially identical, or completely different from the set of the first or second questions mentioned above. In other words, the second question used for training the grading model can also be used as the first or second question for training the answering model or the commenting model; the first question used for the answering model or the second question used for the commenting model can also be used as the third question for training the grading model, thereby achieving full utilization of the question data and improving data utilization.

[0090] In this embodiment, the correction model LM grade Correctness can be calculated, providing a clear feedback signal for reinforcement learning.

[0091] In this embodiment, based on the trained correction model LM grade The process of calculating the correctness and implementing the reinforcement learning of the review model includes:

[0092] Optionally, use a large language model Obtain a second answer corresponding to each second question.

[0093] First, given a large language model Enter the second question q 2 , Large Language Model LM solving Output of the second answer step Then we can infer the complete second answer Then use the correction model LM grade , get the second answer The first accuracy

[0094] Then, based on the review model LM critique Output the second comment Large Language Model For the second answer Correct the existing wrong answering steps to obtain the corrected second answering steps Second revised answer

[0095] Next, we use the same correction model LM grade Get the second corrected answer The second accuracy

[0096] Calculate the difference between the first accuracy and the second accuracy as the second comment The advantage function A is as follows:

[0097]

[0098] Based on the advantage function, the PPO method is used to make the review model LM critique Towards Reward The reward of the review model is as follows:

[0099]

[0100] Among them, E t represents the expectation of the current step, π actor (a t |s t ) represents the current step review model LM critique The likelihood probability, π ref (a t |s t ) represents the previous step review model LM critique The likelihood probability, a t Indicates the currently predicted character, s t Indicates the predicted character a t ,λ is a training parameter used to adjust the proportion of supervised fine-tuning (SFT) loss. represents the expectation of the previous training cycle, and x represents the training data of the previous training process.

[0101] In an overall embodiment, obtaining the answering model and iterative training of the answering model include four stages, including:

[0102] like Figure 4 In the first stage shown, the correction model LM is trained grade Given the third question q 3 , standard answer y, large language model Predict the third question q 3 The corresponding third answer At this stage, the large language model The internal parameters are frozen. Based on the third question q 3 , standard answer y and the third answer Initial correction model Training, during the training process, the initial correction model The internal parameters are fine-tuned, and the training is completed to obtain the corrected model LM grade . Correction Model LM grade Used to calculate the correctness of the answer.

[0103] like Figure 5 In the second stage shown, the review model LM is trained critique Given the second question q 2 , Large Language Model Output the second question q 2 The corresponding second answer step includes Second answer At this stage, the large language model The internal parameters are frozen. 2 and including the second answer step Second answer Initial review model During the training process, the initial review model The internal parameters are fine-tuned, and the training is completed to obtain the review model LM critique . Comments Model LM critique It is used to evaluate the initial answer output by the answering model and generate comments c to point out the incorrect answering steps in the initial answer.

[0104] like Figure 6 In the third stage shown in the figure, a reinforcement learning-based optimization method is used during the review stage training. Based on the second question q 2 The output includes the second answer step Second answer And based on the second comment The corrected answer includes the second corrected answer step Second revised answer Based on the correction model LM grade Calculate the first accuracy and the second accuracy to get the advantage function A, which is used as part of the reinforcement learning reward to improve the initial review model At this stage, the large language model The internal parameters are frozen, and the initial review model The internal parameters can be fine-tuned.

[0105] like Figure 7 The fourth stage shown, for the first question q 1 , large language model Output first initial reply Comment Model LM critique To this first initial response Write a review and get the first review Using a large language model Achieve the first comment First initial reply Correction to get the first answer Then we can get a high-quality first question-answer pair. Based on multiple high-quality first question-answer pairs For large language models Train and get the answer model LM solving Optionally, the answering model LM solving During the use phase, you can collect target questions, target answers and comments, and Update, and for the second question q 2 , Second answer And the second comment Update and re-evaluate the review model LM critique And the answering model LM solving Conduct training to implement the review model LM critique And the answering model LM solving Optimization iterations.

[0106] In the present application, the target question is input into the answering model to obtain the initial answer output by the answering model, wherein the initial answer includes at least one answering step; the target question and the initial answer are input into the comment model to obtain the comment output by the comment model, wherein the comment includes the evaluation information of each answering step in the initial answer; the initial answer is corrected based on the comment to obtain the target answer corresponding to the target question, wherein the answering model is obtained after the large language model is trained based on at least one first question-answer pair, and the first answer in the first question-answer pair includes the answer obtained after correction based on the first comment output by the comment model. In the above process, the first answer in the first question-answer pair used in the training of the answering model includes the answer obtained after correction by the first comment output by the comment model, which improves the correctness of the first answer, thereby improving the performance of the answering model obtained based on the first question-answer pair training, and improving the correctness of the initial answer at the level of the answering model. Furthermore, since the answering model is obtained based on the training of the large language model, the initial answer initially output by the answering model has a certain error rate. This application corrects the results output by the answering model at the level of answering steps based on the comments output by the comment model, and further improves the accuracy of the target answer based on the initial answer.

[0107] Specifically, this application provides an implementation scheme for answering questions using a question answering model based on feedback from a comment model. The scheme is a framework for multi-agent collaboration, including a question answering model LM solving , review model LM critique Specifically, the answering model LM solving Learn from the topic q 1 To the standard answer y. Considering that the answer to the target question has the characteristics of multiple steps, for the answer model LM solving The answering steps in the output initial answer, the comment model LM critique It will give comments to point out the wrong answering steps in the initial answer. Using the step-level comment model, the answering process can be corrected in more detail. Finally, relying on its strong instruction-following ability, according to the comments, the answering model LM solving The error in the initial reply will be corrected to obtain a more accurate target reply. Through reinforcement learning and using the review model to obtain reward signals, the error or bias of human feedback is avoided, which improves the effect of the review model generating comments. critique With the answering model LM solving Iterative optimization can be achieved, answering model LM solving Constantly based on the review model LM critique The output is adjusted continuously based on the feedback, and a more accurate target answer is finally obtained. Multiple iterative optimization strategies are adopted to improve the performance of language model and comment model simultaneously.

[0108] More specifically, during the review model training process, a batch of large language models are sampled Output answer steps Then we get the review model LM critique Training data Initial review model Train the review model LM critique Ability to accurately evaluate the answering model LM at the level of answering steps solving The initial answer is output. At the same time, the review model is optimized based on PPO, using a trained correction model LM grade , computing large language models Initial generated second answer And according to the second comment Modified Second Corrected Answer The accuracy of the first accuracy and the difference between the second accuracy are used as the evaluation model LM critique Generate second comment The advantage function is then used to optimize the review model through PPO to generate higher quality reviews. The reinforcement learning method proposed in this scheme does not rely on human feedback, but adopts a more objective and accurate correction model with strong robustness. critique For large language models The output of the first initial reply First review Build new high-quality first question-answer pairs Complete the large language model training.

[0109] Exemplary Devices

[0110] Accordingly, the present application embodiment also provides a question answering device, such as Figure 8 As shown, the question answering device comprises:

[0111] The initial answering module 801 is used to input the target question into the answering model and obtain the initial answer output by the answering model, wherein the initial answer includes at least one answering step;

[0112] The comment output module 802 is used to input the target question and the initial answer into the comment model to obtain the comment output by the comment model, wherein the comment includes evaluation information of each answering step in the initial answer;

[0113] The answer correction module 803 is used to correct the initial answer based on the comments to obtain the target answer corresponding to the target question, wherein the answer model is obtained by training the large language model based on at least one pair of first question and answer pairs, and the first answer in each pair of first question and answer pairs includes an answer obtained by correcting the first comment output by the comment model.

[0114] In one embodiment, the question answering device also includes a model training module for obtaining a first answer, and the process of obtaining the first answer includes: inputting the first question into a large language model to obtain a first initial reply output by the large language model; inputting the first question and the first initial reply into a comment model to obtain a first comment output by the comment model; judging whether the first comment indicates that there are no incorrect answering steps in the first initial reply, and if not, revising the first initial reply based on the first comment to obtain a first revised reply, re-inputting the first revised reply and the first question into the comment model, and obtaining the first comment output by the comment model again, until the first comment indicates that there are no incorrect answering steps in the first revised reply, and taking the first revised reply as the first answer; if so, taking the first initial reply as the first answer.

[0115] In one embodiment, the comment model is obtained by training the initial comment model based on the second question, the second answer and the second comment;

[0116] The model training module is also used to modify the initial reply based on the comments, and after obtaining the target reply corresponding to the target question, update the target question to the second question, update the target reply to the second answer, and update the comments to the second comments; based on the updated second question, the updated second answer and the updated second comments, the comment model is trained again to obtain an optimized comment model.

[0117] In one embodiment, the model training module is also used to re-obtain a first optimized question-answer pair based on the optimized comment model after obtaining the optimized comment model, wherein the first optimized answer in the first optimized question-answer pair is obtained after correcting the first optimized comment output by the optimized comment model; based on the first optimized question-answer pair, the question-answering model is trained again to obtain an optimized question-answering model.

[0118] In one embodiment, the model training module is also used for training the comment model, and the training process of the comment model includes: correcting the second answer based on the second comment to obtain a second corrected answer; obtaining a first correctness corresponding to the second answer, and obtaining a second correctness corresponding to the second corrected answer; determining a reward in the initial comment model training process based on the difference between the first correctness and the second correctness; with the goal of maximizing the reward, training the initial comment model based on the second question, the second answer and the second comment to obtain the comment model.

[0119] In one embodiment, the model training module is used to input the second question and the second answer into the correction model to obtain the first accuracy output by the correction model; input the second question and the second revised answer into the correction model to obtain the second accuracy output by the correction model; wherein the correction model is obtained by training the initial correction model based on the third question, the standard answer, the third answer and the third accuracy, and the third answer includes an answer obtained based on the large language model or the question answering model.

[0120] In one embodiment, the answer correction module 803 is used to input comments into the question answering model so that the question answering model corrects the initial answer based on the comments to obtain the target answer output by the question answering model.

[0121] The question-solving device provided in this embodiment belongs to the same application concept as the question-solving method provided in the above-mentioned embodiment of this application, and can execute the question-solving method provided in any of the above-mentioned embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in this embodiment, please refer to the specific processing content of the question-solving method provided in the above-mentioned embodiment of this application, and will not be repeated here.

[0122] Exemplary Electronic Devices

[0123] The present application also provides an electronic device, such as Fig. 9 As shown, the electronic device includes: a memory 900 and a processor 901 .

[0124] The memory 900 is connected to the processor 901 and is used to store programs.

[0125] The processor 901 is used to implement the problem-solving method in the above embodiment by running the program stored in the memory 900 .

[0126] Specifically, the electronic device may further include: a communication interface 902 , an input device 903 , an output device 904 and a bus 905 .

[0127] The processor 901, the memory 900, the communication interface 902, the input device 903 and the output device 904 are connected to each other via a bus.

[0128] Bus 905 may include a pathway for transferring information between the various components of the computer system.

[0129] The processor 901 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0130] The processor 901 may include a main processor, and may also include a baseband chip, a modem, and the like.

[0131] The memory 900 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 900 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.

[0132] The input device 903 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.

[0133] Output device 904 may include a device that allows information to be output to a user, such as a display screen, printer, speaker, etc.

[0134] The communication interface 902 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0135] The processor 901 executes the program stored in the memory 900 and calls other devices, which can be used to implement the various steps of the question answering method provided in the above embodiment of the present application.

[0136] Exemplary computer program products and storage media

[0137] In addition to the above-mentioned methods and devices, the embodiments of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the problem-solving method described in the embodiments of the present application.

[0138] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0139] In addition, the embodiments of the present application may also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to execute the steps of the problem-solving method described in the embodiments of the present application.

[0140] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0141] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0142] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0143] The modules and sub-modules in the devices and terminals provided in the embodiments of the present application can be merged, divided and deleted according to actual needs.

[0144] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0145] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.

[0147] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0148] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0149] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0150] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for solving a problem, characterized in that: include: Inputting the target question into the question answering model to obtain an initial answer output by the question answering model, wherein the initial answer includes at least one question answering step; Inputting the target question and the initial answer into a comment model to obtain a comment output by the comment model, wherein the comment includes evaluation information of each of the answering steps in the initial answer; The initial reply is modified based on the comment to obtain a target reply corresponding to the target question, wherein the question answering model is obtained after training a large language model based on at least one first question-answer pair, and the first answer in the first question-answer pair includes an answer obtained by modifying the first comment output by the comment model.

2. The method for answering a question according to claim 1, characterized in that: The process of obtaining the first answer includes: Inputting a first question into the large language model to obtain a first initial answer output by the large language model; Inputting the first question and the first initial answer into the comment model to obtain the first comment output by the comment model; Determine whether the first comment indicates that there are no incorrect answering steps in the first initial reply. If not, revise the first initial reply based on the first comment to obtain a first revised reply, re-input the first revised reply and the first question into the comment model, and obtain the first comment output by the comment model again, until the first comment indicates that there are no incorrect answering steps in the first revised reply, use the first revised reply as the first answer; if so, use the first initial reply as the first answer.

3. The method for answering a question according to claim 1, wherein: The comment model is obtained by training the initial comment model based on the second question, the second answer and the second comment; After the initial answer is modified based on the comment to obtain the target answer corresponding to the target question, the method further includes: updating the target title to the second title, updating the target answer to the second answer, and updating the comment to the second comment; Based on the updated second question, the updated second answer and the updated second comment, the comment model is trained again to obtain an optimized comment model.

4. The method for answering a question according to claim 3, characterized in that: After obtaining the optimized review model, the method further includes: Based on the optimized comment model, reacquire a first optimized question-answer pair, wherein a first optimized answer in the first optimized question-answer pair is obtained by correcting a first optimized comment output by the optimized comment model; Based on the first optimized question-answer pair, the question-answering model is trained again to obtain an optimized question-answering model.

5. The method for answering a question according to claim 3, characterized in that: The training process of the review model includes: Modify the second answer based on the second comment to obtain a second modified answer; Obtaining a first degree of correctness corresponding to the second answer, and obtaining a second degree of correctness corresponding to the second revised answer; determining a reward in the initial review model training process based on a difference between the first accuracy and the second accuracy; With the goal of maximizing the reward, an initial comment model is trained based on the second question, the second answer and the second comment to obtain the comment model.

6. The method for answering a question according to claim 5, characterized in that: The obtaining of the first correctness corresponding to the second answer and the obtaining of the second correctness corresponding to the second revised answer include: Inputting the second question and the second answer into a grading model to obtain the first accuracy output by the grading model; Inputting the second question and the second revised answer into the grading model to obtain the second accuracy output by the grading model; The correction model is obtained by training the initial correction model based on the third question, the standard answer, the third answer and the third accuracy, and the third answer includes an answer based on the large language model or the question-answering model.

7. The method for answering a question according to claim 1, wherein: The step of modifying the initial answer based on the comment to obtain a target answer corresponding to the target question includes: The comments are input into the question answering model so that the question answering model modifies the initial answer based on the comments to obtain the target answer output by the question answering model.

8. A question-solving device, characterized in that: include: An initial answering module, used to input a target question into the answering model and obtain an initial answer output by the answering model, wherein the initial answer includes at least one answering step; A comment output module, used for inputting the target question and the initial answer into a comment model to obtain a comment output by the comment model, wherein the comment includes evaluation information of each answering step in the initial answer; An answer correction module is used to correct the initial reply based on the comment to obtain a target answer corresponding to the target question, wherein the question answering model is obtained after training a large language model based on at least one pair of first question and answer pairs, and the first answer in each pair of the first question and answer pairs includes an answer obtained by correcting the first comment output by the comment model.

9. An electronic device, characterized in that: include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the problem-solving method according to any one of claims 1 to 7 by running the program in the memory.

10. A computer program product, characterized in that includes computer program instructions; When the computer program instructions are executed by a processor, the processor is caused to execute the problem-solving method according to any one of claims 1 to 7.