Question answering task processing model training method, device, equipment and storage medium

By automatically generating labeled data, the problem of relying on manual labeled data for training in the existing technology is solved, which improves the accuracy and training efficiency of the model and reduces costs.

CN119493849BActive Publication Date: 2025-05-13INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510066724.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art relies heavily on manual annotation data when handling model training and optimization of Q&A tasks, resulting in the accuracy of the model being affected by the quality of manual annotation data, and the training process is costly and inefficient.

Method used

By using the initial training data set and the intermediate question-and-answer task processing model, a target training data set containing the steps and labeling information of the model's processing question-and-answer task is generated, and the labeling data is automatically generated, reducing dependence on manual labeling, and improving training efficiency and data scalability.

Benefits of technology

It improves the accuracy and robustness of the Q&A task processing model, reduces training costs and time, reduces the dependence of manually labeled data, and enhances the diversity and coverage of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119493849B_ABST
    Figure CN119493849B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of natural language processing technology, and discloses a question-answering task processing model training method, device, equipment and storage medium. The method comprises: adjusting the question-answering task processing model according to an initial training data set and a first model training algorithm to obtain a first intermediate question-answering task processing model; obtaining a first reasoning step and first annotation information for processing the question-answering task according to the initial training data set and the first intermediate question-answering task processing model to generate a first target training data set; adjusting the question-answering task processing model according to the initial training data set and the second model training algorithm to obtain a second intermediate question-answering task processing model; adjusting the second intermediate question-answering task processing model according to the first target training data set to obtain a target model. The method solves the problem that when training a model for processing question-answering tasks, the model's accuracy is affected by the quality of the manually annotated data, and the training process is costly and inefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a question-answering task processing model training method, device, equipment and storage medium. Background Art

[0002] Currently, various models for natural language processing (NLP) are used to handle question-answering tasks, such as large language models (LLMs). There are many methods to improve the reasoning ability of large language models, including pre-training, fine-tuning, prompt optimization, and reward models, which can improve the performance of large language models in complex question-answering tasks. However, the above methods are highly dependent on manually labeled data for the training and optimization process of large language models. Since manual labeling is not only costly but also inefficient, the cost of training and optimizing large language models is significantly increased, and the task processing speed and flexibility of large language models after training are limited. In addition, manually labeled data may have labeling bias and inaccuracy, which will affect the accuracy of large language models in complex question-answering tasks.

[0003] Therefore, the related technology has the problem of being highly dependent on manually labeled data when training and optimizing models for processing question-answering tasks, which causes the accuracy of the model to be affected by the quality of the manually labeled data, and the training process is costly and inefficient. Summary of the invention

[0004] In view of this, the present invention provides a method, device, equipment and storage medium for training a question-answering task processing model to solve the problem that when training and optimizing the model for processing question-answering tasks, the model's accuracy is highly dependent on manually labeled data, resulting in the model's accuracy being affected by the quality of the manually labeled data, and the training process is costly and inefficient.

[0005] In a first aspect, the present invention provides a method for training a question-answering task processing model, the method comprising:

[0006] According to the initial training data set and the first model training algorithm, the question-answering task processing model to be trained is adjusted to obtain a first intermediate question-answering task processing model, wherein the initial training data set includes a first preset number of question-answering tasks;

[0007] According to the initial training data set and the first intermediate question-answering task processing model, a first reasoning step of the question-answering task processed by the first intermediate question-answering task processing model and first annotation information of the first reasoning step are obtained, and according to the first reasoning step, the first annotation information and the initial training data set, a first target training data set is generated, wherein the first reasoning step is any step in the reasoning process of the question-answering task processed by the first intermediate question-answering task processing model;

[0008] According to the initial training data set and the second model training algorithm, the question-answering task processing model to be trained is adjusted to obtain a second intermediate question-answering task processing model;

[0009] The second intermediate question-answering task processing model is adjusted according to the first target training data set to obtain a target model.

[0010] The question-answering task processing model training method provided in this embodiment uses the initial training data set and the first intermediate question-answering task processing model to generate a first target training data set containing the steps and annotation information of the model processing the question-answering task, and automatically generates the annotation data without the need for manual annotation, thereby improving the efficiency of the training process and enhancing the scalability of the data. In addition, the first target training data set uses the step as the minimum granularity, and uses the first target training data set to adjust the second intermediate question-answering task processing model to obtain the target model, thereby improving the accuracy and robustness of the target model. This solves the problem that when training and optimizing the model for processing question-answering tasks, the model's accuracy is highly dependent on manually annotated data, resulting in the model's accuracy being affected by the quality of the manually annotated data, and the training process is costly and inefficient.

[0011] In a second aspect, the present invention provides a question-answering task processing model training device, the device comprising:

[0012] A first adjustment module is used to adjust the question-answering task processing model to be trained according to the initial training data set and the first model training algorithm to obtain a first intermediate question-answering task processing model, wherein the initial training data set includes a first preset number of question-answering tasks;

[0013] A data set generation module is used to obtain, based on the initial training data set and the first intermediate question-answering task processing model, a first reasoning step of the first intermediate question-answering task processing model for processing a question-answering task and first annotation information of the first reasoning step, and generate a first target training data set based on the first reasoning step, the first annotation information and the initial training data set, wherein the first reasoning step is any step in the reasoning process of the first intermediate question-answering task processing model for processing a question-answering task;

[0014] A second adjustment module is used to adjust the question-answering task processing model to be trained according to the initial training data set and the second model training algorithm to obtain a second intermediate question-answering task processing model;

[0015] The third adjustment module is used to adjust the second intermediate question-answering task processing model according to the first target training data set to obtain a target model.

[0016] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the question and answer task processing model training method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the question-and-answer task processing model training method of the above-mentioned first aspect or any corresponding embodiment thereof.

[0018] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, which are used to enable a computer to execute the question-answering task processing model training method of the above-mentioned first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 is a flowchart of a question-answering task processing model training method according to an embodiment of the present invention;

[0021] Figure 2 is a schematic diagram of a model iterative fine-tuning process according to an embodiment of the present invention;

[0022] Figure 3 is a schematic diagram of a preference data set construction process according to an embodiment of the present invention;

[0023] Figure 4 is a flow chart of an automatic process annotation scoring strategy according to an embodiment of the present invention;

[0024] Figure 5 is a structural block diagram of a question-answering task processing model training device according to an embodiment of the present invention;

[0025] Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0027] Question answering task processing models are various models used for natural language processing, such as large language models (LLMs). Natural language processing includes two main aspects: natural language understanding and natural language generation. It is an intersection of computer science, artificial intelligence and linguistics. Its goal is to enable computers to understand, interpret and generate human natural language. Natural language refers to the language people use in daily life.

[0028] Currently, three types of methods are used to improve the reasoning ability of large language models (LLMs), including pre-training methods, fine-tuning methods, and prompting methods. Pre-training methods mainly improve the reasoning ability of LLMs through self-supervised learning on large-scale datasets. For example, LLMs are pre-trained on large-scale datasets related to reasoning and mathematics. Such datasets not only contain text, but also include math problems, logical reasoning, problem solving, etc., to help LLMs gain a better foundation in reasoning ability. In this way, LLMs can reason more effectively when facing actual reasoning tasks with the help of semantics and reasoning rules learned during the training process. Fine-tuning methods are to further train LLMs on the basis of pre-trained LLMs using task-specific datasets with high-quality annotations, fine-tune LLMs in specific fields, improve the reasoning ability of LLMs, and enable the model to perform better on specific tasks. The fine-tuning process usually involves building and using high-quality question-answer datasets, which contain complex reasoning questions and corresponding answers. By training on these datasets, LLMs can master specific problem-solving strategies and logical reasoning methods in reasoning tasks. The prompt method guides the performance of LLMs in reasoning tasks by designing and optimizing input prompts. The prompt method does not involve updating the parameters of LLMs, but stimulates the reasoning ability of LLMs by optimizing the input. The advantage is that it does not require large-scale training data and has low implementation costs. By designing different types of prompt strategies, such as guiding the model to focus on specific information, using step-by-step reasoning, and providing demonstrations, LLMs can be prompted to generate more accurate reasoning results, but its effect is highly dependent on the quality of the prompt words.

[0029] In addition, iterative optimization methods can also improve the reasoning ability of LLMs. In reasoning tasks, LLMs start from the current reasoning strategy, collect and analyze data, continuously generate new preference data, and adjust and optimize strategies on this basis, so that LLMs can continuously adapt and improve over a long period of time. It can not only improve the accuracy of reasoning tasks, but also reveal how LLMs can better fit the human decision-making process and reasoning complexity. LLMs gradually integrate human values ​​and preferences more closely. One of the core is to use preference data and introduce additional validators to select the best answer from multiple decoding candidate answers. In addition, LLMs are adjusted using reward models, which include: Outcome Reward Model (ORM) and Process Reward Model (PRM). ORM assigns a comprehensive score to the entire solution process to evaluate the quality of the final solution. The focus of ORM is to evaluate whether the final answer is correct, but it lacks a detailed analysis of the reasoning process. PRM independently scores each step in the reasoning process, which can accurately identify errors in the reasoning process and point out the specific location where the problem occurs.

[0030] Although the above methods can effectively improve the performance of LLMs in complex reasoning tasks, they are highly dependent on manually labeled data in the process of training and optimizing LLMs, which not only significantly increases costs, but also limits processing speed and flexibility, resulting in low training and optimization efficiency. In addition, manually labeled data may have bias and inaccuracy, which further affects the effect of model training.

[0031] Based on the above content, an embodiment of the present invention provides a method for training a question-answering task processing model, automatically constructs a process supervision data set, selects positive and negative samples through an automated mechanism, dynamically evaluates the rationality and accuracy of each step in the reasoning task, and automatically generates a high-quality process supervision data set for question-answering task processing model reasoning. The training process scoring model is used to evaluate the performance of the question-answering task processing model at each step in the problem-solving process. And based on the automated process supervision data set, the question-answering task processing model is fine-tuned by RFT (Reinforcement Fine-Tuning) to improve its performance in problem-solving tasks. The RFT model is iteratively generated and optimized to achieve continuous optimization of model performance, improve the generalization ability and adaptability of the model, and enable it to adapt to changing task requirements. The generalization ability and adaptability of the model are improved, so that the model can perform in various complex tasks. In order to achieve the effect of providing efficient and accurate process scoring without relying on manually labeled data sets, thereby helping to optimize the reasoning process of the question-answering task processing model.

[0032] According to an embodiment of the present invention, a question-answering task processing model training embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer device with data processing capabilities, such as a computer, a server, etc., and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0033] In this embodiment, a question-answering task processing model training method is provided, which can be used in the above-mentioned computer device. Figure 1 is a flowchart of a question-answering task processing model training method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0034] Step S101, according to the initial training data set and the first model training algorithm, the question-answering task processing model to be trained is adjusted to obtain a first intermediate question-answering task processing model, wherein the initial training data set includes a first preset number of question-answering tasks.

[0035] Specifically, this embodiment takes the question-answering task processing model as a large language model (LLMs) as an example to illustrate. First, a benchmark data set is constructed, and the reasoning type that needs to be focused on by the large language model is determined, such as: arithmetic reasoning, Python code reasoning, etc. A representative data set under the reasoning type is obtained as a benchmark data set. For example: in terms of arithmetic reasoning type, two representative data sets: the GSM8K data set and the MATH data set are used as benchmark data sets; the GSM8K data set contains elementary school students' mathematics application problems; the MATH data set covers challenging competition mathematics problems. In terms of Python code reasoning type, GPT is used to generate a set of data sets as a benchmark data set. The data set contains code that can pass five unit test cases, and on this basis, the process preference data of code reasoning is generated through the model.

[0036] Taking the benchmark data set as the MATH data set as an example, the MATH data set contains a first preset number of question-and-answer tasks and reference results corresponding to the question-and-answer tasks. The first preset number represents multiple values ​​such as: 500, 1000 or other values ​​that meet actual needs. Based on the above data, RFT data is generated using an automated annotation method, and SFT (Supervised Fine-Tuning) data containing answers corresponding to the inputs is derived based on the above data. SFT data refers to data used for supervised fine-tuning, which usually consists of model inputs and target outputs corresponding to the model inputs. RFT data refers to data used in the reinforcement learning fine-tuning process. RFT is the process of further reinforcement learning fine-tuning on a model trained with SFT fine-tuning. Reinforcement learning fine-tuning usually involves adjusting the model's strategy through environmental feedback so that it can perform better in specific tasks.

[0037] The initial training data set includes the above-mentioned SFT data and RFT data, which can be called as needed.

[0038] The question-answering task processing model to be trained is, for example, LLMs trained using the above-mentioned pre-training method. The first model training algorithm is, for example, an algorithm for training LLMs using the SFT method. The SFT data in the initial training data set is obtained, and the SFT data contains question-answering tasks and reference answers corresponding to the question-answering tasks, i.e., reference results. The question-answering tasks in the SFT are input into the question-answering task processing model to be trained, and the question-answering task processing model to be trained is adjusted using the first model training algorithm according to the process, results, and reference results of processing the question-answering tasks by the first model training algorithm, so as to obtain the initial SFT model as the first intermediate question-answering task processing model. The question-answering task processing model to be trained is adjusted using the SFT method, so that the model can perform accurate knowledge learning based on high-quality labeled data, ensuring that its performance in basic tasks is not biased.

[0039] The above process is as follows Figure 2 As shown, the pre-trained model is adjusted using the SFT data to obtain the SFT fine-tuning model.

[0040] Step S102, based on the initial training data set and the first intermediate question and answer task processing model, obtain the first reasoning step of the first intermediate question and answer task processing model for processing the question and answer task and the first annotation information of the first reasoning step, and generate the first target training data set based on the first reasoning step, the first annotation information and the initial training data set, wherein the first reasoning step is any step in the reasoning process of the first intermediate question and answer task processing model for processing the question and answer task.

[0041] Specifically, the first intermediate question-answering task processing model is used as a generator. The question-answering task in the SFT data contained in the initial training data set is input into the generator, and the generator outputs the reasoning process of processing the question-answering task. Each reasoning process is refined and split according to the steps to obtain the first reasoning step in the reasoning process. Each split first reasoning step is used as a basic unit of data processing for subsequent preference data construction and analysis. The automated process labeling method proposed in this embodiment is used to perform a quality assessment on the first reasoning step, and the first labeling information of the first reasoning step is determined according to the assessment result. For example, the first reasoning step includes and , The first annotation information is a positive sample. The first labeled information is a negative sample. The positive sample indicates that the probability of obtaining the standard answer to the question-answering task through this step is high, and the negative sample indicates that the probability of obtaining the standard answer to the question-answering task through this step is low.

[0042] The Step-DPO dataset contains the steps, preference information, and standard answers of the first intermediate question-answering task processing model when solving the question-answering task. Therefore, according to the first reasoning step, the first annotation information, and the initial training dataset, an initial Step-DPO (Step-wise Direct Preference Optimization) dataset is generated, and the Step-DPO dataset is used as the first target training dataset.

[0043] like Figure 2 As shown, the SFT fine-tuning model is used to automatically generate process supervision data, and Step-DPO data is generated based on the generated process supervision data. In addition, Figure 2 The iterative reasoning (process scoring score) of the process scoring model is used to determine the first annotation information of the first reasoning step.

[0044] Step S103: According to the initial training data set and the second model training algorithm, the language model to be trained is adjusted to obtain a second intermediate question-answering task processing model.

[0045] Specifically, the initial training data set includes the above-mentioned SFT data and RFT data, which can be called according to needs.

[0046] In order to improve the reasoning ability of the model, supervised fine-tuning (SFT) and reinforcement learning fine-tuning (RFT) are combined and applied to further fine-tune the pre-trained LLM. Therefore, the second model training algorithm is, for example: an algorithm for training LLMs by combining the SFT method and the RFT method.

[0047] Obtain the SFT data and RFT data in the initial training data set, where the SFT data contains the question-answering task and the reference answer corresponding to the question-answering task, i.e., the reference result. Input the question-answering task in the SFT into the question-answering task processing model to be trained, and use the first model training algorithm to adjust the question-answering task processing model to be trained according to the process, results, and reference results of the first model training algorithm processing the question-answering task. Use the RFT method to further perform reinforcement learning fine-tuning on the model trained by SFT fine-tuning, and obtain the RFT fine-tuning model as the second intermediate question-answering task processing model.

[0048] The SFT method is used to adjust the question-answering task processing model to be trained, so that LLMs can learn accurate knowledge based on high-quality labeled data, ensuring that their performance in basic tasks is not biased. The RFT method uses a reward mechanism to strengthen the model's reasoning strategy, so that LLMs can better cope with uncertainties and challenges when facing dynamic and complex tasks, thereby improving the model's generalization ability and robustness.

[0049] like Figure 2 As shown, the pre-trained model is adjusted using SFT data + RFT data to obtain the RFT fine-tuning model.

[0050] Step S104, adjusting the second intermediate question-answering task processing model according to the first target training data set to obtain a target model.

[0051] Specifically, the first target training data set is, for example, a Step-DPO data set. Each data entry of the Step-DPO data set includes a comparison of two first reasoning steps, and which of the two first reasoning steps is better can be determined according to the first annotation information of the first reasoning step.

[0052] The Step-DPO data is used to fine-tune the second intermediate question-answering task processing model to obtain the Step-DPO model, and then new Step-DPO data is continuously generated. Based on this, Step-DPO data is iteratively generated and the RFT model is optimized. In each round of iteration, the previously generated Step-DPO data is used to fine-tune the second intermediate question-answering task processing model, and the RFT strategy is continuously updated and optimized to improve the reasoning performance of the second intermediate question-answering task processing model in complex tasks.

[0053] like Figure 2 As shown, in the first iteration, the RFT fine-tuning model is adjusted using the Step-DPO data to obtain the Step-DPO model.

[0054] The question-answering task processing model training method provided in this embodiment uses the initial training data set and the first intermediate question-answering task processing model to generate a first target training data set containing the steps and annotation information of the model processing the question-answering task, and automatically generates the annotation data without the need for manual annotation, thereby improving the efficiency of the training process and enhancing the scalability of the data. In addition, the first target training data set uses the step as the minimum granularity, and uses the first target training data set to adjust the second intermediate question-answering task processing model to obtain the target model, thereby improving the accuracy and robustness of the target model. This solves the problem that when training and optimizing the model for processing question-answering tasks, the model's accuracy is highly dependent on manually annotated data, resulting in the model's accuracy being affected by the quality of the manually annotated data, and the training process is costly and inefficient.

[0055] In some optional implementations, adjusting the second intermediate question-answering task processing model according to the first target training data set to obtain a target model includes: adjusting the second intermediate question-answering task processing model according to the first target training data set to obtain a third intermediate question-answering task processing model; obtaining the second reasoning step of the question-answering task and the second annotation information of the second reasoning step by the third intermediate question-answering task processing model to process the question-answering task according to the initial training data set and the third intermediate question-answering task processing model, and determining the output result of the question-answering task output by the third intermediate question-answering task processing model; obtaining a reference result of the question-answering task in the initial training data set, and obtaining the third intermediate question-answering task processing model according to the reference result and the output result. The accuracy of the intermediate question and answer task processing model is determined; when the accuracy is greater than the first preset threshold, the third intermediate question and answer task processing model is used as the target model; when the accuracy is less than or equal to the first preset threshold, a second target training data set is generated according to the second reasoning step, the second annotation information and the initial training data set, and the current number of iterations is updated; the second target training data set is used as the first target training data set, and subsequent steps are executed starting from adjusting the second intermediate question and answer task processing model according to the first target training data set until the accuracy is greater than the first preset threshold or the number of iterations is equal to the second preset threshold, then the process ends and the third intermediate question and answer task processing model is used as the target model.

[0056] Specifically, the first target training data set is, for example, the Step-DPO data set. The second intermediate question-answering task processing model is, for example, the RFT fine-tuning model. The RFT fine-tuning model is adjusted using the Step-DPO data set to obtain the Step-DPO model as the third intermediate question-answering task processing model. Then continue to generate new Step-DPO data sets. Based on this, the Step-DPO data set is iteratively generated and the RFT fine-tuning model is optimized. In each round of iteration, the second intermediate question-answering task processing model of the current round is fine-tuned using the previously generated Step-DPO data set, and the optimized RFT strategy is continuously updated to improve the model's reasoning performance in complex tasks. The iteration cutoff conditions are determined according to different task requirements, such as: maximum training rounds, convergence conditions, no longer improving performance, computing resource limitations, early termination, training goal achievement, etc.

[0057] The second intermediate question-answering task processing model is adjusted according to the first target training data set to obtain a third intermediate question-answering task processing model. The question-answering task in the SFT data contained in the initial training data set is input into the third intermediate question-answering task processing model, and the generator outputs the reasoning process of processing the question-answering task. Each reasoning process is refined and split according to the steps to obtain the second reasoning step in the reasoning process. The second reasoning step is quality evaluated, and the second annotation information of the second reasoning step is determined according to the evaluation result, and the output result of the question-answering task output by the third intermediate question-answering task processing model is determined.

[0058] The reference results of the question-answering task are obtained in the initial training data set, and the accuracy of the third intermediate question-answering task processing model is obtained based on the reference results and the output results. For example, the third intermediate question-answering task processing model has 100 output results for 100 question-answering tasks, of which 90 output results are the same as the corresponding reference results, and the accuracy is 90 / 100 = 90%. In addition, the performance of the third intermediate question-answering task processing model is evaluated, and its accuracy can be determined to evaluate the performance of the fine-tuned model and improve the model's reasoning ability.

[0059] The first preset threshold is, for example, 99%, 99.5% or other values ​​less than or equal to 100%. When the accuracy is greater than the first preset threshold, it means that the accuracy of the third intermediate question-answering task processing model of the current iteration round is high and can meet the needs of the current question-answering task, and the third intermediate question-answering task processing model is used as the target model.

[0060] When the accuracy is less than or equal to the first preset threshold, it means that the accuracy of the third intermediate question and answer task processing model in the current iteration round is low and needs further adjustment. The Step-DPO dataset contains the steps, preference information and standard answers of the first intermediate question and answer task processing model when solving the question and answer task. Therefore, according to the first reasoning step, the first annotation information and the initial training dataset, a new Step-DPO dataset is generated, and the new Step-DPO dataset is used as the second target training dataset. Update the current number of iterations. During the execution of steps S101 to S104, the number of iterations is 0. When executing the above steps in this embodiment, the number of iterations is 1. Each update of the current number of iterations means adding 1 to the number of iterations.

[0061] The second target training data set is used as the first target training data set, and the third intermediate question and answer task processing model is used as the second intermediate question and answer task processing model. The subsequent steps are executed starting from adjusting the second intermediate question and answer task processing model according to the first target training data set until the accuracy is greater than the first preset threshold or the number of iterations is equal to the second preset threshold. Then the process ends and the third intermediate question and answer task processing model is used as the target model.

[0062] In this embodiment, the question-answering task processing model is continuously iteratively adjusted and optimized to continuously improve the performance of the question-answering task processing model. In each round of iteration, the target training data set is regenerated to further optimize the quality of the preference data, thereby improving the overall performance of the question-answering task processing model. In addition, multiple iterations can also enhance the model's adaptability to new tasks and improve the generalization ability of the question-answering task processing model.

[0063] In some optional embodiments, based on the initial training data set and the first intermediate question and answer task processing model, the first reasoning step of the question and answer task processed by the first intermediate question and answer task processing model and the first annotation information of the first reasoning step are obtained, and based on the first reasoning step, the first annotation information and the initial training data set, a first target training data set is generated, including: inputting the question and answer task in the initial training data set into the first intermediate question and answer task processing model to obtain a second preset number of reasoning processes corresponding to the question and answer task and an output result of the reasoning process; obtaining a reference result of the question and answer task corresponding to the reasoning process in the initial training data set, and generating third annotation information of the reasoning process based on the output result and the reference result, wherein the third annotation information is used to determine whether the output result is equal to the reference result; splitting the reasoning process to obtain a third preset number of first reasoning steps; determining the first annotation information of the first reasoning step based on the third annotation information of the reasoning process where the first reasoning step is located, wherein the first annotation information is used to determine the probability of obtaining the reference result according to the first reasoning step; generating the first target training data set based on the first reasoning step, the first annotation information and the reference result of the question and answer task.

[0064] Specifically, in the result reward scoring mechanism, given a question-answering task and its output result, a real value r is assigned to the output result, and the real value r is used to indicate whether the output result is correct. If the output result is correct, the real value r is 1; if the output result is wrong, the real value r is 0. On this basis, the automated process labeling framework proposed in this embodiment first scores the result of each reasoning process derived from a certain step, and then evaluates the quality of the step. Specifically, by breaking down the reasoning process into discrete steps, each step is represented by a label sequence. Step t Status Defined as a prefix to an inference chain, by adding new inference steps Convert the state to ,in yes and connection, that is In this process, for the intermediate steps By using the "completer", M inference processes are obtained, for example, 4, 5, 6, ... The specific value is set according to actual needs. Steps The quality of is determined by the scores of the M derived answers generated from this step. A step that can derive a reasonable result can be regarded as a good reasoning step, and the score of each derived answer is based on the result score.

[0065] Based on the above content, this embodiment determines the first annotation information of the first reasoning step, and then generates the first target training data set. Figure 3 Provide explanation.

[0066] Input the question-answering task in the SFT data contained in the initial training data set into the first intermediate question-answering task processing model, output a second preset number of reasoning processes for processing the question-answering task, and determine the output result of the question-answering task output by the first intermediate question-answering task processing model after processing the question-answering task according to different reasoning processes. The second preset number is, for example: 100, 200, ... No specific number limit is made here. Figure 3 As shown, the question-answering task containing the standard answer is input into the first intermediate question-answering task processing model, and the complete reasoning process S is generated using the first intermediate question-answering task processing model.

[0067] Obtain the reference results of the question-answering task corresponding to the reasoning process in the initial training data set, and evaluate the correctness of the reasoning process based on the output results and the reference results. If the output result is the same as the reference result, the reasoning process is correct, and the third annotation information of the reasoning process is, for example, r, r=1; if the reasoning process is incorrect, the third annotation information of the reasoning process is, for example, r, r=0.

[0068] The reasoning process is split to obtain a third preset number of first reasoning steps, for example: Figure 3 As shown, the reasoning process is S. Split S and get 、 、 、…、 ,common K The first step of reasoning, K The value is set according to actual needs, without any specific quantity limit.

[0069] The first annotation information is, for example, a positive sample step (chosen) and a negative sample step (reject). A positive sample step indicates that the probability of obtaining the standard answer to the question-answering task through the first reasoning step is high, and a negative sample step indicates that the probability of obtaining the standard answer to the question-answering task through the first reasoning step is low. According to the third annotation information of the reasoning process where the first reasoning step is located, the first annotation information of the first reasoning step is determined, for example: Strategy 1, selects the correct reasoning process S, whose steps (and a random intermediate step, such as ) as a positive sample, using the LLMs model Complete a wrong answer as a negative sample. Strategy 2: Score each step. If The lowest score is taken as a negative sample, and then the LLMs model is used to Finally, a positive sample is selected, and a set of step-level preference data for the problem is obtained.

[0070] A data set about a preferred order of reasoning steps may be obtained according to the first annotation information.

[0071] Each data entry of the Step-DPO dataset includes a comparison of two first reasoning steps. According to the first annotation information of the first reasoning step, it can be determined which of the two first reasoning steps is better. The Step-DPO dataset contains the steps, preference information and standard answers of the first intermediate question-answering task processing model when solving the question-answering task. Therefore, according to the first reasoning step, the first annotation information and the initial training dataset, an initial Step-DPO dataset is generated, and the Step-DPO dataset is used as the first target training dataset.

[0072] In this embodiment, the third annotation information of the reasoning process is generated according to the output results of the reasoning process and the reference results of the question and answer questions, thereby realizing automatic annotation of the steps. It can not only provide high-quality annotation samples for the training of the question and answer task processing model, but also can be used as preference data in large-model reinforcement learning, helping to further optimize the step-level strategy of the reasoning process, thereby improving the reasoning ability and diversity of the model in practical applications.

[0073] In some optional embodiments, the first annotation information of the first reasoning step is determined according to the third annotation information of the reasoning process where the first reasoning step is located, including: when it is determined according to the third annotation information that the output result of the reasoning process where the first reasoning step is located is the same as the reference result, taking the first reasoning step as the first intermediate step; determining the target position of the first intermediate step in the reasoning process; generating the first intermediate reasoning process according to the first intermediate step, the target position, the question and answer task, and the first intermediate question and answer task processing model, wherein the output result of the first intermediate reasoning process is different from the reference result; taking the step where the position of the first intermediate reasoning process is the same as the target position as the second intermediate step corresponding to the first intermediate step; taking the first preset information as the first annotation information of the first intermediate step, wherein the first preset information is used to characterize that the probability of obtaining the reference result according to the first intermediate step is greater than the third preset threshold; taking the second preset information as the first annotation information of the second intermediate step, wherein the second preset information is used to characterize that the probability of obtaining the reference result according to the first intermediate step is less than the fourth preset threshold; taking both the first intermediate step and the second intermediate step as the first reasoning step.

[0074] Specifically, this embodiment performs a quality assessment on each first reasoning step. The strategy adopted in this embodiment for performing quality assessment on the steps is based on the scoring of the problem-solving results. The core idea of ​​this strategy is to score the process through the final result of the reasoning process, and to extract positive sample steps from the reasoning process that can derive the correct answer. Specifically as follows: positive sample extraction, selecting a step of the correct final result as the positive sample; prefix completion: completing the prefix of the selected step to generate a new reasoning process; extracting negative samples: using a hard estimation method to evaluate the correctness of the newly generated reasoning process. If the final answer is wrong, the step is marked as a negative sample and included in the negative sample set. This embodiment uses a hard estimation method to estimate the steps. Quality , as shown in formula (1). As long as this step can generate a wrong answer , then the reasoning step is considered wrong.

[0075] (1)

[0076] in, A collection of answers to all reasoning processes.

[0077] First, according to the third annotation information, the reasoning process that can obtain the correct answer is screened out, and it is determined whether the output result of the reasoning process in the first reasoning step is the same as the reference result. If they are the same, the reasoning process can obtain the correct answer. For example: the third annotation information is r, if r is 1, it means that the output result is the same as the reference result, and if r is 0, it means that the output result is different from the reference result.

[0078] When it is determined according to the third labeling information that the output result of the reasoning process in which the first reasoning step is located is the same as the reference result, the first reasoning step is used as the first intermediate step, and the first intermediate step is a positive sample step (chosen).

[0079] Determine the target position of the first intermediate step in the reasoning process. Input the first intermediate question-answering task processing model according to the first intermediate step, the target position, and the question-answering task, use the first intermediate question-answering task processing model to output multiple reasoning processes corresponding to the first intermediate step, and take the reasoning process whose output result is different from the reference result as the first intermediate reasoning process, and the score of the first intermediate reasoning process is 0.

[0080] The step where the position of the first intermediate reasoning process is the same as the target position is taken as the second intermediate step corresponding to the first intermediate step, and the second intermediate step is a negative sample step (reject). For example: the first intermediate step is , then the first intermediate reasoning process K step as the second intermediate step.

[0081] The first annotation information may be, for example, a positive sample step (chosen) and a negative sample step (reject). The positive sample step indicates that the probability of obtaining the standard answer to the question-answering task through the first reasoning step is greater than the third preset threshold, and the negative sample step indicates that the probability of obtaining the standard answer to the question-answering task through the first reasoning step is less than the fourth preset threshold. The third preset threshold may be, for example, 90%, 91%... or other larger values. The fourth preset threshold may be, for example, 20%, 25%,... or other smaller values.

[0082] The first preset information is, for example, a positive sample step (chosen). The second preset information is, for example, a negative sample step (reject). The first preset information is used as the first annotation information of the first intermediate step, the second preset information is used as the first annotation information of the second intermediate step, and both the first intermediate step and the second intermediate step are used as the first reasoning step. For example, the reasoning process is split into steps by step: , , ...., if the final answer of the reasoning process matches the standard answer, a step is randomly selected as a positive sample. As a positive sample. Then use the question answering task, Generate a new Steps, Evaluation The accuracy of the question-answering task, , Complete and generate 4 reasoning processes. If there is an error in the reasoning process, the newly generated as negative samples.

[0083] In this embodiment, the first reasoning step and the first annotation information of the first reasoning step are determined, and the traditional result-level scoring is decomposed into a more refined step-level scoring, so as to more accurately evaluate and optimize the reasoning process of the question-answering task processing model. In addition, the ability to automatically generate the first annotation information not only reduces the reliance on manually annotated data, but also reduces the cost of data collection and significantly improves the efficiency of data generation.

[0084] In some optional embodiments, a first intermediate reasoning process is generated according to the first intermediate step, the target position, the question and answer task, and the first intermediate question and answer task processing model, including: when the target position is the end position of the reasoning process, taking the previous step of the first intermediate step in the reasoning process as the second intermediate step; taking the first identifier of the second intermediate step as the first prompt information, inputting the question and answer task and the first prompt information into the first intermediate question and answer task processing model to obtain the first intermediate reasoning process; when the target position is a non-starting position of the reasoning process, taking the first identifier as the second prompt information; inputting the question and answer task and the second prompt information into the first intermediate question and answer task processing model to obtain the first intermediate reasoning process.

[0085] Specifically, in this embodiment, the first intermediate question-answering task processing model is set as a generator. Based on the questions in a given data set or the prompt word m with a partial solution process, the generator is used to generate a derived reasoning process.

[0086] When the target position is the end position of the reasoning process, the step before the first intermediate step in the reasoning process is used as the second intermediate step. For example, the first intermediate step is , is at the end of the reasoning process, that is is the last step in the reasoning process. Set as the second intermediate step.

[0087] First identifier, for example: prefix of the second intermediate step . The first prompt information is, for example, prompt word m. The first identifier of the second intermediate step is used as the first prompt information, and the question-answering task and the first prompt information are input into the first intermediate question-answering task processing model to complete the complete reasoning process such as the first intermediate reasoning process. If the output result of the first intermediate reasoning process is different from the reference result, the last step of the first intermediate reasoning process is set as a negative sample step.

[0088] When the target position is not the starting position of the reasoning process, the first intermediate step is the intermediate step of the reasoning process. The first identifier of the first intermediate step is directly used as the second prompt information, and the question-answering task and the second prompt information are input into the first intermediate question-answering task processing model, and the final result is completed to generate a first intermediate reasoning process that is different from the original reasoning process and will obtain an erroneous final answer.

[0089] For example: The last step is different for two samples. - same, Different, because the last step will not continue to generate, so the wrong answer generated under the same prefix That is, negative samples; the intermediate steps are different, such as Different, use question-answering tasks, , Generate an incorrect reasoning process to obtain negative samples .

[0090] In this embodiment, the reasoning process of the first intermediate step is completed to derive multiple reasoning processes. By judging whether a reference result can be obtained according to the derived reasoning process, the first intermediate step is evaluated and the first annotation information is generated. There is no need to manually annotate the steps, which reduces the reliance on manually annotated data, reduces the cost of data collection, and significantly improves the efficiency of data generation.

[0091] In some optional embodiments, after taking the step whose position in the first intermediate reasoning process is the same as the target position as the second intermediate step corresponding to the first intermediate step, the method further includes: determining a first similarity between the first intermediate step and the second intermediate step; and deleting the first intermediate step and the second intermediate step when the first similarity is greater than or equal to a fifth preset threshold.

[0092] Specifically, the first intermediate step and the second intermediate step are compared to determine the first similarity between the first intermediate step and the second intermediate step. The first similarity is, for example, similarity based on word frequency (Bag of Words, BoW), which represents the text as a bag of words model, ignores the word order, and only focuses on the word frequency. Then, cosine similarity can be used to measure the angle between the two vectors. The smaller the angle, the higher the similarity. Jaccard similarity measures similarity by calculating the ratio of the intersection and union of two sets. Word Embeddings maps words to vectors in a high-dimensional space so that semantically similar words are close in the vector space. A pre-trained word vector model can be used to represent the text, and the text similarity can be calculated by averaging word vectors or sentence embedding models (such as Sentence-BERT).

[0093] The fifth preset threshold is, for example, 0.9, 0.91, ... or other larger data within 1. In order to increase the diversity and effectiveness of the data set, this embodiment needs to ensure that there is sufficient difference between the positive sample step and the negative sample step. When the first similarity is greater than or equal to the fifth preset threshold, it means that the similarity between the first intermediate step and the second intermediate step is too high. In order to ensure the diversity and effectiveness of the data set, the first intermediate step and the second intermediate step are deleted.

[0094] In this embodiment, the similarity between the positive sample step and the corresponding negative sample step is calculated, and the positive sample steps and negative sample steps with higher similarity are deleted to avoid incorporating overly similar samples into the target training data set, thereby ensuring the diversity and effectiveness of the target training data set.

[0095] In some optional embodiments, the first annotation information of the first reasoning step is determined according to the third annotation information of the reasoning process where the first reasoning step is located, including: when it is determined according to the third annotation information that the output result of the reasoning process where the first reasoning step is located is the same as the reference result, according to the third preset number of first reasoning steps, the question and answer task and the first intermediate question and answer task processing model, obtaining the output results of the second intermediate reasoning process and the fourth preset number of second intermediate reasoning processes; determining the third intermediate reasoning process among the fourth preset number of second intermediate reasoning processes, wherein the output result of the third intermediate reasoning process is equal to the reference result; determining a first proportion of the third intermediate reasoning process to the fourth preset number of second intermediate reasoning processes, and determining the first reasoning step according to the first proportion. an evaluation parameter; determining a third intermediate step in which the first evaluation parameter is the smallest and the first evaluation parameter is less than a sixth preset threshold value among a third preset number of first reasoning steps; determining a second identifier of the third intermediate step, and determining a fourth intermediate step in which the identifier is the same as the second identifier and the first evaluation parameter is equal to a preset value from the third preset number of first reasoning steps; using the first preset information as the first annotation information of the fourth intermediate step, wherein the first preset information is used to represent that the probability of obtaining a reference result according to the fourth intermediate step is greater than the third preset threshold value; using the second preset information as the first annotation information of the third intermediate step, wherein the second preset information is used to represent that the probability of obtaining a reference result according to the third intermediate step is less than the fourth preset threshold value; and using both the third intermediate step and the fourth intermediate step as the first reasoning step.

[0096] Specifically, this embodiment generates multiple different reasoning processes for the same problem by diversifying the generated reasoning process, and evaluates the score of each problem-solving step. First, evaluate the score of each step in the reasoning process; select the step with the lowest score as a potential negative sample, and verify the final answer of the negative sample. If the final answer of the negative sample is wrong, use the model to complete its prefix to obtain a correct reasoning process, and extract the difference step as a positive sample. If the final answer of the negative sample is correct, it is necessary to select a difference step with a score that is significantly higher than the score of the negative sample. At this time, after completing a correct result from the prefix of the step, the automatic labeling method is used again to score the difference step. Only when the score of the step reaches the preset standard will it be included in the data set as a positive sample. This embodiment uses a hard estimation method to estimate the steps. Quality , as shown in formula (2), the frequency of a step generating a correct answer is set as the quality of the step .

[0097] (2)

[0098] Among them, N is from The number of subsequent inference processes. is the answer to the derived reasoning process.

[0099] According to the design scheme provided by the present invention, the scoring strategy is: first, based on the questions in a given data set or the prompt word m with a partial solution process, a generator is used to generate N derived reasoning processes. Then, an evaluation program is used to score the correctness of each derived process, thereby calculating the score of each step. The fourth preset number is N, and the value of N is set according to actual needs. The first intermediate question-answering task processing model is used as a generator.

[0100] Based on the above, the reasoning process is divided into two categories according to whether the final result of the reasoning process matches the standard answer.

[0101] When the output result of the reasoning process in which the first reasoning step is located is determined to be the same as the reference result based on the third annotation information, the reasoning process is determined to be correct, and each step needs to be completed using a completer until the final answer is obtained, that is, a complete second intermediate reasoning process is obtained, and the number of second intermediate reasoning processes is the fourth preset number. The third preset number of first reasoning steps and question-and-answer tasks are used as input information, and the input information is input into the generator. The generator will complete each first reasoning step until the final answer is obtained, that is, a complete reasoning process is obtained, and a fourth preset number of second intermediate reasoning processes are obtained. In addition, the generator will output the output result of processing the question-and-answer task according to each second intermediate reasoning process.

[0102] For example, the fourth preset number is equal to 4. The third intermediate reasoning process is determined in the 4 second intermediate reasoning processes, and the output result of the third intermediate reasoning process is equal to the reference result. For example, there are 3 third intermediate reasoning processes, and the first proportion of the third intermediate reasoning process in the fourth preset number of second intermediate reasoning processes is determined. The first evaluation parameter of the first reasoning step is determined according to the first proportion. For example, the first proportion is 3 / 4=0.75, and 0.75 can be directly set as the first evaluation parameter of the first reasoning step.

[0103] The above process is as follows Figure 4 As shown in the figure, the question-answering task is "If a+b=8, b+c=-3 and a+c=-5, what is the value of abc" and the reference result is "-120". For more information, see Figure 4 , I will not repeat it here. If the reasoning process S is correct, the splitting steps are , , , ...,right , , ...score each step and use the generator to combine the question-answering task and As the model input, complete the 4 reasoning processes. The final answers (Ans) of three reasoning processes are -120, which is the same as the reference result. The Ans of one reasoning process is 10, which is different from the reference result. The fraction of the step is 3 / 4, and so on , .... are also scored. For the four generated derivative processes, the correctness is evaluated using the result scoring procedure, and the correct frequency in the four processes is used as the score of the step to evaluate its quality.

[0104] The sixth preset threshold is, for example, 0.25, 0.30 or other values. Among the third preset number of first inference steps, the step with the smallest first evaluation parameter and the first evaluation parameter less than the sixth preset threshold is selected as the third intermediate step, and the third intermediate step is a negative sample step. Further, a step with the same prefix as the step and a score of 1 is selected as a positive sample. Therefore, the preset value is, for example, 1. Determine the second identifier of the third intermediate step, and determine the fourth intermediate step whose identifier is the same as the second identifier and the first evaluation parameter is equal to the preset value from the third preset number of first inference steps. The fourth intermediate step is a positive sample step. For example: select the step with the smallest score among all steps and the score is less than or equal to 0.25 as a negative sample. If The score of is the smallest and is 0, then select is a negative sample. Further use of question answering tasks, steps and Generates a new step as model input , for new Rating, if new If the score is 1, then the new As positive samples, thus forming positive and negative sample pairs.

[0105] The first preset information is, for example, a positive sample step (chosen). The second preset information is, for example, a negative sample step (reject). The first preset information is used as the first annotation information of the fourth intermediate step, the second preset information is used as the first annotation information of the third intermediate step, and the third intermediate step and the fourth intermediate step are both used as the first reasoning step.

[0106] In this embodiment, the first annotation information of the first reasoning step is determined to achieve automatic annotation of the steps, which not only improves the quality of the target training data set, but also enhances the data diversity in the target training data set, ensuring that the target training data set can comprehensively cover various problem-solving paths. At the same time, according to the first annotation information, the valid and invalid reasoning steps in the target training data set can be accurately identified.

[0107] In some optional embodiments, after determining a fourth intermediate step whose identifier is the same as the second identifier and whose first evaluation parameter is equal to a preset value from a third preset number of first inference steps, the method further includes: determining a second similarity between the third intermediate step and the fourth intermediate step; and deleting the third intermediate step and the fourth intermediate step when the second similarity is greater than or equal to a seventh preset threshold.

[0108] Specifically, the third intermediate step and the fourth intermediate step are compared to determine the second similarity between the third intermediate step and the fourth intermediate step. The second similarity is, for example, similarity based on word frequency (Bag of Words, BoW), which represents the text as a bag of words model, ignores the word order, and only focuses on the word frequency. Then, cosine similarity can be used to measure the angle between the two vectors. The smaller the angle, the higher the similarity. Jaccard similarity measures similarity by calculating the ratio of the intersection and union of two sets. Word Embeddings maps words to vectors in a high-dimensional space so that semantically similar words are close in the vector space. A pre-trained word vector model can be used to represent the text, and the text similarity can be calculated by averaging word vectors or sentence embedding models (such as Sentence-BERT).

[0109] The seventh preset threshold is, for example, 0.9, 0.91, ... or other larger data within 1. In order to make the data set diverse and effective, this embodiment needs to ensure that there is sufficient difference between the positive sample step and the negative sample step. When the second similarity is greater than or equal to the fifth preset threshold, it means that the similarity between the third intermediate step and the fourth intermediate step is too high. In order to ensure the diversity and effectiveness of the data set, the third intermediate step and the fourth intermediate step are deleted.

[0110] In some optional embodiments, the first annotation information of the first reasoning step is determined according to the third annotation information of the reasoning process in which the first reasoning step is located, including: when it is determined according to the third annotation information that the output result of the reasoning process in which the first reasoning step is located is different from the reference result, according to the third preset number of first reasoning steps, the question and answer task and the first intermediate question and answer task processing model, a fourth preset number of fourth intermediate reasoning processes and the output result of the fourth intermediate reasoning process are obtained; a fifth intermediate reasoning process is determined among the fourth preset number of fourth intermediate reasoning processes, wherein the output result of the fifth intermediate reasoning process is equal to the reference result; a second proportion of the fifth intermediate reasoning process in the fourth preset number of fourth intermediate reasoning processes is determined, and a second evaluation parameter of the first reasoning step is determined according to the second proportion; Determine the fifth intermediate step in which the second evaluation parameter is the smallest and the second evaluation parameter is less than the eighth preset threshold; obtain the sixth intermediate reasoning process according to the fifth intermediate step, the question-answering task and the first intermediate question-answering task processing model, wherein the output result of the sixth intermediate reasoning process is equal to the reference result; determine the third identifier of the fifth intermediate step, and take the step in the sixth intermediate reasoning process whose identifier is equal to the third identifier as the sixth intermediate step; use the first preset information as the first annotation information of the sixth intermediate step, wherein the first preset information is used to characterize that the probability of obtaining the reference result according to the sixth intermediate step is greater than the third preset threshold; use the second preset information as the first annotation information of the fifth intermediate step, wherein the second preset information is used to characterize that the probability of obtaining the reference result according to the fifth intermediate step is less than the fourth preset threshold; take both the fifth intermediate step and the sixth intermediate step as the first reasoning step.

[0111] Specifically, the reasoning process is divided into two categories according to whether the final result of the reasoning process matches the standard answer. The first intermediate question-answering task processing model is used as a generator. The fourth preset number equals 4 as an example for illustration.

[0112] When it is determined according to the third annotation information that the output result of the reasoning process in which the first reasoning step is located is different from the reference result, it is determined that the reasoning process is wrong, and each step needs to be completed using a completer. The third preset number of first reasoning steps and the question-answering task are used as input information, and the input information is input into the generator. The generator will complete each first reasoning step until the final answer is obtained, that is, the complete reasoning process is obtained, and the fourth preset number of fourth intermediate reasoning processes are obtained. In addition, the generator will output the output result of processing the question-answering task according to each fourth intermediate reasoning process.

[0113] The output result of the fifth intermediate reasoning process is equal to the reference result, and the fifth intermediate reasoning process is the correct process. The fifth intermediate reasoning process is determined in the four fourth intermediate reasoning processes. For example, if there is one fifth intermediate reasoning process, the second proportion of the fifth intermediate reasoning process in the fourth preset number of fourth intermediate reasoning processes is 1 / 4=0.25, and 0.25 can be directly set as the second evaluation parameter of the first reasoning step.

[0114] The above process is as follows Figure 4 As shown, if the reasoning process S is wrong, the splitting steps are , , , ..., for each step such as , using the generator to combine the question-answering task and As the model input, 4 results are generated. The correctness of each result is judged using a scoring procedure. If 3 of the 4 completed processes match the standard answer, then The score for is 0.75. , , ...also rated.

[0115] The step with the largest number of errors in the first derived four processes is selected, and the score of the step is less than or equal to 0.25, and it is taken as a negative sample step. Therefore, the eighth preset threshold is, for example, 0.25, 0.30 or other values. The fifth intermediate step with the smallest second evaluation parameter and the second evaluation parameter less than the eighth preset threshold is determined in the third preset number of first reasoning steps, and the fifth intermediate step is a negative sample step.

[0116] At the same time, the prefix of the fifth intermediate step is used as the prompt information, the prompt information and the question-answering task are input into the generator, and a new correct reasoning process is generated as the sixth intermediate reasoning process, and the output result of the sixth intermediate reasoning process is equal to the reference result. The third identifier of the fifth intermediate step is determined, and the step with the identifier equal to the third identifier in the sixth intermediate reasoning process is taken as the sixth intermediate step, and the sixth intermediate step is a positive sample step, thereby constructing positive and negative samples for the incorrect parsing process, namely the fifth intermediate step and the sixth intermediate step.

[0117] The above process is as follows Figure 4 As shown in , the step with the largest number of errors among all steps, i.e. the smallest score, is selected as the negative sample step. The score of is the smallest and is 0, then select is a negative sample step. Further use of question answering tasks, steps and Input generator to generate a new step , for new Rating, if new If the score is 1, then the new As positive samples, thus forming positive and negative sample pairs.

[0118] The first preset information is, for example, a positive sample step (chosen). The second preset information is, for example, a negative sample step (reject). The first preset information is used as the first annotation information of the sixth intermediate step, the second preset information is used as the first annotation information of the fifth intermediate step, and the fifth intermediate step and the sixth intermediate step are both used as the first reasoning step.

[0119] In this embodiment, the first annotation information of the first reasoning step is determined to achieve automatic annotation of the steps, which not only improves the quality of the target training data set, but also enhances the data diversity in the target training data set, ensuring that the target training data set can comprehensively cover various problem-solving paths. At the same time, according to the first annotation information, the valid and invalid reasoning steps in the target training data set can be accurately identified.

[0120] In some optional embodiments, a sixth intermediate reasoning process is obtained according to the fifth intermediate step, the question and answer task, and the first intermediate question and answer task processing model, including: using the third identifier of the fifth intermediate step as the third prompt information; inputting the question and answer task and the third prompt information into the first intermediate question and answer task processing model to obtain a fifth preset number of seventh intermediate reasoning processes; and determining the sixth intermediate reasoning process among the fifth preset number of seventh intermediate reasoning processes.

[0121] Specifically, this embodiment sets the first intermediate question-answering task processing model as a generator. Based on the questions in a given data set or the prompt word m with a partial solution process, the generator is used to generate a derived reasoning process. The third identifier is, for example, the prefix of the fifth intermediate step. .

[0122] Directly use the third identifier of the fifth intermediate step as the second prompt information, input the question and answer task and the third prompt information into the first intermediate question and answer task processing model, complete them to the final result, and generate the fifth preset number of seventh intermediate reasoning processes. The fifth preset number represents multiple, and there is no specific quantity limit here.

[0123] A sixth intermediate reasoning process whose output result is the same as the reference result is determined among the fifth preset number of seventh intermediate reasoning processes, and the sixth intermediate reasoning process is a correct reasoning process.

[0124] In some optional embodiments, before adjusting the question-answering task processing model to be trained based on the initial training data set and the first model training algorithm, the method also includes: obtaining a benchmark data set; extracting question-answering tasks and reference results of question-answering tasks from the benchmark data set; and creating an initial training data set based on the question-answering tasks and the reference results.

[0125] Specifically, we first build a benchmark dataset and determine the reasoning types that the question-answering task processing model needs to focus on, such as arithmetic reasoning, Python code reasoning, etc. We obtain representative datasets under the reasoning type as benchmark datasets. For example, in terms of arithmetic reasoning type, we use two representative datasets: the GSM8K dataset and the MATH dataset, as benchmark datasets; the GSM8K dataset contains elementary school math word problems; the MATH dataset covers challenging competition math problems, using the training sets of the GSM8K and MATH datasets. In terms of Python code reasoning type, we use GPT to generate a set of datasets as benchmark datasets. The dataset contains code that can pass five unit test cases, and on this basis, we generate process preference data for code reasoning through the model.

[0126] Extract question-answering tasks and reference results of question-answering tasks from the benchmark dataset. Taking the benchmark dataset as the MATH dataset as an example, the MATH dataset contains a first preset number of question-answering tasks and reference results corresponding to the question-answering tasks, where the first preset number represents multiple values ​​such as 500, 1000 or other values ​​that meet actual needs.

[0127] According to the question-answering task and the reference results, SFT data and RFT data are created. The initial training data set includes the above SFT data and RFT data, which can be called as needed. SFT data refers to the data used for supervised fine-tuning, which usually consists of model input and target output corresponding to the model input. RFT data refers to the data used in the reinforcement learning fine-tuning process.

[0128] In some optional embodiments, before adjusting the question and answer task processing model to be trained according to the initial training data set and the first model training algorithm, the method also includes: obtaining the initial question and answer task processing model and a sixth preset number of training samples, wherein the training samples include the task to be predicted and the target result of the task to be predicted; inputting the task to be predicted into the initial question and answer task processing model to obtain a prediction result; adjusting the initial question and answer task processing model according to the prediction result and the target result to obtain a pre-trained model, and using the pre-trained model as the question and answer task processing model to be trained.

[0129] Specifically, untrained LLMs are obtained as the initial question-answering task processing model. A large-scale data set is obtained, which includes vocabulary prediction questions, math questions, logical reasoning questions, computational problem solving questions, etc., and also includes the standard answers corresponding to the above questions.

[0130] A sixth preset number of questions are obtained from the large-scale data set as tasks to be predicted, and the standard answers corresponding to the questions are used as target results. The sixth preset number of tasks to be predicted and the corresponding target results constitute a sixth preset number of training samples, and the sixth preset number is, for example, 100, 200 or other values ​​that meet actual needs.

[0131] Input the task to be predicted into the initial question-answering task processing model to obtain the prediction result. According to the prediction result and the target result, adjust the initial question-answering task processing model to obtain a pre-trained model. For example, input the prediction result and the target result into the loss function, calculate the loss value, and adjust the initial question-answering task processing model according to the loss value to obtain the pre-trained model. Use the pre-trained model as the question-answering task processing model to be trained.

[0132] In some optional implementations, the present invention needs to determine the first similarity between the first intermediate step and the second intermediate step, needs to determine the second similarity between the third intermediate step and the fourth intermediate step, and may also determine the similarity between the fifth intermediate step and the sixth intermediate step, and the specific process of determining the similarity may include steps A1 to A6. The calculation of the similarity between the first intermediate step and the second intermediate step is used as an example for description.

[0133] Step A1, performing word segmentation processing on the information of the first intermediate step and the information of the second intermediate step.

[0134] Step A2: Build a Bag-of-Words model.

[0135] Specifically, a list containing all different words is constructed as a word bag. Then, for the first intermediate step and the second intermediate step, the number of times each word appears in the first intermediate step and the second intermediate step is recorded.

[0136] Step A3, calculate term frequency.

[0137] Specifically, for each word, the frequency of each word in the first intermediate step and the second intermediate step is calculated using formula (3) ( ).

[0138] (3)

[0139] Step A4, calculating the inverse document frequency.

[0140] Specifically, we need to know how common this word is in the entire corpus. To do this, we calculate the inverse document frequency ( ), that is, how many different documents the word appears in.

[0141] (4)

[0142] in, is the total number of steps, represents the number of steps containing vocabulary t.

[0143] Step A5, combine and get value.

[0144] Specifically, by combining the above two and Multiply them together to get the final value.

[0145] Step A6, converting the first intermediate step and the second intermediate step into vector form respectively, calculating the cosine similarity between the two vectors, and determining the similarity between the first intermediate step and the second intermediate step according to the calculation result.

[0146] Specifically, once we have all the words in each step After the value is obtained, it can be converted into a vector form at each step. For example, the first intermediate step and the second intermediate step have the same vocabulary set V = {v1, v2, ..., v n}, construct the n-dimensional vector a=(a1, a2, …, a) corresponding to the first intermediate step according to the vocabulary set V n ), construct the n-dimensional vector b=(b1, b2, …, b) corresponding to the second intermediate step n ), vector a and vector b both contain words in vocabulary set V, and the specific value of n is set according to actual needs, such as 10, 11, etc.

[0147] The cosine similarity between the first intermediate step and the second intermediate step is calculated using the vectors of the first intermediate step and the second intermediate step, and the cosine similarity is used as the first similarity.

[0148] In this embodiment, the similarity between the positive sample step and the corresponding negative sample step is calculated, and the positive sample steps and negative sample steps with higher similarity are deleted to avoid incorporating overly similar samples into the target training data set, thereby ensuring the diversity and effectiveness of the target training data set.

[0149] In this embodiment, a question-answering task processing model training device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated hereafter. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0150] This embodiment provides a question-answering task processing model training device, such as Figure 5 As shown, including:

[0151] A first adjustment module 501 is used to adjust the question-answering task processing model to be trained according to the initial training data set and the first model training algorithm to obtain a first intermediate question-answering task processing model, wherein the initial training data set includes a first preset number of question-answering tasks;

[0152] The data set generation module 502 is used to obtain the first reasoning step and the first annotation information of the first reasoning step of the question-answering task processed by the first intermediate question-answering task processing model according to the initial training data set and the first intermediate question-answering task processing model, and generate the first target training data set according to the first reasoning step, the first annotation information and the initial training data set, wherein the first reasoning step is any step in the reasoning process of the question-answering task processed by the first intermediate question-answering task processing model;

[0153] A second adjustment module 503 is used to adjust the question-answering task processing model to be trained according to the initial training data set and the second model training algorithm to obtain a second intermediate question-answering task processing model;

[0154] The third adjustment module 504 is used to adjust the second intermediate question-answering task processing model according to the first target training data set to obtain a target model.

[0155] In some optional implementations, the third adjustment module 504 includes: an adjustment unit, which is used to adjust the second intermediate question and answer task processing model according to the first target training data set to obtain the third intermediate question and answer task processing model; a first determination unit, which is used to obtain the second reasoning step and the second annotation information of the second reasoning step of the question and answer task processed by the third intermediate question and answer task processing model according to the initial training data set and the third intermediate question and answer task processing model, and determine the output result of the question and answer task output by the third intermediate question and answer task processing model; a second determination unit, which is used to obtain the reference result of the question and answer task in the initial training data set, and obtain the accuracy of the third intermediate question and answer task processing model based on the reference result and the output result; the first A judgment unit, used to use the third intermediate question and answer task processing model as the target model when the accuracy is greater than a first preset threshold; a second judgment unit, used to generate a second target training data set according to the second reasoning step, the second annotation information and the initial training data set when the accuracy is less than or equal to the first preset threshold, and update the current number of iterations; a loop unit, used to use the second target training data set as the first target training data set, and execute subsequent steps starting from adjusting the second intermediate question and answer task processing model according to the first target training data set until the accuracy is greater than the first preset threshold or the number of iterations is equal to the second preset threshold, then the process ends and the third intermediate question and answer task processing model is used as the target model.

[0156] In some optional embodiments, the data set generation module 502 includes: an input unit, which is used to input the question and answer task in the initial training data set into the first intermediate question and answer task processing model to obtain a second preset number of reasoning processes and output results of the reasoning processes corresponding to the question and answer task; a first generation unit, which is used to obtain the reference result of the question and answer task corresponding to the reasoning process in the initial training data set, and generate third annotation information of the reasoning process based on the output result and the reference result, wherein the third annotation information is used to determine whether the output result is equal to the reference result; a splitting unit, which is used to split the reasoning process to obtain a third preset number of first reasoning steps; a third determination unit, which is used to determine the first annotation information of the first reasoning step based on the third annotation information of the reasoning process where the first reasoning step is located, wherein the first annotation information is used to determine the probability of obtaining the reference result according to the first reasoning step; a second generation unit, which is used to generate a first target training data set based on the first reasoning step, the first annotation information and the reference result of the question and answer task.

[0157] In some optional embodiments, the data set generation module 502 includes: a third judgment unit, which is used to use the first reasoning step as the first intermediate step when it is determined that the output result of the reasoning process in which the first reasoning step is located is the same as the reference result according to the third annotation information; a fourth determination unit, which is used to determine the target position of the first intermediate step in the reasoning process; a third generation unit, which is used to generate the first intermediate reasoning process according to the first intermediate step, the target position, the question and answer task, and the first intermediate question and answer task processing model, wherein the output result of the first intermediate reasoning process is different from the reference result; a first setting unit, which is used to use the step in which the position of the first intermediate reasoning process is the same as the target position as the second intermediate step corresponding to the first intermediate step; a second setting unit, which is used to use the first preset information as the first annotation information of the first intermediate step, wherein the first preset information is used to represent that the probability of obtaining the reference result according to the first intermediate step is greater than the third preset threshold; a third setting unit, which is used to use the second preset information as the first annotation information of the second intermediate step, wherein the second preset information is used to represent that the probability of obtaining the reference result according to the first intermediate step is less than the fourth preset threshold; and a fourth setting unit, which is used to use both the first intermediate step and the second intermediate step as the first reasoning step.

[0158] In some optional embodiments, the third generation unit includes: a first judgment submodule, which is used to use the previous step of the first intermediate step in the reasoning process as the second intermediate step when the target position is the end position of the reasoning process; a first input submodule, which is used to use the first identifier of the second intermediate step as the first prompt information, and input the question and answer task and the first prompt information into the first intermediate question and answer task processing model to obtain the first intermediate reasoning process; a second judgment submodule, which is used to use the first identifier as the second prompt information when the target position is not the starting position of the reasoning process; and a second input submodule, which is used to input the question and answer task and the second prompt information into the first intermediate question and answer task processing model to obtain the first intermediate reasoning process.

[0159] In some optional embodiments, the data set generation module 502 also includes: a fifth determination unit, used to determine the first similarity between the first intermediate step and the second intermediate step; and a fourth judgment unit, used to delete the first intermediate step and the second intermediate step when the first similarity is greater than or equal to a fifth preset threshold.

[0160] In some optional embodiments, the third determination unit includes: a third judgment submodule, which is used to obtain the output results of the second intermediate reasoning process and the second intermediate reasoning process according to a third preset number of first reasoning steps, question-and-answer tasks, and the first intermediate question-and-answer task processing model when it is determined that the output result of the reasoning process in which the first reasoning step is located is the same as the reference result according to the third annotation information; a first determination submodule, which is used to determine the third intermediate reasoning process in the fourth preset number of second intermediate reasoning processes, wherein the output result of the third intermediate reasoning process is equal to the reference result; a second determination submodule, which is used to determine the first proportion of the third intermediate reasoning process in the fourth preset number of second intermediate reasoning processes, and determine the first evaluation parameter of the first reasoning step according to the first proportion; and a third determination submodule, which is used to determine the output result of the third intermediate reasoning process and the output result of the third intermediate reasoning process in the third preset number of second intermediate reasoning processes. A third intermediate step is determined in the reasoning step, in which the first evaluation parameter is the smallest and the first evaluation parameter is less than a sixth preset threshold value; a fourth determination submodule is used to determine the second identifier of the third intermediate step, and determine the fourth intermediate step whose identifier is the same as the second identifier and the first evaluation parameter is equal to the preset value from the third preset number of first reasoning steps; a first setting submodule is used to use the first preset information as the first annotation information of the fourth intermediate step, wherein the first preset information is used to represent that the probability of obtaining a reference result according to the fourth intermediate step is greater than the third preset threshold value; a second setting submodule is used to use the second preset information as the first annotation information of the third intermediate step, wherein the second preset information is used to represent that the probability of obtaining a reference result according to the third intermediate step is less than the fourth preset threshold value; a third setting submodule is used to use both the third intermediate step and the fourth intermediate step as the first reasoning step.

[0161] In some optional embodiments, the third determination unit also includes: a fifth determination submodule, used to determine the second similarity between the third intermediate step and the fourth intermediate step; and a fourth judgment submodule, used to delete the third intermediate step and the fourth intermediate step when the second similarity is greater than or equal to a seventh preset threshold.

[0162] In some optional embodiments, the third determination unit also includes: a fifth judgment submodule, which is used to obtain a fourth preset number of fourth intermediate reasoning processes and the output result of the fourth intermediate reasoning process according to a third preset number of first reasoning steps, the question and answer task and the first intermediate question and answer task processing model when it is determined that the output result of the reasoning process in which the first reasoning step is located is different from the reference result according to the third annotation information; a sixth determination submodule, which is used to determine the fifth intermediate reasoning process among the fourth preset number of fourth intermediate reasoning processes, wherein the output result of the fifth intermediate reasoning process is equal to the reference result; a seventh determination submodule, which is used to determine the second proportion of the fifth intermediate reasoning process in the fourth preset number of fourth intermediate reasoning processes, and determine the second evaluation parameter of the first reasoning step according to the second proportion; an eighth determination submodule, which is used to determine that the second evaluation parameter is the smallest among the third preset number of first reasoning steps and the second evaluation parameter is less than the eighth A fifth intermediate step with a preset threshold; a ninth determination submodule, used to obtain a sixth intermediate reasoning process according to the fifth intermediate step, the question-and-answer task and the first intermediate question-and-answer task processing model, wherein the output result of the sixth intermediate reasoning process is equal to the reference result; a tenth determination submodule, used to determine the third identifier of the fifth intermediate step, and take the step in the sixth intermediate reasoning process whose identifier is equal to the third identifier as the sixth intermediate step; a fourth setting submodule, used to take the first preset information as the first annotation information of the sixth intermediate step, wherein the first preset information is used to characterize that the probability of obtaining the reference result according to the sixth intermediate step is greater than the third preset threshold; a fifth setting submodule, used to take the second preset information as the first annotation information of the fifth intermediate step, wherein the second preset information is used to characterize that the probability of obtaining the reference result according to the fifth intermediate step is less than the fourth preset threshold; a sixth setting submodule, used to take both the fifth intermediate step and the sixth intermediate step as the first reasoning step.

[0163] In some optional embodiments, the third determination unit also includes: a seventh setting submodule, used to use the third identifier of the fifth intermediate step as the third prompt information; a third input submodule, used to input the question and answer task and the third prompt information into the first intermediate question and answer task processing model to obtain a fifth preset number of seventh intermediate reasoning processes; an eleventh determination submodule, used to determine the sixth intermediate reasoning process among the fifth preset number of seventh intermediate reasoning processes.

[0164] In some optional embodiments, the device also includes: a first acquisition module for acquiring a benchmark data set; an extraction module for extracting question-answering tasks and reference results of question-answering tasks in the benchmark data set; and a creation module for creating an initial training data set based on the question-answering tasks and the reference results.

[0165] In some optional embodiments, the device also includes: a second acquisition module, used to acquire an initial question and answer task processing model and a sixth preset number of training samples, wherein the training samples include the task to be predicted and the target result of the task to be predicted; an input module, used to input the task to be predicted into the initial question and answer task processing model to obtain a prediction result; an adjustment module, used to adjust the initial question and answer task processing model according to the prediction result and the target result to obtain a pre-trained model, and use the pre-trained model as the question and answer task processing model to be trained.

[0166] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0167] The question-and-answer task processing model training device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0168] The embodiment of the present invention also provides a computer device having the above Figure 5 The question-answering task processing model training device shown.

[0169] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.

[0170] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0171] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0172] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0173] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0174] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0175] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0176] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0177] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the present invention.

Claims

1. A question-answering task processing model training method, characterized in that: The method comprises: According to the initial training data set and the first model training algorithm, the question-answering task processing model to be trained is adjusted to obtain a first intermediate question-answering task processing model, wherein the initial training data set includes a first preset number of question-answering tasks; According to the initial training data set and the first intermediate question and answer task processing model, the first reasoning step of the question and answer task processed by the first intermediate question and answer task processing model and the first annotation information of the first reasoning step are obtained, and according to the first reasoning step, the first annotation information and the initial training data set, a first target training data set is generated, wherein the first reasoning step is any step in the reasoning process of the question and answer task processed by the first intermediate question and answer task processing model; wherein, according to the initial training data set and the first intermediate question and answer task processing model, the first reasoning step of the question and answer task processed by the first intermediate question and answer task processing model and the first annotation information of the first reasoning step are obtained. , and generate a first target training data set according to the first reasoning step, the first annotation information and the initial training data set, including: inputting the question-answering task in the initial training data set into the first intermediate question-answering task processing model to obtain a second preset number of reasoning processes corresponding to the question-answering task and an output result of the reasoning process; obtaining a reference result of the question-answering task corresponding to the reasoning process in the initial training data set, and generating third annotation information of the reasoning process according to the output result and the reference result, wherein the third annotation information is used to determine whether the output result is equal to the reference result; splitting the reasoning process to obtain a third preset number of the first reasoning steps; Determine the first annotation information of the first reasoning step according to the third annotation information of the reasoning process where the first reasoning step is located, wherein the first annotation information is used to determine the probability of obtaining the reference result according to the first reasoning step; wherein, determine the first annotation information of the first reasoning step according to the third annotation information of the reasoning process where the first reasoning step is located, including: when it is determined according to the third annotation information that the output result of the reasoning process where the first reasoning step is located is the same as the reference result, take the first reasoning step as the first intermediate step; determine the target position of the first intermediate step in the reasoning process; generate a first intermediate reasoning process according to the first intermediate step, the target position, the question-answering task and the first intermediate question-answering task processing model, wherein the output result of the first intermediate reasoning process is different from the reference result; take the step whose position of the first intermediate reasoning process is the same as the target position as the second intermediate step corresponding to the first intermediate step; take the first intermediate step as the first reasoning step, and the probability of the first intermediate step obtaining the reference result is greater than a third preset threshold, and take the second intermediate step as the first reasoning step, and the probability of the second intermediate step obtaining the reference result is less than a fourth preset threshold; generate the first target training data set according to the first reasoning step, the first annotation information and the reference result of the question-answering task; According to the initial training data set and the second model training algorithm, the question-answering task processing model to be trained is adjusted to obtain a second intermediate question-answering task processing model; The second intermediate question-answering task processing model is adjusted according to the first target training data set to obtain a target model.

2. The method according to claim 1, characterized in that: The step of adjusting the second intermediate question-answering task processing model according to the first target training data set to obtain a target model includes: Adjusting the second intermediate question-answering task processing model according to the first target training data set to obtain a third intermediate question-answering task processing model; According to the initial training data set and the third intermediate question and answer task processing model, obtain the second reasoning step of the question and answer task processed by the third intermediate question and answer task processing model and the second annotation information of the second reasoning step, and determine the output result of the question and answer task output by the third intermediate question and answer task processing model; Obtaining a reference result of the question-answering task in the initial training data set, and obtaining the accuracy of the third intermediate question-answering task processing model based on the reference result and the output result; When the accuracy rate is greater than a first preset threshold, using the third intermediate question-answering task processing model as the target model; When the accuracy is less than or equal to the first preset threshold, generating a second target training data set according to the second reasoning step, the second annotation information and the initial training data set, and updating the current number of iterations; The second target training data set is used as the first target training data set, and subsequent steps are performed starting from adjusting the second intermediate question and answer task processing model according to the first target training data set, until the accuracy is greater than the first preset threshold or the number of iterations is equal to the second preset threshold, then the process ends and the third intermediate question and answer task processing model is used as the target model.

3. The method according to claim 1, characterized in that The determining, according to the third annotation information of the reasoning process in which the first reasoning step is located, the first annotation information of the first reasoning step includes: When it is determined according to the third annotation information that the output result of the reasoning process in which the first reasoning step is located is the same as the reference result, taking the first reasoning step as the first intermediate step; determining a target position of the first intermediate step in the reasoning process; Generate a first intermediate reasoning process according to the first intermediate step, the target position, the question-answering task, and the first intermediate question-answering task processing model, wherein an output result of the first intermediate reasoning process is different from the reference result; Taking the step in which the position of the first intermediate reasoning process is the same as the target position as the second intermediate step corresponding to the first intermediate step; Using first preset information as first annotation information of the first intermediate step, wherein the first preset information is used to indicate that a probability of obtaining the reference result according to the first intermediate step is greater than a third preset threshold; Using second preset information as first annotation information of the second intermediate step, wherein the second preset information is used to indicate that a probability of obtaining the reference result according to the second intermediate step is less than a fourth preset threshold; The first intermediate step and the second intermediate step are both used as the first reasoning step.

4. The method according to claim 3, characterized in that The step of generating a first intermediate reasoning process according to the first intermediate step, the target position, the question-answering task, and the first intermediate question-answering task processing model includes: In a case where the target position is the end position of the reasoning process, taking the step before the first intermediate step in the reasoning process as the second intermediate step; Using the first identifier of the second intermediate step as first prompt information, inputting the question-answering task and the first prompt information into the first intermediate question-answering task processing model, and obtaining the first intermediate reasoning process; When the target position is a non-starting position of the reasoning process, using the first identifier as second prompt information; The question-answering task and the second prompt information are input into the first intermediate question-answering task processing model to obtain the first intermediate reasoning process.

5. The method according to claim 3, characterized in that: After the step of making the first intermediate reasoning process position identical to the target position as a second intermediate step corresponding to the first intermediate step, the method further includes: determining a first similarity between the first intermediate step and the second intermediate step; When the first similarity is greater than or equal to a fifth preset threshold, the first intermediate step and the second intermediate step are deleted.

6. The method according to claim 1, characterized in that The determining, according to the third annotation information of the reasoning process in which the first reasoning step is located, the first annotation information of the first reasoning step includes: When it is determined according to the third annotation information that the output result of the reasoning process in which the first reasoning step is located is the same as the reference result, obtaining the output results of the second intermediate reasoning process and the second intermediate reasoning process according to a third preset number of the first reasoning steps, the question-answering task, and the first intermediate question-answering task processing model; Determining a third intermediate reasoning process among a fourth preset number of the second intermediate reasoning processes, wherein an output result of the third intermediate reasoning process is equal to the reference result; determining a first proportion of the third intermediate reasoning process to a fourth preset number of the second intermediate reasoning processes, and determining a first evaluation parameter of the first reasoning step according to the first proportion; A third intermediate step of determining, in a third preset number of the first reasoning steps, a minimum first evaluation parameter and a first evaluation parameter less than a sixth preset threshold value; Determine a second identifier of the third intermediate step, and determine a fourth intermediate step whose identifier is the same as the second identifier and whose first evaluation parameter is equal to a preset value from a third preset number of the first reasoning steps; Using first preset information as first annotation information of the fourth intermediate step, wherein the first preset information is used to indicate that a probability of obtaining the reference result according to the fourth intermediate step is greater than a third preset threshold; Using second preset information as first annotation information of the third intermediate step, wherein the second preset information is used to represent that the probability of obtaining the reference result according to the third intermediate step is less than a fourth preset threshold; The third intermediate step and the fourth intermediate step are both used as the first reasoning step.

7. The method according to claim 6, characterized in that After the fourth intermediate step of determining from a third preset number of the first reasoning steps that the identifier is the same as the second identifier and the first evaluation parameter is equal to a preset value, the method further includes: determining a second similarity between the third intermediate step and the fourth intermediate step; When the second similarity is greater than or equal to a seventh preset threshold, the third intermediate step and the fourth intermediate step are deleted.

8. The method according to claim 1, characterized in that The determining, according to the third annotation information of the reasoning process in which the first reasoning step is located, the first annotation information of the first reasoning step includes: When it is determined according to the third annotation information that the output result of the reasoning process in which the first reasoning step is located is different from the reference result, obtaining a fourth preset number of fourth intermediate reasoning processes and the output result of the fourth intermediate reasoning process according to a third preset number of the first reasoning steps, the question-answering task, and the first intermediate question-answering task processing model; Determining a fifth intermediate reasoning process from a fourth preset number of the fourth intermediate reasoning processes, wherein an output result of the fifth intermediate reasoning process is equal to the reference result; determining a second proportion of the fifth intermediate reasoning process in a fourth preset number of fourth intermediate reasoning processes, and determining a second evaluation parameter of the first reasoning step according to the second proportion; A fifth intermediate step of determining, in a third preset number of the first reasoning steps, that the second evaluation parameter is the smallest and the second evaluation parameter is less than an eighth preset threshold value; According to the fifth intermediate step, the question-answering task and the first intermediate question-answering task processing model, a sixth intermediate reasoning process is obtained, wherein an output result of the sixth intermediate reasoning process is equal to the reference result; Determine a third identifier of the fifth intermediate step, and take a step in the sixth intermediate reasoning process whose identifier is equal to the third identifier as the sixth intermediate step; Using first preset information as first annotation information of the sixth intermediate step, wherein the first preset information is used to indicate that a probability of obtaining the reference result according to the sixth intermediate step is greater than a third preset threshold; Using second preset information as first annotation information of the fifth intermediate step, wherein the second preset information is used to represent that the probability of obtaining the reference result according to the fifth intermediate step is less than a fourth preset threshold; The fifth intermediate step and the sixth intermediate step are both regarded as the first reasoning step.

9. The method according to claim 8, characterized in that The step of obtaining a sixth intermediate reasoning process according to the fifth intermediate step, the question-answering task, and the first intermediate question-answering task processing model includes: Using the third identifier of the fifth intermediate step as third prompt information; Inputting the question-answering task and the third prompt information into the first intermediate question-answering task processing model to obtain a fifth preset number of seventh intermediate reasoning processes; The sixth intermediate reasoning process is determined among a fifth preset number of the seventh intermediate reasoning processes.

10. The method according to claim 1, characterized in that Before adjusting the question-answering task processing model to be trained according to the initial training data set and the first model training algorithm, the method further includes: Obtain benchmark datasets; Extracting the question-answering task and a reference result of the question-answering task from the benchmark dataset; The initial training data set is created according to the question-answering task and the reference result.

11. The method according to claim 1, characterized in that: Before adjusting the question-answering task processing model to be trained according to the initial training data set and the first model training algorithm, the method further includes: Obtaining an initial question-answering task processing model and a sixth preset number of training samples, wherein the training samples include a task to be predicted and a target result of the task to be predicted; Inputting the task to be predicted into the initial question-answering task processing model to obtain a prediction result; According to the prediction result and the target result, the initial question-answering task processing model is adjusted to obtain a pre-trained model, and the pre-trained model is used as the question-answering task processing model to be trained.

12. A question-answering task processing model training device, characterized in that: The device comprises: A first adjustment module is used to adjust the question-answering task processing model to be trained according to the initial training data set and the first model training algorithm to obtain a first intermediate question-answering task processing model, wherein the initial training data set includes a first preset number of question-answering tasks; A data set generation module is used to obtain the first reasoning step of the question-answering task processed by the first intermediate question-answering task processing model and the first annotation information of the first reasoning step according to the initial training data set and the first intermediate question-answering task processing model, and generate a first target training data set according to the first reasoning step, the first annotation information and the initial training data set, wherein the first reasoning step is any step in the reasoning process of the question-answering task processed by the first intermediate question-answering task processing model; the data set generation module includes: an input unit, which is used to input the question-answering task in the initial training data set into the first intermediate question-answering task processing model, and obtain a second preset number of reasoning processes corresponding to the question-answering task and the output result of the reasoning process; a first generation unit, which is used to obtain the reference result of the question-answering task corresponding to the reasoning process in the initial training data set, and generate the third annotation information of the reasoning process according to the output result and the reference result, wherein the third annotation information is used to determine whether the output result is equal to the reference result; a splitting unit, which is used to split the reasoning process to obtain a third preset number of first reasoning steps; a third determination unit, which is used to determine the first annotation information of the first reasoning step according to the third annotation information of the reasoning process where the first reasoning step is located, wherein the first annotation information is used to determine the probability of obtaining the reference result according to the first reasoning step. ; The third determination unit is also used to use the first reasoning step as the first intermediate step when it is determined that the output result of the reasoning process in which the first reasoning step is located is the same as the reference result according to the third annotation information; determine the target position of the first intermediate step in the reasoning process; generate the first intermediate reasoning process according to the first intermediate step, the target position, the question-answering task and the first intermediate question-answering task processing model, wherein the output result of the first intermediate reasoning process is different from the reference result; use the step in which the position of the first intermediate reasoning process is the same as the target position as the second intermediate step corresponding to the first intermediate step; use the first intermediate step as the first reasoning step, and the probability of the first intermediate step obtaining the reference result is greater than the third preset threshold, and use the second intermediate step as the first reasoning step, and the probability of the second intermediate step obtaining the reference result is less than the fourth preset threshold; generate the first target training data set according to the first reasoning step, the first annotation information and the reference result of the question-answering task; the second generation unit is used to generate the first target training data set according to the first reasoning step, the first annotation information and the reference result of the question-answering task; the second adjustment module is used to adjust the question-answering task processing model to be trained according to the initial training data set and the second model training algorithm to obtain the second intermediate question-answering task processing model; The third adjustment module is used to adjust the second intermediate question and answer task processing model according to the first target training data set to obtain a target model.

13. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the question-answering task processing model training method according to any one of claims 1 to 11 by executing the computer instructions.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the question-answering task processing model training method described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Question and answer model training method and device, electronic equipment, storage medium and product

    CN118153659A

  • Method, apparatus, medium and program product for annotating data

    CN119272729A