A process reward model training method and system

By calculating the confidence score through the generator and completer and screening high-quality generated text, the noise impact existing in the existing technology is solved, the noise problem in the existing technology is optimized, and the training quality of the process reward model and the generated text accuracy of the large language model are improved.

CN120430424BActive Publication Date: 2025-10-03SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510935492.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-03
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The process reward model uses Monte Carlo estimation to generate synthetic data with high noise, which makes it impossible to accurately capture the key features of the text reasoning process, resulting in poor accuracy in text generation by large language models.

Method used

By obtaining a dataset, using the generator to generate multiple generated solutions, using the completer to calculate the confidence score, screening high-confidence positive and negative labeled samples, combining the tolerance distance hyperparameter and binary cross entropy loss, training the process reward model, and optimizing the labeled dataset and training process.

Benefits of technology

It improves the training quality of the process reward model, reduces the impact of noise, improves the accuracy and quality of text generated by large language models, and optimizes the efficiency of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430424B_ABST
    Figure CN120430424B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and in particular to a process reward model training method and system. The present invention calculates the confidence score of the question text in each sample; obtains the target correctness score of each reasoning step of the question text in the labeled sample based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample; trains the process reward model based on the binary cross entropy loss between the correctness prediction score of each reasoning step of the question text in the labeled sample and the target correctness score, and uses the trained process reward model as the target process reward model. The present invention improves the accuracy and reliability of process reward model training and enhances the accuracy of text generation by large-scale language models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a process reward model training method and system. Background Art

[0002] Process Reward Models (PRMs) are a mechanism dedicated to fine-grained evaluation of a model's reasoning process. By providing step-by-step reward signals, they guide the model to engage in slower, deeper reasoning when solving problems. In the application of Large Language Models (LLMs), PRMs can supervise the intermediate steps of text generation. For example, in complex tasks such as mathematical reasoning and logical argumentation, PRMs no longer focus solely on the correctness of the final answer. Instead, they evaluate each step of the reasoning process, the logical coherence of the generated text, and the sufficiency of the supporting evidence. This encourages LLMs to demonstrate a more rigorous chain of thought when generating text, laying the foundation for improving the accuracy of the resulting text.

[0003] Currently, training process reward models primarily relies on data annotation. Professionals meticulously evaluate each step of the model's reasoning and assign corresponding rewards. This approach generates high-quality data, providing precise guidance for process reward model training. However, manual annotation is labor-intensive and time-consuming, making it extremely costly. To address this issue, many studies have explored using Monte Carlo estimation to generate synthetic data. The basic principle is to simulate the model's reasoning process through multiple random samplings by a generator, thereby generating a large amount of training text data. Specifically, when given a text reasoning task, LLMs generate multiple possible reasoning paths and corresponding text content. Then, a completer estimates the correctness of each reasoning step, using this as the true label for each step. This is then used to train the process reward model. This approach leverages the model's inherent generative capabilities and statistical methods to generate data, reducing costs and increasing the data size.

[0004] However, while exploring synthetic data generation through Monte Carlo estimation has alleviated the cost pressure of manual annotation to some extent, this method still has certain drawbacks. Due to the high semantic complexity and diversity of text data, the Monte Carlo data generation process is susceptible to factors such as text semantic ambiguity and contextual changes when simulating reasoning steps. This leads to an inherently high noise ratio in the generated data. This noise manifests itself in problematic text, such as logical jumps in reasoning steps, semantic incoherence in generated text, and irrational reward distribution. This type of noise mainly comes from the inherent capability limitations of the completer used in the annotation process, which may cause the correctness of each reasoning step to be misestimated, resulting in the true label of each reasoning step itself deviating from the actual label. The process reward model trained based on the Monte Carlo estimation method is essentially evaluating the completer's potential ability to deduce the correct answer from the current step, rather than accurately learning the true correctness of the current step. Therefore, using these noisy synthetic data to train the process reward model makes it impossible for the model to accurately judge the quality of reasoning steps when facing real text reasoning tasks. For example, it is unable to identify key features such as "whether the theorem is applied correctly" and "whether the argument fully supports the conclusion" in mathematical text reasoning. As a result, when guiding LLMs to generate text, it is unable to identify logical loopholes in the text generated by LLMs, nor can it supervise the semantic coherence of the text, making it difficult to encourage LLMs to strengthen correct text reasoning patterns and correct incorrect patterns. Ultimately, the text generated by LLMs has poor accuracy and cannot meet the needs of actual applications. Summary of the Invention

[0005] To this end, the technical problem to be solved by the present invention is to overcome the defect that when the process reward model uses Monte Carlo estimation to generate synthetic data for training, the process reward model cannot accurately capture the key features in the text reasoning process due to the high noise in the Monte Carlo estimation synthetic data, which ultimately leads to the poor accuracy of text generation by large language models.

[0006] To solve the above technical problems, the present invention provides a process reward model training method, comprising the following steps:

[0007] Obtain a dataset where each sample includes a question text and its corresponding true answer; use a generator to generate multiple generated solutions for the question text in each sample, including: a predicted answer and multiple reasoning steps;

[0008] Use the completer to generate multiple completion solutions for the question text in each sample, and calculate the confidence score of the question text in each sample based on the true answer in each sample and the predicted answers in the multiple completion solutions;

[0009] For each generated solution of the question text in each sample, the correctness score of each reasoning step in the current generated solution is generated through the completer. The current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score are used as a labeled sample of the labeled dataset;

[0010] Based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample, the target correctness score of each reasoning step of the question text in the labeled sample is obtained;

[0011] The problem text in the labeled sample in the labeled dataset is input into the reward model to generate the solution, and the correctness prediction score of each reasoning step of the problem text in the labeled sample is output;

[0012] Based on the binary cross entropy loss between the correctness prediction scores of each reasoning step of the question text in the labeled samples and the target correctness scores, the process reward model is trained and the trained process reward model is used as the target process reward model.

[0013] Preferably, the labeled data set includes a positive labeled sample set and a negative labeled sample set;

[0014] For each generated solution of the question text in each sample, determine whether its predicted answer is consistent with the true answer corresponding to the question text. If consistent, the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score are taken as the positively labeled sample;

[0015] If they are inconsistent, determine whether the confidence score of the question text in the sample corresponding to the current generated solution is greater than the set threshold. If it is greater, take the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score as negative annotation samples.

[0016] Preferably, the correctness score of each reasoning step in the solution of the positively labeled sample is set to 1;

[0017] The correctness scores of each reasoning step in the solution of negatively labeled samples are obtained through the completer.

[0018] Preferably, the generator and the completer are set to the same model;

[0019] Based on the true answer in each sample and the predicted answers in multiple generated solutions of the question text, the confidence score of the question text in each sample in the generator is calculated as the confidence score of the question text in each sample.

[0020] Preferably, the confidence score of the question text in each sample in the generator is calculated based on the true answer in each sample and the predicted answer in multiple generated solutions of the question text, and the formula is:

[0021] ,

[0022] in, For the Question text in the sample In the generator Medium confidence score, For the The question text in the sample, is the confidence score in the generator, the number of solutions generated for the problem text in the sample, To generate the solution index, For the The first The predicted answer among the generated solutions, For the The true answer of the sample, is the evaluation function, when When consistent, ,when When inconsistent, , is the probability distribution computed by the generator, For the generator, Indicates the Multiple generated solutions to the problem text in samples, express Is a generator In the given problem The probability distribution when The sample obtained from is the sample index.

[0023] Preferably, the target correctness score of each reasoning step of the question text in the labeled sample is obtained based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample, and the formula is:

[0024] ;

[0025] in, The first The target correctness score of each reasoning step, To get the minimum value, To mark the question text in the sample, is the confidence score, To mark the problem text in the sample The confidence score of The first The correctness score of the reasoning steps, is the indicator function, when hour, ,when hour, , is the tolerance distance hyperparameter, is the index of the reasoning step, The first incorrect reasoning step index, that is, satisfying The minimum value, is the probability distribution computed by the completer, The first The true labels for the inference steps, Indicates the first The true label of the reasoning step is correct, Indicates that given the question text and Under the condition of The correctness score of the reasoning steps, Represents the first reasoning step to the first step of the question text in the labeled sample A sequence of reasoning steps.

[0026] Preferably, the binary cross entropy loss between the correctness prediction score and the target correctness score of each reasoning step of the question text in the labeled sample is:

[0027] ,

[0028] in, is the binary cross entropy loss, are the learnable parameters of the process reward model, The first The target correctness score of each reasoning step, is the mathematical expectation, express Obey the labeled dataset The distribution of To label the dataset, Represents the first reasoning step to the first step of the question text in the labeled sample A sequence of inference steps, The first The true labels for the inference steps, For the process reward model, given and When, The correctness prediction score of the reasoning steps, To mark the question text in the sample, Index to the reasoning step.

[0029] Preferably, generating the correctness score of each reasoning step in the current generated solution by the completer includes:

[0030] Each reasoning step in the current generated solution and the corresponding question text of the current generated solution are substituted into the correctness score prompt template and input into the completer to generate the correctness score of each reasoning step in the current generated solution.

[0031] Preferably, the step of generating multiple generated solutions for the question text in each sample by using a generator comprises:

[0032] After substituting the question text in each sample into the generated solution prompt template, it is input into the generator, and multiple generated solutions for the question text in each sample are generated through multiple sampling.

[0033] The present invention also provides a process reward model training system, comprising:

[0034] The data acquisition module is used to obtain a dataset, where each sample includes a question text and its corresponding true answer; the generator is used to generate multiple generated solutions for the question text in each sample, including: a predicted answer and multiple reasoning steps;

[0035] A confidence calculation module is used to generate multiple completion solutions for the question text in each sample using the completer, and calculate the confidence score of the question text in each sample based on the true answer in each sample and the predicted answers in the multiple completion solutions;

[0036] The sample construction module is used to generate the correctness score of each reasoning step in the generated solution of each question text in each sample through the completer, and use the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score as a labeled sample of the labeled dataset;

[0037] A correction module is used to obtain a target correctness score of each reasoning step of the question text in the labeled sample based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample;

[0038] The prediction module is used to input the problem text in the labeled sample in the labeled dataset and the generated solution into the process reward model, and output the correctness prediction score of each reasoning step of the problem text in the labeled sample;

[0039] The training module is used to train the process reward model based on the binary cross entropy loss between the correctness prediction scores of each reasoning step of the question text in the labeled sample and the target correctness score, and use the trained process reward model as the target process reward model.

[0040] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0041] The present invention discloses a process reward model training method and system. The present invention introduces a completer to generate multiple completion solutions. Based on the true answer in each sample and the predicted answer in the generated multiple completion solutions, the confidence score of the question text in each sample is calculated to quantify the reliability of the model output. At the same time, based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample, the true correctness label of the reasoning step is redefined. When the distance between the reasoning step and the first incorrect reasoning step is within the range specified by the tolerance distance hyperparameter, the target correctness score is determined by combining the correctness score of the current reasoning step and the confidence of the question text in the sample. If the confidence of the question text in the sample is high, even if there is a certain question text in the reasoning step, a relatively reasonable target correctness score will be given within the tolerance range. This is more in line with the situation in actual reasoning where occasional local deviations occur but the overall direction is correct. When an inference step exceeds the tolerance distance, the target correctness score will be 1 only if the step is completely correct. This can prevent the impact of the incorrect inference step from continuing to expand. By combining the confidence and the coherence of the inference steps, the calculation method reduces the deviation of the completer and improves the overall data quality. It also improves the training quality of the process reward model, enabling large-scale language models to obtain more reasonable feedback and optimization directions when generating text, effectively reducing logical errors and improving the accuracy of answers, thereby improving the precision and quality of the generated text.

[0042] In addition, traditional data labeling often faces the contradiction between efficiency and cost. It is time-consuming and labor-intensive to comprehensively and meticulously label all samples, and the labeling cost is too high. Moreover, the noise in the high-confidence part of the positively labeled samples is relatively small, and the cost-effectiveness of deep inspection is low. Therefore, the present invention judges whether the predicted answer of each solution to the question text in each sample is consistent with the true answer corresponding to the question text. If consistent, the correctness score of each reasoning step in the current solution is set to 1, and the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score are used as the positively labeled samples of the labeled dataset without further labeling; only the negatively labeled samples with a confidence score greater than the set threshold are gradually labeled for correctness, which effectively avoids the invalid processing of low-quality and low-value negatively labeled samples, because the low-confidence negatively labeled samples themselves have low credibility and limited labeling significance. The high-confidence negatively labeled samples screened out can better reflect the key error information in the model reasoning process. The labeled dataset constructed by the above samples not only ensures the accuracy of the labeled data, but also greatly optimizes the use of computing resources, significantly improves the data synthesis efficiency, and improves the training quality of the process reward model, thereby greatly improving the accuracy of LLMs generated text. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0044] Figure 1 It is a flow chart of a process reward model training method of the present invention.

[0045] Figure 2 This is a comparison of the average scores of different process reward models on different evaluation datasets.

[0046] Figure 3 This is a comparison of various indicators of different process reward models on different evaluation data sets.

[0047] Figure 4 The figure shows the comparison of F1 scores of the process reward model during training using different training methods.

[0048] FIG5 is a comparison of the F1 scores of the process reward model trained with different tolerance distance hyperparameters during the training process.

[0049] Figure 6 Schematic diagram of the ablation experiment results of the present invention. DETAILED DESCRIPTION

[0050] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0051] Reference Figure 1 As shown, this embodiment 1 provides a process reward model training method (Self-Denoising Monte Carlo Annotation, SCAN), including the following steps:

[0052] Step S1: Get the dataset , each sample in the dataset includes question text and its corresponding real answer ; Using the generator Generate the question text in each sample generated solutions , the solution includes: predict the answer and Reasoning steps ; Among them, the question text of each sample in the dataset and its corresponding true answer are both text data, is the total number of samples in the dataset, For the The question text in the sample, For the Question text in the sample The true answer is the number of solutions generated for the problem text in the sample, For the The question text in the sample Generate solutions, , To generate the solution index, For the The first The predicted answer among the generated solutions, For the The question text generated in the sample The generated solution Steps of reasoning, , is the total number of reasoning steps in the solution, is the sample index.

[0053] In this embodiment, specifically, the generating of multiple generated solutions for the question text in each sample by using a generator includes:

[0054] After substituting the question text in each sample into the generated solution prompt template, it is input into the generator, and multiple generated solutions for the question text in each sample are generated through multiple sampling.

[0055] Step S2: Use the completer to generate multiple completion solutions for the question text in each sample, and calculate the confidence score of the question text in each sample based on the true answer in each sample and the predicted answers in the generated multiple completion solutions. ;in, For the Question text in the sample In the completer The confidence score in this embodiment is Question text in the sample The confidence score in the completer, as the confidence score of the question text in each example;

[0056] In this embodiment, specifically, generating the correctness score of each reasoning step in the current generated solution by the completer includes:

[0057] Each reasoning step in the current generated solution and the corresponding question text of the current generated solution are substituted into the correctness score prompt template and input into the completer to generate the correctness score of each reasoning step in the current generated solution.

[0058] In this embodiment, preferably, the generator With the completer Set to the same model, i.e. ;

[0059] Based on the true answer in each sample and the predicted answers in multiple generated solutions of the question text, the confidence score of the question text in each sample in the generator is calculated as the confidence score of the question text in each sample.

[0060] In this embodiment, specifically, the confidence score of the question text in each sample in the generator is calculated based on the true answer in each sample and the predicted answer in multiple generated solutions of the question text, and the formula is:

[0061] ,

[0062] in, For the Question text in the sample In the generator Medium confidence score, For the The question text in the sample, is the confidence score in the generator, the number of solutions generated for the problem text in the sample, To generate the solution index, For the The first The predicted answer among the generated solutions, For the The true answer of the sample, is the evaluation function, when When consistent, ,when When inconsistent, , is the probability distribution computed by the generator, For the generator, Indicates the Multiple generated solutions to the problem text in samples, express Is a generator In the given problem The probability distribution when The sample obtained from is the sample index.

[0063] Similarly, based on the true answer in each sample and the predicted answers in the generated multiple completion solutions, the confidence score of the question text in each sample in the completer is calculated using the formula:

[0064] ,

[0065] in, For the Question text in the sample In the completer The confidence score in , is the number of complete solutions to the problem text in the sample, To complete the solution index, For the The first The predicted answer in the completed solution.

[0066] This embodiment sets the generator and the completer to the same model, and directly calculates the confidence score based on the multiple generated solutions output by the generator, which has significant advantages. If the generator and the completer are deployed separately, multiple generated solutions and completed solutions need to be independently sampled for each question text, resulting in double consumption of computing resources. However, this embodiment avoids the process of repeatedly generating completed solutions through model reuse, and directly uses the generated solution set sampled by the generator to calculate the confidence score without additional calculation. This optimization not only reduces the computational overhead of model inference, but also saves the memory resources required to store intermediate results. Especially when processing large-scale data sets, it can significantly reduce hardware costs and training time. In addition, the solutions and completed solutions generated by the same model are consistent, avoiding evaluation bias caused by model differences, making the confidence score more accurately reflect the reliability of the model itself, thereby improving the efficiency of subsequent labeling and training, and ultimately achieving efficient use of computing resources while ensuring the process reward model training effect.

[0067] Step S3: For each generated solution of the question text in each sample, the completer generates the correctness score of each reasoning step in the current generated solution, and the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score are used as a labeled sample of the labeled dataset;

[0068] In this embodiment, preferably, the labeled data set includes a positive labeled sample set and a negative labeled sample set;

[0069] For each generated solution of the question text in each sample, determine whether its predicted answer is consistent with the true answer corresponding to the question text. If they are consistent, the correctness score of each reasoning step in the current generated solution is set to 1, and the current generated solution is , the correctness score of each reasoning step in generating the solution , question text and its confidence score , as the positive labeled samples of the labeled dataset;

[0070] If they are inconsistent, determine whether the confidence score of the question text in the sample corresponding to the current generated solution is greater than the set threshold If it is greater than, the correctness score of each reasoning step in the current generated solution is obtained through the completer, and the current generated solution is , the correctness score of each reasoning step in generating the solution , question text and its confidence score , as the negative labeled samples of the labeled dataset; based on the positive labeled samples and negative labeled samples of the labeled dataset, construct the labeled dataset ;in, is the correctness score of each reasoning step, For the The correctness score of each reasoning step;

[0071] When constructing the annotated dataset, the present invention adopts a selective annotation strategy to optimize resource utilization efficiency. Specifically, for the multiple generated solutions output by the generator, only the negatively annotated samples whose predicted answers are inconsistent with the true answers and whose confidence scores are higher than the set threshold are selected for further step-by-step annotation, while the positively annotated samples whose predicted answers have been determined to be consistent with the true answers are no longer subject to additional processing. Although in theory there may be noise in the positively annotated samples, such as the case where there are logical loopholes in the reasoning process but the answer is coincidentally correct, the computational cost of filtering such samples through detailed step-by-step correctness checks is extremely high. For example, to fully verify an answer that contains 10 reasoning steps, if each step requires 8 deductions, then a single sample will consume 80 computing resources.

[0072] The present invention observes that the noise content in high-confidence positively labeled samples is extremely low, so it directly uses them as training positive examples without additional labeling. By applying Monte Carlo estimation to generate labeled data only for qualified negatively labeled samples, this solution ensures that all labeled samples are included in the final training set, maximizing sample utilization. This strategy avoids redundant computation of low-noise positively labeled samples while focusing on deep mining of high-value negatively labeled samples, significantly reducing computational overhead while ensuring training data quality.

[0073] Traditional data labeling often faces the contradiction between efficiency and cost. Comprehensive and detailed labeling of all samples is time-consuming and labor-intensive, and the labeling cost is too high. Moreover, the noise in the high-confidence part of the positively labeled samples is relatively small, and the cost-effectiveness of deep inspection is low. Therefore, the present invention judges whether the predicted answer of each generated solution of the question text in each sample is consistent with the true answer. If consistent, the correctness score of each reasoning step in the current generated solution is set to 1, and the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score are used as the positively labeled samples of the labeled data set without further labeling; only the negatively labeled samples with a confidence score greater than the set threshold are gradually labeled for correctness, which effectively avoids the invalid processing of low-quality and low-value negatively labeled samples, because the low-confidence negatively labeled samples themselves have low credibility and limited labeling significance. The screened high-confidence negatively labeled samples can better reflect the key error information in the model reasoning process. The labeled data set constructed by the above samples not only ensures the accuracy of the labeled data, but also greatly optimizes the use of computing resources, significantly improves the data synthesis efficiency, and improves the training quality of the process reward model, thereby greatly improving the accuracy of LLMs generated text.

[0074] Step S4: based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the annotated sample, obtain the target correctness score of each reasoning step of the question text in the annotated sample;

[0075] Step S5: The problem text in the labeled sample in the labeled dataset and the generated solution are input into the process reward model, and the correctness prediction score of each reasoning step of the problem text in the labeled sample is output;

[0076] In this embodiment, preferably, the target correctness score of each reasoning step of the question text in the labeled sample is obtained based on the correctness score, confidence, and tolerance distance hyperparameters of each reasoning step of the question text in the labeled sample, and the formula is:

[0077] ;

[0078] in, The first The target correctness score of each reasoning step, To get the minimum value, To mark the question text in the sample, is the confidence score, To mark the problem text in the sample The confidence score of The first The correctness score of the reasoning steps, is the indicator function, when hour, ,when hour, , is the tolerance distance hyperparameter, is the index of the reasoning step, The first incorrect reasoning step index, that is, satisfying The minimum value, is the probability distribution computed by the completer, The first The true labels for the inference steps, Indicates the first The true label of the reasoning step is correct, Indicates that given the question text and Under the condition of The correctness score of the reasoning steps, Represents the first reasoning step to the first step of the question text in the labeled sample A sequence of reasoning steps.

[0079] In the process reward model training, the completer tends to overestimate the correctness of the current step due to its strong self-correction ability. This overestimation will gradually appear as errors accumulate during the reasoning process, resulting in the correctness score predicted by the completer being higher than the actual correctness score, and the probability of similar deviations appearing near the error location increases significantly. In order to achieve robust learning from noisy labels, this paper proposes a noise-resistant labeling strategy, which labels the steps before the erroneous reasoning step, i.e. Steps, apply soft label processing, within the tolerance distance The strategy is processed within a certain range, and experiments have verified that it can effectively avoid the risk of overfitting in the process reward model (PRM) and improve the learning efficiency of noise labels.

[0080] After denoising, the accuracy of the annotation labels is still significantly limited by the capabilities of the completer model itself. From a probabilistic perspective, the actual reasoning step correctness score , Indicates that the question text and preceding reasoning steps Get the first The true labels for the inference steps For the correct conclusion The probability of the completer model Output accuracy score , which can only be done through the completer To predict, the actual reasoning step correctness score In order to reduce the model dependency bias, the present invention introduces a correction factor , using the confidence score To approximate the correction factor , and reduce the deviation caused by this model, which is based on: different performance models annotate the same sample, a strong model has , a weak model has , stronger models will produce higher accuracy scores, and the present invention currently aims to make the corrected scores consistent across different models, regardless of their inherent strength. Therefore, the present invention uses confidence scores to adjust the estimated accuracy scores as follows:

[0081] ,

[0082] in, To improve the question text by using a more capable completer model (e.g., a model with larger parameters and more comprehensive training) The generated confidence score, To solve the problem text by a less capable completer model (e.g. a model with a slightly smaller parameter size) The generated confidence score, Based on the question text and preceding reasoning steps Get the first The true labels for the inference steps For the correct conclusion probability.

[0083] This method effectively normalizes the interference of different model capabilities on the labeling results, making the labeled data closer to the actual correctness, improving data reliability and unbiasedness. It is particularly suitable for integrating the labeling results of multiple completer models, ensuring the consistency and effectiveness of cross-model labeling, and providing high-quality data support for process reward model training.

[0084] Traditional Monte Carlo estimation methods use the correctness score of each inference step generated by the completer as the true correctness label for the inference step. However, the present invention redefines the true correctness label for an inference step based on the correctness score, confidence score, and tolerance distance hyperparameter of each inference step in the labeled question text. When the distance between an inference step and the first incorrect inference step is within the range specified by the tolerance distance hyperparameter, the target correctness score is determined by combining the correctness score of the current inference step and the confidence of the question text in the sample. If the confidence of the question text in the sample is high, even if the inference step contains some question text, a relatively reasonable target correctness score within the tolerance range will be given. This is more consistent with the situation in actual reasoning where occasional local deviations may occur but the overall direction is correct. When an inference step exceeds the tolerance distance, the target correctness score is only 1 if the step is completely correct. This prevents the impact of the incorrect inference step from continuing to expand. By combining confidence and the coherence of inference steps, this calculation method reduces the completer's bias, improves overall data quality, and enhances the training quality of the process reward model.

[0085] Step S6: Based on the binary cross entropy loss between the correctness prediction scores of each reasoning step of the question text in the labeled sample and the target correctness score, the process reward model is trained, and the trained process reward model is used as the target process reward model.

[0086] In this embodiment, preferably, the binary cross entropy loss between the correctness prediction score and the target correctness score of each reasoning step of the question text in the labeled sample is:

[0087] ,

[0088] in, is the binary cross entropy loss, are the learnable parameters of the process reward model, The first The target correctness score of each reasoning step, is the mathematical expectation, express Obey the labeled dataset The distribution of To label the dataset, Represents the first reasoning step to the first step of the question text in the labeled sample A sequence of inference steps, The first The true labels for the inference steps, For the process reward model, given and When, The correctness prediction score of the reasoning steps, To mark the question text in the sample, Index to the reasoning step.

[0089] To address the performance bottleneck caused by data noise and annotation bias in the training of traditional process reward models, this paper proposes a systematic optimization solution. First, by constructing a basic dataset containing question text-true answer pairs, a generator is used to generate multiple candidate solutions containing predicted answers and reasoning steps. The same model is then reused as a completer to generate completed solutions and calculate the confidence score of the question text based on the true answer and predicted answer, thereby quantifying the reliability of the model output. During the annotation dataset construction stage, a hierarchical screening strategy is adopted: for solutions where the predicted answer is consistent with the true answer, they are directly marked as positively labeled samples and assigned a reasoning step correctness score of 1. For inconsistent solutions, only when the confidence score is higher than the set threshold, the completer evaluates the reasoning step correctness score and incorporates negatively labeled samples, effectively filtering out low-quality data.

[0090] On this basis, the labeled samples are input into the process reward model to obtain the predicted correctness score of the reasoning step. Based on the correctness score, confidence level, and tolerance distance hyperparameters, a dynamic formula is used to correct the target correctness score of the reasoning step, making the training objective more consistent with the actual reasoning scenario. Finally, the optimization objective is constructed using the binary cross-entropy loss between the predicted score and the target score, and the process reward model is iteratively trained. This method achieves data quality screening through confidence metrics. Combined with dynamic target correction and selective labeling strategies, it effectively reduces data noise and labeling bias, improves the model's evaluation accuracy of the reasoning step, and provides a more reliable reward feedback mechanism for large-scale language models to generate high-quality text.

[0091] This paper evaluates the effectiveness of the process reward model from two key dimensions: Best-of-N (BoN) evaluation and step-level error detection.

[0092] In the Best-of-N evaluation, the process reward model acts as a verifier to select the best answer from multiple candidate answers generated by the strategy model. Specifically, the process reward model first assigns a correctness score to each step in the answer, and then aggregates the correctness scores of these steps into an overall reward score for the entire answer, where the aggregation method is to take the lowest correctness score among all steps, and finally select the reasoning step with the highest correctness score as the final answer. The evaluation dataset contains multiple difficulty levels, such as GSM8K (primary school level), MATH (competition level), College Math (university level), and OlympiadBench (Olympiad level). The present invention uses Qwen2.5-Math-7B-Instruct and Llama3.1-8B-Instruct as strategy models, and the number of solutions generated by the problem text in the sample Set to 8. At the same time, we use majority voting as the benchmark and pass@8 as the upper bound of performance. Pass@8 refers to the number of solutions generated by the problem text in the sample. When set to 8, there is at least one correct solution.

[0093] We also delve deeper into the process reward model's ability to accurately detect the location of errors within answers. We use ProcessBench as an evaluation benchmark, which measures the model's ability to identify the first error within a given answer. This evaluation focuses on the model's performance in identifying fully correct examples and accurately identifying errors within incorrect answers. The final F1 score is the harmonic mean of the accuracy rates for correct and incorrect examples.

[0094] The main comparison objects of this invention are 7B-scale process reward models, including Math-Shepherd, RLHFlow-PRM, Skywork-PRM, Math-PSA, and EurusPRM. Models trained using strong supervision (such as Qwen2.5-Math-PRM-7B and UniversalPRM) are also included as reference points. However, these strongly supervised models are not directly compared because they rely on large-scale external criticism models for guidance. It should be noted that these criticism models are relatively powerful in themselves, and their supervision (usually achieved through knowledge distillation) has a significant impact on the performance of the final process reward model. The problem text solved by this invention is different from these methods. It focuses on the denoising problem text of Monte Carlo estimation itself. For Best-of-N evaluation, this invention directly tests the public checkpoints of the comparison models. For the results on ProcessBench, the results are directly derived from the values ​​reported in their research.

[0095] like Figure 2 As shown, Figure 2 Schematic comparison of the average scores of different process reward models on different evaluation datasets. Figure 2 Used to present the performance of different process reward models on multiple difficulty mathematical reasoning datasets (GSM8K, MATH, College Math, Olympiad Bench).

[0096] The Qwen2.5-Math-7B-Scan-Base model is based on the Qwen2.5-Math-7B base model and trained using a process reward model training method proposed in this paper using 101K synthetic samples from the Scan-Base dataset generated by the 1.5B model. It outperforms process reward models trained on a larger synthetic dataset and approaches the performance of process reward models trained on manually annotated data (PRM800K), demonstrating that even with the 1.5B model, it is possible to synthesize data of comparable quality to manually annotated data. Notably, the Qwen2.5-Math-7B-Scan-Pro model, further trained on additional synthetic data generated by the 7B model based on the Qwen2.5-Math-7B-Scan-Base model, surpasses PRM800K in performance.

[0097] Both Qwen2.5-Math-7B-Scan-Base and Qwen2.5-Math-7B-Scan-Pro surpass all other process reward models, including those trained on PRM800K. Notably, their error detection capabilities even surpass those of the 70B-scale critic model Llama-3.3-70B-Instruct.

[0098] like Figure 3 As shown, Figure 3 The following is a comparison of various indicators of different process reward models on different evaluation data sets. Figure 3 It can be seen that through a process reward model training method of the present invention, Qwen2.5-Math-7B-Ins can generate process data for self-training, thereby greatly improving its own error detection capability from 19.9 to 59.1. This verifies the effectiveness of the training framework based on synthetic process data and denoising strategy proposed in the present invention in enhancing the model's error recognition accuracy, and provides a feasible path for the process reward model to break through the data dependence bottleneck and autonomously improve the reasoning quality, highlighting the significant value of this method in enabling model performance evolution.

[0099] The present invention also carried out ablation research to further verify the effectiveness of each component, and the relevant results are as follows Figure 4 As shown, Figure 4 The figure shows a comparison of the F1 scores of process reward models using different training methods during the training process, showing the performance changes of the models during the training process. The horizontal axis represents the amount of training data, and the vertical axis represents the F1 score of the process reward model. Baseline represents the baseline model that does not use the process reward model training method of the present invention. It can be seen that when no denoising strategy is adopted, the model quickly overfits to the noise samples. Scan-Base represents the model that uses the process reward model training method of the present invention. Scan-Pro represents the model that incorporates more data sources based on SCAN-Base. It can be seen that the denoising strategy of the present invention can enable the model to steadily improve performance, avoid overfitting to noise samples, and thus improve model performance. By comparing the results of Scan-Base and Scan-Pro, it can be seen that the integration of additional data sources can enhance data diversity, thereby optimizing the upper bound of model performance.

[0100] like Figure 5 As shown in Figure 5, the F1 score comparison of the process reward model trained with different tolerance distance hyperparameters during the training process is shown in Figure 5. Figure 5 It can be seen that the tolerance distance hyperparameter The choice of is crucial, as too small or too large a tolerance distance may introduce noise. When , hard labels will be formed and serious noise will be generated, which is reflected in the expansion curve; when When , it will become a soft label, which will also introduce significant noise and hinder the improvement of model performance. The experimental results of the value found = 2 can achieve a good balance and reduce overfitting during training.

[0101] like Figure 6 As shown, Figure 6 The figure shows the results of the ablation experiment of the present invention. This demonstrates the effectiveness of the various components of the present invention. +Conf-Based Filtering introduces a confidence filtering strategy based on the baseline, specifically filtering samples with confidence scores greater than a set threshold. Both Best-of-N Accuracy and ProcessBench Overall F1 metrics improve, demonstrating the effectiveness of filtering noisy data. +Tolerance Labeling introduces a noise-tolerant labeling strategy based on the baseline. This strategy uses the correctness scores, confidence scores, and tolerance distance hyperparameters of each reasoning step in the labeled question text to obtain the target correctness scores for each reasoning step. Both Best-of-N Accuracy and ProcessBench Overall F1 metrics improve, demonstrating that soft labels can convey more fine-grained knowledge. +Conf-wise Reweighting (SCAN-Pro) uses a process reward model proposed in the present invention for training based on the baseline, simultaneously introducing both the confidence filtering strategy and the noise-tolerant labeling strategy. At this point, both metrics reach their highest values. In the Best-of-N and ProcessBench evaluations, model performance showed a trend of continuous improvement. Both the confidence filtering strategy and the noise-resistant labeling strategy contributed to the performance improvement. The noise-resistant labeling strategy enhanced the robustness of the model in processing noisy samples, while the confidence filtering strategy helped eliminate the deviation in probability estimation of different models when labeling samples.

[0102] This paper optimizes Monte Carlo estimation from the perspectives of noise analysis and robust learning. The proposed process reward model training method requires no manual labeling or strong model supervision, significantly outperforming existing techniques in terms of training efficiency and data utilization. Table 1 compares the performance and inference speed of different models.

[0103] Table 1

[0104]

[0105] Table 1 details the performance and inference speed differences between the process reward model trained using the method of the present invention and other large-scale evaluation models. This inference speed test was conducted in an environment with four A100-40G GPUs. The experimental results demonstrate that the process reward model trained using the method of the present invention exhibits significant advantages in inference speed compared to traditional evaluation models. In particular, the efficiency bottlenecks of long thought chain evaluation models, such as DeepSeek-R1-Distill-Qwen-7B, are particularly prominent, further highlighting the technical superiority of the present invention in improving model inference efficiency.

[0106] This second embodiment provides a process reward model training system, including:

[0107] The data acquisition module is used to obtain a dataset, where each sample includes a question text and its corresponding true answer; the generator is used to generate multiple generated solutions for the question text in each sample, including: a predicted answer and multiple reasoning steps;

[0108] A confidence calculation module is used to generate multiple completion solutions for the question text in each sample using the completer, and calculate the confidence score of the question text in each sample based on the true answer in each sample and the predicted answers in the multiple completion solutions;

[0109] The sample construction module is used to generate the correctness score of each reasoning step in the generated solution of each question text in each sample through the completer, and use the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score as a labeled sample of the labeled dataset;

[0110] A correction module is used to obtain a target correctness score of each reasoning step of the question text in the labeled sample based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample;

[0111] The prediction module is used to input the problem text in the labeled sample in the labeled dataset and the generated solution into the process reward model, and output the correctness prediction score of each reasoning step of the problem text in the labeled sample;

[0112] The training module is used to train the process reward model based on the binary cross entropy loss between the correctness prediction scores of each reasoning step of the question text in the labeled sample and the target correctness score, and use the trained process reward model as the target process reward model.

[0113] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0114] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0115] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0117] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A process reward model training method, characterized in that: include: Get a dataset where each sample includes a question text and its corresponding true answer; Use the generator to generate multiple generated solutions for the question text in each sample, including: predicted answers and multiple reasoning steps; Use the completer to generate multiple completion solutions for the question text in each sample, and calculate the confidence score of the question text in each sample based on the true answer in each sample and the predicted answers in the multiple completion solutions; For each generated solution of the question text in each sample, the correctness score of each reasoning step in the current generated solution is generated through the completer. The current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score are used as a labeled sample of the labeled dataset; Based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample, the target correctness score of each reasoning step of the question text in the labeled sample is obtained; The problem text in the labeled sample in the labeled dataset is input into the reward model to generate the solution, and the correctness prediction score of each reasoning step of the problem text in the labeled sample is output; Based on the binary cross entropy loss between the correctness prediction scores of each reasoning step of the question text in the labeled samples and the target correctness scores, the process reward model is trained and the trained process reward model is used as the target process reward model.

2. A process reward model training method according to claim 1, characterized in that: The labeled dataset includes a set of positive labeled samples and a set of negative labeled samples; For each generated solution of the question text in each sample, determine whether its predicted answer is consistent with the true answer corresponding to the question text. If consistent, the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score are taken as the positively labeled sample; If they are inconsistent, determine whether the confidence score of the question text in the sample corresponding to the current generated solution is greater than the set threshold. If it is greater, take the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score as negative annotation samples.

3. A process reward model training method according to claim 2, characterized in that: Set the correctness score of each reasoning step in the solution of the positively labeled sample to 1; The correctness scores of each reasoning step in the solution of negatively labeled samples are obtained through the completer.

4. A process reward model training method according to claim 1, characterized in that: Set the generator and completer to the same model; Based on the true answer in each sample and the predicted answers in multiple generated solutions of the question text, the confidence score of the question text in each sample in the generator is calculated as the confidence score of the question text in each sample.

5. A process reward model training method according to claim 4, characterized in that: Based on the true answer in each sample and the predicted answer in multiple generated solutions of the question text, the confidence score of the question text in each sample in the generator is calculated, and the formula is: , in, For the Question text in the sample In the generator Medium confidence score, For the The question text in the sample, is the confidence score in the generator, the number of solutions generated for the problem text in the sample, To generate the solution index, For the The first The predicted answer among the generated solutions, For the The true answer of the sample, is the evaluation function, when When consistent, ,when When inconsistent, , is the probability distribution computed by the generator, For the generator, Indicates the Multiple generated solutions to the problem text in samples, express Is a generator In the given problem The probability distribution when The sample obtained from is the sample index.

6. A process reward model training method according to claim 1, characterized in that: The target correctness score of each reasoning step of the question text in the labeled sample is obtained based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample. The formula is: ; in, The first The target correctness score of each reasoning step, To get the minimum value, To mark the question text in the sample, is the confidence score, To mark the problem text in the sample The confidence score of The first The correctness score of the reasoning steps, is the indicator function, when hour, ,when hour, , is the tolerance distance hyperparameter, is the index of the reasoning step, The first incorrect reasoning step index, that is, satisfying The minimum value, is the probability distribution computed by the completer, The first The true labels for the inference steps, Indicates the first The true label of the reasoning step is correct, Indicates that given the question text and Under the condition of The correctness score of the reasoning steps, Represents the first reasoning step to the first step of the question text in the labeled sample A sequence of reasoning steps.

7. A process reward model training method according to claim 1, characterized in that: The binary cross entropy loss between the correctness prediction score and the target correctness score of each reasoning step of the question text in the labeled sample is: , in, is the binary cross entropy loss, are the learnable parameters of the process reward model, The first The target correctness score of each reasoning step, is the mathematical expectation, express Obey the labeled dataset The distribution of To label the dataset, Represents the first reasoning step to the first step of the question text in the labeled sample A sequence of inference steps, The first The true labels for the inference steps, For the process reward model, given and When, The correctness prediction score of the reasoning steps, To mark the question text in the sample, Index to the reasoning step.

8. A process reward model training method according to claim 1, characterized in that: The correctness score of each reasoning step in the current generated solution is generated by the completer, including: Each reasoning step in the current generated solution and the corresponding question text of the current generated solution are substituted into the correctness score prompt template and input into the completer to generate the correctness score of each reasoning step in the current generated solution.

9. A process reward model training method according to claim 1, characterized in that: The method of using a generator to generate multiple generated solutions for the question text in each sample includes: After substituting the question text in each sample into the generated solution prompt template, it is input into the generator, and multiple generated solutions for the question text in each sample are generated through multiple sampling.

10. A process reward model training system, characterized in that: include: The data acquisition module is used to obtain a data set, where each sample includes a question text and its corresponding true answer; Use the generator to generate multiple generated solutions for the question text in each sample, including: predicted answers and multiple reasoning steps; A confidence calculation module is used to generate multiple completion solutions for the question text in each sample using the completer, and calculate the confidence score of the question text in each sample based on the true answer in each sample and the predicted answers in the multiple completion solutions; The sample construction module is used to generate the correctness score of each reasoning step in the generated solution of each question text in each sample through the completer, and use the current generated solution, the correctness score of each reasoning step in the generated solution, the question text and its confidence score as a labeled sample of the labeled dataset; A correction module is used to obtain a target correctness score of each reasoning step of the question text in the labeled sample based on the correctness score, confidence score, and tolerance distance hyperparameter of each reasoning step of the question text in the labeled sample; The prediction module is used to input the problem text in the labeled sample in the labeled dataset and the generated solution into the process reward model, and output the correctness prediction score of each reasoning step of the problem text in the labeled sample; The training module is used to train the process reward model based on the binary cross entropy loss between the correctness prediction scores of each reasoning step of the question text in the labeled sample and the target correctness score, and use the trained process reward model as the target process reward model.

Citation Information

Patent Citations

  • Reward model training method, answer evaluation method and device

    CN119849635A

  • Mathematical question and answer method and device, storage medium and equipment

    CN120067253A