Large language model reasoning method based on generative step-level reward model

Through the method based on the generative step-level reward model, the large language model is optimized using supervised fine-tuning and direct preference optimization, and the problems of large resource consumption and low data utilization efficiency in the existing technology are solved, and efficient and interpretable step-level reward model reasoning is achieved.

CN120218240APending Publication Date: 2025-06-27NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510288969.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the Monte Carlo tree search method, the reward value and the probability of the placeholder are highly bound, resulting in a difficult time to effectively identify subtle perturbations, and high resource consumption, low data utilization efficiency, and difficult to truly reflect the accuracy of each step.

Method used

Using a method based on a generative step-level reward model, we collect data to form a supervised fine-tuning SFT dataset and direct preference optimization DPO dataset, fine-tuning and optimization of small models to obtain the R-PRM-DPO model, and use this model to evaluate the inference process of the large language model, obtain the reward value of each step, and optimize the inference process through the reward value.

Benefits of technology

It greatly reduces the calculation cost, avoids the dependence of traditional PRM on placeholder probability, improves the accuracy and interpretability of step judgment, and has high data utilization efficiency, which is suitable for resource-constrained scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218240A_ABST
    Figure CN120218240A_ABST
Patent Text Reader

Abstract

The invention provides a large language model reasoning method based on a generative step-level reward model, which comprises the following steps of: 1, collecting data, and forming a supervised fine tuning (SFT) data set and a direct preference optimization (DPO) data set; 2, performing fine tuning and optimization on the small model by using a supervised fine tuning SFT data set and a direct preference optimization DPO data set to obtain an R-PRM-DPO model; step 3, reasoning by using a large language model, and evaluating a reasoning process by using an R-PRM-DPO model to obtain a reward value of each step in the reasoning process; and step 4, optimizing the big language model reasoning process by using the reward value to realize big language model reasoning of the generative step-level reward model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a large language model inference method, in particular to a large language model inference method based on a generative step-level reward model. Background Art

[0002] What is provided in this part is only background information related to the present disclosure, and it is not necessarily prior art.

[0003] The process reward model (PRM) is an important direction in the research of large language model (LLM) inference. Its core goal is to provide refined feedback for the inference process, thereby effectively enhancing the model's inference ability.

[0004] The key element in constructing an efficient PRM lies in the quality of the training data. However, it is difficult to obtain step-level annotation data. OpenAI initially hired professionals to annotate 800,000 pieces of step-level data (PRM800K, reference: Lightman, Hunter, et al. "Let's verify step by step." The Twelfth International Conference on Learning Representations. 2023.), and trained the PRM based on this. The experimental results show that the performance of this model is better than that of the traditional output reward model (ORM). Given the huge resource investment required for manual annotation, the method based on Monte Carlo tree search (MCTS) has recently become an effective alternative.

[0005] However, in the method based on Monte Carlo tree search (MCTS), there is a problem of insufficient persuasiveness in the high binding between the reward value and the probability of placeholders such as "+", resulting in difficulty in effectively identifying subtle perturbations. In addition, using the MCTS method for data annotation not only consumes a large amount of resources but also has low data utilization efficiency. More critically, this method is difficult to truly reflect the accuracy of each step.

[0006] The existing technical solutions still have significant defects. They rely on MCTS for data collection, and MCTS requires a large number of simulation operations to be performed at each step, resulting in extremely high resource consumption. At the same time, in order to improve data utilization, the additional signal filtering mechanism further increases the computational overhead. In addition, the existing solutions have not undergone a fundamental change and still rely on setting placeholders for each step and calculating specific reward values through the placeholders. This method has obvious limitations in identifying small changes and improving interpretability.

[0007] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0008] Object of the Invention: The technical problem to be solved by the present invention is to provide a large language model inference method based on a generative step-level reward model in view of the deficiencies of the prior art.

[0009] To solve the above technical problem, the present invention discloses a large language model inference method based on a generative step-level reward model, including the following steps:

[0010] Step 1, collect data to form a supervised fine-tuning SFT dataset and a direct preference optimization DPO dataset;

[0011] Step 2, use the supervised fine-tuning SFT dataset and the direct preference optimization DPO dataset to fine-tune and optimize a small model to obtain an R-PRM-DPO model;

[0012] Step 3, perform inference using a large language model, use the R-PRM-DPO model to evaluate the inference process, and obtain the reward value for each step in the inference process;

[0013] Step 4, use the reward value to optimize the large language model inference process to achieve large language model inference of the generative step-level reward model.

[0014] Further, the data collection in Step 1 includes:

[0015] Step 1-1, obtain a step-level dataset with accuracy labels annotated for each step;

[0016] Step 1-2, design an accuracy inference template;

[0017] Step 1-3, based on the step-level dataset obtained in Step 1, use the accuracy inference template designed in Step 1-2 to guide the large language model to perform inference. During the inference process, sample A independent inference chains for each step in the answer;

[0018] Step 1-4, process the A independent inference chains using different strategies to obtain a supervised fine-tuning SFT dataset and a direct preference optimization DPO dataset.

[0019] Further, the accuracy inference template in Step 1-2 specifically includes:

[0020] Preliminary step analysis, current step analysis, data source analysis, consistency analysis, and calculation analysis require standardized judgment statements to be output at the end of the template.

[0021] Furthermore, different strategies are adopted in steps 1-4 for processing to obtain the supervised fine-tuning SFT dataset and the direct preference optimization DPO dataset, specifically including:

[0022] Step 1-4-1, randomly retain 1 inference chain consistent with the accuracy label in each step. If all inference chains do not meet the requirements, the data for this step is discarded to obtain the supervised fine-tuning SFT dataset;

[0023] Step 1-4-2, from the A inference chains in each step, select 1 inference chain consistent with the accuracy label and 1 inference chain inconsistent with the label to form a comparison data pair, obtaining the direct preference optimization DPO dataset.

[0024] Furthermore, using the supervised fine-tuning SFT dataset and the direct preference optimization DPO dataset to fine-tune and optimize the small model in step 2 includes:

[0025] Step 2-1, fine-tune the small model to obtain the R-PRM-SFT model;

[0026] Step 2-2, optimize the R-PRM-SFT model to obtain the R-PRM-DPO model.

[0027] Furthermore, the fine-tuning of the small model described in step 2-1 includes:

[0028] Use the supervised fine-tuning SFT dataset to fine-tune the small model to obtain the inference-driven reward model R-PRM-SFT in the supervised fine-tuning version.

[0029] Furthermore, the optimization of the R-PRM-SFT model described in step 2-2 includes:

[0030] Use the direct preference optimization DPO dataset to optimize the inference-driven reward model R-PRM-SFT in the supervised fine-tuning version to obtain the inference-driven reward model R-PRM-DPO in the direct preference optimization version.

[0031] Furthermore, using the R-PRM-DPO model to evaluate the inference process to obtain the reward value for each step in the inference process includes:

[0032] Step 3-1, use the large language model for inference, that is, generate answers step by step;

[0033] Step 3-2: Using the inference-driven reward model R-PRM-DPO in the direct preference optimization version, generate B accurate inference chains for each current step;

[0034] Step 3-3: Sequentially extract, from each inference chain, the probability values of the flag words reflecting the accurate inference process in the standardized judgment statements output according to the accurate inference template, and calculate the average value of all probability values as the reward quantization index for the current step, that is, the reward value for the current step;

[0035] Step 3-4: Calculate the reward values for all steps according to the method in Step 3-3 to obtain the reward value for each step.

[0036] Further, the number A of the independent inference chains described in Step 1 is 4.

[0037] Further, the number B of the accurate inference chains described in Step 3-2 is 10.

[0038] Beneficial effects:

[0039] 1. The present invention abandons the high-resource-consuming Monte Carlo tree search (MCTS), uses an open-source dataset, combines a structured inference template and a high-performance large model to generate inference chains, and through supervised fine-tuning (SFT) and direct preference optimization (DPO), significantly reduces the computational cost. The reward value is quantified based on the "Yes" probability, avoiding the dependence on the placeholder probability of the traditional PRM. At the same time, due to the generation of the inference chains, it further improves the step judgment accuracy and interpretability. Experiments show that even with a small amount of data, this method still performs excellently, and the data utilization efficiency far exceeds the existing solutions, and it has strong interpretability, and users can clearly understand the reasons for step errors.

[0040] 2. By providing a clear chain of thought, the present invention not only clarifies the reasons for errors in each step of reasoning, but also makes the accuracy judgment more reliable due to the support of the chain of thought. In a mathematical reasoning tool, users can intuitively understand the root cause of step errors through the chain of thought, which is convenient for correction; in model development, this method provides more accurate feedback, and the chain of thought itself can be used as high-quality data for further training the model.

[0041] 3. The present invention can easily construct a practical and efficient PRM, which is suitable for resource-constrained scenarios. While reducing costs, it significantly improves the practicality and accuracy of the PRM with clear and credible inference chains, showing broad application potential. Description of the drawings

[0042] The present invention will be further specifically described below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0043] Figure 1 It is a schematic diagram of the overall working process of the method of the present invention.

[0044] Figure 2 It is a schematic diagram of the data collection process in the present invention.

[0045] Figure 3 It is a schematic diagram of the model training process in the present invention.

[0046] Figure 4 It is a schematic diagram of the implementation and application process in the present invention.

[0047] Figure 5 It is a schematic diagram of the accuracy inference template in an embodiment.

[0048] Figure 6 It is a schematic diagram of an R-PRM-DPO output example in an embodiment. Specific embodiments

[0049] The overall idea of the present invention is as follows:

[0050] Considering the deficiencies of the existing technical solutions in terms of data utilization efficiency and PRM interpretability, a new generative step-level reward model modeling method is proposed. This method can make full use of the existing data and provide specific reasons when generating reward values, enhancing the interpretability of the model.

[0051] The technical key point of the present invention is to propose a generative step-level reward model (PRM) modeling method, and its core innovation is to pioneer the combination of PRM and reasoning (Reasoning). Specifically: 1) Generate an accuracy inference chain through a structured inference template, and for the first time integrate the reasoning process into PRM, providing clear reasons for the right or wrong of each step; 2) Quantify the reward value with the probability of "Yes", and optimize it in combination with SFT and DPO to ensure accurate and interpretable judgment; 3) Based on the existing data set (such as PRM800K), abandon MCTS to reduce resource requirements. The point to be protected focuses on this unique integration: the integrated process of generating through the inference chain and PRM modeling significantly improves the interpretability and practicability of the model, which is different from the traditional method.

[0052] In the technical solution of the present invention, as Figure 1 shown, the main process is as follows:

[0053] First, use the existing open-source step-level data set (such as PRM800K) as the basis. Each step in this data set has been labeled with accurate accuracy labels to ensure the reliability and consistency of the data.

[0054] Subsequently, a set of structured accuracy inference templates are pre-designed, including dimensions such as pre-step analysis, current-step analysis, logical analysis, and computational analysis, and a standardized judgment statement is required to be output at the end of the template, such as: "Is the step correct (Yes / No)?", to standardize the presentation of the inference results.

[0055] Next, use the designed accuracy inference template to guide the high-performance large language model (such as Llama-3.3-70B-Instruct) to generate accuracy inference chains based on PRM800K. To make full use of the existing data, four independent inference chains are sampled for each step in the answer, and pre-annotated accuracy labels are used for screening. The screening rule is: randomly retain one inference chain that is consistent with the accuracy label for each step. If all inference chains do not meet the requirements, the data for that step is discarded, thus constructing a high-quality dataset suitable for supervised fine-tuning (SFT, reference: Wei, Jason, et al. "Finetuned language models are zero-shot learners." arXiv preprint arXiv:2109.01652 (2021)), the SFT dataset.

[0056] Subsequently, using the filtered SFT dataset, a small model (such as Qwen2.5-Math-7B-Instruct) is fine-tuned in a supervised manner to obtain the SFT version (R-PRM-SFT) of the Reasoning-Driven Process Reward Model (R-PRM), enabling it to generate similar reasoning chains for step accuracy judgment, not only significantly improving the clarity of step accuracy judgment but also enhancing the interpretability of the model. To further optimize the judgment accuracy and robustness of R-PRM-SFT, the present invention introduces Direct Preference Optimization (DPO, reference: Rafailov, Rafael, et al. "Direct preference optimization: Your language model is secretly a reward model." Advances in Neural Information Processing Systems 36 (2023): 53728-53741.) to refine R-PRM-SFT. The preference dataset is sourced from the aforementioned sampling stage: among the four reasoning chains for each step, one reasoning chain consistent with the accuracy label and one reasoning chain inconsistent with the label are selected to form a comparison data pair for subsequent preference optimization training. Finally, DPO optimization is performed using the preference dataset collected in the previous stage to obtain the target model with optimized performance (R-PRM-DPO). In addition, the present invention extracts the probability value of the word "Yes" in the sentence "Is the step correct (Yes / No)?", specifically, by providing the same prompt and the reasoning chain generated by the PRM, at the stage of "Is the step correct (Yes / No)?", the probability of the next token being "Yes" is obtained through decoding as the reward quantization index for the current step, providing a refined evaluation basis for the model.

[0057] In the specific implementation process, the present invention first requires the model to generate answers step by step. To ensure the accuracy evaluation of each step, using the R-PRM-DPO of the present invention, 10 accuracy reasoning chains are generated for the current step (more can be supported and it has been experimentally verified that increasing the number of reasoning chains can further improve the model performance). Subsequently, the probability value of "Yes" in the sentence "Is the step correct (Yes / No)?" in each reasoning chain is extracted in sequence, and their average value is calculated, and this average value is used as the reward quantization index for the current step.

[0058] Compared with the probability reward model (PRM) constructed by the traditional Monte Carlo tree search (MCTS) method, the present invention makes full use of existing labeled data, effectively reduces resource consumption through multi-stage screening and optimization, and meanwhile shows significant advantages in terms of the interpretability, computational efficiency and judgment accuracy of the inference process.

[0059] The high-performance large language model mentioned in the present invention refers to a large language model with a parameter quantity greater than or equal to 70 billion, while the small model mentioned refers to a large language model with a parameter quantity less than 10 billion.

[0060] Example:

[0061] The workflow diagram is as Figure 1 shown, and its implementation details will be elaborated below and presented in a step-by-step manner.

[0062] I. Collect data, as Figure 2 shown, the specific process is as follows:

[0063] Step 101: Utilize an open-source step-level dataset (such as PRM800K), which contains labeled step-level data, as the basis for data collection.

[0064] Step 102: Predesign a set of accuracy inference templates, including previous step analysis, current step analysis, logical analysis, calculation analysis, etc., and ensure that the template finally outputs in the form of an accuracy inference chain. The inference template is as Figure 5 shown. First, analyze the previous steps, then analyze the current step......, and finally summarize.

[0065] Step 103: Use a powerful large language model (such as Llama-3.3-70B-Instruct) to generate accuracy inference chains on the specified dataset. As Figure 6 shown is an example of an accuracy inference chain. Since the current step is the first step, there is no need to analyze the previous steps. Then, analyze the current step and find useful conditions. The neighbor placed 18 pink flamingos. Then, analyze that the data comes from the problem description..... Finally, summarize that the current step is correct.

[0066] Step 104: To make full use of the existing data, sample 4 different inference chains for each step.

[0067] Step 105: As Figure 1 shown on the left, use the pre-existing labels to screen the generated inference chains. Randomly retain one inference chain that passes the label comparison for each step for subsequent supervised fine-tuning (SFT); if all inference chains do not meet the requirements, discard the data for that step.

[0068] Step 106: Through the above steps, an SFT dataset suitable for supervised fine-tuning is collected.

[0069] Step 107: Similarly, as Figure 1 shown on the left, when generating the reasoning chain, 4 different accuracy reasoning paths are retained for each step, and paired data is collected from them: a reasoning chain consistent with the label and a reasoning chain inconsistent with the label, forming a Direct Preference Optimization (DPO) dataset.

[0070] II. Training the model, as Figure 3 shown, the specific process is as follows:

[0071] Step 201: Use the collected SFT dataset to perform supervised fine-tuning (SFT) on a small model (such as Qwen2.5-Math-7B-Instruct) to enable it to master the specified accuracy reasoning pattern, obtaining a supervised fine-tuning version of the Reasoning Driven Process Reward Model (R-PRM-SFT).

[0072] Step 202: To further improve the process judgment accuracy of the small model, use Direct Preference Optimization (DPO) to optimize the small model after SFT.

[0073] Step 203: Utilize the DPO dataset (containing paired reasoning chain data) collected in Step 107 to perform preference optimization on the small model after SFT, obtaining a Direct Preference Optimization version of the Reasoning Driven Process Reward Model (R-PRM-DPO).

[0074] Step 204: Through preference optimization, obtain the final small model R-PRM-DPO.

[0075] III. Implementation and application, as Figure 4 shown, the specific process is as follows:

[0076] Step 301: Require the large language model to generate answers step by step, and directly output the reasoning content of each step based on the preset reasoning pattern, providing a basis for subsequent accuracy evaluation.

[0077] Step 302: Utilize the final small model R-PRM-DPO designed by the present invention to generate 10 accuracy reasoning chains for the current step (supporting more, and experimental verification shows that increasing the number of reasoning chains can improve performance).

[0078] Step 303: As shown Figure 1 on the right, extract the probability value of "Yes" after the statement "Is the step correct (Yes / No)?" in each inference chain in Step 302.

[0079] Step 304: Calculate the average value of the "Yes" probability values obtained in Step 303 as the reward quantization index for the current step. Subsequently, the obtained reward value can be used for Best of N to guide the search.

[0080] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the content of the invention of a large language model inference method based on a generative step-level reward model and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0081] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the essence of the technical solutions in the embodiments of the present invention, or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. The computer program software product can be stored in a storage medium and includes several instructions to enable a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, an MCU, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present invention.

[0082] The present invention provides an idea and method of a large language model inference method based on a generative step-level reward model. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A large language model reasoning method based on a generative step-level reward model, characterized in that: The following steps are involved: Step 1: Collect data to form supervised fine-tuning SFT dataset and direct preference optimization DPO dataset; Step 2, fine-tune and optimize the small model using the supervised fine-tuning SFT dataset and the direct preference optimization DPO dataset to obtain the R-PRM-DPO model; Step 3: Use the large language model for reasoning, use the R-PRM-DPO model to evaluate the reasoning process, and obtain the reward value of each step in the reasoning process; Step 4: Use the reward value to optimize the large language model reasoning process and implement the large language model reasoning of the generative step-level reward model.

2. The large language model inference method based on the generative step-level reward model according to claim 1, characterized in that: Collect data as described in step 1, including: Step 1-1, obtain a step-level dataset with accuracy labels for each step; Step 1-2, design accuracy reasoning template; Step 1-3: Based on the step-level dataset obtained in step 1, use the accuracy reasoning template designed in step 1-2 to guide the large language model to perform reasoning. During the reasoning process, sample A independent reasoning chains for each step in the answer. In steps 1-4, different strategies are used to process A independent reasoning chains to obtain the supervised fine-tuning SFT dataset and the direct preference optimization DPO dataset.

3. The large language model inference method based on the generative step-level reward model according to claim 2, characterized in that: The accuracy reasoning template described in step 1-2 specifically includes: Previous step analysis, current step analysis, data source analysis, consistency analysis and calculation analysis require the output of standardized judgment statements at the end of the template.

4. The large language model inference method based on the generative step-level reward model according to claim 3, characterized in that: The different strategies described in steps 1-4 are used to obtain the supervised fine-tuning SFT dataset and the direct preference optimization DPO dataset, including: Step 1-4-1: In each step, one inference chain that is consistent with the accuracy label is randomly retained. If all inference chains do not meet the requirements, the data of this step is discarded to obtain the supervised fine-tuning SFT dataset. In step 1-4-2, among the A reasoning chains in each step, select one reasoning chain that is consistent with the accuracy label and one reasoning chain that is inconsistent with the label to form a comparison data pair and obtain the direct preference optimization DPO dataset.

5. The large language model inference method based on the generative step-level reward model according to claim 4, characterized in that: Fine-tune and optimize the small model using the supervised fine-tuning SFT dataset and the direct preference optimization DPO dataset as described in step 2, including: Step 2-1, fine-tune the small model to obtain the R-PRM-SFT model; Step 2-2, optimize the R-PRM-SFT model to obtain the R-PRM-DPO model.

6. The large language model reasoning method based on the generative step-level reward model according to claim 5, characterized in that: Fine-tune the small model as described in step 2-1, including: The small model is fine-tuned using the supervised fine-tuning SFT dataset to obtain the supervised fine-tuning version of the reasoning-driven reward model R-PRM-SFT.

7. The large language model reasoning method based on the generative step-level reward model according to claim 6, characterized in that: The R-PRM-SFT model is optimized as described in step 2-2, including: The supervised fine-tuning version of the reasoning-driven reward model R-PRM-SFT is optimized using the direct preference optimization DPO dataset to obtain the direct preference optimization version of the reasoning-driven reward model R-PRM-DPO.

8. The large language model inference method based on the generative step-level reward model according to claim 7, characterized in that: The R-PRM-DPO model described in step 3 is used to evaluate the reasoning process and obtain the reward value of each step in the reasoning process, including: Step 3-1, use the large language model for reasoning, that is, generate answers one by one according to the steps; Step 3-2, using the direct preference optimization version of the reasoning-driven reward model R-PRM-DPO, generate B accuracy reasoning chains for each current step; Step 3-3, extracting the probability value of the marker word reflecting that the reasoning process is accurate from the standardized judgment statement output by the accuracy reasoning template in each reasoning chain in turn, and calculating the average value of all probability values ​​as the reward quantification indicator of the current step, that is, the reward value of the current step; Step 3-4, complete the calculation of the reward values ​​of all steps according to the method of step 3-3 to obtain the reward value of each step.

9. The large language model inference method based on the generative step-level reward model according to claim 8, characterized in that: The number A of independent reasoning chains described in step 1 is 4.

10. The large language model inference method based on the generative step-level reward model according to claim 9, characterized in that: The number of accuracy reasoning chains, B, described in step 3-2 is 10.