Method for balancing reasoning length and accuracy in large model based on step-by-step exploration and preference optimization
Through the method of step-by-step sampling and parameter merging, a twin model is built to balance the inference length and accuracy of large language models, solving the problem of overthinking large models in the inference process, and achieving more efficient inference effects.
Patent Information
- Application Number
- CN202510754777.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-02
AI Technical Summary
Existing large language models have problems of overthinking in the inference process, resulting in high computational costs and reduced accuracy on simple problems. It is difficult for existing strategies to effectively balance the length and accuracy of reasoning.
A step-by-step sampling strategy is used to generate diversified inference paths, two sets of preference data are constructed to train twin models separately, and the model parameters are optimized through parameter merging to balance the inference length and accuracy, and the model is trained using reward function and direct preference optimization method.
The inference length of the large model is significantly reduced (usually 30-50%), while maintaining or improving the accuracy of the inference, achieving excellent performance on mathematical inference tasks.
Smart Images

Figure CN120579640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of compressing the inference length of large language models, and in particular to a method for balancing the inference length and accuracy in a large model based on gradual exploration and preference optimization. Background Art
[0002] Recent advances in Chain-of-Thought (CoT) methods have significantly improved the reasoning capabilities of large language models (LLMs), inspiring researchers to explore a new scaling paradigm: test-time scaling. This paradigm significantly improves the performance of large models on many challenging reasoning tasks, such as mathematics competitions and doctoral-level question answering, by extending the reasoning trajectory during reasoning, allowing for deeper and iterative thinking. While effective, test-time scaling comes with a high computational cost. Furthermore, recent research suggests that large models may exhibit overthinking behavior on some relatively simple problems, which could undermine the advantages of deep reasoning in some cases. Various strategies have been proposed to address the inefficiencies caused by overthinking.
[0003] In their paper, Ayeong Lee, Ethan Che, and Tianyi Peng. 2025. How well do llms compress their own chain-of-thought? a token complexity approach., they designed 31 different hinting strategies to systematically study the ability of large language models to compress their own chains of thought. These hinting strategies aim to limit the length of inferences generated by the model, thereby evaluating the trade-off between inference length and model accuracy. Hints include: limiting the number of words, limiting the number of characters, limiting the number of inference steps, and removing all punctuation. The study found that regardless of the compression strategy used, a consistent trade-off curve existed between inference length and accuracy, indicating that the length of the inference chain, rather than its specific format, is the primary factor affecting model performance.
[0004] Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025. S1: Simple test-time scaling. They proposed the S1 method, a supervised fine-tuning-based approach that prompts a large model to sample multiple inference paths and selects concise and accurate paths to synthesize supervised fine-tuning data for fine-tuning the large model. This method leverages a more powerful large model (such as ChatGPT) to generate concise and accurate inference paths as training data for supervised fine-tuning.
[0005] Pranjal Aggarwal and Sean Welleck (2025). L1: Controlling how long areasoning model thinks with reinforcement learning. propose the L1 method, a reinforcement learning-based approach to control the length and accuracy of inference trajectories generated by large models. This method leverages the fact that reinforcement learning enables large models to adaptively control the length of inferences. It introduces a length bias penalty and an accuracy reward during training to encourage large models to generate accurate inferences using shorter inference lengths. Summary of the Invention
[0006] The purpose of the present invention is to guide the large model to generate finer-grained and diversified reasoning paths to construct higher-quality reinforcement learning training data by adopting a step-by-step sampling strategy, thereby training the model to achieve the effect of balancing reasoning accuracy and reasoning length. The step-by-step sampling strategy of the present invention can synthesize diverse reasoning paths, from which two sets of different preference data are constructed, and the same model is trained separately, thereby obtaining two twin models with different advantages. By using parameter interpolation to merge the parameters of the two models, the characteristics of the two large models are combined, which not only improves the accuracy of model reasoning, but also further shortens the reasoning process.
[0007] The technical solution of the present invention is a method for balancing inference length and accuracy in a large model based on stepwise exploration and preference optimization, which specifically includes the following steps:
[0008] Step 1: Use a large model and adopt a long-short switching sampling strategy to gradually explore reasoning trajectories while building a pool of reasoning trajectories of different lengths for each reasoning problem;
[0009] Step 2: Construct two different sets of preference data from the inference trajectory pool, one for training the large model to have length advantages and the other for accuracy advantages. Based on the two sets of preference data, use the direct preference optimization method to train the same large model separately, thus obtaining two twin large models with complementary advantages. Each twin large model has advantages in the accuracy and length of the generated inference trajectory.
[0010] Step 3: Use the parameter merging method to merge the parameters of the two twin large models to obtain a merged large model to balance the inference accuracy and inference length.
[0011] The step 1 is specifically as follows:
[0012] Step 1.1: Given a problem , the inference trajectory generated by the large model is recorded as ,in is an intermediate reasoning step, is the final answer; at a certain step , using a long-short switching sampling strategy for sampling; specifically, using two different instructions to prompt the large model to use the currently selected partial optimal inference trajectory Generate subsequent inference traces of varying lengths:
[0013]
[0014]
[0015] Generate long subsequent reasoning trajectories for guiding large models Tips; Generate short subsequent inference trajectories for guiding large models Tips;
[0016] Step 1.2: Partial optimal inference trajectory Merge with each newly generated subsequent reasoning trajectory to form a complete candidate trajectory, and combine the two candidate trajectories and Add to the inference trajectory pool middle:
[0017]
[0018]
[0019] in and is the candidate trajectory after merging;
[0020] Step 1.3: Use the reward function r to evaluate the merged candidate trajectory and Continue to explore new optimal inference trajectories:
[0021] It encourages the correctness of the reasoning trajectory while penalizing its length: when the final answer of the trajectory is correct, the large model is expected to explore shorter trajectories to shorten the length; when the final answer is incorrect, the model tends to explore longer trajectories, because longer reasoning chains help correct wrong answers; the rewards of the two candidate trajectories obtained by comparison are and , select the first candidate trajectory from the candidate trajectory with higher reward The optimal reasoning step :
[0022]
[0023] Finally, the new optimal reasoning trajectory is obtained by merging the optimal reasoning step with the previous optimal reasoning trajectory. :
[0024]
[0025] Based on the new optimal inference trajectory , iteratively repeat the above steps until the maximum number of steps is reached or the generation process is terminated.
[0026] The reward function r is specifically:
[0027] .
[0028] The step 2 is specifically as follows:
[0029] Step 2.1: Based on the candidate pool obtained in step 1 Start constructing preference data pairs; from the candidate pool Choose the correct answer And the trajectory with the shortest inference length is taken as the positive sample:
[0030]
[0031] Select two negative samples, one negative sample has the wrong answer but the longest inference length trajectory, and the other negative sample has the correct answer but the longest inference length trajectory:
[0032] .
[0033] .
[0034] Construct two preference datasets, and ; Contains triples (q, , ) Focus on accuracy; Contains triples (q, ), focusing on length compression;
[0035] Step 2.2: Based on two different sets of preference data and , use direct preference optimization (DPO) to fine-tune the large model separately:
[0036]
[0037]
[0038] The DPO training loss is expressed as follows:
[0039]
[0040] in is a hyperparameter, As a reference model, the parameters are kept frozen during the training process; in this way, a twin large model is obtained and ,in is a length model used to generate shorter inference traces, It is a precision model used to generate more accurate inference trajectories.
[0041] The step 3 is specifically as follows:
[0042] Step 3.1: Use the model parameter merging strategy to merge the length model and precision model Perform parameter interpolation:
[0043]
[0044] Precision Model and length model Parameters; from Sparse selection of some parameters, by Control and use weights Perform weighted and precision model The parameters of the combined model are merged to obtain the combined model ;in Indicates from Select the top x percent of parameters with the largest changes.
[0045] Use the merged large model during inference Make inferences.
[0046] Beneficial effects of the present invention: The present invention proposes a method for balancing the inference length and accuracy in large models based on gradual exploration and preference optimization, which has achieved significant improvement in mathematical reasoning tasks. Experimental results show that this method performs well on two different large language models, Qwen2.5-7b and Llama3.1-8b, and also achieves excellent performance on multiple mathematical reasoning task datasets, including AMC, AME, GSM8K, MATH500 and other mathematical reasoning tasks. Compared with existing baseline methods, the present invention not only significantly reduces the inference length of large models, typically by 30-50%, but also maintains or even further improves the accuracy of reasoning. The present invention is effective in achieving the optimal compromise between balancing the accuracy and efficiency of large model reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Flowchart of a method for balancing inference length and accuracy in large models based on stepwise exploration and preference optimization. DETAILED DESCRIPTION
[0048] A method based on stepwise exploration and preference optimization to balance inference length and accuracy in large models. The specific process is as follows:
[0049] Step 1: Collect question-answer pairs from multiple existing public mathematical datasets. For each problem in the collected dataset, use the large model to employ a long-short switching sampling strategy to gradually generate the optimal reasoning trajectory. While exploring reasoning paths, a pool of reasoning trajectories of varying lengths is adaptively constructed for each reasoning problem.
[0050] Step 2: Construct two different sets of preference data from the inference trajectory pool, one for training the large model with length advantages and the other for accuracy advantages. Based on these two sets of preference data, a direct preference optimization method is used to train the same large model separately, resulting in two twin large models with complementary advantages. Each twin large model has advantages in both accuracy and length of generated inference trajectories, resulting in more accurate and shorter inference trajectories.
[0051] Step 3: Use the DARE-TIES parameter merging method to merge the parameters of the trained twin model to obtain a parameter-merged model to achieve a balance between model inference accuracy and inference length. Furthermore, the step is specifically implemented by executing the following sub-steps:
[0052] The step 1 is specifically as follows:
[0053] Step 1.1: Collect question-answer pairs from existing public mathematical datasets. , the inference trajectory generated by the large model is recorded as ,in is an intermediate reasoning step, is the final answer. ,use and Two different instructions prompt the large model to select the optimal reasoning trajectory based on the current part Generate subsequent inference traces of varying lengths:
[0054]
[0055]
[0056] and are prompts used to guide the large model to generate subsequent reasoning trajectories of different lengths, where Guide the model to generate long subsequent reasoning traces , Generate short subsequent reasoning trajectories .
[0057] Step 1.2: Convert the existing optimal reasoning trajectory Merge with each newly generated subsequent reasoning trajectory to form a complete candidate trajectory, and combine the two trajectories and Added to the candidate pool middle:
[0058]
[0059]
[0060] in and is the merged inference trajectory.
[0061] Step 1.3: Use the designed reward function r to evaluate the merged candidate trajectories and :
[0062]
[0063] It encourages the correctness of the reasoning trajectory while penalizing its length: when the final answer of the trajectory is correct, the large model is expected to explore shorter trajectories to shorten the length; when the final answer is incorrect, the large model is inclined to explore longer trajectories, because longer reasoning chains may help correct incorrect answers. The rewards of the two candidate trajectories obtained by comparison are and , select the first candidate trajectory from the candidate trajectory with higher reward The optimal reasoning step :
[0064]
[0065] Finally, the latest optimal reasoning trajectory is obtained by merging the optimal reasoning step with the previous optimal reasoning trajectory. :
[0066]
[0067] Based on the new optimal inference trajectory , iteratively repeat the above steps until the maximum number of steps is reached or the generation process is terminated.
[0068] The step 2 is specifically as follows:
[0069] Step 2.1: Based on the candidate pool obtained in step 1 Start constructing preference data pairs. From the candidate pool Choose the correct answer And the trajectory with the shortest inference length is taken as the positive sample:
[0070]
[0071] Select two negative samples, one negative sample has the wrong answer but the longest inference length trajectory, and the other negative sample has the correct answer but the longest inference length trajectory:
[0072] .
[0073] .
[0074] Construct two preference datasets, and , the former contains triples (q, , ) focuses on accuracy, the latter contains the triple (q, ), focusing on length compression.
[0075] Step 2.2: Based on two different sets of preference data and , fine-tune the large model separately using Direct Preference Optimization (DPO):
[0076]
[0077]
[0078] The DPO training loss is expressed as follows:
[0079]
[0080] in is a hyperparameter, As a reference model, the parameters remain frozen during the training process. In this way, a twin large model can be obtained and ,in The length model can generate shorter inference traces, The precision model can generate more accurate inference trajectories.
[0081] The specific steps of step 3 are as follows:
[0082] Step 3.1: Use the DARE-TIES model parameter merging strategy to merge the length model and precision model Perform parameter interpolation:
[0083]
[0084] Precision Model and length model Parameters. Sparse selection of some parameters (by control), and use weights Perform weighted and precision model The parameters of the combined model are merged to obtain the combined model .
[0085] Tables 1 and 2 compare the effectiveness of the present invention and existing inference length control methods in mathematical reasoning tasks, where the data values represent the Pass@1 pass rate and inference length #Tok answered by the large model on the relevant inference data. As can be seen from Tables 1 and 2, the present invention achieves the best comprehensive results in both Pass@1 pass rate and inference length #Tok in large model inference scenarios. In mathematical reasoning task scenarios, the accuracy of the present invention not only surpasses traditional fine-tuning methods, but also surpasses some hint-based inference length control methods, such as draft chaining, and also surpasses many reinforcement learning methods, such as L1, fully demonstrating the significant advantages of the present invention in enhancing inference efficiency.
[0086] Table 1 Comparison of the effects of the Qwen2.5-7B large model
[0087]
[0088] Table 2 Comparison of the effects of the Llama-3.1-8B large model
[0089]
Claims
1. A method for balancing inference length and accuracy in large models based on stepwise exploration and preference optimization, characterized by: The specific steps are as follows: Step 1: Use a large model and adopt a long-short switching sampling strategy to gradually explore reasoning trajectories while building a pool of reasoning trajectories of different lengths for each reasoning problem; Step 2: Construct two different sets of preference data from the inference trajectory pool, one for training the large model to have length advantages and the other for accuracy advantages. Based on the two sets of preference data, use the direct preference optimization method to train the same large model separately, thus obtaining two twin large models with complementary advantages. Each twin large model has advantages in the accuracy and length of the generated inference trajectory. Step 3: Use the parameter merging method to merge the parameters of the two twin large models to obtain a merged large model to balance the inference accuracy and inference length.
2. The method for balancing inference length and accuracy in large models based on stepwise exploration and preference optimization according to claim 1, characterized in that: The step 1 is specifically as follows: Step 1.1: Given a problem , the inference trajectory generated by the large model is recorded as ,in is an intermediate reasoning step, is the final answer; At a certain step , using a long-short switching sampling strategy for sampling; specifically, using two different instructions to prompt the large model to use the currently selected partial optimal inference trajectory Generate subsequent inference traces of varying lengths: , Generate long subsequent reasoning trajectories for guiding large models Tips; Generate short subsequent inference trajectories for guiding large models Tips; Step 1.2: Partial optimal inference trajectory Merge with each newly generated subsequent reasoning trajectory to form a complete candidate trajectory, and combine the two candidate trajectories and Add to the inference trajectory pool middle: ,in and is the candidate trajectory after merging; Step 1.3: Use the reward function r to evaluate the merged candidate trajectory and Continue to explore new optimal inference trajectories: The rewards of two candidate trajectories obtained by comparison and , select the first candidate trajectory from the candidate trajectory with higher reward The optimal reasoning step : Finally, the new optimal reasoning trajectory is obtained by merging the optimal reasoning step with the previous optimal reasoning trajectory. : , based on the new optimal inference trajectory , iteratively repeat the above steps until the maximum number of steps is reached or the generation process is terminated.
3. The method for balancing inference length and accuracy in a large model based on stepwise exploration and preference optimization according to claim 2, characterized in that: The reward function r is specifically: 。 4. The method for balancing inference length and accuracy in large models based on stepwise exploration and preference optimization according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1: Based on the candidate pool obtained in step 1 Start constructing preference data pairs; from the candidate pool Choose the correct answer And the trajectory with the shortest inference length is taken as the positive sample: , select two negative samples, one negative sample has the wrong answer but the longest inference length trajectory, and the other negative sample has the correct answer but the longest inference length trajectory: . . Construct two preference datasets, and ; Contains triples (q, , ) Focus on accuracy; Contains triples (q, ), focusing on length compression; Step 2.2: Based on two different sets of preference data and , use direct preference optimization (DPO) to fine-tune the large model separately: 。 5. The method for balancing inference length and accuracy in a large model based on stepwise exploration and preference optimization according to claim 4, characterized in that: The DPO training loss is expressed as follows: ,in is a hyperparameter, As a reference model, the parameters are kept frozen during the training process; in this way, a twin large model is obtained and ,in is a length model used to generate shorter inference traces, It is a precision model used to generate more accurate inference trajectories.
6. The method for balancing inference length and accuracy in large models based on stepwise exploration and preference optimization according to claim 1, characterized in that: The step 3 is specifically as follows: Step 3.1: Use the model parameter merging strategy to merge the length model and precision model Perform parameter interpolation: Precision Model and length model Parameters; from Sparse selection of some parameters, by Control and use weights Perform weighted and precision model The parameters of the combined model are merged to obtain the combined model ;in Indicates from Select the top x percent of parameters with the largest changes.
7. The method for balancing inference length and accuracy in a large model based on stepwise exploration and preference optimization according to claim 6, characterized in that: Use the merged large model during inference Make inferences.