Method for improving effect of reward model through staged training based on reinforcement learning

Through the phased training method based on reinforcement learning, the problems of high cost of manual annotation data and overfitting in reward model training are solved, and the model performance improvement and self-correction ability are enhanced.

CN119939388APending Publication Date: 2025-05-06GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510028908.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing reward model requires a large amount of manually labeled preference data during training, which is prone to overfitting, resulting in uncontrollable behavior or rewardhacking, and the reward function is difficult to explain.

Method used

A staged training method based on reinforcement learning is adopted, and experience values ​​are generated online and spliced ​​with prompt words, reward intervals are defined, rewards are given according to the content and logical correctness of the generated values, and iterative training is carried out to optimize the model.

Benefits of technology

There is no need for additional reward models for large-model preference alignment, effectively avoid overfitting, improve model performance, and improve the model's self-correction ability through reward feedback mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939388A_ABST
    Figure CN119939388A_ABST
Patent Text Reader

Abstract

The invention relates to a method for improving the effect of a reward model through staged training based on reinforcement learning, and the method comprises the following steps: S1, carrying out the first-stage training: S11, inputting a preset cue word into a model, and enabling the model to generate an empirical value online; s12, splicing the generated empirical value and the corresponding cue word to obtain a spliced input parameter, inputting the spliced input parameter into the model, and enabling the model to judge the content and logic of the generated empirical value; s13, defining a reward interval for the model; if the content and logic of the generated empirical value are correct, giving the maximum reward r1 in the reward interval; s14, repeating the steps from S11 to S13 to correct the generation sequence of the model; s2, carrying out second-stage training; s3, generating an iterated model according to the two stages of training; and S4, freezing the optimized and iterated model parameters, and only training a scoring output layer. The method is used for aligning the preferences of the large models and optimizing the models according to rewards.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning, and in particular to a method for improving the effect of a reward model through phased training based on reinforcement learning. Background Art

[0002] Currently, the training of reward models requires a large amount of manually labeled preference data pairs. The cost of manually labeling millions of high-quality preference data is very high. In addition, since the task design of the reward model during training is too simple, it is easy for the reward model to overfit, and may even cause the strategy model to take uncontrollable actions or serious reward hacking to obtain higher rewards. Furthermore, the reward function is usually a black box model, and it is difficult to explain its internal working mechanism.

[0003] Therefore, it is necessary to provide a method based on reinforcement learning to improve the effect of the reward model through phased training, which can be used to align the preferences of large models and optimize the model according to the rewards. Summary of the invention

[0004] The purpose of the present invention is to provide a method for improving the effect of a reward model by staged training based on reinforcement learning, without using an additional reward model to perform preference alignment of a large model, and optimize the model according to the reward.

[0005] In order to solve the problems existing in the prior art, the present invention provides a method for improving the effect of a reward model by staged training based on reinforcement learning, comprising the following steps:

[0006] S1: Conduct the first stage of training. The training method is as follows:

[0007] S11: inputting a preset prompt word into the model, and the model generates an experience value online;

[0008] S12: splicing the generated experience value and the corresponding prompt word to obtain a splicing parameter, and inputting the splicing parameter into the model so that the model judges the content and logic of the generated experience value;

[0009] S13: Define a reward interval for the model; if the content and logic of the generated experience value are correct, the maximum reward r1 in the reward interval is given;

[0010] S14: repeating steps S11 to S13 to correct the generation sequence of the model;

[0011] S2: Conduct the second stage of training. The training method is as follows:

[0012] S21: inputting a preset prompt word into the model, and the model generates an experience value online;

[0013] S22: splicing the generated experience value and the corresponding prompt word to obtain a splicing parameter, and inputting the splicing parameter into the model so that the model judges the content and logic of the generated experience value;

[0014] S23: Define a reward interval for the model; if the content and logic of the generated experience value are correct, the maximum reward r2 in the reward interval is given;

[0015] S3: Generate an iterated model based on two-stage training;

[0016] S4: Scoring output layer of the model after optimization iteration.

[0017] Optionally, in the method for improving the effect of the reward model by staged training based on reinforcement learning,

[0018] The method of judging the content of the generated experience value is as follows: judging whether it is a generated experience value;

[0019] The logic of the generated experience value is judged as follows: whether the generated experience value is correct.

[0020] Optionally, in the method for improving the effect of the reward model through phased training based on reinforcement learning, a reward interval is defined for the model so that the reward interval is -1 to 1.

[0021] Optionally, in the method for improving the effect of the reward model by staged training based on reinforcement learning,

[0022] If the content and logic of the generated experience value are correct, the reward value is 1;

[0023] If the content of the generated experience value is wrong, but the logic of the generated experience value is correct, the reward value is 0.5;

[0024] If the logic of the generated experience value is wrong, the reward value will be -1.

[0025] Optionally, in the method for improving the effect of the reward model through phased training based on reinforcement learning, the generation sequence is formed by multiple experience values.

[0026] Optionally, in the method for improving the effect of the reward model by staged training based on reinforcement learning, the way of optimizing the scoring output layer of the iterative model is as follows:

[0027] Freeze the model after the iteration;

[0028] Obtain the experience value generated by the iterative model, and mark the obtained experience value. The marking method is: if the experience value is correct, it is marked as 1, and if the experience value is wrong, it is marked as -1;

[0029] Use action masking to remove non-generated tokens;

[0030] Get the annotated mean, use the mean square error to construct the loss function, and get the value;

[0031] Based on the obtained value, the activation function uses tanh(x) for numerical mapping;

[0032] The gradient descent method is used to optimize the parameters of the scoring output layer.

[0033] Compared with the prior art, the present invention has the following advantages:

[0034] (1) This method is mainly used in the field of reinforcement learning. It can align the preferences of large models without using additional reward models and optimize the model based on rewards.

[0035] (2) Experiments were conducted on the MATH dataset using the OpenMath2-Llama-3.1-8B and dart-math-mistral-7b-uniform base models. Openmath2-Llama-3.1-8B achieved improvements of 1.64 and 2.62 in the two stages, and dartmath-mistral-7B achieved improvements of 0.82 and 1.08 in the two stages. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A flow chart of phased training provided by an embodiment of the present invention;

[0037] Figure 2 A method flow chart is provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0038] The specific implementation of the present invention will be described in more detail below in conjunction with the schematic diagram. The advantages and features of the present invention will become clearer based on the following description. It should be noted that the drawings are all in a very simplified form and are not in exact proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.

[0039] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0040] Currently, the training of reward models requires a large amount of manually labeled preference data pairs. The cost of manually labeling millions of high-quality preference data is very high. In addition, since the task design of the reward model during training is too simple, it is easy for the reward model to overfit, and may even cause the strategy model to take uncontrollable actions or serious reward hacking to obtain higher rewards. Furthermore, the reward function is usually a black box model, and it is difficult to explain its internal working mechanism.

[0041] In order to solve the problems existing in the prior art, the present invention provides a method for improving the effect of the reward model by staged training based on reinforcement learning, such as Figure 1 and 2 As shown, the method comprises the following steps:

[0042] S1: Conduct the first stage of training. The training method is as follows:

[0043] S11: inputting a preset prompt word into the model, and the model generates an experience value online;

[0044] S12: splicing the generated experience value and the corresponding prompt word to obtain a splicing parameter, and inputting the splicing parameter into the model so that the model judges the content and logic of the generated experience value;

[0045] The method for judging the content of the generated experience value is as follows: judging whether it is a generated experience value; the method for judging the logic of the generated experience value is as follows: judging whether the generated experience value is correct.

[0046] S13: Define a reward range for the model, for example, set the reward range to -1 to 1; if the content and logic of the generated experience value are correct, assign the maximum reward r1 in the reward range; for example, if the content and logic of the generated experience value are correct, assign the reward value to 1; if the content of the generated experience value is wrong, but the logic of the generated experience value is correct, assign the reward value to 0.5; if the logic of the generated experience value is wrong, regardless of whether the content of the generated experience value is correct or wrong, assign the reward value to -1.

[0047] S14: repeating steps S11 to S13 to correct a generation sequence of the model, wherein the generation sequence is formed by a plurality of experience values;

[0048] The rewards are reshaped by the joint optimization goal. In the first stage of training, only the optimization paradigm of "prompt word + experience value content + judgment + judgment result" is given to maximize the reward. This step requires reshaping the rewards of the "output" and "judgment result" of the first stage, that is, to let the model learn whether the "output" is correct or the "judgment result" is correct, and which error has a more serious impact.

[0049] S2: Conduct the second stage of training. The training method is as follows:

[0050] S21: inputting a preset prompt word into the model, and the model generates an experience value online;

[0051] S22: splicing the generated experience value and the corresponding prompt word to obtain a splicing parameter, and inputting the splicing parameter into the model so that the model judges the content and logic of the generated experience value;

[0052] The method for judging the content of the generated experience value is as follows: judging whether it is a generated experience value; the method for judging the logic of the generated experience value is as follows: judging whether the generated experience value is correct.

[0053] S23: Define a reward interval for the model; if the content and logic of the generated experience value are correct, the maximum reward r2 in the reward interval is given;

[0054] Then the total reward for the second stage is: r = r1 + r2 + β (r2 - r1). We only need to optimize the above maximum reward, where β is a hyperparameter, which represents the penalty coefficient for the difference between the two-stage rewards, for example, it can be 10.

[0055] The model is trained under a distribution of self-generated corrected trajectories and uses appropriate regularization to guide the learning process. The scoring model is trained in two stages: in the first stage, it corrects the first generated sequence by training, while using KL divergence to constrain the distribution of the first round to be close to the initialization model; then in the second stage, these two attempts are trained to maximize the reward. Crucially, the second stage of multi-round RL uses reward feedback, which rewards self-correction rather than the correctness of the final generated sequence.

[0056] S3: Generate an iterated model based on two-stage training;

[0057] S4: Optimize the scoring output layer of the iterated model as follows:

[0058] Freeze the model after the iteration;

[0059] Obtain the experience value generated by the iterative model, and mark the obtained experience value. The marking method is: if the experience value is correct, it is marked as 1, and if the experience value is wrong, it is marked as -1;

[0060] Use action masking to remove tokens (proper nouns) that are not generated;

[0061] Get the annotated mean, use the mean square error to construct the loss function, and get the value;

[0062] Based on the obtained value, the activation function uses tanh(x) for numerical mapping;

[0063] The gradient descent method is used to optimize the parameters of the scoring output layer.

[0064] In summary, compared with the prior art, the present invention has the following advantages:

[0065] (1) This method is mainly used in the field of reinforcement learning. It can align the preferences of large models without using additional reward models and optimize the model based on rewards.

[0066] (2) Experiments were conducted on the MATH dataset using the OpenMath2-Llama-3.1-8B and dart-math-mistral-7b-uniform base models. Openmath2-Llama-3.1-8B achieved improvements of 1.64 and 2.62 in the two stages, and dartmath-mistral-7B achieved improvements of 0.82 and 1.08 in the two stages.

[0067] The above is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any technician in the relevant technical field, without departing from the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification to the technical solution and technical content disclosed in the present invention, which does not depart from the content of the technical solution of the present invention and still falls within the protection scope of the present invention.

Claims

1. A method for improving the effect of a reward model by staged training based on reinforcement learning, characterized in that: The following steps are involved: S1: Conduct the first stage of training. The training method is as follows: S11: inputting a preset prompt word into the model, and the model generates an experience value online; S12: splicing the generated experience value and the corresponding prompt word to obtain a splicing parameter, and inputting the splicing parameter into the model so that the model judges the content and logic of the generated experience value; S13: Define a reward interval for the model; if the content and logic of the generated experience value are correct, the maximum reward r1 in the reward interval is given; S14: repeating steps S11 to S13 to correct the generation sequence of the model; S2: Conduct the second stage of training. The training method is as follows: S21: inputting a preset prompt word into the model, and the model generates an experience value online; S22: splicing the generated experience value and the corresponding prompt word to obtain a splicing parameter, and inputting the splicing parameter into the model so that the model judges the content and logic of the generated experience value; S23: Define a reward interval for the model; if the content and logic of the generated experience value are correct, the maximum reward r2 in the reward interval is given; S3: Generate an iterated model based on two-stage training; S4: Scoring output layer of the model after optimization iteration.

2. The method for improving the effect of the reward model by staged training based on reinforcement learning as claimed in claim 1, characterized in that: The method of judging the content of the generated experience value is as follows: judging whether it is a generated experience value; The logic of the generated experience value is judged as follows: whether the generated experience value is correct.

3. The method for improving the effect of the reward model by staged training based on reinforcement learning as claimed in claim 2, characterized in that: A reward interval is defined for the model, so that the reward interval is -1 to 1.

4. The method for improving the effect of the reward model by staged training based on reinforcement learning as claimed in claim 3, characterized in that: If the content and logic of the generated experience value are correct, the reward value is 1; If the content of the generated experience value is wrong, but the logic of the generated experience value is correct, the reward value is 0.5; If the logic of the generated experience value is wrong, the reward value will be -1.

5. The method for improving the effect of the reward model by staged training based on reinforcement learning as claimed in claim 1, characterized in that: The generation sequence is formed from a plurality of experience values.

6. The method for improving the effect of the reward model by staged training based on reinforcement learning as claimed in claim 1, characterized in that: The way to optimize the scoring output layer of the iterative model is as follows: Freeze the model after the iteration; Obtain the experience value generated by the iterative model, and mark the obtained experience value. The marking method is: if the experience value is correct, it is marked as 1, and if the experience value is wrong, it is marked as -1; Use action masking to remove non-generated tokens; Get the annotated mean, use the mean square error to construct the loss function, and get the value; Based on the obtained value, the activation function uses tanh(x) for numerical mapping; The gradient descent method is used to optimize the parameters of the scoring output layer.