Reinforcement learning training method of large language model and related equipment
By introducing reinforcement learning training methods into large language models, using prompt data and preference data to pre-train reward models, and carrying out intensive training of actor models, the problem of poor behavior recognition and feedback generation of large language models in Chinese scenarios is solved, and the model is efficiently trained and stable output in Chinese scenarios is achieved.
Patent Information
- Application Number
- CN202411853165.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-13
AI Technical Summary
The lack of effective reinforcement learning training methods in the prior art, especially in Chinese scenarios, has led to poor performance of large language models in behavior recognition and feedback generation.
A reinforcement learning training method for large language models is proposed. By obtaining prompt data and preference data, the reward model is pre-trained, and the prompt data is input to the actor model for response training, the reward model and commentator model are called to evaluate the response quality and generation strategy, and training feedback is generated, reinforcement training is carried out, monitoring ability improvement indicators and updating commentator model parameters.
It realizes accurate recognition and feedback generation of large language model behaviors, improves the performance and training efficiency of the model in Chinese scenarios, and ensures the quality of model output and the stability of the strategy.
Smart Images

Figure CN119990303A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of electronic technology, and in particular to a reinforcement learning training method, device, server, computer-readable storage medium, and computer program product for a large language model. Background Art
[0002] Large language models (LLMs) have become a hot topic in the field of natural language processing (NLP) research and application. Not only do they significantly outperform existing models in many traditional natural language processing tasks, but they also have strong versatility and integrate multiple capabilities. In addition, large language models have opened up a new paradigm for human-computer interaction, pushing the language, intelligence, and general capabilities of machines to a new level. Using language as a carrier, large language models have expanded their capabilities to different industries, and to some extent have brought the dawn of general artificial intelligence (AGI).
[0003] In the process of training large language models, reinforcement learning from human feedback (RLHF) is particularly important, which further aligns the performance of large language models with human preferences and significantly improves the various capabilities of large language models. Since reinforcement learning belongs to the frontier field, many works such as Instruct-GPT and LLaMA2 are very general in their introduction to reinforcement learning, and the specific technical details contained are not detailed enough. It is difficult to successfully conduct reinforcement learning training based on only a few words in the literature. In recent years, the industry has successively open-sourced a number of reinforcement learning frameworks. However, most of these frameworks are only basic implementations of reinforcement learning, and the effectiveness of their actual training has not been verified. In addition, most of these frameworks are open sourced by foreign units, and they mainly focus on English proficiency, while there are significant differences in language preferences between Chinese and English. Therefore, in Chinese scenarios, there has been a lack of effective methods for reinforcement learning training. Summary of the invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desired to provide a reinforcement learning training method, device, server, computer-readable storage medium and computer program product for a large language model that can capture dynamic behavior information and achieve accurate behavior recognition.
[0005] In a first aspect, an embodiment of the present application provides a reinforcement learning training method for a large language model, comprising:
[0006] Obtain prompt data and corresponding preference data for reinforcement learning training;
[0007] pre-training a reward model based on the preference data;
[0008] Input the prompt data into the actor model for response training, call the reward model to evaluate the response quality, call the critic model to evaluate the response generation strategy, and generate training feedback according to the evaluation results of the reward model and the critic model;
[0009] Performing intensive training on the actor model according to the training feedback, monitoring the capability improvement index during the intensive training, and synchronously updating the parameters of the critic model;
[0010] An iteratively optimized actor model is determined according to the capability improvement index and output as a large language model.
[0011] In one embodiment, performing intensive training on the actor model according to the training feedback and monitoring the capability improvement index during the intensive training include:
[0012] inputting the same prompt data into the actor model and the reference model respectively to generate response data;
[0013] Inputting the prompt data and the response data generated by the actor model and the reference model, respectively, into the reward model for quality evaluation to generate a reward value;
[0014] The mean of the reward values corresponding to the actor model and the reference model are calculated as monitoring indicators.
[0015] In one embodiment, the synchronously updating the parameters of the critic model includes:
[0016] Determine the loss value and gradient value of the actor model and the critic model in the enhanced training, and judge whether the loss value and / or gradient value are abnormal;
[0017] If an exception occurs, the model with the exception is trained separately.
[0018] In one embodiment, performing reinforcement training on the actor model according to the training feedback includes:
[0019] Determining sampling parameters for supervised fine-tuning of the actor model as sampling parameters for the reinforcement training;
[0020] The learning rate and KL divergence weight of the reinforcement training are dynamically adjusted according to the reward value output by the reward model.
[0021] In one embodiment, the obtaining of prompt data and corresponding preference data for RLHF training includes:
[0022] Collect prompt data of each sub-field under the target field according to the preset proportion corresponding to each sub-field; wherein the target field includes a general field and an application scenario field;
[0023] Calling the supervised fine-tuned actor model to generate response data corresponding to the prompt data;
[0024] By comparing the response data, the response data is annotated with preference information and preference strength to generate preference data.
[0025] In one embodiment, pre-training a reward model according to the preference data includes:
[0026] Train the first model based on sentence-level contrastive loss;
[0027] Using the first model to score the output of the actor model, and performing indicator evaluation on the scoring process;
[0028] Train the second model based on the token-level contrastive loss;
[0029] Initialize the critic model using the second model;
[0030] Determine whether the results of the indicator evaluation meet expectations;
[0031] If achieved, the first model is used as a reward model;
[0032] If not reached, execute the step of training the first model according to the sentence-level contrastive loss.
[0033] In a second aspect, the embodiment of the present application provides a reinforcement learning training device for a large language model, including:
[0034] A data acquisition unit, used to acquire prompt data and corresponding preference data for reinforcement learning training;
[0035] A reward model pre-training unit, used for pre-training a reward model according to the preference data;
[0036] A training feedback generating unit, used for inputting the prompt data into the actor model for response training, calling the reward model to evaluate the response quality, calling the critic model to evaluate the response generation strategy, and generating training feedback according to the evaluation results of the reward model and the critic model;
[0037] An enhanced training unit, used to perform enhanced training on the actor model according to the training feedback, monitor the capability improvement index during the enhanced training, and synchronously update the parameters of the critic model;
[0038] The model output unit is used to determine the iteratively optimized actor model according to the capability improvement index and output it as a large language model.
[0039] In a third aspect, an embodiment of the present application provides a server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the embodiment of the present application when executing the program.
[0040] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the embodiment of the present application.
[0041] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method described in the embodiment of the present application.
[0042] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0044] Figure 1 A schematic diagram of a process flow of a reinforcement learning training method for a large language model provided in an embodiment of the present application is shown;
[0045] Figure 2 A schematic diagram of fine-grained splitting and data matching in a general field provided by an embodiment of the present application is shown;
[0046] Figure 3 A schematic diagram of a preference marking system page provided in an embodiment of the present application is shown;
[0047] Figure 4 A schematic diagram of an overall RLHF training process provided in an embodiment of the present application is shown;
[0048] Figure 5 An exemplary structural block diagram of a reinforcement learning training device for a large language model provided in an embodiment of the present application is shown;
[0049] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing a server of an embodiment of the present application is shown. DETAILED DESCRIPTION
[0050] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the invention are shown in the accompanying drawings.
[0051] It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments. Although the embodiments of the present application provide the method operation instruction steps shown in the following embodiments or drawings, more or fewer operation instruction steps may be included in the method based on routine or no creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present application. The method may be executed in the order of the methods shown in the embodiments or drawings or in parallel during the actual processing process or when the device is executed.
[0052] To avoid misunderstanding, the English names corresponding to the Chinese names in this application are uniformly introduced here. Prompt data (prompt), response data (response), reward model (RM model), reward value (reward generated by the reward model), actor model (Actor Model), critic model (Critic Model), reference model (ReferenceModel), RLHF (Reinforcement Learning from Human Feedback, reinforcement learning based on human feedback, referred to as "reinforcement learning" in this application), sentence level (sentence-level), token level (token-level), large language model (LLM, Large Language Model), supervised fine-tuning (SFT, Supervised Fine-tuning), reward hacking (Reward hacking, refers to the phenomenon that during the reinforcement learning training stage, the reward points increase abnormally and the model generation length increases abnormally).
[0053] Please refer to Figure 1 , Figure 1 FIG. 1 is a flow chart showing a method for training a large language model through reinforcement learning according to an embodiment of the present application. Figure 1 As shown, the method includes:
[0054] S101, obtaining prompt data and corresponding preference data for reinforcement learning training;
[0055] The preference data refers to the data reflecting the user's preferences in the response data generated corresponding to the prompt data, which can be achieved by marking the preference of several response data generated by the prompt data. The preference data is mainly used to train the reward model. In this embodiment, there is no limitation on the generation method of the prompt data and the corresponding preference data, and the following embodiment can be referred to.
[0056] S102, pre-training a reward model according to preference data;
[0057] Using preference data to pre-train the reward model can help the model better understand and generate high-quality reward signals. The model structure and specific training process of the reward model are not limited here, and you can refer to the implementation of related technologies.
[0058] S103, inputting the prompt data into the actor model for response training, calling the reward model to evaluate the response quality, calling the critic model to evaluate the response generation strategy, and generating training feedback according to the evaluation results of the reward model and the critic model;
[0059] In this application, the reinforcement learning framework of the Actor-Critic method is used to realize the reinforcement learning of large language models.
[0060] Input the preprocessed prompt data into the actor model to generate responses; select a trained reward model to evaluate the quality of the generated responses; input the responses generated by the actor model into the reward model to obtain a quality score; select a trained critic model to evaluate the effectiveness of the generation strategy; input the responses generated by the actor model and its generation process into the critic model to obtain a strategy score; combine the evaluation results of the reward model and the critic model to generate comprehensive training feedback. Training feedback is mainly used to adjust the parameters of the actor model and optimize the quality and strategy of the generated responses.
[0061] S104, performing intensive training on the actor model according to the training feedback, monitoring the capability improvement indicators during the intensive training, and synchronously updating the parameters of the critic model;
[0062] Obtain the evaluation results of the responses generated by the actor model from the reward model and the critic model, i.e., training feedback, and integrate the training feedback with the original prompt data and the generated responses to form a training dataset. Load or initialize the actor model, generate responses using the prompt data in the training dataset, calculate the loss value of the generated responses based on the loss function, update the model parameters through back propagation, and optimize the generated responses. Test the performance of the model on the validation set to ensure that the model performs well on new data, and then adjust hyperparameters such as learning rate and batch size based on the validation results to further optimize the model. Among them, reinforcement training can adopt different advantage function calculation methods and different reinforcement training methods, which are not limited here.
[0063] In order to determine whether the intensive training is normal, it is necessary to design corresponding monitoring indicators, which is also one of the key factors to ensure the stable progress of intensive training. In this embodiment, there is no limitation on the selection of specific ability improvement indicators. You can refer to the introduction of the following embodiments, which will not be repeated here.
[0064] The actor model is responsible for generating actions, while the critic model evaluates the quality of these actions. In reinforcement training, currently only the actor model is optimized and updated, but updating the actor model alone may lead to strategy instability because the actor model may make large changes in a short period of time. In order to prevent the critic model from gradually losing its ability to accurately assess the environment, this embodiment proposes that during reinforcement training of the actor model, both the actor model and the critic model need to be updated, that is, while updating the actor model, the critic model also needs to be updated in a coordinated manner, and the loss calculations of the two are coupled together.
[0065] Through synchronous updating, the critic model can more accurately evaluate the actions generated by the actor model, thereby providing more effective feedback, accelerating the learning process, and ensuring the coordination between the actor model and the critic model, reducing instability and oscillation caused by model mismatch. In addition, ensuring that the critic model is always in the best state can avoid the actor model from overfitting to a specific evaluation criterion. At the same time, the critic model can better adapt to changes in the environment and provide more accurate evaluation results, thereby helping the actor model converge to the optimal strategy faster.
[0066] S105. Determine the iteratively optimized actor model according to the capability improvement index and output it as the large language model.
[0067] The specified capability improvement index is used to evaluate and guide the iterative optimization process of the actor model. Through multiple iterations, the performance of the actor model is gradually improved so that its performance on specific tasks reaches the expected goal. Finally, the optimized actor model will be output as a large language model.
[0068] It should be noted that although the operations of the method of the present invention are described in a particular order in the drawings, this does not require or imply that the operations must be performed in this particular order or that all illustrated operations must be performed to achieve desired results.
[0069] Based on the above introduction, the method provided in this embodiment obtains the evaluation of the reward model and the critic model for the response training of the actor model as training feedback for reinforcement training. During the reinforcement training of the actor model, while updating the parameters of the actor model, the critic model is also updated in a coordinated manner. This can reduce the instability and oscillation caused by model mismatch, avoid overfitting of the actor model to a specific evaluation standard, ensure coordination between the actor model and the critic model, and at the same time, the critic model can better adapt to changes in the environment and provide more accurate evaluation results, thereby helping the actor model converge to the optimal strategy faster.
[0070] The above embodiments do not limit the specific selection of the ability improvement indicator. Currently, existing methods mostly use actor model loss, commentator model loss, and reward value as monitoring indicators. However, most of these methods observe the model to monitor the changes in indicator values on different batches of data, which are easily affected by data factors and cannot directly show the current status of the model. In order to avoid being affected by data, a monitoring indicator and a corresponding monitoring method are proposed in this embodiment. The process of monitoring the ability improvement indicator in intensive training in step S104 can be specifically referred to the following sub-steps.
[0071] Step S41, inputting the same prompt data into the actor model and the reference model respectively to generate response data;
[0072] The same prompt data {x1,…,x N Input the actor model and the reference model respectively, and use their respective sampling parameters to sample K times to obtain K sampling results, namely:
[0073] actor model:
[0074] reference model:
[0075] Step S42, inputting the prompt data and response data generated by the actor model and the reference model respectively into the reward model for quality evaluation to generate a reward value;
[0076] Send each prompt data and response data pair, i.e. (x, y), into the reward model to calculate the reward value r, and you can get the following results:
[0077] actor model:
[0078] reference model:
[0079] Step S43: Calculate the mean of the reward values corresponding to the actor model and the reference model respectively as a monitoring indicator.
[0080] By taking the average reward value a of the K response data generated by each prompt data, we can get the following result:
[0081] actor model:{a1,…,a N}
[0082]
[0083] Based on the above results, the following monitoring indicator values can be calculated:
[0084]
[0085] The monitoring indicator c reflects the improvement of the actor model relative to the reference model after intensive training, that is, whether the generated results can obtain higher reward values and whether they are more in line with human preferences. The indicator c avoids the influence of different batches of prompt data and improves the stability of the evaluation. As the intensive training proceeds, the value of the indicator c should gradually increase.
[0086] The enhanced training monitoring indicators provided in this embodiment can eliminate the interference caused by data and achieve accurate ability improvement evaluation.
[0087] In addition, the above embodiment does not limit the specific implementation method of the parameter synchronization update strategy of the critic model. In one embodiment, the synchronous update of the parameters of the critic model in step S104 can be implemented by the following steps.
[0088] Step S44, determining the loss value and gradient value of the actor model and the critic model in the reinforcement training;
[0089] Step S45, determine whether the loss value and / or gradient value is abnormal; if abnormal, train the model with the abnormality separately.
[0090] During training, observe the loss and gradient values of the actor model and the critic model. If one side (the actor model or the critic model) has an abnormality, it proves that the abnormal model is "behind" the other side in terms of training progress and cannot adapt and coordinate. At this time, the abnormal model can be trained separately, that is, only a smaller learning rate is used to update the abnormal model during the reinforcement training, and the other model is not updated. After several rounds of training, if the abnormal model can coordinate the training with the other model, normal reinforcement training is started, otherwise the abnormal model continues to be trained separately. If no abnormality occurs, the training process can continue without executing the above-mentioned separate training process, which is not limited here.
[0091] The above embodiments do not limit the setting of hyperparameters in reinforcement training. In order to improve the effectiveness of training, this embodiment proposes a hyperparameter setting rule. Specifically, the sampling parameters of the actor model supervised fine-tuning can be determined as the sampling parameters of the reinforcement training, and then the learning rate and KL divergence weight of the reinforcement training are dynamically adjusted according to the reward value output by the reward model.
[0092] In reinforcement training, the same sampling parameters as the supervised fine-tuning model are used when generating experience. Since supervised fine-tuning has been aligned with instruction fine-tuning, the quality of sampling can be guaranteed, which can ensure the effective implementation of reinforcement training. On the other hand, by using sampling parameters, the sampling space can be expanded, which helps the model explore and discover better strategies and improve the effectiveness of training.
[0093] For the KL divergence weight, in the reinforcement training, the optimization goal is to make the response data generated by the actor model obtain a larger reward R, which can be written as:
[0094] R(y|x)=r θ (y|x)-βKL[π(y|x)||π0(y|x)]
[0095] Where x is the prompt data, y is the response data, r is the output of the reward model, π and π0 represent the actor model and the reference model respectively, KL represents the KL divergence of the two distributions, and β is the weight of the KL divergence. When β is too large, the optimization focus becomes the KL divergence, which constrains the actor model and prevents it from moving away from the critic model. The advantage of this is that it can play a regular role, but the disadvantage is that the actor model will not change much, and the effect of reinforcement training is difficult to show. When β is too small, the optimization focus becomes the reward value. At this time, the reinforcement training effect is more obvious, and the reward value will increase significantly, but it may cause the reward value to increase excessively, resulting in reward value invasion. Therefore, in actual training, the appropriate weight parameter β is set according to the performance of the actor model (reward value changes) to achieve dynamic adjustment.
[0096] Regarding the setting of learning rate, a smaller learning rate helps stabilize the training, but it will significantly slow down the optimization speed. The model needs more rounds of iterative training to achieve reasonable results, which will consume more resources and time. A larger learning rate will significantly shorten the training process, but it will affect the stability of the training and easily lead to reward value invasion, affecting the model training effect. In actual training, you can choose a suitable learning rate according to the change of reward value. For example, if the reward value increases slowly, you can increase the learning rate appropriately; if the reward value rises too fast and is significantly too high, you should lower the learning rate. Specifically, you can set the learning rate to one tenth of the learning rate in the instruction fine-tuning stage, which is not limited here.
[0097] In the construction of prompt data and corresponding preference data, conventional methods can be referred to. In order to construct high-quality and diverse Chinese preference data and prompt data, and effectively ensure the stability and effectiveness of RLHF training, a data construction method is proposed in this embodiment, which specifically includes the following steps:
[0098] Step S11, collecting prompt data of each sub-field under the target field according to the preset proportion corresponding to each sub-field; wherein the target field includes a general field and an application scenario field;
[0099] When constructing prompt data, in order to ensure data diversity, the general field and application scenario field (for the financial security field, it can specifically include the security field and the financial field) are finely divided; prompt data for each sub-field after the split is collected according to a certain magnitude and proportion. This ensures data diversity on the one hand, and on the other hand, it also ensures data magnitude and proportion.
[0100] Among them, a general field of fine-grained splitting and data matching is Figure 2 As shown, it includes Open QA: 22.74%, Creative Writing: 15.16%, Safety: 15.79%, Finance: 8.42%, Information Extraction: 7.57%, Summarization: 7.58%, Rewrite: 7.58%, Translation: 7.58%, Math Computation: 7.58%. The data ratios in other fields can be set accordingly according to the actual application scenarios, which will not be elaborated here.
[0101] Furthermore, to ensure data quality, professionals can be hired to manually clean the data. Manual cleaning can delete or modify data with obvious errors, improve low-quality data, and retain correct data.
[0102] Step S12, calling the supervised fine-tuned actor model to generate response data corresponding to the prompt data;
[0103] To ensure the consistency of the distribution of reward model training data and test data, the model after supervised fine-tuning (SFT) is used to generate response data. The reward model test data is the output of the actor model, and the actor model is initialized by the supervised fine-tuning model.
[0104] Step S13: by comparing the response data, annotating the response data with preference information and preference strength, and generating preference data.
[0105] In practical applications, the current conventional rank labeling method is slow, and the consistency of labeling results from different labelers is low. Such data quality is poor, which is not conducive to the subsequent reward model training. To avoid the above problems, this embodiment proposes a pair labeling method to directly compare two response data and mark which one is more consistent with the prompt data. In addition to marking the preference information, it is also required to mark the preference strength. A preference labeling system page is as follows Figure 3As shown, preference labeling can be performed by directly labeling the reward value by discretizing the reward score into different gears. Of course, it can also be automatically labeled through a large model, which is not limited here.
[0106] A method for generating preference data is introduced above. Based on the above steps, a method for generating prompt data is further proposed, which specifically includes: extracting prompt data from the preference data; generating new prompt data; and combining the prompt data in the preference data with the new prompt data in proportion to generate prompt data.
[0107] The prompt data is mainly used for reinforcement training. The preference data constructed above contains the prompt data. This part of the prompt data can be directly used for reinforcement training. Since the reward model is trained using the preference data, the reward model will score the prompt data in the preference data more accurately (compared with other prompt data). This can avoid a series of problems caused by inaccurate reward model scoring in reinforcement training, such as reward value invasion and training non-convergence. On the basis of extracting the prompt data from the preference data, in order to further improve the generalization ability of the model, new prompt data can be further constructed according to the method described in the above content. The ratio of the prompt data of the preference data to the newly constructed prompt data can be 1:1, which is not limited here.
[0108] At present, the industry mostly uses word-level or sentence-level contrastive loss to train the reward model, and uses the reward model to initialize the critic model. However, when the reward model is actually trained with word-level contrastive loss, the training and testing targets are not consistent, which will cause the generalization of the reward model to decrease to a certain extent. When the reward model is used to initialize the critic model, the reinforcement training will be more stable. When the reward model is trained with sentence-level contrastive loss, the generalization performance of the reward model will be improved. However, when the reward model is used to initialize the critic model, the stability of the reinforcement training will decrease.
[0109] In order to improve the generalization of the reward model and ensure the stability of subsequent reinforcement training, a training method for a reward model is proposed in this embodiment. The process of pre-training the reward model according to the preference data in step S102 can be specifically implemented by referring to the following sub-steps:
[0110] Step S21, training the first model according to sentence-level contrast loss;
[0111] Step S22: using the first model to score the output of the actor model, and performing an indicator evaluation on the scoring process;
[0112] After each round of training, the corresponding reward model is stored and evaluated by indicators. In order to ensure the stability and effectiveness of RLHF training, one indicator is as follows:
[0113] (1) Test accuracy: Test accuracy reflects the rationality and generalization ability of the reward model’s scoring.
[0114] (2) Output scale: If the reward value output by the reward model is too large or too small, it will cause serious numerical problems and affect the stability of reinforcement training and the convergence of the model.
[0115] (3) Score interval: Based on the reward model, the average reward value of positive responses and the average reward value of negative responses in the test set are calculated, and the interval between the two average values is calculated, that is:
[0116]
[0117] Among them, x i is the prompt data in the test set, For positive response data, is the negative response data, r θ (y|x) is the reward model for training. The larger the positive and negative reward intervals, the better the robustness of the reward model, which is more conducive to improving the effect of reinforcement training.
[0118] Step S23, training the second model according to the contrast loss at the word level;
[0119] Step S24: Initialize the critic model using the second model;
[0120] Step S25, determine whether the result of the indicator evaluation meets the expectation; if so, use the first model as the reward model; if not, execute the step of training the first model according to the sentence-level contrast loss.
[0121] The reward model training method provided in this embodiment decouples the training task of the reward model, trains a model (first model) with sentence-level contrast loss, and uses the first model to score the output of the actor model, wherein the first model serves as the reward model; at the same time, trains another model (second model) with word-level contrast loss, and uses the second model to initialize the critic model. In this way, the accuracy of the reward model scoring can be guaranteed, and the stability of the reinforcement training can also be guaranteed, which can significantly improve the stability and effectiveness of the overall RLHF training.
[0122] It should be noted that the reward model can be trained with interval loss to improve robustness, or it can be modeled as a classification model instead of a regression model, which is not limited here.
[0123] Based on the introduction of the above embodiments, this application proposes a method for stable and effective reinforcement learning training in Chinese scenarios, such as Figure 4The figure shows a schematic diagram of the overall training process of RLHF, in which each link of reinforcement learning is optimized, including data construction, reward model training and reinforcement training.
[0124] Among them, data construction includes the composition, collection and cleaning methods of prompt data, the generation method of response data and the construction method of reinforcement training set. By optimizing the data construction process, the richness and quality of data are ensured. The reward model training stage includes the decoupling of reward model tasks and the use of word-level and sentence-level contrast loss, as well as the reward model selection method. The reinforcement training link includes appropriate monitoring indicators to strengthen the training process, coordinate training and hyperparameter setting. Through the optimization of each link, the stability of reinforcement learning training can be significantly improved, and the actual effect of training can be guaranteed, ensuring the stability and effectiveness of large language models in RLHF training.
[0125] Further references Figure 5 , which shows an exemplary structural block diagram of a reinforcement learning training device for a large language model according to an embodiment of the present application, which mainly includes:
[0126] A data acquisition unit 101, used to acquire prompt data and corresponding preference data for reinforcement learning training;
[0127] A reward model pre-training unit 102, used for pre-training a reward model according to preference data;
[0128] The training feedback generating unit 103 is used to input the prompt data into the actor model for response training, and call the reward model to evaluate the response quality, call the critic model to evaluate the response generation strategy, and generate training feedback according to the evaluation results of the reward model and the critic model;
[0129] The reinforcement training unit 104 is used to perform reinforcement training on the actor model according to the training feedback, monitor the ability improvement index during the reinforcement training, and synchronously update the parameters of the critic model;
[0130] The model output unit 105 is used to determine the iteratively optimized actor model according to the capability improvement index and output it as a large language model.
[0131] It should be understood that the units described in the above device are similar to those in the reference Figure 1 The steps in the method described above correspond to each other. Therefore, the operations and features described above for the method are also applicable to the device and the units contained therein, and will not be repeated here. The device can be pre-implemented in the browser or other security application of the server, or loaded into the browser or its security application of the server by downloading or the like. The corresponding units in the device can cooperate with the units in the server to implement the solution of the embodiment of the present application.
[0132] For the several units mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units described above can be embodied in one unit. On the contrary, the features and functions of one unit described above can be further divided into multiple units to be embodied.
[0133] It should be noted that for details not disclosed in the reinforcement learning training device for a large language model in the embodiment of the present application, please refer to the details disclosed in the above embodiments of the present application, which will not be repeated here.
[0134] Reference below Figure 6 , Figure 6 A schematic diagram of the structure of a computer system suitable for implementing a server of an embodiment of the present application is shown.
[0135] like Figure 6 As shown, the computer system includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage part 608 to the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation instructions of the system are also stored. The CPU 601, the ROM 602 and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0136] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage section 608 as needed.
[0137] In particular, according to an embodiment of the present application, the above reference flow chart Figure 1The described process can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-mentioned functions defined in the system of the present application are executed.
[0138] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium such as a computer-readable storage medium that can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0139] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operating instructions of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the aforementioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, the boxes represented by two connections can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operating instruction, or can be implemented with a combination of dedicated hardware and computer instructions.
[0140] The units or modules involved in the embodiments described in the present application may be implemented by software or hardware. The units or modules described may also be arranged in a processor. The names of these units or modules do not, in some cases, constitute limitations on the units or modules themselves.
[0141] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the server described in the above embodiment, or may exist independently without being assembled into the server. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the reinforcement learning training method for the large language model described in the present application.
[0142] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the aforementioned disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A reinforcement learning training method for a large language model, characterized in that: include: Obtain prompt data and corresponding preference data for reinforcement learning training; pre-training a reward model based on the preference data; Input the prompt data into the actor model for response training, call the reward model to evaluate the response quality, call the critic model to evaluate the response generation strategy, and generate training feedback according to the evaluation results of the reward model and the critic model; Performing intensive training on the actor model according to the training feedback, monitoring the capability improvement index during the intensive training, and synchronously updating the parameters of the critic model; An iteratively optimized actor model is determined according to the capability improvement index and output as a large language model.
2. The method according to claim 1, characterized in that Performing intensive training on the actor model according to the training feedback, and monitoring the capability improvement indicators during the intensive training, including: inputting the same prompt data into the actor model and the reference model respectively to generate response data; Inputting the prompt data and the response data generated by the actor model and the reference model, respectively, into the reward model for quality evaluation to generate a reward value; The mean of the reward values corresponding to the actor model and the reference model are calculated as monitoring indicators.
3. The method according to claim 1, characterized in that The synchronously updating the parameters of the critic model includes: Determine the loss value and gradient value of the actor model and the critic model in the enhanced training, and judge whether the loss value and / or gradient value are abnormal; If an exception occurs, the model with the exception is trained separately.
4. The method according to claim 1, characterized in that Performing reinforcement training on the actor model according to the training feedback includes: Determining sampling parameters for supervised fine-tuning of the actor model as sampling parameters for the reinforcement training; The learning rate and KL divergence weight of the reinforcement training are dynamically adjusted according to the reward value output by the reward model.
5. The method according to claim 1, characterized in that The obtaining of prompt data and corresponding preference data for RLHF training includes: Collect prompt data of each sub-field under the target field according to the preset proportion corresponding to each sub-field; wherein the target field includes a general field and an application scenario field; Calling the supervised fine-tuned actor model to generate response data corresponding to the prompt data; By comparing the response data, the response data is annotated with preference information and preference strength to generate preference data.
6. The method according to claim 1, characterized in that Pre-training a reward model based on the preference data includes: Train the first model based on sentence-level contrastive loss; Using the first model to score the output of the actor model, and performing indicator evaluation on the scoring process; Train the second model based on the token-level contrastive loss; Initialize the critic model using the second model; Determine whether the results of the indicator evaluation meet expectations; If achieved, the first model is used as a reward model; If not reached, execute the step of training the first model according to the sentence-level contrastive loss.
7. A reinforcement learning training device for a large language model, characterized in that: include: A data acquisition unit, used to acquire prompt data and corresponding preference data for reinforcement learning training; A reward model pre-training unit, used for pre-training a reward model according to the preference data; A training feedback generating unit, used for inputting the prompt data into the actor model for response training, calling the reward model to evaluate the response quality, calling the critic model to evaluate the response generation strategy, and generating training feedback according to the evaluation results of the reward model and the critic model; An enhanced training unit, used to perform enhanced training on the actor model according to the training feedback, monitor the capability improvement index during the enhanced training, and synchronously update the parameters of the critic model; The model output unit is used to determine the iteratively optimized actor model according to the capability improvement index and output it as a large language model.
8. A server comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Large language model training method and device and electronic equipment
CN120632448A
Large language model inference ability enhancement method based on reward adaptive reinforcement learning exploration
CN120671822A