Reinforcement learning training method and device of large language model, equipment and storage medium

By conducting reinforcement learning training on large language models, introducing generative models as reward models, and generating verb-level supervision information is generated, the problem of insufficient accuracy of pre-trained models in language generation tasks in specific fields is solved, and higher language generation accuracy and training efficiency are achieved.

CN120387495APending Publication Date: 2025-07-29BEIJING DAJIA INTERNET INFORMATION TECH CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510452120.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing pre-trained large language models perform poorly in language generation tasks in specific domains and cannot adapt and complete tasks with high accuracy.

Method used

By introducing a generative model as a reward model, the large language model is strengthened learning and training, and the word level supervision information is generated, the model output is corrected and optimized, and the accuracy of language generation tasks is improved.

Benefits of technology

In reinforcement learning training, the generative model can accurately indicate errors in the text generated by large language models, improve the accuracy of language generation tasks and the utilization efficiency of training data, and enhance the stability of the training process and the adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387495A_ABST
    Figure CN120387495A_ABST
Patent Text Reader

Abstract

The invention provides a reinforcement learning training method and device for a large language model, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps of obtaining first sample data, wherein the first sample data comprises a first question text and a first reply text output by a large language model for the first question text; through the generative model, on the basis of the first sample data, generating supervision information of the first reply text, the supervision information comprising a first corrected text obtained by correcting the first reply text and a reproduction probability of each lexical element in the first reply text, the reproduction probability being used for representing an occurrence probability of the corresponding lexical element in the first corrected text, the accuracy of the first corrected text is higher than that of the first reply text; and based on the first reply text and the supervision information of the first reply text, performing reinforcement learning training on the large language model. According to the technical scheme, reinforcement learning training can be carried out on the large language model, so that the accuracy of executing the language generation task by the large language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and particularly to a method, apparatus, device, and storage medium for reinforcement learning training of a large language model. Background Art

[0002] With the development of artificial intelligence technology, large language models have gradually been applied to language generation tasks in various fields. For example, in the medical field, a large language model can generate a diagnostic report or treatment advice based on the input relevant information. However, currently, a pre-trained large language model alone cannot exhibit excellent performance in a specific field. Therefore, how to enable a pre-trained large language model to adapt to the corresponding field and complete the corresponding language generation task is a technical problem that needs to be solved. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, device, and storage medium for reinforcement learning training of a large language model, which can improve the accuracy of the large language model in language generation tasks in the corresponding field by performing reinforcement learning training on the large language model. The technical solution of the present disclosure is as follows:

[0004] According to one aspect of the embodiments of the present disclosure, a method for reinforcement learning training of a large language model is provided, including:

[0005] Obtaining first sample data, where the first sample data includes a first question text and a first response text output by the large language model for the first question text;

[0006] Generating, by a generative model, supervision information for the first response text based on the first sample data, where the supervision information includes a first corrected text obtained by correcting the first response text and the recurrence probability of each token in the first response text, and the recurrence probability is used to represent the probability that the corresponding token appears in the first corrected text, and the accuracy of the first corrected text is higher than that of the first response text;

[0007] Performing reinforcement learning training on the large language model based on the first response text and the supervision information of the first response text.

[0008] According to another aspect of the embodiments of the present disclosure, a device for reinforcement learning training of a large language model is provided, including:

[0009] An acquisition unit configured to acquire first sample data, where the first sample data includes a first question text and a first response text output by the large language model for the first question text;

[0010] A generation unit, configured to generate supervision information of the first response text based on the first sample data through a generative model, where the supervision information includes a first corrected text obtained by correcting the first response text and the recurrence probability of each token in the first response text, the recurrence probability is used to represent the probability that the corresponding token appears in the first corrected text, and the accuracy rate of the first corrected text is higher than that of the first response text;

[0011] A training unit, configured to perform reinforcement learning training on the large language model based on the first response text and the supervision information of the first response text.

[0012] In some embodiments, the first sample data further includes the correct response to the first question text; the generation unit is configured to input the first prompt information and the first sample data into the generative model, the first prompt information is used to indicate the recurrence probability of each token in the first response text and correct the first response text; through the generative model, determine the recurrence probability of each token in the first response text, and based on the correct response to the first question text, correct the first response text to obtain the first corrected text.

[0013] In some embodiments, the generation unit is further configured to, when the first response text is a non-demand response, clip the recurrence probability of each token in the first response text according to a first clipping coefficient to obtain the reward score of each token, the first clipping coefficient is used to limit the value range of the reward score, and the reward score is used to represent the probability that the token is corrected; or, when the first response text is a demand response, clip the recurrence probability of each token in the first response text according to a second clipping coefficient to obtain the reward score of each token, the second clipping coefficient is used to limit the value range of the reward score; use the reward score of each token and the first corrected text as the supervision information of the first response text.

[0014] In some embodiments, the training unit is further configured to obtain second sample data, where the second sample data includes a second question text, the correct answer to the second question text, and a second answer text output by the large language model for the second question text, and the second answer text includes at least one reasoning step; through a teacher model, based on the second sample data, determine the reasoning step where the earliest error occurs in the second answer text, and start to correct the second answer text at the position where the reasoning step is located according to the minimum editing rule, and the minimum editing rule is used to limit the minimum edit distance between the text before correction and the text after correction; based on the second sample data and the second corrected text, construct training data; and based on the training data, train the generative model.

[0015] In some embodiments, the training unit is further configured to input the training data and second prompt information into the generative model, where the second prompt information is used to indicate determining the reasoning step where the earliest error occurs in the second answer text, and correcting the second answer text according to the minimum editing rule, based on the training data, determine the reasoning step where the earliest error occurs in the second answer text, and start to correct the second answer text at the position where the reasoning step is located according to the minimum editing rule to obtain a third corrected text; based on the difference between the second corrected text and the third corrected text, determine the training loss; and based on the training loss, train the generative model.

[0016] In some embodiments, the training unit is configured to determine an update gradient of the model parameters of the large language model through a reinforcement learning algorithm based on the supervision information of the first answer text; determine the regularization term loss of the large language model based on the first answer text and the first corrected text through a loss function; and update the parameters of the large language model based on the update gradient and the regularization term loss.

[0017] In some embodiments, the training unit is further configured to obtain the third sample data, where the third sample data includes a third question text and a third answer text output by the large language model after reinforcement learning training for the third question text; correct the third answer text in the third sample data through the generative model to obtain multiple fourth corrected texts; determine a target text from the multiple fourth corrected texts, where the edit distance between the target text and the third answer text is the smallest; construct training data based on the third sample data and the target text, and train the generative model based on the constructed training data.

[0018] In some embodiments, the training unit is further configured to, through the trained generative model, generate supervision information for the third response text based on the third sample data, where the supervision information includes a corrected text obtained by correcting the third response text and the recurrence probability of each token in the third response text, and the recurrence probability is used to represent the probability that the corresponding token appears in the corrected text; and based on the third response text and the supervision information of the third response text, perform reinforcement learning training on the large language model again.

[0019] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes:

[0020] One or more processors;

[0021] A memory for storing executable program code by the processor;

[0022] Wherein, the processor is configured to execute the program code to implement the reinforcement learning training method of the large language model as described above.

[0023] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the program code in the computer-readable storage medium is executed by a processor of an electronic device, enabling the electronic device to execute the reinforcement learning training method of the large language model as described above.

[0024] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, which implements the reinforcement learning training method of the large language model as described above when executed by a processor.

[0025] The embodiments of the present disclosure provide a reinforcement learning training method for a large language model, which can introduce a generative model as a reward model in the reinforcement learning training of the large language model. In the reinforcement learning training, the reward model can not only generate token-level supervision information for the response text output by the large language model, but also correct the response text output by the large language model, so as to accurately indicate the token-level errors in the text generated by the large language model, so as to better guide the large language model to complete the language generation task in the reinforcement learning training and improve the accuracy of the large language model in performing the language generation task.

[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.

[0028] Figure 1 It is a schematic diagram of the implementation environment of a reinforcement learning training method for a large language model shown according to an exemplary embodiment.

[0029] Figure 2 It is a flowchart of a reinforcement learning training method for a large language model shown according to an exemplary embodiment.

[0030] Figure 3 It is a flowchart of another reinforcement learning training method for a large language model shown according to an exemplary embodiment.

[0031] Figure 4 It is a schematic diagram of the reinforcement learning training process of a large language model shown according to an exemplary embodiment.

[0032] Figure 5 It is a schematic diagram of the training process of a generative model shown according to an exemplary embodiment.

[0033] Figure 6 It is a block diagram of a reinforcement learning training device for a large language model shown according to an exemplary embodiment.

[0034] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0035] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0037] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the sample data involved in this disclosure is obtained under sufficient authorization.

[0038] The following briefly introduces the terms that appear in the embodiments of this disclosure.

[0039] Large Language Model (LLM): An artificial intelligence model based on deep learning that can understand, generate and reason about natural language. Its core feature is that the model has a huge number of parameters (usually billions to trillions of model parameters) and is trained with a large amount of text data, thus possessing extensive language processing capabilities.

[0040] The key technologies involved in large language models include pre-training, reinforcement learning (RL), etc. Pre-training refers to performing unsupervised or self-supervised learning based on large-scale text data so that the large language model can automatically discover the statistical laws and potential structures in the text, thereby possessing certain language understanding and generation capabilities. Reinforcement learning refers to further training the pre-trained large language model based on the text data in the corresponding field, and introducing a reward mechanism during the training process to improve the performance of the model in language generation tasks in specific fields. The large language model trained by reinforcement learning can show excellent performance in language generation tasks in fields such as medical, legal, financial, or educational.

[0041] Figure 1 It is a schematic diagram of the implementation environment of a reinforcement learning training method for a large language model shown according to an exemplary embodiment. See Figure 1 , and this implementation environment specifically includes: a terminal 101 and a server 102. The terminal 101 can be connected to the server 102 through a wireless network or a wired network.

[0042] In some embodiments, the terminal 101 may be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, and a laptop portable computer. Optionally, the large language model in the embodiments of the present disclosure may be integrated in various electronic devices. For example, the large language model may be integrated in the terminal 101, or the large language model may also be integrated in the server 102. Correspondingly, the reinforcement learning training method for the large language model provided in the embodiments of the present application may be executed by the terminal 101, or may be executed by the server 102, or may also be executed by the terminal 101 and the server 102 in cooperation. The embodiments of the present disclosure do not limit this.

[0043] The terminal 101 may generally refer to one of multiple terminals. In this embodiment, the terminal 101 is taken as an example for illustration. Those skilled in the art can know that the number of the above terminals may be more or less. For example, the above terminals may be several, or the above terminals may be dozens or hundreds, or a larger number. The embodiments of the present disclosure do not limit the number and device type of the terminals.

[0044] The server 102 is at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of the above servers may be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 may further include other functional servers to provide more comprehensive and diverse services. In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing. The server 102 may be connected to the terminal 101 and other terminals through a wireless network or a wired network. Optionally, the number of the above servers may be more or less, and the embodiments of the present disclosure do not limit this.

[0045] Figure 2 is a flowchart of a reinforcement learning training method for a large language model shown according to an exemplary embodiment, as Figure 2 shown. Taking the method being executed by an electronic device as an example, the method includes the following steps:

[0046] In step 201, the electronic device obtains first sample data, where the first sample data includes a first question text and a first reply text output by the large language model for the first question text.

[0047] In the embodiments of the present disclosure, the first sample data may be text data in any field. The first sample data includes at least one first question text and a first reply text corresponding to each first question text. Among them, the first question text is the question text input to the large language model. For example, the first question text in the medical field may be "What are the common symptoms of a cold", and the first question text in the legal field may be "How to determine patent infringement", etc. The first reply text is the answer text generated and output by the large language model for the first question text. Since the pre-trained large language model has not undergone further reinforcement learning training, the first reply text output by the large language model may be a correct answer or an incorrect answer.

[0048] In step 202, the electronic device generates supervision information for the first reply text based on the first sample data through a generative model. The supervision information includes a first corrected text obtained by correcting the first reply text and the recurrence probability of each token in the first reply text.

[0049] In the embodiments of the present disclosure, the generative model is a text generation model and can be trained based on the sequence-to-sequence (Sequence-to-Sequence) method. During the reinforcement learning training process of the large language model, the generative model can be used as a reward model to generate token-level supervision information, and then represent the accuracy of each token output by the large language model through the supervision information. Among them, a token can also be called a Token, which is the smallest semantic unit in a language model.

[0050] In some embodiments, the supervision information includes a first corrected text obtained by correcting the first reply text output by the large language model. The accuracy rate of the first corrected text is higher than that of the first reply text. Among them, the accuracy rate of the reply text can represent the accuracy of the reply text, and the accuracy can reflect at least one of the factual correctness of the reply text, the relevance to the question text, the information completeness, and the logical correctness. Correspondingly, the accuracy rate of the reply text can be determined according to at least one of the above, and the embodiments of the present disclosure do not limit this. In other words, the generative model can correct the reply text with a lower accuracy rate output by the large language model into a reply text with a higher accuracy rate. In addition, the supervision information also includes the recurrence probability of each token in the first reply text, and the recurrence probability is used to represent the probability that the corresponding token appears in the first corrected text. Since correct tokens will be retained and incorrect tokens will be discarded during the process of correcting the reply text. Therefore, for any token, the higher the recurrence probability of the token, the higher the possibility that the token is retained, and the greater the probability that the token is a correct token. It can be seen that this supervision information can represent the accuracy of each token output by the large language model at the token level.

[0051] In step 203, the electronic device performs reinforcement learning training on the large language model based on the first response text and the supervision information of the first response text.

[0052] In the embodiments of the present disclosure, the electronic device performs multiple rounds of reinforcement learning training on the large language model according to the first response text and the supervision information output by the generative model. In each round of reinforcement learning training, the electronic device can determine the training loss of this round according to the first response text output by the large language model in this round of training and the supervision information of this first response text generated by the generative model, and then update the model parameters of the large language model according to the training loss. In reinforcement learning training, the large language model is also a policy model, which is used to generate a response text based on the question text; the generative model is a reward model, which is used to generate token-level supervision information according to the question text and the response text generated by the large language model to quantify the accuracy of each token in the response text output by the large language model. By performing multiple rounds of reinforcement learning training on the large language model, the model parameters of the large language model can be continuously optimized, so that the response text output by the large language model gradually approaches the corrected text in the supervision information. Optionally, when the training end condition is reached, the electronic device terminates the reinforcement learning training. Optionally, the training end condition includes but is not limited to that the supervision information tends to be stable, such as the change in the supervision information in several consecutive rounds is less than the first threshold; or the difference between the corrected text and the response text output by the large language model is less than the second threshold; or the number of iterations reaches a preset number, etc., and the embodiments of the present disclosure do not limit this.

[0053] The embodiments of the present disclosure provide a reinforcement learning training method for a large language model, which can introduce a generative model as a reward model in the reinforcement learning training of the large language model. In reinforcement learning training, the reward model can not only generate token-level supervision information for the response text output by the large language model, but also correct the response text output by the large language model, so as to accurately indicate the token-level errors in the text generated by the large language model, so as to better guide the large language model to complete the language generation task in the reinforcement learning training and improve the accuracy of the large language model in performing the language generation task.

[0054] In some embodiments, the first sample data further includes the correct answer to the first question text; based on the first sample data, the supervision information of the first response text is generated through the generative model, including:

[0055] Input the first prompt information and the first sample data into the generative model, where the first prompt information is used to indicate the recurrence probability of each token in the first response text and correct the first response text;

[0056] Using a generative model, determine the recurrence probability of each token in the first response text, and based on the correct response to the first question text, correct the first response text to obtain a first corrected text.

[0057] In the embodiments of the present disclosure, by introducing a generative model as a reward model, it is possible to generate token-level supervision information. Furthermore, during the reinforcement learning training process, it can accurately guide the large model to recognize errors in the generated content, accelerate the training convergence speed of the large language model, improve the utilization efficiency of training data, and enhance the stability of the training process.

[0058] In some embodiments, the method further includes:

[0059] When the first response text is a non-demand response, clip the recurrence probability of each token in the first response text according to a first clipping coefficient to obtain the reward score of each token. The first clipping coefficient is used to limit the value range of the reward score, and the reward score is used to represent the probability that the token is corrected; or,

[0060] When the first response text is a demand response, clip the recurrence probability of each token in the first response text according to a second clipping coefficient to obtain the reward score of each token. The second clipping coefficient is used to limit the value range of the reward score;

[0061] Use the reward score of each token and the first corrected text as the supervision information of the first response text.

[0062] In the embodiments of the present disclosure, by setting different clipping coefficients for the non-demand responses and demand responses output by the large language model, it is possible to specifically set the corresponding reward scores for each token in the response text, so as to ensure that during the reinforcement learning training process, the large language model pays more attention to the incorrect tokens in the non-demand responses and the correct tokens in the demand responses, that is, pays more attention to the relatively important tokens, in order to improve the utilization efficiency of training data and the effect of reinforcement learning training.

[0063] In some embodiments, the training process of the generative model includes:

[0064] Obtain second sample data, which includes a second question text, the correct response to the second question text, and a second response text output by the large language model for the second question text. The second response text includes at least one reasoning step;

[0065] Through a teacher model, based on the second sample data, determine the earliest reasoning step where an error occurs in the second response text, and start correcting the second response text at the position where the reasoning step is located according to the minimum edit rule. The minimum edit rule is used to limit the minimum edit distance between the text before correction and the text after correction;

[0066] Construct training data based on the second sample data and the second corrected text;

[0067] Train the generative model based on the training data.

[0068] In the embodiments of the present disclosure, by introducing a teacher model, relatively accurate training data can be efficiently generated, thus ensuring the training efficiency of the generative reward model.

[0069] In some embodiments, training the generative model based on the training data includes:

[0070] Input the training data and the second prompt information into the generative model, where the second prompt information is used to indicate the inference step where the error first appears in the second response text, and correct the second response text according to the minimum editing rule;

[0071] Through the generative model, based on the training data, determine the inference step where the error first appears in the second response text, and start to correct the second response text at the position where the inference step is located according to the minimum editing rule to obtain the third corrected text;

[0072] Determine the training loss based on the difference between the second corrected text and the third corrected text;

[0073] Train the generative model based on the training loss.

[0074] In the embodiments of the present disclosure, by training the generative model using the wrong answer rewriting task, the generative model can be made to have the wrong answer rewriting ability. The trained generative model can accurately determine the token-level errors in the response text generated by the large language model, that is, it can efficiently obtain relatively accurate fine-grained supervision information, so as to better guide the reinforcement learning training process of the large language model.

[0075] In some embodiments, the reinforcement learning training of the large language model based on the first response text and the supervision information of the first response text includes:

[0076] Through the reinforcement learning algorithm, determine the update gradient of the model parameters of the large language model based on the recurrence probability of each token in the first response text;

[0077] Through the loss function, determine the regularization loss of the large language model based on the first response text and the first corrected text;

[0078] Update the parameters of the large language model based on the update gradient and the regularization loss.

[0079] In the embodiments of the present disclosure, by performing reinforcement learning training on the large language model according to the supervision information at the token level, it is possible to more precisely guide the large language model to output tokens with a higher degree of accuracy based on the fine-grained supervision information. In addition, by introducing an imitation-based regularization term into the supervision information, it is possible to guide the large language model to learn the correct response text, avoid a large deviation between the response text output by the large language model and the correct response, thereby improving the stability of the reinforcement learning training.

[0080] In some embodiments, the method further includes:

[0081] Obtain third sample data, where the third sample data includes a third question text and a third response text output by the large language model after reinforcement learning for the third question text;

[0082] Use a generative model to correct the third response text in the third sample data to obtain multiple fourth corrected texts;

[0083] Determine a target text from the multiple fourth corrected texts, where the edit distance between the target text and the third response text is the smallest;

[0084] Construct training data based on the third sample data and the target text, and train the generative model based on the constructed training data.

[0085] In the embodiments of the present disclosure, by retraining the generative model using the method of rejection sampling, not only can the problem of lack of labeled data be avoided, but also it can be ensured that the generative model can adapt to the decision distribution of the large language model after reinforcement learning training. Furthermore, the generative model can continue to provide accurate supervision information for the next reinforcement learning training, thereby ensuring the efficiency and accuracy of the reinforcement learning training.

[0086] In some embodiments, the method further includes:

[0087] Use the trained generative model to generate supervision information for the third response text based on the third sample data. The supervision information includes the corrected text obtained by correcting the third response text and the recurrence probability of each token in the third response text. The recurrence probability is used to represent the probability that the corresponding token appears in the corrected text;

[0088] Perform reinforcement learning training on the large language model again based on the third response text and the supervision information of the third response text.

[0089] In the embodiments of the present disclosure, by jointly training the generative model during multiple reinforcement learning training processes of the large language model, it can be ensured that as the capabilities of the large language model gradually improve, the generative model can also adapt to the decision distribution of the large language model, thereby effectively improving the capabilities of the large language model after reinforcement learning training in complex downstream tasks.

[0090] The above-mentioned Figure 2 The following shows the process of a reinforcement learning training method for a large language model according to the present disclosure. Next, the reinforcement learning training solution for the large language model provided by the present disclosure will be further elaborated. Figure 3 It is a flowchart of another reinforcement learning training method for a large language model shown according to an exemplary embodiment. Taking the example that this method is executed by an electronic device, refer to Figure 3 , this method includes the following steps:

[0091] In step 301, the electronic device obtains first sample data, which includes a first question text, the correct answer to the first question text, and a first answer text output by the large language model for the first question text.

[0092] In the embodiment of the present disclosure, the electronic device obtains the first sample data in the same way as in the above step 201, and details are not described herein again.

[0093] In addition, it should be noted that the first sample data may also include the correct answer to the first question text. The correct answer may be an answer text manually labeled, or an answer text output by the teacher model for the first question text. The embodiment of the present disclosure does not limit this. Among them, the teacher model is a model with high performance and rich knowledge, and is often used to guide the learning and training of other language models, so that other language models can converge faster and improve the model generalization ability, etc. For example, the teacher model may be a large language model such as GPT-4, Claude2, etc.

[0094] In step 302, the electronic device inputs the first prompt information and the first sample data into the generative model. The first prompt information is used to indicate the recurrence probability of each token in the first answer text and correct the first answer text.

[0095] In the embodiment of the present disclosure, the generative model serves as the reward model in the reinforcement learning training, and is used to generate token-level supervision information according to the first sample data, so as to quantify the accuracy of the tokens output by the large language model through the supervision information. The first prompt information may be a prompt text, which is used to guide the generative model to generate supervision information according to the first sample data. For example, the first prompt information may be a Prompt (prompt word) input into the generative model.

[0096] Optionally, the first prompt message can guide the generative model to determine the recurrence probability of each token in the first response text, and with reference to the first question text and the correct answer to the first question text, correct the first response text to the correct answer text. Among them, the recurrence probability can also be called the rewriting probability, and the recurrence probability is used to represent the probability of the corresponding token appearing in the corrected response text. Since during the process of correcting the response text, correct tokens will be retained and incorrect tokens will be discarded. Therefore, for any token, the higher the recurrence probability of the token, the higher the possibility that the token will be retained, and the greater the probability that the token is a correct token.

[0097] By introducing a generative model as a reward model, it is possible to generate token-level supervision information. Furthermore, during the process of reinforcement learning training, it can accurately guide the large model to recognize the errors in the generated content, accelerate the training convergence speed of the large language model, improve the utilization efficiency of the training data, and enhance the stability of the training process.

[0098] In step 303, the electronic device determines the recurrence probability of each token in the first response text through the generative model, and corrects the first response text based on the correct answer to the first question text to obtain the supervision information of the first response text.

[0099] In the embodiments of the present disclosure, the supervision information includes the first corrected text and the recurrence probability of each token in the first response text. Among them, the first corrected text is obtained by correcting the first response text. Compared with the first response text, some incorrect tokens are removed from the first corrected text, and the corresponding correct tokens are added. Therefore, the accuracy rate of the first corrected text is higher than that of the first response text.

[0100] For example, the first response text includes 3 reasoning steps: A, B, C. If there is an error in reasoning step C, the generative model can refer to the first question text and its correct answer, generate a relatively accurate reasoning step D based on reasoning steps A and B, and output the first corrected text containing reasoning steps A, B, and D. In addition, the generative model can also generate the recurrence probability of each token in reasoning steps A, B, and C. Since reasoning steps A and B are correct and reasoning step C is incorrect, the recurrence probability of the tokens in reasoning steps A and B is higher than the recurrence probability of the tokens in reasoning step C.

[0101] In some embodiments, the generative model can also clip the recurrence probability of each token to obtain the reward score of each token. Among them, the reward score can also be used as part of the supervision information to guide the reinforcement learning training process of the large language model. The clipping process of the recurrence probability of the token will be described below through the following two situations.

[0102] Case 1: When the first response text is a non-demand response, the generative model clips the recurrence probability of each token in the first response text according to the first clipping coefficient to obtain the reward score for each token. Among them, a non-demand response refers to a response text that does not meet the preset requirements, such as a response text with factual errors or logical errors. The first clipping coefficient is used to limit the value range of the reward score, such as (-0.1, 0). The reward score is used to represent the probability that a token is corrected. The lower the reward score of a token, the lower the degree of return for optimizing the token. In the reinforcement learning training of the large language model, it is not inclined to optimize the token, and the probability that the token is corrected is also smaller.

[0103] For the non-demand response output by the large language model, by setting the first clipping coefficient, when determining the reward score of each token in the non-demand response, after the clipping process, the reward score of the correct token will become smaller or even approach 0, while the reward score of the wrong token will be higher. Then, in the process of reinforcement learning training, the correct token will not be optimized, and the large language model will pay more attention to optimizing the wrong token with a higher reward score.

[0104] Case 2: When the first response text is a demand response, the generative model clips the recurrence probability of each token in the first response text according to the second clipping coefficient to obtain the reward score for each token. Among them, a demand response is also a response that meets the preset requirements, such as a correct response output by the large language model. The second clipping coefficient is also used to limit the value range of the reward score, such as (0, 0.5).

[0105] For the demand response output by the large language model, by setting the second clipping coefficient, when determining the reward score of each token in the demand response, after the clipping process, the reward score of the wrong token will become smaller or even approach 0, while the reward score of the correct token will be higher. Then, in the process of reinforcement learning training, the large language model will not focus on optimizing the wrong token.

[0106] By setting different clipping coefficients for the demand response and non-demand response output by the large language model, it is possible to set corresponding reward scores for each token in the response text in a targeted manner, so as to ensure that in the process of reinforcement learning training, the large language model pays more attention to the wrong tokens in the non-demand response and the correct tokens in the demand response, that is, it pays more attention to the relatively important tokens, so as to improve the utilization efficiency of the training data and the effect of reinforcement learning training.

[0107] It should be noted that after obtaining the reward scores of each token in the first response text, the generative model can output the reward scores of each token and the first corrected text, which will be used as the supervision information of the first response text. In addition, the generative model can also determine whether the first response text is a required response according to the similarity between the first response text and the correct response in the first sample data. For example, when the similarity between the two reaches the preset similarity, it indicates that the first response text is a required response. At this time, the generative model can clip the reproduction probability of each token according to the second clipping coefficient; otherwise, it indicates that the first response text is a non-required response. At this time, the generative model can clip the reproduction probability of each token according to the first clipping coefficient to obtain a more accurate reward score for the token.

[0108] Taking the following formula (1) as an example, the process of the generative model calculating the reward score of a token will be described below.

[0109] Formula (1):

[0110] where q is the first question text, is the first response text output by the large language model for q, and t j is the j-th token in the first response text. s is the correct answer to the first question text. p R is the first prompt information. is the reward score of the j-th token in the first response text. CLIP(*, α, β) is a clipping function, and α, β are clipping coefficients. Through the clipping function, * can be clipped to the range (α, β). is the token t j is the reproduction probability, that is, the probability that the large language model generates the token t R given the prompt p j before the token t j in the first response text, the question q, the correct answer s, and the tokens before the token t

[0111] The above steps mainly introduce the process of generating supervision information for reinforcement learning training through the generative model. The reinforcement learning training process of the large language model will be described below through the following steps 304-306.

[0112] In step 304, the electronic device determines the update gradient of the model parameters of the large language model based on the supervision information of the first response text through a reinforcement learning algorithm.

[0113] In an embodiment of the present disclosure, the reinforcement learning algorithm may be a PPO (Proximal Policy Optimization) algorithm integrated with a reward mechanism. Through this reinforcement learning algorithm, the update gradient of the model parameters of the large language model can be determined based on the reward scores at the token level in the supervision information. Among them, the update gradient is used to represent the change rate of the model parameters in the parameter space, thereby indicating the update direction and amplitude of the model parameters, so as to optimize the model parameters.

[0114] Optionally, for the determination method of the update gradient of the model parameters of the large language model, refer to Formula II below.

[0115] Formula II:

[0116] Where θ is the model parameter of the large language model, is the update gradient of the model parameter θ. n is the number of the first sample data, q i is the first question text in the i-th first sample data, is the first response text for the first question text, t j is the j-th token in the first response text . is the reward score of the token t j . r(q i , t j ) is the importance sampling coefficient related to the first question text q i and the token t j . represents the probability that the large language model with model parameter θ generates the token t i based on the given question q j and the tokens before the token t j in the first response text. represents taking the gradient after taking the logarithm of the probability P θ .

[0117] In step 305, the electronic device determines the regularization term loss of the large language model through a loss function based on the first response text and the first correction text in the supervision information.

[0118] In the embodiments of the present disclosure, since the reinforcement learning training process tends to be unstable, a regularization term loss based on imitation is introduced. The electronic device can calculate the edit distance between the first response text and the first corrected text in the supervision information according to the dynamic programming algorithm, and obtain the regularization term loss of the large language model. Among them, the dynamic programming algorithm is a method for calculating the edit distance between texts. It gradually calculates the minimum number of operations required to convert one text to another by constructing a two-dimensional array (usually called a DP table). Each cell of the DP table represents the solution to a sub-problem, that is, the minimum number of operations required to convert a substring of the source text to a substring of the target text. The edit distance is used to represent the minimum number of operations required to convert one text into another. The operations include but are not limited to inserting, deleting, and replacing characters, etc.

[0119] Since only a small number of tokens in the original answer (the first response text) output by the large language model may cause the final answer to be incorrect, based on the difference between the original answer and the corrected answer (the first corrected text), the tokens deleted and replaced in the original answer can be determined. The above tokens can be considered as incorrect tokens. By weighting the incorrect tokens during the calculation of the regularization term loss, the large language model can be guided to learn relatively important tokens during the reinforcement learning training process. Optionally, the calculation method of the regularization term loss of the large language model is as shown in Formula III below.

[0120] Formula III:

[0121] Where θ is the model parameter of the large language model, is the regularization term loss of the large language model when the model parameter is θ. n is the number of the first sample data, q i is the first question text in the i-th first sample data, is the first response text for the first question text, t j is the j-th token in the first response text . w j is the weight of the token t j , which can be set according to actual needs. When the token t j is an incorrect token, w j is relatively high; when the token t j is a correct token, w j is relatively low.

[0122] In step 306, the electronic device performs reinforcement learning training on the large language model based on the updated gradient and the regularization term loss.

[0123] In the embodiments of the present disclosure, during any round of reinforcement learning training, the electronic device updates the parameters of the large language model through the backpropagation algorithm according to the update gradient and the regularization loss determined in this round of training. If the large language model after parameter update meets the training end condition, the electronic device terminates the reinforcement learning training; otherwise, it continues the next round of reinforcement learning training. Optionally, the training end condition includes, but is not limited to, the supervision information tending to be stable, such as the change in the supervision information in several consecutive rounds being less than the first threshold; or the difference between the corrected text and the response text output by the large language model being less than the second threshold; or the number of iterations reaching the preset number, etc. The embodiments of the present disclosure do not limit this.

[0124] By performing reinforcement learning training on the large language model according to the token-level supervision information, it is possible to more precisely guide the large language model to output tokens with a higher degree of accuracy based on the fine-grained supervision information. In addition, by introducing a regularization term based on imitation into the supervision information, it is possible to guide the large language model to learn the correct response text, avoid a large deviation between the response text output by the large language model and the correct response, thereby improving the stability of the reinforcement learning training.

[0125] The above steps mainly introduce the reinforcement learning training process of the large language model. Figure 4 is a schematic diagram of the reinforcement learning training process of a large language model shown according to an exemplary embodiment. As Figure 4 shown, during any round of reinforcement learning training, the electronic device inputs any first question text into the pre-trained large language model to obtain the first response text output by the large language model. Then, the electronic device inputs the first question text, the correct answer to the first question text, and the first response text as the first sample data into the generative model, and guides the generative model to generate the reward scores of each token in the first response text through the first prompt information, and corrects the first response text based on the correct response to the first question text to obtain the supervision information of the first response text. The supervision information includes the reward scores of each token in the first response text and the first corrected text obtained after correction. Then, the electronic device performs reinforcement learning training on the large language model based on the supervision information of the first response text.

[0126] The above steps mainly introduce the process of performing reinforcement learning training on the large language model according to the supervision information generated by the generative model. The training process of the generative model will be described below through the following steps (1)-(4).

[0127] (1) The electronic device obtains the second sample data. The second sample data includes the second question text, the correct response to the second question text, and the second response text output by the large language model for the second question text. The second response text includes at least one reasoning step.

[0128] Optionally, the first question text and the second question text belong to the same field. For example, in the scenario of training an intelligent interrogation model in the medical field, both the first question text and the second question text belong to the medical field. The large language model is a pre-trained model that has not undergone further reinforcement learning training.

[0129] (2) The electronic device determines, through the teacher model, the inference step where the error first appears in the second response text based on the second sample data, and corrects the second response text starting from the position where the inference step is located according to the minimum editing rule to obtain a second corrected text.

[0130] Among them, the teacher model can be a model with high performance and rich knowledge, and is often used to guide the learning and training of other language models, so that other language models can converge faster and improve the model generalization ability, etc. For example, the teacher model can be large language models such as GPT-4 and Claude2. The minimum editing rule is used to limit the minimum edit distance between the text before correction and the text after correction. Alternatively, the minimum editing rule is used to limit the correction of the second response text with the fewest number of editing times. The second corrected text obtained by correction can be regarded as the correct answer to the second question text.

[0131] (3) The electronic device constructs training data based on the second sample data and the second corrected text.

[0132] (4) The electronic device trains the generative model based on the training data.

[0133] In some embodiments, the electronic device can train the generative model based on the wrong answer rewriting task. Among them, the wrong answer rewriting task aims to correct the wrong tokens in the response text generated by the large language model with as few editing times as possible. Specifically, the wrong answer rewriting task can be decomposed into two subtasks: error localization and answer rewriting. In the error localization subtask, the generative model determines the inference step where the error first appears in the response text, while in the answer rewriting subtask, the generative model rewrites the answer starting from this position under the limitation of the minimum editing rule according to the previously located wrong inference step. Training the generative model based on these two subtasks of error localization and answer rewriting can ensure that the generative model has the ability to correct the response text following the minimum editing rule. The training process of the generative model will be described below.

[0134] The electronic device inputs the training data and the second prompt information into the generative model. Among them, the second prompt information is used to indicate the inference step where the error first appears in the second response text, and correct the second response text according to the minimum editing rule. In other words, the second prompt information can guide the generative model to complete the task of rewriting the wrong answer. Then, based on the training data, the generative model determines the inference step where the error first appears in the second response text, and starts to correct the second response text at the position where the inference step is located according to the minimum editing rule, obtaining the third corrected text. Then, the electronic device determines the training loss based on the difference between the second corrected text and the third corrected text. Optionally, the training loss is positively correlated with the edit distance between the second corrected text and the third corrected text. Then, the electronic device trains the generative model based on the training loss until the generative model meets the training end condition.

[0135] By training the generative model using the wrong answer rewriting task, the generative model can be enabled to have the ability to rewrite wrong answers. The trained generative model can accurately determine the token-level errors in the response text generated by the large language model, that is, it can efficiently obtain fine-grained supervision information with a relatively high degree of accuracy, so as to better guide the reinforcement learning training process of the large language model.

[0136] Figure 5 It is a schematic diagram of the training process of a generative model shown according to an exemplary embodiment. As Figure 5 shown, the electronic device inputs any second response text output by the large language model into the teacher model, and the teacher model corrects it to obtain the second corrected text. Then, the electronic device inputs the second prompt information, the second response text, and the second corrected text into the generative model to guide the generative model to complete the wrong answer rewriting task according to the above input data. Accordingly, the generative model can determine the inference step where the error first appears in the second response text, and start to correct the second response text at the position where the inference step is located according to the minimum editing rule, obtaining the third corrected text. Then, the electronic device trains the generative model based on the difference between the second corrected text and the third corrected text.

[0137] The above steps mainly introduce that first, the response text generated by the pre-trained large language model is used to train the generative model to obtain a reward model that provides fine-grained supervision information in the reinforcement learning training; then, the generative model is used to provide token-level supervision information to perform reinforcement learning training on the large language model.

[0138] It should be noted that, in order to ensure the training effect of the large language model, it is usually necessary to perform at least one reinforcement learning training on the large language model. Each reinforcement learning training includes at least one round of training process. However, after completing one reinforcement learning training, the decision-making of the large language model will change, resulting in the generative model being unable to continue providing accurate supervision information at this time. Therefore, the electronic device can retrain the generative model to adapt it to the decision distribution of the large language model, so as to provide accurate supervision information. The following describes the co-training process of the generative model and the large language model.

[0139] In some embodiments, the electronic device obtains third sample data. Among them, the third sample data includes a third question text and a third answer text output by the large language model after reinforcement learning training for the third question text. Optionally, the third question text belongs to the same field as the first question text. Then, the electronic device corrects the third answer text in the third sample data through the generative model to obtain multiple fourth corrected texts. The electronic device determines a target text from the multiple fourth corrected texts. Among them, the edit distance between the target text and the third answer text is the smallest. In other words, the similarity between the target text and the third answer text is the highest. Then, the electronic device constructs training data based on the third sample data and the target text, that is, uses the target text as the annotation data for the third answer text, and trains the generative model based on the constructed training data.

[0140] The above training method of the generative model can also be called the Rejection Sampling method, that is, using the answer text output by the generative model itself as the annotation data to train the generative model. By retraining the generative model using the Rejection Sampling method, not only can the problem of lack of annotation data be avoided, but also it can ensure that the generative model can adapt to the decision distribution of the large language model after reinforcement learning training. Furthermore, the generative model can continue to provide accurate supervision information for the next reinforcement learning training, thus ensuring the efficiency and accuracy of the reinforcement learning training.

[0141] In some embodiments, the electronic device can continue to provide token-level supervision information for the next reinforcement learning training process of the large language model through the trained generative model. For example, the electronic device can input the third sample data into the generative model, and the generative model generates the supervision information of the third response text. The supervision information includes the corrected text obtained by correcting the third response text and the recurrence probability of each token in the third response text, and the recurrence probability is used to represent the probability that the corresponding token appears in the corrected text. Optionally, the generative model can also clip the recurrence probability of each token to obtain the reward score of each token, and add the reward score of each token to the supervision information. Then, the electronic device performs reinforcement learning training on the large language model again based on the third response text and the supervision information of the third response text. Among them, the process of performing reinforcement learning training again is the same as the above-mentioned reinforcement learning training process, and will not be elaborated here. By jointly training the generative model during multiple reinforcement learning training processes of the large language model, it can be ensured that as the ability of the large language model gradually improves, the generative model can also adapt to the decision distribution of the large language model, thereby effectively improving the ability of the large language model after reinforcement learning training in complex downstream tasks. For example, in the question-and-answer task and the mathematical reasoning task, the reinforcement learning training method provided by the embodiments of the present disclosure can avoid the problem of the granularity of supervision information in traditional reinforcement learning and effectively improve the ability of the large language model to complete related tasks.

[0142] In addition, it should be noted that in the process of performing reinforcement learning training on the large language model in the above manner, an efficient fine-tuning algorithm can be further used to accelerate model training and improve model training efficiency. Optionally, the efficient fine-tuning algorithm includes but is not limited to: algorithms such as LoRA and Adapter. Through the efficient fine-tuning algorithm, the model training speed can be accelerated, and the computational resource overhead in the reinforcement learning training process can be reduced. In addition, in the reinforcement learning training, an instruction fine-tuning algorithm of the large language model can also be used to guide the large language model to imitate and learn to generate high-quality response texts through supervised data, and improve the accuracy of the large language model in downstream tasks in a specific field.

[0143] The embodiments of the present disclosure provide a reinforcement learning training method for a large language model, which can introduce a generative model as a reward model in the reinforcement learning training of the large language model. In the reinforcement learning training, the reward model can not only generate token-level supervision information for the response text output by the large language model, but also correct the response text output by the large language model, so as to accurately indicate the token-level errors in the text generated by the large language model, so as to better guide the large language model to complete the language generation task in the reinforcement learning training and improve the accuracy of the large language model in performing the language generation task.

[0144] Any combination of the above optional technical solutions can form an optional embodiment of the present disclosure, which will not be elaborated one by one here.

[0145] Figure 6 It is a block diagram of a reinforcement learning training device for a large language model shown according to an exemplary embodiment. As Figure 6 shown, the device includes: an acquisition unit 601, a generation unit 602, and a training unit 603.

[0146] The acquisition unit 601 is configured to acquire first sample data, where the first sample data includes a first question text and a first response text output by the large language model for the first question text;

[0147] The generation unit 602 is configured to generate supervision information for the first response text based on the first sample data through a generative model. The supervision information includes a first corrected text obtained by correcting the first response text and the recurrence probability of each token in the first response text. The recurrence probability is used to represent the probability that the corresponding token appears in the first corrected text, and the accuracy rate of the first corrected text is higher than that of the first response text;

[0148] The training unit 603 is configured to perform reinforcement learning training on the large language model based on the first response text and the supervision information of the first response text.

[0149] In some embodiments, the first sample data further includes the correct answer to the first question text; the generation unit 602 is configured to input the first prompt information and the first sample data into the generative model. The first prompt information is used to indicate the recurrence probability of each token in the first response text and correct the first response text; through the generative model, determine the recurrence probability of each token in the first response text, and based on the correct answer to the first question text, correct the first response text to obtain the first corrected text.

[0150] In some embodiments, the generation unit 602 is further configured to, when the first response text is a non-demand response, clip the recurrence probability of each token in the first response text according to a first clipping coefficient to obtain the reward score of each token. The first clipping coefficient is used to limit the value range of the reward score, and the reward score is used to represent the probability that the token is corrected; or, when the first response text is a demand response, clip the recurrence probability of each token in the first response text according to a second clipping coefficient to obtain the reward score of each token. The second clipping coefficient is used to limit the value range of the reward score; use the reward score of each token and the first corrected text as the supervision information of the first response text.

[0151] In some embodiments, the training unit 603 is further configured to obtain second sample data, where the second sample data includes a second question text, a correct answer to the second question text, and a second answer text output by the large language model for the second question text, and the second answer text includes at least one reasoning step; through the teacher model, based on the second sample data, determine the reasoning step where the error first appears in the second answer text, and start to correct the second answer text at the position where the reasoning step is located according to the minimum editing rule, and the minimum editing rule is used to limit the minimum edit distance between the text before correction and the text after correction; based on the second sample data and the second corrected text, construct training data; and based on the training data, train the generative model.

[0152] In some embodiments, the training unit 603 is further configured to input the training data and the second prompt information into the generative model, where the second prompt information is used to indicate determining the reasoning step where the error first appears in the second answer text, and correcting the second answer text according to the minimum editing rule, based on the training data, determine the reasoning step where the error first appears in the second answer text, and start to correct the second answer text at the position where the reasoning step is located according to the minimum editing rule, to obtain a third corrected text; based on the difference between the second corrected text and the third corrected text, determine the training loss; and based on the training loss, train the generative model.

[0153] In some embodiments, the training unit 603 is configured to determine the update gradient of the model parameters of the large language model through a reinforcement learning algorithm based on the supervision information of the first answer text; determine the regularization term loss of the large language model based on the first answer text and the first corrected text through a loss function; and update the parameters of the large language model based on the update gradient and the regularization term loss.

[0154] In some embodiments, the training unit 603 is further configured to obtain third sample data, where the third sample data includes a third question text and a third answer text output by the large language model after reinforcement learning training for the third question text; correct the third answer text in the third sample data through the generative model to obtain multiple fourth corrected texts; determine a target text from the multiple fourth corrected texts, where the edit distance between the target text and the third answer text is the smallest; construct training data based on the third sample data and the target text, and train the generative model based on the constructed training data.

[0155] In some embodiments, the training unit 603 is further configured to generate supervision information of the third response text based on the third sample data through the trained generative model, where the supervision information includes a corrected text obtained by correcting the third response text and the recurrence probability of each token in the third response text, and the recurrence probability is used to represent the probability that the corresponding token appears in the corrected text; and perform reinforcement learning training on the large language model again based on the third response text and the supervision information of the third response text.

[0156] The embodiments of the present disclosure provide a reinforcement learning training method for a large language model, which can introduce a generative model as a reward model in the reinforcement learning training of the large language model. In the reinforcement learning training, the reward model can not only generate token-level supervision information for the response text output by the large language model, but also correct the response text output by the large language model, so as to accurately indicate the token-level errors in the text generated by the large language model, so as to better guide the large language model to complete the language generation task in the reinforcement learning training and improve the accuracy of the large language model in performing the language generation task.

[0157] It should be noted that the reinforcement learning training device for the large language model provided in the above embodiments is only illustrated by the division of the above functional units. In practical applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the reinforcement learning training device for the large language model provided in the above embodiments and the embodiments of the reinforcement learning training method for the large language model belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.

[0158] Regarding the reinforcement learning training device for the large language model in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0159] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. Generally, the electronic device 700 includes a processor 701 and a memory 702.

[0160] The processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 701 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0161] The memory 702 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 702 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one program code, and the at least one program code is used to be executed by the processor 701 to implement the reinforcement learning training method of the large language model provided in the method embodiments of the present disclosure.

[0162] In some embodiments, the electronic device 700 may further optionally include: a peripheral device interface 703 and at least one peripheral device. The processor 701, the memory 702, and the peripheral device interface 703 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, and a power supply 708.

[0163] The peripheral device interface 703 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 701 and the memory 702. In some embodiments, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.

[0164] The radio frequency circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 704 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 704 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 704 can communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 704 may further include a circuit related to NFC (Near Field Communication), and the present disclosure does not limit this.

[0165] The display screen 705 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 705 is a touch display screen, the display screen 705 also has the ability to collect touch signals on or above the surface of the display screen 705. The touch signals can be input to the processor 701 as control signals for processing. At this time, the display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 705, which is provided on the front panel of the electronic device 700; in other embodiments, there can be at least two display screens 705, which are respectively provided on different surfaces of the electronic device 700 or are in a foldable design; in still other embodiments, the display screen 705 can be a flexible display screen, which is provided on the curved surface or the folding surface of the electronic device 700. Even further, the display screen 705 can also be set as an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 705 can be prepared using materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0166] The camera module 706 is used to collect images or videos. Optionally, the camera module 706 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the electronic device, and the rear camera is provided on the back of the electronic device. In some embodiments, there are at least two rear cameras, which can be any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to implement functions such as the combination of the main camera and the depth-of-field camera to achieve the background blurring function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions, or other combined shooting functions. In some embodiments, the camera module 706 can also include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.

[0167] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 701 for processing, or input to the radio frequency circuit 704 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 700. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 707 may further include a headphone jack.

[0168] The power supply 708 is used to supply power to each component in the electronic device 700. The power supply 708 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 708 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0169] Those skilled in the art can understand that Figure 7 the structure shown in

[0170] does not limit the electronic device 700, and may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.

[0171] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned reinforcement learning training method for the large language model.

[0172] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0173] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A reinforcement learning training method for large language models, characterized in that, The method includes: Obtain first sample data, where the first sample data includes a first question text and a first response text output by a large language model for the first question text; Through a generative model, based on the first sample data, generate supervision information for the first response text, where the supervision information includes a first corrected text obtained by correcting the first response text and the recurrence probability of each token in the first response text, and the recurrence probability is used to represent the probability of the corresponding token appearing in the first corrected text, and the accuracy rate of the first corrected text is higher than that of the first response text; Based on the first response text and the supervision information of the first response text, perform reinforcement learning training on the large language model.

2. The reinforcement learning training method for the large language model according to claim 1, wherein The first sample data further includes the correct answer to the first question text; The generating, through the generative model, the supervision information for the first response text based on the first sample data includes: Input a first prompt information and the first sample data into the generative model, where the first prompt information is used to indicate generating the recurrence probability of each token in the first response text and correct the first response text; Through the generative model, determine the recurrence probability of each token in the first response text, and based on the correct answer to the first question text, correct the first response text to obtain the first corrected text.

3. The reinforcement learning training method for the large language model according to claim 2, wherein The method further includes: In the case where the first response text is a non-demand response, clip the recurrence probability of each token in the first response text according to a first clipping coefficient to obtain the reward score of each token, where the first clipping coefficient is used to limit the value range of the reward score, and the reward score is used to represent the probability of the token being corrected; or, In the case where the first response text is a demand response, clip the recurrence probability of each token in the first response text according to a second clipping coefficient to obtain the reward score of each token, where the second clipping coefficient is used to limit the value range of the reward score; Use the reward score of each token and the first corrected text as the supervision information of the first response text.

4. The reinforcement learning training method for the large language model according to claim 1, wherein, The training process of the generative model includes: Obtain second sample data, where the second sample data includes a second question text, the correct answer to the second question text, and a second response text output by the large language model for the second question text, and the second response text includes at least one reasoning step; Through a teacher model, based on the second sample data, determine the earliest reasoning step with an error in the second response text, and start correcting the second response text at the position where the reasoning step is located according to the minimum editing rule, where the minimum editing rule is used to limit the minimum edit distance between the text before correction and the text after correction; Based on the second sample data and the second corrected text, construct training data; Based on the training data, train the generative model.

5. The reinforcement learning training method for the large language model according to claim 4, characterized in that, The training the generative model based on the training data includes: Input the training data and the second prompt information into the generative model, where the second prompt information is used to indicate the inference step where the error first appears in the second response text, and correct the second response text according to the minimum editing rules; Through the generative model, based on the training data, determine the inference step where the error first appears in the second response text, and start correcting the second response text at the position where the inference step is located according to the minimum editing rules to obtain a third corrected text; Based on the difference between the second corrected text and the third corrected text, determine the training loss; Based on the training loss, train the generative model.

6. The reinforcement learning training method for the large language model according to claim 1, wherein The reinforcement learning training of the large language model based on the first response text and the supervision information of the first response text includes: Through the reinforcement learning algorithm, based on the supervision information of the first response text, determine the update gradient of the model parameters of the large language model; Through the loss function, based on the first response text and the first corrected text, determine the regularization term loss of the large language model; Based on the update gradient and the regularization term loss, update the parameters of the large language model.

7. The reinforcement learning training method for the large language model according to claim 1, wherein The method further includes: Obtain the third sample data, where the third sample data includes a third question text and a third response text output by the large language model after reinforcement learning training for the third question text; Through the generative model, correct the third response text in the third sample data to obtain a plurality of fourth corrected texts; Determine a target text from the plurality of fourth corrected texts, where the edit distance between the target text and the third response text is the smallest; Construct training data based on the third sample data and the target text, and train the generative model based on the constructed training data.

8. The reinforcement learning training method for the large language model according to claim 7, wherein The method further includes: Through the trained generative model, based on the third sample data, generate supervision information for the third response text, where the supervision information includes a corrected text obtained by correcting the third response text and the recurrence probability of each token in the third response text, and the recurrence probability is used to represent the probability that the corresponding token appears in the corrected text; Based on the third response text and the supervision information of the third response text, perform reinforcement learning training on the large language model again.

9. A reinforcement learning training device for a large language model, characterized in that, The device includes: An acquisition unit configured to acquire first sample data, where the first sample data includes a first question text and a first response text output by a large language model for the first question text; A generation unit configured to, through a generative model, based on the first sample data, generate supervision information for the first response text, where the supervision information includes a first corrected text obtained by correcting the first response text and the recurrence probability of each token in the first response text, and the recurrence probability is used to represent the probability that the corresponding token appears in the first corrected text, and the accuracy of the first corrected text is higher than that of the first response text; A training unit, configured to perform reinforcement learning training on the large language model based on the first response text and the supervision information of the first response text.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory for storing program code executable by the processor; Wherein, the processor is configured to execute the program code to implement the reinforcement learning training method of the large language model according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the reinforcement learning training method of the large language model according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the reinforcement learning training method of the large language model according to any one of claims 1 to 8.

Citation Information

Cited By

  • Two-stage large model cognitive enhancement method, system and device and storage medium

    CN121390300A

  • Inference model training method and device, equipment, medium and product

    CN122154840A