Reinforcement learning model training method and device, electronic equipment and storage medium

By generating the preference context vector of the review group and constructing an inter-group alignment loss function, the problem of inconsistent scoring by the review group in reinforcement learning model training was solved, and the stability and consistency of training were improved.

CN120633761AActive Publication Date: 2025-09-12北京衔远有限公司 +1

Patent Information

Application Number
CN202511134687.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-09-12
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

The training stability of reinforcement learning models in existing technologies is poor, mainly due to the unreasonable construction of reward models and loss functions, which leads to inconsistent scoring and unstable standards among review groups.

Method used

By obtaining sample data from multiple review groups, we generate preference context vectors for each review group. Based on these vectors, we build a reward model for the reinforcement learning pre-training model, and construct a loss function that includes an inter-group alignment loss term to characterize the differences in evaluations of the same sample data by different review groups.

Benefits of technology

It effectively solves the reward learning bias caused by differences between reward model groups and improves the stability and consistency of reinforcement learning model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633761A_ABST
    Figure CN120633761A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a reinforcement learning model training method and device, electronic equipment and a storage medium. According to the method, N pieces of sample data are obtained firstly, the sample data comprise questions, candidate answers, evaluation of candidate answers by evaluation groups and evaluation group identifiers, then preference context vectors of each evaluation group are generated based on the N pieces of sample data, a reward model of a reinforcement learning pre-training model is constructed at least based on the preference context vectors, and the reward model of the reinforcement learning pre-training model is obtained. A loss function at least comprising inter-group alignment loss items is constructed based on the reward model, so that the loss function can represent evaluation differences of different review groups on the same sample data, and finally the constructed loss function is used to train the reinforcement learning pre-training model to obtain a trained reinforcement learning model. The reward learning deviation caused by the difference between the reward model groups can be effectively solved, and the stability of reinforcement learning model training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a reinforcement learning model training method, device, electronic device and storage medium. Background Art

[0002] In the process of combining reinforcement learning with RLHF (Reinforcement Learning from Human Feedback), reward models are typically trained by assigning preference scores to multiple candidate outputs through subjective human labeling or evaluation. Current reward model training suffers from two issues that affect training stability: 1) Due to cognitive biases, cultural differences, and inconsistent evaluation criteria among different annotators regarding non-deterministic tasks, inconsistent scoring between samples can occur, introducing high variance into the aligned data. This inconsistency directly impacts the preference function learned by the reward model, causing instability or bias in reinforcement learning tasks.

[0003] 2) During reward model training, there are deviations in the scoring standards between different sampling groups. In other words, the reward model lacks a stable standard during different rounds of sampling. This situation will also lead to an imbalance in the reward scores given by the reward model after training.

[0004] Therefore, a method is needed to introduce the review group's scoring information and the model's multi-round scoring context information during the training process to structurally model and weaken the differences between scoring groups. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a reinforcement learning model training method, device, electronic device and storage medium to solve the problem of poor training stability of the reinforcement learning model in the prior art due to the unreasonable construction of the reward model and loss function.

[0006] A first aspect of an embodiment of the present application provides a reinforcement learning model training method, comprising: Obtain N sample data, each sample data including at least a question, a candidate answer to the question from the reinforcement learning pre-trained model, an evaluation of each candidate answer by a target review group, and an identifier of the target review group; the target review group is any one of the M review groups, where M and N are both positive integers; Based on N sample data, generate the preference context vector of each review group; Constructing a reward model for the reinforcement learning pre-trained model based on at least the preference context vectors of each review group; Construct a loss function based on the reward model. The loss function includes at least an inter-group alignment loss term, which is used to characterize the differences in evaluations of the same sample data by different review groups. The reinforcement learning pre-training model is trained based on the loss function to obtain a trained reinforcement learning model.

[0007] A second aspect of the embodiments of the present application provides a reinforcement learning model training device, comprising: an acquisition module configured to acquire N sample data, each sample data including at least a question, a candidate answer to the question by the reinforcement learning pre-trained model, an evaluation of each candidate answer by a target review group, and an identifier of the target review group; the target review group is any one of the M review groups, where M and N are both positive integers; A generation module is configured to generate a preference context vector for each review group based on N sample data; A building module configured to build a reward model of the reinforcement learning pre-training model based on at least the preference context vectors of each review group; The building module is further configured to construct a loss function based on the reward model, where the loss function at least includes an inter-group alignment loss term, where the inter-group alignment loss term is used to characterize the difference in evaluations of the same sample data by different review groups; The training module is configured to train the reinforcement learning pre-training model based on the loss function to obtain a trained reinforcement learning model.

[0008] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0009] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0010] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: the embodiments of the present application first obtain N sample data, which include questions, candidate answers, the review group's evaluation of the candidate answers and the review group identification, and then generate a preference context vector for each review group based on the N sample data, construct a reward model of the reinforcement learning pre-training model at least based on the preference context vector, and then construct a loss function based on the reward model that includes at least an inter-group alignment loss term, so that the loss function can characterize the evaluation differences of different review groups on the same sample data, and finally use the constructed loss function to train the reinforcement learning pre-training model to obtain a trained reinforcement learning model, which can effectively solve the reward learning bias caused by differences between reward model groups and improve the stability of reinforcement learning model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 This is a flow chart of a reinforcement learning model training method provided in an embodiment of the present application.

[0013] Figure 2 4 is a flowchart of a method for generating a preference context vector for each review group based on N sample data provided in an embodiment of the present application.

[0014] Figure 3 This is a flowchart of a method for determining a preference context vector of a target review group based on a token encoding vector provided in an embodiment of the present application.

[0015] Figure 4 It is a flowchart of the method for constructing a reward model provided in an embodiment of the present application.

[0016] Figure 5 This is a flowchart of a method for determining contextual representation of target sample data provided in an embodiment of the present application.

[0017] Figure 6 This is a flowchart of a method for determining an inter-group alignment loss term provided in an embodiment of the present application.

[0018] Figure 7 This is a schematic diagram of a reinforcement learning model training device provided in an embodiment of the present application.

[0019] Figure 8 Schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0021] A reinforcement learning model training method and device according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0022] As mentioned above, current reward model training suffers from issues that affect the stability of reinforcement learning model training. Related technologies attempt to address this issue using methods such as normalization, panel weighting, and Bayesian modeling. While these methods can alleviate the problem to a certain extent, they still suffer from poor generalization and significant reasoning ambiguity in non-deterministic tasks, such as subjective judgments of the quality of answers.

[0023] In view of this, an embodiment of the present application provides a reinforcement learning model training method, which first obtains N sample data, which includes questions, candidate answers, the review group's evaluation of the candidate answers and the review group identification, and then generates a preference context vector for each review group based on the N sample data, and constructs a reward model of the reinforcement learning pre-training model based on at least the preference context vector, and then constructs a loss function based on the reward model that includes at least an inter-group alignment loss term, so that the loss function can characterize the evaluation differences of different review groups on the same sample data. Finally, the reinforcement learning pre-training model is trained using the constructed loss function to obtain a trained reinforcement learning model, which can effectively solve the reward learning bias caused by differences between reward model groups and improve the stability of reinforcement learning model training.

[0024] Figure 1 This is a flow chart of a reinforcement learning model training method provided in the embodiment of the present application. Figure 1 As shown, the method includes the following steps: In step S101, N sample data are obtained, each sample data including at least a question, a candidate answer to the question by the reinforcement learning pre-training model, an evaluation of each candidate answer by the target review group, and an identifier of the target review group.

[0025] The target review group is any review group among the M review groups, and both M and N are positive integers.

[0026] In step S102 , a preference context vector of each review group is generated based on N sample data.

[0027] In step S103, a reward model of the reinforcement learning pre-training model is constructed based on at least the preference context vector of each review group.

[0028] In step S104, a loss function is constructed based on the reward model, where the loss function at least includes an inter-group alignment loss term.

[0029] Among them, the inter-group alignment loss term is used to characterize the evaluation differences of different review groups on the same sample data.

[0030] In step S105, the reinforcement learning pre-training model is trained based on the loss function to obtain a trained reinforcement learning model.

[0031] In some embodiments of the present application, the method may be executed by a server, or by a terminal device with certain computing capabilities.

[0032] In some embodiments of the present application, N sample data may be first obtained, each sample data including at least a question, a candidate answer to the question from a reinforcement learning pre-trained model, an evaluation of each candidate answer by a target review group, and a target review group identifier. The review group refers to an annotation group or annotator who scores or selects a preference for each candidate answer. Different review groups have different style preferences.

[0033] In some embodiments of the present application, a preference context vector for each review group may be generated based on the N sample data. The preference context vector may be used to characterize the style preferences of each review group. Further, the generated preference context vector may be associated with the review group identifier of each review group.

[0034] In certain embodiments of the present application, a reward model for a reinforcement learning pre-trained model may be constructed based at least on the preference context vectors of each review group. In one example, a reinforcement learning pre-trained model may be prepared, and a reward model for the pre-trained model may be constructed based on the preference context vectors of each review group.

[0035] In some embodiments, a loss function can also be constructed based on the reward model. This constructed loss function includes at least an inter-group alignment loss term, which is used to characterize the differences in evaluations of the same sample data by different review groups. The same sample data refers to the same question and the candidate answers to that question from the reinforcement learning pre-trained model.

[0036] After constructing the loss function, the reinforcement learning pre-training model can be trained based on the loss function to obtain the trained reinforcement learning model.

[0037] According to the technical solution provided in the embodiments of the present application, by first obtaining N sample data, the sample data includes questions, candidate answers, the evaluation of the candidate answers by the review group and the review group identification, and then generating a preference context vector for each review group based on the N sample data, constructing a reward model of the reinforcement learning pre-training model at least based on the preference context vector, and then constructing a loss function based on the reward model that includes at least an inter-group alignment loss term, so that the loss function can characterize the evaluation differences of different review groups on the same sample data, and finally using the constructed loss function to train the reinforcement learning pre-training model to obtain a trained reinforcement learning model, which can effectively solve the reward learning bias caused by differences between reward model groups and improve the stability of reinforcement learning model training.

[0038] Figure 2: is a flow chart of a method for generating a preference context vector for each review group based on N sample data provided by an embodiment of the present application. Figure 2 As shown, the method includes the following steps: In step S201 , all sample data of the target review group are obtained from N sample data.

[0039] In step S202, the sample data of the target review group is word-gram encoded to obtain a word-gram encoding vector of the target review group.

[0040] In step S203 , the preference context vector of the target review group is determined based on the word-unit encoding vector.

[0041] In some embodiments of the present application, when generating a preference context vector for each review group based on sample data, the entire sample data for the target review group can be first obtained. The sample data for the target review group can then be tokenized to obtain a token encoding vector for the target review group. Finally, the preference context vector for the target review group can be determined based on the token encoding vector.

[0042] By traversing the M review groups, the preference context vector of each review group can be obtained.

[0043] Figure 3 Schematic diagram of the method for determining the preference context vector of the target review group based on the token encoding vector provided in the embodiment of the present application. Figure 3 As shown, the method includes the following steps: In step S301, embedding encoding is performed on the word element encoding vector to obtain the preference context vector of the target review group.

[0044] In step S302, the historical word element encoding vector of the target review group is trained to obtain the preference context vector of the target review group.

[0045] In some embodiments of the present application, the preference context vector of the target review group can be obtained by embedding the token encoding vector. Alternatively, the preference context vector of the target review group can be obtained by training the historical token encoding vectors of the target review group.

[0046] In one example, a method for dynamically adjusting embedding vectors, such as a learnable embedding table method, can be used to embed the token encoding vector to obtain the target review group's preference context vector. Alternatively, historical sample data corresponding to the target review group can be collected. This historical sample data can be part or all of the N sample data. Sample data in the historical sample data that has the target review group's identifier can then be token encoded separately. Each token encoding vector can then be trained to obtain the target review group's preference context vector.

[0047] Figure 4 This is a flow chart of the method for constructing a reward model provided in the embodiment of the present application. Figure 4 As shown, the method includes the following steps: In step S401 , candidate answers in target sample data and context representation of the target sample data are obtained.

[0048] The target sample data is any sample data.

[0049] In step S402, the candidate answer and the context representation are concatenated to obtain a word element encoding of the concatenated text.

[0050] In step S403, the word unit encoding of the spliced ​​text is spliced ​​with the preference context vector of the review group corresponding to the target sample data to obtain a spliced ​​feature.

[0051] In step S404, the splicing feature is used as an independent variable of the preset network prediction function to obtain a reward model for the target sample data.

[0052] In some embodiments of the present application, when constructing a reward model, the candidate answers in the target sample data and the context representation of the target sample data can be obtained first. In one example, the candidate answers in the target sample data can be represented by r i Indicates that the context of the target sample data is represented by p t Indicates . Where i and t are both positive integers. The value of i is less than or equal to the number of candidate answers in the target sample data. The value of t is determined according to actual needs and represents the data round before the target sample data used for reference.

[0053] In some embodiments, the candidate answer and the context representation of the target sample data may be concatenated and the concatenated text may be tokenized. In one example, the token encoding may be represented as token (r i *p t ), where “*” is a text concatenation character.

[0054] In some embodiments of the present application, the preference context vector of the review group corresponding to the target sample data can also be obtained. For example, the review group identifier in the target sample data can be obtained, and the corresponding preference context vector can be determined based on the review group identifier. In one example, the preference context vector of the review group corresponding to the target sample data can be expressed as g i express.

[0055] The token encoding of the spliced ​​text can be concatenated with the preference context vector of the review group corresponding to the target sample data to obtain the concatenated feature. In one example, the concatenated feature can be expressed as [token (r i *p t );g i ] indicates that “;” is a characteristic concatenation character.

[0056] In some implementations, the splicing features are used as independent variables of a preset network prediction function to obtain a reward model for the target sample data. In one example, if the preset network prediction function is f θ (), then the reward model of the target sample data is S(r i , p t , g i ) = f θ ([token(r i *p t );g i ]).

[0057] Figure 5 FIG is a flow chart of a method for determining the context representation of target sample data provided by an embodiment of the present application. Figure 5 As shown, the method includes the following steps: In step S501 , the evaluation window of t rounds before the target sample data is obtained.

[0058] The evaluation window includes questions from each round, candidate answers to the questions from the reinforcement learning pre-training model, and the evaluation of each candidate answer by the target review group.

[0059] In step S502 , all items in the evaluation window of round t are subjected to text concatenation to obtain a contextual representation of the target sample data.

[0060] In some embodiments of the present application, when determining the contextual representation of the target sample data, the evaluation window of t rounds before the target sample data can be first obtained, and the evaluation window includes the questions of each round, the candidate answers of the reinforcement learning pre-trained model to the questions, and the evaluation of each candidate answer by the target review group.

[0061] Then, all items in the evaluation window of round t are concatenated to obtain the contextual representation of the target sample data.

[0062] Figure 6 : is a flow chart of the method for determining the inter-group alignment loss term provided in the embodiment of the present application. Figure 6 As shown, the method includes the following steps: In step S601 , a first reward value and a second reward value of target sample data are calculated respectively.

[0063] The first reward value and the second reward value correspond to the same context representation of candidate answers and target sample data, and correspond to preference context vectors of different review groups.

[0064] In step S602, a second normal form calculation is performed on the difference between the first reward value and the second reward value to obtain a candidate group alignment loss term.

[0065] In step S603, all reward values ​​of the preference context vectors corresponding to the same candidate answer and context representation of the target sample data and corresponding to different review groups in the target sample data are traversed, and the sum of all candidate inter-group alignment loss items is determined as the inter-group alignment loss item.

[0066] In some embodiments of the present application, when determining the inter-group alignment loss term, the first reward value and the second reward value of the target sample data can be respectively calculated. The first reward value and the second reward value correspond to the same context representation of the candidate answer and the target sample data, and correspond to different preference context vectors of the review groups. For example, the first reward value can be recorded as S(r i , p t , g j ), the second reward value can be recorded as S(r i , p t , g k ), j and k are both positive integers less than or equal to M, and j is not equal to k.

[0067] The difference between the first reward value and the second reward value can be calculated in the second normal form to obtain the alignment loss term between candidate groups. In an example, g j and g k The candidate group alignment loss term is || S (r i , p t , g j )-S(r i , p t , g k )|| 2 .

[0068] Finally, all reward values ​​of the preference context vectors corresponding to the same candidate answer and the context representation of the target sample data and corresponding to different review groups in the target sample data can be traversed to determine the sum of the alignment loss terms between all candidate groups as the inter-group alignment loss term.

[0069] In an example, if the inter-group alignment loss term is recorded as L align ,but .

[0070] In some embodiments of the present application, the loss function may further include an intra-group consistency loss term, which is used to characterize the consistency between different candidate answers to the same question. The intra-group consistency loss term may be determined as follows: Determined using the within-group normalized relative loss method; or Determined using a metric learning algorithm; or Determined using the within-group relative strategy optimization algorithm.

[0071] In other words, the present embodiment effectively addresses the reward learning bias caused by differences in the standards of the reward model groups by incorporating review group scoring information and a model multi-round scoring context-aware modeling approach into reward model training. By jointly modeling the outputs of different review groups and scoring context information, this method guides the reward model to learn a preference function that is consistent within the standard and compatible across groups during the training phase.

[0072] During the data preparation phase, you can collect scoring data from multiple review groups on the same task sample. Each sample data piece can include a question, a candidate answer to the question from the reinforcement learning pre-trained model, the target review group's evaluation of each candidate answer, and the target review group's identifier.

[0073] Next, we can embed the review group to generate the preference context vector g corresponding to the review group identifier. i , the embedding generation can be based on the learnable embedding table (dynamically adjusted embedding) method, or trained from the historical scores of the review group.

[0074] Then build a reward model. The input of the reward model can include the candidate answer r in the target sample data i , context representation p of target sample data t The preference context vector g of the review group corresponding to the target sample data i The reward function is modeled as S(r i , p t , g i ) = f θ ([token(r i *p t );g i]).

[0075] Then we can design the loss function. The loss function can be designed to optimize two objectives at the same time: the intra-group consistency loss L intra and inter-group alignment loss L align . L intra It can be determined using the group normalized relative loss algorithm or the conventional pairwise ranking loss (sorting algorithm). align To characterize the evaluation differences of different review groups for the same sample, we can add regularization terms to encourage the model output to remain semantically close. For example, we can set the margin value in the model not to be too large.

[0076] In a real-time mode, you can set .

[0077] Finally, the joint loss function can be obtained as L=L intra +L align The trained reinforcement learning model can be obtained by using the joint loss function to train the reinforcement learning pre-training model.

[0078] The technical solution provided in the embodiments of this application enhances the adaptability of the reward model in a reinforcement learning model to the scoring context by explicitly incorporating the review group context. Simultaneously, modeling both intra-group accuracy and inter-group alignment objectives effectively reduces scoring conflicts in the training data. This approach is more robust in non-deterministic and subjective tasks (such as content quality and value ranking), does not rely heavily on the accuracy of reviewer identities, and is effective even with crude group identification.

[0079] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0080] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0081] Figure 7 Schematic diagram of a reinforcement learning model training device provided in an embodiment of the present application. Figure 7 As shown, the device includes: The acquisition module 701 is configured to obtain N sample data, each sample data includes at least a question, a candidate answer to the question by the reinforcement learning pre-training model, the evaluation of each candidate answer by the target review group, and the target review group identifier; the target review group is any review group among the M review groups, and M and N are both positive integers.

[0082] The generating module 702 is configured to generate a preference context vector for each review group based on N sample data.

[0083] The construction module 703 is configured to construct a reward model of the reinforcement learning pre-training model based on at least the preference context vector of each review group.

[0084] The construction module 703 is further configured to construct a loss function based on the reward model. The loss function includes at least an inter-group alignment loss term, and the inter-group alignment loss term is used to characterize the difference in evaluation of the same sample data by different review groups.

[0085] The training module 704 is configured to train the reinforcement learning pre-training model based on the loss function to obtain a trained reinforcement learning model.

[0086] According to the technical solution provided in the embodiments of the present application, by first obtaining N sample data, the sample data includes questions, candidate answers, the evaluation of the candidate answers by the review group and the review group identification, and then generating a preference context vector for each review group based on the N sample data, constructing a reward model of the reinforcement learning pre-training model at least based on the preference context vector, and then constructing a loss function based on the reward model that includes at least an inter-group alignment loss term, so that the loss function can characterize the evaluation differences of different review groups on the same sample data, and finally using the constructed loss function to train the reinforcement learning pre-training model to obtain a trained reinforcement learning model, which can effectively solve the reward learning bias caused by differences between reward model groups and improve the stability of reinforcement learning model training.

[0087] In some embodiments, a preference context vector for each review group is generated based on N sample data, including: obtaining all sample data of the target review group from the N sample data; performing word-unit token encoding on the sample data of the target review group to obtain the token encoding vector of the target review group; and determining the preference context vector of the target review group based on the token encoding vector.

[0088] In some embodiments, determining the preference context vector of the target review group based on the token encoding vector includes: embedding encoding the token encoding vector to obtain the preference context vector of the target review group; or training the historical token encoding vectors of the target review group to obtain the preference context vector of the target review group.

[0089] In some embodiments, the reward model is constructed in the following manner: obtaining candidate answers in target sample data and context representation of the target sample data; the target sample data is any sample data; text splicing is performed on the candidate answers and the context representation, and a token encoding of the spliced ​​text is obtained; feature splicing is performed on the token encoding of the spliced ​​text and the preference context vector of the review group corresponding to the target sample data to obtain a spliced ​​feature; the spliced ​​feature is used as the independent variable of the preset network prediction function to obtain a reward model for the target sample data.

[0090] In some embodiments, the contextual representation of the target sample data is determined in the following manner: obtaining the evaluation window of t rounds before the target sample data, the evaluation window including the questions of each round, the candidate answers to the questions by the reinforcement learning pre-trained model, and the evaluation of each candidate answer by the target review group; performing text splicing on all items in the evaluation window of t rounds to obtain the contextual representation of the target sample data.

[0091] In some embodiments, the inter-group alignment loss term is determined in the following manner: a first reward value and a second reward value of the target sample data are calculated separately; wherein the first reward value and the second reward value correspond to the same candidate answer and context representation of the target sample data, and correspond to preference context vectors of different review groups; a second normal form calculation is performed on the difference between the first reward value and the second reward value to obtain the candidate inter-group alignment loss term; all reward values ​​in the target sample data that correspond to the same candidate answer and context representation of the target sample data, and correspond to preference context vectors of different review groups, are traversed, and the sum of all candidate inter-group alignment loss terms is determined as the inter-group alignment loss term.

[0092] In some embodiments, the loss function also includes an intra-group consistency loss term, which is used to characterize the consistency between different candidate answers to the same question; the intra-group consistency loss term is determined in the following ways: using an intra-group normalized relative loss method; or using a metric learning algorithm; or using an intra-group relative strategy optimization algorithm.

[0093] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0094] Figure 8 Schematic diagram of an electronic device provided in an embodiment of the present application. Figure 8As shown, the electronic device 8 of this embodiment includes: a processor 801, a memory 802, and a computer program 803 stored in the memory 802 and executable by the processor 801. When the processor 801 executes the computer program 803, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 801 executes the computer program 803, the functions of the modules / units in the above-described device embodiments are implemented.

[0095] The electronic device 8 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 8 may include but is not limited to a processor 801 and a memory 802. Those skilled in the art will appreciate that Figure 8 This is merely an example of the electronic device 8 and does not limit the electronic device 8 . The electronic device 8 may include more or fewer components than shown in the figure, or different components.

[0096] The processor 801 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0097] The memory 802 can be an internal storage unit of the electronic device 8, such as a hard drive or memory of the electronic device 8. The memory 802 can also be an external storage device of the electronic device 8, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. The memory 802 can also include both an internal storage unit of the electronic device 8 and an external storage device. The memory 802 is used to store computer programs and other programs and data required by the electronic device.

[0098] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0099] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium.

[0100] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A reinforcement learning model training method, characterized in that: include: Obtain N sample data, each sample data including at least a question, a candidate answer to the question by the reinforcement learning pre-trained model, an evaluation of each candidate answer by a target review group, and an identifier of the target review group; the target review group is any one of M review groups, where M and N are both positive integers; Based on the N sample data, generating a preference context vector for each review group; Constructing a reward model for the reinforcement learning pre-trained model based at least on the preference context vectors of each review group; Constructing a loss function based on the reward model, wherein the loss function includes at least an inter-group alignment loss term, and the inter-group alignment loss term is used to characterize the difference in evaluation of the same sample data by different review groups; The reinforcement learning pre-training model is trained based on the loss function to obtain a trained reinforcement learning model.

2. The method according to claim 1, characterized in that Based on the N sample data, a preference context vector for each review group is generated, including: Obtain all sample data of the target review group from the N sample data; Performing word-element token encoding on the sample data of the target review group to obtain a token encoding vector of the target review group; A preference context vector of the target review group is determined based on the token encoding vector.

3. The method according to claim 2, wherein The determining of the preference context vector of the target review group based on the token encoding vector includes: Embedding the token encoding vector to obtain a preference context vector of the target review group; or The historical token encoding vector of the target review group is trained to obtain the preference context vector of the target review group.

4. The method according to claim 3, characterized in that The reward model is constructed as follows: Obtain candidate answers from target sample data and a context representation of the target sample data; the target sample data is any sample data; Concatenate the candidate answer with the context representation and obtain a token encoding of the concatenated text; Perform feature splicing on the token encoding of the spliced ​​text and the preference context vector of the review group corresponding to the target sample data to obtain a spliced ​​feature; The splicing features are used as independent variables of a preset network prediction function to obtain a reward model for the target sample data.

5. The method according to claim 4, characterized in that The context representation of the target sample data is determined in the following manner: Obtaining the evaluation window for the t rounds before the target sample data, wherein the evaluation window includes the questions of each round, the candidate answers of the reinforcement learning pre-trained model to the questions, and the evaluation of each candidate answer by the target review group; All items in the evaluation window of the t rounds are subjected to text concatenation to obtain a contextual representation of the target sample data.

6. The method according to claim 1, wherein The inter-group alignment loss term is determined as follows: Calculating a first reward value and a second reward value for the target sample data respectively; wherein the first reward value and the second reward value correspond to the same context representation of the candidate answer and the target sample data, and correspond to preference context vectors of different review groups; Performing a second normal form calculation on the difference between the first reward value and the second reward value to obtain a candidate group alignment loss term; Traversing the target sample data, all reward values ​​corresponding to the same candidate answer and context representation of the target sample data and corresponding to preference context vectors of different review groups, determine the sum of all candidate inter-group alignment loss items as the inter-group alignment loss item.

7. The method according to claim 1, characterized in that The loss function also includes an intra-group consistency loss term, which is used to characterize the consistency between different candidate answers to the same question; The intra-group consistent loss term is determined as follows: Determined using the within-group normalized relative loss method; or Determined using a metric learning algorithm; or Determined using the within-group relative strategy optimization algorithm.

8. A reinforcement learning model training device, characterized in that: include: an acquisition module configured to acquire N sample data, each sample data including at least a question, a candidate answer to the question by the reinforcement learning pre-trained model, an evaluation of each candidate answer by a target review group, and an identifier of the target review group; the target review group is any one of M review groups, where M and N are both positive integers; A generating module configured to generate a preference context vector for each review group based on the N sample data; A construction module configured to construct a reward model of the reinforcement learning pre-training model based at least on the preference context vectors of each review group; The construction module is further configured to construct a loss function based on the reward model, wherein the loss function includes at least an inter-group alignment loss term, and the inter-group alignment loss term is used to characterize the difference in evaluation of the same sample data by different review groups; The training module is configured to train the reinforcement learning pre-training model based on the loss function to obtain a trained reinforcement learning model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-modal knowledge generation method and device based on feedback enhancement

    CN117035074A

  • Question generation method and device based on reinforcement learning and storage medium

    CN118536585A

  • Reinforcement learning alignment model training method and system based on AI feedback

    CN118735002A

  • Large model scene question and answer optimization method and system based on preference learning

    CN119336960A

  • Large language model training method and device

    CN119443155A

Cited By

  • Open domain process reward model construction method and device based on preference tree

    CN121327084A