Method and device for response evaluation, equipment and storage medium

By obtaining sample query, prediction response and truth response, using the reward model to determine the reward score, using the relative judgment framework and reinforcement learning training question and answer model, the problem of mismatch between paired properties and scalar reward signals in RLHF is solved, and more accurate language model training and output alignment is achieved.

CN120336387APending Publication Date: 2025-07-18BEIJING QINGYANG ZHIWEI TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510422647.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing human feedback reinforcement learning (RLHF) approach faces the problem of paired properties mismatching scalar reward signals and inaccurate initialization of reward models when aligning language models with human preferences, resulting in instability in training and suboptimal alignment.

Method used

By obtaining sample query, predictive response and truth-value response, the reward score is determined using the trained reward model, the question-and-answer model is trained using a relative judgment framework and reinforcement learning, and the strategy update of the language model is optimized.

Benefits of technology

Improve the accuracy of reward scores, ensure that the language model output is more consistent with human preferences, and improves the stability and alignment of training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336387A_ABST
    Figure CN120336387A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for response evaluation, equipment, a storage medium and a program product. The method comprises the steps of obtaining a first sample query, a first prediction response for the first sample query and a first truth value response, wherein the first truth value response corresponds to the first sample query; and determining, using the trained reward model, a reward score for the first predicted response based on the first sample query, the first predicted response, and the first truth response, the reward score indicating whether the first predicted response is superior to the first truth response for the sample query. In this way, the accuracy of the determined reward score may be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and particularly to methods, apparatuses, devices, and computer-readable storage media for response evaluation. Background Art

[0002] In recent years, language models have shown unprecedented capabilities in content generation. However, aligning language models with human preferences to ensure that these models output useful and context-appropriate responses remains a challenge. Human feedback reinforcement learning (RLHF) has become the main method for post-training alignment, enabling language models to learn from human preferences rather than predefined rules. Summary of the Invention

[0003] In a first aspect of the present disclosure, there is provided a method for response evaluation. The method includes: obtaining a first sample query, a first predicted response to the first sample query, and a first ground-truth response corresponding to the first sample query; and using a trained reward model, based on the first sample query, the first predicted response, and the first ground-truth response, to determine a reward score for the first predicted response, the reward score indicating whether the first predicted response is better than the first ground-truth response for the sample query; and training a question-and-answer model through reinforcement learning, the first training objective of the reinforcement learning being configured to increase or maximize the reward score.

[0004] In a second aspect of the present disclosure, there is provided a device for response evaluation. The device includes: an obtaining module configured to obtain a first sample query, a first predicted response to the first sample query, and a first ground-truth response corresponding to the first sample query; and a reward score determination module configured to use a trained reward model, based on the first sample query, the first predicted response, and the first ground-truth response, to determine a reward score for the first predicted response, the reward score indicating whether the first predicted response is better than the first ground-truth response for the sample query.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processing unit, cause the electronic device to execute the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium. A computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program which, when executed by a processor, implements the method of the first aspect.

[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which the embodiments of the present disclosure can be implemented;

[0011] Figure 2 A schematic block diagram showing a response evaluation architecture according to some embodiments of the present disclosure;

[0012] Figure 3 A flowchart showing a method for response evaluation according to some embodiments of the present disclosure;

[0013] Figure 4 A block diagram showing a device for response evaluation according to some embodiments of the present disclosure; and

[0014] Figure 5 A block diagram showing an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term "comprising" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0017] It is understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of the corresponding laws, regulations and related provisions.

[0018] It is understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure shall be informed to the user and the user's authorization shall be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0019] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0020] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0022] As used herein, the term "model" can learn the corresponding association relationship between input and output from training data, so that after training is completed, a corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably in this article.

[0023] "Neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and generally includes an input layer and an output layer, as well as one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications usually include many hidden layers to increase the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.

[0024] Generally, machine learning can roughly include three stages, namely, the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. The testing stage can sometimes be incorporated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values obtained from training and determine the corresponding model output.

[0025] Figure 1 A schematic diagram of an environment 100 in which embodiments of the present disclosure can be implemented is shown. In Figure 1 the environment 100, different stages of the model are shown, including the training stage 102 and the application stage 106. There can also be a testing stage after the training stage 102, which is not shown in the figure.

[0026] In the training stage 102, the model training system 110 is configured to perform the training of the model 105 using the training data set 112. At the beginning of the training, the model can have initial parameter values. The training process is to update the parameter values of the model 105 to the expected values based on the training data.

[0027] In the application stage 106, the obtained model 105 has trained parameter values and can be provided to the model application system 130 for use. In the application stage 106, the model 105 can be used to process the corresponding target input 132 in the actual scenario and provide the corresponding target output 134.

[0028] In Figure 1In this context, the model training system 110 and the model application system 130 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device can refer to any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0029] It should be understood that Figure 1 The components and arrangements in the illustrated environment 100 are merely examples, and the computing systems suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. For example, although shown as separate, the model training system 110 and the model application system 130 may be integrated in the same system or device. The implementations of this disclosure are not limited in this regard.

[0030] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of this disclosure.

[0031] In the field of machine learning, it is often necessary to use a reward model to determine whether a predicted response to a certain query should be given a higher or lower reward. This reward can be used as an incentive signal to motivate the learning of another model, promoting the continuous optimization of the model in the direction of obtaining greater rewards.

[0032] As mentioned previously, as an example of using the reward output by the reward model to motivate the learning of another model, RLHF can enable the language model to learn from human preferences rather than predefined rules. RLHF is generally divided into two stages: in the first stage, a reward model is trained, and the reward model can predict human preferences by making pairwise comparisons of the outputs of the language model; in the second stage, based on these rewards, the language model is optimized using reinforcement learning algorithms.

[0033] Current RLHF methods still face some challenges. The first challenge lies in the mismatch between the pairwise nature of human preference data and the scalar reward signals required for reinforcement learning. Traditional RLHF converts pairwise comparisons into scalar rewards, but these rewards often lack calibration for different prompts and responses. This calibration issue can lead to unstable training dynamics and suboptimal alignment, as the reinforcement learning algorithm may misinterpret the magnitude of the rewards in different situations. The second challenge stems from the initialization of the reward model. Most RLHF initializes the reward model from a generative model (e.g., a pre-trained or supervised fine-tuned language model). However, the reward model is used to perform a discriminative task, i.e., ranking the outputs of the language model according to human preferences, while the generative model is optimized for sequence generation. The above challenges hinder the ability of the reward model to accurately capture human preferences and propagate errors to subsequent reinforcement learning stages.

[0034] To address the problems existing in the above RLHF, in an embodiment of the present disclosure, a method for response evaluation is proposed. Specifically, a first sample query, a first predicted response to the first sample query, and a first ground-truth response are obtained, where the first ground-truth response corresponds to the first sample query. Using a trained reward model, based on the first sample query, the first predicted response, and the first ground-truth response, a reward score for the first predicted response is determined, and the reward score indicates whether the first predicted response is better than the first ground-truth response for the sample query.

[0035] According to the solution of the present disclosure, the reward model can determine a reward score for the first predicted response by evaluating the quality of the first predicted response relative to the first ground-truth response, and the reward score indicates whether the first predicted response is better than the first ground-truth response. In this way, the accuracy of the determined reward score can be improved.

[0036] To better understand the embodiments of the present disclosure, the content related to traditional RLHF is first introduced. The reward model plays a key role in converting human preferences into actionable training signals for the language model. Traditional RLHF methods can train the reward model by fitting the Bradley Terry model on pairwise human preference comparisons to assign scalar rewards to the outputs of the language model. The probability that one response is better than another is usually given by the following formula:

[0037] P(y1>2|q)=σ(r(q,y1)-r(q,y2)) (1) where q represents a prompt sampled from the dataset D (e.g., including the query input to the language model), and y1 and y2 represent the responses generated by the language model for the prompt q.

[0038] The goal of RLHF is to maximize the rewards given by the reward model, which can be defined as follows under the KL constraint:

[0039]

[0040] where T represents the total number of decision-making steps of the language model, r(s t , a t ) represents the token-level reward provided by the reward model, β represents the coefficient controlling the KL regularization strength, and π ref represents the initial policy of the language model.

[0041] A commonly used algorithm to optimize this objective is Proximal Policy Optimization (PPO), which limits the update amplitude of the policy by introducing a clipping mechanism to ensure stable policy updates, thus aligning with the constrained optimization framework inherent in the RLHF setting. PPO uses a clipped surrogate objective to update the policy of the language model, with the principle of restricting policy changes during each update to avoid disrupting large modifications. π θ (a|s) can represent the policy parameterized by θ, representing the policy in the previous iteration. The surrogate objective function for PPO can be defined as follows:

[0042]

[0043] where represents the probability ratio of the new policy to the old policy, represents the advantage estimate at time step t, and ∈ represents the hyperparameter used to control the clipping range.

[0044] Generalized Advantage Estimation (GAE) is used in PPO to calculate a more accurate advantage estimate by integrating multi-step bootstrapping, thereby reducing variance. For a trajectory of length T, the advantage estimate at time step t is calculated as follows:

[0045]

[0046] where γ represents the discount factor, λ ∈ [0, 1] represents the GAE parameter. δ t = r t + γV(s t+1 ) - V(s t ) represents the temporal difference (TD) error, where r t represents the reward at time step t, and V(s) represents the value function. It should be noted that since a discount factor of γ = 1.0 is usually adopted in RLHF, for simplicity of description, γ is ignored in subsequent representations.

[0047] After introducing the content related to RLHF, the embodiments of the present disclosure will be described next with reference to Figure 2 . Figure 2A schematic block diagram of a response evaluation architecture 200 according to some embodiments of the present disclosure is shown. As Figure 2 shown, a sample query 205 (also referred to as a first sample query), a predicted response 215 (also referred to as a first predicted response) for the sample query 205, and a ground truth response 220 (also referred to as a first ground truth response) can be obtained. In some examples, the sample query 205 may include multimodal data, and the predicted response 215 may also include multimodal data. Multimodal data may include images, text, audio, video, and the like. In one example, the sample query 205 may include a picture and a related task description, and the predicted response 215 may include a text description of the picture. In one example, the sample query 205 may include a text description, and the predicted response 215 may include a video or an image corresponding to the text description. In another example, the sample query 205 may be a text-modal question, and the predicted response 215 may be a text-modal answer.

[0048] In some embodiments, a ground truth response 220 for the sample query 205 can be constructed, and the ground truth response 220 can be assumed to be the best response for the sample query 205.

[0049] In some embodiments, the predicted response 215 can be generated based on the sample query 205 using a question-and-answer model 210 to be trained. The question-and-answer model 210 is configured to support interaction in the form of question and answer, where the model input can be considered as a query or a question, and the model output can be considered as a response to the input query. In some examples, the question-and-answer model 210 can be constructed based on a content generation model. For example, the question-and-answer model 210 can be a language model. The sample query 205 can be considered as the input of the question-and-answer model 210, and the predicted response 215 can be considered as the output of the question-and-answer model 210. It can be understood that the input and output of the model can be configured according to the actual task requirements.

[0050] In some embodiments, the predicted response 215 and the ground truth response 220 can be compared pairwise by a reward model 225 to determine a reward score for the predicted response 215. In some examples, the reward model 225 is constructed based on a generative model, and the reward model 225 can also be referred to as a generative reward model.

[0051] Unlike traditional reward models that assign absolute scores to individual responses, the reward model 225 can jointly evaluate two responses (e.g., the predicted response 215 and the ground truth response 220). Specifically, using the trained reward model 225, a reward score 230 for the predicted response 215 can be determined based on the sample query 205, the predicted response 215, and the ground truth response 220. The reward score 230 indicates whether the predicted response 215 is better than the ground truth response 220 for the sample query 205. Thus, the task of the reward model 225 can be transformed from absolute scoring to relative judgment, simplifying the task of the reward model 225. When the data distribution changes, absolute score calibration may fail because the relative judgment framework of the reward model 225 reduces the sensitivity to absolute score calibration, thus reducing the likelihood of problems when the data distribution changes.

[0052] In some embodiments, using the reward model 225, a first prediction output can be generated based on the sample query 205, the predicted response 215, and the ground truth response 220. The first prediction output indicates the probability that the predicted response 215 is better than the ground truth response 220 for the sample query 205. In some examples, the first input can be composed of the sample query 205, the predicted response 215, and the ground truth response 220, and the reward model 225 generates the first prediction output according to the first input. The first input can indicate whether, given the sample query 205, the predicted response 215 is better than the ground truth response 220. The first prediction output can indicate the probability that the reward model 225 answers "yes" for the first input.

[0053] Similarly, using the reward model 225, a second prediction output can be generated based on the sample query 205, the predicted response 215, and the ground truth response 220. The second prediction output indicates the probability that the ground truth response 220 is worse than the predicted response 215 for the sample query 205. Based on the first prediction output and the second prediction output, the reward score 230 can be determined. In some examples, the second input can be composed of the sample query 205, the predicted response 215, and the ground truth response 220, and the reward model 225 generates the second prediction output according to the second input. The second input can indicate whether, given the sample query 205, the ground truth response 220 is better than the predicted response 215. The second prediction output can indicate the probability that the reward model 225 answers "no" for the second input. Next, based on the first prediction output and the second prediction output, the reward score is determined. In this way, through symmetric evaluation (i.e., aggregating the probability of answering "yes" and the probability of answering "no"), the bias that the language model (the base model of the reward model 225) is prone to when predicting "yes" and "no" can be alleviated, any positional artifacts can be eliminated, and it can be ensured that when the predicted response 215 and the ground truth response 220 are consistent, the probability that the predicted response 215 is better than the ground truth response 220 is 0.5.

[0054] In some examples, for a sample query 205 (denoted by q), a predicted response 215 (denoted by γ), and a ground truth response 220 (denoted by y*), generating a reward score 230 using a reward model 225 can be expressed as follows:

[0055]

[0056] where r(y∣y*,q) represents the reward score 230, p grm represents the reward model 225, p grm (yes∣q,y>y*) represents the first predicted output, p grm (no∣q,y*>y) represents the second predicted output. The average of the first predicted output and the second predicted output can be used as the reward score 230.

[0057] In some embodiments, the reward model 225 is constructed based on a language model. To fully utilize the capabilities of the reward model 225 in language understanding, the pairwise comparison task can be formulated as a natural language understanding problem. In some embodiments, a first template prompt can be used to construct a first model input for the reward model 225 based on the sample query 205, the predicted response 215, and the ground truth response 220. The first template prompt guides the reward model 225 to generate the probability that the predicted response 215 is better than the ground truth response 220. For example, the first model input can be constructed as "Given the sample query q and the rule r, is the predicted response y better than the ground truth response y*? Answer yes or no. [MASK]" or its corresponding English or other language form, where the [MASK] token represents a placeholder for the output of the reward model 225. Note that the rule r is usually empty unless a special rule needs to be applied for the comparison. Then, the reward model 225 can generate a first predicted output based on the first model input, for example, the first predicted output indicates the probability that the reward model 225 outputs "yes".

[0058] In some embodiments, a second template prompt can be utilized to construct a second model input for the reward model 225 based on the sample query 205, the predicted response 215, and the ground truth response 220. The second template prompt guides the reward model 225 to generate the probability that the ground truth response 220 is worse than the predicted response 215. For example, the second model input can be "Given sample query q and rule r, is the ground truth response y* better than the predicted response y? Answer yes or no. [MASK]" or its corresponding English or other language forms. The reward model 225 can generate a second prediction output based on the second model input. For example, the second prediction output indicates the probability that the reward model 225 outputs "no". By constructing the task in a natural language manner, the reward model 225 can inherit the context understanding and semantic generalization capabilities of the pre-trained language model, enabling it to perform discriminative tasks using the same components that support its generation ability.

[0059] In some embodiments, after determining the reward score 230, the question-and-answer model 210 can be trained based on the reward score 230 through reinforcement learning. The training objective of reinforcement learning (also referred to as the first training objective) is configured to increase or maximize the reward score. In traditional RLHF, the optimization objective is to maximize the scalar reward assigned by the reward model. This scalar reward usually comes from absolute judgments, resulting in training instability due to poor calibration between different prompts. In the disclosed embodiments, the training objective of reinforcement learning is to maximize the reward score 230, that is, the probability that the predicted response 215 is better than the ground truth response 220 for the sample query 205 (also referred to as the winning probability of the predicted response 215). This optimization objective can be as follows:

[0060]

[0061] where q represents the sample query 205, y represents the predicted response 215, y * represents the ground truth response 220, p(y > y * |q) represents the winning probability, β represents the weight parameter used to balance the reward and the policy complexity, and KL(π(·|s t ) ∥ π ref (·|s t )) represents the KL divergence between the probability distributions of the new policy and the reference policy at state s t .

[0062] In some embodiments, the training objective of the reinforcement learning can also be configured such that the divergence (e.g., the KL divergence in formula (6)) between the output distribution of the question answering model 210 and the output distribution of the reference model for the question answering model 210 is within a predetermined range. The reference model can be a benchmark model of the question answering model 210, and such KL divergence control is to standardize the difference between the question answering model 210 and the reference model. During the training process, the KL divergence between the question answering model 210 and the reference model can be monitored and controlled to ensure that the update of the question answering model 210 is not too large, thereby improving the stability of the training.

[0063] In some embodiments, the training objective of the reinforcement learning includes a PPO objective. In the PPO objective, by introducing a clipping mechanism, it is ensured that the update of the policy (e.g., the set of parameters) of the question answering model 210 is not too large, thereby avoiding instability during the training process.

[0064] In some examples, the winning probability of the predicted response 215 can be modeled as follows:

[0065] p(y > y * |q) = σ(h(q, y, y*)) (7)

[0066] where σ(·) represents the sigmoid function and h(·) represents a comparison function, which is typically the difference in the scalar rewards of y and y* in traditional RLHF. This transformation introduces a key property that when the magnitude of the reward difference becomes large, the sigmoid function saturates, pushing the probability that the predicted response 215 is better than the true response 220 towards 0 or 1. Therefore, the advantage estimate (represented by in the PPO objective will decrease as the reward signal changes in these saturated regions. This saturation effect acts as a dynamic weighting mechanism between sample queries. For example, for a sample query where the predicted response is significantly better than the true response, the winning probability reduces the effective learning signal, preventing the policy from over-optimizing in these cases. For a sample query where the predicted response is comparable to the true response, the winning probability remains sensitive to small reward differences, amplifying the learning signal. Thus, this dynamic weighting mechanism reduces the risk of "reward cheating", which is a phenomenon of exploiting the defects in the reward function to obtain high rewards without truly learning the expected behavior or completing the task. In addition, clipping excessive advantages helps to reduce the upper bound of the KL divergence between the pre-update and post-update policies, thereby controlling the step size of this part of the data.

[0067] In reinforcement learning, a critic model 235 is also introduced. The critic model 235 provides feedback on the optimization direction for the question-and-answer model 210 by predicting the long-term value of the response output by the question-and-answer model 210. The question-and-answer model 210 can adjust its strategy based on the feedback from the critic model 235 to generate responses that better match human preferences. When estimating the reward, the reward model 225 compares the predicted response 215 with the ground-truth response 220. During the sampling process, the question-and-answer model 210 does not have access to the ground-truth response 220, and the training of the critic model 235 will be misaligned with either the reward model 225 or the question-and-answer model 210. If the critic model 235 does not refer to the ground-truth response 220, it is difficult to obtain an accurate critic score. If the critic model 235 refers to the ground-truth response 220, it will be misaligned with the question-and-answer model 210.

[0068] In some embodiments, the critic model 235 can refer to the ground-truth response 220. Based on the sample query 205, the predicted response 215, and the ground-truth response 220, the critic model 235 to be trained can be used to determine the critic score 240 for the currently generated token in the predicted response 215. The critic score 240 indicates the expected reward score when the currently generated token is completed in the generation of the predicted response 215, and the question-and-answer model 210 generates the next token in the predicted response 215 based on the critic score 240. In some examples, the question-and-answer model 210 can be trained through reinforcement learning based on the weighted sum of the reward score 230 and the critic score 240. Specifically, the input of the critic model 235 changes from the sample query 205 and the predicted response 215 to the sample query 205, the predicted response 215, and the ground-truth response 220, which means that the content of the ground-truth is visible during the process of determining the critic score 240. In this way, the input of the critic model 235 is exactly the same as the input of the reward model 225, so that a more accurate critic score 240 can be determined. Next, the critic model 235 can be jointly trained with the question-and-answer model 210.

[0069] In some embodiments, the critic model 235 may not refer to the ground-truth response 220. In this case, the accuracy of the critic score will be affected. One way to compensate for the affected accuracy is to distill a pointwise reward model from the reward model 225 and then use the pointwise reward model to calculate the optimization objective of the critic model 235. Given a sample query q, a predicted response y, and a ground-truth response y * , the reward score of y determined by the reward model 225 can be expressed as r(y|y*,q). The reward score of y determined by the pointwise reward model can be expressed as r p (y|q), and the reward score of y * can be expressed as r p (y *|q). The MSE loss function can be used to distill a point-wise reward model from the reward model 225. By reducing or minimizing the MSE loss, the accuracy of distillation can be increased. The MSE loss function is expressed as follows:

[0070]

[0071] During the process of updating the Q&A model 210, the reward score determined by the reward model 225 can be used to calculate the advantage estimate adv, and this process can be shown as follows:

[0072]

[0073] where represents the temporal difference error at time step t + l, and V(s t ) represents the value function estimate at state s t .

[0074] During the process of updating the critic model 235, the reward score generated by the point-wise reward model can be used to calculate the cumulative reward estimate, and this process can be shown as follows:

[0075]

[0076] Before training the Q&A model 210, the reward model 225 can be trained. In some implementations, the Q&A model 210 can be utilized to generate a second predicted response based on a second sample query. Using the reward model 225 to be trained, based on the second sample query, the second predicted response, and the second ground truth response for the second sample query prompt, a first probability that the second predicted response is better than the second ground truth response can be generated. In some examples, the second sample query, the second predicted response, and the second ground truth response can be constructed as the input of the reward model 225, and this input can be "Given the second sample query q i and the rule r i , is the second predicted response y i better than the second ground truth response ? Answer yes or no. [MASK]". It should be noted that unless special rules need to be applied for this comparison, the rule r i is usually empty. During the training process, the [MASK] token will be replaced by "yes" or "no".

[0077] Then, the reward model 225 can be trained through a second training objective, which is configured to increase or maximize the first probability when the second predicted response is better than the second true response. Whether the second predicted response is better than the second true response can be represented by a label, which can be manually annotated. In some examples, the second training objective can be implemented using a cross-entropy loss function, which can be shown as follows:

[0078]

[0079] where p θ represents the probability distribution generated by the parameters θ of the reward model 225, and yes / no represents the label indicating whether the second predicted response is better than the second true response (yes if better, otherwise no). The loss function of formula (11) can encourage the reward model 225 to increase or maximize the probability that the second predicted response is better than the second true response when the second predicted response is better than the second true response.

[0080] Initializing the reward model 225 with pre-trained weights usually introduces significant positional bias. The reward model 225 is more inclined to responses based on positional order (e.g., preferring the first response in a pair) rather than responses based on their intrinsic quality. To address this issue, in some embodiments, the reward model 225 to be trained can be used to generate a second probability that the second true response is worse than the second predicted response based on the second sample query, the second predicted response, and the second true response for the second sample query prompt. For the second sample query, the second predicted response, and the second true response, two training samples can be generated. In the first training sample, the second predicted response is before the second true response. In the second training sample, the second true response is before the second predicted response.

[0081] In some examples, the first training sample can be "Given the second sample query and rule r i , is the second predicted response better than the second true response? Answer yes or no. [MASK]", where the [MASK] token is replaced by "yes". For the first training sample, the reward model 225 can generate a first probability (i.e., the probability that the second predicted response is better than the second true response). The second training sample can be "Given the second sample query q i and rule r i , the second true response is better than the second predicted response y iYes or no? For the second training sample, the reward model 225 can generate a second probability (i.e., the probability that the second true response is worse than the second predicted response). The second training objective for training the reward model 225 can also be configured to reduce or minimize the difference between the first probability and the second probability, which can be represented by a mean squared error (MSE) loss function, and the loss function can be as follows:

[0082]

[0083] where represents the first probability, represents the second probability. The loss function in Equation (12) encourages the reward model 225 to produce symmetric judgments when the order of the inputs is reversed.

[0084] In some embodiments, the training objective of the reward model 225 can be configured as a weighted sum of the loss function in Equation (11) and the loss function in Equation (12), which can be as follows:

[0085]

[0086] where ζ can be relatively small to retain some differences between different response positions.

[0087] Figure 3 FIG. shows a flowchart of a method 300 for response evaluation according to some embodiments of the present disclosure. The method 300 can be implemented at Figure 1 the model training system 110 or the model application system 130. The method 300 will be described with reference to Figure 1 the environment 100.

[0088] At block 310, the model training system 110 or the model application system 130 obtains a first sample query, a first predicted response for the first sample query, and a first true response, where the first true response corresponds to the first sample query.

[0089] At block 320, the model training system 110 or the model application system 130 uses the trained reward model to determine a reward score for the first predicted response based on the first sample query, the first predicted response, and the first true response, where the reward score indicates whether the first predicted response is better than the first true response for the first sample query.

[0090] In some embodiments, the first predicted response is generated based on the first sample query using the question-and-answer model to be trained. The method 300 further includes: the model training system 110 or the model application system 130 trains the question-and-answer model based on the reward score through reinforcement learning, and the first training objective of the reinforcement learning is configured to increase or maximize the reward score.

[0091] In some embodiments, determining the reward score for the predicted response includes: using the reward model to generate a first predicted output based on the first sample query, the first predicted response, and the first ground-truth response, where the first predicted output indicates the probability that the first predicted response is better than the first ground-truth response for the first sample query; using the reward model to generate a second predicted output based on the first sample query, the first predicted response, and the first ground-truth response, where the second predicted output indicates the probability that the first ground-truth response is worse than the first predicted response for the first sample query; and determining the reward score based on the first predicted output and the second predicted output.

[0092] In some embodiments, the reward model is constructed based on a language model. Generating the first predicted output includes: using the first template prompt word to construct a first model input for the reward model based on the first sample query, the first predicted response, and the ground-truth response, where the first template prompt word guides the reward model to generate the probability that the first predicted response is better than the first ground-truth response; and using the reward model to generate the first predicted output based on the first model input. Generating the second predicted output includes: using the second template prompt word to construct a second model input for the reward model based on the first sample query, the first predicted response, and the first ground-truth response, where the second template prompt word guides the reward model to generate the probability that the first ground-truth response is worse than the first predicted response; and using the reward model to generate the second predicted output based on the second model input.

[0093] In some embodiments, the training of the reward model includes: using the question-and-answer model to generate a second predicted response based on the second sample query; using the reward model to be trained to generate a first probability that the second predicted response is better than the second ground-truth response based on the second sample query, the second predicted response, and the second ground-truth response for the second sample query prompt word; and training the reward model through a second training objective, where the second training objective is configured to increase or maximize the first probability according to the label indicating whether the second predicted response is better than the second ground-truth response.

[0094] In some embodiments, the training of the reward model further includes: using the reward model to be trained to generate a second probability that the second ground-truth response is worse than the second predicted response based on the second sample query, the second predicted response, and the second ground-truth response for the second sample query prompt word; and the second training objective is further configured to reduce or minimize the difference between the first probability and the second probability.

[0095] In some embodiments, method 300 further includes the model training system 110 or the model application system 130 determining, based on the first sample query and the first ground truth response, a criticism score corresponding to a token in the first predicted response currently generated by the question and answer model by using a criticism model to be trained, where the criticism score indicates an expected reward score at the completion of the generation of the currently generated token in the first predicted response, and where the question and answer model generates the next token in the first predicted response based on the criticism score; and co-training the criticism model together with the question and answer model.

[0096] In some embodiments, the first training objective is further configured to keep the divergence between the output distribution of the question and answer model and the output distribution of a reference model for the question and answer model within a predetermined range.

[0097] In some embodiments, the first training objective includes a proximal policy optimization (PPO) objective.

[0098] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 4 Apparatus 400 for response evaluation according to some embodiments of the present disclosure is shown. Apparatus 400 may be implemented as or included in the model training system 110. Each module / component in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0099] As Figure 4 shown, apparatus 400 includes an acquisition module 410 configured to acquire a first sample query, a first predicted response to the first sample query, and a first ground truth response, the first ground truth response corresponding to the first sample query; and a reward score determination module 420 configured to determine a reward score for the first predicted response based on the first sample query, the first predicted response, and the first ground truth response to the first sample query by using a trained reward model, the reward score indicating whether the first predicted response is better than the first ground truth response for the first sample query.

[0100] In some embodiments, the first predicted response is generated based on the first sample query by using a question and answer model to be trained. Apparatus 400 further includes a question and answer model training module configured to train the question and answer model based on the reward score through reinforcement learning, and the first training objective of the reinforcement learning is configured to increase or maximize the reward score.

[0101] In some embodiments, the reward score determination module 420 is further configured to use a reward model to generate a first prediction output based on a first sample query, a first predicted response, and a first ground truth response, where the first prediction output indicates the probability that the first predicted response is better than the first ground truth response for the first sample query; use the reward model to generate a second prediction output based on the first sample query, the first predicted response, and the first ground truth response, where the second prediction output indicates the probability that the first ground truth response is worse than the first predicted response for the first sample query; and determine a reward score based on the first prediction output and the second prediction output.

[0102] In some embodiments, the reward model is constructed based on a language model. The reward score determination module 420 is further configured to use a first template prompt to construct a first model input for the reward model based on the first sample query, the first predicted response, and the ground truth response, where the first template prompt guides the reward model to generate the probability that the first predicted response is better than the first ground truth response; and use the reward model to generate a first prediction output based on the first model input. The reward score determination module 420 is further configured to use a second template prompt to construct a second model input for the reward model based on the first sample query, the first predicted response, and the first ground truth response, where the second template prompt guides the reward model to generate the probability that the first ground truth response is worse than the first predicted response; and use the reward model to generate a second prediction output based on the second model input.

[0103] In some embodiments, the apparatus 400 further includes a reward model training module configured to use a question-answering model to generate a second predicted response based on a second sample query; use the reward model to be trained to generate a first probability that the second predicted response is better than the second ground truth response based on the second sample query, the second predicted response, and a second ground truth response for the second sample query prompt; and train the reward model through a second training objective, where the second training objective is configured to increase or maximize the first probability according to a label indicating whether the second predicted response is better than the second ground truth response.

[0104] In some embodiments, the reward model training module is configured to use the reward model to be trained to generate a second probability that the second ground truth response is worse than the second predicted response based on the second sample query, the second predicted response, and the second ground truth response for the second sample query prompt; where the second training objective is further configured to reduce or minimize the difference between the first probability and the second probability.

[0105] In some embodiments, the apparatus 400 further includes a critic model training module configured to determine, based on a first sample query and a first ground truth response, a critic score corresponding to a token in a first predicted response currently generated by the question answering model by using a critic model to be trained, where the critic score indicates an expected reward score of the currently generated token when the generation of the first predicted response is completed, and where the question answering model generates a next token in the first predicted response based on the critic score; and jointly train the critic model with the question answering model.

[0106] In some embodiments, the first training objective is further configured to keep the divergence between the output distribution of the question answering model and the output distribution of a reference model for the question answering model within a predetermined range

[0107] In some embodiments, the first training objective includes a proximal policy optimization (PPO) objective.

[0108] The units and / or modules included in the apparatus 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units and / or modules in the apparatus 400 can be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0109] It should be understood that one or more steps in the above methods can be performed by a suitable electronic device or a combination of electronic devices. Such an electronic device or a combination of electronic devices can include, for example Figure 1 the model training system 110 or the model application system 130 in

[0110] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 5 the electronic device 500 shown is merely exemplary and should not impose any limitation on the functions and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to implement Figure 1 the model training system 110, the model application system 130 or Figure 4 the apparatus 400 in

[0111] As Figure 5As shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processing units or processors 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processor 510 may be a physical or virtual processor and is capable of performing various processes according to the programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.

[0112] The electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 may be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 may be a removable or non-removable medium and may include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.

[0113] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules that are configured to perform the various methods or actions of the various embodiments of the present disclosure.

[0114] The communication unit 540 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 500 may be implemented by a single computing cluster or multiple computer machines that are capable of communicating through a communication connection. Thus, the electronic device 500 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0115] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed through the communication unit 540. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device (e.g., a network card, a modem, etc.) that enables the electronic device 500 to communicate with one or more other electronic devices. Such communication can be performed via an input / output (I / O) interface (not shown).

[0116] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0117] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0118] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0119] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are performed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions.

[0121] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field of this technology without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled artisans in the field of this technology to understand the various implementations disclosed herein.

Claims

1. A method for response evaluation, comprising: Obtaining a first sample query, a first predicted response to the first sample query, and a first ground truth response, where the first ground truth response corresponds to the first sample query; And Using a trained reward model, based on the first sample query, the first predicted response, and the first ground truth response, determining a reward score for the first predicted response, where the reward score indicates whether the first predicted response is better than the first ground truth response for the first sample query.

2. The method according to claim 1, wherein the first predicted response is generated using a question-and-answer model to be trained, based on the first sample query, and wherein the method further comprises: Training the question-and-answer model based on the reward score through reinforcement learning, where the first training objective of the reinforcement learning is configured to increase or maximize the reward score.

3. The method according to claim 1 or 2, wherein determining the reward score for the predicted response comprises: Using the reward model, based on the first sample query, the first predicted response, and the first ground truth response, generating a first predicted output, where the first predicted output indicates the probability that the first predicted response is better than the first ground truth response for the first sample query; Using the reward model, based on the first sample query, the first predicted response, and the first ground truth response, generating a second predicted output, where the second predicted output indicates the probability that the first ground truth response is worse than the first predicted response for the first sample query; And Based on the first predicted output and the second predicted output, determining the reward score.

4. The method according to claim 3, wherein the reward model is constructed based on a language model, and wherein generating the first predicted output comprises: Using a first template prompt, based on the first sample query, the first predicted response, and the ground truth response, constructing a first model input for the reward model, where the first template prompt guides the reward model to generate the probability that the first predicted response is better than the first ground truth response; And Using the reward model to generate the first predicted output based on the first model input; And wherein generating the second predicted output comprises: Using a second template prompt, based on the first sample query, the first predicted response, and the first ground truth response, constructing a second model input for the reward model, where the second template prompt guides the reward model to generate the probability that the first ground truth response is worse than the first predicted response; And Using the reward model to generate the second predicted output based on the second model input.

5. The method according to claim 2, wherein the training of the reward model comprises: Using the question-and-answer model to generate a second predicted response based on a second sample query; Using the reward model to be trained, based on the second sample query, the second predicted response, and a second ground truth response for the second sample query prompt, generating a first probability that the second predicted response is better than the second ground truth response; And Training the reward model with a second training objective, the second training objective being configured to increase or maximize the first probability based on a label indicating whether a second predicted response is better than the second ground truth response.

6. The method according to claim 5, wherein the training of the reward model further comprises: Using the reward model to be trained, based on the second sample query, the second predicted response, and a second ground truth response for the second sample query prompt, generating a second probability that the second ground truth response is worse than the second predicted response; wherein the second training objective is further configured to reduce or minimize the difference between the first probability and the second probability.

7. The method according to claim 2, the method further comprising: Based on the first sample query and the first ground truth response, using a critic model to be trained to determine a criticism score corresponding to a token in the first predicted response currently generated by the question answering model, wherein the criticism score indicates an expected reward score of the currently generated token at the completion of the generation of the first predicted response, and wherein the question answering model generates the next token in the first predicted response based on the criticism score; and Jointly training the critic model with the question answering model.

8. The method according to claim 2, wherein the first training objective is further configured to keep the divergence between the output distribution of the question answering model and the output distribution of a reference model for the question answering model within a predetermined range.

9. The method according to claim 2, wherein the first training objective includes a proximal policy optimization (PPO) objective.

10. An apparatus for response evaluation, comprising: An acquisition module configured to acquire a first sample query, a first predicted response to the first sample query, and a first ground truth response corresponding to the first sample query; And A reward score determination module configured to use a trained reward model to determine a reward score for the first predicted response based on the first sample query, the first predicted response, and the first ground truth response, the reward score indicating whether the first predicted response is better than the first ground truth response for the first sample query.

11. An electronic device, comprising: At least one processor; And At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the device to perform the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, implementing the method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program, the computer program, when executed by a processor, implementing the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Reinforcement learning model training method, single-layer cloth separation method, device, equipment and medium

    CN120706497A