Reliable reward evaluation method and device combining human preference and verifiable correctness signal

By combining human preferences and verifiable correctness signals, the problem of deviations in the evaluation of responses by existing reward models is solved, and more accurate and reliable reward scores are achieved, which improves the training effect of large language models.

CN120104749AActive Publication Date: 2025-06-06TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202510212277.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-06
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Existing reward models trained based on human preference data are prone to bias when evaluating model responses, tend to choose responses with better language style and longer length, while ignoring the actual content and accuracy of the response.

Method used

Combining human preferences and reliable reward evaluation methods that verifies correctness signals, we obtain the reward score output from each model by obtaining the response of the model to be evaluated and inputting it into the basic reward model and the verified correctness signal reward model respectively, and obtain the final reward score based on the reward score output from each model and the weight of each model.

Benefits of technology

It provides more accurate and reliable reward scores, taking into account verifiable correctness signals such as human preferences and factuality of model responses and instructional compliance, thereby improving the training effect and performance of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104749A_ABST
    Figure CN120104749A_ABST
Patent Text Reader

Abstract

The invention provides a reliable reward evaluation method and device combining human preferences and verifiable correctness signals, and relates to the technical field of artificial intelligence, and the method comprises the steps: inputting a to-be-evaluated model response into a basic reward model and a verifiable correctness signal reward model, and obtaining a reward score outputted by each model; performing weighted summation based on the reward score output by each model and the weight of each model to obtain a final reward score for the response of the to-be-evaluated model; the basic reward model is used for representing preference scores of humans to model responses; a verifiable correctness signal award model is used to characterize correctness of the model response in a particular aspect. According to the reliable reward evaluation method combining the human preference and the verifiable correctness signal, not only can the human preference be considered, but also verifiable correctness signals such as factuality of model response and instruction following can be comprehensively considered, so that more accurate and reliable rewards are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a reliable reward evaluation method and device that combines human preferences with verifiable correctness signals. Background Art

[0002] In the field of modern artificial intelligence, large language models have become the core tools for natural language processing and generation tasks. These models can generate high-quality model responses by learning from large amounts of text data and are widely used in a variety of scenarios such as chatbots, text generation, and machine translation.

[0003] In order to further improve the performance of large language models, reward models are widely used in the training process of large language models. In related technologies, reward models are mainly trained based on human preference data. These models learn how to evaluate and score the quality of model responses by collecting human preference choices for different responses. These preference data are used to train reward models so that they can automatically evaluate the quality of model responses and provide guidance for the training of large language models.

[0004] However, since human preferences are subjective, different people may have completely different evaluations of the same response. Therefore, the reward scores obtained based on human preference data may be affected by subjectivity and may cause the reward model to be biased when evaluating model responses. The model model tends to select responses with better language style and longer length, while ignoring the actual content and accuracy of the model response. Summary of the invention

[0005] The purpose of this application is to provide a reliable reward evaluation method and device that combines human preferences with verifiable correctness signals, which can not only take into account human preferences, but also comprehensively consider verifiable correctness signals such as the factuality of model responses and instruction compliance, thereby providing more accurate and reliable rewards.

[0006] This application provides a reliable reward evaluation method that combines human preferences with verifiable correctness signals, including: Obtain the response of the model to be evaluated, and input the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively to obtain the reward score output by each model; perform weighted summation based on the reward score output by each model and the weight of each model to obtain the final reward score for the response of the model to be evaluated; wherein the basic reward model is used to characterize the preference score of humans for the model response; the verifiable correctness signal reward model is used to characterize the correctness of the model response in specific aspects.

[0007] Optionally, the step of inputting the response of the model to be evaluated into a basic reward model and a verifiable correctness signal reward model respectively to obtain a reward score output by each model, comprising: inputting the response of the model to be evaluated into the basic reward model to obtain a human preference score for the response of the model to be evaluated, and determining the preference score as the reward score output by the basic reward model; wherein the basic reward model is a regression model.

[0008] Optionally, the verifiable correctness signal reward model includes: a router and multiple verification agents; the verification agents include: a factual verification agent and an instruction compliance verification agent; the factual verification agent is used to evaluate whether the factual information contained in the model response is correct; the instruction compliance verification agent is used to evaluate whether the model response satisfies the hard constraints specified in the instruction; the inputting the model response to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively to obtain the reward score of each model output, including: generating corresponding verification instructions based on the model response to be evaluated, and inputting the verification instructions into the router, selecting the verification agent through the router, and filtering out the target agent to be used from the multiple verification agents; the target agent includes at least one verification agent; using the target agent to verify the correctness of the model response to be evaluated to obtain the reward score output by the target agent.

[0009] Optionally, the verification instruction is input into the router, the verification agent is selected by the router, and a target agent to be used is screened out from the multiple verification agents, including: the verification instruction and the description information of each verification agent in the multiple verification agents are input into a natural language model, and the natural language model screens out matching target agents from the multiple verification agents based on the input instruction and the description information of each verification agent; wherein the description information includes: functions and usage conditions of the verification agent.

[0010] Optionally, the target agent includes: the factual verification agent; the model response to be evaluated includes: two different model responses; the use of the target agent to verify the correctness of the model response to be evaluated to obtain a reward score output by the target agent includes: inputting the two different model responses into the factual verification agent, identifying the difference in the claimed facts between the two different model responses, obtaining a difference identification result, and generating a corresponding difference query based on the difference identification result; using the difference query to perform evidence retrieval, and based on the retrieved evidence, factually verifying the difference between the two different model responses, and scoring each of the two different model responses according to the verification result to obtain a score for each model response; based on the score of each model response, determining the reward score of the factual verification agent for the model response to be evaluated; wherein the score of the model response with correct facts is higher than the score of the model response with incorrect facts.

[0011] Optionally, the target agent includes: the instruction compliance verification agent; the use of the target agent to verify the correctness of the response of the model to be evaluated to obtain a reward score output by the target agent, including: extracting at least one constraint condition for the model response from the verification instruction, and generating a corresponding constraint verification script based on the constraint condition; using the constraint verification script to perform constraint verification on the response of the model to be evaluated, judging whether the response of the model to be evaluated satisfies each constraint condition in the at least one constraint condition, and obtaining a constraint verification result; calculating the average value of the score for each constraint condition in the constraint verification result to obtain a reward score for the response of the model to be evaluated by the instruction compliance verification agent.

[0012] Optionally, the using of the target agent to verify the correctness of the response of the model to be evaluated to obtain a reward score output by the target agent includes: performing a weighted summation of the reward scores output by each verification agent in the multiple verification agents to obtain the reward score output by the target agent; wherein the reward score output by the target agent is: the reward score output by the verifiable correctness signal reward model for the response of the model to be evaluated.

[0013] The present application also provides a reliable reward evaluation device that combines human preferences with verifiable correctness signals, including: An acquisition module is used to obtain the response of the model to be evaluated; a reward scoring module is used to input the response of the model to be evaluated into a basic reward model and a verifiable correctness signal reward model respectively to obtain a reward score output by each model; the reward scoring module is also used to perform weighted summation based on the reward score output by each model and the weight of each model to obtain a final reward score for the response of the model to be evaluated; wherein the basic reward model is used to characterize human preference scores for model responses; and the verifiable correctness signal reward model is used to characterize the correctness of model responses in specific aspects.

[0014] Optionally, the reward scoring module is specifically used to input the response of the model to be evaluated into the basic reward model, obtain a human preference score for the response of the model to be evaluated, and determine the preference score as the reward score output by the basic reward model; wherein the basic reward model is a regression model.

[0015] Optionally, the verifiable correctness signal reward model includes: a router and multiple verification agents; the verification agents include: a factual verification agent and an instruction compliance verification agent; the factual verification agent is used to evaluate whether the factual information contained in the model response is correct; the instruction compliance verification agent is used to evaluate whether the model response satisfies the hard constraints specified in the instruction; the reward scoring module is specifically used to generate corresponding verification instructions based on the model response to be evaluated, and input the verification instructions into the router, select the verification agent through the router, and filter out the target agent to be used from the multiple verification agents; the target agent includes at least one verification agent; the reward scoring module is also specifically used to use the target agent to verify the correctness of the model response to be evaluated, and obtain the reward score output by the target agent.

[0016] Optionally, the reward scoring module is specifically used to input the verification instructions and the description information of each verification agent in the multiple verification agents into a natural language model, and the natural language model filters out matching target agents from the multiple verification agents based on the input instructions and the description information of each verification agent; wherein the description information includes: the functions and usage conditions of the verification agent.

[0017] Optionally, the target agent includes: the factual verification agent; the model response to be evaluated includes: two different model responses; the reward scoring module is specifically used to input the two different model responses into the factual verification agent, identify the difference in the claimed facts between the two different model responses, obtain a difference identification result, and generate a corresponding difference query based on the difference identification result; the reward scoring module is specifically used to use the difference query to retrieve evidence, and to factually verify the difference between the two different model responses based on the retrieved evidence, and to score each of the two different model responses according to the verification result to obtain a score for each model response; the reward scoring module is specifically used to determine the reward score of the factual verification agent for the model response to be evaluated based on the score of each model response; wherein the score of the model response with correct facts is higher than the score of the model response with incorrect facts.

[0018] Optionally, the target agent includes: the instruction compliance verification agent; the reward scoring module, which is specifically used to extract at least one constraint condition for the model response from the verification instruction, and generate a corresponding constraint verification script based on the constraint condition; the reward scoring module is specifically used to use the constraint verification script to perform constraint verification on the response of the model to be evaluated, determine whether the response of the model to be evaluated satisfies each constraint condition in the at least one constraint condition, and obtain the constraint verification result; the reward scoring module is specifically used to calculate the average value of the score for each constraint condition in the constraint verification result, and obtain the reward score of the instruction compliance verification agent for the response of the model to be evaluated.

[0019] Optionally, the reward scoring module is specifically used to obtain the reward score output by the target agent by weighted summing the reward scores output by each verification agent in the multiple verification agents; wherein the reward score output by the target agent is: the reward score output by the verifiable correctness signal reward model for the response of the model to be evaluated.

[0020] The present application also provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of any of the above-described reliable reward evaluation methods combining human preferences with verifiable correctness signals.

[0021] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of a reliable reward evaluation method combining human preferences with verifiable correctness signals as described in any one of the above are implemented.

[0022] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described reliable reward evaluation methods combining human preferences with verifiable correctness signals.

[0023] The reliable reward evaluation method and device provided by the present application that combines human preferences with verifiable correctness signals, first, obtain the response of the model to be evaluated, and input the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively, to obtain the reward score output by each model; then, perform weighted summation based on the reward score output by each model and the weight of each model, to obtain the final reward score for the response of the model to be evaluated; wherein, the basic reward model is used to characterize the human preference score for the model response; the verifiable correctness signal reward model is used to characterize the correctness of the model response in specific aspects. In this way, not only human preferences can be considered, but also verifiable correctness signals such as the factuality of the model response and instruction compliance can be comprehensively considered, thereby providing more accurate and reliable rewards. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 It is a schematic diagram of the structure of the reward system provided in the embodiment of the present application; Figure 2 It is a flowchart of a reliable reward evaluation method combining human preferences with verifiable correctness signals provided by the present application; Figure 3 is a schematic diagram of the benchmark test results of the reward model provided by this application; Figure 4 It is a schematic diagram of the best candidate search result of the reward model provided by this application; Figure 5 This is a schematic diagram of the results of using REWARDAGENT to construct training preference pairs and train a large language model provided by this application; Figure 6 It is a schematic diagram of the structure of a reliable reward evaluation device combining human preferences with verifiable correctness signals provided by the present application; Figure 7 It is a structural schematic diagram of the electronic device provided by this application. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0027] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0028] Although the human preference-based reward models in related technologies have achieved certain results in the training of large language models, they also have some significant defects. First, human preferences are subjective in nature, and different people may have different evaluations of the same response. This subjectivity may cause the reward model to be biased when evaluating responses, tending to select those responses with better language style and longer length, while ignoring the actual content and accuracy of the response. Second, existing reward models mainly focus on human preferences, while ignoring key information such as the factuality and instruction following of the response. For example, in a task that requires generating text about a historical event, a response may be fluent in language and long in length, but contain some incorrect historical facts. Existing reward models may give a higher score because of its language style, while ignoring these factual errors. Similarly, in some tasks that require strict instruction following, a response may meet most of the requirements, but miss some key instruction constraints. Existing reward models may not accurately identify these instruction following issues, thus affecting the evaluation of the response. This neglect of factuality and instruction following will not only affect the reliability of the reward model, but may also cause the trained large language models to have incorrect or non-compliant responses in practical applications, affecting the user experience and the credibility of the model.

[0029] In view of the above technical problems existing in the related art, the embodiment of the present application provides a reliable reward evaluation method combining human preferences with verifiable correctness signals, which is applied to Figure 1 The reward evaluation system shown in Figure 1As shown, the system includes a basic reward model and a verifiable correctness signal reward model (including a router and multiple verification agents). After the model response is input into the two models, the reward score output by each model is obtained, and the final reward score is obtained by weighted summation. This method can not only take into account human preferences, but also comprehensively consider verifiable correctness signals such as the factuality of the response and instruction compliance, thereby providing more accurate and reliable rewards. Through this comprehensive evaluation method, the present invention can effectively improve the training effect and performance of large language models, so that they can generate higher quality and more compliant text responses in various natural language processing tasks.

[0030] In conjunction with the accompanying drawings, the reliable reward evaluation method combining human preferences and verifiable correctness signals provided by the embodiment of the present application is described in detail through specific embodiments and their application scenarios.

[0031] like Figure 2 As shown, the embodiment of the present application provides a reliable reward evaluation method that combines human preferences with verifiable correctness signals. The method may include the following steps 201 and 202: Step 201: Obtain the response of the model to be evaluated, and input the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively to obtain the reward score output by each model.

[0032] The basic reward model is used to characterize human preference scores for model responses; the verifiable correctness signal reward model is used to characterize the correctness of model responses in specific aspects.

[0033] Exemplarily, the above-mentioned model response to be evaluated is a text response output by a large language model. The model response to be evaluated needs to be scored by a reward model so that the large language model can be trained according to the scoring results, thereby helping the large language model to learn to generate text responses that are more in line with human expectations.

[0034] For example, in the embodiment of the present application, the traditional reward model based on human preference is combined with verifiable correctness signals from different aspects, and its core idea is shown in the following formula 1: r(x,y)=λ⋅rRM(x,y)+∑i∈Axwi⋅ai(x,y) (Formula 1) Among them, r(x,y) is the final reward score, which represents the comprehensive evaluation of instruction x and response y. λ: The weight of the basic reward model, which is used to adjust the contribution of human preference rewards in the final reward. rRM(x,y): The output of the basic reward model (RewardModel), which represents the human preference score for response y. Ax: The index set of verification agents related to instruction x, indicating the verification agent that needs to be called according to instruction x. wi: The weight of the i-th verification agent, which is used to adjust the contribution of each verification agent in the final reward. ai(x,y): The output of the i-th verification agent, which represents the correctness score of response y in a specific aspect.

[0035] Exemplarily, according to the above formula 1, the reward score in the embodiment of the present application may include the following two parts: ①, basic reward model part: λ⋅rRM(x,y), which represents the contribution of human preference reward. rRM(x,y) is the output of a regression model, usually trained based on human-labeled preference data. It reflects the subjective evaluation of human on response y. The weight λ is used to adjust the importance of this part in the final reward.

[0036] ②. Verifiable correctness signal part: ∑i∈Axwi⋅ai(x,y), this part represents the contribution of the verifiable correctness signal. Ax is a set of verification agents dynamically selected according to instruction x, and each verification agent ai will evaluate the correctness of the response y in a specific aspect. For example, ai can be a factual verification agent or an instruction compliance verification agent. The output ai(x,y) of each verification agent is multiplied by its corresponding weight wi, and then the contributions of all verification agents are added together to get the sum of this part.

[0037] It should be noted that in the specific implementation, the embodiment of the present application designs a reward agent named REWARDAGENT, which integrates human preference rewards and two verifiable signals: factuality and instruction compliance. REWARDAGENT consists of three main modules: Router, Verification Agents and Judger.

[0038] Specifically, in the above step 201, the reward score calculation step for the basic reward model part may include the following steps 201a: Step 201a: input the response of the model to be evaluated into the basic reward model to obtain a preference score of humans for the response of the model to be evaluated, and determine the preference score as the reward score output by the basic reward model.

[0039] Wherein, the basic reward model is a regression model.

[0040] Exemplarily, in the embodiments of the present application, human preferences are still used as part of the reward score. The basic reward model learns how to evaluate and score the quality of the model response by collecting human preference choices for different responses. For example, in a text generation task, human annotators may compare multiple generated text responses and select the responses they think are of higher quality. These preference data are used to train the reward model so that it can automatically evaluate the quality of the response and provide guidance for the training of large language models.

[0041] Exemplarily, for the reward scoring of the above-mentioned verifiable correctness signal part, the above-mentioned verifiable correctness signal reward model includes: a router and multiple verification agents; the verification agents include: a factual verification agent and an instruction compliance verification agent; the factual verification agent is used to evaluate whether the factual information contained in the model response is correct; the instruction compliance verification agent is used to evaluate whether the model response meets the hard constraints specified in the instruction.

[0042] Specifically, in the above step 201, the reward score calculation step for the verifiable correctness signal reward model part may include the following steps 201b and 201c: Step 201b, generating a corresponding verification instruction based on the response of the model to be evaluated, and inputting the verification instruction into the router, selecting a verification agent through the router, and filtering out a target agent to be used from the multiple verification agents; the target agent includes at least one verification agent.

[0043] Step 201c: Use the target agent to verify the correctness of the response of the model to be evaluated, and obtain the reward score output by the target agent.

[0044] Exemplarily, in the embodiment of the present application, the router is the first module of REWARDAGENT, and its main function is to analyze the input instructions and determine which verification agents need to be called to evaluate the response. Different instructions may require different aspects of the response to be evaluated, so dynamically selecting the appropriate verification agent can reduce the cost of reasoning and reduce potential cumulative errors.

[0045] Specifically, in the above step 201b, the step of selecting a verification agent for the router may further include the following step 201b1: Step 201b1: input the verification instruction and the description information of each verification agent in the multiple verification agents into a natural language model, and the natural language model screens out matching target agents from the multiple verification agents based on the input instruction and the description information of each verification agent.

[0046] The description information includes: functions and usage conditions of the verification agent.

[0047] Exemplarily, the router has an existing large language model as its core, and by inputting instructions and descriptions of all verification agents, it prompts the large language model to select an appropriate verification agent. Specifically, we first provide a concise description of each verification agent, explaining its function and usage conditions. Then, the instructions and all agent descriptions are input into the large language model together, and the large language model selects an appropriate verification agent according to the requirements of the instructions. For example, if the instruction requires the generation of a text about a historical event, and it is necessary to ensure that the facts in the text are correct, then the router will select a factual verification agent; if the instruction contains some specific format or length requirements, then the router will select an instruction-compliant verification agent.

[0048] Exemplarily, the verification agent is the core part of REWARDAGENT, which is used to evaluate the correctness of the response in different aspects. In the embodiment of the present application, we designed two main verification agents: a factual verification agent and an instruction compliance verification agent. It should be noted that in addition to the above two verification agents, other agents can be added to the model according to actual conditions. The router can select one or more verification agents according to the input instructions, or it can not select any verification agent.

[0049] In a possible implementation, the target agent includes: the factual verification agent; the model response to be evaluated includes: two different model responses.

[0050] Specifically, in the case where the target agent includes a fact verification agent, the step 201c may further include the following steps 201c1 to 201c3: Step 201c1: input the two different model responses into the factual verification agent, identify the difference in the claimed facts between the two different model responses, obtain a difference identification result, and generate a corresponding difference query based on the difference identification result.

[0051] Step 201c2: Use the difference query to perform evidence retrieval, and based on the retrieved evidence, perform fact verification on the difference between the two different model responses, and score each of the two different model responses according to the verification result to obtain a score for each model response.

[0052] Step 201c3: Based on the score of each model response, determine the reward score of the factual verification agent for the model response to be evaluated.

[0053] Among them, the scores of model responses that are factually correct are higher than the scores of model responses that are factually incorrect.

[0054] For example, the main task of the factual verification agent is to evaluate whether the factual information contained in the response is correct. In order to complete this task efficiently, we designed a factual verification agent based on pairwise comparison. The agent evaluates the factuality of the response through the following four main steps: Difference Proposal: Identify the key differences in the claimed facts between two given responses. For example, if one response claims that the height of Mount Everest is 8,848 meters, and another response claims it is 8,868 meters, then the difference proposal step will identify the difference between the two height values.

[0055] Query Generation: Generate queries based on the identified differences in order to retrieve evidence to distinguish these differences. For example, for the height difference above, the query generation step might generate a query such as "What is the height of Mount Everest?".

[0056] Evidence Generation: Using the generated query, retrieve supporting evidence through an external search engine or the parameter knowledge of a large language model. For example, retrieve the height of Mount Everest is 8,848 meters through a search engine, or use the knowledge base inside a large language model to verify this fact.

[0057] Verification: Based on the collected evidence and the original response, each response is assigned an integer score from 0 to 1. A higher score is given if the facts in the response are verified to be correct, and a lower score is given if the facts are wrong.

[0058] This pairwise comparison-based factual verification agent can effectively capture subtle factual differences between responses while significantly reducing the inference time cost since it only needs to verify the differences between responses instead of all claimed facts.

[0059] In another possible implementation manner, the target agent includes: the instruction compliance verification agent.

[0060] Specifically, in the case where the target agent includes an instruction compliance verification agent, the step 201c may further include the following steps 201c4 to 201c6: Step 201c4: extract at least one constraint condition for the model response from the verification instruction, and generate a corresponding constraint verification script based on the constraint condition.

[0061] Step 201c5: perform constraint verification on the response of the model to be evaluated using the constraint verification script to determine whether the response of the model to be evaluated satisfies each constraint condition in the at least one constraint condition, and obtain a constraint verification result.

[0062] Step 201c6: Calculate the average of the scores for each constraint condition in the constraint verification result to obtain the reward score of the instruction compliance verification agent for the response of the model to be evaluated.

[0063] For example, the main task of the Instruction-Following Verification Agent is to evaluate whether the response meets the hard constraints specified in the instruction. These constraints usually include length restrictions, format requirements, keyword restrictions, etc. The Instruction-Following Verification Agent completes its task through the following three steps: Constraint Parsing: Extracts hard constraints from instructions. For example, if the instruction requires generating a text of no more than 100 words, the constraint parsing step will extract the constraint of "length no more than 100 words".

[0064] Code Generation and Refinement: Generates a Python script that checks whether the response satisfies the extracted constraints. The generated code takes the response as input and returns a Boolean value indicating whether the response satisfies the constraints. For example, for the length constraint above, the generated code might check whether the response has no more than 100 words. If an error occurs when executing the code, the error information and the original code are fed back to the model so that more accurate code can be generated.

[0065] Verification: Execute the generated Python code and obtain a binary score (0 or 1) indicating whether the response satisfies each hard constraint. The final score is the average of all hard constraint scores.

[0066] This instruction compliance verification agent can effectively evaluate whether the response satisfies the hard constraints in the instructions, thereby ensuring that the generated text meets the user's requirements.

[0067] Step 202: Perform a weighted sum based on the reward score output by each model and the weight of each model to obtain a final reward score for the response of the model to be evaluated.

[0068] Exemplarily, after obtaining the reward score output by each model, the final reward score can be calculated by the judge. The judge is the last module of REWARDAGENT, and its main function is to use the above formula 1 to integrate the correctness signal from the verification agent and the human preference score from the reward model to generate the final reward score. In the implementation of this application, we use the weighted sum as the judge, where λ and wi represent the weights of the basic reward model and each verification agent, respectively. These weights can be adjusted according to the specific situation to adapt to different application scenarios. For example, in some scenarios, factuality may be more important, so the weight of the factual verification agent can be increased; in other scenarios, human preferences may be more critical, so the weight of the reward model can be increased. In this way, the judge can comprehensively consider human preferences and verifiable correctness signals to provide a comprehensive and accurate reward score for each response.

[0069] The reliable reward evaluation method combining human preferences and verifiable correctness signals provided in the embodiment of the present application, first, obtain the response of the model to be evaluated, and input the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively, to obtain the reward score output by each model; then, perform weighted summation based on the reward score output by each model and the weight of each model, to obtain the final reward score for the response of the model to be evaluated; wherein, the basic reward model is used to characterize the human preference score for the model response; the verifiable correctness signal reward model is used to characterize the correctness of the model response in specific aspects. In this way, not only human preferences can be considered, but also verifiable correctness signals such as the factuality of the model response and instruction compliance can be comprehensively considered, thereby providing more accurate and reliable rewards.

[0070] By benchmarking existing reward models (such as Figure 3 ) and inference-time best candidate search for real-world downstream tasks (as Figure 4 As shown in Figure 2, REWARDAGENT significantly outperforms traditional reward models. Figure 3 As shown in the figure, in benchmarks such as RM-Bench and JudgeBench, REWARDAGENT's accuracy is significantly higher than that of existing reward models such as ArmoRM, INF-ORM-Llama3.1-70B, etc. In addition, in IFBench, a newly constructed benchmark for evaluating instruction following, REWARDAGENT also performed well, especially in the subset with more hard constraints, where its accuracy is much higher than other models. These results show that REWARDAGENT can more accurately evaluate the quality of responses and select responses that better meet the requirements. In addition, as Figure 5As shown in , using REWARDAGENT to build training preference pairs and train large language models can also achieve better performance than traditional reward models in various natural language processing benchmarks. For example, Figure 5 As shown in the figure, in tasks such as MMLU, MMLU-Pro, TriviaQA, and TruthfulQA, the large language model trained with REWARDAGENT outperforms the large language model trained with the traditional reward model in terms of accuracy and reliability. This further proves the effectiveness of the present invention in improving the training effect and performance of large language models. By comprehensively considering human preferences and verifiable correctness signals, the present invention can effectively improve the performance of the reward model, thereby improving the training effect and application performance of large language models, and has important practical application value and broad application prospects.

[0071] The reliable reward evaluation method combining human preferences with verifiable correctness signals provided in the embodiment of the present application has the key point of proposing a reward system architecture that combines human preferences with verifiable correctness signals. This architecture makes up for the shortcomings of the existing reward model in terms of factuality and instruction compliance by introducing verifiable correctness signals such as factuality and instruction compliance, thereby being able to provide more accurate and reliable rewards. In addition, the present invention also designs a flexible reward agent framework, including a router, a verification agent and a judge, so that the reward system can dynamically select a suitable verification agent according to different instructions, and comprehensively consider multiple aspects of signals to evaluate the quality of the response. This comprehensive evaluation method not only improves the reliability and accuracy of the reward model, but also enhances its applicability and robustness in various natural language processing tasks.

[0072] It should be noted that the reliable reward evaluation method combining human preferences and verifiable correctness signals provided in the embodiments of the present application can be executed by a reliable reward evaluation device combining human preferences and verifiable correctness signals, or a control module in the reliable reward evaluation device combining human preferences and verifiable correctness signals for executing the reliable reward evaluation method combining human preferences and verifiable correctness signals. In the embodiments of the present application, the reliable reward evaluation method combining human preferences and verifiable correctness signals is executed by a reliable reward evaluation device combining human preferences and verifiable correctness signals as an example to illustrate the reliable reward evaluation device combining human preferences and verifiable correctness signals provided in the embodiments of the present application.

[0073] It should be noted that in the embodiments of the present application, the above-mentioned methods shown in the drawings. The reliable reward evaluation method combining human preferences and verifiable correctness signals is illustrated by combining one of the drawings in the embodiments of the present application as an example. In specific implementation, the reliable reward evaluation method combining human preferences and verifiable correctness signals shown in the drawings of the above-mentioned methods can also be implemented in combination with any other drawings that can be combined as shown in the above-mentioned embodiments, which will not be repeated here.

[0074] The following is a description of a reliable reward evaluation device that combines human preferences with verifiable correctness signals provided by the present application. The following description and the above description of a reliable reward evaluation method that combines human preferences with verifiable correctness signals can be referenced to each other.

[0075] Figure 6 A schematic diagram of the structure of a reliable reward evaluation device that combines human preferences with verifiable correctness signals provided in an embodiment of the present application, such as Figure 6 As shown, specifically including: The acquisition module 601 is used to obtain the response of the model to be evaluated; the reward scoring module 602 is used to input the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively to obtain the reward score output by each model; the reward scoring module 602 is also used to perform weighted summation based on the reward score output by each model and the weight of each model to obtain the final reward score for the response of the model to be evaluated; wherein the basic reward model is used to characterize the preference score of humans for the model response; the verifiable correctness signal reward model is used to characterize the correctness of the model response in specific aspects.

[0076] Optionally, the reward scoring module 602 is specifically used to input the response of the model to be evaluated into the basic reward model, obtain a human preference score for the response of the model to be evaluated, and determine the preference score as the reward score output by the basic reward model; wherein the basic reward model is a regression model.

[0077] Optionally, the verifiable correctness signal reward model includes: a router and multiple verification agents; the verification agents include: a factual verification agent and an instruction compliance verification agent; the factual verification agent is used to evaluate whether the factual information contained in the model response is correct; the instruction compliance verification agent is used to evaluate whether the model response satisfies the hard constraints specified in the instruction; the reward scoring module 602 is specifically used to generate corresponding verification instructions based on the model response to be evaluated, and input the verification instructions into the router, select the verification agent through the router, and filter out the target agent to be used from the multiple verification agents; the target agent includes at least one verification agent; the reward scoring module 602 is also specifically used to use the target agent to verify the correctness of the model response to be evaluated, and obtain the reward score output by the target agent.

[0078] Optionally, the reward scoring module 602 is specifically used to input the verification instructions and the description information of each verification agent in the multiple verification agents into a natural language model, and the natural language model filters out matching target agents from the multiple verification agents based on the input instructions and the description information of each verification agent; wherein the description information includes: the functions and usage conditions of the verification agent.

[0079] Optionally, the target agent includes: the factual verification agent; the model response to be evaluated includes: two different model responses; the reward scoring module 602 is specifically used to input the two different model responses into the factual verification agent, identify the difference in the claimed facts between the two different model responses, obtain a difference identification result, and generate a corresponding difference query based on the difference identification result; the reward scoring module 602 is also specifically used to use the difference query to retrieve evidence, and to factually verify the difference between the two different model responses based on the retrieved evidence, and to score each model response in the two different model responses according to the verification result to obtain a score for each model response; the reward scoring module 602 is also specifically used to determine the reward score of the factual verification agent for the model response to be evaluated based on the score of each model response; wherein the score of the model response with correct facts is higher than the score of the model response with incorrect facts.

[0080] Optionally, the target agent includes: the instruction compliance verification agent; the reward scoring module 602, which is specifically used to extract at least one constraint condition for the model response from the verification instruction, and generate a corresponding constraint verification script based on the constraint condition; the reward scoring module 602 is also specifically used to use the constraint verification script to perform constraint verification on the response of the model to be evaluated, determine whether the response of the model to be evaluated satisfies each constraint condition in the at least one constraint condition, and obtain the constraint verification result; the reward scoring module 602 is also specifically used to calculate the average value of the score for each constraint condition in the constraint verification result, and obtain the reward score of the instruction compliance verification agent for the response of the model to be evaluated.

[0081] Optionally, the reward scoring module 602 is specifically used to obtain the reward score output by the target agent by weighted summing up the reward scores output by each verification agent among the multiple verification agents; wherein the reward score output by the target agent is: the reward score output by the verifiable correctness signal reward model for the response of the model to be evaluated.

[0082] The reliable reward evaluation device provided by the present application combines human preferences with verifiable correctness signals. First, the response of the model to be evaluated is obtained, and the response of the model to be evaluated is respectively input into the basic reward model and the verifiable correctness signal reward model to obtain the reward score output by each model; then, a weighted sum is performed based on the reward score output by each model and the weight of each model to obtain the final reward score for the response of the model to be evaluated; wherein, the basic reward model is used to characterize the human preference score for the model response; and the verifiable correctness signal reward model is used to characterize the correctness of the model response in specific aspects. In this way, not only human preferences can be considered, but also verifiable correctness signals such as the factuality of the model response and instruction compliance can be comprehensively considered, thereby providing more accurate and reliable rewards.

[0083] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720 and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute a reliable reward evaluation method combining human preferences with verifiable correctness signals, the method comprising: first, obtaining the response of the model to be evaluated, and inputting the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively, to obtain the reward score output by each model; then, performing weighted summation based on the reward score output by each model and the weight of each model, to obtain the final reward score for the response of the model to be evaluated; wherein the basic reward model is used to characterize the human preference score for the model response; the verifiable correctness signal reward model is used to characterize the correctness of the model response in a specific aspect. This not only takes into account human preferences, but also verifiable correctness signals such as the factuality of the model's responses and instruction compliance, thereby providing more accurate and reliable rewards.

[0084] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0085] On the other hand, the present application also provides a computer program product, the computer program product includes a computer program stored on a computer-readable storage medium, the computer program includes program instructions, when the program instructions are executed by the computer, the computer can execute the reliable reward evaluation method combining human preferences and verifiable correctness signals provided by the above methods, the method comprising: first, obtaining the response of the model to be evaluated, and inputting the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively, to obtain the reward score output by each model; then, based on the reward score output by each model and the weight of each model, a weighted sum is performed to obtain the final reward score for the response of the model to be evaluated; wherein, the basic reward model is used to characterize the human preference score for the model response; the verifiable correctness signal reward model is used to characterize the correctness of the model response in a specific aspect. In this way, not only human preferences can be considered, but also verifiable correctness signals such as the factuality of the model response and instruction compliance can be comprehensively considered, thereby providing more accurate and reliable rewards.

[0086] On the other hand, the present application also provides a computer-readable storage medium having a computer program stored thereon, which is implemented when the processor executes the above-mentioned reliable reward evaluation method combining human preferences and verifiable correctness signals, the method comprising: first, obtaining the response of the model to be evaluated, and inputting the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively, to obtain the reward score output by each model; then, performing weighted summation based on the reward score output by each model and the weight of each model, to obtain the final reward score for the response of the model to be evaluated; wherein the basic reward model is used to characterize the human preference score for the model response; the verifiable correctness signal reward model is used to characterize the correctness of the model response in a specific aspect. In this way, not only human preferences can be considered, but also verifiable correctness signals such as the factuality of the model response and instruction compliance can be comprehensively considered, thereby providing more accurate and reliable rewards.

[0087] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0088] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A reliable reward evaluation method combining human preferences with verifiable correctness signals, characterized in that include: Obtaining the response of the model to be evaluated, and inputting the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively, to obtain the reward score output by each model; Performing a weighted sum based on the reward score output by each model and the weight of each model to obtain a final reward score for the response of the model to be evaluated; The basic reward model is used to characterize human preference scores for model responses; the verifiable correctness signal reward model is used to characterize the correctness of model responses in specific aspects.

2. The method according to claim 1, characterized in that: The step of inputting the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively to obtain the reward score output by each model includes: Inputting the response of the model to be evaluated into the basic reward model to obtain a preference score of a human for the response of the model to be evaluated, and determining the preference score as a reward score output by the basic reward model; Wherein, the basic reward model is a regression model.

3. The method according to claim 1, characterized in that: The verifiable correctness signal reward model includes: a router and a plurality of verification agents; the verification agents include: a factual verification agent and an instruction compliance verification agent; the factual verification agent is used to evaluate whether the factual information contained in the model response is correct; the instruction compliance verification agent is used to evaluate whether the model response meets the hard constraint conditions specified in the instruction; The step of inputting the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively to obtain the reward score output by each model includes: Generate a corresponding verification instruction based on the response of the model to be evaluated, and input the verification instruction into the router, select the verification agent through the router, and screen out the target agent to be used from the multiple verification agents; the target agent includes at least one verification agent; The target agent is used to verify the correctness of the response of the model to be evaluated, and a reward score output by the target agent is obtained.

4. The method according to claim 3, characterized in that The step of inputting the verification instruction into the router, selecting a verification agent through the router, and selecting a target agent to be used from the multiple verification agents includes: Inputting the verification instruction and the description information of each verification agent in the multiple verification agents into a natural language model, and the natural language model screens out a matching target agent from the multiple verification agents according to the input instruction and the description information of each verification agent; The description information includes: functions and usage conditions of the verification agent.

5. The method according to claim 3 or 4, characterized in that: The target agent includes: the factual verification agent; the model response to be evaluated includes: two different model responses; The using the target agent to verify the correctness of the response of the model to be evaluated to obtain the reward score output by the target agent includes: Inputting the two different model responses into the factual verification agent, identifying differences in the claimed facts between the two different model responses, obtaining difference identification results, and generating corresponding difference queries based on the difference identification results; Using the difference query to perform evidence retrieval, and based on the retrieved evidence, perform fact verification on the difference between the two different model responses, and score each of the two different model responses according to the verification result to obtain a score for each model response; Determining a reward score of the factual verification agent for the response of the model to be evaluated based on the score of each model response; Among them, the scores of model responses that are factually correct are higher than the scores of model responses that are factually incorrect.

6. The method according to claim 3 or 4, characterized in that: The target agent includes: the instruction compliance verification agent; The using the target agent to verify the correctness of the response of the model to be evaluated to obtain the reward score output by the target agent includes: Extracting at least one constraint condition for the model response from the verification instruction, and generating a corresponding constraint verification script based on the constraint condition; Using the constraint verification script to perform constraint verification on the response of the model to be evaluated, determining whether the response of the model to be evaluated satisfies each constraint condition in the at least one constraint condition, and obtaining a constraint verification result; An average value of the score for each constraint condition in the constraint verification result is calculated to obtain a reward score for the instruction compliance verification agent's response to the model to be evaluated.

7. The method according to claim 3, characterized in that The using the target agent to verify the correctness of the response of the model to be evaluated to obtain the reward score output by the target agent includes: Obtaining the reward score output by the target agent by weighted summing the reward scores output by each of the multiple verification agents; The reward score output by the target agent is: the reward score output by the verifiable correctness signal reward model for the response of the model to be evaluated.

8. A reliable reward evaluation device that combines human preferences with verifiable correctness signals, characterized in that: The device comprises: An acquisition module is used to obtain the response of the model to be evaluated; A reward scoring module, used to input the response of the model to be evaluated into the basic reward model and the verifiable correctness signal reward model respectively, to obtain a reward score output by each model; The reward scoring module is further used to perform weighted summation based on the reward scores output by each model and the weight of each model to obtain a final reward score for the response of the model to be evaluated; The basic reward model is used to characterize human preference scores for model responses; the verifiable correctness signal reward model is used to characterize the correctness of model responses in specific aspects.

9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the reliable reward evaluation method combining human preference with verifiable correctness signal as claimed in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the reliable reward evaluation method combining human preferences with verifiable correctness signals as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Model evaluation method and device, terminal and storage medium

    CN116578467A

  • Evaluation and verification method and device for perception type model

    CN117493826A

  • Reward model training method, answer evaluation method, device and equipment

    CN117688158A

  • Large language model evaluation method and device, electronic equipment and storage medium

    CN118035807A

  • Method and equipment for optimizing instruction following capability of large language model and medium

    CN119129754A

Cited By

  • Intelligent agent model adaptive optimization method and system based on error feedback information

    CN121434371A

  • Intelligent agent model self-adaptive optimization method and system based on error feedback information

    CN121434371B