Training method, answer evaluation method, device and equipment for reward model
By introducing the target score evaluation of knowledge graphs and sample pairs in the reward model, a reward model that can accurately identify the correctness of knowledge answers to large language models is trained, which solves the problem of low knowledge accuracy assessment in the existing technology and improves the output quality of large language models.
Patent Information
- Application Number
- CN202311828971.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-12-26
AI Technical Summary
The existing reward model is not very accurate when evaluating the knowledge correctness of the output text of the large language model, resulting in insufficient accuracy in identifying factual content by the large language model.
By acquiring multiple sample pairs, using the knowledge graph to determine the target score of the sample answer, and adjusting the model parameters based on the score of the initial reward model, a reward model that can focus on the correctness of the answer knowledge is trained.
It improves the accuracy of knowledge correctness recognition of text output by large language models, ensures that the answers are more in line with human preferences, and is suitable for human-computer interaction, automatic question-and-answer systems, and information retrieval.
Smart Images

Figure CN117688158B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a training method for a reward model, an answer evaluation method, device, and equipment. Background Art
[0002] A knowledge-enhanced reward model can enable a large language model to better combine knowledge correctness in the process of aligning with human needs to output a scoring result, thereby solving the knowledge hallucination problem in the large language model.
[0003] Traditional reward models usually give comprehensive feedback from the semantic level of the text, such as feedback based on aspects such as the fluency, helpfulness, and safety of the text output by the large language model. However, such a reward model has defects in evaluating the knowledge correctness of the text output by the large language model, resulting in a problem of low accuracy in the recognition of factual content by the large language model. Summary of the Invention
[0004] The present invention provides a training method for a reward model, an answer evaluation method, device, and equipment, to solve the defect that the existing reward model has low accuracy in recognizing the knowledge correctness of the text output by the large language model, and to achieve an improvement in the recognition accuracy of the knowledge correctness of the text output by the large language model.
[0005] The present invention provides a training method for a reward model, including:
[0006] Obtain a plurality of sample pairs, each of the sample pairs including a first sample and a second sample, the first sample including a sample question and a first sample answer, the second sample including the sample question and a second sample answer, the target score of the first sample answer being higher than the target score of the second sample answer, the target score being related to knowledge correctness, and the knowledge correctness being determined based on a target answer matching the sample question in a knowledge graph;
[0007] For each of the sample pairs, input the first sample and the second sample in the sample pair into an initial reward model to obtain a first score of the first sample and a second score of the second sample output by the initial reward model;
[0008] Based on the first score and the second score, adjust the model parameters of the initial reward model to obtain a reward model, and the reward model is used to evaluate the answer output by the large language model.
[0009] According to the training method for a reward model provided by the present invention, the adjusting the model parameters of the initial reward model based on the first score and the second score to obtain a reward model includes:
[0010] Determine a first knowledge correctness judgment result of the first sample answer relative to the target answer, and a second knowledge correctness judgment result of the second sample answer relative to the target answer;
[0011] Determine a quantitative gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result;
[0012] Determine loss information based on the first score, the second score, and the quantitative gap;
[0013] Adjust model parameters of the initial reward model based on the loss information to obtain the reward model.
[0014] According to a training method of a reward model provided by the present invention, the method further includes:
[0015] Obtain multiple sample answers corresponding to the sample question;
[0016] Search for the target answer corresponding to the sample question in the knowledge graph;
[0017] Based on the target answer, determine target scores of the sample answers.
[0018] According to a training method of a reward model provided by the present invention, the determining target scores of the sample answers based on the target answer includes:
[0019] Determine target knowledge correctness judgment results of the sample answers relative to the target answer;
[0020] Based on the target knowledge correctness judgment result and target evaluation dimensions of the target answer, determine annotation scores of the sample answers; the target evaluation dimensions include at least one of the following: fluency of the sample answer, user satisfaction with the sample answer, and safety of the sample answer.
[0021] According to a training method of a reward model provided by the present invention, the determining target knowledge correctness judgment results of the sample answers relative to the target answer includes:
[0022] For each of the sample answers, when it is determined that the sample answer includes a first sub - answer and a second sub - answer other than the first sub - answer, the first sub - answer includes an answer corresponding to the sample question;
[0023] Based on the target answer, determine a knowledge correctness judgment result of the first sub - answer and a knowledge correctness judgment result of the second sub - answer respectively;
[0024] Based on the knowledge correctness judgment result of the first sub-answer and the knowledge correctness judgment result of the second sub-answer, determine the target knowledge correctness judgment result corresponding to the sample answer.
[0025] According to a training method for a reward model provided by the present invention, the determining the target knowledge correctness judgment result corresponding to the sample answer based on the knowledge correctness judgment result of the first sub-answer and the knowledge correctness judgment result of the second sub-answer includes:
[0026] When the knowledge correctness judgment result of the first sub-answer is correct and the knowledge correctness judgment result of the second sub-answer is correct, determine that the target knowledge correctness judgment result corresponding to the sample answer is correct;
[0027] When the knowledge correctness judgment result of the first sub-answer is correct and the knowledge correctness judgment result of the second sub-answer is wrong, determine that the target knowledge correctness judgment result corresponding to the sample answer is partially correct;
[0028] When the knowledge correctness judgment result of the first sub-answer is wrong, determine that the target knowledge correctness judgment result corresponding to the sample answer is wrong.
[0029] The present invention provides an answer evaluation method, including:
[0030] Obtain a predicted answer for a target question output by a large language model;
[0031] Input the target question and the predicted answer into a reward model to obtain an evaluation result of the predicted answer output by the reward model, where the reward model is trained according to the training method for the reward model described in any one of the above.
[0032] The present invention further provides a training device for a reward model, including:
[0033] An acquisition module, configured to acquire a plurality of sample pairs, each sample pair including a first sample and a second sample, the first sample including a sample question and a first sample answer, the second sample including the sample question and a second sample answer, the marked score of the first sample answer being higher than the target score of the second sample answer, the target score being related to knowledge correctness, and the knowledge correctness being determined based on a target answer matching the sample question in a knowledge graph;
[0034] An input module, configured to input the first sample and the second sample in each sample pair into an initial reward model for each sample pair, to obtain a first score of the first sample and a second score of the second sample output by the initial reward model;
[0035] An adjustment module, configured to adjust model parameters of the initial reward model based on the first score and the second score to obtain a reward model, where the reward model is used to evaluate answers output by a large language model.
[0036] The present invention further provides an answer evaluation device, including:
[0037] An acquisition module, configured to acquire a predicted answer to a target question output by a large language model;
[0038] An input module, configured to input the target question and the predicted answer into the reward model to obtain an evaluation result of the predicted answer output by the reward model, where the reward model is trained based on the training method of the reward model described in any one of the above.
[0039] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the training method of the reward model described in any one of the above or implements the answer evaluation method described in any one of the above.
[0040] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the training method of the reward model described in any one of the above or implements the answer evaluation method described in any one of the above.
[0041] The present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the training method of the reward model described in any one of the above or implements the answer evaluation method described in any one of the above.
[0042] The training method, answer evaluation method, device and equipment of the reward model provided by the present invention obtain multiple sample pairs. Each sample pair includes a first sample and a second sample. The first sample includes a sample question and a first sample answer, and the second sample includes a sample question and a second sample answer. The target score of the first sample answer is higher than the target score of the second sample answer. The target score is related to the knowledge correctness, and the knowledge correctness is determined based on the target answer matching the sample question in the knowledge graph. For each sample pair, the first sample and the second sample in the sample pair are input into the initial reward model to obtain the first score of the first sample and the second score of the second sample output by the initial reward model, and based on the first score and the second score, the model parameters of the initial reward model are adjusted to obtain the reward model, which is used to evaluate the answers output by the large language model. Since in the process of training the reward model, the knowledge correctness of the first sample answer and the second sample answer in the sample pair is determined based on the target answer in the knowledge graph, and the target answer in the knowledge graph is a knowledge-correct answer, the target score determined based on the knowledge correctness is related to the knowledge correctness of the first sample answer and the second sample answer. The higher the knowledge correctness of the sample answer, the higher the target score. Therefore, when training the reward model based on the sample pair, the first sample answer with a higher target score and the second sample answer with a lower target score can be widened, so that the trained reward model can pay more attention to the knowledge correctness of the answer when evaluating the answers output by the large language model, thereby improving the recognition accuracy of the knowledge correctness of the answer. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 It is a schematic flowchart of the training method of the reward model provided by the embodiment of the present invention;
[0045] Figure 2 It is a schematic diagram of the training process of the reward model provided by the embodiment of the present invention;
[0046] Figure 3 It is a schematic diagram of the training process and fine-tuning process of the reward model provided by the embodiment of the present invention;
[0047] Figure 4 It is a schematic flowchart of the answer evaluation method provided by the embodiment of the present invention;
[0048] Figure 5It is a schematic structural diagram of a training device for a reward model provided by an embodiment of the present invention;
[0049] Figure 6 It is a schematic structural diagram of an answer evaluation device provided by an embodiment of the present invention;
[0050] Figure 7 It exemplifies a schematic structural diagram of an electronic device. Detailed implementation manners
[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.
[0052] The reward model is an important part in the field of deep learning. The reward model is used to measure the quality of the answers generated by the large language model, and provide feedback and guidance for the large language model to improve its performance and output quality.
[0053] Traditional reward models usually give comprehensive feedback based on semantic aspects such as the fluency, helpfulness, and safety of the text output by the large language model. However, this feedback has defects in evaluating the knowledge correctness of the text output by the large language model, which will cause problems with low accuracy in the large language model's recognition of factual content. Therefore, how to more accurately evaluate the output results of the large language model through a knowledge-enhanced reward model without affecting the original learning process and model structure, and how to make full use of the enhanced reward model to align the answers output by the large language model with human needs, so that the answers output by the large language model are more in line with human preferences, remains an urgent problem to be solved.
[0054] In the embodiments of the present invention, in view of the above problems, a training method for a reward model is proposed. In this method, the knowledge correctness of the first sample answer and the second sample answer corresponding to the sample question can be judged based on the target answer matching the sample question in the knowledge graph, and the corresponding target scores can be determined based on this knowledge correctness, so as to obtain a plurality of sample pairs. Each sample pair includes a first sample and a second sample. The first sample includes the sample question and the first sample answer, and the second sample includes the sample question and the second sample answer. Among them, the target score of the first sample answer is higher than the target score of the second sample answer, and the target score is related to the knowledge correctness. The higher the knowledge correctness, the higher the target score. After respectively inputting the obtained sample pairs into the initial reward model, the initial reward model can score the first sample and the second sample, so as to train the initial reward model based on the first score of the first sample and the second score of the second sample to obtain the reward model. Since in the process of training this reward model, the knowledge correctness of the first sample answer and the second sample answer in the sample pair is determined based on the target answer in the knowledge graph, and the target answer in this knowledge graph is a knowledge-correct answer, the target score determined based on the knowledge correctness is related to the knowledge correctness of the first sample answer and the second sample answer. Therefore, when training the reward model based on the sample pair, the gap between the first sample answer with a higher target score and the second sample answer with a lower target score can be widened, so that when the trained reward model evaluates the answer output by the large language model, it can pay more attention to the knowledge correctness of the answer, thereby improving the recognition accuracy of the knowledge correctness of the answer.
[0055] The following will describe Figures 1 to 4 the training method for the reward model provided by the embodiments of the present invention. The embodiments of the present invention can be applied to scenarios where the text output by any model is evaluated, especially to scenarios where the output answers of large language models are evaluated for factual correctness or knowledge correctness. The execution subject of this method can be an electronic device such as a computer, a terminal device, a server, a server cluster, or a training device for a specially designed reward model, or it can be a training device for a reward model set in this electronic device. This training device for a reward model can be implemented by software, hardware, or a combination of both.
[0056] Figure 1 is a schematic flowchart of the training method for the reward model provided by the embodiments of the present invention. As Figure 1 shown, this method includes:
[0057] Step 101: Obtain multiple sample pairs. Each sample pair includes a first sample and a second sample. The first sample includes a sample question and a first sample answer, and the second sample includes a sample question and a second sample answer. The target score of the first sample answer is higher than that of the second sample answer. The target score is related to the knowledge correctness, and the knowledge correctness is determined based on the target answer matched with the sample question in the knowledge graph.
[0058] In this step, the sample question is the question collected by the user to query the large language model, or it can also be other questions collected from the network. By selecting different large language models or changing the model parameters of the same large language model, the collected sample questions are input into each large language model, so as to obtain the sample answers output by each large language model. For example, two different large language models can be selected, and the sample questions are respectively input into each large language model, and a first sample answer output by one large language model and a second sample answer output by another large language model can be obtained. These large language models can be variants developed by different institutions or the same large language model, and the temperature parameter is changed to maximize the diversity. Among them, the large language model can be, for example, the IFlytek Spark or other models capable of knowledge Q&A.
[0059] To enhance the model's ability to judge knowledge correctness, in the embodiment of the present invention, an external knowledge base is introduced, and the knowledge in the external knowledge base is stored in the form of a knowledge graph. The knowledge graph is a structured graph that can represent knowledge with entities and relationships. The present invention utilizes this structural feature of the knowledge graph to strengthen the reward mechanism of the model with the entities and relationships in the graph, so that the reward model can more significantly identify and extract various knowledge concepts and ideas. Therefore, the sample question can be combined with the knowledge graph to extract triples, and the content extracted from the knowledge graph in the triples is the target answer matched with the sample question, and the target answer can be understood as the knowledge-correct answer determined based on the knowledge graph.
[0060] Furthermore, after obtaining the first sample answer and the second sample answer, the knowledge correctness of the first sample answer and the second sample answer can be determined based on the determined target answer, and the target scores of the first sample answer and the second sample answer can be determined based on the knowledge correctness. Among them, the target score of the first sample answer is higher than that of the second sample answer, indicating that the first sample answer is closer to the target answer, or it can also be understood that the factual correctness of the first sample answer is higher than that of the second sample answer, or the knowledge correctness of the first sample answer is higher than that of the second sample answer.
[0061] After determining the target scores of the first sample answer and the second sample answer, the sample question and the first sample answer can be used as the first sample, and the sample question and the second sample answer can be used as the second sample, thereby forming a sample pair.
[0062] Step 102: For each sample pair, input the first sample and the second sample in the sample pair into the initial reward model to obtain the first score of the first sample and the second score of the second sample output by the initial reward model.
[0063] In this step, after obtaining multiple sample pairs, the first sample and the second sample in each sample pair can be input into the initial reward model. Through this initial reward model, the first sample answer will be scored based on the sample question in the first sample to obtain the first score. Similarly, the second sample answer will also be scored based on the sample question in the second sample to obtain the second score. Among them, the sample questions in the first sample and the second sample are the same. The first score represents the evaluation result of the initial reward model on the first sample answer, and the second score represents the evaluation result of the initial reward model on the second sample answer.
[0064] In the specific implementation process, for each sample pair, the sample question and the first sample answer in the first sample can be concatenated, and the sample question and the second sample answer in the second sample can be concatenated, and the concatenated content can be input into the initial reward model to score the first sample and the second sample. In addition, the labeled scores corresponding to the first sample answer and the second sample answer can be converted into a binary label format, and different scores are forced between the two.
[0065] Step 103: Based on the first score and the second score, adjust the model parameters of the initial reward model to obtain a reward model, which is used to evaluate the answers output by the large language model.
[0066] In this step, the loss information can be determined based on the first score and the second score, and then the model parameters of the initial reward model can be adjusted based on this loss information. By continuously iterating the above process until the model converges or the number of iterations reaches the preset number, the finally obtained model is used as the reward model. Among them, the trained reward model can be used to evaluate the answers output by the large language model, and the evaluation result includes the evaluation of the knowledge correctness of the answer.
[0067] Exemplarily, the above loss information can adopt binary ranking loss, specifically as shown in formula (1):
[0068] Loss=-log(σ(r θ (x,y h )-r θ (x,y l))) (1)
[0069] Among them, Loss represents loss information, x represents a sample problem, which includes triples extracted by combining a knowledge graph, y h represents the first sample answer, y l represents the second sample answer, r θ (x, y h ) represents the first score of the first sample of the weight θ of the initial reward model, r θ (x, y l ) represents the second score of the second sample of the weight θ of the initial reward model, and σ(.) represents normalization processing.
[0070] The training method of the reward model provided by the embodiment of the present invention obtains multiple sample pairs through each sample pair including a first sample and a second sample. The first sample includes a sample problem and a first sample answer, and the second sample includes a sample problem and a second sample answer. The target score of the first sample answer is higher than the target score of the second sample answer. The target score is related to the knowledge correctness, and the knowledge correctness is determined based on the target answer matching the sample problem in the knowledge graph. For each sample pair, the first sample and the second sample in the sample pair are input into the initial reward model to obtain the first score of the first sample and the second score of the second sample output by the initial reward model, and based on the first score and the second score, the model parameters of the initial reward model are adjusted to obtain a reward model, which is used to evaluate the answers output by the large language model. Since in the process of training the reward model, the knowledge correctness of the first sample answer and the second sample answer in the sample pair is determined based on the target answer in the knowledge graph, and the target answer in the knowledge graph is a knowledge-correct answer, the target score determined based on the knowledge correctness is related to the knowledge correctness of the first sample answer and the second sample answer. The higher the knowledge correctness of the sample answer, the higher the target score. Therefore, when training the reward model based on the sample pair, the first sample answer with a higher target score and the second sample answer with a lower target score can be made to widen the gap, so that when the trained reward model evaluates the answers output by the large language model, it can pay more attention to the knowledge correctness of the answers, thereby improving the recognition accuracy of the knowledge correctness of the answers.
[0071] Exemplarily, based on the above embodiments, when adjusting the model parameters of the initial reward model based on the first score and the second score to obtain the reward model, the first knowledge correctness judgment result of the first sample answer relative to the target answer and the second knowledge correctness judgment result of the second sample answer relative to the target answer can be determined, and the quantitative gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result can be determined. The loss information is determined based on the first score, the second score, and the quantitative gap, so as to adjust the model parameters of the initial reward model based on the loss information to obtain the reward model.
[0072] In this step, in order to introduce the evaluation of the correctness of knowledge into the reward model, in the embodiments of the present invention, an auxiliary task of knowledge judgment is further added. Specifically, it is a further modification based on the binary ranking loss in the foregoing embodiments to better enable the reward model to learn the ability to judge the correctness of answers by combining knowledge. Therefore, it is necessary to determine the first knowledge correctness judgment result of the first sample answer relative to the target answer and the second knowledge correctness judgment result of the second sample answer relative to the target answer. Among them, the first knowledge correctness judgment result and the second knowledge correctness judgment result are mainly divided into three categories: correct, partially correct, and wrong. Taking the first knowledge correctness judgment result as an example, stricter requirements are imposed on the most core part of the sample question, while the supplementary information output by the large language model can be appropriately relaxed. For example, the first sample answer and the target answer can be compared. When the part of the first sample answer corresponding to the sample question input by the user is answered correctly and the supplementary information added by the large language model is also correct, the first knowledge correctness judgment result can be determined as "correct"; when the part of the first sample answer corresponding to the sample question input by the user is answered correctly, but the supplementary information added by the large language model is wrong, the first knowledge correctness judgment result can be determined as "partially correct"; when the part of the first sample answer corresponding to the sample question input by the user is answered wrong, the first knowledge correctness judgment result can be determined as "wrong". Similarly, the second knowledge correctness judgment result of the second sample answer relative to the target answer can be determined in the above manner.
[0073] Furthermore, the quantitative gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result can be determined. For example, the size of the quantitative gap r can be determined according to the three knowledge labels of "correct", "partially correct", and "wrong". Among them, the value of r is in [0, 2]. For example, if the first knowledge correctness judgment result is "correct" and the second knowledge correctness judgment result is "wrong", then r is 2; if the first knowledge correctness judgment result is "correct" and the second knowledge correctness judgment result is "partially correct", then r is 1; if the first knowledge correctness judgment result is "partially correct" and the second knowledge correctness judgment result is "wrong", then r is 0.8, and so on.
[0074] After determining the quantization gap, the loss information can be determined based on the following formula (2):
[0075] Loss = -log(σ(r θ (x, y h ) - r θ (x, y l ) - m k (r))) (2)
[0076] Among them, the function m k (r) represents a discrete function of the reply correctness deviation, which is used to measure the knowledge correctness gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result, where r represents the quantization gap.
[0077] It can be understood that by adding m k (r) to the loss information, in the process of training the reward model through this multi-task learning method, the model parameters of the initial reward model can be adjusted based on the determined loss information. This method can widen the scores of the sample answers of correct knowledge and wrong knowledge, enabling the trained reward model to more explicitly learn how to combine the knowledge graph to judge the knowledge correctness of the answers output by the large language model.
[0078] In the embodiments of the present invention, by adding the judgment result of knowledge correctness, the reward model can more prominently combine the knowledge graph in the external knowledge base to enhance the knowledge judgment of the reward model. Without relying on the external knowledge base, although the reward model can also make judgments, the reward model is usually small and unstable, and its knowledge storage is often not as good as that of the large model. When encountering unseen or ambiguous knowledge, the model is prone to knowledge hallucinations, such as misattributing or fabricating out of thin air. After adding the external knowledge base, the reward model still needs to learn how to use this knowledge to judge the correctness of the reply. Therefore, after adding the loss of m k (r), the judgment of the auxiliary model can be carried out, and the reward model can more directly combine the three criteria of "correct", "partially correct" and "wrong" to judge the knowledge of the sample answers, which can prevent the reward model from mis-scoring the replies with wrong knowledge and aggravating the knowledge hallucinations of the large language model.
[0079] Furthermore, after training the reward model, the reward model can also be fine-tuned based on the evaluation result of the answers of the large language model. For example, the proximal policy optimization (PPO) method can be used to fine-tune the reward model through reinforcement learning, so as to align the large language model with humans. Using human feedback as the signal for reinforcement learning enables the answers finally generated by the large language model to be optimized not only in terms of relevance to the questions input by the user, but also to better meet the needs and preferences of human users.
[0080] In this embodiment, based on the first knowledge correctness judgment result of the first sample answer relative to the target answer and the second knowledge correctness judgment result of the second sample answer relative to the target answer, after determining the quantitative gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result, the loss information can be determined based on the first score, the second score, and the quantitative gap, and then the model parameters of the initial reward model can be adjusted based on the loss information to obtain the final reward model. Since the quantitative gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result is added to the loss information, it can not only enable the reward model to more explicitly learn how to combine the knowledge graph to judge the correctness of the answers output by the large language model, but also enable the reward model to more directly judge the knowledgeability of the answers based on the three criteria of "correct", "partially correct", and "wrong", and can prevent the phenomenon that the reward model gives a high score to the wrong answer of knowledge, thus aggravating the knowledge hallucination of the large language model.
[0081] In addition, by using the reward model provided in the embodiments of the present invention to evaluate the answers of the large language model, the large language model can be made more in line with human preferences. For example, in fields such as human-computer interaction, automatic question answering systems, and information retrieval, it can play an important role, which is of great significance to improving the application value of artificial intelligence in the current social development.
[0082] It should be understood that the target score of the first sample answer can represent the accuracy of the knowledge of the first sample answer. Therefore, it is very important to determine an accurate target score. Next, the determination method of the target score of the sample answers included in each sample pair will be described in detail. Among them, the sample answers described in the following embodiments can include the first sample answer or the second sample answer, or the sample answers in other sample pairs.
[0083] Exemplarily, when determining the target score of the sample answer, it can be done in the following way: obtain multiple sample answers corresponding to the sample question, and search for the target answer corresponding to the sample question in the knowledge graph, and determine the target score of each sample answer based on the target answer.
[0084] Specifically,Figure 2 Schematic diagram of the training process of the reward model provided by an embodiment of the present invention, as Figure 2 shown, a Prompt can be obtained. For example, it can be obtained through open-source data or self-built collection. Among them, the collected Prompts usually have a wide coverage and diverse forms. After collecting the Prompts, the obtained sample questions and the collected Prompts can be input into different large language models, so as to obtain different sample answers output by each large language model. Among them, for each Prompt, the user can perform preference annotation on the sample answers output by these different large language models to provide learning for the reward model.
[0085] For example, if the sample question is "How many nanometers of chips are installed in the newly released XX mobile phone", after inputting this sample question and the collected Prompt into four different large language models, four sample answers can be obtained, namely A: "The newly released XX mobile phone is equipped with an M1 Pro chip. According to the official data of XX mobile phone, this chip is manufactured by a 3nm process technology...", B: "XX company has not released the XX mobile phone yet. Therefore, it is impossible to determine the nanometer process technology of the chip it is equipped with...", C: "The newly released XX mobile phone is equipped with an M1 Pro or M1 Max chip, and both of these chips are manufactured using TSMC's 5nm process technology.", D: "The A17 Pro chip installed in the newly released XX mobile phone uses a 3nm (nanometer) process, and the design of this chip is very complex...".
[0086] Furthermore, for a sample question, the sample question can be combined with a knowledge graph to extract entities and relationships to form a triple corresponding to the sample question, such as entity-attribute-attribute value. Among them, any method can be used to extract entities and relationships. For example, named entity recognition (NER) technology and relation extraction (RE) technology can be used for extraction. NER is used to identify entities in the sample question, and RE is used to determine the relationships between these entities. It can be understood that the entities and attributes in the above-extracted triples can be extracted from the sample question, and the attribute values are obtained from the knowledge graph. This attribute value can be understood as the target answer corresponding to the sample question. For example, for the sample question "How many nanometers of chips are installed in the newly released XX mobile phone", the triple (XX mobile phone, chip, A17 Pro chip, 3nm) can be extracted by combining the knowledge graph. Among them, XX mobile phone is the entity extracted from the sample question, chip is the attribute extracted from the sample question, and A17 Pro chip and 3nm are the attribute values obtained from the knowledge graph, which are the target answers corresponding to the sample question.
[0087] Furthermore, the target scores of each sample answer can be determined based on the target answer. In this way, the reward model can indirectly learn to distinguish the severity of different knowledge errors, and thus the evaluation scores of the large language model will also change accordingly.
[0088] In this embodiment, the target answer corresponding to the sample question can be found from the knowledge graph, and thus the target scores of each sample answer can be determined based on the target answer. Since the knowledge of the target answer determined from the knowledge graph is correct, the accuracy of the determined target scores can be improved.
[0089] Exemplarily, in the above embodiment, when determining the target scores of each sample answer based on the target answer, the target knowledge correctness judgment results of each sample answer relative to the target answer can be determined, and the target scores of each sample answer can be determined based on the target knowledge correctness judgment result of the target answer and the target evaluation dimension; the target evaluation dimension includes at least one of the following: the fluency of the sample answer, the user's satisfaction with the sample answer, and the security of the sample answer.
[0090] Specifically, the sample answer and the target answer can be compared to determine the target knowledge correctness judgment results of each sample answer relative to the target answer. In one possible implementation, for each sample answer, it can be determined that the sample answer includes a first sub-answer and a second sub-answer other than the first sub-answer, and the first sub-answer includes the answer corresponding to the sample question; based on the target answer, the knowledge correctness judgment results of the first sub-answer and the second sub-answer are respectively determined; based on the knowledge correctness judgment results of the first sub-answer and the second sub-answer, the target knowledge correctness judgment result corresponding to the sample answer is determined.
[0091] Among them, in order to increase the richness of the answer and the fullness of the content, the large language model usually further makes additional supplements while answering the sample question. Therefore, the first sub-answer and the second sub-answer can be determined in each sample answer, where the first sub-answer is the answer corresponding to the sample question, and the second sub-answer is the content additionally supplemented by the large language model. For example, in the sample answer A “The newly released XX mobile phone is equipped with an M1 Pro chip. According to the official data of the XX mobile phone, this chip is a 3nm process technology...” in the above example, the first sub-answer is “3nm”, and the second sub-answer is “The XX mobile phone is equipped with an M1 Pro chip”.
[0092] Furthermore, the first sub-answer and the second sub-answer can be compared with the target answer respectively, so that the knowledge correctness judgment results of the first sub-answer and the second sub-answer can be obtained, and thus the target knowledge correctness judgment result corresponding to the sample answer can be determined according to the knowledge correctness judgment results of these two parts.
[0093] Since the knowledge correctness judgment results of the first sub-answer and the second sub-answer in the sample answer are determined separately to determine the target knowledge correctness judgment result of the sample answer, the determination result of the target knowledge correctness judgment result is made more fine-grained, further improving the accuracy of the target knowledge correctness judgment result. Additionally, this fine-grained target knowledge correctness judgment result enables the ultimately trained reward model to indirectly learn to distinguish the severity of different knowledge errors.
[0094] In addition, when determining the target knowledge correctness judgment result corresponding to the sample answer based on the knowledge correctness judgment results of the first sub-answer and the second sub-answer, it can be that when the knowledge correctness judgment result of the first sub-answer is correct and the knowledge correctness judgment result of the second sub-answer is correct, the target knowledge correctness judgment result corresponding to the sample answer is determined to be correct; when the knowledge correctness judgment result of the first sub-answer is correct and the knowledge correctness judgment result of the second sub-answer is wrong, the target knowledge correctness judgment result corresponding to the sample answer is determined to be partially correct; when the knowledge correctness judgment result of the first sub-answer is wrong, the target knowledge correctness judgment result corresponding to the sample answer is determined to be wrong.
[0095] For example, as Figure 2 shown, when comparing the first sub-answer and the second sub-answer in the above sample answers A, B, C, and D with the target answers A17 Pro chip and 3nm in the triple, it can be determined that the target knowledge correctness judgment result of sample answer A is "partially correct", the target knowledge correctness judgment result of sample answer B is "wrong", the target knowledge correctness judgment result of sample answer C is "wrong", and the target knowledge correctness judgment result of sample answer D is "correct".
[0096] By separately determining the knowledge correctness judgment results of the first sub-answer and the second sub-answer, a more accurate target knowledge correctness judgment result can be obtained. In this way, when training the reward model, adding the target knowledge correctness judgment result to the loss information will enable the trained reward model to more accurately evaluate the knowledge correctness of the answers output by the large language model.
[0097] Furthermore, after determining the target knowledge correctness judgment result of the sample answer, the annotation score of the sample answer can also be determined based on the target knowledge correctness judgment result of the target answer and the target evaluation dimension, so as to comprehensively score and rank the sample answers in combination with various standards. Among them, this process is an important process for aligning the large language model with human preferences. Therefore, comprehensive standards need to be formulated. In the embodiments of the present invention, in addition to the target knowledge correctness judgment result, other target evaluation dimensions mainly include the fluency of the sample answer, the user's satisfaction with the sample answer, and the safety of the sample answer.
[0098] Among them, the target knowledge correctness judgment result mainly characterizes the knowledgeability of the sample answer. This knowledgeability evaluation standard mainly focuses on whether the sample answer generated by the large language model is accurate, relevant, and has practical reference value. For example, if the target knowledge correctness judgment result of the sample answer is incorrect, the score of the sample answer regarding knowledgeability will not be higher than that of the sample answer with a correct target knowledge correctness judgment result.
[0099] The user's satisfaction with the sample answer can also be understood as the helpfulness of the sample answer. The helpfulness evaluation standard mainly focuses on whether the sample answer generated by the large language model can meet the user's needs, help the user solve problems, or provide useful suggestions. Under this standard, it will be evaluated whether the response of the large language model is closely related to the sample question proposed by the user and whether it provides a practical and feasible solution for the user.
[0100] The safety of the sample answer can also be understood as the security of the sample answer. The security evaluation standard focuses on whether the sample answer generated by the large language model contains content that may lead to adverse consequences such as misleading, discrimination, or provocation. To ensure the security of the large language model, the output sample answers will be screened to exclude information containing potential risks and mark the bad content for corresponding adjustment during the training process.
[0101] As Figure 2 shown, in the foregoing example, after determining the target scores of sample answers A, B, C, and D based on the target knowledge correctness judgment result and the target evaluation dimension, the ranking result obtained based on the target score is D > A > B = C.
[0102] After determining the ranking result among the sample answers, the first sample answer with a high target score and the second sample answer with a low target score can be determined based on the ranking result, so as to form the first sample and the second sample in combination with the sample question, and then train the initial reward model to obtain the final reward model.
[0103] In this embodiment, based on the target evaluation dimension, a target knowledge correctness judgment result for characterizing the knowledgeability of the target answer is further added. Therefore, the trained reward model can more accurately evaluate the knowledge correctness of the answers output by the large language model, effectively improving the performance of the large language model during application, enhancing its efficiency and effectiveness. Especially when dealing with complex and variable human language information, it can have stronger adaptability and response capabilities. Additionally, since the target score of the sample answer can be determined by combining the target evaluation dimension and the target knowledge correctness judgment result, and this target score is a comprehensive score, it can not only avoid giving high scores to sample answers with knowledge errors, thereby alleviating the knowledge hallucination problem that may be aggravated in the large language model alignment scheme, but also retain the effect of evaluation based on the target evaluation dimension, and can significantly improve the judgment of knowledge correctness, alleviating the knowledge hallucination problem of the large language model during the alignment process, and having good effectiveness and practicality.
[0104] It should be noted that the foregoing annotation process, that is, the process of determining the target score, can also be executed by annotators.
[0105] Figure 3 This is a schematic diagram of the training process and fine-tuning process of the reward model provided by the embodiment of the present invention. As Figure 3 shown, this process mainly includes an annotation process, a process of training the reward model, and a process of reinforcement learning fine-tuning.
[0106] For the annotation process, this stage mainly includes four steps: First, collect data of sample questions and generate several sample answers for comparison. Among them, the sample questions can be questions preferred by users; second, extract triples in the sample questions based on the knowledge graph; third, judge the correctness of the sample answers based on the extracted triples; fourth, comprehensively sort the sample answers. Among them, the above process can be executed by an electronic device or by annotators.
[0107] Reinforcement Learning with Human Feedback (RLHF) is a very important part in the current training of large language models, and it is particularly important in aligning human and model preferences and instructions. Specifically speaking, RLHF is a model training method. First, it is necessary to sample data that can represent human preferences, and then annotators choose which of the two model outputs they like. This human feedback is then used to train a reward model, which can automatically make preference decision scoring on behalf of humans after aligning with human preferences. It can be seen that the effect of RLHF highly depends on the reward model, and the scoring mechanism of the reward model must be consistent with human preferences. Therefore, the annotation process is the most important and also the most complex part of the whole process.
[0108] For the process of training the reward model, at this stage, the sample answers collected and sorted in the previous stage can be used as training samples to train the reward model. Through training optimization, the large language model can calculate the optimal answer when obtaining new questions.
[0109] In addition, to improve the accuracy of the reward model, reinforcement learning fine-tuning can also be performed. At this stage, the reinforcement learning method can be used to fine-tune and optimize the system in combination with the trained reward model. Using human feedback as the signal for reinforcement learning enables the large language model to finally generate answers that are not only optimized in terms of relevance to the user's input question but also better meet the needs and preferences of human users.
[0110] Figure 4 The flowchart of the answer evaluation method provided by the embodiment of the present invention is shown as Figure 4 shown, and the method includes:
[0111] Step 401: Obtain the predicted answer for the target question output by the large language model.
[0112] In this step, after obtaining the target question input by the user, the target question can be input into the large language model, so as to obtain the predicted answer output by the large language model. This predicted answer can be understood as the answer to the target question output by the large language model.
[0113] Step 402: Input the target question and the predicted answer into the reward model to obtain the evaluation result of the predicted answer output by the reward model.
[0114] Among them, the reward model is trained based on the training method of the reward model described in any of the above embodiments. The specific training process can refer to any of the foregoing embodiments, and the specific training process will not be elaborated here.
[0115] In this step, the predicted answer output by the large language model may have knowledge errors. Therefore, it is necessary to input the target question and the obtained predicted answer into the reward model, and the reward model judges the predicted answer to determine whether it has knowledge errors, so as to obtain the evaluation result output by the reward model. The evaluation result includes the judgment result on the knowledge correctness of the predicted answer, such as correct, partially correct, or wrong.
[0116] The answer evaluation method provided by the embodiments of the present invention obtains a predicted answer for a target question output by a large language model, and inputs the target question and the predicted answer into a reward model, so as to obtain an evaluation result of the predicted answer output by the reward model. Since in the training process of the reward model, the knowledge correctness of the first sample answer and the second sample answer in the sample pair is determined based on the target answer in the knowledge graph, and the target answer in the knowledge graph is a knowledge-correct answer, the target score determined based on the knowledge correctness is related to the knowledge correctness of the first sample answer and the second sample answer. The higher the knowledge correctness of the sample answer, the higher the target score. Therefore, when training the reward model based on the sample pair, the first sample answer with a higher target score and the second sample answer with a lower target score can be made to have a greater gap, so that the trained reward model can pay more attention to the knowledge correctness of the predicted answer when evaluating the predicted answer output by the large language model. Thus, the recognition accuracy of the knowledge correctness of the predicted answer can be improved.
[0117] The training device of the reward model provided by the present invention will be described below. The training device of the reward model described below can be mutually corresponding and referenced to the training method of the reward model described above.
[0118] Figure 5 is a schematic structural diagram of the training device of the reward model provided by the embodiments of the present invention. Refer to Figure 5 As shown, the training device 500 of the reward model includes:
[0119] An acquisition module 501, configured to acquire a plurality of sample pairs, each sample pair includes a first sample and a second sample, the first sample includes a sample question and a first sample answer, the second sample includes the sample question and a second sample answer, the target score of the first sample answer is higher than the target score of the second sample answer, and the target score is related to knowledge correctness, and the knowledge correctness is determined based on the target answer in the knowledge graph that matches the sample question;
[0120] An input module 502, configured to input the first sample and the second sample in the sample pair into an initial reward model for each sample pair, so as to obtain a first score of the first sample and a second score of the second sample output by the initial reward model;
[0121] An adjustment module 503, configured to adjust the model parameters of the initial reward model based on the first score and the second score to obtain a reward model, and the reward model is used to evaluate the answer output by the large language model.
[0122] In an exemplary embodiment, the adjustment module 503 is specifically configured to:
[0123] Determine a first knowledge correctness judgment result of the first sample answer relative to the target answer and a second knowledge correctness judgment result of the second sample answer relative to the target answer;
[0124] Determine a quantitative gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result;
[0125] Determine loss information based on the first score, the second score, and the quantitative gap;
[0126] Adjust model parameters of the initial reward model based on the loss information to obtain the reward model.
[0127] In one exemplary embodiment, the apparatus further includes a search module and a determination module, where:
[0128] The obtaining module 501 is further configured to obtain a plurality of sample answers corresponding to the sample question;
[0129] The search module is configured to search for a target answer corresponding to the sample question in the knowledge graph;
[0130] The determination module is configured to determine a target score for each of the sample answers based on the target answer.
[0131] In one exemplary embodiment, the determination module is specifically configured to:
[0132] Determine a target knowledge correctness judgment result of each of the sample answers relative to the target answer;
[0133] Determine a target score for each of the sample answers based on the target knowledge correctness judgment result and target evaluation dimensions of the target answer; the target evaluation dimensions include at least one of the following: fluency of the sample answer, user satisfaction with the sample answer, and safety of the sample answer.
[0134] In one exemplary embodiment, the determination module is specifically configured to:
[0135] For each of the sample answers, determine that the sample answer includes a first sub-answer and a second sub-answer other than the first sub-answer, where the first sub-answer includes an answer corresponding to the sample question;
[0136] Based on the target answer, respectively determine a knowledge correctness judgment result of the first sub-answer and a knowledge correctness judgment result of the second sub-answer;
[0137] Based on the knowledge correctness judgment result of the first sub-answer and the knowledge correctness judgment result of the second sub-answer, determine a target knowledge correctness judgment result corresponding to the sample answer.
[0138] In an exemplary embodiment, the determination module is specifically configured to:
[0139] When the knowledge correctness judgment result of the first sub - answer is correct and the knowledge correctness judgment result of the second sub - answer is correct, determine that the target knowledge correctness judgment result corresponding to the sample answer is correct;
[0140] When the knowledge correctness judgment result of the first sub - answer is correct and the knowledge correctness judgment result of the second sub - answer is wrong, determine that the target knowledge correctness judgment result corresponding to the sample answer is partially correct;
[0141] When the knowledge correctness judgment result of the first sub - answer is wrong, determine that the target knowledge correctness judgment result corresponding to the sample answer is wrong.
[0142] The device of this embodiment can be used to execute the method of any one of the method embodiments on the training side of the reward model. Its specific implementation process and technical effects are similar to those in the method embodiments on the training side of the reward model. For details, reference can be made to the detailed introduction in the method embodiments on the training side of the reward model, which will not be elaborated here.
[0143] Figure 6 is a schematic structural diagram of an answer evaluation device provided by an embodiment of the present invention. Refer to Figure 6 As shown, the answer evaluation device 600 includes:
[0144] An acquisition module 601, configured to acquire a predicted answer for a target question output by a large - language model;
[0145] An input module 602, configured to input the target question and the predicted answer into a reward model to obtain an evaluation result of the predicted answer output by the reward model. The reward model is trained based on the training method of the reward model described in any one of the above embodiments.
[0146] The device of this embodiment can be used to execute the method of any one of the method embodiments on the answer evaluation side. Its specific implementation process and technical effects are similar to those in the method embodiments on the answer evaluation side. For details, reference can be made to the detailed introduction in the method embodiments on the answer evaluation side, which will not be elaborated here.
[0147] Figure 7 Illustrates a schematic structural diagram of an electronic device, as Figure 7As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logical instructions in the memory 730 to execute the training method of the reward model. The method includes: obtaining a plurality of sample pairs, each of the sample pairs including a first sample and a second sample. The first sample includes a sample question and a first sample answer, and the second sample includes the sample question and a second sample answer. The target score of the first sample answer is higher than the target score of the second sample answer. The target score is related to the knowledge correctness, and the knowledge correctness is determined based on the target answer matching the sample question in the knowledge graph; for each of the sample pairs, inputting the first sample and the second sample in the sample pair into the initial reward model to obtain a first score of the first sample and a second score of the second sample output by the initial reward model; based on the first score and the second score, adjusting the model parameters of the initial reward model to obtain a reward model, and the reward model is used to evaluate the answers output by the large language model.
[0148] In addition, the processor 710 may also call the logical instructions in the memory 730 to execute the answer evaluation method. The method includes: obtaining a predicted answer for a target question output by the large language model; inputting the target question and the predicted answer into the reward model to obtain an evaluation result of the predicted answer output by the reward model, and the reward model is trained based on the training method of the reward model described in any of the above embodiments.
[0149] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0150] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method of the reward model provided by each of the above methods. The method includes: obtaining a plurality of sample pairs, each sample pair including a first sample and a second sample. The first sample includes a sample question and a first sample answer, and the second sample includes the sample question and a second sample answer. The target score of the first sample answer is higher than the target score of the second sample answer. The target score is related to the knowledge correctness, and the knowledge correctness is determined based on the target answer matching the sample question in the knowledge graph; for each sample pair, input the first sample and the second sample in the sample pair into the initial reward model to obtain a first score of the first sample and a second score of the second sample output by the initial reward model; based on the first score and the second score, adjust the model parameters of the initial reward model to obtain a reward model, and the reward model is used to evaluate the answers output by the large language model.
[0151] When the computer program is executed by a processor, the computer can also execute the answer evaluation method provided by each of the above methods. The method includes: obtaining a predicted answer to a target question output by the large language model; inputting the target question and the predicted answer into the reward model to obtain an evaluation result of the predicted answer output by the reward model, and the reward model is trained based on the training method of the reward model described in any one of the above embodiments.
[0152] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the training method of the reward model provided by each of the above methods. The method includes: obtaining a plurality of sample pairs, each sample pair including a first sample and a second sample. The first sample includes a sample question and a first sample answer, and the second sample includes the sample question and a second sample answer. The target score of the first sample answer is higher than the target score of the second sample answer. The target score is related to the knowledge correctness, and the knowledge correctness is determined based on the target answer matching the sample question in the knowledge graph; for each sample pair, input the first sample and the second sample in the sample pair into the initial reward model to obtain a first score of the first sample and a second score of the second sample output by the initial reward model; based on the first score and the second score, adjust the model parameters of the initial reward model to obtain a reward model, and the reward model is used to evaluate the answers output by the large language model.
[0153] In addition, when the computer program is executed by a processor, it implements an answer evaluation method provided by the above-described various methods. The method includes: obtaining a predicted answer to a target question output by a large language model; inputting the target question and the predicted answer into a reward model to obtain an evaluation result of the predicted answer output by the reward model, where the reward model is trained based on the training method of the reward model described in any of the previous embodiments.
[0154] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0155] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a reward model, characterized in that Including: Obtain multiple sample pairs, each of the sample pairs including a first sample and a second sample. The first sample includes a sample question and a first sample answer, and the second sample includes the sample question and a second sample answer. The target score of the first sample answer is higher than the target score of the second sample answer. The target score is related to the knowledge correctness, and the knowledge correctness is determined based on the target answer matched with the sample question in the knowledge graph. For each of the sample pairs, input the first sample and the second sample in the sample pair into the initial reward model to obtain a first score of the first sample and a second score of the second sample output by the initial reward model. Based on the first score and the second score, adjust the model parameters of the initial reward model to obtain a reward model, which is used to evaluate the answers output by the large language model. The adjusting the model parameters of the initial reward model based on the first score and the second score to obtain a reward model includes: Determine a first knowledge correctness judgment result of the first sample answer relative to the target answer and a second knowledge correctness judgment result of the second sample answer relative to the target answer. Determine the quantitative gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result. Determine loss information based on the first score, the second score, and the quantitative gap. Adjust the model parameters of the initial reward model based on the loss information to obtain the reward model.
2. The training method of the reward model according to claim 1, characterized in that The method further includes: Obtain multiple sample answers corresponding to the sample question. Search for the target answer corresponding to the sample question in the knowledge graph. Based on the target answer, determine the target scores of the sample answers.
3. The training method of the reward model according to claim 2, wherein The determining the target scores of the sample answers based on the target answer includes: Determine the target knowledge correctness judgment results of the sample answers relative to the target answer. Based on the target knowledge correctness judgment result and the target evaluation dimension of the target answer, determine the target scores of the sample answers. The target evaluation dimension includes at least one of the following: the fluency of the sample answer, the user satisfaction with the sample answer, and the safety of the sample answer.
4. The training method of the reward model according to claim 3, characterized in that The determining the target knowledge correctness judgment results of the sample answers relative to the target answer includes: For each of the sample answers, determine that the sample answer includes a first sub-answer and a second sub-answer other than the first sub-answer, and the first sub-answer includes the answer corresponding to the sample question. Based on the target answer, respectively determine the knowledge correctness judgment result of the first sub-answer and the knowledge correctness judgment result of the second sub-answer. Based on the knowledge correctness judgment result of the first sub-answer and the knowledge correctness judgment result of the second sub-answer, determine the target knowledge correctness judgment result corresponding to the sample answer.
5. The training method of the reward model according to claim 4, wherein Determining the target knowledge correctness judgment result corresponding to the sample answer based on the knowledge correctness judgment result of the first sub-answer and the knowledge correctness judgment result of the second sub-answer includes: When the knowledge correctness judgment result of the first sub-answer is correct and the knowledge correctness judgment result of the second sub-answer is correct, determining that the target knowledge correctness judgment result corresponding to the sample answer is correct; When the knowledge correctness judgment result of the first sub-answer is correct and the knowledge correctness judgment result of the second sub-answer is wrong, determining that the target knowledge correctness judgment result corresponding to the sample answer is partially correct; When the knowledge correctness judgment result of the first sub-answer is wrong, determining that the target knowledge correctness judgment result corresponding to the sample answer is wrong.
6. A method for evaluating answers, characterized in that, Including: Obtaining a predicted answer for the target question output by the large language model; Inputting the target question and the predicted answer into the reward model to obtain an evaluation result of the predicted answer output by the reward model, where the reward model is trained based on the training method of the reward model according to any one of claims 1-5.
7. A training device for a reward model, characterized in that Including: An acquisition module for acquiring a plurality of sample pairs, each sample pair including a first sample and a second sample. The first sample includes a sample question and a first sample answer, and the second sample includes the sample question and a second sample answer. The target score of the first sample answer is higher than the target score of the second sample answer, and the target score is related to knowledge correctness, and the knowledge correctness is determined based on a target answer matching the sample question in the knowledge graph; An input module for inputting the first sample and the second sample in each sample pair into the initial reward model for each sample pair to obtain a first score of the first sample and a second score of the second sample output by the initial reward model; An adjustment module for adjusting the model parameters of the initial reward model based on the first score and the second score to obtain a reward model for evaluating the answer output by the large language model; The adjustment module is specifically used for: Determining a first knowledge correctness judgment result of the first sample answer relative to the target answer and a second knowledge correctness judgment result of the second sample answer relative to the target answer; Determining a quantitative gap between the first knowledge correctness judgment result and the second knowledge correctness judgment result; Determining loss information based on the first score, the second score, and the quantitative gap; Adjusting the model parameters of the initial reward model based on the loss information to obtain the reward model.
8. An answer evaluation device, characterized in that, Including: An acquisition module for acquiring a predicted answer for the target question output by the large language model; An input module for inputting the target question and the predicted answer into the reward model to obtain an evaluation result of the predicted answer output by the reward model, where the reward model is trained based on the training method of the reward model according to any one of claims 1-5.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the training method of the reward model according to any one of claims 1-5 or implements the answer evaluation method according to claim 6.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the training method of the reward model according to any one of claims 1-5 or implements the answer evaluation method according to claim 6.
Citation Information
Patent Citations
Method and device for training generative large language model based on knowledge base feedback
CN117009490A
Large language model training method and device and text processing method and device
CN117149989A