Reward model acquisition method and device, equipment and storage medium

By generating a representation vector set and adjusting the reward model parameters, the similarity of the same answer is enhanced and the similarity of different answers is reduced. The problem of insufficient robustness of the existing reward model is solved, and higher robustness and practicality are achieved.

CN120258139APending Publication Date: 2025-07-04DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343399.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing reward model is poorly robust, resulting in insufficient practicality and general applicability, failure to effectively consider the characteristics shared by good and bad answers, and prone to overfitting.

Method used

By obtaining the first and second answers of the user sample, the initial double reward model is used to generate a representation vector set, and the model parameters are adjusted based on the preset similarity loss calculation formula to enhance the similarity of the same answer and reduce the similarity of different answers, the target reward model is obtained.

Benefits of technology

The representation learning ability and robustness of the target reward model is significantly enhanced, allowing it to adapt to a variety of Q&A scenarios, and improving practicality and universal applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258139A_ABST
    Figure CN120258139A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of artificial intelligence, and discloses a reward model obtaining method and device, equipment and a storage medium, and the method comprises the steps: obtaining a first sample of a user; inputting the first sample into an initial double-reward model to obtain a first representation vector set and a second representation vector set of the first sample; processing the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, the preset similarity calculation formula being used for increasing the similarity of data belonging to the same representation vector set and reducing the similarity of data belonging to different representation vector sets; and adjusting parameters of the initial double-reward model based on the target loss value to obtain a target reward model, the target reward model being used for determining an optimal answer conforming to the question asked by the user. By enhancing the similarity of the same answer and reducing the similarity of different answers, the representation learning ability of the target reward model can be significantly enhanced, and the robustness of the target reward model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and storage medium for obtaining a reward model. Background Art

[0002] In the field of artificial intelligence technology, the reward model is an important part of reinforcement learning based on human feedback, used to reflect human preferences and determine the forward direction of the subsequent proximal policy optimization algorithm. However, in the prior art, the reward model is mainly trained based on the Bradley-Terry model (BT-model, probability model). Specifically, in the process of training the reward model, only the good answers and bad answers of a single piece of data are sorted, and the common features of the good answers and bad answers are not considered. When using the good answers to optimize the reward model, only the high scores of the good answers are considered, and the gap between the scores of the good answers and bad answers is not considered. The robustness of the reward model is not considered, which easily leads to the problem of overfitting of the reward model. Therefore, the robustness of the reward model in the prior art is poor, resulting in poor practicality of the reward model and lack of universal applicability. Summary of the Invention

[0003] The purpose of the present invention is to provide at least one method, apparatus, device, and storage medium for obtaining a reward model, which can at least solve the technical problem that the robustness of the reward model in the prior art is poor, resulting in poor practicality of the reward model and lack of universal applicability, and can at least achieve the technical effect of improving the robustness of the reward model, making the reward model more practical and having universal applicability.

[0004] To solve the above technical problems, at least one embodiment of the present application provides a method for obtaining a reward model, including: obtaining a first sample of a user, where the first sample includes a first question, a first answer, and a second answer to the first question; inputting the first sample into an initial double-reward model to obtain a first representation vector set and a second representation vector set of the first sample, where the first and second are used to represent different representation vector sets for different answers, and each representation vector set stores multiple representation vectors for the same answer to the same question, and the representation vector refers to a feature vector used to understand and process the input sample; processing the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, where the preset similarity calculation formula is used to increase the similarity of data belonging to the same representation vector set and decrease the similarity of data belonging to different representation vector sets; adjusting the parameters of the initial double-reward model based on the target loss value to obtain a target reward model, where the target reward model is used to determine the optimal answer that meets the user's question.

[0005] At least one embodiment of the present application further provides an acquisition device for a reward model, including: an acquisition module, configured to acquire a first sample of a user, where the first sample includes a first question, a first answer and a second answer to the first question; an input module, configured to input the first sample into an initial double-reward model to obtain a first representation vector set and a second representation vector set of the first sample, where the first and the second are used to represent different representation vector sets for different answers, and each representation vector set stores multiple representation vectors for the same answer to the same question, and the representation vector refers to a feature vector for understanding and processing the input sample; a processing module, configured to process the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, where the preset similarity calculation formula is used to increase the similarity of data belonging to the same representation vector set and decrease the similarity of data belonging to different representation vector sets; an adjustment module, configured to adjust the parameters of the initial double-reward model based on the target loss value to obtain a target reward model, where the target reward model is used to determine the optimal answer that conforms to the question asked by the user.

[0006] At least one embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned acquisition method of the reward model.

[0007] At least one embodiment of the present application further provides a computer-readable storage medium, storing a computer program, where the computer program realizes the above-mentioned acquisition method of the reward model when executed by a processor.

[0008] The method for obtaining a reward model provided by an embodiment of the present application includes: obtaining a first sample of a user, where the first sample includes a first question, a first answer, and a second answer to the first question; inputting the first sample into an initial dual-reward model to obtain a first representation vector set and a second representation vector set of the first sample, where the first and second are used to represent different representation vector sets for different answers, and each representation vector set stores multiple representation vectors for the same answer to the same question, and the representation vector refers to a feature vector for understanding and processing the input sample; processing the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, where the preset similarity calculation formula is used to increase the similarity of data belonging to the same representation vector set and decrease the similarity of data belonging to different representation vector sets; adjusting the parameters of the initial dual-reward model based on the target loss value to obtain a target reward model, where the target reward model is used to determine the optimal answer that conforms to the user's question. By enhancing the similarity of the same answer and reducing the similarity of different answers, the representation learning ability of the target reward model can be significantly enhanced, and the robustness of the target reward model is improved. The target reward model can be adapted to various question-and-answer scenarios, making the target reward model have universal applicability and improving the practicality of the target reward model.

[0009] In some alternative embodiments, the initial dual-reward model includes two identical reward models, namely a first reward model and a second reward model. The step of inputting the first sample into the initial dual-reward model to obtain the first representation vector set and the second representation vector set of the first sample includes: using the first question and the first answer in the first sample as a first sub-sample, and using the first question and the second answer in the first sample as a second sub-sample; inputting the first sub-sample and the second sub-sample into the first reward model to obtain a first sub-representation vector and a second sub-representation vector; inputting the first sub-sample and the second sub-sample into the second reward model to obtain a third sub-representation vector and a fourth sub-representation vector; using the first sub-representation vector and the third sub-representation vector as the first representation vector set, and using the second sub-representation vector and the fourth sub-representation vector as the second representation vector set. By setting two identical reward models, the same sub-sample can be output twice, which not only ensures the stability of the output result but also enables the acquisition of the same features through comparative analysis of the two output results. This allows the reward model to learn the relevant features between the question and the answer, improves the learning ability of the reward model, and makes the reward model more accurate.

[0010] In some alternative embodiments, processing the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value includes: calculating a first similarity value between the first sub - representation vector and the third sub - representation vector; calculating a second similarity value between the second sub - representation vector and the fourth sub - representation vector; calculating a plurality of third similarity values between the first sub - representation vector and the third sub - representation vector and the fourth sub - representation vector respectively; calculating a plurality of fourth similarity values between the second sub - representation vector and the third sub - representation vector and the fourth sub - representation vector respectively; obtaining a first sum of the first similarity value and the second similarity value; obtaining a second sum of the plurality of third similarity values and the plurality of fourth similarity values; and taking the ratio of the first sum to the second sum as the target loss value. By obtaining the ratio of the first sum and the second sum, it is possible to increase the similarity of the same answers and decrease the similarity of different answers, so that the target reward model can better compare different answers and the same answers, and better learn the correlation between the answers and the questions. The robustness of the target reward model is improved.

[0011] In some alternative embodiments, the method further includes: obtaining a first label of the first answer, where the first label is used to represent the degree of compliance of the first answer with the standard of the first question; obtaining a second label of the second answer; inputting the first sub - sample and the second sub - sample into the first reward model to obtain a first score and a second score, where the first score is a score for predicting the degree of compliance between the answer and the question in the first sub - sample; and obtaining a scoring loss value based on the first label, the second label, the first score, and the second score. By using the first label and the second label, the deviation existing in the initial dual - reward model scoring can be determined, so as to adjust the parameters of the initial reward model to improve the accuracy of the target reward model.

[0012] In some alternative embodiments, obtaining the scoring loss value based on the first label, the second label, the first score, and the second score includes: if the first label is greater than the second label, determining the scoring loss value using the difference between the first score and the second score. By comparing the magnitudes of the first label and the second label, it is possible to determine which answer is a better answer, so as to determine how to calculate the loss value using the prediction results, and thus accurately determine how to adjust the model parameters. The model can better learn the relationship between the answers and the questions, and the accuracy of the target reward model is improved.

[0013] In some alternative embodiments, adjusting the parameters of the initial dual-reward model based on the target loss value to obtain a target reward model includes: obtaining a preset hyperparameter; fusing the scoring loss value and the target loss value by using the preset hyperparameter to obtain a model loss value; and adjusting the parameters of the initial dual-reward model by using the model loss value to obtain a target reward model. By fusing the scoring loss value and the target loss value, more accurate model parameters can be obtained, improving the accuracy of the trained target reward model.

[0014] In some alternative embodiments, fusing the scoring loss value and the target loss value by using the preset hyperparameter to obtain a model loss value includes: taking the sum of the product of the preset hyperparameter and the target loss value and the scoring loss value as the model loss value. By using the preset hyperparameter, the target loss value and the scoring loss value can be better fused, avoiding the situation where the proportion of the target loss value is too large or too small, resulting in inaccurate target reward models. The accuracy of the target reward model is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] One or more embodiments are illustrated by way of example in the accompanying drawings, which do not constitute a limitation on the embodiments.

[0016] Figure 1 is a schematic flowchart of a method for obtaining a reward model provided by an embodiment of the present application;

[0017] Figure 2 is a schematic structural diagram of a reward model provided by an embodiment of the present application;

[0018] Figure 3 is a schematic structural diagram of a method for obtaining a representation vector between answers provided by another embodiment of the present application;

[0019] Figure 4 is a schematic structural diagram of an apparatus for obtaining a reward model provided by another embodiment of the present application;

[0020] Figure 5 is a schematic structural diagram of an electronic device provided by another embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will elaborate on each embodiment of this application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of this application, many technical details are presented to help readers better understand this application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined and cross-referenced with each other on the premise of no contradiction.

[0022] It should be noted that the acquisition or use of the data in the embodiments of this application requires user consent. Relevant data can only be obtained after the user authorizes and permits it, and the acquisition or use of the data complies with the provisions of relevant laws and regulations.

[0023] To facilitate the understanding of some terms in this solution, relevant terms are explained here:

[0024] Representation learning: It is used to indicate that the reward model automatically learns an effective representation of the data so that the reward model can better complete various tasks.

[0025] Alignment effect: It refers to the degree of consistency between the results given by the reward model and the results expected by humans.

[0026] Bradley-Terry model (BT-model, probability model): It is a classical probability model mainly used to predict and model the preference relationship between two competitors.

[0027] In the field of artificial intelligence technology, the reward model is an important part of reinforcement learning based on human feedback, used to reflect human preferences and determine the forward direction of the subsequent proximal policy optimization algorithm. However, in the prior art, the reward model is mainly trained based on the Bradley-Terry model (BT-model, probability model). Specifically, in the process of training the reward model, only the good answers and bad answers of a single piece of data are sorted, without considering the common features of good answers and bad answers. When using good answers to optimize the reward model, only the fact that good answers have high scores is considered, without considering the gap between the scores of good answers and bad answers. The robustness of the reward model is not considered, which easily leads to the problem of overfitting of the reward model. Therefore, the robustness of the reward model in the prior art is poor, resulting in poor practicality of the reward model and lack of general applicability. Moreover, the importance of representation learning is ignored, resulting in a poor alignment effect of the reward model.

[0028] To solve the technical problem in the prior art that the robustness of the reward model is poor, resulting in poor practicality of the reward model and lack of universal applicability, the present invention proposes a method for obtaining a reward model. The following details the implementation of the method for obtaining the reward model in this embodiment. The following content is only for facilitating understanding of the implementation details and is not necessary for implementing the solution.

[0029] Embodiment 1:

[0030] The method for obtaining the reward model in this embodiment can be applied to an electronic device with communication, computing, and data storage capabilities. Its specific process can be as Figure 1 shown, including:

[0031] Step 101, obtain a first sample of the user, where the first sample includes a first question, a first answer, and a second answer to the first question.

[0032] Specifically, the first answer and the second answer to the first question can be generated based on a large model. For example, input the first question into the large model, and the large model can give two answers based on the first question, namely the first answer and the second answer.

[0033] Specifically, the large model can receive a set of user requests, and the set of requests includes at least one first question. The large model can, through the inference method of a preset sampling algorithm, give multiple different answers to each question in the set of requests. Specifically, the large model can adopt the Top-P algorithm and perform sampling by setting different temperature coefficients and probability values during sampling to give multiple answers to the questions in the set of requests. Among them, for the Top-P algorithm, reference can be made to the prior art, and details will not be elaborated here.

[0034] Exemplary one, if there is only one first question in the set of user requests, denote this first question as x1; the large model gives two answers to x1, denote the first answer as y1 and the second answer as y2. Then the first sample is denoted as (x1, y1, y2).

[0035] Exemplary two, if the set of user requests includes multiple questions, here take two questions as an example for illustration. Denote the second question as x2 and the third question as x3; the large model gives two answers to x2, denote the third answer as y3 and the fourth answer as y4. The large model gives two answers to x3, denote the fifth answer as y5 and the sixth answer as y6. Then the second sample is denoted as (x2, y3, y4), and the third sample is denoted as (x3, y5, y6).

[0036] Among them, there can be multiple first samples, and there can also be multiple first questions. For the sake of easy understanding, the second sample and the third sample in this second exemplary example can be regarded as multiple first samples, and the second question and the third question in this second exemplary example can be regarded as multiple first questions.

[0037] In some examples, the method further includes: the first answer and the second answer can be labeled based on manual annotation, and through the annotation, it can be learned which answer in the first answer and the second answer is more in line with the user's question.

[0038] Step 102, input the first sample into the initial dual-reward model to obtain a first representation vector set and a second representation vector set of the first sample. Among them, the first and the second are used to represent different representation vector sets for different answers. Each representation vector set stores multiple representation vectors of the same answer to the same question. The representation vector refers to the feature vector used to understand and process the input sample.

[0039] In some examples, in the aforementioned step 102, the initial dual-reward model includes two identical reward models, namely the first reward model and the second reward model. The step of inputting the first sample into the initial dual-reward model to obtain the first representation vector set and the second representation vector set of the first sample includes: taking the first question and the first answer in the first sample as the first sub-sample, and taking the first question and the second answer in the first sample as the second sub-sample; inputting the first sub-sample and the second sub-sample into the first reward model to obtain a first sub-representation vector and a second sub-representation vector; inputting the first sub-sample and the second sub-sample into the second reward model to obtain a third sub-representation vector and a fourth sub-representation vector; taking the first sub-representation vector and the third sub-representation vector as the first representation vector set, and taking the second sub-representation vector and the fourth sub-representation vector as the second representation vector set.

[0040] In some examples, reference can be made to Figure 2 , the first reward model uses llama3 (llama3 refers to a model in the LLM model) as the backbone network of the reward model. Specifically, the LM_head layer of llama3 is replaced by the value head layer. The Embedding layer, GQA layer, and MLP layer of llama3 are retained. The value head layer is a linear layer, and the input is the feature of the second-to-top layer of llama3, and the output is a one-dimensional reward score. Among them, the reward score is used to represent the score of the degree of conformity between the input answer and the question. The second reward model has the same structure as the first reward model.

[0041] Continuing with the foregoing Example 1, the first sub-sample is (x1, y1), and the second sub-sample is (x1, y2); the first sub-sample and the second sub-sample are input into the first reward model to obtain a first sub-representation vector R1(x1, y1) and a second sub-representation vector R1(x1, y2); the first sub-sample and the second sub-sample are input into the second reward model to obtain a third sub-representation vector R2(x1, y1) and a fourth sub-representation vector R2(x1, y2).

[0042] Continuing with the foregoing Example 2, the second sample (x2, y 3, y4) is divided into sub-sample 1 (x2, y3) and sub-sample 2 (x2, y4), and the third sample (x3, y5, y6) is divided into sub-sample 3 (x3, y5) and sub-sample 4 (x3, y6). Sub-sample 1, sub-sample 2, sub-sample 3, and sub-sample 4 are input into the first reward model to obtain sub-representation vector 1 R1(x2, y3), sub-representation vector 2 R1(x2, y4), sub-representation vector 3 R1(x3, y5), and sub-representation vector 4 R1(x3, y6). Sub-sample 1, sub-sample 2, sub-sample 3, and sub-sample 4 are input into the second reward model to obtain sub-representation vector 5 R2(x2, y3), sub-representation vector 6 R2(x2, y4), sub-representation vector 7 R2(x3, y5), and sub-representation vector 8 R2(x3, y6).

[0043] In some examples, the method further includes: obtaining a first label of the first answer, where the first label is used to represent the degree of compliance of the first answer with the standard of the first question; obtaining a second label of the second answer; inputting the first sub-sample and the second sub-sample into the first reward model to obtain a first score and a second score, where the first score is the score of the degree of compliance of the answer in the first sub-sample with the question predicted; obtaining a scoring loss value based on the first label, the second label, the first score, and the second score.

[0044] In some examples, obtaining the scoring loss value based on the first label, the second label, the first score, and the second score includes: if the first label is greater than the second label, determining the scoring loss value using the difference between the first score and the second score.

[0045] Continuing with the foregoing Example 1, the first label of the first answer is 1, and the second label of the second answer is 0. The first score is S(x1, y1), and the second score is S(x1, y2). Since the first label is greater than the second label, the scoring loss value can be seen in the following formula:

[0046]

[0047] Following the foregoing Example 2, the label of the third answer is 2, the label of the fourth answer is 3, the label of the fifth answer is 4, and the label of the sixth answer is 2. The score of sub-sample 1 is S(x2, y3) and the score of sub-sample 2 is S(x2, y4), the score of sub-sample 3 is S(x3, y5) and the score of sub-sample 4 is S(x3, y6).

[0048]

[0049] Step 103: Process the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, where the preset similarity calculation formula is used to increase the similarity of data belonging to the same representation vector set and reduce the similarity of data belonging to different representation vector sets.

[0050] In some examples, processing the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value can be seen in the following formula:

[0051]

[0052] where cl_loss refers to the target loss value, N is the number of representation vector sets, (R1(x i ,y i ),R2(x i ,y i )) represents two sub-representation vectors belonging to the same representation vector set, R1(x i ,y i ) is obtained through the first reward model, R2(x i ,y i ) is obtained through the second reward model, (R1(x i ,y j ),R2(x x ,y y )) represents two sub-representation vectors belonging to the same or different representation vector sets, where R1(x i ,y i ) is obtained through the first reward model, and R2(x x ,y y ) is obtained through the second reward model.

[0053] In some examples, in the foregoing step 103, processing the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value includes: calculating a first similarity value between the first sub-representation vector and the third sub-representation vector; calculating a second similarity value between the second sub-representation vector and the fourth sub-representation vector; calculating a plurality of third similarity values between the first sub-representation vector and the third sub-representation vector and the fourth sub-representation vector respectively; calculating a plurality of fourth similarity values between the second sub-representation vector and the third sub-representation vector and the fourth sub-representation vector respectively; obtaining a first sum of the first similarity value and the second similarity value; obtaining a second sum of the plurality of third similarity values and the plurality of fourth similarity values; and taking the ratio of the first sum to the second sum as the target loss value.

[0054] Continuing with the foregoing Exemplary One, processing the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, the following formula can be referred to:

[0055]

[0056] Continuing with the foregoing Exemplary Two, processing the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, the following formula can be referred to:

[0057]

[0058]

[0059] Step 104, adjusting the parameters of the initial dual-reward model based on the target loss value to obtain a target reward model, where the target reward model is used to determine the optimal answer that meets the user's question.

[0060] In some examples, in the foregoing step 104, adjusting the parameters of the initial dual-reward model based on the target loss value to obtain a target reward model includes: obtaining a preset hyperparameter; using the preset hyperparameter to fuse the scoring loss value and the target loss value to obtain a model loss value; and using the model loss value to adjust the parameters of the initial dual-reward model to obtain a target reward model.

[0061] In some examples, using the preset hyperparameter to fuse the scoring loss value and the target loss value to obtain a model loss value includes: taking the sum of the product of the preset hyperparameter and the target loss value and the scoring loss value as the model loss value.

[0062] In some examples, the sum of the product of the preset hyperparameter and the target loss value and the scoring loss value is used as the model loss value. See the following formula:

[0063] Model loss value = Scoring loss value + Preset hyperparameter × cl_loss

[0064] For ease of understanding this solution, see Figure 3 , the first sub-sample and the second sub-sample are both processed by the first reward model and the second reward model to obtain a first sub-representation vector, a second sub-representation vector, a third sub-representation vector, and a fourth sub-representation vector. Through contrastive learning, the similarity between the first sub-representation vector and the third sub-representation vector is increased, and the similarity between the first sub-representation vector and the third sub-representation vector and the fourth sub-representation vector is decreased. Similarly, the similarity between the second sub-representation vector and the fourth sub-representation vector is increased, and the similarity between the second sub-representation vector and the first sub-representation vector and the third sub-representation vector is decreased, thereby improving the representation learning ability of the target reward model. To improve the robustness of the target reward model, specifically, the similarity between the first sub-representation vector and the third sub-representation vector can be increased by increasing the cosine similarity between the first sub-representation vector and the third sub-representation vector.

[0065] In summary, this application proposes to obtain a first sample of the user, where the first sample includes a first question, a first answer and a second answer to the first question; input the first sample into the initial dual-reward model to obtain a first representation vector set and a second representation vector set of the first sample, where the first and second are used to represent different representation vector sets for different answers, and each representation vector set stores multiple representation vectors for the same answer to the same question, and the representation vector refers to the feature vector used to understand and process the input sample; based on the preset similarity loss calculation formula, the first representation vector set and the second representation vector set are processed to obtain a target loss value, where the preset similarity calculation formula is used to increase the similarity of the data belonging to the same representation vector set and decrease the similarity of the data belonging to different representation vector sets; based on the target loss value, the parameters of the initial dual-reward model are adjusted to obtain a target reward model, where the target reward model is used to determine the solution for obtaining the optimal answer that conforms to the user's question. By enhancing the similarity of the same answer and reducing the similarity of different answers, the representation learning ability of the target reward model can be significantly enhanced, and the robustness of the target reward model is improved. The target reward model can be adapted to a variety of question-and-answer scenarios, making the target reward model have general applicability and improving the practicality of the target reward model.

[0066] Specifically, the solution proposed in this program trains the target reward model through two identical reward models. By comparing the performance vectors between the same answers and different answers, the representation learning ability of the target reward model is significantly enhanced, and the alignment effect of the target reward model is improved. The performance ceiling of the target reward model is significantly increased. Based on the traditional BT-model, by adding contrastive learning technology, the information inside good answers and bad answers is explicitly modeled. The robustness of the current reward model is significantly enhanced, and it has strong practicality and universality. Among them, a good answer refers to an answer that meets the user's expectations, and a bad answer refers to an answer that does not meet the user's expectations.

[0067] Embodiment 2:

[0068] Another embodiment of the present application relates to an acquisition device for a reward model. The implementation details of the acquisition device for the reward model in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this program. The schematic diagram of the acquisition device for the reward model in this embodiment can be as Figure 4 shown, including an acquisition module 401, an input module 402, a processing module 403, and an adjustment module 404.

[0069] The acquisition module 401 is used to acquire the first sample of the user, where the first sample includes a first question, a first answer, and a second answer to the first question;

[0070] The input module 402 is used to input the first sample into the initial dual reward model to obtain a first representation vector set and a second representation vector set of the first sample. Here, the first and second are used to represent different representation vector sets for different answers. Each representation vector set stores multiple representation vectors for the same answer to the same question. The representation vector refers to the feature vector used to understand and process the input sample;

[0071] The processing module 403 is used to process the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, where the preset similarity calculation formula is used to increase the similarity of data belonging to the same representation vector set and decrease the similarity of data belonging to different representation vector sets;

[0072] The adjustment module 404 is used to adjust the parameters of the initial dual reward model based on the target loss value to obtain a target reward model, where the target reward model is used to determine the optimal answer that meets the user's question.

[0073] In some examples, the initial dual reward model in the device includes two identical reward models, namely the first reward model and the second reward model. When the device is used to input the first sample into the initial dual reward model to obtain the first representation vector set and the second representation vector set of the first sample, it is specifically used for: taking the first question and the first answer in the first sample as the first subsample, and taking the first question and the second answer in the first sample as the second subsample; inputting the first subsample and the second subsample into the first reward model to obtain the first sub-representation vector and the second sub-representation vector; inputting the first subsample and the second subsample into the second reward model to obtain the third sub-representation vector and the fourth sub-representation vector; taking the first sub-representation vector and the third sub-representation vector as the first representation vector set, and taking the second sub-representation vector and the fourth sub-representation vector as the second representation vector set.

[0074] In some examples, when the device is used to process the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, it is specifically used for: calculating a first similarity value between the first sub-representation vector and the third sub-representation vector; calculating a second similarity value between the second sub-representation vector and the fourth sub-representation vector; calculating a plurality of third similarity values between the first sub-representation vector and the third sub-representation vector and the fourth sub-representation vector respectively; calculating a plurality of fourth similarity values between the second sub-representation vector and the third sub-representation vector and the fourth sub-representation vector respectively; obtaining a first sum of the first similarity value and the second similarity value; obtaining a second sum of the plurality of third similarity values and the plurality of fourth similarity values; taking the ratio of the first sum to the second sum as the target loss value.

[0075] In some examples, the device is further used for: obtaining a first label of the first answer, where the first label is used to represent the degree of compliance of the first answer with the first question; obtaining a second label of the second answer; inputting the first subsample and the second subsample into the first reward model to obtain a first score and a second score, where the first score is the score of the degree of compliance between the predicted answer and the question in the first subsample; obtaining a scoring loss value based on the first label, the second label, the first score, and the second score.

[0076] In some examples, when the device is used to obtain the scoring loss value based on the first label, the second label, the first score, and the second score, it is specifically used for: if the first label is greater than the second label, determining the scoring loss value using the difference between the first score and the second score.

[0077] In some examples, when the device is used to adjust the parameters of the initial dual-reward model based on the target loss value to obtain the target reward model, it is specifically used for: obtaining a preset hyperparameter; fusing the scoring loss value and the target loss value by using the preset hyperparameter to obtain a model loss value; and adjusting the parameters of the initial dual-reward model by using the model loss value to obtain the target reward model.

[0078] In some examples, when the device is used to fuse the scoring loss value and the target loss value by using the preset hyperparameter to obtain a model loss value, it is specifically used for: taking the sum of the product of the preset hyperparameter and the target loss value and the scoring loss value as the model loss value.

[0079] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or implemented as a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0080] Embodiment Three:

[0081] Another embodiment of the present application relates to an electronic device, as Figure 5 shown, including: at least one processor 901; and a memory 902 communicatively connected to the at least one processor 901; wherein, the memory 902 stores instructions executable by the at least one processor 901, and the instructions are executed by the at least one processor 901 so that the at least one processor 901 can execute the method for obtaining the reward model in the above embodiments.

[0082] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0083] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor when executing operations.

[0084] Embodiment 4:

[0085] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0086] That is, those skilled in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., which can store program codes.

[0087] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.

Claims

1. A method for obtaining a reward model, characterized in that, Including: Obtain a first sample of the user, where the first sample includes a first question, a first answer, and a second answer to the first question; Input the first sample into an initial dual-reward model to obtain a first set of representation vectors and a second set of representation vectors for the first sample, where the first and second are used to represent different sets of representation vectors for different answers, and each set of representation vectors stores multiple representation vectors for the same answer to the same question. The representation vector refers to a feature vector used to understand and process the input sample; Process the first set of representation vectors and the second set of representation vectors based on a preset similarity loss calculation formula to obtain a target loss value, where the preset similarity calculation formula is used to increase the similarity of data belonging to the same set of representation vectors and decrease the similarity of data belonging to different sets of representation vectors; Adjust the parameters of the initial dual-reward model based on the target loss value to obtain a target reward model, where the target reward model is used to determine the optimal answer that conforms to the user's question.

2. The method for obtaining the reward model according to claim 1, characterized in that The initial dual-reward model includes two identical reward models, namely a first reward model and a second reward model. The step of inputting the first sample into the initial dual-reward model to obtain the first set of representation vectors and the second set of representation vectors for the first sample includes: Use the first question and the first answer in the first sample as a first sub-sample, and use the first question and the second answer in the first sample as a second sub-sample; Input the first sub-sample and the second sub-sample into the first reward model to obtain a first sub-representation vector and a second sub-representation vector; Input the first sub-sample and the second sub-sample into the second reward model to obtain a third sub-representation vector and a fourth sub-representation vector; Use the first sub-representation vector and the third sub-representation vector as the first set of representation vectors, and use the second sub-representation vector and the fourth sub-representation vector as the second set of representation vectors.

3. The method for obtaining the reward model according to claim 2, wherein The step of processing the first set of representation vectors and the second set of representation vectors based on a preset similarity loss calculation formula to obtain a target loss value includes: Calculate a first similarity value between the first sub-representation vector and the third sub-representation vector; Calculate a second similarity value between the second sub-representation vector and the fourth sub-representation vector; Calculate multiple third similarity values between the first sub-representation vector and the third sub-representation vector and the fourth sub-representation vector respectively; Calculate multiple fourth similarity values between the second sub-representation vector and the third sub-representation vector and the fourth sub-representation vector respectively; Obtain the first sum of the first similarity value and the second similarity value; Obtain the second sum of the multiple third similarity values and the multiple fourth similarity values; Use the ratio of the first sum to the second sum as the target loss value.

4. The method for obtaining the reward model according to claim 2, characterized in that, The method further includes: Obtain a first label for the first answer, where the first label is used to represent the degree of compliance of the first answer with the standard of the first question; Obtain a second label for the second answer; Input the first sub-sample and the second sub-sample into the first reward model to obtain a first score and a second score, where the first score is the score of the degree of conformity between the answer and the question in the predicted first sub-sample. Obtain a scoring loss value based on the first label, the second label, the first score, and the second score.

5. The method for obtaining the reward model according to claim 4, wherein, The obtaining the scoring loss value based on the first label, the second label, the first score, and the second score includes: If the first label is greater than the second label, determine the scoring loss value using the difference between the first score and the second score.

6. The method for obtaining the reward model according to claim 4, wherein The adjusting the parameters of the initial dual reward model based on the target loss value to obtain a target reward model includes: Obtain a preset hyperparameter. Fuse the scoring loss value and the target loss value using the preset hyperparameter to obtain a model loss value. Adjust the parameters of the initial dual reward model using the model loss value to obtain a target reward model.

7. The method for obtaining the reward model according to claim 6, wherein, The fusing the scoring loss value and the target loss value using the preset hyperparameter to obtain a model loss value includes: Use the sum of the product of the preset hyperparameter and the target loss value and the scoring loss value as the model loss value.

8. An acquisition device for a reward model, characterized in that Includes: An obtaining module, configured to obtain a first sample of a user, where the first sample includes a first question, a first answer, and a second answer to the first question. An input module, configured to input the first sample into an initial dual reward model to obtain a first representation vector set and a second representation vector set of the first sample, where the first and second are used to represent different representation vector sets for different answers, and each representation vector set stores multiple representation vectors for the same answer to the same question, and the representation vector refers to a feature vector used to understand and process the input sample. A processing module, configured to process the first representation vector set and the second representation vector set based on a preset similarity loss calculation formula to obtain a target loss value, where the preset similarity calculation formula is used to increase the similarity of data belonging to the same representation vector set and decrease the similarity of data belonging to different representation vector sets. An adjusting module, configured to adjust the parameters of the initial dual reward model based on the target loss value to obtain a target reward model, where the target reward model is used to determine the optimal answer that conforms to the user's question.

9. An electronic device, characterized in that, Includes: At least one processor; And, A memory communicatively connected to the at least one processor; where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for obtaining a reward model according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the method for obtaining a reward model according to any one of claims 1 to 7.