Reward model optimization method and device, computer equipment and storage medium

By obtaining diverse training data and optimizing the parameters of the reward model, the problem of low evaluation performance of the reward model is solved, and the evaluation effect is improved and the training process is accelerated.

CN120197723APending Publication Date: 2025-06-24SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311776016.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, the training data of the reward model is insufficient, resulting in low evaluation performance, which in turn affects the evaluation effect.

Method used

By obtaining multiple sets of training data, the data is improved, and the original reward model is used to score and label the data, determine the optimization function value, optimize the model parameters until the convergence conditions are met, and the target reward model is obtained.

Benefits of technology

The evaluation performance of the reward model is improved, making its evaluation effect close to manual annotation, saving human resources, and accelerating the alignment training process of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197723A_ABST
    Figure CN120197723A_ABST
Patent Text Reader

Abstract

The invention discloses a reward model optimization method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining training data which comprises cue words and training replies; receiving an annotation result corresponding to the training data, and determining a first reward score of each piece of training data based on the annotation result; scoring and marking the training data by adopting the original reward model, and determining a second reward score of each training data; determining an optimization function value corresponding to the original reward model based on the first reward score and the second reward score of the plurality of training data corresponding to the same prompt word; when the optimization function value does not meet the convergence condition, optimizing model parameters of the original reward model; and when the optimization function value meets a convergence condition, taking the original reward model as a target reward model. According to the method, the evaluation effect of the reward model can be close to the evaluation effect of manual annotation, and the purpose of improving the evaluation performance of the original reward model is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a method, device, computer device, and storage medium for optimizing a reward model. Background Art

[0002] Currently, the reward model plays an important role in various algorithm models. Especially in the alignment training of large language models, the reward model usually needs to replace manual annotation to score and evaluate the text content generated by the large language model, so as to reduce the iteration cycle of the alignment training of the large language model, accelerate the alignment training process of the large language model, and save human resources. However, since the scoring effect of the reward model needs to conform to the model characteristics of the large language model, during the training stage of the large language model, it is necessary to optimize the reward model to improve the evaluation performance of the reward model, so that the reward model has a better evaluation effect on the text content output by the large language model.

[0003] In the prior art, during the optimization process of the reward model, since the training data of the reward model comes from the output of the large language model corresponding to the reward model, for the same prompt word, the differences between multiple training responses in the training data are not significant, and the diversity of the training data is insufficient, resulting in a low evaluation performance of the reward model, and thus the evaluation effect is not ideal. Moreover, in the prior art, during the optimization process of the reward model, usually the overall evaluation score of the training responses in the training data is obtained, or the training responses in the training data are sorted according to the preference order. The way of obtaining the overall evaluation score and the sorting method cannot comprehensively reflect the effect of the training responses, which will also lead to a low evaluation performance of the reward model and an unsatisfactory evaluation effect.

[0004] In summary, how to improve the evaluation performance of the reward model to improve the evaluation effect of the reward model in practical applications is a technical problem that needs to be solved urgently. Summary of the Invention

[0005] Embodiments of the present invention provide a method, device, computer device, and storage medium for optimizing a reward model to solve the technical problem of how to improve the evaluation performance of the reward model.

[0006] A method for optimizing a reward model includes:

[0007] Obtain training data, where the training data includes prompt words and training responses;

[0008] Receive the annotation result corresponding to the training data, and based on the annotation result, determine the first reward score of each piece of training data;

[0009] Use the original reward model to score and annotate the training data, and determine the second reward score of each piece of training data;

[0010] Determine an optimization function value corresponding to the original reward model based on a first reward score and a second reward score of a plurality of the training data corresponding to the same prompt;

[0011] When the optimization function value does not meet the convergence condition, optimize the model parameters of the original reward model;

[0012] When the optimization function value meets the convergence condition, use the original reward model as the target reward model.

[0013] Preferably, the obtaining of the training data includes:

[0014] Perform sampling processing on the source data to obtain at least one prompt;

[0015] Perform multi-dimensional answer processing on each of the prompts to obtain at least one original reply corresponding to each of the prompts;

[0016] Use a target intent injection model to process the original reply corresponding to each of the prompts to obtain at least one training reply corresponding to each of the prompts.

[0017] Preferably, the target intent injection model includes a refusal to answer intent injection model;

[0018] Preferably, the using a target intent injection model to process the original reply corresponding to each of the prompts to obtain at least one training reply corresponding to each of the prompts includes:

[0019] Perform sensitive word recognition on each of the prompts to determine whether the prompt is a sensitive word;

[0020] If the prompt is a sensitive word, use the refusal to answer intent injection model to process the original reply corresponding to the sensitive word to obtain a first training reply corresponding to the sensitive word;

[0021] If the prompt is not a sensitive word, perform optimization processing on the original reply corresponding to the prompt to obtain at least one second training reply corresponding to the prompt.

[0022] Preferably, the using a refusal to answer intent injection model to process the original reply corresponding to the prompt to obtain a first training reply corresponding to the prompt includes:

[0023] If the prompt is a sensitive word, determine all the original replies corresponding to the sensitive word as invalid replies;

[0024] Call the preset reply template corresponding to the sensitive word in the refusal to answer intention injection model, and re-edit the invalid reply to obtain the first training reply corresponding to the sensitive word.

[0025] Preferably, the target intention injection model includes an output format intention injection model and an instruction following intention injection model;

[0026] Preferably, the optimizing the original reply corresponding to the prompt word to obtain at least one second training reply corresponding to the prompt word includes:

[0027] Using the output format intention injection model to process the output format of the original reply corresponding to the prompt word to obtain at least one second training reply corresponding to the prompt word;

[0028] Using the instruction following intention injection model to fuse at least two original replies corresponding to the prompt word to obtain at least one second training reply corresponding to the prompt word.

[0029] Preferably, the receiving the annotation result corresponding to the training data and determining the first reward score of each training data based on the annotation result includes:

[0030] Receiving the initial reward score corresponding to each sentence in the training reply of the training data;

[0031] Statistically processing the initial reward scores corresponding to each sentence in the training reply to obtain the first reward score corresponding to the training reply.

[0032] Preferably, the determining the optimization function value corresponding to the original reward model based on the first reward scores and second reward scores of multiple training data corresponding to the same prompt word includes:

[0033] Determining the first difference corresponding to the same prompt word based on the first reward scores and second reward scores of multiple training data corresponding to the same prompt word;

[0034] Determining the optimization function value corresponding to the original reward model based on the first differences corresponding to multiple prompt words.

[0035] A reward model optimization device includes:

[0036] A training data acquisition module for acquiring training data, where the training data includes a prompt word and a training reply;

[0037] A first reward score determination module for receiving the annotation result corresponding to the training data and determining the first reward score of each training data based on the annotation result;

[0038] The second reward score determination module is used to score and label the training data by using the original reward model, and determine the second reward score of each piece of training data;

[0039] The optimization function value determination module determines the optimization function value corresponding to the original reward model based on the first reward scores and the second reward scores of multiple pieces of training data corresponding to the same prompt;

[0040] The model parameter optimization module is used to optimize the model parameters of the original reward model when the optimization function value does not meet the convergence condition;

[0041] The target reward model determination module is used to use the original reward model as the target reward model when the optimization function value meets the convergence condition.

[0042] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above reward model optimization method is implemented.

[0043] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above reward model optimization method is implemented.

[0044] For the above reward model optimization method, device, computer device and storage medium, obtaining multiple groups of training data can improve the diversity of training data, facilitate subsequent precise optimization of the original reward model through the training data, and improve the evaluation performance of the reward model. Receiving the annotation results corresponding to the training data and determining the first reward score of each group of training data based on the annotation results facilitate subsequent optimization and update of the reward model according to the first reward score, so that the evaluation effect of the reward model is close to the evaluation effect of manual annotation and the evaluation performance of the reward model is improved. Determining the optimization function value corresponding to the original reward model based on the first reward scores and the second reward scores of multiple pieces of training data corresponding to the same prompt facilitates subsequent update of the original reward model according to the optimization function value and improves the evaluation effect of the original reward model. When the optimization function value does not meet the convergence condition, optimizing the model parameters of the original reward model, and when the optimization function value meets the convergence condition, determining the original reward model corresponding to the optimization function value that meets the convergence condition as the target reward model can make the evaluation effect of the reward model relatively close to the evaluation effect of manual annotation and achieve the purpose of improving the evaluation performance of the original reward model. Description of the Drawings

[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 is a flowchart of a method for optimizing a reward model according to an embodiment of the present invention;

[0047] Figure 2 is another flowchart of a method for optimizing a reward model according to an embodiment of the present invention;

[0048] Figure 3 is another flowchart of a method for optimizing a reward model according to an embodiment of the present invention;

[0049] Figure 4 is another flowchart of a method for optimizing a reward model according to an embodiment of the present invention;

[0050] Figure 5 is another flowchart of a method for optimizing a reward model according to an embodiment of the present invention;

[0051] Figure 6 is another flowchart of a method for optimizing a reward model according to an embodiment of the present invention;

[0052] Figure 7 is a schematic diagram of a reward model optimization device according to an embodiment of the present invention;

[0053] Figure 8 is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed implementation manners

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0055] The embodiments of the present invention provide a method for optimizing a reward model for the purpose of improving the evaluation performance of the reward model. The method for optimizing the reward model can be applied to a computer device.

[0056] In one embodiment, as Figure 1 shown, a method for optimizing a reward model is provided. Taking the application of this method in a Figure 8 computer device as an example for illustration, the method includes the following steps:

[0057] S101: Obtain training data, where the training data includes prompt words and training responses;

[0058] S102: Receive the annotation results corresponding to the training data, and based on the annotation results, determine the first reward score for each piece of training data;

[0059] S103: Use the original reward model to score and annotate the training data, and determine the second reward score for each piece of training data;

[0060] S104: Based on the first reward scores and the second reward scores of multiple pieces of training data corresponding to the same prompt word, determine the optimization function value corresponding to the original reward model;

[0061] S105: When the optimization function value does not meet the convergence condition, optimize the model parameters of the original reward model;

[0062] S106: When the optimization function value meets the convergence condition, use the original reward model as the target reward model.

[0063] Among them, the training data refers to the data used to optimize and train the original reward model. The prompt word is used as the basis for obtaining the training response. The training response refers to the response data to the prompt word.

[0064] As an example, in step S101, the computer device obtains at least one prompt word, inputs each prompt word into different models for processing, obtains at least one training response corresponding to each prompt word, and uses one training response corresponding to each prompt word as a set of training data. For example, for the same prompt word Q1, if its corresponding training responses are A 11 , A 12 , …, A 1a , then there are a sets of training data: {Q1, A 11}, {Q2, A 12}, …, {Q1, A 1a}. Similarly, for the same prompt word Q2, if its corresponding training responses are A 21 , A 22 , …, A 2b , then there are b sets of training data: {Q2, A 21}, {Q2, A 22}, …, {Q2, A 2b}. The prompt words can come from specific application scenarios. For example, the prompt words can come from application scenarios such as text summarization scenarios or foreign language translation. In this example, the training data can be obtained according to the application scenario of the large language model corresponding to the reward model. For example, if the application scenario of the large language model is a foreign language translation scenario, the training data can come from a foreign language database.

[0065] In this example, obtaining multiple sets of training data can improve the diversity of the training data, facilitating subsequent precise optimization of the original reward model using the training data, enhancing the evaluation effect of the original reward model, and obtaining a target reward model with improved evaluation performance.

[0066] Among them, the annotation result refers to the score obtained by annotating the training data. The first reward score refers to the score obtained by processing the annotation result corresponding to the training data.

[0067] As an example, in step S102, the computer device receives the annotation result corresponding to each set of training data, processes the annotation result, and obtains the first reward score for each set of training data. Among them, the annotation result corresponding to the training data can come from the client. For example, the annotation result can be the result of manual annotation of the training data obtained by the client. Another example is that the annotation result can also be the result of annotating each set of training data by the client. It can be understood that the annotation result corresponding to the training data can be obtained through manual annotation or client annotation. Among them, manual annotation means that a human (annotator) judges the response quality of the training response to the prompt word according to the prompt word in each set of training data, and scores each set of training data according to the response quality. For example, if the human (annotator) believes that the response quality of the training response in a certain set of training data is high, a higher evaluation score, such as 10 points, is given to this set of training data; if the human (annotator) believes that the response quality of the training response in a certain set of training data is average, an average evaluation score, such as 1 point, is given to this set of training data; if the human (annotator) believes that the response quality of the training response in a certain set of training data is poor, a lower evaluation score, such as -10 points, is given to this set of training data. In this example, the computer device obtains the annotation result of scoring each set of training data sent by the client, determines the score of each set of training data, and processes the score of each set of training data to obtain the first reward score corresponding to each set of training data. For example, the computer device obtains the evaluation score of manual annotation corresponding to each training data, and processes the evaluation score of manual annotation corresponding to each training data to obtain the first reward score corresponding to each training data.

[0068] Understandably, when evaluating the score of the output text of a large language model through manual annotation, it usually has a good evaluation effect. However, since the manual annotation method usually wastes a large amount of human resources, it is necessary to use a reward model to replace manual annotation to evaluate the score of the output text of the large language model. Therefore, during the optimization process of the reward model, when the score evaluated by the reward model for the training data is close to the score evaluated by manual annotation for the training data, it is considered that the reward model has better evaluation performance and can produce better evaluation effects. In this example, the annotation results corresponding to the training data are received, and based on the annotation results, the first reward score for each group of training data is determined, which is convenient for subsequent optimization and update of the reward model according to the first reward score, so that the evaluation effect of the reward model is close to that of manual annotation and the evaluation performance of the reward model is improved.

[0069] Among them, the original reward model refers to the reward model before optimization. The second reward score refers to the score obtained by using the original reward model to evaluate the training data. Among them, the reward model refers to a model used to describe and calculate the reward score of behavior in reinforcement learning. In this example, the last decoding output layer of the fine-tuned large language model is replaced with a linear fully connected layer to obtain a reward model. Among them, the large language model can include but is not limited to the chatGPT model and the T5 (Text ToText Transfer Transformer) model. This reward model is used to replace manual annotation and score and annotate the text content generated by the large language model it applies to, so as to accelerate the alignment training process of the large language model.

[0070] As an example, in step S103, the computer device uses the original reward model to score and annotate each group of training data to obtain the score of each group of training data, and determines the score of each group of training data as the second reward score corresponding to the group of training data. In this example, the computer device uses the original reward model to determine the response quality of the training response in each training data relative to the corresponding prompt word, and scores each group of training data according to the response quality corresponding to each training data to obtain the second reward score of each group of training data. Understandably, since the original reward model has not been optimized, it usually has low evaluation performance and poor evaluation effects, and it is necessary to optimize the original reward model. In this example, the second reward score obtained by the original reward model for scoring and annotating each training data is obtained, which is convenient for subsequent determination of the optimization function value corresponding to the original reward model based on the first reward score determined by manual annotation and the second reward score determined by the original reward model annotation, so as to achieve the purpose of optimizing the original reward model.

[0071] In this example, the original reward model can be applied to the large language model to be aligned, used to evaluate and score the text data output by the large language model to be aligned, and accelerate the alignment process of the large language model to be aligned. The large language model to be aligned can be applied to scenarios such as text summarization, sentiment analysis, and translation to implement the function of replying to the prompt words for the corresponding scenarios.

[0072] Among them, the optimized function value refers to the function value used to optimize the original reward model determined according to the first reward score and the second reward score of the training data.

[0073] As an example, in step S104, after the computer device obtains the first reward score and the second reward score corresponding to each training data, according to the first reward score and the second reward score corresponding to the same prompt word and the same training reply, it determines the optimized function value corresponding to the original reward model. For example, the computer device directly obtains the difference value between the first reward score and the second reward score corresponding to the same prompt word and the same training reply, and obtains the sum of the above difference values corresponding to all the same prompt words and the same training replies, to obtain the optimized function value corresponding to the original reward model.

[0074] It can be understood that the same prompt word can correspond to at least one training reply. A prompt word and a training reply form a set of training data. For any two sets of training data, they may have the same prompt word and different training replies, or they may have different prompt words and different training replies. In this example, the computer device classifies the training data according to different prompt words, obtains the first reward score and the second reward score of multiple training data corresponding to the same prompt word, and determines the optimized function value corresponding to the original reward model based on the first reward score and the second reward score of multiple training data corresponding to the same prompt word, which is convenient for subsequent optimization and update of the original reward model according to the optimized function value to improve the evaluation effect of the original reward model.

[0075] Among them, the convergence condition includes that the optimized function value is not higher than the preset function value, or the difference between the optimized function values determined twice in succession is not greater than the preset difference.

[0076] As an example, in step S105, after the computer device determines the optimization function value corresponding to the original reward model, it determines whether the optimization function value meets the convergence condition. When it determines that the optimization function value does not meet the convergence condition, it further optimizes the model parameters of the original reward model and continues to execute step S101 until the optimization function value meets the convergence condition. Understandably, if the optimization function value does not meet the convergence condition, it indicates that the evaluation performance of the original reward model is low and the evaluation effect does not reach the expected effect. It is necessary to further optimize the parameters of the original reward model and continue to execute steps S101 to S104 until it is determined that the optimization function value meets the convergence condition, indicating that the evaluation performance of the original reward model corresponding to the optimization function value is better and the evaluation effect can reach the expected effect. In this example, when the computer device determines that the optimization function value is higher than the preset function value, or the difference between the currently determined optimization function value and the previously determined optimization function value is greater than the preset difference, it determines that the optimization function value does not meet the convergence condition. In this example, when the optimization function value does not meet the convergence condition, optimizing the model parameters of the original reward model can enable the original reward model to have better evaluation performance and improve the evaluation effect of the original reward model.

[0077] Among them, the target reward model refers to the optimized original reward model.

[0078] As an example, in step S106, after the computer device determines the optimization function value corresponding to the original reward model, it determines whether the optimization function value meets the convergence condition. When it determines that the optimization function value meets the convergence condition, it determines the original reward model corresponding to the optimization function value as the target reward model. Understandably, if it is determined that the optimization function value meets the convergence condition, it indicates that the evaluation performance of the original reward model corresponding to the optimization function value is high and the evaluation effect can reach the expected effect. There is no need to optimize the original reward model again, and the original reward model that meets the convergence condition can be directly determined as the target reward model. Moreover, the evaluation effect of this target reward model is relatively close to the manual annotation effect, improving the evaluation performance of the original reward model. In this example, when the computer device determines that the optimization function value is not higher than the preset function value, or the difference between the currently determined optimization function value and the previously determined optimization function value is not greater than the preset difference, it determines that the optimization function value meets the convergence condition. In this example, when the optimization function value meets the convergence condition, determining the original reward model corresponding to the optimization function value that meets the convergence condition as the target reward model can achieve the purpose of improving the evaluation performance of the original reward model.

[0079] In this embodiment, obtaining multiple sets of training data can improve the diversity of the training data, facilitating subsequent more accurate optimization of the original reward model using the training data and enhancing the evaluation performance of the reward model. Receiving the annotation results corresponding to the training data and determining the first reward score for each training data based on the annotation results facilitate subsequent optimization and update of the reward model according to the first reward score, making the evaluation effect of the reward model close to the evaluation effect of manual annotation and enhancing the evaluation performance of the reward model. Determining the optimization function value corresponding to the original reward model based on the first and second reward scores of multiple training data corresponding to the same prompt word facilitates subsequent update of the original reward model according to the optimization function value and improves the evaluation effect of the original reward model. When the optimization function value does not meet the convergence condition, optimizing the model parameters of the original reward model; when the optimization function value meets the convergence condition, determining the original reward model corresponding to the optimization function value that meets the convergence condition as the target reward model can make the evaluation effect of the reward model relatively close to the evaluation effect of manual annotation, achieving the purpose of enhancing the evaluation performance of the original reward model.

[0080] In another embodiment, after step S106, the reward model optimization method further includes:

[0081] S1061: Input the prompt word corresponding to the target scenario into the language model to be aligned, and obtain the target text data output by the language model to be aligned;

[0082] S1062: Input the target text data into the target reward model to obtain the target score corresponding to the target text data.

[0083] Herein, the target scenario refers to the source scenario of the prompt word. For example, the prompt word can come from scenarios such as text or images. The language model to be aligned refers to the language model to be aligned. For example, the language model to be aligned can be a large language model to be aligned. Understandably, the language model can generate response text data according to the prompt word of the corresponding scenario to achieve text summarization, sentiment analysis, and translation of the text scenario corresponding to the prompt word, or image analysis and recognition of the image scenario corresponding to the prompt word. Usually, it is necessary to perform alignment training on the language model to be aligned according to the prompt word of the target scenario so that the trained language model has a better response effect on the prompt word of the target scenario. The target text data includes the prompt word and the response text data output by the language model to be aligned.

[0084] As an example, in step S1061, the computer device obtains the prompt words corresponding to the target scenario, inputs the prompt words corresponding to the target scenario into the language model to be aligned, outputs the response text data corresponding to the prompt words, and determines the prompt words corresponding to the target scenario and the response text data as the target text data, which is convenient for subsequently using the target reward model to score and label the target text data to obtain the target score.

[0085] Among them, the target score refers to the score of the target text data by the target reward model. Understandably, during the alignment training process of the language model to be aligned, the target reward model scores and labels the target text data generated by processing the prompt words of the target scenario by the language model to be aligned, generates the target score, which is used to reflect the application effect of the language model to be aligned in the target scenario, so as to determine whether to continue training the language model to be aligned according to this application effect, and realize the alignment training process of the language model to be aligned. Therefore, during the alignment process of the language model to be aligned, it is necessary to score and label the target text data output by the language model to be aligned through the target reward model to generate the target score.

[0086] As an example, in step S1062, the computer device inputs the target text data into the target reward model, and uses the target reward model to score and label the target text data to obtain the target score. In this example, the target reward model is used to replace manual scoring and labeling of the target text data, and the target reward model has an evaluation effect that is relatively close to manual labeling and has excellent evaluation performance.

[0087] In this embodiment, the optimized target reward model is used to score and label the target text data of the language model to be aligned, and determine whether the response text data in the target text data of the language model to be aligned has a good response effect. Instead of manual scoring and labeling, it has an evaluation effect that is relatively close to manual labeling, which can not only save human resources, but also speed up the alignment process of the language model to be aligned.

[0088] In one embodiment, as Figure 2 shown, step S101, that is, obtaining training data, includes:

[0089] S201: Perform sampling processing on the source data to obtain at least one prompt word;

[0090] S202: Perform multi-dimensional answer processing on each prompt word to obtain at least one original response corresponding to each prompt word;

[0091] S203: Use the target intention injection model to process the original response corresponding to each prompt word to obtain at least one training response corresponding to each prompt word.

[0092] Among them, the source data refers to the dataset used to obtain the prompt words.

[0093] As an example, in step S201, the computer device determines the source data for obtaining the prompt words, samples the source data, and obtains at least one prompt word. In this example, the source data comes from the scenarios applied by the language model to be aligned corresponding to the original reward model. For example, the source data comes from text scenarios corresponding to text summarization, sentiment analysis, or translation, etc. The language model to be aligned can be the large language model to be aligned. In this example, sampling the prompt words from the source data makes it feasible to determine the original response according to the prompt words subsequently.

[0094] Among them, the original response refers to the response data obtained after performing multi-dimensional response processing on the prompt words. Multi-dimensional response processing refers to performing response processing on the prompt words in multiple ways.

[0095] As an example, in step S202, after the computer device obtains the prompt words, it performs multi-dimensional response processing on each prompt word to obtain at least one original response corresponding to each prompt word. In this example, the computer device uses response methods such as artificial response, existing large language model response, and large language model response to be aligned, etc., to perform response processing on each prompt word to obtain at least one original response corresponding to each prompt word. Among them, when using the two methods of existing large language model response and large language model response to be aligned to obtain the original response corresponding to the prompt word, at least one original response corresponding to each prompt word can be obtained by adjusting the model parameters of the existing large language model and the large language model to be aligned. For example, the model parameters such as the temperature coefficient, Top P parameter, and / or TopK parameter in the existing large language model and the large language model to be aligned can be adjusted. In this example, performing multi-dimensional response processing on each prompt word to obtain at least one original response corresponding to each prompt word is convenient for subsequently obtaining at least one training response corresponding to each prompt word according to the at least one original response corresponding to each prompt word, improving the diversity and difference of the training data, and facilitating subsequent more accurate optimization of the original reward model according to the training data corresponding to the training response. In particular, by adjusting the model parameters such as the temperature coefficient, Top P parameter, and / or Top K parameter of the existing large language model and the large language model to be aligned, at least one original response corresponding to each prompt word is obtained, which is convenient for subsequently obtaining more diverse training responses, reducing the common defects between the training responses corresponding to the same prompt word, increasing the difference and diversity between the training responses corresponding to the same prompt word, and enabling more accurate optimization of the original reward model according to the training data corresponding to the training response subsequently.

[0096] Among them, the target intent injection model is used to convert the original response corresponding to each prompt into the training response corresponding to each prompt. In this example, the target intent injection model is constructed according to the intent expected by the developer, and it is a model for converting the original response. For example, if the developer hopes that the computer device processes at least one obtained original response so that the original response outputs the training response according to the intent expected by the developer, then a model corresponding to the intent expected by the developer is pre-constructed to form the target intent injection model, which is stored in the system database of the computer device. So that when the computer device needs to process the original response, it calls the target intent injection model in the system database to perform conversion processing on the original response and outputs the obtained training response.

[0097] In this example, the target intent injection model includes a refusal to answer intent injection model, an output format intent injection model, and an instruction following intent injection model.

[0098] Among them, the refusal to answer intent injection model is a model for answering sensitive words. Understandably, if the prompt word corresponding to the original response is a sensitive word, then the refusal to answer intent injection model needs to be used to answer the sensitive word to obtain the training response. The developer needs to pre-construct a model that can identify sensitive words and output the training response corresponding to the sensitive words to form the refusal to answer intent injection model, and pre-store the refusal to answer intent injection model in the computer database to answer the sensitive words and obtain the training response corresponding to the sensitive words.

[0099] The output format intent injection model is a model that processes the output format of the original response to obtain the training response, so that the training response has an output format that meets the expectations of the developer. Understandably, if the developer hopes that the computer device processes at least one obtained original response so that the original response outputs the training response according to the output format expected by the developer, then a model corresponding to the output format expected by the developer needs to be pre-constructed to form the target intent injection model, which is stored in the system database of the computer device. So that when the computer device needs to process the original response, it calls the target intent injection model in the system database to perform conversion processing on the original response and outputs the obtained training response with an output format that meets the expectations of the developer. For example, if the expected output format of the developer is to output the original response in the subject-verb-object order, then the model code is edited according to the standard of recognizing and sorting the word order of the original response to obtain an output format intent injection model that can recognize and sort the word order for output.

[0100] The instruction-following intent injection model is a model that performs fusion processing on the original response to obtain a training response with relatively complex logic. Understandably, if the prompt has complex logic, a training response with relatively complex logic needs to be output to improve the response quality of the training response. Therefore, the instruction-following intent injection model is required to perform fusion processing on the original response corresponding to the prompt with complex logic to obtain a training response with relatively complex logic. Developers need to pre-construct a model that can identify complex logic in prompts and perform fusion processing on the original response to form an instruction-following intent injection model and store it in the system database of the computer device for performing fusion processing on the original response corresponding to the prompt with complex logic.

[0101] As an example, in step S203, after the computer device obtains the original responses with diversity for each prompt, it uses the target intent injection model to process all the original responses corresponding to each prompt to obtain at least one training response corresponding to each prompt. In this example, after the computer device obtains the original responses with diversity for each prompt, it calls the reject response intent injection model, output format intent injection model, and instruction-following intent injection model stored in the system database to process each prompt and the original response corresponding to each prompt, and respectively obtains the training responses corresponding to each prompt output by the reject response intent injection model, output format intent injection model, and instruction-following intent injection model. In this example, using the target intent injection model to process the original response corresponding to the prompt to obtain at least one training response corresponding to the prompt can not only maintain the diversity and difference of the training responses corresponding to each prompt, but also facilitate subsequent scoring and annotation of each prompt and the training responses corresponding to each prompt.

[0102] In this embodiment, performing multi-dimensional response processing on each prompt to obtain at least one original response corresponding to each prompt facilitates subsequent obtaining of at least one training response corresponding to each prompt based on the at least one original response corresponding to each prompt, improving the diversity and difference of the training data, and facilitating subsequent more accurate optimization of the original reward model based on the training data corresponding to the training response. Using the target intent injection model to process the original response corresponding to the prompt to obtain at least one training response corresponding to the prompt can not only maintain the diversity and difference of the training responses corresponding to each prompt, but also facilitate subsequent scoring and annotation of each prompt and the training responses corresponding to each prompt.

[0103] In one embodiment, the target intent injection model in step S203 includes a reject response intent injection model. Among them, the reject response intent injection model is used to process the original response corresponding to the prompt with sensitive words.

[0104] In one embodiment, as Figure 3 shown, step S203, that is, using the target intention injection model to process the original response corresponding to each prompt word to obtain at least one training response corresponding to each prompt word, includes:

[0105] S301: Identify sensitive words in each prompt word to determine whether the prompt word is a sensitive word;

[0106] S302: If the prompt word is a sensitive word, use the reject answer intention injection model to process the original response corresponding to the sensitive word to obtain the first training response corresponding to the sensitive word;

[0107] S303: If the prompt word is not a sensitive word, optimize the original response corresponding to the prompt word to obtain at least one second training response corresponding to the prompt word.

[0108] Among them, sensitive word recognition is used to identify sensitive words in the prompt word.

[0109] As an example, in step S301, the computer device identifies sensitive words in each prompt word to determine whether the prompt word is a sensitive word and / or an illegal word. In this example, identifying sensitive words in the prompt word to determine whether the prompt word is a sensitive word facilitates subsequent classification processing of the original response corresponding to the prompt word to obtain the training data corresponding to the prompt word.

[0110] Among them, the first training response is the training response corresponding to the prompt word that is a sensitive word.

[0111] As an example, in step S302, when the computer device determines that the prompt word is a sensitive word, it uses the reject answer intention injection model to process the original response corresponding to the prompt word that is a sensitive word to obtain the first training response corresponding to the prompt word that is a sensitive word. In this example, when it is determined that the prompt word is a sensitive word, using the reject answer intention injection model to process the original response corresponding to the sensitive word to obtain the first training response corresponding to the sensitive word facilitates subsequent optimization of the original reward model according to the first training response.

[0112] Among them, the second training response is the training response corresponding to the prompt word that is not a sensitive word.

[0113] As an example, in step S303, when the computer device determines that the prompt word is not a sensitive word, it further optimizes the original reply corresponding to the prompt word to obtain at least one second training reply corresponding to each prompt word. In this example, optimizing the original reply corresponding to the prompt word that is not a sensitive word to obtain at least one second training reply makes it feasible to score and label the training data corresponding to the second training reply in different ways later. Understandably, the original reply corresponding to the prompt word may have situations where the output format, output word order, and output logic are not standardized. Therefore, it is necessary to further optimize the original reply corresponding to the prompt word that does not contain sensitive words to obtain a second training reply with a more standardized output format, output word order, and output logic for each prompt word, which is convenient for scoring the training data corresponding to the second training reply later.

[0114] In this embodiment, sensitive word recognition is performed on the prompt word, and according to whether the prompt word is a sensitive word, the original reply corresponding to the prompt word is processed to obtain the corresponding training reply, which is convenient for scoring the training data corresponding to the training reply later and achieving the purpose of optimizing the original reward model.

[0115] In one embodiment, as Figure 4 shown, step S302, that is, using the refusal to answer intent injection model to process the original reply corresponding to the sensitive word to obtain the first training reply corresponding to the sensitive word, includes:

[0116] S401: If the prompt word is a sensitive word, then all the original replies corresponding to the sensitive word are determined to be invalid replies;

[0117] S402: Call the preset reply template corresponding to the sensitive word in the refusal to answer intent injection model to re-edit the invalid reply to obtain the first training reply corresponding to the sensitive word.

[0118] Among them, the invalid reply refers to the original reply corresponding to the prompt word that is a sensitive word.

[0119] As an example, in step S401, when the computer device determines that the prompt word is a sensitive word, all the original replies corresponding to the prompt word that is a sensitive word are determined to be invalid replies. Understandably, for the prompt word that is a sensitive word, it usually involves illegal or sensitive topics. Usually, during the alignment training process of the language model, the language model is trained to refuse to reply to such prompt words. Therefore, in this example, the original reply corresponding to the prompt word that is a sensitive word is determined to be an invalid reply, which is convenient for subsequently calling the preset reply template corresponding to the sensitive word in the refusal to answer intent injection model to re-edit the invalid reply to obtain the first training reply corresponding to the sensitive word.

[0120] Among them, the preset reply template is a template used to re-edit the invalid reply.

[0121] As an example, in step S402, after the computer device determines the original reply corresponding to the prompt word that is a sensitive word as an invalid reply, it calls the preset reply template corresponding to the sensitive word in the refusal to answer intent injection model to re-edit the invalid reply, obtaining the first training reply corresponding to the prompt word that is a sensitive word. By repeating the above process of calling the preset reply template, the invalid reply corresponding to each sensitive word is re-edited, obtaining the first training reply corresponding to each prompt word that is a sensitive word. In this example, based on the refusal to answer intent injection model, the first training reply corresponding to each prompt word that is a sensitive word is obtained, facilitating subsequent scoring and annotation of the first training reply in different ways, and more accurately determining the optimization function value corresponding to the original reward model.

[0122] In this embodiment, determining the original reply corresponding to the prompt word that is a sensitive word as an invalid reply facilitates subsequent calling of the preset reply template corresponding to the sensitive word in the refusal to answer intent injection model to re-edit the invalid reply, obtaining the first training reply corresponding to the prompt word that is a sensitive word. Based on the refusal to answer intent injection model, the first training reply corresponding to each prompt word that is a sensitive word is obtained, facilitating subsequent scoring of the first training reply in different ways, and more accurately determining the optimization function value corresponding to the original reward model.

[0123] In one embodiment, the target intent injection model in step S203 includes an output format intent injection model and an instruction following intent injection model.

[0124] Among them, the output format intent injection model is used to process the output format of the original reply corresponding to the prompt word. The instruction following intent injection model is used to perform a fusion process on the original reply corresponding to the prompt word.

[0125] In one embodiment, step S303, that is, optimizing the original reply corresponding to the prompt word to obtain at least one second training reply corresponding to the prompt word, includes:

[0126] S3031: Using the output format intent injection model to process the output format of the original reply corresponding to the prompt word, obtaining at least one second training reply corresponding to the prompt word;

[0127] S3032: Using the instruction following intent injection model to perform a fusion process on at least two original replies corresponding to the prompt word, obtaining at least one second training reply corresponding to the prompt word.

[0128] Among them, the output format refers to the output word order of the sentence.

[0129] As an example, in step S3031, after the computer device determines that the prompt word is not a sensitive word, it uses an output format intent injection model to perform output format processing on at least one original reply corresponding to each prompt word, and obtains at least one second training reply corresponding to each prompt word. In this example, the computer device can use the output format intent injection model to split and combine multiple original answers corresponding to each prompt word, construct a training reply with an output format that meets the expectations of developers, and determine the training reply as the second training reply. In this example, using the output format intent injection model to perform output format processing on the original reply corresponding to each prompt word to obtain the second training reply makes it more convenient to score by manual annotation in the follow-up, can improve the efficiency of manual annotation, and can also reduce the recognition difficulty of the original reward model and improve the scoring and annotation efficiency of the original reward model.

[0130] As an example, in step S3032, after the computer device determines that the prompt word is not a sensitive word, it uses an instruction following intent injection model to perform fusion processing on at least two original replies corresponding to the same prompt word, obtains at least one second training reply corresponding to the prompt word, and processes the original reply corresponding to each prompt word in the above manner to obtain at least one second training reply corresponding to each prompt word. In this example, the computer device judges the logical complexity of the prompt word. When it determines that the prompt word has complex logic, it uses the instruction following intent injection model to perform at least one fusion processing on at least two original replies corresponding to the prompt word with complex logic according to a certain weight, realizes the rearrangement and combination of the original replies, and obtains at least one second training reply corresponding to the same prompt word.

[0131] For example, for M original replies corresponding to the same prompt word, different numbers of original replies can be selected multiple times from the M original replies, and the selected original replies can be fused according to different weights to obtain at least one second training reply corresponding to the agreed prompt word. In this example, using the instruction following intent injection model to perform fusion processing on at least two original replies corresponding to the same prompt word to obtain at least one second training reply corresponding to the same prompt word can improve the efficiency of manual annotation of the second training reply by humans in the follow-up, and can also reduce the recognition difficulty of the original reward model and improve the scoring and annotation efficiency of the original reward model.

[0132] In this embodiment, after determining that the prompt word is not a sensitive word, the output format intent injection model and the instruction following intent injection model are respectively used to optimize the original reply corresponding to the prompt word. While improving the efficiency of manual annotation of the second training reply by humans in the follow-up, it can also reduce the recognition difficulty of the original reward model, improve the scoring and annotation efficiency of the original reward model, and thus improve the optimization efficiency of the original reward model.

[0133] In one embodiment, as Figure 5 shown, step S102, that is, receiving the annotation result corresponding to the training data, and determining the first reward score of each training data based on the annotation result, includes:

[0134] S501: Receiving the initial reward score corresponding to each sentence in the training response of the training data;

[0135] S502: Performing statistical processing on the initial reward scores corresponding to each sentence in the training response to obtain the first reward score corresponding to the training response.

[0136] Wherein, the initial reward score refers to the reward score received for the annotation of each sentence in the training response.

[0137] As an example, in step S501, the computer device receives the annotation result corresponding to the training response of each training data, determines the score corresponding to each sentence in the training response in the annotation result, and determines the score as the initial reward score corresponding to each sentence. For example, for the k-th sentence in the j-th training response corresponding to the i-th prompt word, the annotated initial reward score is g ijk . In this example, the computer device can receive the score of each sentence in each training response obtained by manual scoring annotation sent by the client, and this score is the initial reward score corresponding to each sentence. For example, during manual scoring annotation, if it is determined that the response quality of the k-th sentence in the j-th training response corresponding to the i-th prompt word is good, then the score of g ijk is given a value of 10; if it is determined that the response quality of the k-th sentence in the j-th training response corresponding to the i-th prompt word is average, then the score of g ijk is given a value of 1; if it is determined that the response quality of the k-th sentence in the j-th training response corresponding to the i-th prompt word is poor, then the score of g ijk is given a value of -10.

[0138] Understandably, since sorting the training responses according to the human preference order and overall scoring and annotating each training response cannot focus on the response quality of each sentence in the training response. For example, since there are a small number of sentences with high response quality in the training response while the response quality of the remaining majority of sentences is low, when overall scoring and annotating, the majority of sentences with low response quality will not be considered. The small number of sentences with high response quality will deceive and cause the training response to obtain a high initial reward score, while the training response does not have a response quality matching the obtained initial reward score, resulting in the initial reward score being unable to focus on the response quality of local sentences and not conforming to the actual situation. In this example, by scoring and annotating each sentence of each training response to obtain the initial reward score corresponding to each sentence, rather than sorting the training responses according to the human preference order or overall scoring and annotating each training response, this method of scoring and annotating each sentence of each training response can overcome the defect that the above methods cannot focus on the response quality of each sentence in the training response, pay more attention to the response quality of each sentence, that is, focus on the local response quality of the training response, which is convenient for subsequently paying attention to the overall response quality of the training response by obtaining the initial reward scores of each sentence in the training response, and is more in line with the actual situation. In this example, obtaining the initial reward scores for scoring and annotating each sentence of the training response corresponding to each prompt word is convenient for realizing the attention to the local response quality of the training response, and can overcome the defects that sorting the training responses according to the human preference order and overall scoring and annotating each training response cannot focus on the response quality of each sentence in the training response.

[0139] Among them, the first reward score refers to the reward score corresponding to the training response obtained by processing the initial reward scores corresponding to each sentence in the training response.

[0140] As an example, in step S502, the computer device statistically processes the initial reward scores corresponding to each sentence in the training response to obtain the first reward score corresponding to the training response. In this example, for N prompt words, each prompt word corresponds to M training responses, and each training response corresponds to Z sentences. For the j-th training response corresponding to the i-th prompt word, the average score S ij The calculation formula of is:

[0141]

[0142] Among them, g ijk is the initial reward score annotated for the k-th sentence in the j-th training response corresponding to the i-th prompt word. The average score S corresponding to the j-th training response ijDetermine the first reward score corresponding to the j-th training response. In the above manner, obtain the first reward score of each training response corresponding to each prompt word, and this first reward score is the first reward score of the training data corresponding to the training response. In this example, by statistically processing the initial reward scores corresponding to each sentence in the training response, the first reward score corresponding to the training response is obtained, which can realize the attention to the overall response quality of the training response by focusing on the partial response quality of the training response. While paying attention to the overall response quality corresponding to the training response, also pay attention to the partial response quality corresponding to the training response, which is convenient for subsequent more accurate optimization of the original reward model based on the first reward score.

[0143] In this embodiment, obtain the initial reward scores for scoring and annotating each sentence of the training responses corresponding to each prompt word, which is convenient for realizing the attention to the partial response quality of the training responses. By statistically processing the initial reward scores corresponding to each sentence in the training response, the first reward score corresponding to the training response is obtained, which can realize the attention to the overall response quality of the training response by focusing on the partial response quality of the training response. While paying attention to the overall response quality corresponding to the training response, also pay attention to the partial response quality corresponding to the training response, which is convenient for subsequent more accurate optimization of the original reward model based on the first reward score, and achieve the purpose of improving the evaluation performance of the original reward model.

[0144] In one embodiment, as Figure 6 shown, step S104, that is, based on the first reward scores and second reward scores of multiple training data corresponding to the same prompt word, determine the optimization function value corresponding to the original reward model, including:

[0145] S601: Based on the first reward scores and second reward scores of multiple training data corresponding to the same prompt word, determine the first difference corresponding to the same prompt word;

[0146] S602: Based on the first differences corresponding to multiple prompt words, determine the optimization function value corresponding to the original reward model.

[0147] Among them, the first difference refers to the difference of multiple training responses corresponding to the same prompt word determined according to the difference between the first reward score and the second reward score of the same training response corresponding to the same prompt word.

[0148] As an example, in step S601, the computer device obtains the difference between the first reward score and the second reward score of each training data corresponding to each prompt word, and performs calculation processing on the difference between the first reward score and the second reward score of each training data corresponding to the same prompt word to obtain the first difference corresponding to the same prompt word. For example, for the j-th training response corresponding to the i-th prompt word, its corresponding first reward score is S ij, the second reward score is [P ij , then calculate the difference between the first reward score and the second reward score of the j-th training response corresponding to the i-th prompt word as (S ij - [P ij ) 2 , if there are M training responses for the i-th prompt word, the first difference corresponding to the i-th prompt word is:

[0149]

[0150] In the above manner, obtain the first difference corresponding to each prompt word. In this example, based on the first reward scores and the second reward scores of multiple training data corresponding to the same prompt word, determine the first difference corresponding to the same prompt word, which is convenient for calculating the optimization function value corresponding to the original optimization model according to the first difference later.

[0151] As an example, in step S602, the computer device calculates and processes the first difference corresponding to each prompt word to obtain the optimization function value corresponding to the original reward model. In this example, if there are N prompt words and each prompt word corresponds to M training responses, the optimization function value corresponding to the original reward model is:

[0152]

[0153] In this example, sum up the first differences corresponding to each prompt word, and use the minimize function to constrain the sum of the first differences to obtain the optimization function value corresponding to the original reward model, so that in the process of optimizing the original reward model according to the optimization function value, the evaluation effect of the original reward model on each training data corresponding to each prompt word is infinitely close to the manual annotation effect, achieving the purpose of improving the evaluation performance of the original reward model. This optimization function value reflects whether the evaluation effect of the original reward model on each training data corresponding to each prompt word is close to the manual annotation effect, that is, this optimization function value is used to reflect the evaluation performance of the original reward model. In this example, determining the optimization function value corresponding to the original reward model is convenient for more accurately judging the evaluation performance of the original reward model later.

[0154] In this embodiment, based on the first reward scores and the second reward scores of multiple training data corresponding to the same prompt, the first difference corresponding to the same prompt is determined, which can more intuitively and comprehensively reflect the difference between the evaluation effect of the original reward model on each prompt and the evaluation effect of manual scoring and annotation on each prompt. Based on the first difference corresponding to each prompt, the optimization function value corresponding to the original reward model is determined, and based on this optimization function value, the difference between the evaluation effect of the original reward model and the evaluation effect of manual scoring and annotation can be determined, which is convenient for further determining whether it is necessary to optimize the original reward model, improving the evaluation effect of the original reward model, and obtaining a target reward model with better evaluation performance.

[0155] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0156] In one embodiment, a reward model optimization device is provided, and this reward model optimization device corresponds one-to-one with the reward model optimization method in the above embodiment. As Figure 7 shown, this reward model optimization device includes a training data acquisition module 701, a first reward score determination module 702, a second reward score determination module 703, an optimization function value determination module 704, a model parameter optimization module 705, and a target reward model determination module 706. The detailed description of each functional module is as follows:

[0157] The training data acquisition module 701 is used to acquire training data, and the training data includes prompts and training responses;

[0158] The first reward score determination module 702 is used to receive the annotation results corresponding to the training data, and based on the annotation results, determine the first reward score of each training data;

[0159] The second reward score determination module 703 is used to score and annotate the training data by using the original reward model, and determine the second reward score of each training data;

[0160] The optimization function value determination module 704 determines the optimization function value corresponding to the original reward model based on the first reward scores and the second reward scores of multiple training data corresponding to the same prompt;

[0161] The model parameter optimization module 705 is used to optimize the model parameters of the original reward model when the optimization function value does not meet the convergence condition;

[0162] The target reward model determination module 706 is used to use the original reward model as the target reward model when the optimization function value meets the convergence condition.

[0163] In one embodiment, the training data acquisition module 701 includes:

[0164] A prompt word acquisition sub-module, configured to perform sampling processing on the source data to obtain at least one prompt word;

[0165] An original response acquisition sub-module, configured to perform multi-dimensional response processing on each prompt word to obtain at least one original response corresponding to each prompt word;

[0166] A training response acquisition sub-module, configured to use a target intent injection model to process the original response corresponding to each prompt word to obtain at least one training response corresponding to each prompt word.

[0167] In one embodiment, the training response acquisition sub-module includes:

[0168] A sensitive word recognition unit, configured to perform sensitive word recognition on each prompt word to determine whether the prompt word is a sensitive word;

[0169] A first sensitive word determination unit, configured to, if the prompt word is a sensitive word, use a refusal to answer intent injection model to process the original response corresponding to the sensitive word to obtain a first training response corresponding to the sensitive word;

[0170] A second sensitive word determination unit, configured to, if the prompt word is not a sensitive word, perform optimization processing on the original response corresponding to the prompt word to obtain at least one second training response corresponding to the prompt word.

[0171] In one embodiment, the first sensitive word determination unit includes:

[0172] An invalid response determination sub-unit, configured to, if the prompt word is a sensitive word, determine all the original responses corresponding to the sensitive word as invalid responses;

[0173] A first training response determination sub-unit, configured to call a preset response template corresponding to the sensitive word in the refusal to answer intent injection model to re-edit the invalid response to obtain a first training response corresponding to the sensitive word.

[0174] In one embodiment, the second sensitive word determination unit includes:

[0175] An output format processing sub-unit, configured to use an output format intent injection model to process the output format of the original response corresponding to the prompt word to obtain at least one second training response corresponding to the prompt word;

[0176] A fusion processing sub-unit, configured to use an instruction following intent injection model to perform fusion processing on at least two original responses corresponding to the prompt word to obtain at least one second training response corresponding to the prompt word.

[0177] In one embodiment, the first reward score determination module 702 includes:

[0178] An initial reward score determination sub-module, configured to receive the initial reward score corresponding to each sentence in the training response of the training data;

[0179] A first reward score determination sub-module, configured to perform statistical processing on the initial reward scores corresponding to each sentence in the training response to obtain the first reward score corresponding to the training response.

[0180] In one embodiment, the optimization function value determination module 704 includes:

[0181] A first difference determination sub-module, configured to determine a first difference corresponding to the same prompt word based on the first reward scores and second reward scores of multiple training data corresponding to the same prompt word;

[0182] An optimization function value determination sub-module, configured to determine the optimization function value corresponding to the original reward model based on the first differences corresponding to multiple prompt words.

[0183] For the specific limitations of the reward model optimization device, reference can be made to the limitations of the reward model optimization method in the foregoing text, which will not be elaborated here. Each module in the above reward model optimization device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or independent of it, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0184] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 8 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the data used or generated during the execution of the reward model optimization method. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a reward model optimization method.

[0185] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the reward model optimization method in the above embodiment. For example Figure 1As shown in S101-S106, or Figures 3 to 6 as shown in Figures 3 to 6 . To avoid repetition, it will not be elaborated here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the reward model optimization device. For example Figure 7 the functions of the training data acquisition module 701, the first reward score determination module 702, the second reward score determination module 703, the optimization function value determination module 704, the model parameter optimization module 705, and the target reward model determination module 706 as shown in Figure 7 . To avoid repetition, it will not be elaborated here.

[0186] In one embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the reward model optimization method in the above embodiment. For example Figure 1 as shown in S101-S106, or Figures 3 to 6 as shown in Figures 3 to 6 . To avoid repetition, it will not be elaborated here. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the above reward model optimization device. For example Figure 7 the functions of the training data acquisition module 701, the first reward score determination module 702, the second reward score determination module 703, the optimization function value determination module 704, the model parameter optimization module 705, and the target reward model determination module 706 as shown in Figure 7 . To avoid repetition, it will not be elaborated here. The computer-readable storage medium can be non-volatile or volatile.

[0187] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0188] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0189] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A method for optimizing a reward model, characterized in that, Including: Obtain training data, where the training data includes prompt words and training responses; Receive the annotation results corresponding to the training data, and based on the annotation results, determine the first reward score for each piece of training data; Use the original reward model to score and annotate the training data, and determine the second reward score for each piece of training data; Based on the first reward scores and second reward scores of multiple pieces of training data corresponding to the same prompt word, determine the optimization function value corresponding to the original reward model; When the optimization function value does not meet the convergence condition, optimize the model parameters of the original reward model; When the optimization function value meets the convergence condition, use the original reward model as the target reward model.

2. The reward model optimization method according to claim 1, wherein The obtaining of the training data includes: Perform sampling processing on the source data to obtain at least one prompt word; Perform multi-dimensional answer processing on each prompt word to obtain at least one original response corresponding to each prompt word; Use the target intent injection model to process the original responses corresponding to each prompt word to obtain at least one training response corresponding to each prompt word.

3. The reward model optimization method according to claim 2, characterized in that, The target intent injection model includes a refusal to answer intent injection model; The using of the target intent injection model to process the original responses corresponding to each prompt word to obtain at least one training response corresponding to each prompt word includes: Perform sensitive word recognition on each prompt word to determine whether the prompt word is a sensitive word; If the prompt word is a sensitive word, use the refusal to answer intent injection model to process the original response corresponding to the sensitive word to obtain the first training response corresponding to the sensitive word; If the prompt word is not a sensitive word, perform optimization processing on the original response corresponding to the prompt word to obtain at least one second training response corresponding to the prompt word.

4. The reward model optimization method according to claim 3, characterized in that The using of the refusal to answer intent injection model to process the original response corresponding to the prompt word to obtain the first training response corresponding to the prompt word includes: If the prompt word is a sensitive word, determine all the original responses corresponding to the sensitive word as invalid responses; Call the preset response template corresponding to the sensitive word in the refusal to answer intent injection model to re-edit the invalid responses to obtain the first training response corresponding to the sensitive word.

5. The reward model optimization method according to claim 3, wherein The target intent injection model includes an output format intent injection model and an instruction following intent injection model; The performing of optimization processing on the original response corresponding to the prompt word to obtain at least one second training response corresponding to the prompt word includes: Use the output format intent injection model to process the output format of the original response corresponding to the prompt word to obtain at least one second training response corresponding to the prompt word; Use the instruction following intent injection model to perform fusion processing on at least two original responses corresponding to the prompt word to obtain at least one second training response corresponding to the prompt word.

6. The reward model optimization method according to claim 1, wherein The receiving of the annotation results corresponding to the training data and determining the first reward score for each piece of training data based on the annotation results includes: The initial reward score corresponding to each sentence in the training response that receives the training data Perform statistical processing on the initial reward scores corresponding to each sentence in the training response to obtain the first reward score corresponding to the training response 7. The reward model optimization method according to claim 1, characterized in that Determining the optimization function value corresponding to the original reward model based on the first reward scores and second reward scores of multiple pieces of the training data corresponding to the same prompt includes Based on the first reward scores and second reward scores of multiple pieces of the training data corresponding to the same prompt, determine the first difference corresponding to the same prompt Based on the first differences corresponding to multiple prompts, determine the optimization function value corresponding to the original reward model 8. An apparatus for optimizing a reward model, characterized in that, including A training data acquisition module for acquiring training data, where the training data includes a prompt and a training response A first reward score determination module for receiving the annotation result corresponding to the training data and determining the first reward score of each piece of the training data based on the annotation result A second reward score determination module for using the original reward model to score and annotate the training data and determining the second reward score of each piece of the training data An optimization function value determination module for determining the optimization function value corresponding to the original reward model based on the first reward scores and second reward scores of multiple pieces of the training data corresponding to the same prompt A model parameter optimization module for optimizing the model parameters of the original reward model when the optimization function value does not meet the convergence condition A target reward model determination module for using the original reward model as the target reward model when the optimization function value meets the convergence condition 9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the reward model optimization method according to any one of claims 1 to 7 10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the reward model optimization method according to any one of claims 1 to 7

Citation Information

Cited By

  • Large language model optimization method and optimization device

    CN120633740A

  • Script generation model optimization method and device, equipment and storage medium

    CN121280832A