Language model training method, reward model training method, device and electronic equipment

By obtaining the reply generated by the language model, using the target tool to obtain reference information and generate target evaluation information, the problem of insufficient LLM training effect is solved, and the training effect and reply quality of the model are improved.

CN117273117BActive Publication Date: 2025-08-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311281419.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-08-22
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

The training effect of existing large-scale language models (LLM) is insufficient, affecting the user experience, and it is difficult to generate high-quality responses.

Method used

By obtaining the reply generated by the language model, calling the target tool to obtain reference information, generating target evaluation information based on the reference information, and adjusting model parameters during the reinforcement learning stage, and using the target tool to timely update information to improve the accuracy and timeliness of the evaluation.

Benefits of technology

The training effect of the language model has been improved, the quality of the generated reply is higher, the user experience is improved, and the model training process is more accurate and objective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117273117B_ABST
    Figure CN117273117B_ABST
Patent Text Reader

Abstract

The present disclosure provides a language model training method, a reward model training method, a device and an electronic device, which relate to the field of artificial intelligence technology, specifically the technical fields of deep learning, natural language processing, large models, etc., and can be applied to interactive scenarios based on artificial intelligence. The specific implementation scheme is: obtaining the response generated by the language model based on the target question; according to the target question, calling the target tool to obtain the reference information of the target question, and the reference information is determined based on the return data of the target tool; according to the reference information of the target question, generating target evaluation information for the response; according to the target evaluation information, adjusting the model parameters of the language model in the reinforcement learning stage. The reference information obtained by the target tool contains information related to the expected response to the target question. Based on the reference information, the response can be accurately and objectively evaluated. Therefore, adjusting the model parameters of the language model according to the target evaluation information can improve the training effect of the language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically deep learning, natural language processing, large models and other technical fields, and can be applied to interactive scenarios based on artificial intelligence, especially to language model training methods, reward model training methods, devices and electronic equipment. Background Art

[0002] Large Language Models (LLMs), also known as large models, are usually used as dialogue systems. Users first ask questions as the interaction context, and then the model gives appropriate responses.

[0003] The quality of LLM's responses is an important factor affecting user experience, and the quality of LLM's responses is related to the training effect of LLM. Therefore, how to improve the training effect of LLM is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] The present disclosure provides a language model training method, a reward model training method, a device, and an electronic device.

[0005] According to a first aspect of the present disclosure, a language model training method is provided, comprising:

[0006] Get the response generated by the language model based on the target question;

[0007] According to the target problem, call the target tool to obtain the reference of the target problem

[0008] Information, the reference information is determined based on the return data of the target tool;

[0009] Generate a target evaluation for the response based on the reference information of the target question

[0010] information;

[0011] According to the target evaluation information, model parameters of the language model are adjusted in the reinforcement learning stage.

[0012] According to a second aspect of the present disclosure, a reward model training method is provided, comprising:

[0013] Obtain training samples, wherein the training samples at least include

[0014] target question, the response output by the language model for the target question, and the

[0015] Expected evaluation information for reply;

[0016] The target question and the answer are input into a reward model, wherein the reward model

[0017] The target tool is used to call the target tool according to the target problem and obtain the parameters of the target problem.

[0018] Reference information, wherein the reference information is obtained based on return data of the target tool;

[0019] Generating target evaluation information of the reply according to the reference information;

[0020] The reward model is adjusted in terms of model parameters based on at least a difference between the target evaluation information and the expected evaluation information.

[0021] According to a third aspect of the present disclosure, a language model training device is provided, comprising:

[0022] A first acquisition module is used to obtain a response generated by the language model based on the target question;

[0023] The tool calling module is used to call the target tool according to the target problem and obtain

[0024] Reference information of the target problem, the reference information is based on the return of the target tool

[0025] Return data to confirm;

[0026] The first evaluation module is used to evaluate the target problem based on the reference information.

[0027] The above responses generate target evaluation information;

[0028] The first adjustment module is used to adjust the model parameters of the language model in the reinforcement learning stage according to the target evaluation information.

[0029] According to a fourth aspect of the present disclosure, there is provided a reward model training device, comprising:

[0030] The second acquisition module is used to acquire training samples, and the training samples at least include

[0031] A target question for inputting a language model, the language model inputting a target question for the target question

[0032] The responses given, and the expected evaluation information of the responses;

[0033] An input module is used to input the target question and the response into a reward model,

[0034] In the example, the reward model is used to call the target tool according to the target problem to obtain the

[0035] Reference information of the target problem, wherein the reference information is based on the return of the target tool

[0036] Data obtained;

[0037] The second evaluation module is used to generate the target of the reply based on the reference information.

[0038] Evaluation information;

[0039] The second adjustment module is configured to adjust model parameters of the reward model at least according to a difference between the target evaluation information and the expected evaluation information.

[0040] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0041] at least one processor; and

[0042] a memory communicatively connected to the at least one processor; wherein,

[0043] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the language model training method described in the first aspect or the reward model training method described in the second aspect.

[0044] According to the sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the language model training method described in the first aspect or the reward model training method described in the second aspect.

[0045] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the steps of the language model training method described in the first aspect or the reward model training method described in the second aspect.

[0046] The language model training method, reward model training method, device, and electronic device provided by the present disclosure have the following beneficial effects:

[0047] In the disclosed embodiment, for the reply generated by the language model based on the target question, the target tool is called according to the target question to obtain the reference information of the target question, and based on the reference information of the target question, target evaluation information is generated for the reply, and then the model parameters of the language model are adjusted in the reinforcement learning stage according to the target evaluation information. The reference information obtained by the target tool contains information related to the expected reply to the target question. Based on the reference information, the reply can be accurately and objectively evaluated, so that when the model parameters of the language model are adjusted according to the target evaluation information, the training process of the language model can be accurately and objectively guided, thereby improving the training effect of the language model. In addition, due to the timeliness of the target tool updating information, the time lag between the reference information and the target evaluation information is small. Therefore, the timeliness of the language model trained based on the target evaluation information is better.

[0048] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0050] Figure 1 1 is a flow chart of a language model training method provided according to an embodiment of the present disclosure;

[0051] Figure 2 is a flowchart of a language model training method provided according to another embodiment of the present disclosure;

[0052] Figure 3 is a flowchart of a language model training method provided according to another embodiment of the present disclosure;

[0053] Figure 4 1 is a flow chart of a reward model training method according to an embodiment of the present disclosure;

[0054] Figure 5 is a structural diagram of a language model training device provided according to an embodiment of the present disclosure;

[0055] Figure 6 is a schematic diagram of the structure of a reward model training device provided according to an embodiment of the present disclosure;

[0056] Figure 7 It is a block diagram of an electronic device used to implement the language model training method or reward model training method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0057] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0058] The following describes the language model training method, reward model training method, device and electronic device of the embodiments of the present disclosure with reference to the accompanying drawings.

[0059] It should be noted that the executor of the language model training method of this embodiment is the language model training device, and the executor of the reward model training method of this embodiment is the reward model training device. The language model training device and the reward model training device can be implemented by software and / or hardware, and can be configured in electronic devices. The electronic devices may include but are not limited to terminals, server terminals, etc.

[0060] Figure 1 It is a flowchart of a language model training method provided according to an embodiment of the present disclosure.

[0061] like Figure 1 As shown, the language model training method includes:

[0062] Step 101: Obtain a response generated by a language model based on a target question.

[0063] Optionally, the language model may refer to a model that can interact with the user, that is, a model that can give corresponding responses to the user's questions. As an example and not a limitation, the language model may refer to an LLM.

[0064] Optionally, the target question may refer to a question randomly selected from a question library, or may refer to a question input by a user.

[0065] Step 102: According to the target problem, the target tool is called to obtain reference information of the target problem. The reference information is determined based on the return data of the target tool.

[0066] For specific questions, such as calculation questions or code execution questions, the tool's response is more accurate than the language model's. Therefore, the present disclosure uses the tool to obtain the correct answer to the target question, that is, the second answer.

[0067] Different questions may require different tools. For example, a question about the weather might require a weather forecast tool. A question about calculation results, such as multiplication and division, might require a calculation tool, such as a calculator. An information query, such as "What is a banana?", might require a search engine. Specifically, tools may include calculators, code interpreters, translation systems, search engines, knowledge bases, weather forecast tools, calendars, and the like, though this embodiment does not limit these tools.

[0068] As an optional implementation, based on the target question, the target tool is called to obtain reference information of the target question, including: sending the target question to the target tool, or sending keywords randomly extracted from the target question to the target tool; receiving return data from the target tool to obtain reference information of the target question.

[0069] Some tools may not have parameter extraction capabilities or have poor parameter extraction capabilities. Therefore, the method of sending the target problem to the target tool to obtain reference information in the above embodiment may result in tool invocation failure and inability to obtain reference information due to the tool's lack of parameter extraction capabilities, or may result in slow reference information acquisition due to the tool's poor parameter extraction capabilities. To address the above issues, the present disclosure proposes implementing the target tool invocation through the following embodiment.

[0070] As another optional implementation, based on the target problem, the target tool is called to obtain reference information of the target problem, including: extracting the input parameters required by the target tool from the target problem; sending the input parameters to the target tool through a calling interface corresponding to the target tool, and receiving return data from the target tool to obtain the reference information.

[0071] Different tools require different input parameters. The present invention first extracts the input parameters required by the target tool from the target problem, and then sends the required input parameters to the target tool to obtain reference information. This method not only ensures that the target tool can be successfully called, but also improves the speed of obtaining reference information.

[0072] Step 103: Generate target evaluation information for the reply based on the reference information of the target question.

[0073] Optionally, the target evaluation information may be used to indicate the accuracy of the response, or may be used to indicate the degree of matching between the response and the target question, that is, to indicate whether the response does not answer the question.

[0074] As an example and not a limitation, suppose the target question is "How is the weather today?", the response is "It's sunny today", and the reference information is "XX year XX month XX day, XX region, cloudy weather, precipitation XX mm, temperature XX degrees Celsius". In this case, the corresponding target evaluation information may indicate that the accuracy of the response is low, or it may indicate that the response does not irrelevant to the question.

[0075] Assume that the target question is "How is the weather today?", the response is "Today is Monday", and the reference information is "XX year XX month XX day, XX region, cloudy, precipitation XX mm, temperature XX degrees Celsius". In this case, the corresponding target evaluation information may indicate that the accuracy of the response is very low, or may indicate that the response does not answer the question.

[0076] As an optional implementation, target evaluation information is generated for the reply based on the reference information of the target question, including: determining the difference between the reference information and the reply; scoring the reply based on at least the difference between the reference information and the reply to obtain the target evaluation information.

[0077] The difference between the reference information and the response output by the language model can be understood as the difference between the response output by the language model and the expected response. Therefore, by obtaining target evaluation information based on the difference between the reference information and the response output by the language model, the response output by the language model can be trained to be close to the expected response, thereby improving the response quality of the language model.

[0078] The reference information may include information that is irrelevant to the target question. In order to obtain accurate target evaluation information, information related to the target question can be extracted from the reference information, and the difference between the response output by the language model and the information related to the target question in the reference information can be determined to determine the target evaluation information.

[0079] As an example and not a limitation, assume that the target question is "How is the weather today?", the response output by the language model is "It's sunny today", and the reference information is "XX year XX month XX day, XX region, cloudy weather, precipitation XX mm, temperature XX degrees Celsius", then the information related to the target question extracted from the reference information can be "XX year XX month XX day, cloudy weather".

[0080] Optionally, the score and difference can be set to be negatively correlated, that is, the higher the score, the smaller the difference between the reply and the reference information, and the lower the score, the greater the difference between the reply and the reference information; the score and difference can also be set to be positively correlated, that is, the higher the score, the greater the difference between the reply and the reference information, and the lower the score, the smaller the difference between the reply and the reference information.

[0081] Optionally, the similarity between the reply and the reference information can be calculated to obtain the difference between the reference information and the reply. The difference between the reference information and the reply can also be determined by character comparison. As an example and not a limitation, the number of different characters in the reply and the reference information can be counted to obtain the difference between the reference information and the reply.

[0082] Optionally, determining the difference between the reference information and the response includes: in the case of multiple calls to the target tool, determining the difference between the reference information and the response obtained in each call.

[0083] Alternatively, keywords may be randomly extracted from the target question multiple times, and the target tool may be called multiple times based on the keywords extracted each time. Alternatively, the target tool may be called multiple times based on the target question or the input parameters required by the target tool extracted from the target question. The reference information obtained from each call to the target tool may be different.

[0084] A single call to the target tool may fail or return incorrect reference information due to accidental factors. The present disclosure determines the difference between the reference information and the response obtained from each call, which can avoid evaluation deviations caused by accidental factors and ensure the accuracy and objectivity of the target evaluation information.

[0085] Optionally, when the target tool is called multiple times, the replies are scored according to the difference between the reference information and the replies to obtain target evaluation information, including: determining the target evaluation information according to the difference between the reference information and the replies obtained in each call.

[0086] Specifically, the average difference is calculated based on the difference between the reference information and the response obtained from each call, and the target evaluation information is determined based on the average difference.

[0087] Taking the difference as similarity as an example, the similarity between the reference information and the reply obtained in each call is calculated, and the similarity mean is obtained based on multiple similarities to obtain the target evaluation information.

[0088] As another optional embodiment, target evaluation information is generated for the response based on the reference information of the target question, including: inputting the response and the reference information into an evaluation model, and obtaining target evaluation information output by the evaluation model. The target evaluation information can be a numerical value or a character such as a letter that can represent accuracy or matching.

[0089] It should be noted that for the same question, the reference information corresponding to the question may change at any time. For example, for weather information, compared with the language model, the weather information in the weather forecast tool is updated more promptly. Therefore, based on the reference information obtained through the target tool, the target evaluation information obtained has a smaller time lag and better timeliness.

[0090] Step 104: Adjust the model parameters of the language model in the reinforcement learning phase according to the target evaluation information.

[0091] Optionally, a first loss function may be constructed based on the target evaluation information, and model parameters of the language model may be adjusted in the reinforcement learning phase based on the first loss function.

[0092] Among them, when the target evaluation information is a character, the character can be mapped to a numerical value first, and then the loss is calculated based on the mapped numerical value.

[0093] Taking the target evaluation information as an example, when the score and the difference are negatively correlated, the score and the loss value of the first loss function are negatively correlated. The higher the score, the smaller the difference, and the smaller the loss value of the first loss function. The lower the score, the greater the difference, and the greater the loss value of the first loss function.

[0094] In the disclosed embodiment, for the reply generated by the language model based on the target question, the target tool is called according to the target question to obtain the reference information of the target question, and based on the reference information of the target question, target evaluation information is generated for the reply, and then the model parameters of the language model are adjusted in the reinforcement learning stage according to the target evaluation information. The reference information obtained by the target tool contains information related to the expected reply to the target question. Based on the reference information, the reply can be accurately and objectively evaluated, so that when the model parameters of the language model are adjusted according to the target evaluation information, the training process of the language model can be accurately and objectively guided, thereby improving the training effect of the language model. In addition, due to the timeliness of the target tool updating information, the time lag between the reference information and the target evaluation information is small. Therefore, the timeliness of the language model trained based on the target evaluation information is better.

[0095] Figure 2 It is a flowchart of a language model training method provided according to another embodiment of the present disclosure.

[0096] like Figure 2 As shown, the language model training method includes:

[0097] Step 201: Obtain a response generated by a language model based on a target question.

[0098] The specific implementation of step 201 can refer to the detailed description in other embodiments of the present disclosure and will not be described in detail here.

[0099] Step 202 : According to at least one intent corresponding to the target problem, a target tool corresponding to each intent is called, and reference information of the target problem is obtained based on return data of the target tool corresponding to each intent.

[0100] Optionally, the target question corresponds to at least one intent.

[0101] When there is an intent corresponding to the target problem, a target tool corresponding to the intent is called, and the return data of the target tool is reference information.

[0102] In the case where the target question corresponds to multiple intents, the target tools corresponding to each intent are called, and the return data of each target tool constitutes the reference information. It should be noted that there can be one or more target tools corresponding to each intent.

[0103] Each intent has a target tool that needs to be called. By obtaining the return data corresponding to each intent through the target tool corresponding to each intent, the accuracy and timeliness of the return data corresponding to each intent can be guaranteed, thereby ensuring the accuracy and timeliness of the reference information.

[0104] As an example and not a limitation, assume the target question is "What's the weather like today, and what are the license plate numbers of vehicles that are restricted today?" This target question corresponds to two intents, and the target tools corresponding to the two intents can be a weather forecast tool and a search engine tool, respectively. The return data of the weather forecast tool and the search engine tool constitute the reference information. For example, if the return data of the weather forecast tool is "xx / xx / xx, xx region, cloudy, precipitation xx mm, temperature xx degrees Celsius," and the return data of the search engine tool is "xx / xx / xx, restricted vehicles with license plate numbers 1 and 6, restricted hours from xx:xx to xx:xx," then the reference information is "xx / xx / xx, xx region, cloudy, precipitation xx mm, temperature xx degrees Celsius, restricted vehicles with license plate numbers 1 and 6, restricted hours from xx:xx to xx:xx."

[0105] Among them, if the search engine tool has a weather query function, the target tool corresponding to the two intentions of the target question can be the search engine tool.

[0106] Optionally, if there are multiple target tools corresponding to each intent, in order to make the questions and responses correspond one-to-one, the return data of each target tool can be sorted according to the arrangement order of the multiple intents corresponding to the target questions to obtain reference information of the target questions.

[0107] Step 203: Generate target evaluation information for the reply based on the reference information of the target question.

[0108] The specific implementation of step 203 can refer to the detailed description in other embodiments of the present disclosure and will not be described in detail here.

[0109] Step 204: Adjust the model parameters of the language model in the reinforcement learning phase according to the target evaluation information.

[0110] The training phase of the language model includes the pre-training phase, the supervised training phase, the reward modeling phase, and the reinforcement learning phase.

[0111] Reinforcement learning refers to fine-tuning the language model in the supervised training phase. Since the language model in the supervised training phase has already undergone a supervised training process, the probability distribution it obtains for the target problem is close to the expected probability distribution. In order to make the probability distribution of the language model in the reinforcement learning phase for the target problem also close to the expected probability distribution, the present disclosure also combines the language model in the supervised training phase to adjust the model parameters of the language model in the reinforcement learning phase. The specific process of model parameter adjustment is as follows:

[0112] According to the target evaluation information, the model parameters of the language model are adjusted in the reinforcement learning stage, including: obtaining a first probability distribution obtained by the language model in the reinforcement learning stage for the target question; wherein the first probability distribution is used to indicate the probability of each word in the word list for each word position of the reply; obtaining a second probability distribution obtained by the language model in the supervised training stage for the target question; wherein the second probability distribution is used to indicate the probability of each word in the word list for each word position of the third reply, and the third reply is the reply of the language model in the supervised training stage to the target question; comparing the probability distribution difference between the first probability distribution and the second probability distribution; and adjusting the model parameters of the language model in the reinforcement learning stage according to the target evaluation information and the probability distribution difference.

[0113] When a language model outputs a response to a question, for each word position in the response, each word in the vocabulary table has a corresponding probability of occurrence at that word position. Based on the probability of occurrence of each word in the vocabulary table at each word position and context information, the language model can predict the word that appears in each word position and generate a response. The vocabulary table can also be called a dictionary.

[0114] Optionally, a second loss function may be constructed based on the target evaluation information and the probability distribution difference, and model parameters of the language model in the reinforcement learning stage may be adjusted based on the second loss function.

[0115] There is a positive correlation between the probability distribution difference and the loss value of the second loss function. The smaller the probability distribution difference, the smaller the loss value of the second loss function, and the larger the probability distribution difference, the larger the loss value of the second loss function.

[0116] Optionally, the KL divergence between the first probability distribution and the second probability distribution may be calculated to obtain a probability distribution difference between the first probability distribution and the second probability distribution.

[0117] In the disclosed embodiment, based on at least one intent corresponding to a target question, a target tool corresponding to each intent is invoked. Reference information for the target question is obtained based on the return data from the target tool corresponding to each intent. Based on the reference information for the target question, target evaluation information is generated for the response. Based on the target evaluation information, model parameters of the language model are adjusted during the reinforcement learning phase. The disclosed embodiment can invoke the target tool corresponding to each intent, obtain the correct response corresponding to each intent, and thus obtain the correct response to the target question, ensuring the accuracy and objectivity of the target evaluation information.

[0118] Figure 3 It is a flowchart of a language model training method provided according to another embodiment of the present disclosure.

[0119] like Figure 3 As shown, the language model training method includes:

[0120] Step 301: Obtain a response generated by the language model based on the target question.

[0121] The specific implementation of step 301 can refer to the detailed description in other embodiments of the present disclosure and will not be described in detail here.

[0122] Step 302: Identify whether a tool call is required based on the target problem.

[0123] Optionally, the target question does not necessarily require a tool call, that is, it is not necessary to obtain the target evaluation information of the response with the help of reference information obtained based on the target tool. Therefore, in order to reduce unnecessary calculations, the present disclosure first determines whether the target question requires a tool call.

[0124] As an optional implementation, based on the target question, identifying whether a tool call is required includes: querying the mapping relationship between intent and tool based on the intent corresponding to the target question; and determining that a tool call is required when the target tool corresponding to the intent of the target question is queried.

[0125] The mapping relationship between intentions and tools can be pre-set.

[0126] Based on the mapping relationship between intent and tool, the method of determining whether a tool call is needed according to whether the target tool corresponding to the intent of the target problem is queried does not involve complex algorithms, so it is possible to identify whether a tool call is needed for the target problem without performing complex calculations.

[0127] An intent can correspond to one or more tools; wherein, multiple tools corresponding to an intent implement the same function. As an example and not a limitation, an intent to ask about the weather can correspond to multiple weather forecast tools.

[0128] As another optional implementation, identifying whether a tool call is required based on the target problem includes: classifying the target problem based on its semantics; and determining whether a tool call is required based on the classification result.

[0129] The semantic vector of the target question can be obtained and input into a classification model to obtain a classification result, which is used to indicate whether a tool call is required. The classification model can be a binary classification model.

[0130] In the above implementation method of determining whether a tool call is required by querying the mapping relationship between the intent and the tool, the query speed will decrease as the number of mapping relationships increases. When there are many mapping relationships, the speed of determining whether a tool call is required by this method is slow and inefficient.

[0131] The method of determining whether a tool call is needed is based on the classification results. The main operation process in this method is the classification process. The current classification algorithm is relatively mature and has a fast operation speed. Therefore, the speed of identifying whether a tool call is needed in this method is not affected by the number of mapping relationships, and the recognition speed is relatively fast.

[0132] Step 303: When it is identified that a tool call is required, a target tool is called according to the target problem to obtain reference information of the target problem.

[0133] Step 304: Generate target evaluation information for the reply based on the reference information of the target question.

[0134] Step 305: Adjust the model parameters of the language model in the reinforcement learning phase according to the target evaluation information.

[0135] The specific implementation of steps 303 to 305 can refer to the detailed description in other embodiments of the present disclosure and will not be described in detail here.

[0136] In the embodiment of the present disclosure, whether a tool call is required is first identified based on the target problem, and then, if it is identified that a tool call is required, the target tool is called according to the target problem to obtain reference information of the target problem, thereby reducing unnecessary calculations.

[0137] Figure 4 1 is a flow chart of a reward model training method provided according to an embodiment of the present disclosure.

[0138] like Figure 4 As shown, the reward model training method includes:

[0139] Step 401: Acquire a training sample, where the training sample includes at least a target question for input into a language model, a response output by the language model to the target question, and expected evaluation information of the response.

[0140] Alternatively, the desired evaluation information may be determined by the user.

[0141] Optionally, multiple training samples may be obtained, and the target questions included in each training sample may be different.

[0142] In step 402, the target question and the response are input into a reward model, wherein the reward model is used to call a target tool according to the target question and obtain reference information of the target question, wherein the reference information is obtained based on the return data of the target tool.

[0143] Optionally, the target question and the response in at least one training sample may be input into the reward model.

[0144] For the relevant content of calling the target tool in step 402, please refer to the detailed description in other embodiments of the present disclosure, and will not be repeated here.

[0145] Step 403: Generate reply target evaluation information based on the reference information.

[0146] The specific implementation of step 403 can refer to the detailed description in other embodiments of the present disclosure and will not be described in detail here.

[0147] The reference information obtained through the target tool contains information related to the expected response to the target question. Based on the reference information, accurate and objective target evaluation information about the response can be obtained. Therefore, adjusting the model parameters of the reward model according to the target evaluation information can provide a clear and objective training and learning direction for the reward model, enabling the reward model to accurately and objectively evaluate the response.

[0148] Step 404 : Adjusting model parameters of the reward model based on at least the difference between the target evaluation information and the expected evaluation information.

[0149] Optionally, a third loss function can be constructed based on the difference between the target evaluation information and the expected evaluation information. The difference between the target evaluation information and the expected evaluation information is positively correlated with the loss value of the third loss function. The smaller the difference between the target evaluation information and the expected evaluation information, the smaller the loss value of the third loss function. The larger the difference between the target evaluation information and the expected evaluation information, the larger the loss value of the third loss function.

[0150] Optionally, the training sample also includes calling information of the desired tool corresponding to the target problem, and the calling information of the desired tool includes at least one of information indicating whether a tool call is required, identification information of the desired tool, and input parameter information of the desired tool.

[0151] Optionally, the information indicating whether a tool call is required may be “a tool call is required” or “a tool call is not required”.

[0152] Optionally, the identification information may be information such as a tool name, a tool calling interface name, etc.

[0153] Optionally, the input parameter information may be information such as attributes and formats of input parameters required by the tool. Taking a weather forecast tool as an example, its input parameter information may be date and region.

[0154] Adjust the model parameters of the reward model based on at least the difference between the target evaluation information and the expected evaluation information, including:

[0155] Acquiring call information of a target tool generated by the reward model, where the call information of the target tool includes at least one of information indicating whether a tool call is required, identification information of the target tool, and input parameter information of the target tool;

[0156] The reward model is parameterized based on the difference between the target evaluation information and the expected evaluation information, and the difference between the expected tool call information and the target tool call information.

[0157] There is a correlation between the call information of the target tool and the training effect of the reward model. The smaller the difference between the call information of the target tool and the call information of the expected tool, the better the training effect of the reward model. Conversely, the better the training effect of the reward model, the smaller the difference between the call information of the target tool and the call information of the expected tool.

[0158] Therefore, when adjusting the model parameters of the reward model, introducing the difference between the call information of the expected tool and the call information of the target tool can provide another training and learning direction for the reward model, improve the training effect of the reward model, and enable the reward model to not only accurately and objectively evaluate the responses, but also achieve precise tool calls.

[0159] In the disclosed embodiment, a training sample consisting of a target question for input into a language model, a response output by the language model for the target question, and expected evaluation information for the response is obtained. The target question and response are then input into a reward model, and a target tool is called to obtain reference information for the target question. Based on the reference information, target evaluation information for the response is generated, and the model parameters of the reward model are adjusted based on the difference between the target evaluation information and the expected evaluation information. The reference information obtained by the target tool contains information related to the expected response to the target question. Based on the reference information, accurate and objective target evaluation information for the response can be obtained, providing a clear and objective training and learning direction for the reward model, enabling the reward model to accurately and objectively evaluate the response.

[0160] Figure 5 Schematic diagram of the structure of a language model training device provided according to an embodiment of the present disclosure.

[0161] like Figure 5 As shown, the language model training device includes:

[0162] A first acquisition module 51 is used to obtain a response generated by the language model based on the target question;

[0163] A tool calling module 52 is used to call a target tool according to a target problem and obtain reference information of the target problem, where the reference information is determined based on the return data of the target tool;

[0164] A first evaluation module 53 is used to generate target evaluation information for the reply based on the reference information of the target question;

[0165] The first adjustment module 54 is used to adjust the model parameters of the language model in the reinforcement learning stage according to the target evaluation information.

[0166] Optionally, the first evaluation module 53 is configured to:

[0167] Identify discrepancies between reference information and responses;

[0168] At least based on the difference between the reference information and the reply, the reply is scored to obtain target evaluation information.

[0169] Optionally, the first evaluation module 53 is configured to:

[0170] In the case of multiple calls to the target tool, the differences between the reference information and the replies obtained from each call are determined.

[0171] Optionally, the tool calling module 52 is configured to:

[0172] According to at least one intention corresponding to the target problem, calling a target tool corresponding to each intention;

[0173] Based on the return data of the target tool corresponding to each intention, reference information of the target problem is obtained.

[0174] Optionally, the tool calling module 52 is configured to:

[0175] Based on the target problem, identify whether tool calls are needed;

[0176] When it is identified that a tool call is required, the target tool is called according to the target problem to obtain reference information of the target problem.

[0177] Optionally, the tool calling module 52 is configured to:

[0178] Based on the intent corresponding to the target question, query the mapping relationship between intent and tools;

[0179] When the target tool corresponding to the intent of the target question is found, it is determined that the tool needs to be called.

[0180] Optionally, the tool calling module 52 is configured to:

[0181] Classify according to the semantics of the target question;

[0182] Based on the classification results, determine whether a tool call is needed.

[0183] Optionally, the tool calling module 52 is configured to:

[0184] Extract the input parameters required by the target tool from the target problem;

[0185] Through the calling interface corresponding to the target tool, input parameters are sent to the target tool, and return data from the target tool is received to obtain reference information of the target problem.

[0186] Optionally, the first adjustment module 54 is configured to:

[0187] Obtaining a first probability distribution obtained by the language model in the reinforcement learning phase for the target question; wherein the first probability distribution is used to indicate the probability of each word in the vocabulary for each word position in the response;

[0188] Obtaining a second probability distribution obtained by the language model in the supervised training phase for the target question; wherein the second probability distribution is used to indicate the probability of taking each word in the vocabulary for each word position in the third response, and the third response is the response of the language model in the supervised training phase to the target question;

[0189] comparing a probability distribution difference between the first probability distribution and the second probability distribution;

[0190] According to the target evaluation information and probability distribution differences, the model parameters of the language model in the reinforcement learning stage are adjusted.

[0191] It should be noted that the above explanation of the language model training method is also applicable to the language model training device of this embodiment and will not be repeated here.

[0192] Figure 6 2 is a schematic diagram of the structure of a reward model training device provided according to an embodiment of the present disclosure.

[0193] like Figure 6 As shown, the reward model training device includes:

[0194] A second acquisition module 61 is configured to acquire a training sample, wherein the training sample includes at least a target question input into the language model, a response output by the language model to the target question, and expected evaluation information of the response;

[0195] An input module 62 is configured to input the target question and the response into the reward model, wherein the reward model is configured to call the target tool according to the target question and obtain reference information of the target question, wherein the reference information is obtained based on the return data of the target tool;

[0196] The second evaluation module 63 is used to generate target evaluation information for reply based on the reference information;

[0197] The second adjustment module 64 is configured to adjust model parameters of the reward model at least according to the difference between the target evaluation information and the expected evaluation information.

[0198] Optionally, the training sample further includes call information of a desired tool corresponding to the target problem, where the call information of the desired tool includes at least one of information indicating whether a tool call is required, identification information of the desired tool, and input parameter information of the desired tool. The second adjustment module 64 is configured to:

[0199] Acquiring call information of the target tool generated by the reward model, the call information of the target tool including at least one of information indicating whether a tool call is required, identification information of the target tool, and input parameter information of the target tool;

[0200] The reward model is adjusted in terms of model parameters according to the difference between the target evaluation information and the expected evaluation information, and the difference between the calling information of the expected tool and the calling information of the target tool.

[0201] It should be noted that the above explanation of the reward model training method is also applicable to the reward model training device of this embodiment and will not be repeated here.

[0202] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0203] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0204] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 702 or a computer program loaded from a storage unit 705 into a RAM (Random Access Memory) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An I / O (Input / Output) interface 705 is also connected to the bus 704.

[0205] Various components in device 700 are connected to I / O interface 705, including: input unit 706, such as a keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 705, such as a magnetic disk, optical disk, etc.; and communication unit 709, such as a network card, modem, wireless communication transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0206] The computing unit 701 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the language model training method or the reward model training method. For example, in some embodiments, the language model training method or the reward model training method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 705. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the language model training method or the reward model training method described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the language model training method or the reward model training method in any other appropriate manner (e.g., by means of firmware).

[0207] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0208] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0209] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0210] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0211] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0212] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0213] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0214] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0215] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A language model training method, comprising: Get the response generated by the language model based on the target question; According to the target problem, calling a target tool to obtain reference information of the target problem, wherein the reference information is determined based on return data of the target tool; generating target evaluation information for the response based on the reference information of the target question; constructing a first loss function according to the target evaluation information, and adjusting model parameters of the language model in a reinforcement learning phase based on the first loss function; Among them, the target evaluation information includes a score. When the score and the difference are negatively correlated, the score and the loss value of the first loss function are negatively correlated. The higher the score, the smaller the difference, and the smaller the loss value of the first loss function; the lower the score, the greater the difference, and the greater the loss value of the first loss function.

2. The method according to claim 1, wherein Generating target evaluation information for the reply based on the reference information of the target question includes: determining a difference between the reference information and the response; The reply is scored at least based on the difference between the reference information and the reply to obtain the target evaluation information.

3. The method according to claim 2, wherein: The determining of the difference between the reference information and the response includes: In the case of multiple calls to the target tool, the difference between the reference information and the reply obtained in each call is determined.

4. The method according to claim 1, wherein The step of calling a target tool according to the target problem to obtain reference information of the target problem includes: According to at least one intention corresponding to the target problem, calling a target tool corresponding to each intention; The reference information is obtained based on the return data of the target tool corresponding to each of the intentions.

5. The method according to claim 1, wherein The step of calling a target tool according to the target problem to obtain reference information of the target problem includes: Based on the target problem, identifying whether a tool call is needed; When it is identified that a tool call is required, the target tool is called according to the target problem to obtain the reference information.

6. The method according to claim 5, wherein: The step of identifying whether a tool call is required based on the target problem includes: According to the intent corresponding to the target question, the mapping relationship between query intent and tools; When a target tool corresponding to the intent of the target question is found, it is determined that a tool call needs to be performed.

7. The method according to claim 5, wherein: The step of identifying whether a tool call is required based on the target problem includes: Classify according to the semantics of the target problem; Based on the classification results, determine whether a tool call is needed.

8. The method according to claim 1, wherein The step of calling a target tool according to the target problem to obtain reference information of the target problem includes: Extracting input parameters required by the target tool from the target problem; The reference information is obtained by sending the input parameters to the target tool through a calling interface corresponding to the target tool and receiving return data from the target tool.

9. The method according to claim 1, wherein The adjusting of model parameters of the language model in the reinforcement learning stage according to the target evaluation information includes: Obtaining a first probability distribution obtained by the language model of the reinforcement learning stage for the target question; wherein the first probability distribution is used to indicate the probability of each word in the vocabulary for each word position of the response; Obtaining a second probability distribution obtained by the language model in the supervised training phase for the target question; wherein the second probability distribution is used to indicate the probability of taking each word in the vocabulary for each word position in the third response, the third response being the response of the language model in the supervised training phase to the target question; comparing a probability distribution difference between the first probability distribution and the second probability distribution; According to the target evaluation information and the probability distribution difference, model parameters of the language model in the reinforcement learning stage are adjusted.

10. A reward model training method, the method further comprising: Acquire a training sample, wherein the training sample includes at least a target question for input into a language model, a response output by the language model to the target question, and expected evaluation information of the response; Inputting the target question and the response into a reward model, wherein the reward model is used to call a target tool according to the target question and obtain reference information of the target question, wherein the reference information is obtained based on the return data of the target tool; Generating target evaluation information of the reply according to the reference information; adjusting model parameters of the reward model based on at least a difference between the target evaluation information and the expected evaluation information; The training sample also includes call information of an expected tool corresponding to the target problem, the call information of the expected tool including at least one of information indicating whether a tool call is required, identification information of the expected tool, and input parameter information of the expected tool. Adjusting the model parameters of the reward model based at least on the difference between the target evaluation information and the expected evaluation information includes: Acquiring call information of the target tool generated by the reward model, the call information of the target tool including at least one of information indicating whether a tool call is required, identification information of the target tool, and input parameter information of the target tool; The reward model is adjusted in terms of model parameters according to the difference between the target evaluation information and the expected evaluation information, and the difference between the calling information of the expected tool and the calling information of the target tool.

11. A language model training device, comprising: A first acquisition module is used to obtain a response generated by the language model based on the target question; A tool calling module is used to call a target tool according to the target problem and obtain reference information of the target problem, wherein the reference information is determined based on the return data of the target tool; A first evaluation module, configured to generate target evaluation information for the response based on the reference information of the target question; a first adjustment module, configured to construct a first loss function according to the target evaluation information, and adjust model parameters of the language model in a reinforcement learning phase based on the first loss function; Among them, the target evaluation information includes a score. When the score and the difference are negatively correlated, the score and the loss value of the first loss function are negatively correlated. The higher the score, the smaller the difference, and the smaller the loss value of the first loss function; the lower the score, the greater the difference, and the greater the loss value of the first loss function.

12. The device according to claim 11, wherein The first evaluation module is used to: determining a difference between the reference information and the response; The reply is scored at least based on the difference between the reference information and the reply to obtain the target evaluation information.

13. The device according to claim 12, wherein The first evaluation module is used to: In the case of multiple calls to the target tool, the difference between the reference information and the reply obtained in each call is determined.

14. The device according to claim 11, wherein The tool calling module is used to: According to at least one intention corresponding to the target problem, calling a target tool corresponding to each intention; The reference information is obtained based on the return data of the target tool corresponding to each of the intentions.

15. The device according to claim 11, wherein The tool calling module is used to: Based on the target problem, identifying whether a tool call is needed; When it is identified that a tool call is required, the target tool is called according to the target problem to obtain the reference information.

16. The device according to claim 15, wherein The tool calling module is used to: According to the intent corresponding to the target question, the mapping relationship between query intent and tools; When a target tool corresponding to the intent of the target question is found, it is determined that a tool call needs to be performed.

17. The device according to claim 15, wherein The tool calling module is used to: Classify according to the semantics of the target problem; Based on the classification results, determine whether a tool call is needed.

18. The device according to claim 11, wherein The tool calling module is used to: Extracting input parameters required by the target tool from the target problem; The reference information is obtained by sending the input parameters to the target tool through a calling interface corresponding to the target tool and receiving return data from the target tool.

19. The device according to claim 11, wherein The first adjustment module is configured to: Obtaining a first probability distribution obtained by the language model of the reinforcement learning stage for the target question; wherein the first probability distribution is used to indicate the probability of each word in the vocabulary for each word position of the response; Obtaining a second probability distribution obtained by the language model in the supervised training phase for the target question; wherein the second probability distribution is used to indicate the probability of taking each word in the vocabulary for each word position in the third response, the third response being the response of the language model in the supervised training phase to the target question; comparing a probability distribution difference between the first probability distribution and the second probability distribution; According to the target evaluation information and the probability distribution difference, model parameters of the language model in the reinforcement learning stage are adjusted.

20. A reward model training device comprising: A second acquisition module is configured to acquire a training sample, wherein the training sample includes at least a target question input into a language model, a response output by the language model to the target question, and expected evaluation information of the response; an input module, configured to input the target question and the response into a reward model, wherein the reward model is configured to call a target tool according to the target question and obtain reference information of the target question, wherein the reference information is obtained based on return data of the target tool; A second evaluation module is used to generate target evaluation information of the reply based on the reference information; a second adjustment module, configured to adjust model parameters of the reward model based at least on a difference between the target evaluation information and the expected evaluation information; The training sample also includes calling information of a desired tool corresponding to the target problem, and the calling information of the desired tool includes at least one of information indicating whether a tool call is required, identification information of the desired tool, and input parameter information of the desired tool. The second adjustment module is configured to: Acquiring call information of the target tool generated by the reward model, the call information of the target tool including at least one of information indicating whether a tool call is required, identification information of the target tool, and input parameter information of the target tool; The reward model is adjusted in terms of model parameters according to the difference between the target evaluation information and the expected evaluation information, and the difference between the calling information of the expected tool and the calling information of the target tool.

21. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9 or 10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9 or 10.

23. A computer program product comprising a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 9 or 10.

Citation Information

Patent Citations

  • Automatic evaluation method for retrieval self-module in question-answer system

    CN107301226A

  • Dynamic sampling dialogue generation model training method and device, equipment and medium

    CN116719920A