Training method for dialogue model, answer information generation method, apparatus and medium
The combined training method for dialogue models using supervised fine-tuning and reinforcement learning with user satisfaction feedback maintains task accuracy and improves user intention understanding, resulting in better response generation.
Patent Information
- Application Number
- JP2024098979
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-06-30
- Filing Date
- 2024-06-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-06-19
Smart Images

Figure 0007704331000012 
Figure 0007704331000013 
Figure 0007704331000014
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to the fields of natural language processing and intelligent dialogue technology, and specifically relates to a training method for a dialogue model, a method, apparatus, electronic device, computer-readable storage medium, and computer program for generating response information realized based on the dialogue model.
Background Art
[0002] Artificial intelligence is a subject that studies how to simulate some human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) on a computer, including both hardware technologies and software technologies. The hardware technologies of artificial intelligence generally include technologies such as sensors, artificial intelligence dedicated chips, cloud computing, distributed storage, and big data processing. The artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] The task-based dialogue generation technology based on large-scale language models is currently one of the research focuses in the field of artificial intelligence. This technology can utilize the natural language generation ability of large-scale language models to link to the specific needs of task-based dialogue and generate dialogue content that meets the requirements of specific tasks.
[0004] The methods described in this section are not necessarily the methods previously assumed or adopted. Unless otherwise specified, none of the methods described in this section should be considered as prior art merely because they are included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered as those approved by the prior art.
Summary of the Invention
[0005] The present disclosure provides a method for training a dialogue model, a method, apparatus, electronic device, computer-readable storage medium, and computer program for generating response information realized based on the dialogue model.
[0006] According to one aspect of the present disclosure, there is provided a method for training a dialogue model, including obtaining a first sample data set including at least one first sample data and at least one second sample data, where each of the at least one first sample data includes a first question text and a first answer text, and each of the at least one second sample data includes a second question text; using the first sample data set to train the dialogue model by respectively inputting at least one first question text corresponding to at least one first sample data into the dialogue model to obtain at least one corresponding first answer prediction result output by the dialogue model, inputting the second question text for each of the at least one second sample data into the dialogue model to obtain a second answer prediction result output by the dialogue model, inputting the second answer prediction result into a reward model obtained by training based on at least one sample question, a plurality of answer texts corresponding to each of the at least one sample question, and labels of each of the plurality of answer texts to obtain a score of the second answer prediction result output by the reward model, performing an operation where the label indicates the user satisfaction of the corresponding answer text; determining a comprehensive loss based on at least one first answer prediction result, each first answer text in at least one first sample data, and the score corresponding to each of the at least one second sample data, and adjusting at least one parameter of the dialogue model based on the comprehensive loss, that is, executing a first training process.
[0007] According to one aspect of the present disclosure, there is provided a method for generating response information implemented based on a dialogue model, including obtaining a user's question text, and inputting the question text into the dialogue model to obtain a response text generated by the dialogue model, wherein the dialogue model is trained according to the training method of the above-mentioned dialogue model.
[0008] According to one aspect of the present disclosure, there is provided a training device for a dialogue model, including a first acquisition unit configured to acquire a first sample data set including at least one first sample data and at least one second sample data, wherein each of the at least one first sample data includes a first question text and a first response text, and each of the at least one second sample data includes a second question text; using the first sample data set, for each of the at least one first question texts corresponding to the at least one first sample data, inputting them into the dialogue model respectively to obtain at least one corresponding first response prediction result output by the dialogue model; for each of the second question texts in the at least one second sample data, inputting the second question text into the dialogue model to obtain a second response prediction result output by the dialogue model, inputting the second response prediction result into a reward model obtained by training based on at least one sample question, a plurality of response texts corresponding to each of the at least one sample questions, and labels of each of the plurality of response texts to obtain a score of the second response prediction result output by the reward model, where the label indicates the user satisfaction of the corresponding response text; performing an operation of determining a comprehensive loss based on at least one first response prediction result, each of the first response texts in the at least one first sample data, and the scores corresponding to each of the at least one second sample data, and adjusting at least one parameter of the dialogue model based on the comprehensive loss, and a first training unit configured to execute the first training process.
[0009] According to another aspect of the present invention, there is provided an answer information generation device realized based on a dialogue model, including an acquisition unit configured to acquire a user's question text, and a generation unit configured to input the question text into the dialogue model and acquire the answer text generated by the dialogue model. Here, the dialogue model is trained according to the training method of the above-mentioned dialogue model.
[0010] According to another aspect of the present disclosure, there is provided an electronic device, including at least one processor and a memory communicatively connected to the at least one processor. Here, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the above-mentioned dialogue model or the answer information generation method realized based on the dialogue model.
[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the training method of the above-mentioned dialogue model or the answer information generation method realized based on the dialogue model.
[0012] According to another aspect of the present disclosure, there is provided a computer program that, when executed by a processor, realizes the training method of the above-mentioned dialogue model or the answer information generation method realized based on the dialogue model.
[0013] According to one or more embodiments of the present disclosure, in the reinforcement learning training stage based on the artificial feedback of the dialogue model, by introducing the loss of supervised fine-tuning training, the ability to solve the dialogue tasks learned during the supervised fine-tuning training is not forgotten during the reinforcement learning stage, the fact accuracy of the dialogue model and the ability to understand user intentions are improved, and thereby the overall answer information generation effect of the dialogue model can be improved.
[0014] It should be understood that the content described in this part is not intended to identify the key points or important features of the embodiments of the present disclosure, nor is it for limiting the protection scope of the present disclosure. Other features of the present disclosure will be easily understood from the following description.
Brief Description of the Drawings
[0015] The drawings illustrate embodiments by way of example, form part of the description, and are used together with the written description to explain exemplary embodiments of the present disclosure. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to elements that are similar but not necessarily identical.
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0017] Exemplary embodiments of the present disclosure will be described below with reference to the drawings. The various details in the embodiments of the present disclosure included therein are for assisting in understanding, and they should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and brevity, descriptions of known functions and structures are omitted in the following description.
[0018] In the present disclosure, unless otherwise specified, terms such as "first" and "second" used to describe various elements are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same example of the element, and in some cases, they may refer to different examples based on the description of the context.
[0019] The terms used in the description of various examples of the present disclosure are only for the purpose of describing a specific example and are not intended to be limiting. Unless otherwise clearly indicated in the context, if the number of elements is not particularly limited, an element may be one or more. Note that the term "and / or" used in the present disclosure covers any and all possible combinations of the listed items.
[0020] Exemplary embodiments of the present disclosure will be described in detail below with reference to the drawings.
[0021] FIG. 1 shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented, according to an embodiment of the present disclosure. Referring to FIG. 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.
[0022] In an embodiment of the present disclosure, the server 120 can execute one or more services or software applications that enable the execution of the training method of the above-described dialogue model or the response information generation method implemented based on the dialogue model.
[0023] In some embodiments, the server 120 can also provide other services or software applications that can include a non-virtual environment and a virtual environment. In some embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 in a software as a service (SaaS) model.
[0024] In the configuration shown in FIG. 1, server 120 may include one or more assemblies that implement the functions executed by server 120. These assemblies may include software assemblies, hardware assemblies, or combinations thereof that can be executed by one or more processors. A user operating client devices 101, 102, 103, 104, 105, and / or 106 can interact with server 120 by sequentially using one or more client applications to utilize the services provided by these assemblies. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, FIG. 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0025] The user can input interactive text using client devices 101, 102, 103, 104, 105, and / or 106. The client device can provide an interface for the user of the client device to interact with the client device. The client device can also output information to the user via the interface. Although only six client devices are illustrated in FIG. 1, as will be understood by those skilled in the art, the present disclosure can support any number of client devices.
[0026] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices such as portable handheld devices, general-purpose computers (e.g., personal computers and notebook computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, game systems, thin clients, various messaging devices, sensors, or other detection devices. These computer devices can run various types and versions of software applications and operating systems such as MICROSOFT Windows, APPLE iOS, UNIX-like (registered trademark) operating systems, Linux (registered trademark), or Linux-like (registered trademark) operating systems (e.g., GOOGLE Chrome OS), and can include various mobile operating systems such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android, etc. Portable handheld devices may include mobile phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (e.g., smart glasses) and other devices. Game systems may include various handheld game devices, Internet-enabled game devices, etc. Client devices can run, for example, Internet-related applications, communication applications (e.g., email applications), short message service (SMS) applications, and can execute various applications and use various communication protocols.
[0027] Network 110 may be any type of network known to those skilled in the art, and it can use any one of a plurality of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth (registered trademark), WIFI), and / or any combination of these and / or other networks.
[0028] Server 120 may include one or more general-purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX (registered trademark) servers, midrange servers), blade servers, mainframe computers, server clusters, or any other suitable configuration and / or combination. Server 120 may include one or more virtual machines that execute a virtual operating system, or other computing architectures related to virtualization (e.g., one or more flexible pools of virtualized logical storage devices to maintain virtual storage devices of the server). In various embodiments, server 120 can execute one or more services or software applications that provide the functions described below.
[0029] The computing unit in server 120 can execute one or more operating systems including any of the above-described operating systems and any commercial server operating systems. Server 120 can also execute any one of various additional server applications and / or middleware applications, such as an HTTP server, an FTP server, a CGI server, a JAVA (registered trademark) server, a database server, etc.
[0030] In some embodiments, server 120 may include one or more applications for analyzing and synthesizing data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0031] In some embodiments, server 120 may be a server of a distributed system or a server incorporating a blockchain. Server 120 may be a cloud server, or an intelligent cloud computing server or an intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product in a cloud computing service system, which solves the defects of high management difficulty and weak business scalability existing in conventional physical hosts and virtual private server (VPS) services.
[0032] System 100 may also include one or more databases 130. In some embodiments, these databases can be used to store data and other information. For example, one or more of the databases 130 can be used to store information such as audio files and video files. The databases 130 can be configured at various locations. For example, the database used by the server 120 may be local to the server 120, or it may be away from the server 120 and communicate with the server 120 via a network or a dedicated connection. The databases 130 can be of various types. In some embodiments, the database used by the server 120 may be a relational database. One or more of these databases can store, update, and retrieve data from the database in response to instructions.
[0033] In some embodiments, one or more of the databases 130 can be used by an application and can also store the data of the application. The databases used in the application can be of various types, such as a key-value repository, an object repository, or a general-purpose repository supported by a file system.
[0034] The system 100 in FIG. 1 can be configured and operated in various ways so as to apply the various methods and apparatuses described based on the present disclosure.
[0035] According to an embodiment of the present disclosure, as shown in FIG. 2, a method for training a dialogue model is provided, obtaining a first sample data set including at least one first sample data and at least one second sample data, each of the at least one first sample data including a first question text and a first answer text, and each of the at least one second sample data including a second question text, step S201; using the first sample data set to train a dialogue model, Step S2021 of inputting at least one first question text corresponding to at least one first sample data into the dialogue model respectively and obtaining at least one corresponding first answer prediction result output by the dialogue model; For each of the second question texts for each of the at least one second sample data, Step S2022 of inputting the second question text into the dialogue model and obtaining a second answer prediction result output by the dialogue model; Inputting the second answer prediction result into a reward model obtained by training based on at least one sample question, a plurality of answer texts corresponding to each of the at least one sample question, and labels of each of the plurality of answer texts, and obtaining the score of the second answer prediction result output by the reward model, where the label indicates the user satisfaction of the corresponding answer text; Step S2023; Step S2024 of determining a comprehensive loss based on at least one first answer prediction result, each first answer text in at least one first sample data, and the score corresponding to each of the at least one second sample data; Including Step S202 of executing a first training process including Step S2025 of adjusting at least one parameter of the dialogue model based on the comprehensive loss.
[0036] Thereby, in the reinforcement learning training stage based on the artificial feedback of the dialogue model, by introducing the loss of teacher-assisted fine-tuning training, the ability to solve the dialogue tasks learned during teacher-assisted fine-tuning training is not forgotten during the reinforcement learning stage, the factual accuracy of the dialogue model and the user's intention understanding ability are improved, and thereby the overall answer information generation effect of the dialogue model can be improved.
[0037] In some embodiments, the first sample dataset can include two types of sample data, where each first sample data includes a first question text and a corresponding first answer text, and each second sample data includes one second question text.
[0038] In some embodiments, the second question text may be the same as a certain first question text in the first sample dataset. In some embodiments, the first question text and the second question text in the first sample dataset may be different from each other.
[0039] In some embodiments, each first question text and the second question text can be input into the current dialogue model, and the corresponding first answer prediction result and second answer prediction result can be obtained respectively. Then, each second answer prediction result and the corresponding second question text are input into a pre-trained reward model to obtain the score of the second answer prediction result output by the reward model. Here, the score of the second answer prediction result can be used to indicate the user's satisfaction with the answer information.
[0040] In some embodiments, the above reward model can be obtained by training in the following manner. First, obtain one or more sample questions, and input each sample question into the current dialogue model in turn, thereby generating a plurality of answer texts for each sample question, and obtaining the label of each answer text based on manual marking. Here, the label can indicate the user satisfaction of the corresponding answer text. Then, each answer text and the corresponding sample question are input into the initial reward model to obtain the prediction result of the model. Subsequently, until the reward model converges, the loss can be calculated based on the prediction result and the corresponding label, and the model parameters can be adjusted.
[0041] In some embodiments, the above reward model may be established based on an architecture such as a multi-layer perceptron or a neural network.
[0042] In some embodiments, based on at least one first answer prediction result, each first answer text in at least one first sample data, and the corresponding score in at least one second sample data, the calculation of the comprehensive loss can be performed. For example, when the first question text and the second question text are the same, after calculating the difference between each first answer prediction result and the corresponding first answer text, based on the corresponding score, the weight coefficient of the difference is determined, and the comprehensive loss can be calculated based on the difference multiplied by one or more weight coefficients.
[0043] In some embodiments, as shown in FIG. 3, determining the comprehensive loss based on at least one first answer prediction result, each first answer text in at least one first sample data, and the corresponding score in at least one second sample data, step S301 of determining the first loss based on each first answer text and the corresponding first answer prediction result in at least one first sample data, step S302 of determining the second loss based on at least one score corresponding to at least one second sample data, and may include step S303 of determining the comprehensive loss based on the first loss and the second loss.
[0044] Thereby, by calculating the losses of the two parts respectively and obtaining the comprehensive loss, the joint modeling of the two-stage model of the teacher-guided fine-tuning training stage and the reinforcement learning stage with artificial feedback for the dialogue task orientation is realized, promoting each other, improving the modeling ability of the model for user preferences, maintaining the understanding and satisfaction ability of the model for user instructions, and further overall enhancing the answer information generation effect of the dialogue model.
[0045] In some embodiments, the first loss can be determined based on the difference between the first answer prediction result for each in at least one first sample data and the corresponding first answer text.
[0046] In some embodiments, the above difference can be measured based on the cross - entropy error function or the mean squared error function.
[0047] In some embodiments, the second loss may be the average or expected value of at least one score corresponding to at least one second answer prediction result.
[0048] In some embodiments, the total loss can be obtained by combinations such as a mixed loss and an alternating minimization loss.
[0049] In some embodiments, determining the total loss based on the first loss and the second loss can include weighting the first loss and the second loss based on a first predetermined weight corresponding to the first loss and a second predetermined weight corresponding to the second loss, and obtaining the total loss.
[0050] Thereby, by weighting the two losses, the influence of the two training methods on the overall training of the model by the weights is controlled, and the training effect is guaranteed.
[0051] In some embodiments, in the training stage of the dialogue model, a plurality of first sample data sets can be sequentially obtained, and a plurality of rounds of model training can be performed based on the plurality of first sample data sets.
[0052] In some embodiments, as the number of training rounds increases, in a situation where the total is guaranteed to be constant, the first predetermined weight and the second predetermined weight are correspondingly adjusted to adjust the ratio of the loss function in different training rounds, thereby controlling the process of model training. For example, one variation factor is set, and as the number of training rounds increases, while gradually decreasing the second predetermined weight by the variation factor, the first predetermined weight is gradually increased, thereby improving the training efficiency of the model while ensuring the training effect of the model.
[0053] In some embodiments, as shown in FIG. 4, determining the second loss based on at least one score corresponding to at least one second sample data includes step S401 of determining the average and variance of at least one score based on at least one score; step S402 of normalizing each score of at least one score based on the average and variance to obtain an updated score; and step S403 of determining the second loss based on at least one updated score.
[0054] Thereby, by calculating the loss after normalizing the score, the score distribution in the training process can be further optimized, thereby avoiding the problem that the influence of the reinforcement learning loss caused by introducing other losses in the reinforcement learning stage on the model is weakened, improving the stability of reinforcement learning, and ensuring the effect of the reinforcement learning loss on model training under the framework of joint optimization.
[0055] In the stage of reinforcement learning, by introducing an extra loss to perform joint training, the increase in the loss in the reinforcement learning part can be suppressed, and furthermore, the role that reinforcement learning should originally play is weakened.
[0056] In some embodiments, before calculating the second loss, the average and variance can be calculated for the scores corresponding to all the second response prediction results in the first sample data set of the current round. Then, all the scores in this round are normalized based on the average and variance to obtain a normalized score, and the second loss can be calculated based on each normalized score.
[0057] In some exemplary embodiments, for the score r i the normalization operation can be represented by the following formula.
Number
Number
Number
[0058] Thereby, the scores within the round can be brought under a single dynamic standard normal distribution, thereby enhancing the stability of the reinforcement learning and ensuring that under the framework of the coalition optimization, the loss of the reinforcement learning increases normally.
[0059] In some embodiments, the second loss may be the average or expected value of the normalized scores.
[0060] In some embodiments, the method for training the dialogue model includes obtaining a pre-trained language model and a second sample data set including at least one third sample data, where each in the at least one third sample data includes a third question text and a third answer text, the pre-trained language model is obtained by training based on a predetermined amount of unlabeled sample corpora, and before training the dialogue model using the first sample data set, an initial dialogue model is obtained. Based on each third sample data in the second sample data set, until the pre-trained language model converges, the third question text corresponding to the third sample data is input into the pre-trained language model, and a third answer prediction result output by the pre-trained language model is obtained. Based on the third answer prediction result and the third answer text corresponding to the third sample data, the parameters of the pre-trained language model are adjusted to update the pre-trained language model, and the training operation for the pre-trained language model can be further included by repeating the execution.
[0061] Thereby, before performing the first training process, first, supervised fine-tuning training is performed based on the pre-trained language model to obtain an initial dialogue model. Then, further, based on the initial dialogue model, joint training of supervised fine-tuning training and reinforcement learning training based on artificial feedback is performed, thereby enabling the model to have the ability to solve dialogue tasks and further obtaining the ability to predict user preferences, thereby improving the overall performance of the model.
[0062] FIG. 5 shows a flowchart of a method for training a dialogue model according to an embodiment of the present disclosure.
[0063] According to some embodiments, as shown in FIG. 5, the training process of the dialogue model includes: step S501 of training a general pre-trained language model based on a corpus of a predetermined scale without a teacher to obtain a general pre-trained language model; step S502 of performing supervised fine-tuning training on the general pre-trained language model. Specifically, first, a second sample dataset is obtained, and the third sample data therein is applied to perform supervised fine-tuning training on the general pre-trained language model, so as to obtain an initial dialogue model that can understand the intention included in the question or command input by the user and has a relatively high-quality answering ability based on this intention; step S503 of training the initial dialogue model based on reinforcement learning to obtain a final dialogue model. Here, in this step, based on the training method described above, a plurality of batches of first sample datasets are obtained, and in each round, two different sample data in the first sample dataset are applied to jointly train the model to obtain a final dialogue model.
[0064] In some embodiments, the third sample data in the second sample dataset and the first sample data in the first sample dataset can be obtained from the same pre-prepared sample dataset, and each sample data in the sample dataset has both a sample question and corresponding answer information. Thereby, while enabling the model to better predict the user's preferences, the model's understanding ability for the questions or commands input by the user and its ability to generate high-quality answer information can be maintained.
[0065] In some embodiments, the dialogue model is obtained through at least one first training process based on an initial dialogue model. The training method of the dialogue model further includes inputting the second question text into the initial dialogue model to obtain a fourth answer prediction result output by the initial dialogue model. Here, determining the second loss based on at least one score corresponding to at least one second sample data includes determining the second loss based on at least one score, the second question text corresponding to each in at least one second sample data, the second answer prediction result, and the fourth answer prediction result.
[0066] Thereby, in the second loss calculation process, by introducing a normalization term calculated based on the second question text, the second answer prediction result, and the fourth answer prediction result, the stability of model training can be further improved.
[0067] In some embodiments, the normalization term can be determined based on the difference between the second question text of each second sample data and the second answer prediction result and the fourth answer prediction result. Then, the second loss is determined based on the normalization term and the score (or normalized score) corresponding to the second answer prediction result.
[0068] In some embodiments, the above normalization term can be calculated based on KL divergence.
[0069] In some embodiments, the total loss can be represented by the following formula.
Number
Number
Number
Number
Number
Number
Number
Number
[0070] In some embodiments, both the first number of at least one first sample data and the second number of at least one second sample data are plural, and the first number and the second number conform to a predetermined ratio.
[0071] Thereby, by controlling the ratio of the two types of sample data, in the training process, the degree to which the two training methods in the joint training affect the model can be controlled, and the training effect of the model can be optimized as a whole.
[0072] In some embodiments, the predetermined ratio between the first number and the second number may be, for example, 1:7.
[0073] In some embodiments, for a plurality of first sample data sets in a plurality of rounds, the ratio of the first number to the second number can be increased as the number of training rounds increases. For example, the ratio occupied by the first number can be gradually increased. Thereby, while ensuring the training effect of the model, the training efficiency of the model can be improved.
[0074] In some exemplary embodiments, the dialogue model can be established based on, for example, a knowledge-expanded large language model for dialogue (such as ERNIE bot, etc.).
[0075] In some embodiments, as shown in FIG. 6, a method for generating answer information realized based on a dialogue model is provided, including step S601 of obtaining a user's question text, and step S602 of inputting the question text into the dialogue model and obtaining the answer text generated by the dialogue model. Here, the dialogue model is trained according to the training method of the above-mentioned dialogue model.
[0076] Thereby, by using the dialogue model trained by the above training method, it is possible to have better accuracy in fact understanding and the ability to understand user intentions, and thereby generate answer information that better meets the user's expectations.
[0077] In some embodiments, as shown in FIG. 7, a training device 700 for a dialogue model is provided. A first acquisition unit 710 configured to acquire a first sample data set including at least one first sample data and at least one second sample data, wherein each of the at least one first sample data includes a first question text and a first answer text, and each of the at least one second sample data includes a second question text. Using the first sample data set to train the dialogue model. Input at least one first question text corresponding to at least one first sample data into the dialogue model respectively, and obtain the corresponding at least one first answer prediction result output by the dialogue model. For each second question text for each of at least one second sample data, input the second question text into the dialogue model, and obtain the second answer prediction result output by the dialogue model. Input the second answer prediction result into the reward model obtained by training based on at least one sample question, a plurality of answer texts corresponding to each of the at least one sample question, and labels of each of the plurality of answer texts, and obtain the score of the second answer prediction result output by the reward model, and perform an operation that the label indicates the user satisfaction of the corresponding answer text. Based on at least one first answer prediction result, each first answer text in at least one first sample data, and each score corresponding to each of at least one second sample data, determine a comprehensive loss. Including a first training unit 720 configured to execute a first training process of adjusting at least one parameter of the dialogue model based on the comprehensive loss.
[0078] In some embodiments, determining the comprehensive loss based on at least one first answer prediction result, each first answer text in at least one first sample data, and each score corresponding to each of at least one second sample data can include: determining a first loss based on each first answer text and the corresponding first answer prediction result in at least one first sample data; determining a second loss based on at least one score corresponding to at least one second sample data; and determining the comprehensive loss based on the first loss and the second loss.
[0079] In some embodiments, determining the second loss based on at least one score corresponding to at least one second sample data may include determining the average and variance of at least one score based on at least one score, normalizing the score based on the average and variance for each score in the at least one score to obtain an updated score, and determining the second loss based on at least one updated score.
[0080] In some embodiments, determining the overall loss based on the first loss and the second loss may include weighting the first loss and the second loss based on a first predetermined weight corresponding to the first loss and a second predetermined weight corresponding to the second loss to obtain the overall loss.
[0081] In some embodiments, the training device may further include a second acquisition unit configured to acquire a pre-trained language model and a second sample data set including at least one third sample data. Each of the at least one third sample data includes a third question text and a third answer text. The pre-trained language model is trained based on a predetermined number of unlabeled sample corpora. Before training the dialogue model using the first sample data set, based on each third sample data in the second sample data set, until the pre-trained language model converges, input the third question text corresponding to the third sample data into the pre-trained language model, obtain the third answer prediction result output by the pre-trained language model, and based on the third answer prediction result and the third answer text corresponding to the third sample data, adjust the parameters of the pre-trained language model to update the pre-trained language model, and further include a second training unit configured to repeatedly execute the training operation on the pre-trained language model.
[0082] In some embodiments, the dialogue model is obtained through at least one first training process based on an initial dialogue model. The training device further includes a third acquisition unit configured to input the second question text into the initial dialogue model and acquire a fourth answer prediction result output by the initial dialogue model. Here, determining the second loss based on at least one score corresponding to at least one second sample data includes determining the second loss based on at least one score, the second question text corresponding to each in at least one second sample data, the second answer prediction result, and the fourth answer prediction result.
[0083] In some embodiments, the first number of at least one first sample data and the second number of at least one second sample data are each plural, and the first number and the second number conform to a predetermined ratio.
[0084] In some embodiments, as shown in FIG. 8, an answer information generation device 800 realized based on the dialogue model is further provided, including an acquisition unit 810 configured to acquire a user's question text, and a generation unit 820 configured to input the question text into the dialogue model and acquire an answer text generated by the dialogue model. Here, the dialogue model is trained according to the training method of the above-mentioned dialogue model.
[0085] According to the embodiments of the present disclosure, an electronic device, a readable storage medium, and a computer program are further provided.
[0086] Referring to FIG. 9, a block diagram of an electronic device 900 that functions as a server or a client of the present disclosure will be described. This is an example of a hardware device applicable to each aspect of the present disclosure. The electronic device may represent various forms of digital electronic computers, such as laptop computers, desktop computers, tablets, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may further represent various forms of mobile devices, such as personal digital processors, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown in this specification, their connection relationships, and their functions are merely exemplary and do not limit the implementation of the present disclosure described and / or claimed in this specification.
[0087] As shown in FIG. 9, the electronic device 900 includes a computing unit 901 that can execute various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Also, various programs and data necessary for the operation of the electronic device 900 may be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0088] In the electronic device 900, a plurality of components including an input unit 906, an output unit 907, a memory unit 908, and a communication unit 909 are connected to the I / O interface 905. The input unit 906 may be any type of device capable of inputting information into the electronic device 900. The input unit 906 may receive input numerical or character information, generate key signal inputs related to user settings and / or function controls of the electronic device, and include, but is not limited to, a mouse, a keyboard, a touch screen, a track board, a track ball, an operation lever, a microphone, and / or a remote control. The output unit 907 may be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The memory unit 908 may include, but is not limited to, a magnetic disk and an optical disk. The communication unit 909 enables the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset (e.g., a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like).
[0089] The computing unit 901 may be various general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes various methods and processes described above, such as the training method of the above-mentioned dialogue model or the answer information generation method realized based on the dialogue model. For example, in some embodiments, the training method of the above-mentioned dialogue model or the answer information generation method realized based on the dialogue model can be realized as a computer software program tangibly embodied in a machine-readable medium such as the storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed into the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, the training method of the above-mentioned dialogue model or the answer information generation method realized based on the dialogue model can be executed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the training method of the above-mentioned dialogue model or the answer information generation method realized based on the dialogue model in any other suitable manner (e.g., by firmware).
[0090] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be implemented in one or more computer programs, which may be executed and / or interpreted in a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, that receives data and instructions from a memory system, at least one input device, and at least one output device, and transmits the data and instructions to the memory system, the at least one input device, and the at least one output device.
[0091] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to the processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, it implements the functions / operations specified in the flowchart and / or block diagram. The program code may be executed entirely by a machine, partially by a machine, partially by a machine as an independent software package and partially by a remote machine, or entirely by a remote machine or server.
[0092] In the context of the present disclosure, a machine-readable medium may be a tangible medium and may comprise or store a program for use in or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections through one or more leads, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0093] To provide for interaction with a user, a computer may implement the systems and techniques described herein, the computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball), whereby the user may provide input to the computer through the keyboard and the pointing device. Other types of devices may be further provided for interacting with the user, for example, any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback) may be provided to the user, and input from the user may be received in any form (including voice input, speech input, or tactile input).
[0094] The systems and techniques described herein may be implemented in a computing system that includes a back-end member (e.g., as a data server), a computing system that includes a middleware member (e.g., an application server), a computing system that includes a front-end member (e.g., a user computer having a graphical user interface and a web browser, through which a user can realize an interaction with embodiments of those systems and techniques), or a computing system consisting of any combination of those back-end members, middleware members, or front-end members. The members of the system may be interconnected by digital data communication in any form or medium (e.g., a communication network). An example of a communication network includes a local area network (LAN), a wide area network (WAN), and the Internet.
[0095] A computer system may include a client side and a server. The client side and the server are generally far apart from each other and usually interact via a communication network. The relationship between the client side and the server is generated by operating a computer program corresponding to a computer having a client-side / server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.
[0096] It should be understood that the steps may be reordered, increased, or deleted again using the various forms of flows described above. For example, each step described in the present disclosure may be executed in parallel, sequentially, or in a different order, and the text is not limited to this as long as the technical solutions disclosed in the present disclosure can achieve the desired results.
[0097] Embodiments or examples of the present disclosure have been described with reference to the drawings. However, the above methods, systems, and apparatuses are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but it should be understood that it is limited only by the scope of the claims after authorization and their equivalent scope. Various elements of the embodiments or examples may be omitted or replaced by their equivalent elements. Note that each step may be executed in an order different from the order described in the present disclosure. Furthermore, various elements of the embodiments or examples may be combined in various ways. What is important is that with the evolution of technology, many of the elements described here can be replaced by equivalent elements that appear after the present disclosure.
Claims
1. A method for training a dialogue model, comprising: obtaining a first sample data set including at least one first sample data and at least one second sample data, wherein each of the at least one first sample data includes a first question text and a first answer text, and each of the at least one second sample data includes a second question text; training the dialogue model using the first sample data set; inputting each of the at least one first question text corresponding to the at least one first sample data into the dialogue model to obtain at least one corresponding first answer prediction result output by the dialogue model; for each second question text in the at least one second sample data, inputting the second question text into the dialogue model to obtain a second answer prediction result output by the dialogue model; inputting the second answer prediction result into a reward model obtained by training based on at least one sample question, a plurality of answer texts corresponding to each of the at least one sample question, and labels of each of the plurality of answer texts, and obtaining a score of the second answer prediction result output by the reward model, and performing an operation that the label indicates the user satisfaction of the corresponding answer text; determining a comprehensive loss based on the at least one first answer prediction result, each first answer text in the at least one first sample data, and the score corresponding to each of the at least one second sample data; executing a first training process of adjusting at least one parameter of the dialogue model based on the comprehensive loss. A method for training a dialogue model comprising the above steps.
2. Determining a comprehensive loss based on the at least one first answer prediction result, each first answer text in the at least one first sample data, and the score corresponding to each of the at least one second sample data includes: determining a first loss based on each first answer text and the corresponding first answer prediction result in the at least one first sample data; Determining a second loss based on at least one score corresponding to the at least one second sample data; Determining the comprehensive loss based on the first loss and the second loss, the method according to claim 1.
3. Determining a second loss based on at least one score corresponding to the at least one second sample data includes: Determining an average and a variance of the at least one score based on the at least one score; Normalizing each score in the at least one score based on the average and the variance to obtain an updated score; Determining the second loss based on at least one updated score, the method according to claim 2.
4. Determining the comprehensive loss based on the first loss and the second loss includes: Weighting the first loss and the second loss based on a first predetermined weight corresponding to the first loss and a second predetermined weight corresponding to the second loss to obtain the comprehensive loss, the method according to claim 3.
5. Obtaining a pre-trained language model and a second sample data set including at least one third sample data, each of the at least one third sample data including a third question text and a third answer text, and the pre-trained language model being obtained by training based on a predetermined amount of unlabeled sample corpora; Before training the dialogue model using the first sample data set, obtaining an initial dialogue model, and based on each third sample data in the second sample data set, repeating a training operation on the pre-trained language model until the pre-trained language model converges; Inputting the third question text corresponding to the third sample data into the pre-trained language model to obtain a third answer prediction result output by the pre-trained language model; Further including repeating and executing a training operation on the pre-trained language model of adjusting parameters of the pre-trained language model based on the third answer prediction result and the third answer text corresponding to the third sample data to update the pre-trained language model, the method according to claim 4.
6. The dialogue model is obtained through at least one of the first training processes based on the initial dialogue model, and the method includes: further including inputting the second question text into the initial dialogue model to obtain a fourth answer prediction result output by the initial dialogue model; determining a second loss based on at least one score corresponding to the at least one second sample data, The method according to claim 5, wherein determining the second loss includes determining the second loss based on the at least one score, the second question text corresponding to each of the at least one second sample data, the second answer prediction result, and the fourth answer prediction result.
7. The method according to claim 6, wherein the first number of the at least one first sample data and the second number of the at least one second sample data are each a plurality, and the first number and the second number conform to a predetermined ratio.
8. An answer information generation method realized based on a dialogue model, including: obtaining a user's question text; inputting the question text into the dialogue model to obtain an answer text generated by the dialogue model, wherein the dialogue model is trained according to the training method described in any one of claims 1 to 7.
9. A training device for a dialogue model, including: a first acquisition unit configured to acquire a first sample data set including at least one first sample data and at least one second sample data, wherein each of the at least one first sample data includes a first question text and a first answer text, and each of the at least one second sample data includes a second question text; training the dialogue model using the first sample data set; inputting each of the at least one first question texts corresponding to the at least one first sample data into the dialogue model to obtain at least one corresponding first answer prediction result output by the dialogue model; for each second question text for each of the at least one second sample data, Input the second question text into the dialogue model to obtain the second answer prediction result output by the dialogue model, input the second answer prediction result into an incentive model obtained by training based on at least one sample question, a plurality of answer texts corresponding to each of the at least one sample question, and labels of each of the plurality of answer texts, and obtain the score of the second answer prediction result output by the incentive model, and execute an operation that the label indicates the user satisfaction of the corresponding answer text, determine a comprehensive loss based on the at least one first answer prediction result, each first answer text in the at least one first sample data, and the scores corresponding to each of the at least one second sample data, a first training unit configured to execute a first training process of adjusting at least one parameter of the dialogue model based on the comprehensive loss, including a training device for the dialogue model.
10. Determining the comprehensive loss based on the at least one first answer prediction result, each first answer text in the at least one first sample data, and the scores corresponding to each of the at least one second sample data includes: determining a first loss based on each first answer text and the corresponding first answer prediction result in the at least one first sample data; determining a second loss based on at least one score corresponding to the at least one second sample data; and determining the comprehensive loss based on the first loss and the second loss. The device according to claim 9.
11. Determining the second loss based on at least one score corresponding to the at least one second sample data includes: determining the average and variance of the at least one score based on the at least one score; normalizing each score in the at least one score based on the average and the variance to obtain an updated score; and determining the second loss based on at least one updated score. The device according to claim 10.
12. Determining the comprehensive loss based on the first loss and the second loss includes: The apparatus according to claim 11, comprising: weighting the first loss and the second loss based on a first predetermined weight corresponding to the first loss and a second predetermined weight corresponding to the second loss, and obtaining the comprehensive loss.
13. A second acquisition unit configured to acquire a pre-trained language model and a second sample data set including at least one third sample data, each of the at least one third sample data including a third question text and a third answer text, and the pre-trained language model being trained based on a predetermined amount of unlabeled sample corpus. Before training the dialogue model using the first sample data set, based on each third sample data in the second sample data set, until the pre-trained language model converges, an initial dialogue model is obtained. Inputting the third question text corresponding to the third sample data into the pre-trained language model, and obtaining a third answer prediction result output by the pre-trained language model. The apparatus according to claim 12, further comprising: a second training unit configured to repeatedly execute a training operation on the pre-trained language model, wherein based on the third answer prediction result and the third answer text corresponding to the third sample data, the parameters of the pre-trained language model are adjusted to update the pre-trained language model.
14. The dialogue model is obtained through at least one of the first training processes based on the initial dialogue model, and the apparatus further comprises: A third acquisition unit configured to input the second question text into the initial dialogue model and obtain a fourth answer prediction result output by the initial dialogue model. Determining the second loss based on at least one score corresponding to the at least one second sample data includes: The apparatus according to claim 13, wherein determining the second loss based on the at least one score, the second question text corresponding to each of the at least one second sample data, the second answer prediction result, and the fourth answer prediction result.
15. The first number of the at least one first sample data and the second number of the at least one second sample data are each plural, and the first number and the second number conform to a predetermined ratio. The apparatus according to claim 14.
16. An answer information generation apparatus realized based on a dialogue model, An acquisition unit configured to acquire a user's question text, A generation unit configured to input the question text into the dialogue model and acquire the answer text generated by the dialogue model. The dialogue model is trained according to the training method described in any one of claims 1 to 7. An answer information generation apparatus realized based on a dialogue model.
17. An electronic device, At least one processor, A memory communicatively connected to the at least one processor, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of claims 1 to 7. An electronic device.
18. A non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method described in any one of claims 1 to 7. A non-transitory computer-readable storage medium.
19. A computer program that, when executed by a processor, executes the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Question and answer model training method, question and answer model training device, question and answer method, and question and answer device
CN113961686A
Answer selection device, model learning device, method for selecting answer, method for learning model, and program
JP2019192073A
Language model pre-training method, apparatus, device, and storage medium
JP2023012493A