Dialogue model training method and device, equipment and medium
By optimizing the rating of the answer text generated by the dialogue model of the reward model, the problem of unreasonable rating of the reward model is solved, and the training accuracy and scoring accuracy of the dialogue model are improved.
Patent Information
- Application Number
- CN202311850584.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the distribution of the predicted quality scores of different input texts is unreasonable, resulting in a low training accuracy of the dialogue model.
The pre-trained reward model is used to evaluate the answer text generated by the trained dialogue model, obtain the output probability of each phrase, and optimize it with the answer score, build the target reward model for scoring, and finally train the dialogue model.
It improves the training accuracy of the dialogue model, avoids answering questions of unreasonable distribution of text scores, and enhances the accuracy of the score.
Smart Images

Figure CN120258127A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and medium for training a dialogue model. Background Art
[0002] With the development of natural language processing technology, task-based dialogue generation technology based on ultra-large language models has emerged. This technology utilizes the natural language generation ability of large language models and combines the specific requirements of task-based dialogue to generate dialogue content that meets specific task requirements. During the learning process of task-based dialogue generation technology, it is generally guided by a reward model. Therefore, the accuracy of the reward model determines the performance ceiling of the dialogue model. In the prior art, the distribution of predicted quality scores of the reward model for different input texts often does not meet the ideal situation. In many cases, for response texts of different qualities to the same instruction text, although the quality scores given by the reward model are correctly sorted according to quality, the interval of their score distribution is often unreasonable, resulting in low accuracy of the predicted quality scores determined by the reward model. Consequently, the training accuracy of the dialogue model trained based on the reward model is low. Therefore, how to improve the training accuracy of the dialogue model has become an urgent problem to be solved. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, device, equipment and medium for training a dialogue model to solve the problem of low training accuracy of the dialogue model.
[0004] In a first aspect, an embodiment of the present invention provides a method for training a dialogue model, the training method comprising:
[0005] Using the dialogue model to be trained, answering the obtained instruction text data to obtain N answer texts, where N is an integer greater than zero;
[0006] Using the pre-trained reward model to evaluate each answer text to obtain the answer score of each answer text;
[0007] Obtaining the output probability of each phrase in each answer text generated by the dialogue model to be trained, and optimizing the answer score of each answer text using the output probability of each phrase in each answer text to obtain the optimized score result of each answer text;
[0008] Training the dialogue model to be trained according to the optimized score result of each answer text to obtain a trained dialogue model.
[0009] In a second aspect, an embodiment of the present invention provides a device for training a dialogue model, the training device comprising:
[0010] An answering module, configured to use a dialogue model to be trained to answer the obtained instruction text data, obtaining N answer texts, where N is an integer greater than zero;
[0011] A scoring module, configured to use a pre-trained reward model to evaluate each answer text, obtaining an answer score for each answer text;
[0012] An optimization module, configured to obtain the output probability of each phrase in each answer text generated by the dialogue model to be trained, and use the output probability of each phrase in each answer text to optimize the answer score of each answer text, obtaining an optimized score result for each answer text;
[0013] A training module, configured to train the dialogue model to be trained according to the optimized score result of each answer text, obtaining a trained dialogue model.
[0014] In a third aspect, an embodiment of the present invention provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the training method described in the first aspect is implemented.
[0015] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the training method described in the first aspect is implemented.
[0016] The beneficial effects of the present invention compared with the prior art are as follows:
[0017] Using a dialogue model to be trained to answer the obtained instruction text data, obtaining N answer texts, where N is an integer greater than zero, using a pre-trained reward model to evaluate each answer text, obtaining an answer score for each answer text, obtaining the output probability of each phrase in each answer text generated by the dialogue model to be trained, using the output probability of each phrase in each answer text to optimize the answer score of each answer text, obtaining an optimized score result for each answer text, and training the dialogue model to be trained according to the optimized score result of each answer text, obtaining a trained dialogue model. In this application, after the dialogue model generates a high-quality answer text, the self-scoring model of the dialogue model is used to self-evaluate the generated high-quality answer text, which can avoid the problem of unreasonable distribution intervals of the answer text scores for the same instruction text data. Using the target reward model constructed by the self-scoring model and the preset reward model for scoring can improve the accuracy of scoring. Training the dialogue model to be trained according to the scored answer text can improve the training accuracy of the dialogue model. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 FIG. is a schematic diagram of an application environment of a method for training a dialogue model provided by an embodiment of the present invention;
[0020] Figure 2 FIG.
[0018] is a schematic flowchart of a method for training a dialogue model provided by an embodiment of the present invention;
[0021] Figure 3 FIG. is a schematic structural diagram of a device for training a dialogue model provided by an embodiment of the present invention;
[0022] Figure 4 FIG.
[0019] is a schematic structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0024] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0025] It should be understood that when used in the specification and claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0026] It should also be understood that the term "and / or" as used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0027] As used in the specification of the present invention and the appended claims, the term "if" may be construed as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be construed as meaning "once determined", "in response to determining", "once [described condition or event] is detected", or "in response to detecting [described condition or event]" depending on the context.
[0028] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for differentiating descriptions and cannot be construed as indicating or implying relative importance.
[0029] The reference to "one embodiment" or "some embodiments" or the like described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0030] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results.
[0031] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0032] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or posterior, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0033] To illustrate the technical solution of the present invention, specific embodiments will be used for illustration below.
[0034] A training method for a dialogue model provided by an embodiment of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server. Among them, the client includes but is not limited to computer devices such as palm computers, desktop computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, cloud computer devices, personal digital assistants (PDAs), etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0035] See Figure 2 , which is a schematic flowchart of a training method for a dialogue model provided by an embodiment of the present invention. The above-mentioned training method for the dialogue model can be applied to the server in Figure 1 such as Figure 2 shown. The training method for the dialogue model may include the following steps.
[0036] S201: Use the dialogue model to be trained to answer the obtained instruction text data to obtain N answer texts, where N is an integer greater than zero.
[0037] In step S201, use the dialogue model to be trained to answer the obtained instruction text data to obtain answer texts. Among them, the dialogue model to be trained is a dialogue model with a certain self-evaluation ability. The dialogue model to be trained can be a large language model, that is, the dialogue model to be trained is a pre-trained model after one or more trainings. Among them, the large language model can refer to a deep neural network model with millions or billions of parameters. Such a model can process large-scale data and tasks and has achieved remarkable results in the fields of natural language processing, computer vision, speech recognition, etc., and is used to answer the input instruction text data. The answer text is the text output by the dialogue model to be trained.
[0038] In this embodiment, the dialogue model can be a large language model, which can process large-scale natural language data and has good language understanding ability. As an example, one or a combination of large language models such as ChatGPT (Chat Generative Pre-trained Transformer) and BLOOM (BigScience Large Open-science Open-access Multilingual Language Model) can be used as the pre-trained large language base model, which is not limited in this embodiment.
[0039] In this embodiment, the dialogue model to be trained can adopt a natural language processing model based on the self-attention mechanism. Optionally, the general dialogue model uses the bidirectional attention mechanism for feature processing, so that each word in the input instruction text data can pay attention to all other words. For the output answer text, the general dialogue model uses the unidirectional attention mechanism for feature processing, so that each word in the output answer text can only pay attention to the words before that word and cannot pay attention to the words after that word. By adopting the above method, the general dialogue model can better understand the semantic information of the input instruction text data. Moreover, the natural language processing model has strong modeling ability and good scalability. For example, the semantic features are obtained by extracting the features of the dialogue information through the BERT (Bidirectional Encoder Representations from Transformer) model. The BERT model is a language representation model based on the bidirectional encoder of Transformer. The MLM (masked language model) is used to pre-train the bidirectional Transformers so as to generate deep bidirectional language representations. The goal of the BERT model is to obtain the semantic features containing rich semantic information of the text by training with a large-scale unlabeled corpus. The main model structure of the BERT model is the stacking of Transformers. Transformer is the core module of the BERT model, and the attention mechanism is the most critical part of Transformer. The main role of the attention mechanism is to let the neural network focus on a part of the input, that is, to distinguish the influence of different parts of the input on the output.
[0040] The obtained instruction text data is a task or question expressed in natural language. Among them, the instruction text data can be the question text in a single round, or one of the texts in multi-round questions. For example, in the single-round question, the instruction text data "Please list five of the most popular tourist attractions in Guilin and briefly describe their historical and cultural values." has no subsequent interaction; in the multi-round question, the instruction text data "Question: Is platinum the same as gold? Answer: Platinum and gold are two different metals, but they are related. Platinum is made by adding rhodium to ordinary metals. Gold is made by adding copper to ordinary metals. Basically, platinum contains more silver than gold. Question: What about rose gold?"
[0041] When obtaining the instruction text data, it can be obtained through an open-source dataset, or through an existing large language model to generate the corresponding instruction text data, or through manual writing. This embodiment does not make a limitation.
[0042] Using the dialogue model to be trained, answer the obtained instruction text data to obtain answer texts. It should be noted that multiple answer texts are obtained, and each answer text includes one or more phrases.
[0043] In this embodiment, using the dialogue model to be trained, answer the obtained instruction text data to obtain answer texts. The dialogue model to be trained is a dialogue model with a certain self-evaluation ability, so that the dialogue model to be trained can self-evaluate the output answer texts to obtain self-evaluation results, and use the self-evaluation results as a factor in reward evaluation, thereby improving the accuracy of reward evaluation.
[0044] Optionally, before using the dialogue model to be trained to answer the obtained instruction text data to obtain answer texts, it further includes:
[0045] Using the initial dialogue model, answer the obtained instruction text data to obtain N initial answer texts;
[0046] Using the pre-trained reward model, evaluate the initial answer texts to obtain the initial answer scores of each initial answer text;
[0047] According to the initial answer scores of each initial answer text, calculate the initial loss of the initial dialogue model;
[0048] Use the initial loss to update the parameters of the initial dialogue model to obtain an updated dialogue model, and determine the updated dialogue model as the dialogue model to be trained.
[0049] In this embodiment, the reward model refers to a model used to describe and calculate the reward value of behaviors in reinforcement learning. In reinforcement learning, an agent continuously interacts with the environment and obtains a certain reward value therefrom. The reward model can describe and calculate the reward value obtained by the agent in each interaction. Based on these reward values, the agent can learn how to make better decisions to obtain a higher cumulative reward value. That is, for the reward of each output answer text, the goal of the reward model is that the scalar score corresponding to the answer text with a higher ranking should be higher than the scalar score corresponding to the answer text with a lower ranking, and the higher the better. The reward model can be constructed based on architectures such as multi-layer perceptrons and neural networks.
[0050] Using the initial dialogue model, answer the obtained instruction text data to get multiple initial answer texts. Among them, the initial dialogue model is a large language model, and one or a combination of models such as the ChatGPT large language model and the BLOOM large language model can be used as the pre-trained large language base model, which is not limited in this embodiment.
[0051] Considering the application method of the reward model, a good reward model should satisfy a Gaussian distribution with a mean of 0 for the scores of answer texts of different qualities. The scores obtained for answer texts of better quality should be positive and increase as the quality improves. Similarly, the scores for poorer-quality responses should be negative and decrease as the quality deteriorates. For a given input and its corresponding answer text, the reward model will predict a score to represent its quality.
[0052] Use the pre-trained reward model to evaluate each initial answer text to obtain an initial answer score. It should be noted that before using the pre-trained reward model to evaluate each initial answer text, the initial reward model needs to be trained. During training, use the trained dialogue model for training. The trained dialogue model can be a trained dialogue model obtained by training based on the instruction text data. For example, obtain the instruction text data, input the instruction text data into the trained dialogue model, output the answer text corresponding to the instruction text data, obtain the answer evaluation value of the answer text. The answer evaluation value can reflect the quality of the answer text. Train the initial reward model according to the answer evaluation value of the answer text and the answer text. The obtained reward model can output a reward value according to the answer text, and the reward value matches the answer evaluation value obtained for the answer text, so as to evaluate the quality of the output answer text through the reward value.
[0053] It should be noted that when training the reward model, a sample training set is obtained. Each sample in the sample training set includes an instruction text, a positive example response text for the instruction text, and a negative example response text for the instruction text. The matching degree between the positive example response text and the instruction text is pre-annotated as higher than the matching degree between the negative example response text and the instruction text. The sample training set is used to train the reward model to be trained until the target loss value corresponding to the reward model to be trained meets the preset loss condition, at which point the training ends, and the reward model to be trained at the end of the training is determined as the trained reward model.
[0054] Among them, the target loss value is a loss value determined based on a first loss value and a second loss value. The first loss value is a loss value determined based on the difference between a first predicted quality score and a second predicted quality score. The first predicted quality score is a predicted quality score determined by the reward model to be trained based on the instruction text and the positive example response text in the input sample. The second predicted quality score is a predicted quality score determined by the reward model to be trained based on the instruction text and the negative example response text in the input sample. The second loss value is a loss value determined based on the difference between the first predicted quality score and a pre-obtained first target score and the difference between the second predicted quality score and a pre-obtained second target score.
[0055] Through the above method, the target loss value corresponding to the reward model to be trained introduces the second loss value. The second loss value is a loss value determined based on the difference between the first predicted quality score and a pre-obtained first target score and the difference between the second predicted quality score and a pre-obtained second target score. It can be understood that by introducing the second loss value, the difference between the predicted quality score determined by the reward model and the target score is reduced, making the fluctuation range of the predicted quality score determined by the reward model more concentrated, reducing the probability that the predicted quality score determined by the reward model appears as an extremely large or extremely small predicted score, and making the result output by the reward model accurate. In this way, the result output by the large language model using the reward model is also accurate. Thus, the technical effect of making the result output by the reward model accurate is achieved.
[0056] Sort each initial response text from largest to smallest according to the initial response score to obtain a sorting result. The initial response text with a higher initial response score is ranked at the front, and the initial response text with a lower initial response score is ranked at the end. Calculate the initial loss of the initial dialogue model according to the sorting result. The calculation formula of the initial loss is as follows:
[0057]
[0058] Among them,
[0059]
[0060]
[0061]
[0062] is the initial response score of the k-th response text in the sorted result of N response texts, x is the instruction text data, y k is the k-th response text in the sorted result of the response texts, is the t-th phrase in the k-th response text, is the probability of answering the t-th phrase based on the instruction text data and the first (t - 1) phrases in the k-th response text. is the scaling factor calculated based on the initial response score of the k-th response text in the sorted result of the response texts, is the scaling factor calculated based on the initial response scores of the (k + 1)-th to N-th response texts in the sorted result of the response texts, where N is the number of initial response texts.
[0063] Use the initial loss to update the parameters of the initial dialogue model to obtain an updated dialogue model, and determine the updated dialogue model as the dialogue model to be trained, and determine the updated dialogue model as the dialogue model to be trained.
[0064] In this embodiment, a pre-trained reward model is used to evaluate each initial response text to obtain the initial response score of each initial response text, and the initial loss of the initial dialogue model is calculated based on the initial response score of each initial response text, so as to improve the calculation accuracy of the initial loss, thereby improving the accuracy of the updated dialogue model.
[0065] Optionally, use the dialogue model to be trained to answer the obtained instruction text data to obtain N response texts, including:
[0066] Preprocess the instruction text data to obtain preprocessed instruction text data;
[0067] Perform word segmentation on the preprocessed instruction text data to obtain a word segmentation result;
[0068] Perform encoding on the text in the word segmentation result to obtain an encoding result;
[0069] Input the encoding result into the dialogue model to be trained, and output N response texts generated by the dialogue model to be trained according to the encoding result.
[0070] In this embodiment, the instruction text data is preprocessed to obtain the preprocessed instruction text data. During the preprocessing, it includes converting the instruction text data into a unified format, filtering sensitive words in the instruction text data, removing duplicates from the instruction text data, and manually sampling and removing unreasonable data. For example, when converting the instruction text data into a unified format, the instruction text data can be converted into data with a unified encoding format. For example, common text formats include Office, PDF, webmail, XML, etc. Different instruction text data has different formats. When processing these instruction text data, all the instruction text data is converted into a unified format. It can be converted into the UTF-8 (8-bit Unicode Transformation Format) format. The UTF-8 format is a variable-length character encoding for Unicode, also known as the Universal Character Set. Since the Universal Character Set can represent any character in the Unicode standard and the first byte in its encoding is still compatible with ASCII, software that originally processed ASCII characters can continue to be used without or with only minor modifications. When filtering sensitive words in the instruction text data, according to the pre-constructed sensitive word library, the sensitive word multi-level filtering algorithm is used to filter the instruction text data according to the priority of the sensitive words. The sensitive word multi-level filtering algorithm in this embodiment adopts the WM (Wu-Manber) algorithm, which makes it efficient in multiple pattern string matching. When removing duplicates from the instruction text data, the minhash duplicate removal algorithm can be used.
[0071] The preprocessed instruction text data is subjected to word segmentation processing to obtain the word segmentation processing result. The text in the word segmentation processing result is subjected to encoding processing to obtain the encoding result. When performing word segmentation processing, the instruction text data can be input into a pre-trained word segmenter. The pre-trained word segmenter includes a preset dictionary, and the preset dictionary includes multiple pre-segmented words. The pre-segmented words in the instruction text data are obtained through the preset dictionary. Multiple word segmentation paths corresponding to the text are obtained according to the pre-segmented words in the instruction text data, and one word segmentation path is selected from the multiple word segmentation paths as the word segmentation result. The word segmentation processing result of the instruction text data is obtained according to the word segmentation result. The position encoding of the character is generated according to the first position information of the character in the word segmentation in the instruction text data and the second position information of the word segmentation where the character is located in the instruction text data.
[0072] Among them, the preset dictionary may include pre-segmented words learned through a large amount of external annotated data and expert knowledge. The external annotated data and expert knowledge may cover knowledge such as syntactic information, semantic knowledge, part of speech, entity recognition, semantic role annotation, syntactic parsing, etc., which is beneficial for the encoding method of the present application to introduce external knowledge through the segmented word information. After obtaining the pre-segmented words in the instruction text data through the preset dictionary, a word segmentation path with the shortest word segmentation or the highest probability of word segmentation can be selected from multiple word segmentation paths as the word segmentation processing result, and the word segmentation information of the instruction text data can be obtained according to the word segmentation processing result. In a specific example, the pre-trained word segmenter can be Hanlp word segmenter, Jieba word segmenter, etc. Thus, a large amount of external annotated data and expert knowledge can be introduced into the word segmentation information of the present application, which is beneficial for a large amount of external annotated data and expert knowledge to improve the accuracy of word segmentation processing, and further improves the encoding quality of the encoding result.
[0073] Encode the text in the word segmentation processing result to obtain an encoding result. Among them, when encoding, each segmented word can be numerically encoded. When performing numerical encoding, the corresponding encoding result can be determined by looking up the position of each segmented word in the preset word library. The preset word library is a word library containing all the segmented words in the instruction text. For example, the word segmentation result is: "I", "am", "training", "large", "language", "model", and the corresponding encoding results can be 43, 56, 52, 345, 9752, 46351. Among them, the position of "I" in the preset word library is 43, the position of "am" in the preset word library is 56, the position of "training" in the preset word library is 52, the position of "large" in the preset word library is 345, the position of "language" in the preset word library is 9752, and the position of "model" in the preset word library is 46351. Input the encoding result into the dialogue model to be trained, and output the response text generated by the dialogue model to be trained according to the encoding result.
[0074] In this embodiment, when performing word segmentation processing, the instruction text data can be input into the pre-trained word segmenter. The pre-trained word segmenter includes a preset dictionary, and the preset dictionary may include pre-segmented words learned through a large amount of external annotated data and expert knowledge. A large amount of external annotated data and expert knowledge can be introduced into the word segmentation information of the present application, which is beneficial for a large amount of external annotated data and expert knowledge to improve the accuracy of word segmentation processing, and further improves the encoding quality of the encoding result.
[0075] S202: Use the pre-trained reward model to evaluate each response text to obtain the response score of each response text.
[0076] In step S202, the pre-trained reward model is used to evaluate each response text, obtaining the response score of each response text. Among them, the response score of each response text is used as a factor for the final score of each response text.
[0077] In this embodiment, the pre-trained reward model is used to evaluate each response text, obtaining the response score of each response text. The reinforcement learning environment feedback is generated according to the response score of each response text. Then, the dialogue model to be trained updates the parameters of the dialogue model to be trained according to the reinforcement learning environment feedback. That is, the trained reward model is used to score the output of the dialogue model to be trained and thereby promote the iteration of the dialogue model to be trained.
[0078] In this embodiment, the trained reward model is used to evaluate each response text, obtaining the response score of each response text, so as to use the response score of each response text as a part of the final reward score, improving the training accuracy of the dialogue model to be trained.
[0079] S203: Obtain the output probability of each phrase in each response text generated by the dialogue model to be trained, and use the output probability of each phrase in each response text to optimize the response score of each response text, obtaining the optimized score result of each response text.
[0080] In step S203, when the dialogue model to be trained outputs each response text, it simultaneously outputs the output probability of each phrase in each response text. The output probability of each phrase in each response text generated by the dialogue model to be trained can be obtained, and the output probability of each phrase in each response text is used to optimize the response score of each response text. During optimization, the output probability of each phrase in each response text generated by the dialogue model to be trained is fused into the response score, obtaining the optimized score result of each response text.
[0081] In this embodiment, for each instruction text data, multiple response texts can be obtained. The output probability of each phrase in each response text generated by the dialogue model to be trained is obtained, the evaluation result of the self-evaluation of the model to be trained is determined according to the output probability, and the evaluation result of the self-evaluation is fused with the response score obtained by using the trained reward model for scoring, obtaining the optimized score result of each response text.
[0082] When optimizing the response score of each response text using the output probability of each phrase in each response text, the output probability can be used as the weight of the response score for optimization. For the dialogue model to be trained, the higher the quality of the output response text, the greater the corresponding output probability, and the higher the response score obtained using the trained reward model. Therefore, using the output probability as the weight value of the response score increases the importance of the response text with better quality and can highlight the response text with better quality.
[0083] For example, for the instruction text data, N response texts are output, and the output probabilities of the N response texts are obtained. The output probability of one response text is P, and the corresponding response score is R. The output probability is used to optimize the response score, that is, the output probability P is multiplied by the response score R to obtain the optimized score result of the response text.
[0084] In another embodiment, when optimizing the response score of each response text using the output probability of each phrase in each response text, the output probability can also be used as a weight value to be fused with the corresponding response score to obtain the fused response score, and the fused response score is used as the optimized score result of each response text. For example, when the response score of the response text is relatively high, generally the output probability of the dialogue model is also relatively high. The output probability is used as a weight value to be fused with the response score of the response text, thereby increasing the accuracy of the corresponding response text and increasing the score gap between different response texts. When optimizing the response score of each response text using the output probability of each phrase in each response text, other methods can also be used for optimization, which is not limited in this embodiment.
[0085] In this embodiment, the output probability of each phrase in each response text generated by the dialogue model to be trained is obtained, and the response score of each response text is optimized according to the output probability of each phrase in each response text, avoiding the problem that the response score of the reward model for complex questions and answers is inaccurate and improving the accuracy of the response text scoring.
[0086] Optionally, optimizing the response score of each response text using the output probability of each phrase in each response text to obtain the optimized score result of each response text includes:
[0087] According to the output probability of each phrase in each response text, calculate the probability mean of the output probabilities of all phrases in each response text;
[0088] Fuse the probability mean with the response score of each response text to obtain the fused score result, and determine the fused score result as the optimized score result of each response text.
[0089] In this embodiment, according to the output probability of each phrase in each answer text, the probability mean value of the output probabilities of all phrases in each answer text is calculated. Here, the probability mean value of the output probabilities is the mean value of the output probabilities of all phrases in the answer text. The calculation formula for the probability mean value of the output probabilities is as follows:
[0090]
[0091] where is the probability mean value of the output probability corresponding to the k-th answer text regarding the instruction text x, y k is the k-th answer text regarding the instruction text x, t is the t-th phrase in the k-th answer text, is the probability of answering the t-th phrase based on the instruction text data and the first (t - 1) phrases in the k-th answer text.
[0092] The probability mean value is fused with the answer score to obtain the fused score result, and the fused score result is determined as the optimized score result for each answer text. When fusing, the probability mean value and the answer score can be averaged to obtain an average value, and the average value is used as the optimized score result for each answer text. Here, the optimized score result for each answer text is the final optimized score result for the k-th answer text. The calculation formula for the optimized score result for each answer text is as follows:
[0093]
[0094] where S k (x, y k ) is the optimized score result for each answer text of the k-th answer text in the sorted result after sorting the answer texts, is the probability mean value of the output probability corresponding to the k-th answer text regarding the instruction text x, is the initial answer score of the k-th answer text in the sorted result after sorting the answer texts.
[0095] In this embodiment, according to the output probability of the answer text, the probability mean value of the output probability is calculated. According to the output probability of each phrase in each answer text, the probability mean value of the output probabilities of all phrases in each answer text is calculated. The probability mean value is fused with the answer score to obtain the fused score result, and the fused score result is determined as the optimized score result for each answer text. Taking the probability of the answer text output by the dialogue model as one of the scoring factors enables the dialogue model to self-evaluate the output answer text, taking into account the factors of the dialogue model itself, and improves the scoring accuracy of the answer text.
[0096] Optionally, fusing the probability mean value with the answer score of each answer text to obtain the fused score result includes:
[0097] Normalize the response scores for each response text to obtain the normalized response scores.
[0098] Perform weighted fusion of the probability mean and the normalized response scores to obtain the fused score results for each response text.
[0099] In this embodiment, since the probability values output by the dialogue model range from 0 to 1, therefore, normalize the response scores to obtain the normalized response scores, so that the values of the normalized response scores range from 0 to 1, which is convenient for performing weighted fusion of the probability mean and the normalized response scores to obtain the fused score results.
[0100] It should be noted that when performing weighted fusion of the probability mean and the normalized response scores, different weight values are set for the probability mean and the normalized response scores respectively. When setting the weight values, dynamic setting can be performed, that is, based on the response scores of the response text, different weight values are set for the probability mean and the normalized response scores respectively. In the early stage of training the dialogue model, the self-evaluation effect of the dialogue model is poor. The response scores obtained by the trained reward model can be used as the main scores. A higher weight can be set for the normalized response scores and a smaller weight can be set for the probability mean, that is, the weight of the normalized response scores is greater than one-half, and the weight of the probability mean is less than one-half, and the sum of the two is 1. As the training of the dialogue model progresses, the self-evaluation effect of the dialogue model gets better and better, and the weight of the probability mean can be gradually increased to one-half. Specific embodiments can be set according to the actual training situation and are not limited in this embodiment.
[0101] S204: Train the dialogue model to be trained according to the optimized score results of each response text to obtain the trained dialogue model.
[0102] In step S204, train the dialogue model to be trained according to the optimized score results of each response text to obtain the trained dialogue model. During training, perform unsupervised training according to the optimized score results of each response text to obtain the trained dialogue model.
[0103] In this embodiment, calculate the corresponding loss according to the optimized score results of each response text, and train the dialogue model to be trained to obtain the trained dialogue model.
[0104] It should be noted that when training the dialogue model to be trained, it is possible to determine whether to stop training according to the preset number of iterations. When the number of iterations reaches the preset number of iterations, training stops. When the number of iterations does not reach the preset number of iterations, training continues. It is also possible to determine whether to stop training according to the convergence degree of the loss value. When the loss value converges, training stops. When the loss value does not converge, training continues.
[0105] In this embodiment, the dialogue model to be trained is trained according to the optimized scoring results of each response text, and a trained dialogue model is obtained. The optimized scoring results of each response text have a high scoring accuracy, so that a loss value with higher accuracy can be calculated. When training based on the loss value, the training accuracy of the dialogue model is improved.
[0106] After obtaining the trained dialogue model, the trained dialogue model can be used to answer the text to be answered, and a response text with higher answer quality can be obtained.
[0107] Optionally, training the dialogue model to be trained according to the optimized scoring results of each response text to obtain a trained dialogue model includes:
[0108] Calculating the optimized loss of the dialogue model to be trained according to the optimized scoring results of each response text;
[0109] Training the dialogue model to be trained according to the optimized loss to obtain a trained dialogue model.
[0110] In this embodiment, the optimized loss of the dialogue model to be trained is calculated according to the optimized scoring results of each response text. Among them, the optimized loss includes the self-evaluation loss of the dialogue model. During the training process, both the output accuracy of the dialogue model and the self-evaluation accuracy of the dialogue model are improved. The improvement of the self-evaluation accuracy can also accelerate the training of the dialogue model and improve the training efficiency of the dialogue model.
[0111] Training the dialogue model to be trained according to the optimized loss to obtain a trained dialogue model. When training, it is possible to determine whether to stop training according to the preset number of iterations. When the number of iterations reaches the preset number of iterations, training stops. When the number of iterations does not reach the preset number of iterations, training continues. It is also possible to determine whether to stop training according to the convergence degree of the loss value. When the loss value converges, training stops. When the loss value does not converge, training continues.
[0112] Optionally, calculating the optimized loss of the dialogue model to be trained according to the optimized scoring results of each response text includes:
[0113] Sort the N response texts according to the optimized scoring results of each response text to obtain the sorting results of the N response texts;
[0114] According to the sorting results, calculate the optimized loss of the dialogue model to be trained.
[0115] In this embodiment, before calculating the optimized loss of the dialogue model to be trained, re - sort the N response texts according to the optimized scoring results of each response text, place the response text with a larger optimized scoring result at the front, and place the response text with a smaller optimized scoring result at the end. According to the sorting results, calculate the optimized loss of the dialogue model to be trained. The calculation formula of the optimized loss is as follows:
[0116]
[0117] Among them,
[0118]
[0119]
[0120] S k (x,y k ) is the initial response score of the k - th response text in the sorting results after sorting the N response texts based on the optimized scoring results of the N response texts. x is the instruction text data, and y k is the k - th response text in the sorting results after sorting based on the optimized scoring results of each response text, is the t - th phrase in the k - th response text, is the probability of answering the t - th phrase based on the instruction text x and the first (t - 1) phrases in the k - th response text. is the scaling factor calculated from the optimized scoring result of the k - th response text in the sorting results after sorting the N response texts based on the optimized scoring results of each response text, is the scaling factor calculated from the optimized scoring results of the (k + 1) - th to the N - th response texts in the sorting results after sorting the N response texts based on the optimized scoring results of each response text. N is the number of response texts.
[0121] Using the dialogue model to be trained, answer the obtained instruction text data to obtain N answer texts, where N is an integer greater than zero. Use the pre-trained reward model to evaluate each answer text to obtain the answer score of each answer text. Obtain the output probability of each phrase in each answer text generated by the dialogue model to be trained, and use the output probability of each phrase in each answer text to optimize the answer score of each answer text to obtain the optimized score result of each answer text. According to the optimized score result of each answer text, train the dialogue model to be trained to obtain a trained dialogue model. In this application, after the dialogue model generates a high-quality answer text, use the self-evaluation model of the dialogue model to self-evaluate the generated high-quality answer text, which can avoid the problem of unreasonable interval of the answer text score distribution for the same instruction text data. Using the target reward model constructed by the self-evaluation model and the preset reward model for scoring can improve the accuracy of scoring. Train the dialogue model to be trained according to the scored answer text, thereby improving the training accuracy of the dialogue model.
[0122] Please refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of a training device for a dialogue model provided by an embodiment of the present invention. In this embodiment, each unit included in the terminal is used to execute Figure 2 the corresponding steps in the corresponding embodiment. Specifically, please refer to Figure 2 the relevant descriptions in the corresponding embodiment. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 3 , the training device 30 includes: an answering module 31, a scoring module 32, an optimization module 33, and a training module 34.
[0123] The answering module 31 is used to use the dialogue model to be trained to answer the obtained instruction text data to obtain N answer texts, where N is an integer greater than zero;
[0124] The scoring module 32 is used to use the pre-trained reward model to evaluate each answer text to obtain the answer score of each answer text;
[0125] The optimization module 33 is used to obtain the output probability of each phrase in each answer text generated by the dialogue model to be trained, and use the output probability of each phrase in each answer text to optimize the answer score of each answer text to obtain the optimized score result of each answer text;
[0126] The training module 34 is used to train the dialogue model to be trained according to the optimized score result of each answer text to obtain a trained dialogue model.
[0127] Optionally, the above training device 30 further includes:
[0128] A obtaining module, configured to use an initial dialogue model to answer the obtained instruction text data, and obtain N initial answer texts.
[0129] An initial scoring module, configured to use a pre-trained reward model to evaluate each initial answer text, and obtain an initial answer score for each initial answer text.
[0130] A calculation module, configured to calculate an initial loss of the initial dialogue model according to the initial answer scores of each initial answer text.
[0131] An update module, configured to update the parameters of the initial dialogue model using the initial loss, obtain an updated dialogue model, and determine the updated dialogue model as the dialogue model to be trained.
[0132] Optionally, the above-mentioned answer module 31 includes:
[0133] A preprocessing unit, configured to preprocess the instruction text to obtain a preprocessed instruction text;
[0134] A word segmentation unit, configured to perform word segmentation on the preprocessed instruction text to obtain a word segmentation result.
[0135] An encoding unit, configured to encode the text in the word segmentation result to obtain an encoding result.
[0136] An output unit, configured to input the encoding result into the dialogue model to be trained, and output N answer texts generated by the dialogue model to be trained according to the encoding result.
[0137] Optionally, the above-mentioned optimization module 33 includes:
[0138] A calculation unit, configured to calculate a probability mean value of the output probabilities of all phrases in each answer text according to the output probability of each phrase in each answer text.
[0139] A fusion unit, configured to fuse the probability mean value with the answer score of each answer text to obtain a fused score result, and determine the fused score result as the optimized score result of each answer text.
[0140] Optionally, the above-mentioned optimization module 33 includes:
[0141] A normalization subunit, configured to perform normalization processing on the answer score of each answer text to obtain a normalized answer score.
[0142] A weighting subunit, configured to perform weighted fusion on the probability mean value and the normalized answer score to obtain a fused score result of each answer text.
[0143] Optionally, the above training module 34 includes:
[0144] An optimization loss calculation unit, configured to calculate the optimization loss of the dialogue model to be trained according to the optimized scoring results of each response text.
[0145] A training unit, configured to train the dialogue model to be trained according to the optimization loss to obtain a trained dialogue model.
[0146] Optionally, the above optimization loss calculation unit includes:
[0147] A sorting subunit, configured to sort the response texts according to the optimized scoring results of each response text to obtain a sorting result of the response texts.
[0148] A calculation subunit, configured to sort the N response texts according to the optimized scoring results of each response text to obtain a sorting result of the N response texts.
[0149] It should be noted that for the information interaction, execution process, etc. between the above modules, units, and subunits, since they are based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0150] Figure 4 This is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 4 shown, the computer device of this embodiment includes: at least one processor ( Figure 4 only one is shown in ), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps in any of the above method embodiments for training the dialogue model are implemented.
[0151] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 4 merely examples of computer devices are provided, which do not constitute a limitation on computer devices. A computer device may include more or fewer components than those shown in the figure, or combine certain components, or different components.
[0152] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0153] The memory includes a readable storage medium, internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders, data, and other programs, etc., and the other programs such as the program code of the computer program. The memory may also be used to temporarily store the data that has been output or will be output.
[0154] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working process of the units and modules in the above device can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiment can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0155] All or part of the processes in the above method embodiments of the present invention can also be completed by a computer program product. When the computer program product runs on a computer device, it enables the computer device to execute and implement the steps in the above method embodiments.
[0156] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0157] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0158] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0159] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0160] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A training method for a dialogue model, characterized in that, The training method includes: Using the dialogue model to be trained to answer the obtained instruction text data, obtaining N answer texts, where N is an integer greater than zero; Using the pre-trained reward model to evaluate each answer text, obtaining the answer score of each answer text; Obtaining the output probability of each phrase in each answer text generated by the dialogue model to be trained, and using the output probability of each phrase in each answer text to optimize the answer score of each answer text, obtaining the optimized score result of each answer text; Training the dialogue model to be trained according to the optimized score result of each answer text, obtaining a trained dialogue model.
2. The training method according to claim 1, wherein Before using the dialogue model to be trained to answer the obtained instruction text data and obtaining answer texts, it further includes: Using the initial dialogue model to answer the obtained instruction text data, obtaining N initial answer texts; Using the pre-trained reward model to evaluate each initial answer text, obtaining the initial answer score of each initial answer text; Calculating the initial loss of the initial dialogue model according to the initial answer score of each initial answer text; Updating the parameters of the initial dialogue model using the initial loss, obtaining an updated dialogue model, and determining the updated dialogue model as the dialogue model to be trained.
3. The training method according to claim 1, characterized in that Using the dialogue model to be trained to answer the obtained instruction text data and obtaining N answer texts includes: Preprocessing the instruction text data to obtain preprocessed instruction text data; Performing word segmentation on the preprocessed instruction text data to obtain a word segmentation result; Performing encoding on the text in the word segmentation result to obtain an encoding result; Inputting the encoding result into the dialogue model to be trained, and outputting N answer texts generated by the dialogue model to be trained according to the encoding result.
4. The training method according to claim 1, wherein Using the output probability of each phrase in each answer text to optimize the answer score of each answer text and obtaining the optimized score result of each answer text includes: Calculating the probability mean of the output probabilities of all phrases in each answer text according to the output probability of each phrase in each answer text; Fusing the probability mean with the answer score of each answer text to obtain a fused score result, and determining the fused score result as the optimized score result of each answer text.
5. The training method according to claim 4, characterized in that Fusing the probability mean with the answer score of each answer text to obtain a fused score result includes: Performing normalization processing on the answer score of each answer text to obtain a normalized answer score; Performing weighted fusion on the probability mean and the normalized answer score to obtain the fused score result of each answer text.
6. The training method according to claim 1, wherein, Training the dialogue model to be trained according to the optimized score result of each answer text and obtaining a trained dialogue model includes: Calculating the optimized loss of the dialogue model to be trained according to the optimized score result of each answer text; Train the to-be-trained dialogue model according to the optimized loss to obtain a trained dialogue model.
7. The training method according to claim 6, wherein Calculating the optimized loss of the to-be-trained dialogue model according to the optimized scoring results of each response text includes: Sorting the N response texts according to the optimized scoring results of each response text to obtain a sorting result of the N response texts; Calculating the optimized loss of the to-be-trained dialogue model according to the sorting result.
8. A training device for a dialogue model, characterized in that, The training device includes: A response module, configured to use the to-be-trained dialogue model to respond to the obtained instruction text data to obtain N response texts, where N is an integer greater than zero; A scoring module, configured to use a pre-trained reward model to evaluate each response text to obtain a response score for each response text; An optimization module, configured to obtain the output probability of each phrase in each response text generated by the to-be-trained dialogue model, and optimize the response score of each response text by using the output probability of each phrase in each response text to obtain an optimized scoring result for each response text; A training module, configured to train the to-be-trained dialogue model according to the optimized scoring results of each response text to obtain a trained dialogue model.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the training method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the training method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Enhanced feedback-based medical interactive large model training method and system
CN120809166A
Power task processing method and device, storage medium and electronic equipment
CN121660073A