Dialogue response generation model training method and device and dialogue response generation method
Through the training method of the dialogue response generation model, combined with the quadratic related document selection strategy, the problem of insufficient reply quality in the existing technology is solved, and high-quality and highly logical dialogue response generation is achieved.
Patent Information
- Application Number
- CN202210361738.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-04-07
AI Technical Summary
The existing methods of directly generating replies cannot meet the requirements of high-quality replies, the generalization performance of the model is insufficient, and the generated replies lack logic and flexibility.
A training method for dialogue response generation model is adopted. By obtaining preset sample data and document library, the dialogue response generation model is used to traverse the dialogue process data, and response information is generated based on dialogue history data and document library, and model parameters are optimized through loss functions. In the document selection stage, a secondary correlation method is used to maintain the status based on the current topic, and select the most relevant documents from the document library.
The quality of dialogue response is improved, so that the response information generated by the model can not only correctly reflect document knowledge, but also meet context-related requirements, and enhance the logic and flexibility of dialogue responses.
Smart Images

Figure CN114706955B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to human-computer dialogue technology, and in particular to a training method and device for a dialogue response generation model and a dialogue response generation method. Background Art
[0002] Human-computer dialogue systems are an important research direction in the field of natural language processing. With the development of deep neural network technology, systems such as Alexa and XiaoIce can already have fluent conversations with people in some fields. In order to make the dialogue system adaptable to more fields, inspired by human-to-human dialogue, researchers began to explore adding external knowledge to the dialogue system. Among these external knowledge, documents are easy to obtain and contain rich information, and document-driven dialogue tasks have been proposed. In order to make the dialogue closer to human-to-human dialogue in actual scenarios, high-quality responses need to meet the requirements of fluent language, context relevance, and correct reflection of document knowledge.
[0003] There are many types of document-driven conversations. Based on the user's needs for responses, they can be divided into knowledge-oriented responses and business-oriented responses. Based on the way documents are obtained, they can be divided into two forms: fixed single document driven and multi-document driven with free document selection.
[0004] Multi-document driven requires the dialogue system to find the document that should be relied on from the document library before each round of output response, which is closer to the real-life human dialogue; knowledge-oriented response gives the dialogue system the ability to explain or communicate knowledge to users, which has a high application prospect. Therefore, multi-document driven dialogue response generation for knowledge response is the focus of current research. For this type of problem, there are currently two main methods: retrieving preset responses and directly generating responses.
[0005] The method of directly generating responses is currently the focus of academic research. It uses related text generation technologies to directly generate responses. This type of method usually includes two stages. In the first stage, several encoders are used to encode the conversation context information and candidate documents, and the correlation between the two is calculated with the help of a neural network and the retrieval results are output, thereby retrieving the target document from the candidate documents. The second stage is based on the "encoder-decoder" architecture. Several encoders are used to encode the conversation context information and the documents selected in the first stage. After the two types of encoded information are effectively integrated, they are passed to the decoder to generate a response.
[0006] In the process of implementing the present invention, the inventor found that the existing method of directly generating responses cannot meet the requirements of high-quality responses. In view of this problem, the inventor found the following reasons after research and analysis:
[0007] Existing methods for directly generating responses completely rely on supervisory signals at the decoder end, and the distribution of responses generated is highly correlated with the training corpus used to train the model. These training corpuses are often only a small subset of real data in actual applications, so the generalization performance of the model is insufficient, and it tends to output generic responses such as "OK." and "Yes." These generic responses can neither correctly reflect document knowledge nor meet context-related requirements.
[0008] In addition, the existing method of directly generating responses encodes and integrates the conversation context information and the document selected in the first stage, and then directly passes the integration result to the decoder to generate a response. This often leads to the lack of logic in the generated response, which is prone to the problem of incoherent sentences. Summary of the invention
[0009] In view of this, the main purpose of the present invention is to provide a method and device for training a dialogue response generation model and a dialogue response generation method, which can improve the quality of dialogue responses.
[0010] In order to achieve the above object, the technical solution proposed in the present invention is:
[0011] A method for training a dialogue response generation model, comprising:
[0012] Obtaining preset sample data and a document library, wherein the sample data includes conversation process data and correct document tags and topic retention tags for each round of conversation;
[0013] Using the dialogue response generation model, each round of dialogue corresponding to the dialogue process data is traversed, and based on the dialogue history data generated before the response information of this round of dialogue and the document library, the response information of this round of dialogue is generated, and based on the corresponding labels in the sample data, the loss function value is calculated, and the parameters of the dialogue response generation model are optimized and adjusted using the loss function value; wherein, when performing the generation, a quadratic correlation method is adopted, and based on the current topic retention status, the most relevant document for generating the response information is selected from the document library.
[0014] The embodiment of the present invention further provides a method for generating a dialogue response, comprising:
[0015] During the conversation process, when a conversation response needs to be generated, the conversation response generation model is used to generate and output the conversation response based on the currently generated conversation history data and a preset document library;
[0016] Wherein, the dialogue response generation model is obtained based on the training method described above.
[0017] The embodiment of the present invention also provides a training device for a dialogue response generation model, comprising a processor and a memory;
[0018] The memory stores an application program executable by the processor, which is used to enable the processor to execute the training method of the dialogue response generation model as described above.
[0019] In summary, the above technical solution proposed in the embodiment of the present invention, when using the dialogue response generation model to generate dialogue responses based on multi-document drive, for the document selection stage, a secondary correlation method is adopted to select the most relevant document for generating the response information from the document library based on the current topic retention status. In this way, the accuracy of selecting the document for generating the response information can be improved, so that the response information generated by the model can not only correctly reflect the document knowledge, but also meet the context-related requirements. In addition, the sample data not only contains real response information, but also contains correct document labels and topic retention labels. In this way, when training the model, the supervisory signal is no longer limited to the response information of the dialogue, and the model can also be trained based on document selection and topic retention status to improve the generation quality of the response information. Therefore, the embodiment of the present invention can effectively improve the reply quality of the dialogue. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A flowchart of a method for training a dialogue response generation model according to an embodiment of the present invention;
[0021] Figure 2 A schematic diagram of training a text encoder in an embodiment of the present invention;
[0022] Figure 3 A schematic diagram of determining the relevance score of conversation history data to each document in an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of determining the final probability of each candidate document being the most relevant document for the current round of dialogue in an embodiment of the present invention:
[0024] Figure 5 The figure is a flow chart of generating dialogue response information by using a response generation network according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] Figure 1 Schematic diagram of the method flow of an embodiment of the present invention, as shown in Figure 1 As shown, the training method of the dialogue response generation model implemented in this embodiment mainly includes:
[0027] Step 101: Obtain preset sample data and document library.
[0028] The sample data includes conversation process data and the correct document labels and topic retention labels for each round of conversation. In this way, when training the model, the supervision signal is no longer limited to the response information. The model can also be trained based on document selection and topic retention to improve the generation quality of response information.
[0029] Step 102: traverse each round of dialogue corresponding to the dialogue process data using the dialogue response generation model, generate response information for the dialogue round based on the dialogue history data generated before the response information for the dialogue round and the document library, calculate a loss function value based on the corresponding labels in the sample data, and optimize and adjust the parameters of the dialogue response generation model using the loss function value; wherein, when performing the generation, a quadratic correlation method is adopted to select the most relevant document for generating the response information from the document library based on the current topic retention status.
[0030] In this step, it is necessary to traverse each round of dialogue corresponding to the dialogue process data in the sample data (i.e., each round of dialogue included in the corresponding dialogue process), and the model generates the response information of this round of dialogue, so as to tune the model according to the generation of the response information, so that the model can generate high-quality response information. Different from the existing scheme, in the process of generating dialogue responses, it is necessary to adopt a secondary correlation method in the document selection link, based on the current topic retention status, to select the most relevant document for generating the response information from the document library, so that the accuracy of selecting the document for generating the response information can be improved, so that the response information generated by the model can not only correctly reflect the document knowledge, but also meet the context-related requirements.
[0031] In one implementation, for each round of dialogue in the dialogue process, the dialogue response generation model may specifically use the following method to generate response information for the round of dialogue:
[0032] Step 201: Encode the conversation history data using a pre-trained text encoder, and perform average pooling on the obtained encoded representation to obtain a vector representation of the conversation history data; obtain a vector representation of each document in the document library; wherein the vector representation of the document is obtained by encoding the documents separately using the text encoder, and performing average pooling on the obtained encoded representation.
[0033] Here, for the conversation history data corresponding to the current round of conversation and each document in the document library, it is necessary to use a pre-trained text encoder to generate corresponding vector representations, so that the most relevant documents can be selected for the current round of conversation based on these vector representations.
[0034] In one implementation, the method may be specifically Figure 2 The process described above trains the text encoder, and the specific method includes the following steps a1 to a4:
[0035] Step a1: Obtain preset coded sample data, wherein the coded sample data includes the conversation history data C and the document d related to the conversation history data C. + and unrelated documents d-.
[0036] In practical applications, we can randomly select a document from the document library except for the relevant document d + Documents other than d- are regarded as irrelevant documents.
[0037] Step a2: using the text encoder to respectively encode the conversation history data C and the related document d in the encoded sample data. + and the irrelevant document d-, and perform average pooling on each of the obtained encoded representations to obtain the vector representation C of the conversation history data C vec , the vector representation of the relevant document d+ and the vector representation of the irrelevant document d- The specific implementation is achieved using the following formula:
[0038] C vec =Average(TransformerEncoder(C));
[0039]
[0040]
[0041] TransformerEncoder() is the text encoder chosen here.
[0042] Step a3: Based on the conversation history data C and the related document d + The vector representation of , calculate the cosine similarity (i.e., correlation distribution) between the vectors, and obtain the conversation history data C and the related document d + The relevance score relavance(C,d + ); Based on the conversation history data C and the irrelevant document d - The vector representation of the conversation history data C and the irrelevant document d are obtained by calculating the cosine similarity between the vectors. - The relevance score relavance(C,d - ), which is implemented using the following formula:
[0043]
[0044]
[0045] Step a4: Based on the relevance score (i.e., relavance (C, d + )、relavance(C,d - )) Using the hinge loss function, we calculate the encoding loss function value Loss(C,d + ,d - ; m); using the encoding loss function value, the parameters of the text encoder are optimized and adjusted, and the following formula is specifically used to calculate the encoding loss function value:
[0046] Loss(C,d + ,d - ;m)=max(0,m-(relavance(C,d + )-relavance(C,d - )))
[0047] Among them, m is a hyperparameter.
[0048] Here, by using the hinge loss function to calculate the encoding loss function value, the model's ability to distinguish between relevant and irrelevant documents can be improved.
[0049] Step 202: Based on the vector representation, using a document selection network and a quadratic relevance method, the most relevant document for the current round of conversation is determined based on the current topic retention status.
[0050] This step is used to determine the most relevant document for the current round of dialogue traversed, so as to generate corresponding response information based on the most relevant document.
[0051] It should be noted here that the document selection method used in the traditional solution only uses the conversation history and the text matching method to calculate the relevance between the current conversation content and the candidate documents, and selects the most relevant document as the selection result. However, the inventor found in the process of implementing the present invention that in the same conversation, the documents selected between different rounds have a certain correlation. Taking a conversation in the preset sample data as an example, the sequence of its text selection is "Interstellar - Interstellar - Interstellar - Christopher Nolan - Christopher Nolan - Christopher Nolan". Among them, the topic transfer occurred in the first and fourth rounds, and the previous topics were maintained in the remaining rounds. Based on this, in the embodiment of the present invention, by modeling this correlation characteristic, that is, when selecting documents, a secondary correlation method is used to consider the current topic retention status (that is, whether the topic of the current round of conversation is consistent with the topic of the previous round of conversation), and determine the most relevant document of the current round of conversation, the accuracy of selecting documents for generating response information can be effectively improved.
[0052] In one implementation, in step 202, the most relevant document for the current round of conversation may be determined by specifically using the following steps 2021 to 2024:
[0053] Step 2021: Based on the vector representation, by calculating the cosine similarity between the vectors, the relevance score between the conversation history data and each of the documents is obtained.
[0054] like Figure 3 As shown, in this step, it is necessary to calculate the relevance score between the conversation history data corresponding to the current round of conversation and each document based on the vector representation of the conversation history data obtained in step 201 and the vector representation of each document in the document library, so as to select several documents most relevant to the conversation history data as candidate documents based on the relevance score in the subsequent steps.
[0055] The method for calculating the cosine similarity between vectors is known to those skilled in the art and will not be described in detail here.
[0056] Step 2022: Select the first N documents with the largest relevance scores from the document library as candidate documents; N is a preset integer greater than 1.
[0057] Specifically, in this step, according to the relevance scores obtained in step 2021, the documents in the document library are sorted in descending order of the relevance scores, and the first N documents are selected as candidate documents from the sorting results.
[0058] The N is a preset number of candidate documents, and in practical applications, those skilled in the art can set a suitable value of N according to actual needs. For example, N can be 4, but is not limited thereto.
[0059] Step 2023: Based on the most relevant document d selected when generating the response information of the previous round of dialogue last , determine the final probability that the candidate document is the most relevant document for the current round of dialogue.
[0060] Here, based on the most relevant document selected when generating the response information of the previous round of dialogue, it is possible to predict whether the current topic is consistent with the previous round of dialogue. In this way, the most relevant document of the current round of dialogue can be accurately determined in combination with the current topic retention status. Therefore, in this step, a second correlation calculation is performed based on the most relevant document selected when generating the response information of the previous round of dialogue to determine the final probability for each candidate document obtained in step 2022, so that the subsequent steps can accurately select the most relevant document for the current round of dialogue based on the final probability of each candidate document.
[0061] In one embodiment, if Figure 4 As shown, in step 2023, the following steps 20231 to 20234 may be specifically used to determine, for each candidate document obtained in step 2022, the final probability that it is the most relevant document in the current round of dialogue:
[0062] Step 20231: Based on the vector representation of the candidate document and the vector representation of the conversation history data, a dot product calculation method is used to predict the probability that the candidate document is the most relevant document in the current round of conversation, and the direct relevance probability of the candidate document is obtained. Specifically, the following formula can be used to obtain the direct relevance probability P(c, h i |relative direct ):
[0063] P(C,d i |relative direct )=softmax(c·h i )
[0064] in,
[0065] c=Average(TransformerEncoder(C));
[0066] h i =Average(TransformerEncoder(d i ));
[0067] C is the conversation history data of the current round of conversation; d i For alternative documents.
[0068] Step 20232: Based on document d lastThe vector representation of the conversation history data and the vector representation of the conversation history data are used to predict the document d last is the probability P of the most relevant document in the current round of dialogue keep Specifically, the Sigmoid function can be used to implement the following formula:
[0069] P keep =sigmoid(MLP(c;h last ))
[0070] in,
[0071] h last For document d last The vector representation of ;
[0072] h last =Average(TransformerEncoder(d last )).
[0073] Step 20233: When the document d last When the document is one of the candidate documents, the probability P is used keep , the direct relevance probability of the candidate document is corrected; and the corrected probability is normalized to obtain the final probability that the candidate document is the most relevant document in the current round of dialogue.
[0074] Here, when the document d last When it is a document in the candidate documents, it means that the topic of the current round of dialogue and the topic of the previous round of dialogue may be the same. At this time, it is necessary to combine the probability of the previous round of topic being maintained in this round to determine the final probability that the candidate document is the most relevant document in the current round of dialogue, so that the correlation between adjacent dialogues can be fully considered when selecting documents, thereby improving the accuracy of selecting dialogue-related documents.
[0075] Specifically, the following method can be used in this step, using the probability P keep , the direct relevance probability of the candidate document is modified:
[0076] If the candidate document is the document d last , then calculate the direct relevance probability P(c,h i |relative direct ) and the P keep The sum of the direct related probability of the candidate document is obtained by correcting the direct related probability of the candidate document P(C,d i |relative combined ), otherwise, calculate the direct relevance probability P(c,h i|relative direct ) and △, and obtain the result P(C,d i |relative combined ).
[0077] Wherein, the △=1-P keep .
[0078] For example, suppose there are 4 device selection documents, and document d last is one of them, then:
[0079] If the alternative document d i is the document selected in the previous round of dialogue d last , its directly related probability P(c,h i |relative direct ) is corrected to obtain the corrected probability P(C,d i |relative combined ) is:
[0080] P(C,d i |relative combined )=P(C,d i |relative direct )+P keep , i is the candidate document number;
[0081] Otherwise, the alternative document d i Not the document selected in the previous round of conversation last , its directly related probability P(c,h i |relative direct ) is corrected as follows:
[0082] P(C,d i |relative combined )=P(C,d i |relative direct )+(1-P keep ).
[0083] The above method is used to obtain the corrected probability P(C,d i |relative combined ), and then normalize according to the following formula to obtain each candidate document d i is the final probability P(C,d i |relative):
[0084]
[0085] Step 20234: When the document d last When the document does not belong to the candidate documents, the direct relevance probability of the candidate document is normalized to obtain the final probability that the candidate document is the most relevant document in the current round of dialogue.
[0086] Here, when the document d last If it is not any candidate document, it means that the topic has shifted. Therefore, there is no need to revise the direct relevance probability of the candidate documents. Instead, the direct relevance probability of the candidate documents is normalized to obtain the final probability that each candidate document is the most relevant document in the current round of dialogue.
[0087] Step 2024: Select the candidate document with the highest final probability as the most relevant document for the current round of dialogue.
[0088] Step 203: Based on the encoded representation of the most relevant document and the encoded representation of the conversation history data, using a response generation network, generate response information for the current round of conversation.
[0089] In this step, the response information for the current round of dialogue will be generated based on the response generation network in the dialogue response generation model.
[0090] Considering that the existing scheme adopts the method of cross-encoding the conversation context with the selected document and then directly generating a response, when generating a word at each time step, it is easily affected by the distribution of training data and tends to generate words that appear more frequently in the training corpus, resulting in the final generated responses being mostly low-quality general responses.
[0091] In order to overcome the impact of the above-mentioned problems on the quality of response information, in one implementation, the generation process of response information can be further optimized and divided into two parts. First, the initial response text is generated based on the words with higher frequency in the corpus, and then the key information with lower frequency (i.e., target knowledge below) is retrieved from the selected document to modify the initial response text. In this way, on the one hand, the generation of response information can be not limited to the training corpus, thereby improving the generalization of the model, and on the other hand, the semantics of the response information can be made smooth, thereby further improving the quality of dialogue responses. Taking the response information "I have seen it, it is a movie released in 2014." as an example, the initial response text is "I have seen it, it is a movie released by __knowledge__.", and the key information is "2014". Replacing the special word "__knowledge__" in the initial response text with the key information will give the response information "I have seen it, it is a movie released in 2014."
[0092] Based on the above, specifically, the following steps 2031 to 2033 may be adopted to generate response information of the current round of dialogue using the response generation network:
[0093] Step 2031: Based on the encoded representation of the most relevant document and the encoded representation of the conversation history data, a dot product attention mechanism is used to obtain a cross representation of the most relevant document and the conversation history data.
[0094] Here, the encoding of the most relevant document is d encoded and the encoding representation of the conversation history data C encoded , obtained based on the above text encoder. Specifically, the following formula can be used to obtain the cross representation CaD (ContextAwareDocument) of the two:
[0095]
[0096] Step 2032: Based on the cross representation and the preset vocabulary, a pre-trained text decoder is used to generate an initial response text for the current round of dialogue using a copy mechanism; the vocabulary is generated based on the documents and dialogue texts in the corpus used for model training, and contains special words to be replaced by target knowledge.
[0097] In one implementation, the following steps x1 to x2 may be specifically used to generate an initial response text for the current round of dialogue:
[0098] Step x1: At each time step of the text decoder, traverse each word w in the vocabulary j , determine the word t generated at the current time step i For the word w j The final probability of t is obtained, and from the word list, the word with the largest final probability is selected as the word t generated at the current time step i .
[0099] In one implementation, the following method may be used to determine the word t generated at the current time step: i For word w j The final probability is:
[0100] Based on all the words currently generated by the text decoder and the cross representation CaD, determine the word t generated at the current time step i For the word w j The initial probability P generate (t i =w j ). It can be realized by the following formula:
[0101] P generate (t i =wj )=TransformerDecoder(CaD,t0,...,t i-1 );
[0102] where t0, ..., t i-1 For all the words currently generated, TransformerDecoder() represents the text decoder selected here.
[0103] Based on the encoding representation of the most relevant document, a multi-layer perceptron is used to determine when the word t i For the word w j The word t is directly copied from the most relevant document i The probability P copy (t i =w j ). Specifically, we can use the softmax function and use the following formula to implement it:
[0104] P copy (t i =w j )=softmax(MLP(d encoded )).
[0105] Based on the P generate (t i =w j ), using a multi-layer perceptron, determine when the word t i For the word w j The word t is directly copied from the word list i The probability P(copy) can be realized by using the sigmoid function and the following formula:
[0106] P(copy)=sigmoid(MLP(TransformerDecoder(CaD,t0,...,t i-1 ))).
[0107] Based on the P generate (t i =w j ), said P copy (t i =w j ) and P(copy), according to P(t i =w j )=(1-P(copy)×P copy (t i =w j )+P(copy)×P generate (t i =wj ), determine the t i For the word w j The final probability P(t i =w j ). It can be realized by the following formula:
[0108] P(t i =w j )=(1-P(copy))*P copy (t i =w j )+P(copy)*P generate (t i =w j ).
[0109] Step x2: Sequentially obtain the words t obtained in all the time steps i Concatenate to obtain the initial text of the response.
[0110] Step 2033: If the initial response text contains the special word, the target knowledge is retrieved from the most relevant document, and the special word in the initial response text is replaced by the target knowledge to obtain the response information of the current round of dialogue; otherwise, the initial response text is used as the response information of the current round of dialogue.
[0111] After the above step 2022, the initial response text can be obtained. If the initial response text contains specified special words (such as __knowledge__), it is necessary to further retrieve the target knowledge from the most relevant documents, so as to use the retrieved target knowledge to replace the special words in the initial response text and obtain the final response information. Otherwise, the initial response text is directly used as the final reply.
[0112] In one implementation, the following steps y1 to y3 may be specifically used to retrieve target knowledge from the most relevant document:
[0113] Step y1: Encode the initial response text using the text encoder to obtain the encoded representation of the initial response text. The specific operation can be expressed by the following formula:
[0114] PR encoded =TransformerEncoder(PR).
[0115] Among them, PR (PrototypeResponse) is the initial response text.
[0116] PR = {t0, ..., t n-1}, n represents the number of words in the initial response text.
[0117] Step y2: Based on the encoded representation of the initial response text and the cross representation CaD, a dot product attention mechanism is used to obtain the cross representation RaD (ResponseAwareDocument) of the initial response text and the cross representation CaD. The specific operation can be implemented using the softmax function and the following formula:
[0118] RaD=softmax(PR encodea CaD T )CaD
[0119] It should be noted that in order to avoid the problem of incoherent sentences when the target knowledge is subsequently combined with the initial response text, an improvement in the attention interaction between the initial response text and the document is introduced here. That is, it is necessary to adopt a dot product attention mechanism to cross-represent the encoded representation of the initial response text and the cross representation CaD. In this way, compared with traditional methods, higher quality replies can be generated.
[0120] Step y3, based on the cross representation RaD, using a multilayer perceptron, for each word in the most relevant document, predict the probability that the word is the starting point of the target knowledge and the probability that the word is the end point of the target knowledge; use the word with the highest probability of the starting point as the starting point of the target knowledge, and use the word with the highest probability of the end point as the end point of the target knowledge, and based on the starting point and end point of the target knowledge, obtain the target knowledge from the most relevant document.
[0121] Specifically, in this step, after a linear layer is processed, the entire sequence is processed by softmax to obtain the probability that the position index i of each word in the most relevant document is the starting point or end point of the target knowledge. Specifically, the following formula can be used to obtain the probability P(start=i) that the position index i is the starting point of the target knowledge and the probability P(end=i) that it is the end point of the target knowledge:
[0122] P(start=i)=softmax(MLP(RaD));
[0123] P(end=i)=softmax(MLP(RaD)).
[0124] Figure 5 A schematic diagram of a preferred example process of generating response information for the current round of dialogue using the method in step 203 is given, as shown in FIG. Figure 5As shown, in this process, a copy mechanism is introduced to generate an initial response text. Afterwards, when there are special words in the initial response text, the initial response text is used to extract knowledge fragments from the document, that is, to retrieve the target knowledge. Finally, based on the combination of the initial response text and the target knowledge, the response information of the current round of dialogue is generated. In this way, the generation of response information is not limited to the training corpus, which effectively improves the generalization of the model. At the same time, it can ensure the fluency of the response information, thereby effectively improving the quality of dialogue responses.
[0125] In step 102, after generating response information for the current round of dialogue, the corresponding loss function value can be calculated in combination with the label data in the sample data to tune the network parameters of the model. Preferably, in order to improve the effect of model training, in one embodiment, the sample data can further include: the initial text label of the response information of each round of dialogue and the starting label and the end label of the target knowledge. In this way, when calculating the loss function, the initial text of the response and the target knowledge can be used as supervisory signals, so that the quality of the response information generated by the model can be effectively improved based on the diversification of the supervisory signals.
[0126] Accordingly, in step 102, the following steps z1 to z2 may be specifically used to calculate the loss function value:
[0127] Step z1: Based on the final probability, the P keep , the correct document label and topic retention label of the current round of dialogue, using the cross entropy loss function, calculate the first loss function value Loss(θ, keep, d golden ;λ); The first loss function value is used to optimize and adjust the parameters of the document selection network.
[0128] This step is implemented using the following formula:
[0129] Loss(θ, keep, d golden ;λ)=λ*CE(P keep , keep)+(1-λ)*CE(P(C,d i |relative),d golden )
[0130] Among them, CE() represents the cross entropy loss function; θ is the model parameter, keep represents the topic retention label, and d golden represents the correct document label, and λ is a hyperparameter.
[0131] Step z2: based on the word t i For the word w jThe final probability of the starting point and the end point, the initial text label of the response, and the starting point label and the end point label of the target knowledge are calculated using the cross entropy loss function to calculate the second loss function value Loss(θ, PR golden , start golden , end golden ), the second loss function value is used to optimize and adjust the parameters of the response generation network.
[0132] This step is implemented using the following formula:
[0133]
[0134] Among them, θ is the model parameter, PR golden Indicates the response initial text label, It is PR golden The word corresponding to time step i, start golden The starting label and end label of the target knowledge golden Represents the endpoint label of the target knowledge, and λ is a hyperparameter.
[0135] It can be seen from the above embodiments that the above scheme builds a dialogue response generation model based on the changing characteristics of the dialogue topic in the person-to-person dialogue, optimizes the document selection network in the model, and adopts a quadratic correlation method to select the most relevant documents for each round of dialogue based on the state of topic retention and transfer to generate response information. In this way, compared with the traditional algorithm, it can effectively improve the effect of the document selection stage in the generation of multi-document driven dialogue responses for knowledge replies. At the same time, since the model explicitly outputs the probability of topic retention and final selection, the generation of response information has a certain degree of explainability.
[0136] In addition, in view of the widespread "general reply" situation in text generation, the above embodiment further optimizes the process of generating response information based on the selected relevant documents, and splits the generation process of response information into two stages: generation of initial response text and retrieval of target knowledge. The words with higher frequency in the response information are generated by generating the initial response text, and the key information in the document that should be reflected in the response information is retrieved by retrieving the target knowledge. In addition, in view of the situation where the combination of the two pieces of information leads to incoherent sentences, the attention interaction improvement of the initial response text and the document is introduced, so that higher quality replies can be generated compared with traditional methods.
[0137] Based on the above-mentioned training method embodiment of the dialogue response generation model, accordingly, the present embodiment also provides a dialogue response generation method, which includes:
[0138] During the conversation process, when a conversation response needs to be generated, the conversation response generation model is used to generate and output the conversation response based on the currently generated conversation history data and a preset document library.
[0139] The dialogue response generation model is trained based on the embodiment of the training method for the dialogue response generation model.
[0140] As described above, by using the above training method embodiment, the dialogue response generation model can improve the quality of response information generation. Therefore, compared with the existing method, the above dialogue response generation method embodiment can effectively improve the quality of replies during the dialogue process and enhance the intelligence of machine replies.
[0141] Based on the above training method embodiment, the embodiment of the present invention also provides a training device for a dialogue response generation model, including a processor and a memory; the memory stores an application program executable by the processor, which is used to enable the processor to execute the training method for the dialogue response generation model as described above. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code for implementing the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium. In addition, the operating system or the like operating on the computer can also be made to complete part or all of the actual operations through instructions based on the program code. The program code read from the storage medium can also be written to a memory provided in an expansion board inserted into the computer or to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU or the like installed on the expansion board or the expansion unit is made to execute part or all of the actual operations, thereby realizing the functions of any of the above embodiments in the training method for the dialogue response generation model.
[0142] The memory may be implemented as various storage media such as an electrically erasable programmable read-only memory (EEPROM), a flash memory (Flash memory), and a programmable program read-only memory (PROM). The processor may be implemented as including one or more central processing units or one or more field programmable gate arrays, wherein the field programmable gate array integrates one or more central processing unit cores. Specifically, the central processing unit or the central processing unit core may be implemented as a CPU or an MCU.
[0143] It should be noted that not all steps and modules in the above processes and structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The division of each module is only for the convenience of describing the functional division adopted. In actual implementation, a module can be implemented by multiple modules, and the functions of multiple modules can also be implemented by the same module. These modules can be located in the same device or in different devices.
[0144] The hardware modules in each embodiment can be implemented mechanically or electronically. For example, a hardware module may include a specially designed permanent circuit or logic device (such as a dedicated processor, such as an FPGA or ASIC) for performing a specific operation. The hardware module may also include a programmable logic device or circuit (such as a general-purpose processor or other programmable processor) temporarily configured by software to perform a specific operation. As for whether to implement the hardware module mechanically, or using a dedicated permanent circuit, or using a temporarily configured circuit (such as configured by software), it can be decided based on cost and time considerations.
[0145] In this article, "schematic" means "serving as an example, instance or explanation", and any diagram or implementation method described as "schematic" in this article should not be interpreted as a more preferred or more advantageous technical solution. In order to make the drawings concise, only the parts related to the present invention are schematically shown in each figure, and do not represent the actual structure of the product. In addition, in order to make the drawings concise and easy to understand, in some figures, only one of the parts with the same structure or function is schematically drawn, or only one of them is marked. In this article, "one" does not mean that the number of the relevant parts of the present invention is limited to "only one", and "one" does not mean that the number of the relevant parts of the present invention is "more than one". In this article, "upper", "lower", "front", "back", "left", "right", "inside", "outside", etc. are only used to indicate the relative position relationship between the relevant parts, rather than to limit the absolute position of these relevant parts.
[0146] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for training a dialogue response generation model, characterized in that: include: Obtaining preset sample data and a document library, wherein the sample data includes conversation process data and correct document tags and topic retention tags for each round of conversation; Using the dialogue response generation model, traverse each dialogue round corresponding to the dialogue process data, generate the response information of the dialogue round based on the dialogue history data generated before the response information of the dialogue round and the document library, and calculate the loss function value based on the corresponding label in the sample data, and optimize and adjust the parameters of the dialogue response generation model by using the loss function value; wherein, when performing the generation, a quadratic correlation method is adopted to select the most relevant document for generating the response information from the document library based on the current topic retention status; The response information generated for this round of dialogue includes: Encoding the conversation history data using a pre-trained text encoder, and performing average pooling processing on the obtained encoded representation to obtain a vector representation of the conversation history data; Obtaining a vector representation of each document in the document library; wherein the vector representation of the document is obtained by encoding the document respectively using the text encoder and performing average pooling processing on the obtained encoded representation; Based on the vector representation, using a document selection network, using a quadratic relevance method, and based on the current topic retention status, determining the most relevant document for the current round of conversation; Based on the encoded representation of the most relevant document and the encoded representation of the conversation history data, using a response generation network, generating response information for the current round of conversation; The most relevant documents for determining the current round of dialogue include: Based on the vector representation, by calculating the cosine similarity between the vectors, a relevance score between the conversation history data and each of the documents is obtained; Selecting the first N documents with the largest relevance scores from the document library as candidate documents; N is a preset integer greater than 1; Based on the most relevant document d selected when generating the response information of the previous round of dialogue last Whether it belongs to the candidate document, and determining the final probability that the candidate document is the most relevant document in the current round of dialogue; The candidate document with the highest final probability is selected as the most relevant document for the current round of dialogue.
2. The method according to claim 1, characterized in that The training of the text encoder includes: Obtain preset coded sample data, the coded sample data including conversation history data C, and documents related to the conversation history data d + and unrelated documents - ; The text encoder is used to respectively encode the conversation history data C and the related document d in the encoded sample data. + and the irrelevant document d - Encode, and perform average pooling processing on each encoded representation to obtain the conversation history data C and the related document d + and the irrelevant document d - The respective vector representations; Based on the conversation history data C and the related document d + The vector representation of the conversation history data C and the related document d are obtained by calculating the cosine similarity between the vectors. + The relevance score of Based on the conversation history data C and the irrelevant document d - The vector representation of the conversation history data C and the irrelevant document d are obtained by calculating the cosine similarity between the vectors. - The relevance score of Based on the relevance score, a hinge loss function is used to calculate an encoding loss function value; and the encoding loss function value is used to optimize and adjust the parameters of the text encoder.
3. The method according to claim 2, characterized in that The final probability of determining that the candidate document is the most relevant document for the current round of dialogue includes: Based on the vector representation of the candidate document and the vector representation of the conversation history data, a dot product calculation method is used to predict the probability that the candidate document is the most relevant document in the current round of conversation, and a direct relevance probability of the candidate document is obtained; Based on the document d last The vector representation of the conversation history data and the vector representation of the conversation history data are used to predict the document d last is the probability P of the most relevant document in the current round of dialogue keep ; When the document d last When it belongs to the candidate document, the probability P is used keep , correcting the direct relevance probability of the candidate document; and normalizing the corrected probability to obtain a final probability that the candidate document is the most relevant document for the current round of dialogue; When the document d last When the document does not belong to the candidate documents, the direct relevance probability of the candidate document is normalized to obtain the final probability that the candidate document is the most relevant document in the current round of dialogue; Wherein, the probability P is used keep , modifying the direct relevance probability of the candidate document includes: If the candidate document is the document d last , then calculate the direct correlation probability of the candidate document and the P keep The sum of the direct relevance probability of the candidate document is obtained, otherwise, the sum of the direct relevance probability of the candidate document and △ is calculated to obtain the result of correcting the direct relevance probability of the candidate document; wherein △ = 1-P keep .
4. The method according to claim 3, characterized in that The step of using the response generation network to generate response information for the current round of dialogue includes: Based on the encoded representation of the most relevant document and the encoded representation of the conversation history data, a dot product attention mechanism is used to obtain a cross representation of the most relevant document and the conversation history data; Based on the cross representation and the preset vocabulary, using a pre-trained text decoder, and adopting a copy mechanism, an initial response text is generated for the current round of dialogue; the vocabulary is generated based on documents and dialogue texts in a corpus used for model training, and includes special words to be replaced by target knowledge; If the initial response text contains the special word, the target knowledge is retrieved from the most relevant document, and the special word in the initial response text is replaced by the target knowledge to obtain the response information of the current round of dialogue; otherwise, the initial response text is used as the response information of the current round of dialogue.
5. The method according to claim 4, characterized in that Generating the initial response text for the current round of dialogue includes: At each time step of the text decoder, traverse each word w in the vocabulary j , based on all the words currently generated by the text decoder and the cross representation, determine the word t generated at the current time step i For the word w j The initial probability P generate (t i =w j ), based on the encoding representation of the most relevant document, using a multi-layer perceptron, determine when the word t i For the word w j The word t is directly copied from the most relevant document i The probability P copy (t i =w j ); Based on the P generate (t i =w j ), using a multi-layer perceptron, determine when the word t i For the word w j The word t is directly copied from the word list i The probability P(copy); based on the P generate (t i =w j ), said P copy (t i =w j ) and P(copy), according to P(t i =w j )=(1-P(copy)×P copy (t i =w j )+P(copy)×P generate (t i =w j ), determine the t i For the word w j The final probability P(t i =w j );Select the word with the highest final probability from the word list as the word generated in the current time step t i ; Sequentially get all the words t obtained in the time step i Concatenate to obtain the initial response text; The retrieving target knowledge from the most relevant document comprises: Using the text encoder, encoding the initial response text to obtain an encoded representation of the initial response text; Based on the encoded representation of the initial response text and the cross representation, a dot product attention mechanism is used to obtain a cross representation RaD of the initial response text and the cross representation; Based on the cross representation RaD, a multilayer perceptron is used to predict, for each word in the most relevant document, the probability that the word is the starting point of the target knowledge and the probability that the word is the end point of the target knowledge; the word with the highest probability of the starting point is used as the starting point of the target knowledge, and the word with the highest probability of the end point is used as the end point of the target knowledge; based on the starting point and end point of the target knowledge, the target knowledge is obtained from the most relevant document.
6. The method according to claim 5, characterized in that The sample data further includes the response initial text label of the response information of each round of dialogue and the starting label and the ending label of the target knowledge; The calculation of the loss function value includes: Based on the final probability, the P keep , using the correct document label and topic retention label of the current round of dialogue, and using the cross entropy loss function to calculate a first loss function value; the first loss function value is used to optimize and adjust the parameters of the document selection network; Based on the word t i For the word w j The final probability of the starting point, the probability of the end point, the initial text label of the response, and the starting point label and the end point label of the target knowledge are used to calculate the second loss function value by using the cross entropy loss function; the second loss function value is used to optimize and adjust the parameters of the response generation network.
7. A method for generating a dialogue response, characterized in that: include: During the conversation process, when a conversation response needs to be generated, the conversation response generation model is used to generate and output the conversation response based on the currently generated conversation history data and a preset document library; Wherein, the dialogue response generation model is trained based on any training method described in claims 1 to 6.
8. A training device for a dialogue response generation model, characterized in that: including a processor and a memory; The memory stores an application program executable by the processor, which is used to enable the processor to execute the training method of the dialogue response generation model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Document multi-label classification method and device
CN112183655A
Model training method, and task type visual dialogue problem generation method and device
CN112579759A