Open-domain Dialogue Generation Method, System, Device and Medium Based on Dialogue Relationship

By introducing dialogue relationships and focus vectors into the open domain dialogue generation model, the problem of single content generation of traditional models is solved, and a richer and more diverse reply generation is achieved.

CN115563254BActive Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211137397.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2025-07-11
Estimated Expiration
2042-09-19

AI Technical Summary

Technical Problem

The traditional open domain dialogue generation model is insufficient in generating replies, and the end-to-end model cannot effectively adapt to the phenomenon of one question and multiple answers, resulting in a single content generation.

Method used

By annotating the dialogue relationship between the replies of the training samples and historical dialogues, the natural language model Transformer generates focus vectors, and guides the generation of reply contents with dialogue relationship identifiers to enhance diversity.

Benefits of technology

It improves the diversity and fluency of generating replies, enhances the interpretability of the model, and generates content more in line with user expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115563254B_ABST
    Figure CN115563254B_ABST
Patent Text Reader

Abstract

The present invention discloses an open-domain dialogue generation method, system, device and medium based on dialogue relationships. The method includes: obtaining a corpus, obtaining training samples according to the corpus, obtaining the focus of attention of the historical dialogue of each training sample, and obtaining the dialogue relationship between the reply and the focus of attention; training a natural language model according to the training samples for dialogue generation tasks; after the natural language model is trained, calculating the focus of attention of the current historical dialogue according to the dialogue relationship identifier; generating a reply content by combining the historical dialogue, the focus of attention, the reply and the dialogue relationship of the focus of attention. By annotating the dialogue relationship between the reply of each training sample and the historical dialogue it targets, and according to the dialogue relationship, generating the focus of attention of the encoder and the starting generation identifier of the decoder, and integrating them into the generation process of the dialogue model, the generated reply is made more rich and diverse, and can be widely applied to the field of natural language processing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to an open-domain dialogue generation method, system, device and medium based on dialogue relationship. Background Art

[0002] The development of artificial intelligence has brought great convenience to the use of electronic devices by humans. From keyboard and mouse operations to gesture and voice inputs, the interaction between humans and computers has become more natural. Computers have gradually started to actively "understand" information from simply accepting specific information passively. Dialogue is the main way of human communication, and how to build an ideal dialogue system has always been one of the important research directions in artificial intelligence.

[0003] Generally speaking, according to the construction method and application scenario, dialogue systems can be mainly divided into closed-domain dialogue systems and open-domain dialogue systems. Compared with closed-domain dialogue systems, open-domain dialogue systems are mainly used for casual conversations, and the text content they can handle is more extensive. When interacting with an open-domain dialogue system, the input content is not restricted, and users can speak freely. However, the responses given by the dialogue system cannot be controlled and are also unpredictable.

[0004] The methods of open-domain dialogue systems can be mainly divided into two categories: (1) Retrieval-based methods: When a user asks a question, the system searches for a batch of question-answer pairs in a corpus that are closest to the user's question. Through a sorting algorithm, the best question-answer pair is selected, and finally, the answer of the question-answer pair is output as the corresponding response. The question-answer pairs mainly come from real human conversations. Therefore, when studying retrieval-based methods, more attention is paid to semantic representation, similarity measurement, and sorting methods, etc. (2) Generation-based methods: Use neural networks to establish a mapping relationship between historical conversations and reply content. The system encodes the content input by the user into a vector form, and then the model calculates the probability of generating the next word one by one based on these vectors until the entire sentence reply is obtained. Open-domain dialogue systems have good application prospects. An ideal open-domain dialogue model can not only improve the usage experience of intelligent dialogue systems such as AI customer service, but also be more suitable for aspects such as home care, early education, and training of autistic children.

[0005] Traditional dialogue generation models have the following disadvantages: (i) Ordinary dialogue models cannot generate responses for a specific historical dialogue sentence. There are many samples in the training of dialogue models that are not responses to the last sentence, but most of the training samples' responses are for the last historical dialogue sentence. The model will overly focus on the last round of the historical dialogue, greatly reducing the diversity of the generated content. (ii) End-to-end dialogue models cannot effectively adapt to the phenomenon of one question with multiple answers. During the training process, if there are two samples with similar contexts but different responses, the parameters of the dialogue model will converge in different directions. These two training samples have a competitive relationship, and finally the model tends to generate the safer solution that appears more frequently. Summary of the Invention

[0006] To at least to some extent solve one of the technical problems existing in the prior art, an object of the present invention is to provide an open-domain dialogue generation method, system, device and medium based on dialogue relationships.

[0007] The technical solution adopted by the present invention is as follows:

[0008] An open-domain dialogue generation method based on dialogue relationships, comprising the following steps:

[0009] Obtain a corpus, obtain training samples according to the corpus, obtain the focus of attention of the historical dialogue of each training sample, and obtain the dialogue relationship between the response and the focus of attention;

[0010] Train the natural language model Transformer for dialogue generation tasks according to the training samples; wherein, the encoder generates a focus vector according to the focus of attention of the training samples, and integrates the focus vector into the encoding process of the context semantic vector; the decoder starts from the start identifier, combines the semantic vector given by the encoder, and generates the response content word by word, and finally fits the mapping relationship between the historical dialogue and the response content through an optimizer;

[0011] After the natural language model Transformer is trained, set the starting dialogue relationship identifier, and calculate the focus of attention of the current historical dialogue according to the dialogue relationship identifier;

[0012] Generate the response content by combining the historical dialogue, the focus of attention, the response and the dialogue relationship of the focus of attention.

[0013] Further, the training samples are obtained by splitting and parsing the dialogue corpus, and specifically include:

[0014] Retrieve each long conversation from the corpus. Starting from the second sentence of the long conversation, sequentially obtain the conversation sentences as the response sentences of the training samples, and use the several rounds of conversations before each response sentence as the historical conversations. The specific number of rounds needs to be set manually according to the characteristics of the dataset; Parse the samples through a conversation relationship analysis tool to obtain the position of the sentence in the historical conversation that the response targets, that is, the focus of attention, and the conversation relationship between the response and that sentence; Each training sample consists of the following parts:

[0015] (1) Several sentences of historical conversation C = {U1, …, U m}, where each sentence of historical conversation consists of several words

[0016] (2) The response conversation Y = {y1, …, y k} containing several words;

[0017] (3) The position t of the sentence in the historical conversation that the response targets, t ∈ [1, m];

[0018] (4) The conversation relationship R between the targeted historical conversation and the response.

[0019] Furthermore, the encoder generates an intermediate semantic vector of the historical conversation based on the historical conversation of the training sample and the focus of attention, including:

[0020] Add a start identifier before each sentence of historical conversation to separate the historical conversations;

[0021] Initialize the word vector space, word position vector space, turn vector space, and focus vector space; For each word, calculate the word vector according to the number of the word in the word list, calculate the position vector according to the position of the word in the sentence, calculate the turn vector according to the turn number of the sentence where the word is located, and calculate the focus vector according to whether the word is in the sentence that the response conversation targets;

[0022] Add the word vector, position vector, turn vector, and focus vector together to obtain a set of hidden layer vectors containing semantic information and position information, corresponding one-to-one to each word in the historical conversation;

[0023] Input the hidden layer vectors of the words in the historical conversation into an N-layer encoder; Each layer of the encoder consists of a multi-head attention mechanism network with a residual addition mechanism and a fully connected layer; Each head of the multi-head attention mechanism network contains W q 、W k 、W vThree matrices, the matrix composed of the hidden layer vectors of words in the historical conversation is multiplied by these three matrices one by one to obtain three matrices Q, K, and V respectively. Subsequently, the Q matrix is multiplied by the transposed K matrix, and through the softmax layer, the attention degree S of each word to other words is obtained:

[0024]

[0025] Among them, d head is the dimension of the hidden layer vector of the word after linear transformation by the three matrices W q , W k , W v :

[0026] The hidden layer vectors of words are combined according to the weights in the attention degree S to form a new vector representing the word. Finally, all the outputs of the multi-head attention mechanism network are combined, and through the linear layer, the intermediate semantic vector of each word in the historical conversation is obtained as the intermediate semantic vector of the historical conversation.

[0027] Furthermore, the decoder starts from the dialogue relationship identifier between the reply content and the historical conversation it targets, combines the semantic vector given by the encoder, and generates the reply content word by word. Finally, the mapping relationship between the historical conversation and the reply content is fitted through the optimizer, including:

[0028] Through the word embedding layer and the position embedding layer, the hidden layer vectors corresponding to the start identifier and the generated words are obtained, and these hidden layer vectors are input into an N-layer decoder;

[0029] Among them, each layer of the decoder consists of two multi-head attention networks with a residual addition mechanism and a fully connected layer; the first multi-head attention mechanism network encodes the generated reply content; the second multi-head attention mechanism network obtains the Q matrix through the hidden layer vector of the generated reply content, obtains the K matrix through the context semantic vector, and calculates the attention degree of the words in the generated reply content to the words in the historical conversation. Finally, the context semantic vector is combined and passed through the linear layer to obtain the output of this layer of the decoder;

[0030] After obtaining the vector matrix of the decoder output, the vector of the last word is obtained, and through the linear layer and the softmax layer, the probability distribution of the next word in the dictionary is obtained; repeat the above decoding process until the generated word is the end identifier, then the reply sentence generation is completed;

[0031] The generation probability of the reply sentence of each training sample is expressed as:

[0032]

[0033] In the formula, C represents the historical conversation, y represents the word in the reply sentence, p(y l|y <l , C) represents the probability distribution of the next word calculated based on the previous l-1 words and the historical conversation;

[0034] The loss function during the training process is:

[0035] L = -logp(y1,…y k |C)

[0036] Furthermore, in practical applications, since the next sentence of a multi-round conversation is unpredictable and it is impossible to independently predict the historical conversation and conversation relationship matching the next round of conversation, therefore, before generating a response, the conversation relationship and the focus of attention on the historical conversation are given;

[0037] Among them, the method of giving the conversation relationship is:

[0038] Train to obtain a personalized model, and calculate the conversation relationship corresponding to the current response according to the personalized conversation habit; or,

[0039] Generate the distribution probability of the conversation relationship through the content of the historical conversation, and obtain the conversation relationship corresponding to the maximum distribution probability as the conversation relationship of the next sentence.

[0040] Furthermore, during the decoding process, the conversation relationship R is used as the starting identifier.

[0041] Another technical solution adopted by the present invention is:

[0042] An open-domain conversation generation system based on conversation relationship, comprising:

[0043] A data acquisition module, configured to acquire a corpus, obtain training samples according to the corpus, obtain the focus of attention of the historical conversation of each training sample, and obtain the conversation relationship between the response and the focus of attention;

[0044] A model training module, configured to train a natural language model Transformer for conversation generation tasks according to the training samples; wherein, the encoder generates a focus of attention vector according to the focus of attention of the training samples and integrates the focus of attention vector into the encoding process of the context semantic vector; the decoder starts from the starting identifier, combines the semantic vector given by the encoder, and generates the response content word by word, and finally fits the mapping relationship between the historical conversation and the response content through an optimizer;

[0045] An actual application module, configured to, after the natural language model Transformer is trained, set the starting conversation relationship identifier, calculate the focus of attention of the current historical conversation according to the conversation relationship identifier; generate the response content by combining the historical conversation, the focus of attention, the response, and the conversation relationship of the focus of attention.

[0046] Another technical solution adopted by the present invention is:

[0047] An open-domain dialogue generation device based on dialogue relationships, comprising:

[0048] At least one processor;

[0049] At least one memory for storing at least one program;

[0050] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned method.

[0051] Another technical solution adopted by the present invention is:

[0052] A computer-readable storage medium storing a program executable by a processor, the program executable by the processor being used to execute the method as described above when executed by the processor.

[0053] The beneficial effects of the present invention are as follows: By annotating the dialogue relationship between the response of each training sample and the historical dialogue it targets, and according to the dialogue relationship, generating the attention focus of the encoder and the starting generation identifier of the decoder, and integrating them into the generation process of the dialogue model, the generated responses are made more diverse. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 is a flowchart of the steps of an open-domain dialogue generation method based on dialogue relationships in an embodiment of the present invention;

[0056] Figure 2 is a flowchart of an embodiment of an open-domain dialogue generation method based on dialogue relationships in an embodiment of the present invention;

[0057] Figure 3 is a model framework diagram of an embodiment of an open-domain dialogue generation method based on dialogue relationships in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention. For the step numbers in the following embodiments, they are only set for convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0059] In the description of the present invention, it should be understood that with regard to the orientation description, such as the upper, lower, front, rear, left, right, etc., the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present invention.

[0060] In the description of the present invention, the meaning of "several" is one or more, the meaning of "multiple" is two or more, and understandings such as "greater than", "less than", "exceeding", etc. do not include the recited number, and understandings such as "above", "below", "within", etc. include the recited number. If the first and second are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0061] In the description of the present invention, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0062] Traditional open-domain dialogue models cannot identify which part of the historical dialogue a current sample's response is generated for. When training an open-domain dialogue generation model, there is a one-to-many phenomenon between the historical dialogues and responses in the corpus, that is, similar historical dialogues may correspond to different responses. There is a competitive relationship between samples with different responses but similar historical dialogues, resulting in a poor training effect of the model, making the model tend to generate safe solutions and reducing the diversity of dialogue generation.

[0063] In the process of the end-to-end open-domain dialogue model generating responses verbatim, not only the semantics of the context but also the previously generated content, that is, the sentences that are not fully generated, need to be considered. In the experiment, we found that although the content generated by the traditional dialogue model does not exactly correspond to the historical dialogue, the grammatical accuracy of the sentences generated by the model is very high. Thus, it is speculated that the partially ungenerated sentences have a greater impact on the subsequent generated content. At the same time, there are significant differences in the beginning parts of the sentences for different dialogue behaviors. For example, responses as answers (question-answer_pair) are more commonly "yes" or "no", and responses as questions (clarification question) are more commonly "what" or "how", etc. Therefore, the present invention uses the dialogue relationship as the starting tag for generation and marks the sentences of interest in the historical dialogue to guide the model to converge in different directions in different training samples.

[0064] See Figure 1 、 Figure 2 and Figure 3 , this embodiment provides an open-domain dialogue generation method based on dialogue relationship, which guides the generation of response content through the dialogue relationship and the focus of attention, enhancing the generation diversity and interpretability of the traditional end-to-end dialogue model. The method specifically includes the following steps:

[0065] S1. Obtain a corpus, obtain training samples according to the corpus, obtain the focus of attention of the historical dialogue for each training sample, and obtain the dialogue relationship between the response and the focus of attention.

[0066] Preprocess the corpus and divide the training samples. For each sample, obtain the focus of attention of the historical dialogue through the dialogue relationship analysis model, that is, the position of the utterance targeted by the response in the historical dialogue, and the dialogue relationship between the response and the focus of attention:

[0067] In this embodiment, the dialogue dataset uses the DailyDialog corpus, which consists of dialogue content, dialogue topics, dialogue acts, and sentiment labels. The original data comes from the daily conversations of English learners, and the sentiment labels are manually annotated. This corpus not only has rich labels, but also has more formal dialogue grammar and more suitable number of turns for each dialogue to train the model. The total number of characters in this corpus is about 100 million, and the total number of dialogue sentences included is 7 million. During the process of dividing the training samples, for each dialogue in the corpus, starting from the second sentence, each sentence is used as a reply, and at most three sentences in front of it are used as the historical dialogue to form the dialogue content of the training sample. After dividing the dialogue samples from the corpus, the Deep Sequential dialogue relationship analysis model trained by the STAC corpus is used to analyze the reply content of each sample, the corresponding historical dialogue content, and the dialogue relationship between them. Finally, 16 types of dialogue relationships such as question-answer pair, clarification question, comment, contrast, continuation, and explanation are obtained. Retain the 6 types of dialogue relationships with the largest quantity, and merge the remaining 10 types of dialogue relationships accounting for about 10% into the other category.

[0068] S2. Train the natural language model Transformer for the dialogue generation task. Among them, the encoder generates a focus vector according to the focus of the training sample, and integrates the focus vector into the encoding process of the context semantic vector; the decoder starts from the start identifier, combines the semantic vector given by the encoder, and generates the reply content word by word, and finally fits the mapping relationship between the historical dialogue and the reply content through the optimizer.

[0069] Train the natural language model Transformer for the dialogue generation task. The encoder generates a focus vector according to the position of the historical dialogue targeted by the current reply in the training sample, and integrates it into the encoding process of the context semantic vector. The decoder starts from the dialogue relationship identifier between the reply content and the targeted historical dialogue, combines the semantic vector given by the encoder, and generates the reply content word by word. Finally, the optimizer fits the mapping relationship between the historical dialogue and the reply content. Specifically as follows:

[0070] Each training sample includes: (1) several historical dialogues C = {U1,...U m}, where each sentence consists of several words (2) a reply dialogue Y = {y1,...y k}; (3) The position t of the statement the response targets in the historical dialogue, where t ∈ [1, m]; (4) The dialogue relationship R between the historical dialogue being targeted and the response. Before encoding the historical dialogue, it is also necessary to preprocess the historical dialogue: First, add the start identifier "[CLS]" before each sentence of the historical dialogue to segment the historical dialogue. Subsequently, initialize the word vector space, word position vector space, turn vector space, and focus vector space. For each word, calculate the word vector Word Embedding according to its number in the word table, calculate the position vector Position Embedding according to its position in the sentence, calculate the turn vector Turn Embedding according to the turn number of the sentence it is in, and calculate the focus vector Focus Embedding according to whether it is in the sentence the response targets. Add the four vectors together to obtain a set of hidden layer vectors containing semantic information and position information, corresponding one-to-one to each word in the historical dialogue.

[0071] After obtaining the initial hidden layer state of the words in the historical dialogue, input it into an N-layer encoder. Each layer consists of a multi-head attention mechanism network with a residual addition mechanism and a fully connected layer. Each head of the multi-head attention mechanism network contains W q 、W k 、W v Three matrices. The matrix composed of the hidden layer vectors of the words in the historical dialogue is multiplied by these three matrices one by one to obtain three matrices Q, K, and V respectively. Subsequently, the Q matrix is multiplied by the transposed K matrix, and after passing through the softmax layer, the attention degree S of each word to other words is obtained:

[0072]

[0073] where, d head is the dimension of the hidden layer vector of the word after linear transformation by the three matrices W q 、W k 、W v . Subsequently, combine the hidden layer vectors of the words according to the weights in S to form a new vector representing a certain word. Finally, combine all the outputs of the multi-head attention mechanism network, pass through a linear layer, and input it into the next layer. The output matrix finally obtained is the intermediate semantic vector of each word in the historical dialogue.

[0074] In the process of the end-to-end dialogue model generating responses word by word, it is necessary to consider both the context semantic vector and the previously generated content, that is:

[0075]

[0076] Therefore, when the model calculates the distribution probability of the next word in the dictionary, the previously generated content also plays a very important role. After completing the encoding, during the decoding process, different from the traditional end-to-end dialogue model that uniformly uses "[CLS]" as the start identifier for each sample, this method uses the dialogue relationship R as the start identifier to increase the diversity of the generated reply content. During decoding, according to the already generated reply words, after word embedding and position embedding, their hidden layer vectors are input into an N-layer decoder. Each layer consists of two multi-head attention networks with a residual addition mechanism and a fully connected layer. Similar to the encoding process, the first multi-head attention mechanism network encodes the already generated reply content, while the second multi-head attention mechanism network obtains the Q matrix through the hidden layer vectors of the already generated reply content, obtains the K matrix through the context semantic vector, and calculates the attention degree of the words in the already generated reply content to the historical dialogue words. Finally, the context semantic vector is combined and passed through a linear layer to obtain the output of this layer of the decoder.

[0077] After obtaining the vector matrix of the decoder output, take the vector of the last word, and through a linear layer and a softmax layer, obtain the probability distribution of the next word in the dictionary. The loss function during the training process is:

[0078] L = -logp(y1,…y k |C)

[0079] S3. After training the natural language model Transformer, set the starting dialogue relationship identifier, and calculate the focus of attention of the current historical dialogue according to the dialogue relationship identifier.

[0080] S4. Generate the reply content by combining the historical dialogue, the focus of attention, the reply, and the dialogue relationship of the focus of attention.

[0081] In actual use, it is necessary to give the starting dialogue relationship identifier through a personalized module or heuristic setting, and then calculate the focus of attention of the current historical dialogue according to the dialogue relationship identifier. Finally, generate the reply content by combining the historical dialogue, the focus of attention, the reply, and the dialogue relationship of the focus of attention.

[0082] In practical applications, since the next sentence of a multi-round conversation is unpredictable, the model cannot autonomously predict which historical conversation the next round of conversation is targeted at and the conversation relationship between them. Therefore, before the model generates a response, the conversation relationship and the model's focus of attention on the historical conversation need to be given. The methods for giving the conversation relationship are as follows: (1) Train a personalized model to calculate which conversation relationship the current response should be according to personalized conversation habits; (2) Generate the distribution probability of the conversation relationship of the next sentence through the content of the historical conversation and select the most likely one. Since there is a strong corresponding relationship between the conversation relationship and the focus of attention, after obtaining the conversation relationship of the next sentence, it is possible to obtain which historical conversation the next sentence is targeted at. For example, in the "question-answer pair" relationship, the response is targeted at the last historical conversation.

[0083] In this embodiment, experiments on the dialogue generation task were carried out on the DailyDialog dataset. The experimental results were compared with multiple language models, namely: ReCoSa, Transformer, and GATE. As shown in Table 1, Table 1 shows the PPL, BLEU-2, and Dist-2 score results of the method of the present invention and three baseline models on the DailyDialog dataset. Through the comparison of the experimental results in Table 1, we found that the results of the present invention have been improved compared with the baseline models, and the generated responses are stronger than other baseline models in terms of fluency, accuracy, and diversity. Among them, the two indicators of BLEU-2 and Dist-2 are particularly obvious. This result shows that the responses generated by the present invention in the open-domain dialogue generation task have stronger diversity.

[0084] Table 1

[0085] PPL BLEU-2 Dist-2 ReCoSa 19.846 20.538 16.611 Transformer 18.275 19.519 17.381 GATE(δ=1) 18.405 19.142 17.742 This invention* 18.002* 22.985* 19.499*

[0086] In summary, compared with the prior art, this embodiment has the following advantages and beneficial effects:

[0087] (1) The responses generated by the present invention are stronger than traditional end-to-end dialogue models in terms of fluency, accuracy, and diversity, and have a better performance in the open-domain dialogue generation task.

[0088] (2) When generating responses, the present invention can guide the generation direction of the model through a personalized module or heuristic setting, and the interpretability of the model is stronger.

[0089] This embodiment also provides an open-domain dialogue generation system based on conversation relationships, including:

[0090] A data acquisition module for acquiring a corpus, obtaining training samples according to the corpus, obtaining the focus of attention of the historical conversation of each training sample, and obtaining the conversation relationship between the response and the focus of attention;

[0091] A model training module for training a natural language model Transformer for dialogue generation tasks according to training samples; wherein, the encoder generates a focus vector based on the focus of the training samples and incorporates the focus vector into the encoding process of the context semantic vector; the decoder starts from the start identifier, combines the semantic vector given by the encoder, generates the response content word by word, and finally fits the mapping relationship between the historical dialogue and the response content through an optimizer;

[0092] An actual application module for setting the starting dialogue relationship identifier after the natural language model Transformer is trained, calculating the focus of the current historical dialogue according to the dialogue relationship identifier; generating the response content by combining the four elements of the historical dialogue, the focus, the response, and the dialogue relationship of the focus.

[0093] An open-domain dialogue generation system based on dialogue relationship in this embodiment can execute an open-domain dialogue generation method provided by the method embodiment of the present invention, can execute any combination of the implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0094] This embodiment also provides an open-domain dialogue generation device based on dialogue relationship, including:

[0095] At least one processor;

[0096] At least one memory for storing at least one program;

[0097] When the at least one program is executed by the at least one processor, the at least one processor is caused to implement Figure 1 The method shown.

[0098] An open-domain dialogue generation device based on dialogue relationship in this embodiment can execute an open-domain dialogue generation method provided by the method embodiment of the present invention, can execute any combination of the implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0099] This embodiment also provides a storage medium storing instructions or programs that can execute an open-domain dialogue generation method provided by the method embodiment of the present invention. When the instructions or programs are run, any combination of the implementation steps of the method embodiment can be executed, and the corresponding functions and beneficial effects of the method are achieved.

[0100] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially concurrently or the blocks may sometimes be executed in reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.

[0101] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, the scope of which is determined by the full scope of the appended claims and their equivalents.

[0102] If the described functions are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0103] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered a definitional sequence of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.

[0104] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0105] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0106] In the foregoing description of this specification, the descriptions with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0107] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0108] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. An open-domain dialogue generation method based on dialogue relationships, characterized in that, It includes the following steps: Obtain a corpus, obtain training samples according to the corpus, obtain the focus of attention of the historical conversation of each training sample, and obtain the conversation relationship between the response and the focus of attention; Train the natural language model Transformer for the dialogue generation task according to the training samples; among them, the encoder generates a focus vector according to the focus of attention of the training samples and integrates the focus vector into the encoding process of the context semantic vector; the decoder starts from the start identifier, combines the semantic vector given by the encoder, generates the response content word by word, and finally fits the mapping relationship between the historical conversation and the response content through the optimizer; After the natural language model Transformer is trained, set the starting conversation relationship identifier and calculate the focus of attention of the current historical conversation according to the conversation relationship identifier; Generate the response content by combining the historical conversation, the focus of attention, the response, and the conversation relationship of the focus of attention; The encoder generates an intermediate semantic vector of the historical conversation according to the historical conversation and the focus of attention of the training samples, including: Add a start identifier before each sentence of the historical conversation to segment the historical conversation; Initialize the word vector space, word position vector space, turn vector space, and focus vector space; for each word, calculate the word vector according to the number of the word in the word list, calculate the position vector according to the position of the word in the sentence, calculate the turn vector according to the turn number of the sentence where the word is located, and calculate the focus vector according to whether the word is in the sentence targeted by the response conversation; Add the word vector, position vector, turn vector, and focus vector to obtain a set of hidden layer vectors containing semantic information and position information, corresponding one by one to each word in the historical conversation; Input the hidden layer vectors of the words in the historical dialogue into an N-layer encoder; each layer of the encoder consists of a multi-head attention mechanism network with a residual addition mechanism and a fully connected layer; each head of the multi-head attention mechanism network contains W q , W k , W v Three matrices, the matrix composed of the hidden layer vectors of the words in the historical dialogue is multiplied by these three matrices one by one to obtain three matrices Q, K, and V respectively. Subsequently, the Q matrix is multiplied by the transposed K matrix, and the attention degree S of each word to other words is obtained through the softmax layer: Among them, d head is the dimension of the hidden layer vector of the word after being linearly transformed by the three matrices W q , W k , and W v . Combine the hidden layer vectors of the words according to the weights in the attention degree S to form a new vector representing the word, and finally combine all the outputs of the multi-head attention mechanism network, pass through the linear layer, and obtain the intermediate semantic vector of each word in the historical conversation as the intermediate semantic vector of the historical conversation; The decoder starts from the conversation relationship identifier between the response content and the targeted historical conversation, combines the semantic vector given by the encoder, generates the response content word by word, and finally fits the mapping relationship between the historical conversation and the response content through the optimizer, including: Through the word embedding layer and the position embedding layer, obtain the hidden layer vectors corresponding to the start identifier and the generated words, and input these hidden layer vectors into an N-layer decoder; Among them, each layer of the decoder consists of two multi-head attention networks with a residual addition mechanism and a fully connected layer; the first multi-head attention mechanism network encodes the generated response content; the second multi-head attention mechanism network obtains the Q matrix through the hidden layer vector of the generated response content, obtains the K matrix through the context semantic vector, and calculates the attention degree of the words in the generated response content to the words in the historical conversation, and finally combines the context semantic vector and passes through the linear layer to obtain the output of this layer of the decoder; After obtaining the vector matrix output by the decoder, obtain the vector of the last word, and obtain the probability distribution of the next word in the dictionary through the linear layer and the softmax layer; The generation probability of the response sentence for each training sample is expressed as: Where C represents the historical conversation, y represents the words in the response sentence, and p(y l |y <l , C) represents the probability distribution of the next word calculated based on the previous l - 1 words and the historical conversation; The loss function during the training process is: L = -log p(y1, … y k |C).

2. The open domain dialogue generation method based on dialogue relationship according to claim 1, wherein The training samples are obtained by segmenting and parsing the dialogue corpus, and specifically include: Obtain each long dialogue from the corpus. Starting from the second sentence of the long dialogue, sequentially obtain the dialogue sentences as the response sentences of the training samples, and the previous several rounds of dialogue before each response sentence as the historical dialogue; Parse the samples through a dialogue relationship analysis tool to obtain the position of the sentence in the historical dialogue that the response targets, that is, the focus of attention, and the dialogue relationship between the response and this sentence; Each training sample consists of the following parts: (1) A number of historical conversations C = {U1, …, U m}, where each historical conversation consists of a number of words (2) The reply dialogue Y containing several words = {y1, … y k}; (3) The position t of the sentence that the response targets in the historical dialogue, t ∈ [1, m]; (4) The dialogue relationship R between the targeted historical dialogue and the response.

3. The open-domain dialogue generation method based on dialogue relationship according to claim 1, wherein In practical applications, since the next sentence of a multi-round dialogue is unpredictable, it is impossible to independently predict the historical dialogue and the dialogue relationship that the next round of dialogue matches. Therefore, before generating a response, the dialogue relationship and the focus of attention on the historical dialogue are given; Among them, the method of giving the dialogue relationship is: Train a personalized model, and calculate the dialogue relationship corresponding to the current response according to the personalized dialogue habit; or, generate the distribution probability of the dialogue relationship through the content of the historical dialogue, and obtain the dialogue relationship corresponding to the maximum distribution probability as the dialogue relationship of the next sentence.

4. The open-domain dialogue generation method based on dialogue relationship according to claim 1, wherein During the decoding process, use the dialogue relationship R as the starting identifier.

5. An open-domain dialogue generation system based on dialogue relationships, characterized in that, Including: A data acquisition module for obtaining the corpus, obtaining training samples according to the corpus, obtaining the focus of attention of the historical dialogue of each training sample, and obtaining the dialogue relationship between the response and the focus of attention; A model training module for training the natural language model Transformer for the dialogue generation task according to the training samples; Among them, the encoder generates a focus vector according to the focus of attention of the training sample, and integrates the focus vector into the encoding process of the context semantic vector; The decoder starts from the starting identifier, combines the semantic vector given by the encoder, and generates the response content word by word. Finally, the optimizer fits the mapping relationship between the historical dialogue and the response content; A practical application module for setting the starting dialogue relationship identifier after the natural language model Transformer is trained, and calculating the focus of attention of the current historical dialogue according to the dialogue relationship identifier; Generate the response content by combining the historical dialogue, the focus of attention, the response, and the dialogue relationship of the focus of attention; The encoder generates an intermediate semantic vector of the historical dialogue according to the historical dialogue and the focus of attention of the training sample, including: Add a start identifier before each historical dialogue sentence to segment the historical dialogue; Initialize the word vector space, word position vector space, round vector space, and focus vector space; For each word, calculate the word vector according to the number of the word in the word list, calculate the position vector according to the position of the word in the sentence, calculate the round vector according to the round number of the sentence where the word is located, and calculate the focus vector according to whether the word is in the sentence that the response dialogue targets; Add the four types of vectors, namely word vectors, position vectors, turn vectors, and focus vectors, to obtain a set of hidden layer vectors containing semantic and position information, which correspond one-to-one to each word in the historical dialogue; Input the hidden layer vectors of words in the historical dialogue into an N-layer encoder; where each layer of the encoder consists of a multi-head attention mechanism network with a residual addition mechanism and a fully connected layer; each head of the multi-head attention mechanism network contains W q 、W k 、W v Three matrices, the matrix composed of the hidden layer vectors of words in the historical dialogue is multiplied by these three matrices one by one to obtain three matrices Q, K, and V respectively. Subsequently, the Q matrix is multiplied by the transposed K matrix, and the attention degree S of each word to other words is obtained through the softmax layer: Among them, d head is the dimension of the hidden layer vector of the word after linear transformation by three matrices W q , W k , W v . Combine the hidden layer vectors of words according to the weights in the attention degree S to form new vectors representing words. Finally, combine all the outputs of the multi-head attention mechanism network, and pass through a linear layer to obtain the intermediate semantic vectors of each word in the historical dialogue as the intermediate semantic vector of the historical dialogue; The decoder starts from the dialogue relationship identifier between the reply content and the targeted historical dialogue, combines the semantic vectors given by the encoder, and generates the reply content word by word. Finally, the optimizer fits the mapping relationship between the historical dialogue and the reply content, including: Through the word embedding layer and the position embedding layer, obtain the hidden layer vectors corresponding to the start identifier and the generated words, and input these hidden layer vectors into an N-layer decoder; Among them, each layer of the decoder consists of two multi-head attention networks with a residual addition mechanism and a fully connected layer; the first multi-head attention mechanism network encodes the generated reply content; the second multi-head attention mechanism network obtains the Q matrix through the hidden layer vectors of the generated reply content, obtains the K matrix through the context semantic vectors, and calculates the attention degree of the words in the generated reply content to the words in the historical dialogue. Finally, combine the context semantic vectors and pass through a linear layer to obtain the output of this layer of the decoder; After obtaining the vector matrix of the decoder output, obtain the vector of the last word, and pass through a linear layer and a softmax layer to obtain the probability distribution of the next word in the dictionary; The generation probability of the reply sentence of each training sample is expressed as: where C represents the historical conversation, y represents the words in the response sentence, and p(y l |y <l , C) represents the probability distribution of the next word calculated based on the previous l - 1 words and the historical conversation; The loss function during the training process is: L = -log p(y1,…y k |C).

6. An open-domain dialogue generation device based on dialogue relationships, characterized in that, Including: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-4.

7. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor is used to execute the method according to any one of claims 1-4 when executed by the processor.

Citation Information

Patent Citations

  • Training method and device of dialogue generation model and dialogue generation method and device

    CN114547272A