A topic-driven knowledge selection method in knowledge dialogue generation

By constructing a topic-driven knowledge selection method and using topic information to guide knowledge selection and generation, the problem of inconsistent replies in the dialogue generation system is solved, and more accurate knowledge selection and more meaningful response generation are achieved.

CN115186149BActive Publication Date: 2025-07-18TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210901140.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-07-18
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

The existing dialogue generation system lacks topic matching when generating responses, resulting in inconsistent responses and inability to effectively utilize common sense knowledge, and the generated responses lack meaningful content.

Method used

Construct a context encoder and knowledge encoder that integrates topic information, select appropriate knowledge through topic-driven knowledge selectors, and generate replies using knowledge-aware generation decoder, and minimize the distance between the prior distribution and posterior distribution in the training stage to use prior distribution to select knowledge in the test stage.

Benefits of technology

The generated replies are consistent with the context, improving the accuracy of knowledge selection and the quality of response generation, and enhancing the topic matching capabilities of the dialogue system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115186149B_ABST
    Figure CN115186149B_ABST
Patent Text Reader

Abstract

The present invention discloses a topic-driven knowledge selection method in knowledge dialogue generation. First, a context encoder and a knowledge encoder integrating topic information are constructed. The context encoder and the knowledge encoder are used to extract the context information \(x^{t}\), the true response information \(y^{t}\), the topic information \(f^{t}\), and the external knowledge information \(k^{t}\) in the \(t\)-th round of dialogue in the training corpus. Then, a topic-driven knowledge selector is constructed. The main purpose of the topic-driven knowledge selector is to select appropriate knowledge from candidate knowledge sentences according to the topic information and the context information. The knowledge selector also pays attention to the knowledge information selected in the previous round and models it as a latent variable, so as to perform joint inference on multi-round knowledge selection and response generation. Finally, a knowledge-aware generation decoder is constructed. The main purpose of the knowledge-aware generation decoder is to generate the response \(y^{t}\) of the \(t\)-th round of dialogue according to the knowledge information, the context information \(x^{t}\), and the topic information \(f^{t}\) selected by the knowledge selector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of natural language processing and dialogue system, and specifically to a topic-driven knowledge selection method in knowledge dialogue generation. Background Art

[0002] Open-domain dialogue generation systems show great potential, which can endow machines with the ability to converse with humans in natural language. However, in many real scenarios, the performance of response generation is still far from satisfactory. For example, in a dialogue system, only conditioned on the input text, the text generation system usually produces general or meaningless responses composed of frequent words or phrases in the corpus. For example, for a given input text "My skin is too dry", it usually generates "Me too" or "Oh, my god!" Compared with human responses containing rich knowledge, these responses lack meaningful content. General or meaningless responses are usually caused by the common sense knowledge gap between machines and humans. Common sense knowledge provides information on the relationships between concepts in the context, thus potentially guiding humans to capture the implicit logic between concepts. In this way, humans in real conversations usually connect the context with their common sense knowledge to make informative responses. To solve this problem, researchers have introduced some large-scale knowledge bases to enhance the results of dialogue generation. Many knowledge corpora have also greatly advanced the development of this research topic. The form of knowledge can be unstructured knowledge texts, structured knowledge graphs, or their hybrid representations.

[0003] Knowledge-based dialogue generation aims to generate informative responses based on the discourse context and external knowledge. It mainly focuses on solving two research problems: (1) knowledge selection: selecting appropriate knowledge according to the dialogue context and previously selected knowledge; (2) knowledge-aware generation: injecting the required knowledge to generate meaningful and informative responses. Selecting appropriate knowledge is a prerequisite for the success of a knowledge-based dialogue system. Existing knowledge selection models generally fall into two categories. The first category is non-sequential selection, which only captures the relationship between the current context and background knowledge. To further utilize the dialogue history in terms of both context and knowledge, research tends to use the second category of sequential selection. That is, the model not only depends on the current environment but also uses the previously selected knowledge to facilitate knowledge selection.

[0004] However, previous studies mostly ignored the guiding role of topic information in knowledge selection and only used it for knowledge-aware generation. This may lead to a topic mismatch between the entire dialogue and the selected knowledge, resulting in inconsistent responses with the context. The dialogue system can better select knowledge that matches the context and generate more appropriate responses under the guidance of the dialogue topic information. Therefore, it is necessary to consider topic information in knowledge selection. Summary of the Invention

[0005] The object of the present invention is to overcome the deficiencies in the prior art and provide a topic-driven knowledge selection method in knowledge dialogue generation. This method explores the guiding role of topic information in knowledge selection. Specifically, under the guidance of topic information, we minimize the distance between the prior distribution and the posterior distribution of knowledge selection in the training stage by using the Kullback-Leibler divergence loss (KL loss), so that when there is no posterior distribution in the test stage, the prior distribution can be used to select knowledge. The results have achieved better results in terms of ACC, BLEU-2, BLEU-4, and ROUGE-2 metrics compared to existing models.

[0006] The object of the present invention is achieved through the following technical solutions: A topic-driven knowledge selection method in knowledge dialogue generation, comprising the following steps:

[0007] (1) Construct a context encoder and a knowledge encoder that fuse topic information:

[0008] The context encoder and the knowledge encoder are used to extract the context information x t of the t-th round of dialogue in the training corpus, the true reply information y t the topic information f t and the external knowledge information k t ; it not only makes full use of the advantages of pre-training with a large amount of data, but also fuses and encodes the topic information and the context information through a bidirectional attention mechanism to obtain

[0009] (2) Construct a topic-driven knowledge selector:

[0010] The main purpose of the topic-driven knowledge selector is to select appropriate knowledge from candidate knowledge sentences according to the topic information and the context information. To select more appropriate knowledge, the knowledge selector we constructed also pays attention to the knowledge information selected in the previous round and models it as a latent variable, so as to jointly infer multi-round knowledge selection and reply generation. Specifically, given the encoded topic information f t the context information x t the knowledge k t the previously selected knowledge and the true reply information y t (only in the training stage), the knowledge selector will make full use of this information to select appropriate

[0011] (3) Construct a knowledge-aware generation decoder:

[0012] The main purpose of the knowledge-aware generation decoder is to generate a reply according to the knowledge information selected by the knowledge selector the context information x tand topic information f t Generate the response y for the t-th round of conversation t 。

[0013] Beneficial effects

[0014] 1. The present invention utilizes the guiding role of topic information in knowledge selection, and through the topic matching between the entire conversation and the selected knowledge, the generated response is consistent with the context.

[0015] 2. The dialogue system can better select knowledge that matches the context and generate more appropriate responses under the guidance of the dialogue topic information. Brief description of the drawings

[0016] Figure 1 It is a framework diagram of the topic-driven knowledge selection method in knowledge dialogue generation. Detailed implementation manners

[0017] The present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] Taking the Wizard of Wikipedia knowledge dialogue dataset as an example, the implementation method of the present invention is given. The entire system algorithm process includes 4 steps: processing of the dataset, construction of a context encoder and a knowledge encoder that integrate topic information, construction of a topic-driven knowledge selector, construction of a knowledge-aware generation decoder, and response generation.

[0019] The specific steps are as follows:

[0020] (1) Processing of the dataset

[0021] The Wizard of Wikipedia dataset includes 18,430 training conversations, 1,948 validation conversations, and 1,933 test conversations. The test set is divided into two subsets, visible test and invisible test. The visible test contains 965 conversations whose topics appear in the training set or the validation set, and the invisible test contains 968 conversations whose topics never appear in the training set and the test set.

[0022] (2) Construction of a context encoder and a knowledge encoder that integrate topic information

[0023] The context encoder and the knowledge encoder are composed of a transformer-based bert encoder, which consists of N layers of multi-head attention mechanisms and a feed-forward neural network. It takes the context information x t , external knowledge information k t and the true response information y t and respectively combines them with the topic information f tPerform encoding to obtain the corresponding output. The specific calculation formula is as follows:

[0024]

[0025]

[0026]

[0027]

[0028] (3) Construct a topic-driven knowledge selector

[0029] During the training phase, the true response information y t is used as pseudo-labels to help select appropriate knowledge information. At this time, the selection of knowledge is based on the posterior distribution, denoted as

[0030]

[0031] where is the hidden state of the GRU network, is the hidden state of another GRU network. The initial state W post is a trainable parameter, and subsequently is sampled from the following formula.

[0032]

[0033] During the test phase, since the true response information is unknown, the selection of knowledge at this time is based on the prior knowledge distribution, denoted as

[0034]

[0035]

[0036] where W prior is a trainable parameter. During the test phase, the selection of knowledge is obtained from the following formula:

[0037]

[0038] Immediately afterwards, the sampled knowledge information is fed into the knowledge-aware generation decoder for generating responses.

[0039] (4) Construct a knowledge-aware generation decoder:

[0040] Using GRU as the basic network structure of the decoder, at each time step, the decoder can generate two types of words, namely the words in the vocabulary and the copied words. At the beginning, the decoder first updates the decoder state according to the following formula.

[0041] s t = GRU D (s t-1 , [o(y t-1 ) ; h k ; h f )

[0042] s0 = W D [h c ; h k ; h f + b D

[0043] Here, W D and b D are learnable parameters, and o(y t-1 ) represents the encoding representation of the last step. The words in the vocabulary are sampled by the following formula:

[0044] p v (y t = w) = softmax(w T (W V s t + b V ))

[0045] W V and b V are learnable parameters, and w is the one - hot vector of the vocabulary W.

[0046] The copied words are sampled by the following formula:

[0047]

[0048] Here, H is a fully - connected layer with the activation function tanh. The above two distributions are flexibly combined by the following formula:

[0049] p(y t = w) = (1 - α) * p v (y t = w) + α * p c (y t = w)

[0050] Here, α is a hyperparameter and is set to 0.5. Then we get the generated word according to the following formula.

[0051] y t = arg max w∈vp(y t = w)

[0052] To narrow the gap between the real response and the model-generated response, the NLL loss is used to accomplish this function.

[0053]

[0054] Meanwhile, since the real response is known during the training phase, the posterior distribution can be used to select knowledge. However, the real response is unknown during the testing phase, so only the prior distribution can be used to select knowledge. To make the prior distribution approximate the posterior distribution, the KL loss is used during training.

[0055]

[0056] The overall loss function is:

[0057]

[0058] where λ represents a parameter for response generation.

[0059] Table 1 shows the results of this model (TopicKS), other baseline models (MemNet, PostKS, DiffKS) on the WizardofWiki pedia dataset and evaluation metrics (ACC, BLEU-2, BLEU-4, ROUGE-2). To study the effectiveness of the topic-driven idea on general models, the experimental results after adding topic information to the baseline models are also included in Table 1.

[0060] Table 1 Automatic evaluation results of the Wizard of Wikipedia dataset

[0061]

[0062] The comparative experimental algorithms in the table are described as follows:

[0063] MemNet: This method uses a memory network to store knowledge and selects knowledge based on semantic similarity. We also evaluated a variant (MemNet+Topic_info), where knowledge selection is guided by other topic information.

[0064] PostKS: This method uses the posterior knowledge distribution as the pseudo-label for knowledge selection, but this method ignores the knowledge selection information of the previous round and the importance of topic information for the current round of knowledge selection. Therefore, we also evaluated a variant that uses topic information to assist knowledge selection (PostKS+Topic_info).

[0065] DiffKS: This method uses the difference information between the selected knowledge in knowledge-based multi-turn conversations for knowledge selection. It focuses on the knowledge information in the previous rounds while ignoring the importance of topic information in knowledge selection.

[0066] As can be seen from the experimental results in Table 1, after comparing the performance of different methods on the WoW dataset, the experimental results of the test set and the unseen test set show that the topic-driven knowledge selection method has significant advantages. It can not only improve the accuracy of knowledge selection but also improve the quality of response generation.

[0067] Compared with the baseline model, this method also shows stronger generalization ability from in-domain (test visible) to out-of-domain data (test unseen). The variable experimental results of MemNet and PostKS can further prove the importance of topic information in knowledge selection.

[0068] Introducing topic information in knowledge selection can not only reduce the errors in knowledge selection but also improve the BLUE-2 / 4 and ROUGE-2 metrics in the response generation process.

[0069] Table 2 Ablation experiments on the Wizard of Wikipedia dataset

[0070]

[0071] The ablation experiment results are shown in Table 2. To verify the effectiveness of the topic-driven knowledge selection method, an ablation experiment was conducted on the WOW dataset. The topic information was removed from the model, and the model was trained under this condition. The experimental results are shown in Table 2. The results show that the topic information plays a guiding role in knowledge selection. Under the global guidance of the topic information, not only can the accuracy of knowledge selection be improved, but also the quality of the generated responses can be improved.

[0072] The present invention is not limited to the embodiments described above. The above description of the specific embodiments is intended to describe and illustrate the technical solutions of the present invention. The above specific embodiments are merely illustrative and not restrictive. Without departing from the spirit of the present invention and the scope protected by the claims, those of ordinary skill in the art can make many specific transformations in various forms under the inspiration of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A topic-driven knowledge selection method in knowledge dialogue generation, characterized in that, It includes the following steps: (1) Construct a context encoder and a knowledge encoder that fuse topic information The context encoder and the knowledge encoder are used to extract the context information x of the t-th round of conversation in the training corpus t , the real reply information y t , the topic information f t and the external knowledge information k t ; The topic information and context information are fused and encoded through a bidirectional attention mechanism to obtain (2) Construct a topic-driven knowledge selector The main purpose of the topic-driven knowledge selector is to select appropriate knowledge from candidate knowledge sentences according to topic information and context information; the constructed knowledge selector pays attention to the knowledge information selected in the previous round and models it as a latent variable, so as to jointly reason about multi-round knowledge selection and response generation; Specifically, given the encoded topic information f t , the context information x t , the external knowledge information k t , the previously selected knowledge and the true reply information y t , the knowledge selector will make full use of this information to select appropriate (3) Construct a knowledge-aware generation decoder The main purpose of the knowledge-aware generation decoder is to generate the true response information y for the t-th round of conversation based on the knowledge information selected by the knowledge selector context information x t and topic information f t to generate the true response information y for the t-th round of conversation t ; In step (1), each conversation in the training corpus consists of context, response, and given knowledge information; Among them, the context information in each conversation is represented as the i-th word in the context information of the conversation; the real response information in each conversation is represented as the j-th word in the response information; The candidate knowledge information is represented as First, the context encoder and the knowledge encoder are composed of a Transformer-based BERT encoder, which consists of N layers of multi-head attention mechanisms and feed-forward neural networks; It takes the context information x t , the external knowledge information k t and the true response information y t and encodes them with the topic information f t respectively to obtain the corresponding outputs. The specific calculation formula is as follows:

2. The method for theme-driven knowledge selection in knowledge dialogue generation according to claim 1, characterized in that, In step (2), the goal of the knowledge selector is to select appropriate knowledge from candidate knowledge sentences according to topic information and context information, pay attention to the previously selected knowledge information, and model it as a latent variable, so as to jointly reason about multi-round knowledge selection and response generation; Specifically, given the encoded topic information f t , the context information x t , the external knowledge information k t , the previously selected knowledge and the true response y only in the training phase t , the knowledge selection method will make full use of this information to select the appropriate During the training phase, the true response information y t is used as a pseudo-label to help select appropriate knowledge information. At this time, the selection of knowledge is based on the posterior distribution, denoted as Among them is the hidden state of the GRU network, is the hidden state of another GRU network, with the initial state W post is a trainable parameter, and subsequently is sampled from the following formula; In the test phase, since the true response information is unknown, the knowledge selection at this time is based on the prior knowledge distribution, denoted as W prior are trainable parameters. During the test phase, the selection of knowledge is obtained from the following formula: Immediately afterwards, the sampled knowledge information is fed into the knowledge-aware generation decoder to generate a response.

3. The method for generating a dialogue integrating explicit and implicit personalized information according to claim 1, wherein In step (3), according to the knowledge information selected by the knowledge selector, as well as the topic information and context information, the response information for the t-th round is generated; Use GRU as the basic network structure of the decoder. At each moment, the decoder can generate two types of words, namely the words in the vocabulary and the copied words. At the beginning, the decoder first updates the decoder state according to the following formula, s t = GRU D (s t-1 , [o(y t-1 ) ; h k ; h f ) s0 = W D [h c ; h k ; h f +b D W D and b D are learnable parameters, and o(y t-1 ) represents the encoded representation at the last step. Words in the vocabulary are sampled by the following formula: p v (y t = w) = softmax(w T (W V s t + b V )) W V and b V are learnable parameters, w is the one-hot vector of the vocabulary W, and the copy word is sampled by the following formula: H is a fully connected layer, and the activation function is tanh; The above two distributions are flexibly combined by the following formula: p(y t = w) = (1 - α) * p v (y t = w) + α * p c (y t = w) α is a hyperparameter and is set to 0.5; Then, the generated word is obtained according to the following formula: y t = arg max w∈v p(y t = w) To narrow the gap between the true response and the model-generated response, the NLL loss is used to complete it: At the same time, since the true response is known in the training phase, the posterior distribution can be used to select knowledge, but the true response is unknown in the test phase. Therefore, only the prior distribution can be used to select knowledge. To make the prior distribution approximate the posterior distribution, the KL loss is used during training: The overall loss function is: Among them, λ represents a parameter.

Citation Information

Patent Citations

  • Open domain dialogue generation method and model with strong generalization knowledge selection

    CN112463935A

  • Method, apparatus, device, and storage medium for training model and generating dialog

    US20210342551A1