A multi-turn empathetic dialogue generation method integrating multi-source common sense knowledge

CN118211662BActive Publication Date: 2026-09-25HUAZHONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410349774.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2026-09-25
Estimated Expiration
2044-03-26

AI Technical Summary

Technical Problem

尽管这些研究工作很大程度上推动了该领域的进步,但大多数研究仍然局限于将单一形式的知识集成到对话模型中

Benefits of technology

[0055]1.本发明提供一种整合多源常识知识的多轮共情对话生成方法,这种方法的核心目的是实现多源知识的有效融合,并将其注入到预训练的对话模型中,通过这种创新性的方法,解决目前单源知识依赖的对话模型在共情能力上的不足,改善了现有对话生成模型面对多源知识拓展性差的问题,显著提高了模型面对多种知识表示形式时的扩展性和灵活性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118211662B_ABST
    Figure CN118211662B_ABST
Patent Text Reader

Abstract

The application provides a multi-turn empathetic dialogue generation method integrating multi-source common sense knowledge, comprising the following steps: S1, inputting historical multi-turn dialogue content as dialogue context into COMET, LLama2 and ConceptNet knowledge base respectively, so as to obtain three types of common sense knowledge: phrase type reasoning knowledge, document type fact knowledge and triple type ontology knowledge; S2, using word embedding technology, a Transformer encoder and a graph neural network (GNN) to encode and process the obtained reasoning knowledge, fact knowledge and ontology knowledge respectively, to generate corresponding reasoning knowledge feature vectors, fact knowledge feature vectors and ontology graph feature vectors; S3, inputting the encoded knowledge feature vectors and dialogue context into a pre-trained model, generating an empathetic reply through processing and generation of the model; the application effectively solves the shortcomings of single source knowledge in coverage and field knowledge restriction by integrating multi-source common sense knowledge. This method significantly enhances the model's ability in semantic understanding and emotional perception, making the generated empathetic reply more accurate, rich and diverse, and improving the overall performance of the dialogue system and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of natural language processing, specifically relating to a method for generating multi-turn empathetic dialogues that integrates multi-source common sense knowledge. Background Technology

[0002] Against the backdrop of current social development and technological progress, humanity's need for emotional communication is growing. Empathic dialogue tasks, as an important research area in the field of artificial intelligence in recent years, aim to leverage AI technology to achieve emotional communication and interaction between computers and humans. Research in this field strives to help people understand and express emotions more deeply, in order to meet humanity's ever-increasing emotional needs.

[0003] With the rapid development of internet technology, numerous knowledge bases have been built, such as ConceptNet and ATOMIC, accumulating a vast amount of knowledge information. In recent years, work on empathic responses has focused on how to integrate information from these knowledge bases into dialogue models to enhance their empathic capabilities. While these studies have significantly advanced the field, most remain limited to integrating single forms of knowledge into dialogue models. Therefore, the limitation of relying on a single knowledge source may restrict the model's ability to generate empathic responses. To overcome this problem, it is necessary to explore how to integrate multiple forms of knowledge to improve the empathic capabilities and coverage of dialogue systems. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention aims to provide a multi-turn empathetic dialogue generation method that integrates multi-source common-sense knowledge. The core objective of this method is to effectively fuse multi-source knowledge and inject it into a pre-trained dialogue model. This innovative approach overcomes the deficiencies in empathetic capabilities of current dialogue models that rely on single-source knowledge, and significantly improves the model's scalability and flexibility when faced with various knowledge representations. The implementation of this method will greatly enhance the dialogue system's ability to understand and generate responses, especially when handling complex, multi-dimensional emotional communication scenarios.

[0005] To achieve the above objectives, this invention provides a method for generating empathic dialogues that integrates multi-source knowledge, comprising the following steps:

[0006] S1. Use the historical multi-turn dialogue content as the dialogue context and input it into COMET, LLama2 and ConceptNet respectively to obtain three types of common sense knowledge: phrase-type reasoning knowledge, document-type factual knowledge and triple-type ontology knowledge.

[0007] S2. For different forms of knowledge, three encoding methods are used to transform textual knowledge into feature vectors for subsequent injection into the pre-trained model. For phrase-type reasoning knowledge, word embedding technology is used to transform the semantic information of words or phrases into vectors with fixed dimensions; for document-type factual knowledge, a Transformer encoder is used for feature extraction to capture the contextual information within sentences; for triple ontology knowledge, triples are transformed into graph structures, and graph neural networks are used to capture entity relationship information in triples.

[0008] S3. Input the inference knowledge feature vector, factual knowledge feature vector, and ontology graph feature vector into the feedforward network model of the pre-trained model BART based on the Transformer architecture. These feature vectors are used as additional inputs and are processed by the feedforward network together with the hidden representation of the BART encoder. Through the processing and generation of this model, an empathic response is generated.

[0009] In each layer of the BART (Bidirectional and Auto-Regressive Transformers) model's feedforward network module, inference knowledge feature vectors, factual knowledge feature vectors, and ontology graph feature vectors are injected. These feature vectors serve as additional inputs, processed along with the model's original inputs by the feedforward network. During subsequent fine-tuning, the parameter matrix in the feedforward network is continuously adjusted to enhance the model's utilization of the knowledge feature vectors.

[0010] Preferably, step 1 specifically includes:

[0011] S11. For phrase-based reasoning knowledge, input both the dialogue context D and the predefined common sense type r into the COMET knowledge base to obtain the reasoning knowledge set:

[0012] KG s =COMET(r,D)

[0013] Where r∈{xWant,xNeed,xIntent,xEffect,xReact}, xWant,xNeed,xIntent,xEffect,xReact are the edge types required for generating reasoning knowledge in the COMET knowledge base, with meanings of "someone thinks", "someone needs", "someone desires", "impact on someone", and "someone's reaction", respectively. KG s For each context, there are five categories of reasoning knowledge, each containing multiple phrases, such as: [[“become sociable”, “avoid conflict”, “become friendly”], [“talk about someone”, “get a hug”, “get a friend”], ..., [“happy”, “friendly”, “satisfied”]].

[0014] S12. For document-based factual knowledge, provide LLama2 with a specific prompt to guide the model in generating document-based knowledge relevant to context D:

[0015] KG d =llama2(P,D)

[0016] Where P represents the Prompt statement, which reads: "You act as a common sense knowledge base. I provide a dialogue, and you generate five document-based facts that are most relevant to the dialogue." KG d It is a collection of document-based factual knowledge. For each context, there are five document knowledge items, such as: ["Safe driving requires attention to the road and other vehicles", "You must be responsible for your actions when driving", "Running over others on the road is dangerous and may cause serious injury or death", "If someone is at fault in an accident, they should take responsibility and apologize", "Road safety must be prioritized and aggressive or reckless driving behavior should be avoided"];

[0017] S13. For triple-based ontology knowledge, firstly, stop words are removed from the dialogue context, then sentiment words in the context are selected using the NRC sentiment lexicon, and finally, sentiment-related triples from ConceptNet are selected to construct an ontology set:

[0018] KG t =ConceptNet(x)

[0019] Where x represents the contextual sentiment word, KG t This is a set of triples. Each triple is represented as (x, r, t), where x and t are entities, and r is an entity relation, such as ("birth", "related", "happy").

[0020] Preferably, step S2 specifically includes:

[0021] S21. For phrase-based reasoning knowledge, KG s Each type of reasoning knowledge in the sequence is pieced together to form a knowledge sequence CS. s Subsequently, CS was obtained through word embedding technology. s Corresponding word embedding vector Finally, the average word embedding vector yields the feature representation of the reasoning knowledge.

[0022]

[0023] S22. For document-based factual knowledge, KG d Each document in the document is preceded by a start marker [CLS], forming a sequence CS. dNext, the word embedding vectors corresponding to these sequences are input into the Transformer encoder to extract the hidden vector corresponding to the last layer [CLS] position, which is used as the feature vector corresponding to the factual knowledge.

[0024]

[0025] Enc stands for BART encoder; E W [0] represents the word embedding layer; [0] represents the indexing operation, indicating the extraction of the hidden representation corresponding to the [CLS] position;

[0026] S23. For triple ontology knowledge, since triples mainly represent the relationships between entities, we consider transforming them into graph networks and using graph neural networks to capture the entity relationship information in triples.

[0027] First, we need to construct a graph G = (V, E), where V is the node set and E is the node relation. To do this, we treat the entities in the triples as nodes and the entity relations as node relations. Next, we construct an adjacency matrix A using the node sets and node relations. The values ​​in A are either 1 or 0, for example, A0 = 0. ij =1 indicates that there is an edge between node i and node j. To avoid losing its own information during feature extraction, the diagonal lines of the adjacency matrix are set to 1, resulting in the self-adjacency matrix.

[0028] However, the adjacency matrix alone cannot fully represent the graph because each node has its own unique features. Therefore, we use the word vectors corresponding to each node to represent those features, resulting in the feature matrix H corresponding to the adjacency matrix. Then, the feature representation of each node is calculated using the following formula:

[0029]

[0030] in Let W be the degree matrix of the self-connected adjacency matrix, W be the trainable weight matrix, Wl represent the trainable parameter matrix of the l-th layer of the GCN, and Hl represent the degree matrix of the self-connected adjacency matrix. l The features extracted after passing through the l-th GCN layer are σ, which represents a non-linear activation function, such as ReLU.

[0031] Finally, to obtain the features of the entire graph, MaxPooling is performed on the node features extracted from the last layer of GCN to obtain the ontology graph features.

[0032]

[0033] Preferably, step S3 specifically includes:

[0034] S31. First, the extracted reasoning knowledge feature vector, factual knowledge feature vector, and ontology graph feature vector are concatenated to form a comprehensive knowledge vector H. CS :

[0035]

[0036] Next, vector H CS With two trainable parameter matrices W respectively k and W v Multiplying them yields two new eigenvectors H. k and H v :

[0037] H k =W k H CS

[0038] H v =W v H CS

[0039] During the fine-tuning process, matrix W k and W v It is constantly being adjusted to emphasize the importance of different types of knowledge.

[0040] Then, the obtained H k and H v The following calculations are performed on the hidden states calculated using self-attention:

[0041] FFN(H l )=f(H l .[H k :K l ]).[H v :V l ]

[0042] Where FFN is the feedforward neural network of the BART encoder, Kl and Vl are the trainable parameters in the l-th layer of the feedforward neural network of the BART encoder, l represents the l-th layer of the BART encoder, and H l For the hidden state calculated by the l-th self-attention module, K and V are the parameter matrices in the original BARTLayer. The calculation results are passed to the upper layers for similar calculations until the hidden state H of the last layer of the BART encoder is reached. ctx It was calculated.

[0043] S32. To enhance the model's emotion perception ability, H is used. ctx The hidden vector corresponding to the [CLS] position represents the entire context, and is then fed into a fully connected layer. The probability distribution P of the contextual sentiment is further calculated using the softmax function.emo :

[0044] P emo =Softmax(W e H ctx [0])

[0045] During the fine-tuning process, this probability distribution P is continuously updated. emo This approximates the true probability distribution of emotions, thereby accurately perceiving the emotions expressed in the context of the dialogue. We represents the trainable parameter matrix used to train H... ctx [0] Perform a linear transformation.

[0046] S33, Set the hidden layer state H of the last layer of the BART encoder. ctx The input is fed into the BART decoder to generate an empathetic response.

[0047] This invention also provides a multi-turn empathic dialogue generation device that integrates multi-source common-sense knowledge, comprising:

[0048] The knowledge acquisition module uses historical multi-turn dialogue content as dialogue context and inputs it into COMET, LLama2 and ConceptNet respectively to obtain three types of common sense knowledge: phrase-type reasoning knowledge, document-type factual knowledge and triple-type ontology knowledge.

[0049] The feature vector generation module uses three encoding methods to transform three different forms of common sense knowledge into corresponding feature vectors for subsequent injection into the pre-trained model. For phrase-type reasoning knowledge, word embedding technology is used to transform the semantic information of words or phrases into vectors with fixed dimensions. For document-type factual knowledge, a Transformer encoder is used for feature extraction to capture the contextual information within sentences. For triple ontology knowledge, triples are transformed into graph structures, and graph neural networks are used to capture entity relationship information within triples.

[0050] The empathic response generation module inputs the inference knowledge feature vector, factual knowledge feature vector, and ontology graph feature vector into the feedforward network model of the pre-trained BART model based on the Transformer architecture. These feature vectors, as additional inputs, are processed by the feedforward network together with the hidden representations of the BART encoder. Through the processing and generation of this model, an empathic response is generated.

[0051] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor runs the computer program, it performs the steps of the multi-turn empathic dialogue generation method integrating multi-source common sense knowledge as described above.

[0052] The present invention also provides a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the multi-turn empathic dialogue generation method integrating multi-source common sense knowledge as described above.

[0053] The present invention also provides a computer program product, comprising a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the multi-turn empathic dialogue generation method integrating multi-source common sense knowledge as described above.

[0054] In summary, compared with the prior art, the present invention has the following beneficial technical effects:

[0055] 1. This invention provides a multi-turn empathic dialogue generation method that integrates multi-source common sense knowledge. The core purpose of this method is to achieve effective fusion of multi-source knowledge and inject it into a pre-trained dialogue model. Through this innovative method, the shortcomings of current dialogue models that rely on single-source knowledge in terms of empathic ability are solved, the problem of poor scalability of existing dialogue generation models in the face of multi-source knowledge is improved, and the scalability and flexibility of the model in the face of various knowledge representation forms are significantly improved.

[0056] 2. The present invention provides a method for generating multi-turn empathetic dialogues that integrates multi-source common sense knowledge. This method incorporates knowledge into the feedforward network layer of a pre-trained model. This method only requires adding a small number of parameters to effectively inject common sense knowledge into the pre-trained model and participate in fine-tuning.

[0057] 3. The present invention provides a multi-turn empathic dialogue generation method that integrates multi-source common sense knowledge. The implementation of this method will greatly improve the ability of dialogue systems to understand and generate responses, especially when dealing with complex and multi-dimensional emotional communication scenarios. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of the multi-turn empathic dialogue generation method that integrates multi-source common sense knowledge provided by the present invention;

[0059] Figure 2 This is a schematic diagram illustrating the principle of the multi-turn empathic dialogue generation model that integrates multi-source common sense knowledge provided by the present invention.

[0060] Figure 3 This is a schematic diagram of the empathic response generated by the multi-turn empathic dialogue generation model that integrates multi-source common sense knowledge provided by the present invention. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0062] Figure 1 This is a flowchart illustrating the multi-turn empathic dialogue generation method integrating multi-source common sense knowledge as described in this invention. The method includes the following steps:

[0063] S1. Use the content of historical multi-turn dialogues as the dialogue context and input them into the COMET, LLama2 and ConceptNet knowledge bases respectively to obtain three types of common sense knowledge: phrase-type reasoning knowledge, document-type factual knowledge and triple-type ontology knowledge;

[0064] S2. The obtained text knowledge is encoded using word embedding technology, Transformer encoder and graph neural network GCN respectively to generate corresponding inference knowledge feature vector, fact knowledge feature vector and ontology graph feature vector;

[0065] S3. Input the encoded knowledge vector and dialogue context into the pre-trained model. Through the processing and generation of the model, an empathetic response is generated.

[0066] Figure 2 A schematic diagram illustrating the principle of a multi-turn empathic dialogue generation model that integrates multi-source common-sense knowledge for the invention:

[0067] S1. Input the dialogue context into COMET, LLama2 and ConceptNet respectively to obtain phrase-type reasoning knowledge, document-type fact knowledge and triple-type ontology knowledge;

[0068] S2. For different forms of knowledge, three specific encoding methods are used to transform textual knowledge into feature vectors for subsequent injection into the pre-trained model. For phrase-type reasoning knowledge, word embedding technology is used to transform the semantic information of words or phrases into vectors with fixed dimensions; for document-type factual knowledge, a Transformer encoder is used for feature extraction to capture the contextual information within sentences; for triple ontology knowledge, triples are transformed into graph structures, and graph neural networks are used to capture entity relationship information in triples.

[0069] S3. In each layer of the feedforward network module of the BART model, inference knowledge feature vectors, factual knowledge feature vectors, and ontology graph feature vectors are injected. These feature vectors serve as additional inputs and are processed by the feedforward network along with the hidden representations of the BART encoder. During subsequent fine-tuning, the parameter matrix in the feedforward network is continuously adjusted to enhance the model's utilization of knowledge feature vectors.

[0070] Step S1 includes:

[0071] S11. For phrase-based reasoning knowledge, input both the dialogue context D and the predefined common sense type r into the COMET knowledge base to obtain a phrase knowledge set:

[0072] KG s =COMET(r,D)

[0073] Where r∈{xWant,xNeed,xIntent,xEffect,xReact}, xWant,xNeed,xIntent,xEffect,xReact are the edge types required for generating reasoning knowledge in the COMET knowledge base, with meanings of "someone thinks", "someone needs", "someone desires", "impact on someone", and "someone's reaction", respectively. KG s For each context, there are five categories of reasoning knowledge, each containing multiple phrases, such as: [[“become sociable”, “avoid conflict”, “become friendly”], [“talk about someone”, “get a hug”, “get a friend”], ..., [“happy”, “friendly”, “satisfied”]].

[0074] S12. For document-based factual knowledge, provide LLama2 with specific prompts to guide the model in generating document-based factual knowledge relevant to context D:

[0075] KG d =llama2(P,D)

[0076] Where P represents the Prompt statement, which reads: "You act as a common sense knowledge base. I provide a dialogue, and you generate five document-based facts that are most relevant to the dialogue." KG d For each context, there are five pieces of factual knowledge, such as: ["Safe driving requires attention to the road and other vehicles", "You must be responsible for your actions when driving", "Running over others on the road is dangerous and may cause serious injury or death", "If someone is at fault in an accident, they should take responsibility and apologize", "Road safety must be prioritized and aggressive or reckless driving behavior should be avoided"];

[0077] S13. For triple-based ontology knowledge, firstly, stop words are removed from the dialogue context. Then, sentiment words in the context are selected using the NRC sentiment lexicon. Finally, sentiment-related triples from ConceptNet are filtered to construct a tuple set.

[0078] KG t =ConceptNet(x)

[0079] Where x represents the contextual sentiment word, KG tThis is a set of triples. Each triple is represented as (x, r, t), where x and t are entities, and r is an entity relation, such as ("birth", "related", "happy").

[0080] Step S2 includes:

[0081] S21. For phrase-based reasoning knowledge, KG s Each type of reasoning knowledge in the sequence is pieced together to form a knowledge sequence CS. s Subsequently, CS was obtained through word embedding technology. s Corresponding word embedding vector Finally, the average word embedding vector yields the feature representation of phrase knowledge.

[0082]

[0083] S22. For document-based factual knowledge, KG d Each document in the document is preceded by a start marker [CLS], forming a sequence CS. d Next, the word embedding vectors corresponding to these sequences are input into the Transformer encoder to extract the hidden vector corresponding to the last layer [CLS] position, which is used as the feature vector corresponding to the document knowledge.

[0084]

[0085] Enc stands for BART encoder; E W [0] represents the word embedding layer; [0] represents the indexing operation, indicating the extraction of the hidden representation corresponding to the [CLS] position;

[0086] S23. For triple ontology knowledge, since triples mainly represent the relationships between entities, we consider transforming them into graph networks and using graph neural networks to capture the entity relationship information in triples.

[0087] First, we need to construct a graph G = (V, E), where V is the node set and E is the node relation. To do this, we treat the entities in the triples as nodes and the entity relations as node relations. Next, we construct an adjacency matrix A using the node sets and node relations. The values ​​in A are either 1 or 0, for example, A0 = 0. ij =1 indicates that there is an edge between node i and node j. To avoid losing its own information during feature extraction, the diagonal lines of the adjacency matrix are set to 1, resulting in the self-adjacency matrix.

[0088] However, the adjacency matrix alone cannot fully represent the graph because each node has its own unique features. Therefore, we use the word vectors corresponding to each node to represent those features, resulting in the feature matrix H corresponding to the adjacency matrix. Then, the feature representation of each node is calculated using the following formula:

[0089]

[0090] in Let W be the degree matrix of the self-connected adjacency matrix, W be the trainable weight matrix, Wl represent the trainable parameter matrix of the l-th layer of the GCN, and Hl represent the degree matrix of the self-connected adjacency matrix. l These are the features extracted after passing through the l-th GCN layer;

[0091] Finally, to obtain the features of the entire graph, MaxPooling is performed on the node features extracted from the last layer of GCN to obtain the ontology graph features.

[0092]

[0093] Step S3 includes:

[0094] S31. First, the extracted reasoning knowledge feature vector, factual knowledge feature vector, and ontology graph feature vector are concatenated to form a comprehensive knowledge vector H. CS :

[0095]

[0096] Next, vector H CS With two trainable parameter matrices W respectively k and W v Multiplying them yields two new eigenvectors H. k and H v :

[0097] H k =W k H CS

[0098] H v =W v H CS

[0099] During the fine-tuning process, matrix W k and W v It is constantly being adjusted to emphasize the importance of different types of knowledge.

[0100] Then, the obtained H k and H v The following calculations are performed on the hidden states calculated using self-attention:

[0101] FFN(H l )=f(H l .[H k :K l ]).[H v :V l]

[0102] Where FFN is the feedforward neural network of the BART encoder, Kl and Vl are the trainable parameters in the feedforward neural network of the l-th layer of the BART encoder, l represents the l-th layer of the BART encoder, and H l For the hidden state calculated by the l-th layer self-attention module, K and V are the parameter matrices in the original BARTLayer. The calculation results are passed to the upper layers for similar calculations until the hidden state H of the last layer of the BART encoder. ctx It was calculated.

[0103] S32. To enhance the model's emotion perception ability, H is used. ctx The hidden vector corresponding to the [CLS] position represents the entire context, and is then fed into a fully connected layer. The probability distribution P of the contextual sentiment is further calculated using the softmax function. emo :

[0104] P emo =Softmax(W e H ctx [0])

[0105] During the fine-tuning process, this probability distribution P is continuously updated. emo This approximates the true probability distribution of emotions, thereby accurately perceiving the emotions expressed in the context of the dialogue. We represents the trainable parameter matrix used to train H... ctx [0] Perform a linear transformation.

[0106] S33, Set the hidden layer state H of the last layer of the BART encoder. ctx The input is fed into the BART decoder to generate an empathetic response.

[0107] Figure 3 This diagram illustrates the empathic response generated by a multi-turn empathic dialogue generation model that integrates multi-source common-sense knowledge, as described in this invention. The diagram reveals how the model effectively infers the user's emotional state from common-sense knowledge when generating empathic responses, providing reasonable suggestions and solutions while comforting the user. This result clearly demonstrates that the present technical solution achieves its intended effect: enhancing the empathic understanding and empathic response generation capabilities of the dialogue model using multi-source common-sense knowledge.

[0108] This invention also provides a multi-turn empathic dialogue generation device that integrates multi-source common-sense knowledge, comprising:

[0109] The knowledge acquisition module uses historical multi-turn dialogue content as dialogue context and inputs it into COMET, LLama2 and ConceptNet respectively to obtain three types of common sense knowledge: phrase-type reasoning knowledge, document-type factual knowledge and triple-type ontology knowledge.

[0110] The feature vector generation module uses three encoding methods to transform three different forms of common sense knowledge into corresponding feature vectors for subsequent injection into the pre-trained model. For phrase-type reasoning knowledge, word embedding technology is used to transform the semantic information of words or phrases into vectors with fixed dimensions. For document-type factual knowledge, a Transformer encoder is used for feature extraction to capture the contextual information within sentences. For triple ontology knowledge, triples are transformed into graph structures, and graph neural networks are used to capture entity relationship information within triples.

[0111] The empathic response generation module inputs the inference knowledge feature vector, factual knowledge feature vector, and ontology graph feature vector into the feedforward network model of the pre-trained BART model based on the Transformer architecture. These feature vectors, as additional inputs, are processed by the feedforward network together with the hidden representations of the BART encoder. Through the processing and generation of this model, an empathic response is generated.

[0112] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor runs the computer program, it performs the steps of the multi-turn empathic dialogue generation method integrating multi-source common sense knowledge as described above.

[0113] The present invention also provides a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the multi-turn empathic dialogue generation method integrating multi-source common sense knowledge as described above.

[0114] The present invention also provides a computer program product, comprising a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the multi-turn empathic dialogue generation method integrating multi-source common sense knowledge as described above.

[0115] The above description is merely a preferred embodiment of the present invention and is not intended to limit this application. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating multi-turn empathetic dialogues that integrates multi-source common-sense knowledge, characterized in that, Includes the following steps: S1. Use the historical multi-turn dialogue content as the dialogue context and input it into COMET, LLama2 and ConceptNet respectively to obtain three types of common sense knowledge: phrase-type reasoning knowledge, document-type factual knowledge and triple-type ontology knowledge. S2. For different forms of common sense knowledge, three encoding methods are used to transform three different types of common sense knowledge into corresponding feature vectors for subsequent injection into the pre-trained model: For phrase-type reasoning knowledge, word embedding technology is used to transform the semantic information of words or phrases into vectors with fixed dimensions; for document-type fact knowledge, a Transformer encoder is used for feature extraction to capture the contextual information within the sentence; for triple ontology knowledge, triples are transformed into graph structures, and graph neural networks are used to capture entity relationship information in triples. S3. Input the inference knowledge feature vector, factual knowledge feature vector, and ontology graph feature vector into the feedforward network model of the pre-trained model BART based on the Transformer architecture. These feature vectors are used as additional inputs and are processed by the feedforward network together with the hidden representation of the BART encoder. Through the processing and generation of this model, an empathic response is generated.

2. The method according to claim 1, characterized in that, Step S1 includes the following steps: S11. For phrase-based reasoning knowledge, input both the dialogue context D and the predefined common sense type r into the COMET knowledge base to obtain the reasoning knowledge set: KG s =COMET(r,D) Where r∈{xWant,xNeed,xIntent,xEffect,xReact}, KG s It is a set of reasoning knowledge. For each context, it contains the above five categories of reasoning knowledge, and each category contains multiple phrases. S12. For document-based factual knowledge, provide a prompt for the LLama2 large model to guide the model in generating document-based factual knowledge relevant to the context D: KG d =llama2(P,D) Where P represents the Prompt statement; KG d It is a document-based set of factual knowledge, with five factual knowledge points corresponding to each context; S13. For triple-based ontology knowledge, firstly, stop words are removed from the dialogue context, then sentiment words in the context are selected using the NRC sentiment lexicon, and finally, sentiment-related triples from ConceptNet are selected to construct an ontology set: KG t =ConceptNet(x) Where x represents the contextual sentiment word, KG t Construct an ontology set for triples, where each triple is represented as (x, r, t), x and t are entities, and r is an entity relation.

3. The method according to claim 2, characterized in that, Step S2 includes the following steps: S21. For phrase-based reasoning knowledge, KG s The words in each type of reasoning knowledge are concatenated to form a knowledge sequence CS. s Subsequently, CS was obtained through word embedding technology. s Corresponding word embedding vector Finally, the average word embedding vector yields the feature representation of the reasoning knowledge. S22. For document-based factual knowledge, KG d Each document in the document is preceded by a start marker [CLS], forming a sequence CS. d Next, the word embedding vectors corresponding to these sequences are input into the Transformer encoder to extract the hidden vector corresponding to the last layer [CLS] position, which serves as the feature vector corresponding to the factual knowledge. Enc stands for BART encoder; E W [0] represents the word embedding layer; [0] represents the indexing operation, indicating the extraction of the hidden representation corresponding to the [CLS] position; S23. For triple ontology knowledge, since triple ontology knowledge focuses on representing the relationship between entities, we consider transforming it into a graph network and using a graph neural network to capture the entity relationship information in the triple. First, we need to construct a graph G = (V, E), where V is the set of nodes and E is the node relationship; Therefore, the entities in the triples are treated as nodes, and the entity relations are treated as node relations; then, an adjacency matrix A is constructed using the node set and node relations, where the values ​​in A are either 1 or 0. ij =1 indicates that there is an edge between node i and node j; to avoid losing its own information during the feature extraction process, the diagonal lines of the adjacency matrix are set to 1, resulting in the self-adjacency matrix. However, the adjacency matrix alone cannot fully represent the graph because each node has its own features. Therefore, we use the word vectors corresponding to the nodes to represent the node features, obtaining the feature matrix H corresponding to the adjacency matrix. Then, we calculate the feature representation of each node using the following formula: in Let W be the degree matrix of the self-connected adjacency matrix, and W be the trainable weight matrix. l H represents the trainable parameter matrix of the l-th layer of GCN. l The features extracted after passing through the l-th GCN layer are σ, which represents the nonlinear activation function. Finally, to obtain the features of the entire graph, MaxPooling is performed on the node features extracted from the last layer of GCN to obtain the ontology graph features.

4. The method according to claim 3, characterized in that, Step S3 includes the following steps: S31. First, the extracted reasoning knowledge feature vector, factual knowledge feature vector, and ontology graph feature vector are concatenated to form a comprehensive knowledge vector H. CS : Next, vector H CS With two trainable parameter matrices W respectively k and W v Multiplying them yields two new eigenvectors H. k and H v : H k =W k H CS H v =W v H CS During the fine-tuning process, matrix W k and W v It is constantly being adjusted to emphasize the importance of different types of knowledge; Then, the obtained H k and H v The following calculations are performed on the hidden states calculated using self-attention: FFN(H l )=f(H l .[H k :TO l ]).[H v :IN l ] Where FFN is the feedforward neural network of the BART encoder, Kl and Vl are the trainable parameters in the feedforward neural network of the l-th layer of the BART encoder, l represents the l-th layer of the BART encoder, and H l For the hidden state calculated by the l-th layer self-attention module, K and V are the parameter matrices in the original BARTLayer. The calculation results are passed to the upper layers for similar calculations until the hidden state H of the last layer of the BART encoder. ctx Calculated; S32. To enhance the model's emotion perception ability, H is used. ctx The hidden vector corresponding to the [CLS] position represents the entire context, and is then fed into a fully connected layer. The probability distribution P of the contextual sentiment is further calculated using the softmax function. emo : P emo =Softmax(W e H ctx [0]) During the fine-tuning process, this probability distribution P is continuously updated. emo This makes it approximate the real emotion probability distribution, thereby accurately perceiving the emotions expressed in the dialogue context; We represents the trainable parameter matrix, used to train H ctx [0] Perform a linear transformation; S33. Finally, set the hidden layer state H of the last layer of the BART encoder. ctx The input is fed into the BART decoder to generate an empathetic response.

5. A multi-turn empathic dialogue generation device integrating multi-source common-sense knowledge, characterized in that, include: The knowledge acquisition module uses historical multi-turn dialogue content as dialogue context and inputs it into COMET, LLama2 and ConceptNet respectively to obtain three types of common sense knowledge: phrase-type reasoning knowledge, document-type factual knowledge and triple-type ontology knowledge. The feature vector generation module uses three encoding methods to transform three different forms of common sense knowledge into corresponding feature vectors for subsequent injection into the pre-trained model. For phrase-type reasoning knowledge, word embedding technology is used to transform the semantic information of words or phrases into vectors with fixed dimensions. For document-type factual knowledge, a Transformer encoder is used for feature extraction to capture the contextual information within sentences. For triple ontology knowledge, triples are transformed into graph structures, and graph neural networks are used to capture entity relationship information within triples. The empathic response generation module inputs the inference knowledge feature vector, factual knowledge feature vector, and ontology graph feature vector into the feedforward network model of the pre-trained BART model based on the Transformer architecture. These feature vectors, as additional inputs, are processed by the feedforward network together with the hidden representations of the BART encoder. Through the processing and generation of this model, an empathic response is generated.

6. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor runs the computer program, it performs the steps of the multi-turn empathic dialogue generation method integrating multi-source common-sense knowledge as described in any one of claims 1-4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-turn empathic dialogue generation method that integrates multi-source common sense knowledge as described in any one of claims 1-4.

8. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the multi-turn empathic dialogue generation method that integrates multi-source common-sense knowledge as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Conversation emotion recognition method based on common sense perception and hierarchical multi-task learning

    CN114722838A

  • Common-condition reply generation method and device, terminal and storage medium

    CN115934909A