Multi-modal dialogue generation method and device based on cognitive common situation and storage medium

Through a multimodal dialogue generation method based on cognitive empathy, the shortcomings of the existing technology in generating emotional resonance and context coherence dialogue are solved, and natural and empathetic text responses are realized, and the empathetic expression ability of the dialogue system is improved.

CN120146065AInactive Publication Date: 2025-06-13DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510631489.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is underperformance in generating conversations with emotional resonance and contextual coherence, especially in scenarios where it is necessary to understand the user's emotional state and respond empathically.

Method used

A multimodal dialogue generation method based on cognitive empathy is proposed. By obtaining historical speeches, determining input text sequences, performing embedding operations and context encoding, emotional and cognitive sequences are generated, and combining common sense knowledge and memetic data, the accuracy of the dialogue is calculated to generate the final conversation.

Benefits of technology

It realizes the generation of more natural and empathetic text responses, improves the empathetic expression ability of the dialogue system, and enhances the effects of dialogue empathetic expression and meme retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146065A_ABST
    Figure CN120146065A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal dialogue generation method and device based on cognitive emotion sharing and a storage medium, which utilize external common knowledge to deeply understand the situation and feeling of a user so as to generate a more natural and emotion sharing text response and accurately select an internet momentum conforming to the emotion of the generated text. According to the method, text and image information are effectively integrated, the aim of effectively integrating the common-situation expression in multi-modal dialogue generation is achieved by establishing a novel model framework and an evaluation method, the common-situation expression ability of a dialogue system is improved, and the remarkable advantages of the method in the aspects of enhancing the common-situation expression of the dialogue and memetic retrieval are displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dialogue generation, and in particular, to a multi-modal dialogue generation method, device, and storage medium based on cognitive empathy. Background Art

[0002] With the rapid development of artificial intelligence technology, the natural language processing (NLP) field has made remarkable progress in text generation, dialogue systems, etc. However, these methods often perform poorly in generating conversations with emotional resonance and context coherence, especially in scenarios where it is necessary to understand the user's emotional state and make an empathic response. The text generation method is an important branch in the natural language processing (NLP) field. Traditional text generation methods mainly include rule-based templates and statistic-based methods. The rule-based template method requires designing a lot of manual templates, which can generate accurate but limited coverage; while the statistic-based method builds a statistical model based on data, but cannot understand the semantic features behind words, and the generation system is complex, requiring a large amount of manual feature engineering. Without sufficient data, it usually cannot generate natural and interesting texts, and there are certain limitations in empathy. Summary of the Invention

[0003] Based on this, it is necessary to address the above problems and propose a multi-modal dialogue generation method, device, and storage medium based on cognitive empathy.

[0004] A multi-modal dialogue generation method based on cognitive empathy, where the dialogue consists of multiple texts or emoticons, each text contains at least one word, and the target text is generated from two aspects: emotion and words. The method includes: Obtain each historical speech, and determine the input text sequence according to the historical speech; determine the embedding sequence according to the input text sequence; Perform a context encoding operation on the embedding sequence to obtain a context representation; Add different relationship tags and COMET operations to the last speech in the historical speech to obtain the first sequence, the second sequence, the third sequence, the fourth sequence, and the fifth sequence; Determine the cognitive sequence according to the second sequence, the third sequence, the fourth sequence, and the fifth sequence; Determine the context representation of the emotion sequence according to the first sequence; determine the context representation of the cognitive sequence according to the cognitive sequence; Determine the emotion fusion representation according to the context representation and the context representation of the emotion sequence; Determine the cognitive fusion representation according to the context representation and the context representation of the cognitive sequence; Determine the probability of the emotion category distribution according to the emotion fusion representation and the cognitive fusion representation; Obtain the true label and determine the cross-entropy loss according to the true label and the sentiment category distribution; Obtain the text embedding sequence corresponding to the generated text, and determine the probability of the generated text according to the text embedding sequence and the combined context representation of common sense knowledge; Determine the word similarity between the words in the generated text and the target word according to the probability; Obtain the meme data and the predicted meme, and determine the meme similarity according to the meme data, the predicted meme, and the input text sequence; the meme data includes: the embedding vector of the meme corresponding to the pre-stored meme, the context marker of the word in the embedding sequence, the meme prediction sample, the total number of sentiment categories, and the meme category; Determine the accuracy of multiple target dialogues according to the word similarity and the meme similarity, and the target dialogue with the highest accuracy is the finally generated dialogue.

[0005] In one embodiment, The determining the embedding sequence according to the input text sequence includes: Perform word embedding operation on the input text sequence to obtain a word embedding sequence, perform position embedding operation on the input text sequence to obtain a position embedding sequence, and perform dialogue embedding operation on the input text sequence to obtain a dialogue embedding sequence; Determine the embedding sequence according to the word embedding sequence, the position embedding sequence, and the dialogue embedding sequence; The determining the cognitive sequence according to the second sequence, the third sequence, the fourth sequence, and the fifth sequence includes: Concatenate the second sequence, the third sequence, the fourth sequence, and the fifth sequence to obtain a cognitive sequence; The determining the context representation of the sentiment sequence according to the first sequence; the determining the context representation of the cognitive sequence according to the cognitive sequence includes: Perform sentiment sequence embedding operation on the first sequence to obtain a sentiment sequence embedding representation; perform sentiment encoding operation on the sentiment sequence embedding representation to obtain a context representation of the sentiment sequence; Perform embedding operation on the cognitive sequence to obtain a cognitive sequence embedding representation; perform cognitive encoding operation on the cognitive sequence embedding representation to obtain a context representation of the cognitive sequence.

[0006] In one embodiment, The determining the sentiment fusion representation according to the context representation and the context representation of the sentiment sequence includes: Perform average hidden operation on the context representation of the sentiment sequence to obtain an average hidden representation; Concatenate the context representation and the average hidden representation to obtain a sentiment fusion representation; Determining the cognitive fusion representation according to the context representation and the cognitive sequence context representation includes: Performing a classification and hiding operation on the cognitive sequence context representation to obtain a classification and hidden representation; Concatenating the context representation and the classification and hidden representation to obtain a cognitive fusion representation.

[0007] In one embodiment, Determining the probability of the emotion category distribution according to the emotion fusion representation and the cognitive fusion representation; includes: Performing emotion refinement encoding on the emotion fusion representation to obtain an emotion refinement context representation; Performing cognitive refinement encoding on the cognitive fusion representation to obtain a cognitive refinement context representation; Concatenating a plurality of the cognitive refinement context representations to obtain a cognitive refinement context concatenated representation; Concatenating the cognitive refinement context concatenated representation and the emotion refinement context representation to obtain a common sense refinement context representation; Determining a common sense knowledge combination context representation according to the common sense refinement context representation; Determining an emotion refinement hidden representation according to the emotion refinement context representation; Determining the probability of the emotion category distribution according to the emotion refinement hidden representation.

[0008] In one embodiment, the multi-modal dialogue generation method based on cognitive empathy, Is implemented through the following expression: =E 1 +E 2 +E 3 Determining the cognitive sequence according to the second sequence, the third sequence, the fourth sequence, and the fifth sequence is implemented through the following expression: Determining the emotion sequence context representation according to the first sequence; determining the cognitive sequence context representation according to the cognitive sequence is implemented through the following expression: Wherein, Is each historical speech, C is the input text sequence, E 1 Is the word embedding sequence, E 2 Is the position embedding sequence, E 3is a dialogue embedding sequence, is an embedding sequence, is a context representation, is a context encoder, is a cognitive sequence, is a first sequence, is a second sequence, is a third sequence, is a fourth sequence, is a fifth sequence; is an emotional sequence context representation, is an embedding representation of the emotional sequence, is an emotion encoder, is a cognitive sequence context representation, is an embedding representation of the cognitive sequence, is a cognitive encoder.

[0009] In one embodiment, The average hidden operation on the emotional sequence context representation to obtain the average hidden representation is achieved by the following expression: Where, is the average hidden representation, is the emotional sequence context representation; The classification hidden operation on the cognitive sequence context representation to obtain the classification hidden representation is achieved by the following expression: Where, is the cognitive sequence context representation, is the classification hidden representation; The concatenation of the context representation and the average hidden representation to obtain the emotional fusion representation is achieved by the following expression: Where, is the average hidden representation, is the context representation, is the emotional fusion representation; The concatenation of the context representation and the classification hidden representation to obtain the cognitive fusion representation is achieved by the following expression: Where, is the classification hidden representation, is the context representation, is the cognitive fusion representation.

[0010] In one embodiment, Performing emotion refinement encoding on the emotion fusion representation to obtain an emotion refinement context representation; performing cognitive refinement encoding on the cognitive fusion representation to obtain a cognitive refinement context representation is achieved through the following expressions: where, is the emotion refinement context representation, is the cognitive refinement context representation, is the emotion refinement encoder, is the cognitive refinement encoder, is the emotion fusion representation, is the cognitive fusion representation; Performing concatenation on multiple cognitive refinement context representations to obtain a concatenated cognitive refinement context representation; concatenating the concatenated cognitive refinement context representation with the emotion refinement context representation to obtain a common sense refinement context representation is achieved through the following expressions: where, is the concatenated cognitive refinement context representation, , is the common sense refinement context representation, is the emotion refinement context representation.

[0011] In one embodiment, Determining a common sense knowledge combination context representation according to the common sense refinement context representation; determining an emotion refinement hidden representation according to the emotion refinement context representation; determining the probability of an emotion category distribution through the following expressions: where, ∈ is the common sense knowledge combination context representation, is the importance score, is the common sense refinement context representation, is the perception operation, ⊙ represents element-wise multiplication, ∈ is the emotion refinement hidden representation; is the emotion refinement context representation; ∈ is the probability of the emotion category distribution, ∈ is the weight vector of the linear layer.

[0012] A computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps: Obtain each historical speech and determine an input text sequence according to the historical speech; determine an embedding sequence according to the input text sequence; Perform a context encoding operation on the embedding sequence to obtain a context representation; Add different relationship tags and COMET operations to the last speech in the historical speech respectively to obtain a first sequence, a second sequence, a third sequence, a fourth sequence, and a fifth sequence; Determine a cognitive sequence according to the second sequence, the third sequence, the fourth sequence, and the fifth sequence; Determine an emotional sequence context representation according to the first sequence; determine a cognitive sequence context representation according to the cognitive sequence; Determine an emotional fusion representation according to the context representation and the emotional sequence context representation; Determine a cognitive fusion representation according to the context representation and the cognitive sequence context representation; Determine the probabilities of the emotional category distribution according to the emotional fusion representation and the cognitive fusion representation; Obtain a true label and determine a cross-entropy loss according to the true label and the emotional category distribution; Obtain a text embedding sequence corresponding to the generated text, and determine the probability of the generated text according to the text embedding sequence and the common sense knowledge combined context representation; Determine the word similarity between the words in the generated text and the target word according to the probability; Obtain meme data and a predicted meme, and determine a meme similarity according to the meme data, the predicted meme, and the input text sequence; the meme data includes: an embedding vector of the meme corresponding to the pre-stored meme, a context tag of the word in the embedding sequence, a meme prediction sample, the total number of emotional categories, and a meme category; Determine the accuracy of the target dialogue according to the word similarity and the meme similarity.

[0013] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor performs the following steps: Obtain each historical speech and determine an input text sequence according to the historical speech; determine an embedding sequence according to the input text sequence; Perform a context encoding operation on the embedding sequence to obtain a context representation; Add different relationship tags and COMET operations to the last speech in the historical speech respectively to obtain the first sequence, the second sequence, the third sequence, the fourth sequence and the fifth sequence; Determine the cognitive sequence according to the second sequence, the third sequence, the fourth sequence and the fifth sequence; Determine the context representation of the emotion sequence according to the first sequence; Determine the context representation of the cognitive sequence according to the cognitive sequence; Determine the emotion fusion representation according to the context representation and the context representation of the emotion sequence; Determine the cognitive fusion representation according to the context representation and the context representation of the cognitive sequence; Determine the probability of the emotion category distribution according to the emotion fusion representation and the cognitive fusion representation; Obtain the true label, and determine the cross-entropy loss according to the true label and the emotion category distribution; Obtain the text embedding sequence corresponding to the generated text, and determine the probability of the generated text according to the text embedding sequence and the context representation combined with common sense knowledge; Determine the word similarity between the words in the generated text and the target word according to the probability; Obtain the meme data and the predicted meme, and determine the meme similarity according to the meme data, the predicted meme and the input text sequence; The meme data includes: the embedding vector of the meme corresponding to the pre-stored meme, the context label of the word in the embedding sequence, the meme prediction sample, the total number of emotion categories, the meme category; Determine the accuracy of the target dialogue according to the word similarity and the meme similarity.

[0014] The present invention utilizes external common sense knowledge to deeply understand the user's situation and feelings, so as to generate more natural and empathetic text responses, and accurately select Internet memes that match the emotion of the generated text. It effectively integrates text and image information, and realizes the goal of effectively integrating empathetic expression in multi-modal dialogue generation by establishing a novel model framework and evaluation method, improves the empathetic expression ability of the dialogue system, and also shows its significant advantages in enhancing dialogue empathetic representation and meme retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Among them: Figure 1 Flowchart of a multimodal dialogue generation method based on cognitive empathy in an embodiment; Figure 2 ECME overall multitask training framework in an embodiment; Figure 3 Cognitive empathy dialogue example in an embodiment; Figure 4 Multimodal dialogue example in the MOD dataset in an embodiment; Figure 5 A successful case of meme retrieval in an embodiment; Figure 6 A failed case of meme retrieval in an embodiment; Figure 7 An appropriate case of meme retrieval in an embodiment; Figure 8 Block diagram of a computer device in an embodiment. Detailed implementation manners

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] To solve the technical problems in the background art, as Figure 1 shown, in an embodiment, a multimodal dialogue generation method based on cognitive empathy is provided. The dialogue is composed of multiple texts or emoticons, and each text contains at least one word. The target text is generated from two aspects: emotion and words. This method can be applied to both terminals and servers. In this embodiment, an example of applying it to a terminal is given. In the cognitive empathy stage, as Figure 3As shown in the figure, it is an example of a cognitive empathy dialogue. In the present invention, an embedding sequence of the input text is obtained by adding word embedding, position embedding, and dialogue state embedding, and is processed successively by a context encoder, a cognitive encoder, and an emotion encoder to obtain a context representation. In the emotional empathy stage, the present invention performs emotional classification on the context representation, generates an emotional category distribution through a linear layer and a softmax operation, and uses cross-entropy loss (CE) to optimize the weights. In the stage of generating the response text, the present invention generates the target response Y by generating the tokens of the decoder one by one. Secondly, in the meme retrieval process, it is judged whether the candidate meme is appropriate according to the dialogue context. The present invention connects the embedded dialogue context and the meme embedding and inputs them into the BERT model. At the same time, the present invention improves the model's understanding ability of multi-modal input through three auxiliary tasks: masked context prediction, meme emotion classification, and meme semantic prediction. In the masked context prediction task, the present invention uses the standard masked language modeling task (MLM), and at the same time introduces a fusion mechanism of meme embedding to prompt the model to better understand and reconstruct the context. In the meme emotion classification task, the present invention extracts the hidden state corresponding to the meme by a deep neural network and inputs it into an improved classifier to comprehensively consider the visual content analysis of the meme and the emotional connotation of the context, and designs a loss function based on the cross-entropy principle to quantify the model performance and guide the training process. In the meme semantic prediction task, the present invention designs an innovative semantic label prediction task to train the model and predict the semantic labels of the memes. As Figure 2 shown, it specifically includes the following steps: The multi-modal dialogue generation method based on cognitive empathy, as Figure 1 shown, specifically includes the following steps: S10: Obtain each historical speech, and determine an input text sequence according to the historical speech; determine an embedding sequence according to the input text sequence; S20: Perform a context encoding operation on the embedding sequence to obtain a context representation; S30: Add different relationship tags and COMET operations to the last speech in the historical speech respectively to obtain the first sequence, the second sequence, the third sequence, the fourth sequence, and the fifth sequence; S40: Determine a cognitive sequence according to the second sequence, the third sequence, the fourth sequence, and the fifth sequence; S50: Determine an emotional sequence context representation according to the first sequence; determine a cognitive sequence context representation according to the cognitive sequence; S60: Determine an emotional fusion representation according to the context representation and the emotional sequence context representation; S70: Determine a cognitive fusion representation according to the context representation and the cognitive sequence context representation; S80: Determine the probability of the emotional category distribution based on the emotional fusion representation and the cognitive fusion representation; S90: Obtain the true label, and determine the cross-entropy loss based on the true label and the emotional category distribution; S100: Obtain the text embedding sequence corresponding to the generated text, and determine the probability of the generated text based on the text embedding sequence and the context representation determined by combining common sense knowledge; S110: Determine the word similarity between the words in the generated text and the target word based on the probability; S120: Obtain meme data and predicted memes, and determine the meme similarity based on the meme data, the predicted memes, and the input text sequence; the meme data includes: the embedding vector of the meme corresponding to the pre-stored meme, the context tags of the words in the embedding sequence, meme prediction samples, the total number of emotional categories, and meme categories; S130: Determine the accuracy of multiple target dialogues based on the word similarity and the meme similarity, and the target dialogue with the highest accuracy is the finally generated dialogue.

[0019] In one embodiment, the determining the embedding sequence according to the input text sequence in step S10 includes: S101: Perform word embedding operation on the input text sequence to obtain a word embedding sequence, perform position embedding operation on the input text sequence to obtain a position embedding sequence, and perform dialogue embedding operation on the input text sequence to obtain a dialogue embedding sequence; S102: Determine the embedding sequence based on the word embedding sequence, the position embedding sequence, and the dialogue embedding sequence; The determining the cognitive sequence according to the second sequence, the third sequence, the fourth sequence, and the fifth sequence in step S40 includes: S401: Concatenate the second sequence, the third sequence, the fourth sequence, and the fifth sequence to obtain a cognitive sequence; The determining the emotional sequence context representation according to the first sequence in step S50; the determining the cognitive sequence context representation according to the cognitive sequence includes: S501: Perform emotional sequence embedding operation on the first sequence to obtain an emotional sequence embedding representation; perform emotional coding operation on the emotional sequence embedding representation to obtain an emotional sequence context representation; S502: Perform embedding operation on the cognitive sequence to obtain a cognitive sequence embedding representation; perform cognitive coding operation on the cognitive sequence embedding representation to obtain a cognitive sequence context representation.

[0020] In one embodiment, the determining the emotion fusion representation according to the context representation and the emotion sequence context representation in step S60 includes: S601: performing an average hidden operation on the emotion sequence context representation to obtain an average hidden representation; S602: concatenating the context representation and the average hidden representation to obtain an emotion fusion representation; The determining the cognitive fusion representation according to the context representation and the cognitive sequence context representation in step S70 includes: S701: performing a classification hidden operation ([CLS]) on the cognitive sequence context representation to obtain a classification hidden representation; S702: concatenating the context representation and the classification hidden representation to obtain a cognitive fusion representation.

[0021] In one embodiment, the determining the emotion category distribution probabilities according to the emotion fusion representation and the cognitive fusion representation in step S80 includes: S801: performing emotion refinement encoding on the emotion fusion representation to obtain an emotion refinement context representation; S802: performing cognitive refinement encoding on the cognitive fusion representation to obtain a cognitive refinement context representation; S803: concatenating a plurality of the cognitive refinement context representations to obtain a cognitive refinement context concatenated representation; S804: concatenating the cognitive refinement context concatenated representation and the emotion refinement context representation to obtain a common sense refinement context representation; S805: determining a common sense knowledge combination context representation according to the common sense refinement context representation; S806: determining an emotion refinement hidden representation according to the emotion refinement context representation; S807: determining the emotion category distribution probabilities according to the emotion refinement hidden representation.

[0022] Specifically, for steps S101 - S107: In the cognitive empathy task, the present invention represents the historical conversation as , where represents that the i-th utterance consists of T words. Connecting the utterances and adding a special token [CLS] in front to obtain the input text sequence C: (1) The present invention uses the final hidden representation of [CLS] as the representation of the entire input sequence, where is each historical utterance, C is the input text sequence, and ⊕ represents the concatenation operation; The input text sequence C undergoes word embedding to obtain the word embedding sequence E 1 , undergoes position embedding to obtain the position embedding sequence E 2 , undergoes dialogue embedding to obtain the dialogue embedding sequence E 3 , due to the historical dialogue which contains both user utterances and model utterances, the dialogue state embedding here can be used to distinguish between user and model utterances. Add the word embedding sequence E1, the position embedding sequence E 2 and the dialogue embedding sequence E 3 to obtain the embedding sequence of the input text sequence C and the context representation as shown in the following formula: = E 1 + E 2 + E 3 (2) (3) where, is the embedding sequence, is the embedding sequence of the context representation, L is the length of the sequence, and d is the hidden size of the context encoder; representing the input sequence from semantic information, position information, and context information; is the context encoder; Attach five relation tags ([xReact], [xWant], [xNeed], [xIntent], [xEffect]) to the last utterance in the input text sequence C and simultaneously use COMET to generate common sense judgments for each relation. After connection, obtain the cognitive sequence. For the cognitive sequence the calculation formula is as follows: (4) where, is the cognitive sequence, is the first sequence, is the second sequence, is the third sequence, is the fourth sequence, is the fifth sequence; Divide the relations into two groups: emotion and cognition. The mean of the hidden representation of [xReact] represents the emotion sequence, and the mean of the hidden representations of other relations represents the cognitive sequence. Input the embedding representations of the corresponding sequences into independent cognitive and emotion encoders: (5) (6) (7) Among them, is the emotional sequence context representation, is the embedding representation of the emotional sequence, is the emotional encoder, is the cognitive sequence context representation, is the embedding representation of the cognitive sequence, is the cognitive encoder.

[0023] Specifically, the specific implementation steps of step S601 are as follows: The average hidden operation on the emotional sequence context representation to obtain the average hidden representation is achieved through the following expression: (8) Among them, is the average hidden representation, is the emotional sequence context representation; The classification hidden operation on the cognitive sequence context representation to obtain the classification hidden representation is achieved through the following expression: (9) Among them, is the cognitive sequence context representation, is the classification hidden representation; The concatenation of the context representation and the average hidden representation to obtain the emotional fusion representation is achieved through the following expression: (10) Among them, is the average hidden representation, is the context representation, is the emotional fusion representation; The concatenation of the context representation and the classification hidden representation to obtain the cognitive fusion representation is achieved through the following expression: (11) Among them, is the classification hidden representation, is the context representation, is the cognitive fusion representation.

[0024] Specifically, in steps S801 and S802, the emotional refinement encoding of the emotional fusion representation to obtain the emotional refinement context representation; the cognitive refinement encoding of the cognitive fusion representation to obtain the cognitive refinement context representation is achieved through the following expressions: (12) (13) Among them, is the emotional refinement context representation, is the cognitive refinement context representation, is the emotional refinement encoder, is the cognitive refinement encoder, is the emotional fusion representation, is the cognitive fusion representation; The splicing of multiple cognitive refinement context representations to obtain a cognitive refinement context splicing representation; the splicing of the cognitive refinement context splicing representation and the emotional refinement context representation to obtain a common sense refinement context representation is realized through the following expression: (14) (15) Among them, is the cognitive refinement context splicing representation, , is the common sense refinement context representation, is the emotional refinement context representation.

[0025] Specifically, in steps S805, S806, and S807, determining the common sense knowledge combination context representation according to the common sense refinement context representation; determining the emotional refinement hidden representation according to the emotional refinement context representation; determining the probability of the emotional category distribution according to the emotional refinement hidden representation is realized through the following expression: Applying the Sigmoid function ( ) to measure the importance of each common sense refinement context representation for response generation, and then using the obtained common sense refinement context representation multiplied by the corresponding importance score , and finally passing the obtained representation through a multi-layer perceptron (MLP) with the ReLU activation function to learn how to mix common sense knowledge of different relationships into a combined context representation: (16) (17) (18) In the emotional empathy task, in the hypothesis of the present invention, an emotional label is provided for each dialogue , using the hidden representation of the [CLS] token of the emotionally refined context representation obtained in the first step ( ) for emotion classification. The present invention passes the emotional refinement hidden representation ( ) through a linear layer, and after a Softmax operation, generates an emotional category distribution ∈ where q is the number of available sentiment categories. Among them, ∈ is the combined context representation of common sense knowledge, is the importance score, is the refined context representation of common sense, is the perceptual operation, and ⊙ represents element-wise multiplication, ∈ is the refined hidden representation of sentiment; is the refined context representation of sentiment; ∈ is the probability distribution of sentiment categories, ∈ is the weight vector of the linear layer.

[0026] In one embodiment, for steps S90 - S130, the specific implementation steps are as follows: The present invention optimizes the weights by minimizing the cross - entropy loss (Cross - Entropy, CE) between the sentiment category distribution probability and the sentiment label ; (19) Among them, is the cross - entropy loss, and the weight matrix is optimized through backpropagation and gradient descent to maximize the prediction probability of the model for the true category, is the sentiment category distribution probability; is the sentiment label; The smaller it is, the smaller the sentiment gap between the text and the target text.

[0027] In the target response for the response text generation task, with a length of T, represents the T words in the target response. Given , the probability distribution of generating the next word is obtained by generating the decoder tokens one by one. Assuming that the training has been completed, the model will generate a word list based on the training data as a repository for word selection in subsequent text generation. Finally, the word with the highest probability in the word list is selected as as shown in formula (20): (20) Among them, t is a variable , , t starts from 1 and increments by 1 each time until T, and w 1 takes [BOS] as a common starting token as the start signal of the decoder, and then is obtained through formula (20), , Represents the embedded sequence of generated tokens. Meanwhile, the cross-attention output by the encoder is modified to calculate attention on the common-sense refined context representation that integrates context and common-sense reasoning information. is a word; The research on text generation of cognitive empathy adopts the standard Negative Log Likelihood (NLL) loss on the target response W, making the generated text and the target text as close as possible when training the model by punishing low-probability predictions: the smaller the approximation, the closer the predicted word is to the target word.

[0028] (21) The main task of meme retrieval is to judge whether a candidate meme is appropriate according to the dialogue context, such as Figures 5 - 7 , which are the successful cases, failure cases, and appropriate cases of meme retrieval respectively. To complete this task, this study concatenates the embedded dialogue context and the meme embedding as the input into the BERT model. Then a binary classification layer is applied on the hidden state of the [CLS] ( [CLS] is a special token added by BERT at the beginning of the input sequence) token. To enhance the model's ability to understand multimodal inputs, this study designs three auxiliary tasks: 1) Masked context prediction, which is used to improve the model's understanding of the dialogue context; 2) Meme sentiment classification, aiming to enable the model to better understand the sentiment of the meme; 3) Meme semantic prediction, which infuses the semantic information of the meme into the model.

[0029] Suppose there is a multi-turn dialogue context , a set of candidate memes , where represents the i-th candidate meme, represents the i-th utterance in the dialogue, N is the number of utterances in the dialogue, M represents the number of candidate memes. In this work, the present invention assumes that there is only one appropriate meme and and belong to the same speaker. The purpose is to train a model to be able to select the correct meme from all candidates S given the dialogue history U .

[0030] The empathy-based meme retrieval model described above includes three auxiliary tasks: masked context prediction, meme sentiment classification, meme semantic prediction, and a core binary classification task.

[0031] First, in the masked context detection task, the present invention combines the meme embedding vector closely related to the context with the text embedding vector Combined and input into the BERT model together to perform the masked language modeling (MLM) task, enhancing the effect of multimodal interaction between text and memes. Define the loss function of this task as: (22) where is the i-th masked context token in the text, , represents the first word in the k-th response denotes the set of all previous context tokens, m represents the embedding vector of the meme, and N is the total number of masked tokens; In the meme sentiment classification task, since the same meme may carry different sentiment meanings in different dialogue contexts, the present invention comprehensively understands the sentiment by using text and meme information while training the model. Capturing the deep features of the meme by a deep neural network is the sample O, and using a classifier with a softmax layer to map the features to sentiment categories to obtain ; Based on the principle of cross-entropy and combined with the softmax function, the loss function of this task is defined as: (23) where A is the total number of sentiment categories, and the deep features of the meme output by the deep neural network constitute the prediction sample O, used to indicate the category of the meme, is the true label of category a (if a = 1, it is the first in A), and when the prediction sample O (sample 1234) belongs to category a has a value of 1, is the softmax probability that the prediction sample O belongs to category a; The loss of meme sentiment classification.

[0032] In the meme semantic prediction task, the present invention designs an innovative semantic label prediction task to predict and utilize the semantic meaning of the meme to improve the performance of the meme retrieval task. In this task, the input of the model is modified, and a fixed-length [MASK] token sequence is inserted after the dialogue text (the dialogue sample containing both text C and the meme ) as the input of the model. Train the model to recover the most appropriate semantic label from the [MASK] hidden labels based on the context and meme content; define the loss function: (24) where is the i-th semantic token, is the total number of semantic tokens, is the probability that the model predicts the i-th token, given the context C and the meme Meanwhile, since there are no real semantic tags in the dataset, the present invention uses the text information extracted from memes by an OCR (Optical Character Recognition) tool as semantic tags. This loss function is only applicable to memes from which text has been successfully recognized by the OCR tool. For memes without text or with unrecognizable text, their training depends on the loss functions of other auxiliary tasks. In the core binary classification task, the present invention aims to judge the suitability of candidate memes for the dialogue context. The present invention regards each dialogue-meme pair in the dataset as a positive sample, and randomly samples an equal number of memes to form negative samples, simulating the selection of inappropriate memes. The binary label s represents the suitability of each meme in a specific dialogue context C. If the meme is suitable, i.e., a positive sample, then g = 1; conversely, if it is a negative sample, then g = 0. The goal of the model is to predict the correct label g. When performing meme retrieval, the larger it is, the more preferentially the meme is selected.

[0033] In addition, the present invention defines a cross-entropy loss function: (25) where is the probability that the model predicts the suitability of meme in the context ; The training of the meme retrieval model is completed by comprehensively considering the losses of all auxiliary tasks and the main task. The final loss function is the weighted sum of these losses, in the following form: (26) where λ1, λ2, and λ3 are hyperparameters, and the selection of the final hyperparameters is adjusted through experiments to obtain the optimal results; The total loss function of the present invention is as follows: (27) The meme retrieval task is regarded as an auxiliary task and is trained together with the main text generation task. As Figure 2 shown, it constitutes a multi-task learning framework. Among them, λ is a hyperparameter used to balance the contributions of the two loss functions.

[0034] The present invention utilizes external common sense knowledge to deeply understand the user's situation and feelings, so as to generate more natural and empathetic text responses, and accurately select Internet memes that match the emotion of the generated text. It effectively integrates text and image information, and through establishing a novel model framework and evaluation method, achieves the goal of effectively integrating empathetic expressions in multi-modal dialogue generation, improves the empathetic expression ability of the dialogue system, and also shows its significant advantages in enhancing dialogue empathy representation and meme retrieval.

[0035] Combined with the solution of the present invention, the experimental analysis is as follows: Pre-training tasks 1) Internet meme feature extraction: Most existing convolutional neural networks (CNNs), including EfficientNet, are built on real-world photos. Therefore, it is not feasible to directly apply these networks to Internet memes for feature extraction. In the dataset, each sticker is assigned an emoji tag indicating its general emotion. Considering the correlation between Internet memes and stickers, this study adopted a pre-training classification task to help the model effectively understand memes. Specifically, this study used the features output by the CNN to predict the emoji attached to the corresponding sticker. This study added an additional MLP layer for research on empathy-based multi-modal dialogue generation and used cross-entropy loss as the optimization function.

[0036] 2) Cross-modal emotion modeling: The initial parameters of the ECME model of the present invention are loaded from a model trained only on a Chinese text corpus, and this model lacks cross-modal knowledge construction. Therefore, this study utilized the additional emotion labels included in the dataset to introduce emotion analysis into the model to help the model better process Internet meme content. Specifically, given the dialogue history, the system aims to predict the emotion label when using an Internet meme in the last utterance, which can also be regarded as a classification problem. The present invention resampled the first 100 emotion annotations to avoid training bias.

[0037] Comparison models and evaluation metrics So far, the research on the MOD task is insufficient, and there are few existing comparison models. The following models are selected for comparison: SRS (meme response) uses the self-attention mechanism to learn the representation of each utterance in the text history and extracts the feature representation of the meme through a convolutional neural network (CNN). Further, SRS adopts a deep interaction network to simulate the complex dependencies between utterances and memes to accurately predict the target meme.

[0038] CDial-GPT (text response) is a 12-layer Transformer decoder based on DialoGPT, designed specifically for Chinese dialogue generation, capable of generating smooth and natural dialogue text.

[0039] MHERAD (text response and meme response) is a multimodal hierarchical encoder-decoder model that combines visual features. It integrates these features into a basic hierarchical recurrent network (HRED) and is optimized for task-oriented dialogue in the retail domain.

[0040] MOD-GPT (text response and meme response) is the baseline model for the MOD dataset. This model can generate conversations containing Internet memes within a simple and effective framework, as Figure 4 shown.

[0041] Since the output of the present invention can be pure text, pure meme, or a combination of both, text evaluation metrics and meme evaluation metrics are set: The present invention uses perplexity (PPL) and Distinct-n (Dist-n) as the main automatic evaluation metrics. These metrics are widely used in the field of natural language processing and can effectively evaluate the performance of the generation model. Perplexity is used to evaluate the overall quality of the generated response. Distinct-n is used to evaluate the diversity of the generation.

[0042] Perplexity is an important metric for measuring the quality of a language model. It reflects the confidence of the model in a given dataset. The higher the confidence, the lower the perplexity. The specific calculation formula is as follows: (28) where N is the number of words in the test dataset, is the probability of predicting the i-th word when the previous i - 1 words in the k-th response of the model are known, The lower the perplexity, the more accurate the model predicts the next word in the sequence, indicating that the overall quality of the response generated by the model is higher.

[0043] Distinct-n is an important metric for measuring the diversity of the generated response. This metric calculates the proportion of unique n-grams in the generated text. n-gram , which represents the words in the generated text. The value of Distinct-n reflects the diversity of the generated text. A high Distinct-n value indicates that there are more different n-grams in the generated text, reducing repetition and enhancing the richness and naturalness of the dialogue. Especially in a dialogue system, maintaining the diversity of the generated responses is crucial for avoiding monotonous and repetitive answers, which significantly improves the user experience. The commonly used values of n are 1 and 2, namely Distinct-1 and Distinct-2. The specific calculation formulas are as follows: (29) Among them, "Unique n-grams" represents the number of unique n-grams in the generated text, while "Total n-grams" represents the total number of all n-grams in the generated text.

[0044] The present invention uses as an evaluation metric The metric is a ranking-based evaluation method used to evaluate the prediction ability of a model in a given candidate set. The specific calculation process is as follows: For each test sample, first generate a set containing n candidate memes. Then, the model uses to score and rank these candidate memes, that is, calculate the probability of each meme being appropriate for the context. The higher the P value, the higher the ranking. If the true meme is among the top k after sorting, then this test sample is counted as a correct prediction. Finally, by counting the number of correct predictions in all test samples and calculating the proportion of the total number of test samples, the value is obtained, where k is set to 1, 2, 5, and n is set to 10. The specific representation is as follows: (30) Performance Evaluation 1) Text Evaluation: In the evaluation of text, the present invention uses two different schemes: automatic evaluation and manual evaluation. Table 1 summarizes the performance of the baseline model and the ECME model in text automatic evaluation. The ECME model benefits from the meme's empathy ability while introducing cognitive empathy, and its evaluation scores exceed all baseline models.

[0045] Table 1 ECME Text Automatic Evaluation Results In the manual evaluation, the present invention conducted pairwise preference tests based on aspects. That is, for a given context, the present invention paired the responses of the ECME model and the responses of the baseline model, and asked annotators to select better responses according to the context, coherence (Coh.), empathy (Emp.), and informativeness (Inf.). The present invention randomly selected 100 response pairs and assigned three human annotators to conduct annotation evaluations. The results are shown in Table 2.

[0046] Table 2 Manual Evaluation Results of ECME Texts In the table, κ represents the inter-annotator agreement measured by Fleiss’s kappa, where 0.4 < κ < 0.6 indicates moderate agreement. †,‡ represent significant improvements, with p-values < 0.1 / 0.05 (sign test).

[0047] ECME outperformed the baseline model in all three aspects of the manual evaluation. The ECME model performed excellently in terms of coherence, empathy, and informativeness. Compared with MHERAD, CDial-GPT, and MOD-GPT, the ECME model had more advantages in generating responses. Examples of the texts generated by the ECME model and the baseline model are shown in Table 3.

[0048] 2) Meme evaluation: Since the generated text involves pure text, pure memes, and combinations of text and memes, a prediction task for whether to use memes is required. In this task, the accuracy of binary classification is 88%, which also indicates that the ECME model performs well in judging whether to use memes.

[0049] Secondly, there is the meme retrieval task when using memes. The present invention only considers conversations ending with Internet meme responses. The model will obtain n Internet memes, where only one is the real meme and the others are driver samples. The model needs to rank the memes according to the relevance of the multimodal context and then determine whether the real meme is within the top k memes. The evaluation results are shown in Table 4: Table 4 Automatic Evaluation Results of ECME Memes Among them, ctx, emo, and sem correspond to the three auxiliary tasks of masked context prediction, meme emotion classification, and meme semantic prediction in the meme retrieval task respectively. The complete model of the present invention (ECME + ctx + emo + sem) outperformed all baseline models on the test set and achieved the best performance in almost all settings.

[0050] At the same time, the experimental results also show that MOD-GPT is a powerful baseline model with good generalization ability on the test set. However, the complete model of the present invention can outperform MOD-GPT through multi-task learning. It is also concluded that the multi-task learning method of the present invention can improve sticker selection by explicitly guiding the model to understand multi-modal information.

[0051] 3) Ablation study: The present invention also conducted ablation experiments to verify the effects of each auxiliary task. As shown in Table 4, as each auxiliary training task is added to the ECME model, the performance of the model gradually improves, which also verifies the effectiveness of the auxiliary task settings in the present invention. At the same time, it is concluded that introducing semantic information can improve the generalization ability of the model. It is also found that the complete model of the present invention achieved an accuracy of 60% on the validation set in the auxiliary sticker sentiment classification task, where there are a total of 52 sentiment labels, which also confirms that the model of this study can learn from auxiliary tasks.

[0052] When using ablation experiments to verify the effectiveness of components, the present invention designed two variants of the model: w / o Aff: Removed the sentiment and sentiment modifier encoders (Equations 5, 12), ignored the sentiment representation in the commonsense modifier representation (Equation 15), and used the hidden representation of the [CLS] token in the encoding context for sentiment classification (Equation 3); w / o Cog: Removed the cognitive and cognitive modifier encoders (Equations 6 and 13), ignored the cognitive representation in the commonsense modifier representation (Equation 15), and replaced the MLP with a linear layer (Equation 12).

[0053] The experimental results shown in Table 1 indicate that sentiment and cognitive information have a significant impact on sentiment classification accuracy, which shows that correctly identifying the user's emotions and situational information is necessary for accurately judging their feelings. In addition, this study also observed that removing the diversity loss leads to a lower Dist-n score, which shows that this loss is effective in generating more diverse responses.

[0054] The present application also provides a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps: Obtain each historical speech and determine an input text sequence according to the historical speech; determine an embedding sequence according to the input text sequence; Perform a context encoding operation on the embedding sequence to obtain a context representation; Add different relationship tags and COMET operations to the last speech in the historical speech respectively to obtain the first sequence, the second sequence, the third sequence, the fourth sequence, and the fifth sequence; Determine a cognitive sequence according to the second sequence, the third sequence, the fourth sequence, and the fifth sequence; Determine a context representation of an emotion sequence according to the first sequence; determine a context representation of a cognitive sequence according to the cognitive sequence; Determine an emotion fusion representation according to the context representation and the context representation of the emotion sequence; Determine a cognitive fusion representation according to the context representation and the context representation of the cognitive sequence; Determine the probabilities of emotion category distributions according to the emotion fusion representation and the cognitive fusion representation; Obtain a true label, and determine a cross-entropy loss according to the true label and the emotion category distribution; Obtain a text embedding sequence corresponding to a generated text, and determine the probability of the generated text according to the text embedding sequence and a context representation combined with common sense knowledge; Determine the word similarity between a word in the generated text and a target word according to the probability; Obtain meme data and a predicted meme, and determine a meme similarity according to the meme data, the predicted meme, and the input text sequence; the meme data includes: an embedding vector of a meme corresponding to a pre-stored meme, a context label of a word in the embedding sequence, a meme prediction sample, the total number of emotion categories, and a meme category; Determine the accuracy of a target conversation according to the word similarity and the meme similarity.

[0055] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the following steps: Obtain each historical utterance, and determine an input text sequence according to the historical utterance; determine an embedding sequence according to the input text sequence; Perform a context encoding operation on the embedding sequence to obtain a context representation; Add different relationship tags and COMET operations to the last utterance in the historical utterance respectively to obtain a first sequence, a second sequence, a third sequence, a fourth sequence, and a fifth sequence; Determine a cognitive sequence according to the second sequence, the third sequence, the fourth sequence, and the fifth sequence; Determine a context representation of an emotion sequence according to the first sequence; determine a context representation of a cognitive sequence according to the cognitive sequence; Determine an emotion fusion representation according to the context representation and the context representation of the emotion sequence; Determine a cognitive fusion representation according to the context representation and the context representation of the cognitive sequence; Determine the probabilities of emotion category distributions according to the emotion fusion representation and the cognitive fusion representation; Obtain the true label, and determine the cross-entropy loss according to the true label and the sentiment category distribution; Obtain the text embedding sequence corresponding to the generated text, and determine the probability of the generated text according to the text embedding sequence and the combined context representation of common sense knowledge; Determine the word similarity between the words in the generated text and the target word according to the probability; Obtain the meme data and the predicted meme, and determine the meme similarity according to the meme data, the predicted meme and the input text sequence; the meme data includes: the embedding vector of the meme corresponding to the pre-stored meme, the context label of the word in the embedding sequence, the meme prediction sample, the total number of sentiment categories, and the meme category; Determine the accuracy of the target dialogue according to the word similarity and the meme similarity.

[0056] Figure 8 The internal structure diagram of a computer device in an embodiment is shown. The computer device can specifically be a terminal or a server. As Figure 8 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can implement the multi-modal dialogue generation method based on cognitive empathy. The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute the multi-modal dialogue generation method based on cognitive empathy. Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0057] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0058] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0059] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A multimodal dialogue generation method based on cognitive empathy, wherein the dialogue consists of multiple texts or emoticons, each text contains at least one word, and the target text is generated from both emotion and word aspects, characterized in that: The method comprises: Obtain each historical speech, and determine an input text sequence according to the historical speech; determine an embedding sequence according to the input text sequence; Performing a context encoding operation on the embedded sequence to obtain a context representation; Add different relationship tags and COMET operations to the last speech in the historical speeches to obtain a first sequence, a second sequence, a third sequence, a fourth sequence and a fifth sequence; Determine a recognition sequence according to the second sequence, the third sequence, the fourth sequence and the fifth sequence; Determine an emotional sequence context representation according to the first sequence; determine a cognitive sequence context representation according to the cognitive sequence; Determine an emotion fusion representation according to the context representation and the emotion sequence context representation; determining a cognitive fusion representation according to the context representation and the cognitive sequence context representation; Determine a probability of emotion category distribution according to the emotion fusion representation and the cognitive fusion representation; Obtaining a true label, and determining a cross entropy loss according to the true label and the sentiment category distribution; Obtaining a text embedding sequence corresponding to the generated text, and determining the probability of the generated text according to the text embedding sequence and the context representation of the common sense knowledge combination; Determining the word similarity between the word in the generated text and the target word according to the probability; Acquire meme data and predicted memes, and determine meme similarity according to the meme data, the predicted memes, and the input text sequence; the meme data includes: an embedding vector of a meme corresponding to a pre-stored meme, context tags of words in the embedding sequence, a meme prediction sample, a total number of emotion categories, and a meme category; The accuracy of multiple target dialogues is determined according to the word similarity and the meme similarity, and the target dialogue with the highest accuracy is the final generated dialogue.

2. The method for generating multimodal dialogue based on cognitive empathy according to claim 1, characterized in that: Determining the embedding sequence according to the input text sequence comprises: Perform a word embedding operation on the input text sequence to obtain a word embedding sequence, perform a position embedding operation on the input text sequence to obtain a position embedding sequence, and perform a dialogue embedding operation on the input text sequence to obtain a dialogue embedding sequence; Determine an embedding sequence based on the word embedding sequence, the position embedding sequence, and the conversation embedding sequence; Determining the recognition sequence according to the second sequence, the third sequence, the fourth sequence and the fifth sequence comprises: The second sequence, the third sequence, the fourth sequence and the fifth sequence are spliced ​​together to obtain a recognition sequence; The determining of the emotional sequence context representation according to the first sequence; and the determining of the cognitive sequence context representation according to the cognitive sequence include: Performing an emotion sequence embedding operation on the first sequence to obtain an emotion sequence embedding representation; performing an emotion encoding operation on the emotion sequence embedding representation to obtain an emotion sequence context representation; An embedding operation is performed on the cognitive sequence to obtain an embedding representation of the cognitive sequence; and a cognitive encoding operation is performed on the embedding representation of the cognitive sequence to obtain a context representation of the cognitive sequence.

3. The method for generating multimodal dialogue based on cognitive empathy according to claim 2, characterized in that: The determining of the emotion fusion representation according to the context representation and the emotion sequence context representation comprises: Performing an average hidden operation on the context representation of the emotion sequence to obtain an average hidden representation; Concatenating the context representation with the average hidden representation to obtain a sentiment fusion representation; Determining the cognitive fusion representation according to the context representation and the cognitive sequence context representation comprises: Performing a classification hiding operation on the cognitive sequence context representation to obtain a classification hiding representation; The context representation is concatenated with the classification hidden representation to obtain a cognitive fusion representation.

4. The method for generating multimodal dialogue based on cognitive empathy according to claim 3, characterized in that: Determining the probability of emotion category distribution according to the emotion fusion representation and the cognitive fusion representation comprises: Performing emotion refinement encoding on the emotion fusion representation to obtain an emotion refinement context representation; Performing cognitive refinement encoding on the cognitive fusion representation to obtain a cognitive refinement context representation; splicing a plurality of the cognitive refinement context representations to obtain a cognitive refinement context splicing representation; Concatenating the cognitive refined context concatenated representation with the emotional refined context representation to obtain a common sense refined context representation; Determine a common sense knowledge combined context representation according to the common sense refined context representation; Determining a sentiment-refined hidden representation according to the sentiment-refined contextual representation; The emotion category distribution probabilities are determined based on the emotion refined hidden representation.

5. The method for generating multimodal dialogue based on cognitive empathy according to claim 2, characterized in that: Determining the embedding sequence according to the input text sequence is achieved by the following expression: <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> =E1+E2+E3 The determination of the recognition sequence according to the second sequence, the third sequence, the fourth sequence and the fifth sequence is achieved by the following expression: The determining of the emotional sequence context representation according to the first sequence and the determining of the cognitive sequence context representation according to the cognitive sequence are implemented by the following expressions: in, For each historical speech, C is the input text sequence, E1 is the word embedding sequence, E2 is the position embedding sequence, and E3 is the dialogue embedding sequence. is the embedding sequence, For context, is the context encoder, For cognitive sequence, For the first sequence, For the second sequence, For the third sequence, For the fourth sequence, It is the fifth sequence; is the context representation of the sentiment sequence, is the embedding representation of the sentiment sequence, is the emotion encoder, is the cognitive sequence context representation, is the embedding representation of the cognitive sequence, Cognitive encoder.

6. The method for generating multimodal dialogue based on cognitive empathy according to claim 5, characterized in that: The average hidden representation obtained by performing an average hidden operation on the context representation of the emotion sequence is realized by the following expression: in, is the average hidden representation, Contextual representation for sentiment sequences; The classification and hiding operation of the cognitive sequence context representation to obtain the classification and hiding representation is implemented by the following expression: in, is the cognitive sequence context representation, Hidden representation for classification; The emotion fusion representation obtained by concatenating the context representation with the average hidden representation is realized by the following expression: in, is the average hidden representation, For context, For emotional integration, it is indicated; The concatenation of the context representation and the classification hidden representation to obtain the cognitive fusion representation is achieved by the following expression: in, is the classification hidden representation, For context, It is represented by cognitive fusion.

7. The method for generating multimodal dialogue based on cognitive empathy according to claim 6, characterized in that: The emotion fusion representation is subjected to emotion refinement coding to obtain the emotion refinement context representation; the cognition fusion representation is subjected to cognition refinement coding to obtain the cognition refinement context representation through the following expressions: in, Refine the contextual representation for sentiment, To refine the context representation for cognition, is the sentiment refinement encoder, For cognitive refinement encoder, For emotional integration, It is represented by cognitive fusion; The step of concatenating the plurality of cognitive refinement context representations to obtain a cognitive refinement context concatenation representation; and concatenating the cognitive refinement context concatenation representation with the emotion refinement context representation to obtain a common sense refinement context representation is implemented by the following expression: in, for Cognitively refine contextual concatenated representations, , Refine contextual representations for common sense, Refining contextual representations for sentiment.

8. The method for generating multimodal dialogue based on cognitive empathy according to claim 7, characterized in that: The determining of the common sense knowledge combination context representation according to the common sense refined context representation; the determining of the emotion refined hidden representation according to the emotion refined context representation; and the determining of the emotion category distribution probability according to the emotion refined hidden representation are implemented by the following expressions: in, ∈ Combine contextual representations for commonsense knowledge, is the importance score, Refine contextual representations for common sense, is the sensing operation, ⊙ represents element-by-element multiplication, ∈ Hidden representation for sentiment refinement; Refining contextual representations for sentiment; ∈ is the probability distribution of sentiment categories, ∈ is the weight vector of the linear layer.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Emotion dialogue generation method based on multi-resolution emotion and multi-type knowledge

    CN114372135A

  • Topic prediction and emotional reasoning integrated emotional dialogue generation method

    CN114385802A