A multimodal empathy dialogue generation method integrating personality traits
By integrating personality traits and multimodal data into the empathic dialogue generation method, the problems of insufficient emotion recognition and cognitive empathy in existing technologies are solved, more accurate user emotional understanding and personalized dialogue responses are achieved, and the satisfaction and efficiency of human-computer interaction are improved.
Patent Information
- Application Number
- CN202411966445.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing multimodal empathic dialogue generation methods have deficiencies in emotion recognition and cognitive empathy, and lack consideration of user personalized information, resulting in a lack of personalization and consistency in empathic responses, affecting the effectiveness and satisfaction of human-computer interaction.
A multimodal empathetic dialogue generation method that integrates personality traits is constructed. By combining image and text modal data, a bag-of-words model is used to generate feature vectors, personality traits and emotion classifiers are constructed, and empathetic responses are generated using a pre-trained language model. A small language model is used to reduce resource consumption.
It improves the accuracy of emotion recognition and understanding of the user's actual situation, generates more personalized and empathetic responses that fit the user's personality, improves the satisfaction of human-computer interaction, and is efficiently deployed in resource-constrained environments.
Smart Images

Figure CN119783691B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal empathy dialogue generation, and in particular to a multimodal empathy dialogue generation method integrating personality characteristics. Background Art
[0002] In interpersonal communication, empathy is closely linked to individual personality traits, which can be described using the Myers-Briggs Type Indicator (MBTI). This tool categorizes individuals into 16 different personality types using four dichotomous dimensions: Extraversion (E) vs. Introversion (I), Sensing (S) vs. Intuition (N), Thinking (T) vs. Feeling (F), and Judging (J) vs. Perceiving (P). During conversations, people not only rely on their habitual empathic expressions but also adjust their responses based on the other person's personality traits. For example, someone with a Thinking (T) orientation might respond more logically, while someone with a Feeling (F) orientation might prioritize the other person's feelings and offer a more empathetic response. Similarly, someone with a Perceiving (P) orientation might show a greater interest in openness and flexibility, while someone with a Judging (J) orientation might prioritize structure and decision-making. Therefore, understanding MBTI personality types can help empathic dialogue systems better understand and adapt to others, promoting harmonious and effective human-computer interactions.
[0003] Multimodal empathic dialogue generation is a cutting-edge research effort that combines natural language processing, computer vision, and speech processing technologies to achieve more emotionally understanding and humanistic conversational interactions. Existing empathic dialogue generation methods fall into three main technical approaches: The first focuses on identifying user emotions in conversational text and integrating these identified emotional features into the generation of empathic responses. However, given that human emotional expression is a multimodal process encompassing multiple aspects such as language symbols, voice intonation, facial expressions, and body posture, relying solely on text data for emotion recognition may limit its accuracy. Furthermore, empathy itself involves not only emotional resonance (emotional empathy) but also the cognitive process of understanding the other person's perspective (cognitive empathy). Therefore, solely incorporating user emotions may result in empathic dialogue systems failing to achieve deep cognitive empathy.
[0004] The second approach is to identify user emotions based on conversation text, derive commonsense knowledge from the conversation context with the help of an external knowledge graph, and integrate this inferred knowledge with the identified emotion labels into the generation process of empathetic responses. However, this approach still fails to fully consider the multimodal nature of human communication and may therefore lack accuracy in emotion recognition. In addition, relying solely on knowledge graphs for text reasoning to obtain commonsense knowledge related to users has certain limitations. This not only limits a comprehensive understanding of the user's specific context, but also affects the validity and reliability of the reasoning results due to the potential for inaccurate or incomplete knowledge graph content.
[0005] The third approach is to use a large multimodal language model as a conversation agent. This type of model can receive multimodal data input related to user interaction and generate corresponding output content based on it. However, it is worth noting that most existing large multimodal language models are not specifically designed for the task of empathic conversation generation. Although these models demonstrate excellent capabilities in language reasoning and understanding, they generally perform worse than specially fine-tuned small models in empathic understanding and expression. In addition, large multimodal language models have relatively high requirements for computing resources, which limits their large-scale deployment and application on end-side devices, especially in resource-constrained environments.
[0006] Furthermore, the three aforementioned approaches generally fail to adequately consider the user's individual information when generating empathic responses. This results in highly consistent and generalized empathic responses, failing to tailor responses to individual user characteristics. This lack of targeted empathy can undermine the effectiveness of human-computer interaction and, in turn, reduce satisfaction with the experience. Summary of the Invention
[0007] In order to address the shortcomings of the above-mentioned existing technologies, the present invention proposes a method for generating multimodal empathic dialogues that integrates personality characteristics, in order to construct a deep learning network that generates empathic dialogues that are consistent with the user's personality characteristics and expression habits in a multimodal environment, thereby improving the ability of the multimodal empathic dialogue system in expressing emotional resonance and enhancing the user's satisfaction with the human-computer interaction experience.
[0008] In order to achieve the above-mentioned object, the present invention adopts the following technical solutions:
[0009] The multimodal empathy dialogue generation method integrating personality traits of the present invention is characterized by being performed according to the following steps:
[0010] Step 1: Build a multimodal data set containing personality trait labels and emotion labels ;in, Indicates the Multimodal information of each user, Indicates the The visual modality data of each user, Indicates the Text modal data after word segmentation of each user, Indicates the The personality trait labels of each user, Indicates the The user's emotion label, is the total number of users; , Indicates the first Text modal data The words, Indicates the The length of the text modal data;
[0011] Step 2: Generate using the bag-of-words model The set of feature vectors ; Using the feature vector set generate The likelihood probability ,in, Represents a set of feature vectors The vector in ; thus, the personality feature classifier is constructed using formula (1) Optimization goal, and train the personality trait classifier :
[0012] (1)
[0013] In formula (1), express The posterior probability of Indicates direct proportion; express The likelihood probability of ; express The prior probability of
[0014] Step 3: Identify user s’s personality trait labels and its vector representation ;
[0015] Step 4: Build a sentiment classifier and and Processing, prediction User's emotion label ; Thus, the cross entropy loss function of the emotion classifier is constructed using formula (2) :
[0016] (2)
[0017] Step 5: Use the language model to Process the multimodal information of each user to obtain The probability distribution of empathetic responses of each user , thus using formula (3) to construct the negative logarithmic maximum likelihood loss of the language model :
[0018] (3)
[0019] In formula (3), Expressing hope;
[0020] Step 6: Use formula (4) to construct the total loss function of the multimodal empathy dialogue generation model that integrates personality features and is composed of the emotion classifier and the language model. :
[0021] (4)
[0022] In formula (4), and are 2 hyperparameters;
[0023] Step 7: Based on the total loss function , use the Adam optimizer to train the multimodal empathy dialogue generation model that integrates personality features and calculate the total loss function , update the network parameters according to the back propagation and gradient descent method until the number of iterations reaches the maximum or the total loss function When it no longer decreases, the training step is stopped, thereby obtaining a multimodal empathy dialogue generation model that optimally integrates personality features and is used to generate empathy dialogues.
[0024] The method for generating multimodal empathy dialogue integrating personality traits according to the present invention is also characterized in that step 3 is performed as follows:
[0025] Step 3.1: Obtain the text modal data of any user s and input it into the trained personality trait classifier after word segmentation. Make predictions and get the user Character traits tags ;
[0026] Step 3.2, according to , get the personality trait labels A text description is given. After adding a global tag at the beginning of the text description, it is input into the pre-trained language model BERT for processing to obtain the text description embedding vector and the embedding at the position of the global tag is As a personality trait label Vector representation of .
[0027] Furthermore, step 4 is performed as follows:
[0028] Step 4.1: Extract using the pre-trained language model GPT-2 Text modality features ; Extracted using pre-trained visual language model BLIP Visual modality features ;in, Represents the dimension of the feature;
[0029] Step 4.2: The emotion classifier classifies the visual modality features By linear transformation, it is mapped to the query and key value vector space, and then the formula (5) is used to calculate Self-attention :
[0030] (5)
[0031] In formula (5), Respectively represent the visual modality features The three parameters to be learned are mapped to the query and key-value vector space, is the dimension of the hidden layer; represents the softmax function;
[0032] Step 4.3: Calculate text modal features according to the steps in step 4.2 Self-attention ;
[0033] Step 4.4, the sentiment classifier will As the query vector, As key and value vectors, calculate the Cross-modal attention of a user ; Then cross-modal attention After processing by the feedforward layer and the normalization layer, the output vector of the hidden layer is obtained ; Thus, using formula (6) we can get The emotion prediction label of each user :
[0034] (6)
[0035] In formula (6), represents a linear classification layer, yes The parameters to be learned in ; Indicates the number of categories of sentiment labels.
[0036] Further, the step 5 is to Multimodal information of a user Input into the language model and use formula (7) to predict the The probability distribution of empathetic responses of each user ,Will pass After function processing, we get User's empathetic response:
[0037] (7)
[0038] In formula (7), represents the next word predicted by the language model, Indicates the first Text modal data The words, Represents the parameters to be learned in the language model.
[0039] The electronic device of the present invention includes a memory and a processor, and is characterized in that the memory is used to store a program that supports the processor to execute the multimodal empathy dialogue generation method integrating personality features, and the processor is configured to execute the program stored in the memory.
[0040] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, executes the steps of the method for generating a multimodal empathy dialogue integrating personality features.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] 1. By combining image and text data, this invention not only significantly improves the accuracy of emotion recognition compared to empathic dialogue generation methods based on pure text, but also provides a more comprehensive perspective for understanding the user's actual situation, thereby enhancing the empathic dialogue system's ability to perceive the user's emotions and real situation.
[0043] 2. By incorporating the user's individual characteristics into the dialogue generation process, this invention can more accurately identify and adapt to user personality differences compared to traditional empathic dialogue generation methods. This personalized adjustment enables the empathic dialogue system to generate empathic responses that are more consistent with individual personality, significantly improving the level of personalization of the dialogue. By considering the user's personality characteristics, the empathic dialogue system can adjust the depth and style of empathic expression, enhancing the fit and naturalness of the dialogue, thereby improving user satisfaction and interactive experience.
[0044] 3. The small language model used in this invention has higher computational efficiency and lower resource consumption compared to large language models. While maintaining good performance, the system can significantly reduce the model's storage requirements and computational burden, thereby enabling more efficient deployment in environments with limited hardware resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 Flow chart of the method of the present invention;
[0046] Figure 2 This is a network structure diagram of the emotion classifier of the present invention;
[0047] Figure 3 This is a network structure diagram of the feedforward layer of the present invention;
[0048] Figure 4 This is a network structure diagram of the empathy dialogue generator of the present invention. DETAILED DESCRIPTION
[0049] In this embodiment, a method for generating multimodal empathy dialogues integrating personality traits mainly uses the user personality trait labels and emotion labels identified from multimodal dialogue data to guide the generation of empathy dialogues, such as Figure 1 As shown, the method is carried out in the following steps:
[0050] Step 1: Build a multimodal data set containing personality trait labels and emotion labels ;in, Indicates the Multimodal information of each user, Indicates the The visual modality data of each user, Indicates the Text modal data after word segmentation of each user, Indicates the The personality trait labels of each user, Indicates the The user's emotion label, is the total number of users; , Indicates the first Text modal data The words, the word segmentation step involves removing stop words, noise words and punctuation marks, Indicates the The length of the text modal data; In this embodiment, the MELD and MEDIC datasets are used to train and test the model. The MELD dataset contains more than 13,000 dialogues from the TV series "Friends", and the MEDIC dataset includes 771 video clips from psychological counseling scenarios. Both datasets provide rich multimodal dialogue data and are pre-divided into training sets and test sets.
[0051] Step 2: Generate using the bag-of-words model The set of feature vectors ; Using the feature vector set calculate The posterior probability ,in, Represents a set of feature vectors The vector in ; thus, the personality feature classifier is constructed using formula (1) Optimization goal, and train the personality trait classifier :
[0052] (1)
[0053] Formula (1) uses Bayes' theorem to describe how to Calculate personality trait labels under the condition The probability of knowing the personality trait label is calculated by hour Likelihood probability and personality trait labels The prior probability of is used to infer the posterior probability of each personality trait label; in formula (1), Indicates that given input text Under the condition of The posterior probability of Indicates direct proportion; Indicates the known personality trait label In the case of The likelihood probability of ; Characteristic label The prior probability of , which reflects the probability of the personality trait label in the entire dataset.
[0054] Step 3: Identify user s’s personality trait labels and its vector representation ;
[0055] Step 3.1: Obtain the text modal data of any user s and input it into the trained personality trait classifier after word segmentation. Make predictions and get the user Character traits tags ;
[0056] Step 3.2, according to , get the personality trait labels A text description is given. After adding a global tag at the beginning of the text description, it is input into the pre-trained language model BERT for processing to obtain the text description embedding vector and the embedding at the position of the global tag is As a personality trait label The global mark added in this embodiment is , and extract the output vector of the last hidden layer of the pre-trained language model BERT as the vector representation of the entire text description.
[0057] Step 4: Build a sentiment classifier and and Processing, prediction User's emotion label ;
[0058] like Figure 2 As shown, the emotion classifier includes a visual feature purification module, a cross-modal fusion module, and a linear classification layer;
[0059] Visual feature extraction module After processing, a more consistent and expressive feature representation can be obtained, which is beneficial to the subsequent modal fusion process; the visual feature purification module includes a self-attention layer with 4 attention heads, 2 residual normalization layers and 1 feedforward layer; the output of the last residual normalization layer of the visual feature purification module will participate in the calculation of cross-modal attention as the key and value vectors.
[0060] The cross-modal fusion module includes a self-attention layer with 4 attention heads, 3 residual normalization layers, 1 cross-modal attention layer and 1 feed-forward layer. In view of the fact that text representation contains rich emotion-related information, in this embodiment, text features are used as query vectors in the cross-modal attention layer. The output vector of the last residual normalization layer in the cross-modal fusion module is processed by the linear classification layer to obtain the predicted emotion label. The linear classification layer includes a dimension of A linear layer, a ReLU activation function layer, a Dropout layer with a dropout rate of 0.3, and a dimension of Linear output layer;
[0061] like Figure 3 As shown, all feedforward layers in this embodiment contain two linear layers and use ReLU activation function between the linear layers, where the dimension of linear layer 1 is , the dimension of linear layer 2 is .
[0062] Step 4.1: Extract using the pre-trained language model GPT-2 Text modality features ; In this embodiment, the output vector of the last hidden layer of the pre-trained language model GPT-2 is used as Feature representation; extracted using the pre-trained visual language model BLIP Visual modality features ; In this embodiment, the output vector of the last hidden layer of the pre-trained visual language model BLIP is used as The feature representation of Indicates the dimension of the feature; in this embodiment Set to 768.
[0063] Step 4.2: The emotion classifier classifies the visual modality features By linear transformation, it is mapped to the query and key value vector space, and then the formula (2) is used to calculate Self-attention :
[0064] (2)
[0065] In formula (2), Respectively represent the visual modality features The three parameters to be learned that map to the query and key-value vector spaces are randomly initialized; is the dimension of the hidden layer; in this embodiment is 768; Represents the softmax function.
[0066] Step 4.3: Calculate text modal features according to the steps in step 4.2 Self-attention ;
[0067] Step 4.4, the sentiment classifier will As the query vector, As key and value vectors, calculate the Cross-modal attention of a user ;
[0068] use The reason for using text as a query vector is that text usually carries direct clues of emotions, which enables the cross-modal fusion module to adjust the attention allocation of the visual modality based on text information when processing visual modality features;
[0069] As key and value vectors, the cross-modal fusion module can better focus on relevant visual features under the guidance of specific text information;
[0070] Cross-modal attention After processing by the residual normalization layer and the feedforward layer, the output vector is obtained ; As the input of the linear classification layer, we can use formula (3) to get the first The emotion prediction label of each user :
[0071] (3)
[0072] In formula (3), represents a linear classification layer, yes The parameters to be learned in ; Indicates the number of types of emotion tags; in this embodiment, Set to 7.
[0073] Step 4.5: Use formula (4) to construct the cross entropy loss function of the emotion classifier :
[0074] (4)
[0075] Step 5: Use the language model to Process the multimodal information of each user to obtain Empathetic responses from users;
[0076] like Figure 4 As shown, the language model constructed in this embodiment is a decoder architecture, including an input layer, a Transformer layer, and an output layer;
[0077] The input layer includes three embedding layers with a dimension of 768: the word embedding layer maps the input text into a vector, the position embedding layer models the positional relationship between input sequences, and the tag category embedding layer is used to distinguish special tags and text word tags in the input sequence;
[0078] The Transformer layer has a total of 12 layers, each of which includes a self-attention layer with 12 attention heads, 2 residual normalization layers and 1 feed-forward layer;
[0079] The linear layer maps the vector of dimension 768 to the dimension of the vocabulary, obtaining the predicted distribution of words at each position, and the dimension of the vocabulary is 50265.
[0080] Step 5.1, Multimodal information of a user Input into the language model, and As a control signal, it represents different factors that express empathy, allowing the model to guide the generation of empathetic responses from the perspectives of the user's personality traits and emotional state;
[0081] Using formula (5) to predict the The probability distribution of empathetic responses of each user ;Will pass After function processing, we get User's empathetic response:
[0082] (5)
[0083] In formula (5), represents the next word predicted by the language model, Indicates the first Text modal data The words, Indicates the parameters to be learned in the language model; in this embodiment, the maximum length of the input sequence is 1024. Input sequences exceeding the maximum length limit are truncated before being input.
[0084] Step 5.2: Use formula (6) to construct the negative logarithmic maximum likelihood loss of the language model :
[0085] (6)
[0086] In formula (6), Indicates the desired operation.
[0087] Step 6: Use formula (7) to construct the total loss function of the multimodal empathy dialogue generation model that integrates personality features and is composed of the emotion classifier and the language model. :
[0088] (7)
[0089] In formula (7), and Are two hyperparameters; in this embodiment is 0.5, It is 1, which is used to balance the contribution of the two loss functions when training the model.
[0090] Step 7: Based on the total loss function , use the Adam optimizer to train the multimodal empathy dialogue generation model that integrates personality features and calculate the total loss function , update the network parameters according to the back propagation and gradient descent method until the number of iterations reaches the maximum or the total loss function When it no longer decreases, the training step is stopped, thereby obtaining a multimodal empathy dialogue generation model that optimally integrates personality features, which is used to generate empathy dialogues. In this embodiment, the maximum number of iterations is 10,000, the batch size of the training process is 8, and a learning rate warm-up strategy is adopted. This strategy helps stabilize the training process and avoid instability caused by an excessively high learning rate in the early stages of training. The initial learning rate is set to 0, and the learning rate increases linearly to 0.001 in the first 3,000 iterations. After 3,000 iterations, the learning rate begins to decay linearly until the total number of iterations reaches the maximum number of iterations or the training loss no longer decreases. In the testing phase, a kernel sampling strategy with a cumulative probability threshold of 0.8 is used. This strategy helps to improve the diversity of generated responses. The batch size is 1 and the maximum number of decoding steps is 30.
[0091] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0092] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are executed.
Claims
1. A multimodal empathy dialogue generation method integrating personality characteristics, characterized by: The steps are as follows: Step 1: Build a multimodal data set containing personality trait labels and emotion labels ;in, Indicates the Multimodal information of each user, Indicates the The visual modality data of each user, Indicates the Text modal data after word segmentation of each user, Indicates the The personality trait labels of each user, Indicates the The user's emotion label, is the total number of users; , Indicates the first Text modal data The words, Indicates the The length of the text modal data; Step 2: Generate using the bag-of-words model The set of feature vectors ; Using the feature vector set generate The likelihood probability ,in, Represents a set of feature vectors The vector in ; thus, the personality feature classifier is constructed using formula (1) Optimization goal, and train the personality trait classifier : (1) In formula (1), express The posterior probability of Indicates direct proportion; express The likelihood probability of ; express The prior probability of Step 3: Identify user s’s personality trait labels and its vector representation ; Step 4: Build a sentiment classifier and and Processing, prediction User's emotion label ; Thus, the cross entropy loss function of the emotion classifier is constructed using formula (2) : (2) Step 4.1: Extract using the pre-trained language model GPT-2 Text modality features ; Extracted using pre-trained visual language model BLIP Visual modality features ;in, Represents the dimension of the feature; Step 4.2: The emotion classifier classifies the visual modality features By linear transformation, it is mapped to the query and key value vector space, and then the formula (5) is used to calculate Self-attention : (5) In formula (5), Respectively represent the visual modality features The three parameters to be learned are mapped to the query and key-value vector space, is the dimension of the hidden layer; represents the softmax function; Step 4.3: Calculate text modal features according to the steps in step 4.2 Self-attention ; Step 4.4, the sentiment classifier will As the query vector, As key and value vectors, calculate the Cross-modal attention of a user ; Then cross-modal attention After processing by the feedforward layer and the normalization layer, the output vector of the hidden layer is obtained ; Thus, using formula (6) we can get The emotion prediction label of each user : (6) In formula (6), represents a linear classification layer, yes The parameters to be learned in ; Indicates the number of types of emotion labels; Step 5: Use the language model to Process the multimodal information of each user to obtain The probability distribution of empathetic responses of each user , thus using formula (3) to construct the negative logarithmic maximum likelihood loss of the language model : (3) In formula (3), Expressing hope; Step 6: Use formula (4) to construct the total loss function of the multimodal empathy dialogue generation model that integrates personality features and is composed of the emotion classifier and the language model. : (4) In formula (4), and are 2 hyperparameters; Step 7: Based on the total loss function , use the Adam optimizer to train the multimodal empathy dialogue generation model that integrates personality features and calculate the total loss function , update the network parameters according to the back propagation and gradient descent method until the number of iterations reaches the maximum or the total loss function When it no longer decreases, the training step is stopped, thereby obtaining a multimodal empathy dialogue generation model that optimally integrates personality features and is used to generate empathy dialogues.
2. The method for generating multimodal empathy dialogue integrating personality characteristics according to claim 1, characterized in that: Described step 3 is carried out as follows: Step 3.1: Obtain the text modal data of any user s and input it into the trained personality trait classifier after word segmentation. Make predictions and get the user Character traits tags ; Step 3.2, according to , get the personality trait labels A text description is given. After adding a global tag at the beginning of the text description, it is input into the pre-trained language model BERT for processing to obtain the text description embedding vector and the embedding at the position of the global tag is As a personality trait label Vector representation of .
3. The method for generating multimodal empathy dialogue integrating personality characteristics according to claim 1, characterized in that: The step 5 is to Multimodal information of a user Input into the language model and use formula (7) to predict the The probability distribution of empathetic responses of each user ,Will pass After function processing, we get User's empathetic response: (7) In formula (7), represents the next word predicted by the language model, Indicates the first Text modal data The words, Represents the parameters to be learned in the language model.
4. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the multimodal empathy dialogue generation method integrating personality features as described in any one of claims 1-3, and the processor is configured to execute the program stored in the memory.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating a multimodal empathy dialogue integrating personality features according to any one of claims 1 to 3 are executed.
Citation Information
Patent Citations
Chat robot and dialogue method based on chat robot
CN117390136A
Large language model high-context common-situation enhancement reply generation method and device
CN117668201A