A Multimodal Emotion Prediction Method and System Based on Interactive Robot Dialogue

By constructing multimodal features and adjusting and fusion of timing windows, the accuracy problem of single modal sentiment analysis is solved, multimodal sentiment prediction is realized, and the emotion recognition ability of dialogue robots is improved.

CN114254096BActive Publication Date: 2025-07-25COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111591253.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-23
Publication Date
2025-07-25
Estimated Expiration
2041-12-23

AI Technical Summary

Technical Problem

Traditional emotional analysis algorithms for text, voice and video are based on a single mode, which leads to the reduction in the accuracy of emotional information expressed by users and is unable to effectively integrate multiple sensory information.

Method used

Build multimodal features, including speech, dialogue context and video modal features, and input neural network models for emotional category prediction through modal timing window adjustment and fusion.

Benefits of technology

It improves the accuracy of identifying emotional changes in the conversation between users and customer service, improves user satisfaction, and creates a warm and emotional dialogue robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114254096B_ABST
    Figure CN114254096B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal emotion prediction method and system based on the dialogue of an interactive robot. The method includes: constructing multi-modal features based on the dialogue between a user and an interactive robot; adjusting the modal time series window of the multi-modal features; fusing the multi-modal features after the adjustment of the modal time series window; and inputting the fused multi-modal features into a trained neural network model to perform emotion category prediction. The present invention fuses and identifies the three modalities of text, speech, and video generated in the dialogue, and can better identify the emotional changes during the dialogue interaction between the user and the customer service, providing knowledge support for improving user satisfaction and creating a warm and emotional dialogue robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of emotion prediction, and in particular to a multi-modal emotion prediction method and system based on the dialogue of an interactive robot. Background Art

[0002] Emotion is a psychological reaction of humans when they encounter external things and receive information. In psychology courses, it is considered that "emotions and feelings are the attitude experiences of people towards objective things". In daily life, people convey individual emotions through signals such as facial expressions, actions, speech, and intonation expressions. In the 20th century, Ekman et al. classified human emotions into six basic emotions: anger, disgust, fear, happiness, sadness, and surprise. In the subsequent work of researchers, not only the basic emotions were given, but also emotions from the second level to the third level were given. Different scholars have different classification criteria, and there is no unified specification for emotion classification. Generally, there are the following two basic viewpoints: the discrete mode (categorical emotion states, CES) and the continuous mode (dimensional emtion space, DES). The two modes have different classification systems. Therefore, for the emotion analysis task, valuable emotion categories can be screened under the emotion analysis of the application scenario. The so-called "modal" is a biological concept proposed by the German physiologist Helmholtz, that is, the channels through which organisms receive information by relying on sensory organs and experience, such as the visual, auditory, tactile, gustatory, olfactory and other modalities of humans. "Multi-modal" is a way of fusing information obtained from multiple senses, such as sound, body language, information carriers (text, pictures, audio, video), etc. Traditional emotion analysis algorithms for text, speech, and video are all based on a single-modal training set combined with machine learning or deep learning methods, which are relatively single, and the emotional information expressed by users is isolated and split into a single modality, resulting in a decrease in accuracy. Summary of the Invention

[0003] In view of the above problems, the present invention provides a multi-modal emotion prediction method and system based on the dialogue of an interactive robot.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A multi-modal emotion prediction method based on the dialogue of an interactive robot, comprising:

[0006] Constructing multi-modal features based on the dialogue between the user and the interactive robot; the multi-modal features include: speech-modal emotion features, dialogue context-modal features, and video-modal emotion features;

[0007] Adjusting the modal time series window of the multi-modal features;

[0008] Fuse the multi-modal features after adjusting the modal temporal window;

[0009] Input the fused multi-modal features into the trained neural network model for sentiment category prediction.

[0010] Optionally, constructing multi-modal features based on the conversation between the user and the interactive robot specifically includes:

[0011] Extract the acoustic sentiment features in the conversation speech through an acoustic feature toolkit, splice and reduce the dimension to construct speech modal sentiment features;

[0012] Input the text features of the current sentence and the previous three sentences in the conversation into the BERT model to generate the user conversation context feature vector and construct the conversation context modal features;

[0013] Identify the facial expressions of the user in the conversation video and construct video modal sentiment features.

[0014] Optionally, the identifying the facial expressions of the user in the conversation video and constructing video modal sentiment features specifically includes:

[0015] Extract frames to identify the face area of the user in the conversation video and segment out the user;

[0016] Obtain the facial sentiment features of the user through the FACET facial expression analysis system;

[0017] Perform pooling operation on the facial sentiment features to obtain the video modal sentiment features within the current conversation interval.

[0018] Optionally, the adjusting the modal temporal window of the multi-modal features specifically includes:

[0019] Extreme the sentiment intensity of the current conversation through the sentiment words in the current conversation text;

[0020] Adjust the modal temporal window of the multi-modal features according to the sentiment intensity.

[0021] Optionally, the extreme the sentiment intensity of the current conversation through the sentiment words in the current conversation text specifically includes:

[0022] Segment the current conversation text;

[0023] Determine the sentiment words in the segmentation based on the sentiment dictionary;

[0024] Determine the sentiment intensity of the current conversation according to the number of negative sentiment words and positive sentiment words in the sentiment words.

[0025] Optionally, use the publicly available multi-modal dataset MOSEI as the training data to train the neural network model.

[0026] The present invention also provides a multi-modal emotion prediction system based on an interactive robot dialogue, comprising:

[0027] A multi-modal feature construction module, configured to construct multi-modal features based on a dialogue between a user and an interactive robot; the multi-modal features include: speech-modal emotion features, dialogue context-modal features, and video-modal emotion features;

[0028] An adjustment module, configured to adjust the modal time series window of the multi-modal features;

[0029] A fusion module, configured to fuse the multi-modal features after the adjustment of the modal time series window;

[0030] An emotion category prediction module, configured to input the fused multi-modal features into a trained neural network model to perform emotion category prediction.

[0031] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:

[0032] The present invention provides a multi-modal emotion prediction method based on an interactive robot dialogue, comprising: constructing multi-modal features based on a dialogue between a user and an interactive robot; adjusting the modal time series window of the multi-modal features; fusing the multi-modal features after the adjustment of the modal time series window; inputting the fused multi-modal features into a trained neural network model to perform emotion category prediction. The present invention fuses and identifies the three modalities of text, speech, and video generated in the dialogue, and can better identify the emotional changes between the user and the customer service dialogue interaction, providing knowledge support for improving user satisfaction and creating a warm and emotional dialogue robot. Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0034] Figure 1 It is a flowchart of the multi-modal emotion prediction method based on an interactive robot dialogue in the embodiment of the present invention;

[0035] Figure 2 It is a schematic diagram of the multi-modal emotion prediction method based on an interactive robot dialogue in the embodiment of the present invention. Detailed Embodiments

[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0038] As Figure 1-2 shown, a multi-modal emotion prediction method based on the dialogue of an interactive robot provided by the present invention includes the following steps:

[0039] Step 101: Based on the dialogue between the user and the interactive robot, construct multi-modal features; the multi-modal features include: speech modal emotion features, dialogue context modal features, and video modal emotion features.

[0040] (1) Speech (Audio Input): Extract acoustic emotion features such as MFCC, Mel spectral energy dynamic coefficients, and speech rate from the wav-format speech through the Librosa acoustic feature toolkit, and splice and reduce the dimensions to construct speech modal emotion features.

[0041] Embedding A = Librosa(Audio Input)

[0042] (2) Construct the dialogue context vector between the robot and the user (Context Embedding):

[0043] Considering that in the dialogue scenario between the robot and the user, the influence of the upper and lower sentences in the speech and video modalities is directly less relevant. The current sentence and the text features of the first three sentences for predicting emotion discrimination need to be input into the BERT model (the content range of the first three sentences includes the customer service conversation sentences) to generate the user dialogue context feature vector, which better represents the text features with context sentences effective for emotion analysis and replaces the dialogue sentence vector of the text model as the input to the model.

[0044]

[0045] (3) Video (Vedio Input): First, extract frames to identify the speaking face area in the dialogue, segment the interlocutors, and use the FACET facial expression analysis system to obtain facial emotion features and perform pooling operations on the frame-level features to calculate the emotion features of the video modality within the dialogue sentence interval.

[0046] EmbeddingV = FACET(Visual input)

[0047] Step 102: Adjust the modal time series window of the multi-modal features. Specifically, it includes: determining the emotional intensity of the current conversation through the emotional words in the current conversation text; adjusting the modal time series window of the multi-modal features according to the emotional intensity.

[0048] Since the scope of the emotional impact of the three modalities on different sentences in the conversation is different, speech and video are mostly strongly associated with the current sentence, and text is strongly associated with the previous two sentences. Therefore, the default window for selecting the modality of the current sentence is: text - the previous two sentences and the current sentence, video - the current sentence, speech - the current sentence. Calculate the intensity of the emotion of the sentence by obtaining the number of emotional words (matching positive and negative in the emotion dictionary). If it exceeds the threshold of the set emotional intensity, the emotional features of the previous sentence will be added to the feature vectors of the video and speech modalities, and the modal features of the previous sentence will be added to the speech and video modalities that originally only store single sentences, endowing the video and speech modalities with the context of the scene.

[0049] Calculation process:

[0050] Segment the sentence using iieba to split the sentence into individual words.

[0051] Combine with the emotion dictionary to obtain the set of emotional words (senti-word) of the current sentence. Adjust the threshold Score of the emotional intensity according to different scenarios used in the emotion analysis algorithm. For example, in this algorithm mainly used in customer service, when there are more than two emotional words, it can be determined that the user's emotion fluctuates greatly. Count the number of negative and positive emotional words in the sentence and summarize n is the number of words matching the emotion dictionary.

[0052] Let the threshold be S. When Score > S, the video and speech modalities of the current sentence also select the modal features of the previous sentence and splice them with the features of the current sentence and then pool to obtain the emotional features when the emotion fluctuates strongly.

[0053] Step 103: Integrate the multi-modal features after adjusting the modal time series window.

[0054] Step 104: Input the integrated multi-modal features into the trained neural network model for emotion category prediction.

[0055] The three current sentence modal sentiment features of the user are concatenated to obtain the fusion features of the three modalities. The final multi-modal fusion sentiment features are used as the final sentiment Embedding input. Using public multi-modal datasets such as MOSEI as training data, the softmax function loss function is selected for training, so that the model can learn and predict sentiment categories.

[0056] The present invention also provides a multi-modal sentiment prediction system based on an interactive robot dialogue, including:

[0057] A multi-modal feature construction module for constructing multi-modal features based on the dialogue between the user and the interactive robot; the multi-modal features include: speech modal sentiment features, dialogue context modal features, and video modal sentiment features;

[0058] An adjustment module for adjusting the modal time series window of the multi-modal features;

[0059] A fusion module for fusing the multi-modal features after adjusting the modal time series window;

[0060] A sentiment category prediction module for inputting the fused multi-modal features into a trained neural network model for sentiment category prediction.

[0061] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section.

[0062] Specific examples are used in this article to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A multi-modal emotion prediction method based on the dialogue of an interactive robot, characterized in that, Including: Construct multi-modal features based on the conversation between the user and the interactive robot; the multi-modal features include: speech modality emotion features, dialogue context modality features, and video modality emotion features; Adjust the modality time series window of the multi-modal features; specifically include: calculate the emotion intensity of the current conversation through the emotion words in the current conversation text; adjust the modality time series window of the multi-modal features according to the emotion intensity; specifically, the default modal selection window for predicting the current sentence is: text - the previous two sentences and the current sentence, video - the current sentence, speech - the current sentence; obtain the emotion intensity of this sentence by calculating the number of emotion words in the text, if it exceeds the set emotion intensity threshold, the emotion features of the previous sentence will be added to the feature vectors of the video and speech modalities; Fuse the multi-modal features after adjusting the modality time series window; Input the fused multi-modal features into the trained neural network model for emotion category prediction.

2. The multimodal emotion prediction method based on the interaction robot dialogue according to claim 1, characterized in that, The constructing multi-modal features based on the conversation between the user and the interactive robot specifically includes: Extract the acoustic emotion features in the conversation speech through an acoustic feature toolkit, splice and reduce the dimension to construct speech modality emotion features; Input the text features of the current sentence and the previous three sentences in the conversation into the BERT model to generate user dialogue context feature vectors and construct dialogue context modality features; Identify the facial expressions of the user in the conversation video and construct video modality emotion features.

3. The multimodal emotion prediction method based on interactive robot dialogue according to claim 2, wherein, The identifying the facial expressions of the user in the conversation video and constructing video modality emotion features specifically includes: Extract frames to identify the face area of the user in the conversation video and segment the user; Obtain the facial emotion features of the user through the FACET facial expression analysis system; Perform pooling operation on the facial emotion features to obtain video modality emotion features within the current conversation interval.

4. The multimodal emotion prediction method based on interactive robot dialogue according to claim 1, wherein The calculating the emotion intensity of the current conversation through the emotion words in the current conversation text specifically includes: Segment the current conversation text; Determine the emotion words in the segmentation based on the emotion dictionary; Determine the emotion intensity of the current conversation according to the number of negative emotion words and positive emotion words in the emotion words.

5. The multimodal emotion prediction method based on interactive robot dialogue according to claim 1, characterized in that, Use the publicly available multi-modal dataset MOSEI as training data to train the neural network model.

6. A multimodal emotion prediction system based on interactive robot dialogue, characterized in that, Including: A multi-modal feature construction module for constructing multi-modal features based on the conversation between the user and the interactive robot; the multi-modal features include: speech modality emotion features, dialogue context modality features, and video modality emotion features; An adjustment module for adjusting the modality time series window of the multi-modal features; specifically include: calculate the emotion intensity of the current conversation through the emotion words in the current conversation text; adjust the modality time series window of the multi-modal features according to the emotion intensity; specifically, the default modal selection window for predicting the current sentence is: text - the previous two sentences and the current sentence, video - the current sentence, speech - the current sentence; obtain the emotion intensity of this sentence by calculating the number of emotion words in the text, if it exceeds the set emotion intensity threshold, the emotion features of the previous sentence will be added to the feature vectors of the video and speech modalities; A fusion module for fusing the multi-modal features after adjusting the modality time series window; The emotion category prediction module is used to input the fused multi-modal features into the trained neural network model for emotion category prediction.

Citation Information

Patent Citations

  • Emotion recognition method based on multi-modal dialogue text

    CN113609289A

  • Method and System for Sentiment Analysis

    US20200159826A1