A Multimodal Emotion Recognition Method for Medical Care Robots

By adopting multimodal emotion recognition method in medical care robots, combining self-attention, mutual attention and graph convolutional neural network technology, the problem of medical care robots lacking multimodal emotion recognition and context information is solved, and higher emotional recognition accuracy and humanized communication ability are achieved.

CN114724224BActive Publication Date: 2025-06-17ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210399065.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2025-06-17
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

Existing medical care robots lack multimodal emotion recognition capabilities, and multimodal emotion information lacks context information, resulting in low accuracy of emotional recognition and inability to meet the application of emotional interaction.

Method used

Multimodal emotion recognition method is adopted to extract multimodal emotional features of facial expressions, behavioral actions, and voice signals, multimodal context emotional information is introduced, and emotional features are fusion using self-attention mechanism and mutual attention mechanism, and context emotional features are extracted through graph convolution neural network, and emotional classification recognition and voice interaction are finally carried out.

Benefits of technology

It improves the accuracy of emotional recognition, enhances the humanized communication ability of medical care robots, improves the real-time and accuracy of patient user experience and condition feedback, and reduces the work intensity of doctors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724224B_ABST
    Figure CN114724224B_ABST
Patent Text Reader

Abstract

A multi-modal emotion recognition method for a medical care robot, comprising: collecting multi-modal emotion information, collecting video information and audio information of a patient; extracting expression self-attention emotion features and action self-attention emotion features according to the video information, and extracting speech self-attention emotion features and text self-attention emotion features according to the audio information; performing emotion feature fusion based on a mutual attention mechanism on the 4 types of self-attention emotion features to obtain complete multi-modal emotion features; performing context emotion feature extraction based on a graph convolutional neural network on the multi-modal emotion features to obtain multi-modal emotion features containing context information; performing emotion classification and recognition on the multi-modal emotion features containing context information to obtain an emotion label result; and performing speech interaction and display according to the emotion label result. The present invention can improve the accuracy of emotion recognition of people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of emotion recognition, and particularly to a multi-modal emotion recognition method for a medical care robot. Background Art

[0002] With the continuous development of human-computer interaction technology, emotional interaction has become a major focus of research in human-computer interaction. Among them, emotion recognition is the key to emotional interaction, and its purpose is to enable machines to perceive the emotional states of humans, making robots more user-friendly. The multi-modal emotion recognition technology has broad application prospects in the field of medical care robots. By using various sensors carried by medical care robots, multi-modal signals containing emotional features such as patients' facial expressions, behavioral actions, and voice signals are obtained. Feature extraction and fusion are realized by using methods such as deep learning to analyze and predict the emotions of patients, making the medical care robot more user-friendly, with stronger emotion recognition capabilities, enhancing the real-time and accuracy of patients' condition feedback while improving the user experience of patients when using the medical care robot, and reducing the work intensity of doctors.

[0003] Most of the existing medical care robots lack the function of emotion recognition. Some robots with emotion recognition capabilities can only achieve simple emotion recognition functions based on a single modality, such as emotion recognition technologies based on expressions or voices. Such robots do not consider the complementarity between the emotional information obtained from different modalities. In the case where the emotional information is affected by noise or the emotional information is not obtained sufficiently, the emotion recognition accuracy is low and cannot meet the application of emotional interaction.

[0004] The existing multi-modal emotion analysis methods mainly focus on the research of extracting appropriate single-modal emotional features and constructing a robust multi-modal emotion analysis model, ignoring the context relationship existing between the upper and lower frames of the video. In fact, the emotion at the current moment is often closely related to the emotions expressed at the previous and subsequent moments, and this relationship cannot be ignored.

[0005] In summary, developing a method based on multi-modal feature acquisition and expression, using multi-modal emotional information of facial expressions, behavioral actions, voices, and text information, overcoming the difficulty of low efficiency of existing single-modal emotion recognition, introducing multi-modal context emotional information, and constructing a multi-modal emotion recognition system suitable for medical care robots has become an urgent problem to be solved by those skilled in the art in this research field. Summary of the Invention

[0006] To overcome the problems that existing medical care robots lack multi-modal emotion recognition ability and multi-modal emotion information lacks context information, the present invention provides a multi-modal emotion recognition method for medical care robots that utilizes multi-modal emotion features of facial expressions, behavioral actions, and voice signals and introduces multi-modal context emotion information.

[0007] The technical solution adopted by the present invention is as follows:

[0008] The present invention provides a multi-modal emotion recognition method for medical care robots, including multi-modal emotion information collection, expression self-attention emotion feature extraction, action self-attention emotion feature extraction, voice self-attention emotion feature extraction, text self-attention emotion feature extraction, emotion feature fusion based on cross-attention mechanism, context emotion feature extraction based on graph convolutional neural network, emotion classification and recognition, voice interaction, and display. The specific process includes the following steps:

[0009] 1. Perform multi-modal emotion information collection, and collect video information and audio information of the patient.

[0010] 2. According to the video information, perform expression self-attention emotion feature extraction and action self-attention emotion feature extraction, and according to the audio information, perform voice self-attention emotion feature extraction and text self-attention emotion feature extraction.

[0011] 3. Perform emotion feature fusion based on the cross-attention mechanism on the 4 types of self-attention emotion features to obtain complete multi-modal emotion features.

[0012] 4. Perform context emotion feature extraction based on graph convolutional neural network on the multi-modal emotion features to obtain multi-modal emotion features containing context information.

[0013] 5. Perform emotion classification and recognition on the multi-modal emotion features containing context information to obtain emotion label results.

[0014] 6. Perform voice interaction and display according to the emotion label results.

[0015] Further, the expression self-attention emotion feature extraction is used to extract the emotion feature vector of the patient's expression according to the video information and transform it into expression self-attention emotion features through the self-attention mechanism.

[0016] Further, the specific steps of extracting the emotion feature vector of the patient's expression include:

[0017] First, a pre-trained model and a combination network are used to extract video features. At the same time, a facial expression recognition library is used to detect the key points of the human face in the framed pictures. Then, by calculating the center point and the distance from each key point to the center point, the features of the key points are obtained. Finally, the two parts of the features are spliced together to form the complete expression emotion features.

[0018] Further, the conversion into expression self-attention emotion features through the self-attention mechanism specifically includes:

[0019] Taking the obtained expression emotion features as the input of the self-attention mechanism, the expression emotion feature vectors are converted into I groups of feature vectors according to the video frames corresponding to the video information. The size of each group of feature vectors is where I is the number of video frames and E is the dimension of the expression emotion feature vector. The expression self-attention emotion features obtained through the self-attention mechanism are as follows:

[0020]

[0021]

[0022] where is the weight coefficient of the i-th group of feature vectors, represents the i-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W E is the trainable linear transformation parameter vector, and F E is the expression self-attention emotion feature through the self-attention mechanism.

[0023] Further, the extraction of the action self-attention emotion features is used to extract the emotion feature vectors of the patient's actions according to the video information and convert them into action self-attention emotion features through the self-attention mechanism;

[0024] Further, the extraction of the emotion feature vectors of the patient's actions specifically includes:

[0025] First, a pre-trained model and a combination network are used to extract video features. At the same time, a human body pose detection library is used to detect the joint points of the human body in the framed pictures. Then, by calculating the center of gravity of the human body and the distance and angle from each joint point to the center of gravity, the features of the human body joint points are obtained. Finally, the two parts of the features are spliced together to form the complete action emotion features.

[0026] Further, the conversion into action self-attention emotion features through the self-attention mechanism specifically includes:

[0027] Taking the obtained action emotion features as the input of the self-attention mechanism, the action emotion feature vectors are converted into J groups of feature vectors according to the video frames corresponding to the video information. The size of each group of feature vectors is Where J is the number of video frames and A is the dimension of the action emotion feature vector. The action self-attention emotion feature obtained through the self-attention mechanism is as follows:

[0028]

[0029]

[0030] Where is the weight coefficient of the j-th group of feature vectors, represents the j-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W A is the trainable linear transformation parameter vector, and F A is the action self-attention emotion feature obtained through the self-attention mechanism.

[0031] Furthermore, the speech self-attention emotion feature extraction is used to extract the emotion feature vector of the patient's speech according to the audio information and transform it into the speech self-attention emotion feature through the self-attention mechanism;

[0032] Furthermore, the specific steps of extracting the emotion feature vector of the patient's speech include:

[0033] Preprocess the collected audio signal and draw a spectrogram, then construct and train a convolutional neural network, and finally use the trained network to extract the speech emotion feature.

[0034] Furthermore, the specific steps of transforming it into the speech self-attention emotion feature through the self-attention mechanism include:

[0035] Use the obtained speech emotion feature as the input of the self-attention mechanism, and convert the speech emotion feature vector into K groups of feature vectors according to the number of speech frames of each audio information. The size of each group of feature vectors is Where K is the number of audio frames and V is the dimension of the expression emotion feature vector. The expression self-attention emotion feature obtained through the self-attention mechanism is as follows:

[0036]

[0037]

[0038] Where is the weight coefficient of the k-th group of feature vectors, represents the k-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W V is the trainable linear transformation parameter vector, and F V is the speech self-attention emotion feature obtained through the self-attention mechanism.

[0039] Further, the text self-attention emotion feature extraction is used to extract the emotion feature vector of the patient's text according to the audio information, and transform it into the text self-attention emotion feature through the self-attention mechanism;

[0040] Further, the extraction of the emotion feature vector of the patient's text specifically includes:

[0041] First, use an end-to-end ASR system to extract the audio signal into text information, then use a pre-trained model to extract the word vector features in the text information, then add the word vectors of each word in each sentence to obtain a sentence vector, and at the same time use a pre-trained model to extract the sentence vector of each sentence. Finally, combine and splice the sentence vectors extracted from the two parts to obtain the complete text emotion feature.

[0042] Further, the transformation into the text self-attention emotion feature through the self-attention mechanism specifically includes:

[0043] Take the obtained text emotion feature as the input of the self-attention mechanism, and convert the text emotion feature vector into L groups of feature vectors according to the number of words in the text. The size of each group of feature vectors is where L is the number of audio frames, and X is the dimension of the text emotion feature vector. The text self-attention emotion feature obtained through the self-attention mechanism is as follows:

[0044]

[0045]

[0046] where is the weight coefficient of the l-th group of feature vectors, represents the l-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W X is a trainable linear transformation parameter vector, and F X is the speech self-attention emotion feature through the self-attention mechanism.

[0047] Further, the emotion feature fusion based on the mutual attention mechanism is used to combine the above self-attention emotion features in pairs to obtain the mutual attention emotion feature, and the mutual attention emotion feature is fused through concatenation to obtain the complete multi-modal emotion feature.

[0048] Further, the combination of the above self-attention emotion features in pairs to obtain the mutual attention emotion feature specifically includes:

[0049] Taking the combination of expression-action as an example to illustrate the mutual attention mechanism. First, take the expression self-attention emotion feature and the action self-attention emotion feature as the input of the expression-action mutual attention mechanism, and obtain the expression-action mutual attention mechanism as follows:

[0050]

[0051]

[0052]

[0053] where softmax represents the normalized exponential function, and are trainable parameter vector matrices, Concat(·) represents vector concatenation, F A_E represents the relative action self-attention emotion feature, the feature after weighting the expression self-attention emotion feature, F E_A represents the relative expression self-attention emotion feature, the feature after weighting the action self-attention emotion feature, F EA represents the expression-action mutual attention emotion feature.

[0054] According to the above method, similarly, the expression-speech mutual attention emotion feature F EV , the expression-text mutual attention emotion feature F EX , the action-speech mutual attention emotion feature F AV , the action-text mutual attention emotion feature F AX , and the speech-text mutual attention emotion feature F VX can be obtained.

[0055] Furthermore, the specific process of obtaining the complete multi-modal emotion feature by cascading and fusing the mutual attention emotion features includes:

[0056] F = Concat(F EA , F EV , F EX , F AV , F AX , F VX ) (12)

[0057] where F represents the obtained complete multi-modal emotion feature.

[0058] Furthermore, the extraction of context emotion features based on the graph convolutional neural network is used to obtain multi-modal emotion features containing context information through the graph convolutional neural network for the multi-modal emotion features.

[0059] Furthermore, the specific process of obtaining multi-modal emotion features containing context information through the graph convolutional neural network includes:

[0060] First, the above multi-modal emotion features are used as the node set G of the graph structure v ∈R N×f, where N is the number of adjacent samples, f is the dimension of the multi-modal sentiment features extracted using the attention mechanism, and the similarity formula shown below is used to construct the adjacency matrix A to define the edge set information G between adjacent samples e .

[0061]

[0062] Among them, v i and v j represent the multi-modal sentiment feature vectors of the i-th and j-th adjacent samples, ‖·‖ represents the modulus operation, sim represents the cosine similarity. When sim ≥ 0.75, the element a i,j of the adjacency matrix A is 1; when sim < 0.75, the element a i,j of the adjacency matrix A is 0.

[0063] After that, the graph structure is used as the input of the graph convolutional neural network, and graph convolutional operations are performed on the graph structure. The graph convolutional operation is shown as follows

[0064]

[0065] Among them, the degree matrix a i,j represents the element in the i-th row and j-th column of the adjacency matrix A, H l represents the output of the l-th layer of the graph convolutional layer, and H 0 = F, where F is the feature vector output after the attention mechanism, W l is the trainable linear transformation parameter, l represents the number of layers of graph convolution, and σ(·) represents the ReLU activation function.

[0066] Furthermore, the sentiment classification and recognition is used to load the above multi-modal sentiment features containing context information into a pre-constructed and trained sentiment classification model for classification and recognition, and obtain multiple discrete sentiment label recognition results.

[0067] Furthermore, the multiple discrete sentiment labels specifically include pain, sadness, calmness, nausea, comfort, and happiness.

[0068] Furthermore, the voice interaction and display are used to perform real-time voice and display feedback of the sentiment analysis results to the patient and the doctor, and perform corresponding voice interactions with the patient and the doctor according to the obtained sentiment label recognition results.

[0069] Compared with the prior art, the present invention has the following advantages:

[0070] 1. Traditional single-modal emotion recognition systems have the disadvantages of incomplete emotion information and being vulnerable to noise interference, resulting in inaccurate emotion recognition. The present invention makes full use of patients' facial expressions, behavioral actions, voices, and text information for emotion analysis. Through multi-modal information fusion, emotion features are enriched, and the accuracy of emotion recognition is improved.

[0071] 2. The multi-modal context emotion feature extraction method based on graph convolutional neural network designed in the present invention obtains more abundant information and stronger representation ability in the multi-modal emotion features, which can effectively improve the accuracy of emotion analysis.

[0072] 3. The voice interaction mode designed in the present invention enables the medical care robot to communicate with patients and doctors more humanely according to the recognized emotion label results, providing a better experience for patients during medical care. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 It is a principle block diagram of a multi-modal emotion recognition method for a medical care robot provided in an embodiment of the present invention;

[0074] Figure 2 It is a principle block diagram of an expression self-attention emotion feature extraction method provided in an embodiment of the present invention;

[0075] Figure 3 It is a principle block diagram of an action self-attention emotion feature extraction method provided in an embodiment of the present invention;

[0076] Figure 4 It is a principle block diagram of a voice self-attention emotion feature extraction method provided in an embodiment of the present invention;

[0077] Figure 5 It is a principle block diagram of a text self-attention emotion feature extraction method provided in an embodiment of the present invention;

[0078] Figure 6 It is a principle block diagram of an emotion feature method based on mutual attention mechanism provided in an embodiment of the present invention;

[0079] Figure 7 It is a principle block diagram of a context emotion feature extraction method based on graph convolutional neural network provided in an embodiment of the present invention;

[0080] Figure 8 It is a principle block diagram of an emotion classification and recognition method provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] To illustrate the technical solutions and advantages of the embodiments of the present invention in detail, the technical solutions in the embodiments of the present invention will be clearly and completely described here in conjunction with the accompanying drawings in the embodiments of the present invention.

[0082] The following descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. All equivalent process changes made using the contents of the present invention specification are included in the patent protection scope of the present invention.

[0083] Embodiments of the present invention:

[0084] like Figure 1 As shown, this embodiment provides a multimodal emotion recognition method for a medical care robot, including multimodal emotion information collection, expression self-attention emotion feature extraction, action self-attention emotion feature extraction, voice self-attention emotion feature extraction, text self-attention emotion feature extraction, emotion feature fusion based on mutual attention mechanism, context emotion feature extraction based on graph convolutional neural network, emotion classification recognition, voice interaction and display. The specific process includes the following steps:

[0085] 1. Multimodal emotional information collection, collecting video and audio information of patients,

[0086] 2. Perform facial expression self-attention emotion feature extraction and action self-attention emotion feature extraction based on the video information, and perform speech self-attention emotion feature extraction and text self-attention emotion feature extraction based on the audio information.

[0087] 3. The four self-attention sentiment features are fused based on the mutual attention mechanism to obtain a complete multimodal sentiment feature.

[0088] 4. The multimodal sentiment features are subjected to contextual sentiment feature extraction based on a graph convolutional neural network to obtain multimodal sentiment features containing contextual information.

[0089] 5. The multimodal sentiment features containing contextual information are used for sentiment classification and recognition to obtain sentiment label results.

[0090] 6. Perform voice interaction and display based on the emotion label results

[0091] Each step is described in detail below.

[0092] 1. Multimodal emotional information collection, using multiple sensors to collect multimodal emotional data of patients served by the medical care robot. In this embodiment, a Kinect camera is used as a video collection device, and a microphone built into the Kinect camera is used as an audio collection device.

[0093] 2. Expression self-attention emotion feature extraction;

[0094] Figure 2 This is the principle block diagram of an expression self-attention emotion feature extraction method provided by the present invention, and its specific steps are as follows:

[0095] (1) First, use the opencv library to read video frame data, process the complete video information into the data form required by VGG-19, then use the pre-trained VGG-19 model on ImageNet to extract the feature representations of each frame of image, and then use LSTM (Long Short-Term Memory Network) to add temporal information through training. The output after passing through the fully connected layer is used as the video feature vector. At the same time, use the face detection function provided by the Dlib library to detect 68 key point information of the human face, such as eyes, eyebrows, nose, mouth, and face contour. Secondly, calculate the center point coordinates according to the coordinates of all key points, then use the cosine theorem to calculate the distance from each key point coordinate to the center point, and finally combine the distance information to obtain the expression key point features of the human face.

[0096] (2) Concatenate the above video feature vector and the expression key point feature vector to obtain the expression emotion feature vector.

[0097] (3) Use the obtained expression emotion feature as the input of the self-attention mechanism, and convert the expression emotion feature vector into I groups of feature vectors according to the number of video frames corresponding to the video information. The size of each group of feature vectors is , where I is the number of video frames and E is the dimension of the expression emotion feature vector. The expression self-attention emotion feature obtained through the self-attention mechanism is as follows:

[0098]

[0099]

[0100] Among them is the weight coefficient of the i-th group of feature vectors, represents the i-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W E is the trainable linear transformation parameter vector, and F E is the expression self-attention emotion feature obtained through the self-attention mechanism.

[0101] 3. Action self-attention emotion feature extraction;

[0102] Figure 3 This is the principle block diagram of an action self-attention emotion feature extraction method provided by the present invention, and its specific steps are as follows:

[0103] (1) First, use a combined network of the VGG-19 pre-trained model and LSTM to extract video feature vectors. At the same time, use the human action detection function provided by OpenPose to detect 25 body key point information. Secondly, calculate the center point coordinates according to all key point coordinates. Then, use the cosine theorem to calculate the distance from each key point coordinate to the center point. Finally, combine the distance information to obtain the pose key point feature vector of the human body.

[0104] (2) Concatenate the above video feature vectors and the pose key point feature vectors of the human body to obtain the expression emotion feature vector.

[0105] (3) Use the obtained action emotion feature as the input of the self-attention mechanism. According to the number of video frames corresponding to the video information, convert the action emotion feature vector into J groups of feature vectors, and the size of each group of feature vectors is where J is the number of video frames and A is the dimension of the action emotion feature vector. The action self-attention emotion feature obtained through the self-attention mechanism is as follows:

[0106]

[0107]

[0108] where is the weight coefficient of the j-th group of feature vectors, represents the j-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W A is the trainable linear transformation parameter vector, and F A is the action self-attention emotion feature obtained through the self-attention mechanism.

[0109] 4. Speech self-attention emotion feature extraction;

[0110] Figure 4 The principle block diagram of a speech self-attention emotion feature extraction method provided by the present invention is as follows. The specific steps are as follows:

[0111] (1) Use a first-order high-pass digital filter to perform signal compensation on the original audio signal. Then, divide the original audio signal into multiple speech frames, and then use a triangular band-pass filter to perform windowing on the signal.

[0112] (2) First, calculate the short-time power spectrum of the preprocessed speech signal through the short-time Fourier transform. Then, splice the power spectra of each frame pair in chronological order to generate a complete spectrogram.

[0113] (3) Construct and train a convolutional neural network, and use the trained network to extract speech emotion feature vectors.

[0114] (4) Use the obtained speech emotion features as the input of the self-attention mechanism, and convert the speech emotion feature vectors into K groups of feature vectors according to the number of speech frames of each audio message. The size of each group of feature vectors is where K is the number of audio frames, and V is the dimension of the expression emotion feature vector. The expression self-attention emotion features obtained through the self-attention mechanism are as follows:

[0115]

[0116]

[0117] where is the weight coefficient of the k-th group of feature vectors, represents the k-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W V is the trainable linear transformation parameter vector, and F V is the speech self-attention emotion feature through the self-attention mechanism.

[0118] 5. Text self-attention emotion feature extraction;

[0119] Figure 5 This is the principle block diagram of a method for extracting text self-attention emotion features provided by the present invention. The specific steps are as follows:

[0120] (1) First, use an end-to-end ASR system to recognize the audio signal into text information.

[0121] (2) Use the Word2Vec pre-trained model to extract the word vector features in the text information. Then, add the word vectors of each word in each sentence to obtain an 80-dimensional sentence vector. At the same time, use the Bert pre-trained model to extract the sentence vector of each sentence.

[0122] (3) Combine and splice the sentence vectors extracted from the above two parts to obtain a complete text emotion feature vector.

[0123] (4) Use the obtained text emotion features as the input of the self-attention mechanism, and convert the text emotion feature vectors into L groups of feature vectors according to the number of words in the text. The size of each group of feature vectors is where L is the number of audio frames, and X is the dimension of the text emotion feature vector. The text self-attention emotion features obtained through the self-attention mechanism are as follows:

[0124]

[0125]

[0126] where is the weight coefficient of the l-th group of feature vectors, represents the l-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, and W X is a trainable linear transformation parameter vector, and F X is the speech self-attention emotion feature through the self-attention mechanism.

[0127] 6. Emotion feature fusion based on the mutual attention mechanism;

[0128] Figure 6 This is the principle block diagram of an emotion feature fusion method based on the mutual attention mechanism provided by the present invention, and the specific steps are as follows:

[0129] Taking the combined expression-action as an example to illustrate the mutual attention mechanism. First, the expression self-attention emotion feature and the action self-attention emotion feature are used as the inputs of the expression-action mutual attention mechanism, and the expression-action mutual attention mechanism is obtained as follows:

[0130]

[0131]

[0132]

[0133] where softmax represents the normalized exponential function, and are trainable parameter vector matrices, Concat(·) represents vector concatenation, and F A_E represents the relative action self-attention emotion feature, the feature after adding weights to the expression self-attention emotion feature, and F E_A represents the relative expression self-attention emotion feature, the feature after adding weights to the action self-attention emotion feature, and F EA represents the expression-action mutual attention emotion feature.

[0134] According to the above method, similarly, the expression-speech mutual attention emotion feature F EV , the expression-text mutual attention emotion feature F EX , the action-speech mutual attention emotion feature F AV , the action-text mutual attention emotion feature F AX , and the speech-text mutual attention emotion feature F VX can be obtained.

[0135] Furthermore, the specific process of obtaining the complete multi-modal emotion feature by cascading and fusing the mutual attention emotion features includes:

[0136] F = Concat(F EA,F EV ,F EX ,F AV ,F AX ,F VX ) (12)

[0137] Among them, F represents the obtained complete multi-modal sentiment feature.

[0138] 7. Context sentiment feature extraction based on graph convolutional neural network;

[0139] Figure 7 The following is the principle block diagram of the context sentiment feature extraction method based on graph convolutional neural network provided by the present invention, and the specific steps are as follows:

[0140] (1). First, use the above-obtained multi-modal sentiment feature as the node set G of the graph structure v ∈R N×f , where N is the number of adjacent samples, f is the dimension of the multi-modal sentiment feature extracted using the attention mechanism, and the similarity formula shown below is used to construct the adjacency matrix A to define the edge set information G between adjacent samples e .

[0141]

[0142] Among them, v i and v j represent the multi-modal sentiment feature vectors of the i-th and j-th adjacent samples, ‖·‖ represents the modulus operation, sim represents the cosine similarity, when sim≥0.75, the element a i,j of the adjacency matrix A = 1; when sim<0.75, the element a i,j of the adjacency matrix A = 0.

[0143] (2). Use the graph structure as the input of the graph convolutional neural network and perform graph convolution operations on the graph structure. The graph convolution operation is shown as follows,

[0144]

[0145] Among them, the degree matrix a i,j represents the element in the i-th row and j-th column of the adjacency matrix A, H l represents the output of the l-th graph convolutional layer, and H 0 = F, F is the feature vector output through the attention mechanism, W l is the trainable linear transformation parameter, l represents the number of graph convolution layers, and σ(·) represents the ReLU activation function.

[0146] 8. Sentiment classification and recognition;

[0147] Figure 8 The principle block diagram of an emotion classification and recognition method provided by the present invention is as follows:

[0148] (1) Normalize the obtained multi-modal context emotion features.

[0149] (2) Select the RBF kernel function as the kernel function of the SVM classifier, use cross-validation to find the optimal parameter γ in the RBF kernel function, and use the optimal parameter γ to train the SVM model.

[0150] (3) Transmit the multi-modal data into the trained SVM model to obtain emotion labels, specifically including pain, sadness, calmness, nausea, comfort, and happiness.

[0151] 9. Voice interaction and display;

[0152] The voice interaction and display perform corresponding voice interactions according to the obtained emotion label results, and at the same time display the recognition results.

[0153] Implement a system for a multi-modal emotion recognition method for a medical care robot according to the present invention. This embodiment provides a multi-modal emotion recognition system for a medical care robot, including a multi-modal emotion information acquisition module, an expression self-attention emotion feature extraction module, an action self-attention emotion feature extraction module, a voice self-attention emotion feature extraction module, a text self-attention emotion feature extraction module, an emotion feature fusion module based on the mutual attention mechanism, a context emotion feature module based on a graph convolutional neural network, an emotion classification and recognition module, and a voice interaction and display module. The specific process is as follows:

[0154] 1. Use the multi-modal emotion information acquisition module on the medical care robot to obtain the video and audio information of the patient.

[0155] 2. Implement the preprocessing of multi-modal information and the extraction of self-attention emotion features in different modal self-attention emotion feature extraction modules.

[0156] 3. The obtained self-attention emotion features of different modalities are fused through the mutual attention mechanism emotion feature fusion module to achieve multi-modal emotion feature fusion.

[0157] 4. The obtained multi-modal emotion features are output as complete multi-modal emotion features through the context emotion feature module.

[0158] 5. Obtain emotion classification labels through the emotion classification and recognition module.

[0159] 6. Output the multi-modal emotion recognition results through the voice interaction and display module.

[0160] The following is a specific description of each module.

[0161] 1. The multi-modal emotion information acquisition module includes a video acquisition device and an audio acquisition device, and uses multiple sensors to collect multi-modal emotion data of the patients served by the medical care robot. In this embodiment, a Kinect camera is used as the video acquisition device, and the microphone built in the Kinect camera is used as the audio acquisition device.

[0162] 2. Facial expression self-attention emotion feature extraction module

[0163] Figure 2 The principle block diagram of the facial expression self-attention emotion feature extraction module provided by the present invention is as follows. The specific steps are as follows:

[0164] (1). First, use the opencv library to read the video frame data, process the complete video information into the data form required by VGG-19, then use the pre-trained VGG-19 model on ImageNet to extract the representation features of each frame of image, and then use LSTM (Long Short-Term Memory Network) to add temporal information after training. The output after passing through the fully connected layer is used as the video feature vector. At the same time, use the face detection function provided by the Dlib library to detect 68 key point information of the human face, such as eyes, eyebrows, nose, mouth, and face contour. Secondly, calculate the center point coordinates according to the coordinates of all key points, then use the cosine theorem to calculate the distance from each key point coordinate to the center point, and finally combine the distance information to obtain the expression key point features of the human face.

[0165] (2). Concatenate the above video feature vector and the expression key point feature vector to obtain an expression emotion feature vector.

[0166] (3). Use the obtained expression emotion feature as the input of the self-attention mechanism, and convert the expression emotion feature vector into I groups of feature vectors according to the number of video frames corresponding to the video information. The size of each group of feature vectors is where I is the number of video frames, and E is the dimension of the expression emotion feature vector. The expression self-attention emotion feature obtained through the self-attention mechanism is as follows:

[0167]

[0168]

[0169] where is the weight coefficient of the i-th group of feature vectors, represents the i-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W E is a trainable linear transformation parameter vector, and F E is the expression self-attention emotion feature obtained through the self-attention mechanism.

[0170] 3. Action self-attention emotion feature extraction module

[0171] Figure 3 The following is the principle block diagram of the action self-attention emotion feature extraction module provided by the present invention, and the specific steps are as follows:

[0172] (1). First, use a combined network of the VGG-19 pre-trained model and LSTM to extract video feature vectors. At the same time, use the human action detection function provided by OpenPose to detect 25 body key point information. Secondly, calculate the center point coordinates according to all key point coordinates. Then, use the cosine theorem to calculate the distances from each key point coordinate to the center point. Finally, combine the distance information to obtain the human pose key point feature vector.

[0173] (2). Concatenate the above video feature vectors and human pose key point feature vectors to obtain the expression emotion feature vector.

[0174] (3). Use the obtained action emotion feature as the input of the self-attention mechanism. According to the video frames corresponding to the video information, convert the action emotion feature vector into J groups of feature vectors, and the size of each group of feature vectors is where J is the number of video frames and A is the dimension of the action emotion feature vector. The action self-attention emotion feature obtained through the self-attention mechanism is as follows:

[0175]

[0176]

[0177] where is the weight coefficient of the j-th group of feature vectors, represents the j-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W A is the trainable linear transformation parameter vector, and F A is the action self-attention emotion feature through the self-attention mechanism.

[0178] 4. Speech self-attention emotion feature extraction module

[0179] Figure 4 The following is the principle block diagram of the speech self-attention emotion feature extraction module provided by the present invention, and the specific steps are as follows:

[0180] (1). Use a first-order high-pass digital filter to compensate the original audio signal. Then, divide the original audio signal into multiple speech frames, and then use a triangular band-pass filter to perform a windowing operation on the signal.

[0181] (2) First, calculate the short-time power spectrum of the preprocessed speech signal through short-time Fourier transform, and then splice the power spectra of each pair of frames in chronological order to generate a complete spectrogram.

[0182] (3) Construct and train a convolutional neural network, and use the trained network to extract speech emotion feature vectors.

[0183] (4) Use the obtained speech emotion features as the input of the self-attention mechanism, and convert the speech emotion feature vectors into K groups of feature vectors according to the number of speech frames of each audio message. The size of each group of feature vectors is where K is the number of audio frames, and V is the dimension of the expression emotion feature vector. The expression self-attention emotion features obtained through the self-attention mechanism are as follows:

[0184]

[0185]

[0186] where is the weight coefficient of the k-th group of feature vectors, represents the k-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W V is a trainable linear transformation parameter vector, and F V is the speech self-attention emotion feature obtained through the self-attention mechanism.

[0187] 5. Text self-attention emotion feature extraction module

[0188] Figure 5 The principle block diagram of the text self-attention emotion feature extraction module provided by the present invention is as follows. The specific steps are as follows:

[0189] (1) First, use an end-to-end ASR system to recognize the audio signal into text information.

[0190] (2) Use the Word2Vec pre-trained model to extract the word vector features in the text information, and then add the word vectors of each word in each sentence to obtain an 80-dimensional sentence vector. At the same time, use the Bert pre-trained model to extract the sentence vector of each sentence.

[0191] (3) Combine and splice the sentence vectors extracted from the above two parts to obtain a complete text emotion feature vector.

[0192] (4) Use the obtained text emotion features as the input of the self-attention mechanism, and convert the text emotion feature vectors into L groups of feature vectors according to the number of words in the text. The size of each group of feature vectors is Where L is the number of audio frames and X is the dimension of the text emotion feature vector. The text self-attention emotion feature obtained through the self-attention mechanism is as follows:

[0193]

[0194]

[0195] Where is the weight coefficient of the l-th group of feature vectors, represents the l-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W X is the trainable linear transformation parameter vector, and F X is the speech self-attention emotion feature through the self-attention mechanism.

[0196] 6. Emotion Feature Fusion Module Based on Cross-Attention Mechanism

[0197] Figure 6 The following is the principle block diagram of the emotion feature fusion module based on the cross-attention mechanism provided by the present invention, and the specific steps are as follows:

[0198] Taking the combined expression-action as an example to illustrate the cross-attention mechanism. First, the expression self-attention emotion feature and the action self-attention emotion feature are used as the input of the expression-action cross-attention mechanism to obtain the expression-action cross-attention mechanism as follows:

[0199]

[0200]

[0201]

[0202] Where softmax represents the normalized exponential function, and are the trainable parameter vector matrices, Concat(·) represents vector concatenation, F A_E represents the relative action self-attention emotion feature, the feature after weighting the expression self-attention emotion feature, F E_A represents the relative expression self-attention emotion feature, the feature after weighting the action self-attention emotion feature, F EA represents the expression-action cross-attention emotion feature.

[0203] According to the above method, similarly, the expression-speech cross-attention emotion feature F EV , the expression-text cross-attention emotion feature F EX , the action-speech cross-attention emotion feature F AV, the action-text mutual attention emotion feature F AX , the speech-text mutual attention emotion feature F VX .

[0204] Further, the obtaining of the complete multi-modal emotion feature by cascading and fusing the mutual attention emotion features specifically includes:

[0205] F = Concat(F EA , F EV , F EX , F AV , F AX , F VX ) (12)

[0206] where F represents the obtained complete multi-modal emotion feature.

[0207] 7. Context Emotion Feature Module Based on Graph Convolutional Neural Network

[0208] Figure 7 The following is the principle block diagram of the context emotion feature module based on the graph convolutional neural network provided by the present invention, and its specific steps are as follows:

[0209] (1). First, take the above-obtained multi-modal emotion feature as the node set G of the graph structure v ∈R N×f , where N is the number of adjacent samples, and f is the dimension of the multi-modal emotion feature extracted using the attention mechanism. Use the similarity formula shown below to construct the adjacency matrix A to define the edge set information G between adjacent samples e .

[0210]

[0211] where, v i and v j represent the multi-modal emotion feature vectors of the i-th and j-th adjacent samples, ‖·‖ represents the modulus operation, sim represents the cosine similarity. When sim≥0.75, the element a i,j of the adjacency matrix A = 1; when sim<0.75, the element a i,j of the adjacency matrix A = 0.

[0212] (2). Take the graph structure as the input of the graph convolutional neural network, and perform graph convolution operation on the graph structure. The graph convolution operation is shown as follows

[0213]

[0214] where, the degree matrix a i,j represents the element of the i-th row and j-th column of the adjacency matrix A, Hl represents the output of the l-th layer graph convolutional layer, and H 0 = F, where F is the feature vector output by the attention mechanism, and W l are trainable linear transformation parameters, l represents the number of layers of graph convolution, and σ(·) represents the ReLU activation function.

[0215] 8. Sentiment Classification and Recognition Module

[0216] Figure 8 The following is a schematic block diagram of a sentiment classification and recognition module provided by the present invention, and the specific steps are as follows:

[0217] (1) Normalize the obtained multi-modal context sentiment features.

[0218] (2) Select the RBF kernel function as the kernel function of the SVM classifier, use cross-validation to find the optimal parameter γ in the RBF kernel function, and use the optimal parameter γ to train the SVM model.

[0219] (3) Transmit the multi-modal data into the trained SVM model to obtain sentiment labels, specifically including pain, sadness, calm, nausea, comfort, and happiness.

[0220] 9. Voice Interaction and Display Module

[0221] The voice interaction and display module performs corresponding voice interactions according to the obtained sentiment label results and simultaneously displays the recognition results.

Claims

1. A multi-modal emotion recognition method for a medical care robot, comprising the following steps:

1. Perform multi-modal emotion information collection to collect the patient's video information and audio information.

2. Extract facial self-attention emotion features and motion self-attention emotion features based on the video information, and extract speech self-attention emotion features and text self-attention emotion features based on the audio information. The extraction of facial self-attention emotion features is to extract the emotion feature vector of the patient's facial expression according to the video information and transform it into facial self-attention emotion features through the self-attention mechanism. The extraction of the emotion feature vector of the patient's facial expression specifically includes: First, use a pre-trained model and a combined network to extract video features. At the same time, use a facial expression recognition library to detect the key points of the human face in the frame-divided pictures. Then, calculate the center point and the distance from each key point to the center point to obtain the features of the key points. Finally, splice the two parts of the features to form a complete facial expression emotion feature. The transformation into facial self-attention emotion features through the self-attention mechanism specifically includes: Take the obtained expression and emotion features as the input of the self-attention mechanism, and convert the expression and emotion feature vectors into I groups of feature vectors according to the video frames corresponding to the video information. The size of each group of feature vectors is where I is the number of video frames and E is the dimension of the expression and emotion feature vector. The expression self-attention emotion features obtained by the self-attention mechanism are as follows: wherein is the weight coefficient of the i-th group of feature vectors, represents the i-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, and W E is the trainable linear transformation parameter vector, and F E is the expression self-attention emotion feature through the self-attention mechanism; The extraction of motion self-attention emotion features is to extract the emotion feature vector of the patient's motion according to the video information and transform it into motion self-attention emotion features through the self-attention mechanism. The extraction of the emotion feature vector of the patient's motion specifically includes: First, use a pre-trained model and a combined network to extract video features. At the same time, use a human pose detection library to detect the joint points of the human body in the frame-divided pictures. Then, calculate the center of gravity of the human body and the distance and angle from each joint point to the center of gravity to obtain the features of the human body joint points. Finally, splice the two parts of the features to form a complete motion emotion feature. The transformation into motion self-attention emotion features through the self-attention mechanism specifically includes: Take the obtained action emotion features as the input of the self-attention mechanism, and convert the action emotion feature vectors into J groups of feature vectors according to the number of video frames corresponding to the video information. The size of each group of feature vectors is where J is the number of video frames and A is the dimension of the action emotion feature vector; the action self-attention emotion features obtained by the self-attention mechanism are as follows: wherein is the weight coefficient of the j-th group of feature vectors, represents the j-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W A is a trainable linear transformation parameter vector, F A is the action self-attention emotion feature through the self-attention mechanism; The extraction of speech self-attention emotion features is to extract the emotion feature vector of the patient's speech according to the audio information and transform it into speech self-attention emotion features through the self-attention mechanism. The extraction of the emotion feature vector of the patient's speech specifically includes: preprocess the collected audio signal and draw a spectrogram, then construct and train a convolutional neural network, and finally use the trained network to extract speech emotion features. The conversion into speech self-attention emotion features through the self-attention mechanism specifically includes: using the obtained speech emotion features as the input of the self-attention mechanism, and converting the speech emotion feature vectors into K groups of feature vectors according to the number of speech frames of each audio message, where the size of each group of feature vectors is where K is the number of audio frames and V is the dimension of the expression emotion feature vector; the expression self-attention emotion features obtained through the self-attention mechanism are as follows: wherein is the weight coefficient of the k-th group of eigenvectors, represents the k-th group of eigenvectors, exp represents the exponential function with the natural constant e as the base, W V is the trainable linear transformation parameter vector, F V is the speech self-attention emotion feature through the self-attention mechanism; The extraction of text self-attention emotion features is used to extract the emotion feature vector of the patient's text according to the audio information and transform it into text self-attention emotion features through the self-attention mechanism. The extraction of the emotion feature vector of the patient's text specifically includes: first, use an end-to-end ASR system to extract the audio signal into text information, then use a pre-trained model to extract the word vector features in the text information, then add the word vectors of each word in each sentence to obtain a sentence vector, and at the same time use a pre-trained model to extract the sentence vector of each sentence. Finally, combine and splice the two parts of the extracted sentence vectors to form a complete text emotion feature. The conversion into text self-attention sentiment features through the self-attention mechanism specifically includes: using the obtained text sentiment features as the input of the self-attention mechanism, converting the text sentiment feature vector into L groups of feature vectors according to the number of words in the text, and the size of each group of feature vectors is where L is the number of audio frames, and X is the dimension of the text sentiment feature vector; the text self-attention sentiment features obtained through the self-attention mechanism are as follows: wherein is the weight coefficient of the l-th group of feature vectors, represents the l-th group of feature vectors, exp represents the exponential function with the natural constant e as the base, W X is a trainable linear transformation parameter vector, F X is the speech self-attention emotion feature through the self-attention mechanism; 3. Perform emotion feature fusion based on the mutual attention mechanism on the 4 types of self-attention emotion features to obtain complete multi-modal emotion features.

4. Perform context emotion feature extraction based on a graph convolutional neural network on the multi-modal emotion features. Obtain multi-modal emotion features containing context information.

5. Perform sentiment classification and recognition on the multi-modal sentiment features containing context information to obtain sentiment label results; 6. Perform voice interaction and display according to the sentiment label results.

2. The multi-modal emotion recognition method for a medical care robot according to claim 1, characterized in that: The sentiment feature fusion based on the mutual attention mechanism described in step 3 is used to combine the self-attention sentiment features in pairs to obtain mutual attention sentiment features, and the mutual attention sentiment features are fused through concatenation to obtain complete multi-modal sentiment features; The specific process of combining the above self-attention sentiment features in pairs to obtain mutual attention sentiment features includes: Taking the combination of expression-action as an example to illustrate the mutual attention mechanism. First, take the expression self-attention sentiment feature and the action self-attention sentiment feature as the input of the expression-action mutual attention mechanism, and obtain the expression-action mutual attention mechanism as shown below: where softmax represents the normalized exponential function, F AI =(F A ) T W A , F EI =(F E ) T W E , and are trainable parameter vector matrices, Concat(·) represents vector concatenation, F A_E represents the relative action self-attention emotion feature, the feature after weighting the expression self-attention emotion feature, F E_A represents the relative expression self-attention emotion feature, the feature after weighting the action self-attention emotion feature, F EA represents the expression-action mutual attention emotion feature; According to the above method, the expression-speech mutual attention emotion feature F can be obtained in the same way EV , the expression-text mutual attention emotion feature F EX , the action-speech mutual attention emotion feature F AV , the action-text mutual attention emotion feature F AX , the speech-text mutual attention emotion feature F VX ; The specific process of fusing the mutual attention sentiment features through concatenation to obtain complete multi-modal sentiment features includes: F = Concat(F EA , F EV , F EX , F AV , F AX , F VX )#(12) Where F represents the obtained complete multi-modal sentiment feature.

3. The multimodal emotion recognition method for a medical care robot according to claim 1, wherein: The context sentiment feature extraction based on the graph convolutional neural network described in step 4 is used to obtain multi-modal sentiment features containing context information by passing the multi-modal sentiment features through the graph convolutional neural network; The specific process of obtaining multi-modal sentiment features containing context information by passing through the graph convolutional neural network includes: First, the above-mentioned multi-modal sentiment features are used as the node set \(G\) of the graph structure v \(\in\mathbb{R}\) N×f , where \(N\) is the number of adjacent samples, \(f\) is the dimension of the multi-modal sentiment features extracted using the attention mechanism, and the similarity formula shown below is used to construct the adjacency matrix \(A\) to define the edge set information \(G\) between adjacent samples e ; where, v i and v j represent the multi-modal sentiment feature vectors of the i-th and j-th adjacent samples, ||·|| represents the modulus operation, sim represents the cosine similarity. When sim ≥ 0.75, the element a i,j of the adjacency matrix A is 1; when sim < 0.75, the element a i,j of the adjacency matrix A is 0; After that, take the graph structure as the input of the graph convolutional neural network and perform graph convolutional operations on the graph structure; the graph convolutional operation is shown in the following formula, Among them, the degree matrix a i,j represents the element at the i-th row and j-th column of the adjacency matrix A, and H l represents the output of the l-th layer of the graph convolutional layer, and H 0 = F, where F is the feature vector output after the attention mechanism, and W l is a trainable linear transformation parameter, l represents the number of layers of graph convolution, and σ(·) represents the ReLU activation function.

4. The multimodal emotion recognition method for a medical care robot according to claim 1, wherein: The sentiment classification and recognition described in step 5 is used to load the above multi-modal sentiment features containing context information into a pre-constructed and trained sentiment classification model for classification and recognition to obtain multiple discrete sentiment label recognition results; The multiple discrete sentiment labels specifically include pain, sadness, calmness, nausea, comfort, and happiness.

5. The multimodal emotion recognition method for a medical care robot according to claim 1, wherein: The voice interaction and display described in step 6 are used to provide real-time voice and display feedback on the sentiment analysis results to the patient and the doctor, and perform corresponding voice interactions with the patient and the doctor according to the obtained sentiment label recognition results.

Citation Information

Patent Citations

  • Attention mechanism-based multi-modal emotion feature learning and recognition method

    CN111753549A

  • Multi-modal emotion recognition method based on attention enhancing mechanism

    CN112489635A