An Emotional Recognition Method, Device, Storage Medium and Equipment for an Escort Robot
By combining multi-signal fusion method with images, speech, text and pulse waves, the problem of errors in single signal source recognition in escort robots is solved, and accurate recognition of human emotions is achieved, especially the recognition of masked emotions is achieved.
Patent Information
- Application Number
- CN202310018201.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-01-06
AI Technical Summary
The existing emotional recognition methods of escort robots mainly rely on a single signal source, resulting in a high error rate of emotion recognition and the inability to accurately identify emotions covered by humans.
The multi-signal fusion method is adopted, combining the image emotion recognition model, speech emotion recognition model, text emotion recognition model and pulse wave emotion recognition model, and the convolution network and probability processing formulas combine face images, speech, text and pulse wave characteristics to identify the emotions of the target person.
It improves the accuracy of emotional recognition, can identify the emotions expressed by the target character and intentionally concealed emotions, and enhances the comprehensiveness and accuracy of emotional recognition.
Smart Images

Figure CN116257816B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application belong to the technical field of artificial intelligence, and particularly relate to a method, device, storage medium and equipment for emotion recognition of a companion robot. Background Art
[0002] Affective computing is one of the key technologies for a companion robot to recognize human emotions. Currently, the existing methods include non-physiological signal recognition methods such as face micro-expression recognition methods and speech intonation recognition methods. However, these methods are all applications of a single signal source in the companion robot scenario, such as only recognizing emotions through face micro-expressions or only recognizing emotions through the speech intonation of a person speaking. Since humans are very complex, they may cover up their actual inner emotions for some reason, resulting in incorrect emotion recognition. Summary of the Invention
[0003] The present application provides a method, device, storage medium and equipment for emotion recognition of a companion robot, which can improve the accuracy of emotion recognition.
[0004] The specific technical solutions are as follows:
[0005] In a first aspect, an embodiment of the present application provides a method for emotion recognition of a companion robot, the method including:
[0006] Performing feature extraction on a face image of a target person based on an image emotion recognition model to obtain face image emotion features;
[0007] Performing feature extraction on the speech information and mouth image of the target person based on a speech emotion recognition model to obtain speech emotion features;
[0008] Performing feature extraction on the text information corresponding to the speech information based on a text emotion recognition model to obtain text emotion features;
[0009] Fusing the face image emotion features, the speech emotion features and the text emotion features to obtain first fused emotion features;
[0010] Performing feature extraction on the pulse wave of the target person based on a pulse wave emotion recognition model to obtain pulse wave emotion features, and fusing the pulse wave emotion features with the first fused emotion features to obtain second fused emotion features;
[0011] Obtain the image emotion recognition result obtained by the image emotion recognition model for the facial image emotion features, the speech emotion recognition result obtained by the speech emotion recognition model for the speech emotion features, the text emotion recognition result obtained by the text emotion recognition model for the text emotion features, and the pulse wave emotion recognition result obtained by the pulse wave emotion recognition model for the second fused emotion features respectively, where the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result all include the probabilities of each emotion category;
[0012] Determine the target emotion of the target person according to the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result.
[0013] In one implementation manner, the fusing the facial image emotion features, the speech emotion features, and the text emotion features to obtain the first fused emotion feature includes:
[0014] Process the facial image emotion features, the speech emotion features, and the text emotion features respectively according to a preset convolutional network formula to obtain a first convolutional emotion feature corresponding to the facial image emotion features, a second convolutional emotion feature corresponding to the speech emotion features, and a third convolutional emotion feature corresponding to the text emotion features;
[0015] Concatenate the first convolutional emotion feature and the second convolutional emotion feature to obtain a first concatenated emotion feature, and process the concatenated first concatenated emotion feature according to the preset convolutional network formula to obtain a fourth convolutional emotion feature;
[0016] Concatenate the third convolutional emotion feature and the fourth convolutional emotion feature to obtain a second concatenated emotion feature, and process the second concatenated emotion feature according to the preset convolutional network formula to obtain the first fused emotion feature;
[0017] Wherein, the preset convolutional network formula includes: Y = F(X) + X, where Y represents the calculation result of the preset convolutional network formula, X represents the emotion feature to be calculated, and F(X) represents a function determined according to the weight layer and the rectified linear unit (Relu) function in the convolutional network.
[0018] In one implementation manner, the fusing the pulse wave emotion feature and the first fused emotion feature to obtain the second fused emotion feature includes:
[0019] Process the pulse wave emotion features according to the preset convolutional network formula to obtain fifth convolutional emotion features;
[0020] Concatenate the fifth convolutional emotion features and the first fused emotion features to obtain the second fused emotion features.
[0021] In one implementation, the feature extraction of the speech information and mouth image of the target person based on the speech emotion recognition model to obtain speech emotion features includes:
[0022] Based on the speech emotion recognition model, perform feature extraction on the speech information and the mouth image respectively to obtain speech sub-emotion features corresponding to the speech information and mouth image emotion features corresponding to the mouth image;
[0023] Based on the speech emotion recognition model, perform convolutional processing on the concatenated speech sub-emotion features and the mouth image emotion features to obtain the speech emotion features.
[0024] In one implementation, the determination of the target emotion of the target person according to the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result includes:
[0025] Determine the target emotion w of the target person according to a preset probability processing formula;
[0026] The preset probability processing formula includes:
[0027]
[0028] Among them,
[0029]
[0030]
[0031]
[0032]
[0033] The λ represents the weight for adjusting the emotion recognition result based on the non-physiological signal model. The non-physiological signal model includes the image emotion recognition model, the speech emotion recognition model, and the text emotion recognition model. The represents the probability of the i-th emotion in the pulse wave emotion recognition result. The n represents the total number of emotion categories. The P image represents the image emotion recognition result. The Indicates the probability of the first emotion to the nth emotion in the image emotion recognition result, the P voice Indicates the speech emotion recognition result, the Indicates the probability of the first emotion to the nth emotion in the speech emotion recognition result, the P text Indicates the text emotion recognition result, the Indicates the probability of the first emotion to the nth emotion in the text emotion recognition result.
[0034] In a second aspect, an embodiment of the present application provides an emotion recognition device for a companion robot, the device includes:
[0035] A first extraction unit, configured to extract features from a face image of a target person based on an image emotion recognition model to obtain face image emotion features;
[0036] A second extraction unit, configured to extract features from the speech information and mouth image of the target person based on a speech emotion recognition model to obtain speech emotion features;
[0037] A third extraction unit, configured to extract features from the text information corresponding to the speech information based on a text emotion recognition model to obtain text emotion features;
[0038] A first fusion unit, configured to fuse the face image emotion features, the speech emotion features, and the text emotion features to obtain first fusion emotion features;
[0039] A fourth extraction unit, configured to extract features from the pulse wave of the target person based on a pulse wave emotion recognition model to obtain pulse wave emotion features;
[0040] A second fusion unit, configured to fuse the pulse wave emotion features with the first fusion emotion features to obtain second fusion emotion features;
[0041] An acquisition unit, configured to respectively acquire an image emotion recognition result obtained by the image emotion recognition model for emotion recognition of the face image emotion features, a speech emotion recognition result obtained by the speech emotion recognition model for emotion recognition of the speech emotion features, a text emotion recognition result obtained by the text emotion recognition model for emotion recognition of the text emotion features, and a pulse wave emotion recognition result obtained by the pulse wave emotion recognition model for emotion recognition of the second fusion emotion features, wherein the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result all include the probabilities of each emotion category;
[0042] A determination unit for determining the target emotion of the target person according to the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result.
[0043] In one implementation, the first fusion unit includes:
[0044] A first calculation module for processing the face image emotion feature, the speech emotion feature, and the text emotion feature respectively according to a preset convolutional network formula to obtain a first convolutional emotion feature corresponding to the face image emotion feature, a second convolutional emotion feature corresponding to the speech emotion feature, and a third convolutional emotion feature corresponding to the text emotion feature;
[0045] A first splicing module for splicing the first convolutional emotion feature and the second convolutional emotion feature to obtain a first spliced emotion feature;
[0046] A second calculation module for processing the spliced first spliced emotion feature according to the preset convolutional network formula to obtain a fourth convolutional emotion feature;
[0047] A second splicing module for splicing the third convolutional emotion feature and the fourth convolutional emotion feature to obtain a second spliced emotion feature;
[0048] A third calculation module for processing the second spliced emotion feature according to the preset convolutional network formula to obtain the first fusion emotion feature;
[0049] Wherein, the preset convolutional network formula includes: Y = F(X) + X, where Y represents the calculation result of the preset convolutional network formula, X represents the emotion feature to be calculated, and F(X) represents a function determined according to the weight layer and the rectified linear unit (Relu) function in the convolutional network.
[0050] In one implementation, the second fusion unit includes:
[0051] A fourth calculation module for processing the pulse wave emotion feature according to the preset convolutional network formula to obtain a fifth convolutional emotion feature;
[0052] A third splicing module for splicing the fifth convolutional emotion feature and the first fusion emotion feature to obtain the second fusion emotion feature.
[0053] In one implementation, the second extraction unit includes:
[0054] An extraction module, configured to respectively extract features from the speech information and the mouth image based on the speech emotion recognition model, to obtain a speech sub-emotion feature corresponding to the speech information and a mouth image emotion feature corresponding to the mouth image;
[0055] A convolution module, configured to perform convolution processing on the spliced speech sub-emotion feature and the mouth image emotion feature based on the speech emotion recognition model to obtain the speech emotion feature.
[0056] In one implementation, the determination unit is configured to determine the target emotion w of the target person according to a preset probability processing formula;
[0057] The preset probability processing formula includes:
[0058]
[0059] Wherein,
[0060]
[0061]
[0062]
[0063]
[0064] The λ represents a weight for adjusting the emotion recognition result based on the non-physiological signal model. The non-physiological signal model includes the image emotion recognition model, the speech emotion recognition model, and the text emotion recognition model. The represents the probability of the i-th emotion in the pulse wave emotion recognition result. The n represents the total number of emotion categories. The P image represents the image emotion recognition result. The represents the probabilities of the first emotion to the n-th emotion in the image emotion recognition result. The P voice represents the speech emotion recognition result. The represents the probabilities of the first emotion to the n-th emotion in the speech emotion recognition result. The P text represents the text emotion recognition result. The represents the probabilities of the first emotion to the n-th emotion in the text emotion recognition result.
[0065] In a third aspect, an embodiment of the present application provides a storage medium, on which executable instructions are stored. When the instructions are executed by a processor, the processor implements the method according to any implementation of the first aspect.
[0066] Fourthly, an embodiment of the present application provides an electronic device, including:
[0067] One or more processors;
[0068] A storage device for storing one or more programs,
[0069] wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any implementation manner of the first aspect.
[0070] As can be seen from the above, the emotional recognition method, device, storage medium and equipment of the accompanying robot provided by the embodiments of the present application can not only perform emotional recognition on the emotional features after the fusion of non-physiological signal features (including facial image emotional features, speech emotional features, text emotional features) and physiological signal features (i.e., pulse wave emotional features) based on the pulse wave emotional recognition model to obtain the pulse wave emotional recognition result, but also can determine the final target emotion comprehensively according to the image emotional recognition result obtained by performing emotional recognition on the facial image emotional features by the image emotional recognition model, the speech emotional recognition result obtained by performing emotional recognition on the speech emotional features by the speech emotional recognition model, the text emotional recognition result obtained by performing emotional recognition on the text emotional features by the text emotional recognition model, and the pulse wave emotional recognition result. Therefore, compared with performing emotional recognition only through a single non-physiological signal feature, the embodiments of the present application can implement emotional recognition by fusing multiple signals of images, speech, content, and pulse waves, so that not only the emotions shown on the surface of the target person can be recognized, but also the emotions deliberately concealed by the target person can be recognized, and thus the accuracy of emotional recognition can be improved. Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages at the same time.
[0071] The innovative points of the embodiments of the present application include:
[0072] 1. The embodiments of the present application can implement emotional recognition by fusing multiple signals of images, speech, content, and pulse waves, so that not only the emotions shown on the surface of the target person can be recognized, but also the emotions deliberately concealed by the target person can be recognized, and thus the accuracy of emotional recognition can be improved.
[0073] 2. When the embodiments of the present application fuse multiple non - physiological signal features with physiological signal features, the multiple non - physiological signal features can be first fused according to the high - low levels of different feature expressions, and then the fused non - physiological signal features are fused with the physiological signal features, thereby improving the accuracy of emotion feature fusion and further improving the accuracy of the pulse - wave emotion recognition model in recognizing masked emotions. Among them, when fusing multiple non - physiological signal features, first use the preset convolutional network formula to calculate the first convolutional emotion feature corresponding to the facial image emotion feature, the second convolutional emotion feature corresponding to the speech emotion feature, and the third convolutional emotion feature corresponding to the text emotion feature respectively. Then, splice the first convolutional emotion feature and the second convolutional emotion feature to obtain the first spliced emotion feature, and process the spliced first spliced emotion feature according to the preset convolutional network formula to obtain the fourth convolutional emotion feature. Finally, after splicing the third convolutional emotion feature and the fourth convolutional emotion feature to obtain the second spliced emotion feature, process the second spliced emotion feature according to the preset convolutional network formula to obtain the first fused emotion feature; when fusing the fused non - physiological signal features with the physiological signal features, it is also possible to first process the pulse - wave emotion feature according to the preset convolutional network formula to obtain the fifth convolutional emotion feature, and after splicing the fifth convolutional emotion feature and the first fused emotion feature to obtain the third spliced emotion feature, obtain the second fused emotion feature. It can be seen from this that the embodiments of the present application can implement a feature fusion method of sequential cascaded residuals.
[0074] 3. When the embodiments of the present application extract speech emotion features based on the speech emotion recognition model, it does not simply extract the speech sub - emotion features in the speech information, but fuses the speech sub - emotion features in the speech information with the mouth image emotion features in the mouth image, thereby improving the accuracy of the speech emotion features and further improving the accuracy of recognizing speech emotions based on the speech emotion recognition model.
[0075] 4. When the embodiments of the present application determine the target emotion of the target person according to the preset probability processing formula, this preset probability processing formula combines the importance of each emotion recognition result to the target emotion, thereby improving the accuracy of the target emotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0077] Figure 1Schematic flowchart of an emotion recognition method for a companion robot provided by an embodiment of the present application;
[0078] Figure 2 Composition example diagram of an F(X) provided by an embodiment of the present application;
[0079] Figure 3 Example diagram of emotion feature fusion provided by an embodiment of the present application;
[0080] Figure 4 Block diagram of the composition of an emotion recognition device for a companion robot provided by an embodiment of the present application. Detailed implementation manners
[0081] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0082] It should be noted that the terms "including" and "having" in the embodiments of the present application and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0083] Figure 1 Schematic flowchart of an emotion recognition method for a companion robot provided by an embodiment of the present application. This method can be applied to a terminal, such as a companion robot, or to a server. The method may include the following steps:
[0084] S110: Extract features from the face image of the target person based on the image emotion recognition model to obtain the face image emotion features.
[0085] Among them, the image emotion recognition model is trained according to multiple face sample images and the emotion category annotation information of each face sample image. The emotion category annotation information of the face sample image can be manually annotated, that is, a person manually determines the emotion category represented by each face sample image by viewing the micro-expressions in the face sample image. The dimension of the face image emotion features can be N, and the face image emotion features can be represented by V1.
[0086] Emotional categories include a variety of discrete emotional states. For example, academically, there are usually 8 emotions, including four positive emotions: happiness, trust, surprise, and anticipation, and four negative emotions: anger, sadness, disgust, and fear.
[0087] S120: Extract features from the speech information and mouth image of the target person based on the speech emotion recognition model to obtain speech emotion features.
[0088] Among them, the speech emotion recognition model is trained based on multiple speech sample information, the corresponding mouth sample image for each speech sample information, and the emotion category annotation information for the speech sample information and the mouth sample image. The emotion category annotation information for the speech sample information and the mouth sample image can be manually annotated, that is, manually determine the represented emotion category by combining each speech sample information and its corresponding mouth image.
[0089] Whether it is the training process of the speech emotion recognition model or the model application process after training, the specific implementation method of extracting features from the speech information and mouth image of the target person based on the speech emotion recognition model to obtain speech emotion features can include: extracting features from the speech information and mouth image respectively based on the speech emotion recognition model to obtain the speech sub-emotion features corresponding to the speech information and the mouth image emotion features corresponding to the mouth image; performing convolution processing on the concatenated speech sub-emotion features and mouth image emotion features based on the speech emotion recognition model to obtain speech emotion features.
[0090] The speech information of the target person includes M discrete speech waveform points, that is, the original continuous speech information is sampled to obtain M discrete speech waveform points. The M discrete speech waveform points can be used as input and sent to a one-dimensional convolutional network for feature extraction to obtain N-dimensional speech sub-emotion features, which can be represented by V2. While the target person emits speech information, an image of the mouth area can be collected, normalized and then sent to a convolutional network to extract mouth image emotion features V3 with a dimension of N. After concatenating the speech sub-emotion features V2 and the mouth image emotion features V3, [V2, V3] is obtained, and [V2, V3] is sent to a convolutional network to obtain the top-level output speech emotion feature V4 of the speech emotion recognition model.
[0091] S130: Extract features from the text information corresponding to the speech information based on the text emotion recognition model to obtain text emotion features.
[0092] Among them, the text emotion recognition model is trained based on multiple text sample information and the emotion category annotation information of each text sample information. The emotion category annotation information of the text sample information can be manually annotated. The text sample information is the text information corresponding to the speech sample information, that is, the text information converted from the speech sample information using AI (Artificial Intelligence) technology is used as the text sample information.
[0093] When inputting the text information into the text emotion recognition model, the text information can be first converted into word vectors through tools such as bag of word or word2vec, and then the word vectors are input into the text emotion recognition model for feature extraction and emotion recognition. In addition, the dimension of the text emotion feature in this step can be N, and the text emotion feature can be represented by V5.
[0094] S140: Fuse the facial image emotion feature, the speech emotion feature, and the text emotion feature to obtain the first fused emotion feature.
[0095] When fusing the facial image emotion feature, the speech emotion feature, and the text emotion feature, they can be fused according to the high and low levels of different feature expressions. The levels of the facial image emotion feature and the speech emotion feature are higher than those of the text emotion feature, that is, the facial image emotion feature and the speech emotion feature can better reflect the emotions shown by the target person. Therefore, the facial image emotion feature and the speech emotion feature can be first fused, and then the fused feature and the text emotion feature are fused to finally obtain the first fused emotion feature.
[0096] The specific implementation method includes steps A1 - A3:
[0097] A1. Process the facial image emotion feature, the speech emotion feature, and the text emotion feature respectively according to the preset convolutional network formula to obtain the first convolutional emotion feature corresponding to the facial image emotion feature, the second convolutional emotion feature corresponding to the speech emotion feature, and the third convolutional emotion feature corresponding to the text emotion feature.
[0098] Among them, the preset convolutional network formula includes: Y = F(X) + X, where Y represents the calculation result of the preset convolutional network formula, X represents the emotion feature to be calculated, and F(X) represents a function determined according to the weight layer in the convolutional network and the rectified linear unit (Relu) function, such as Figure 2As shown in the figure, F(X) is obtained by passing a weight layer in the convolutional network through the ReLU function and then connecting another weight layer. In addition, in practical applications, there may be differences in the convolutional parameters used for different emotion features during convolutional calculations. Therefore, when different emotion features are fused, the actual F(X) used may vary. To accurately distinguish different F(X), the form Fi(X) is used for differentiation and representation below.
[0099] As Figure 3 shown in the figure, when the facial image emotion feature, speech emotion feature, and text emotion feature are represented by V1, V4, and V5 respectively, the first convolutional emotion feature V7 = F1(V1) + V1, the second convolutional emotion feature V8 = F2(V4) + V4, and the third convolutional emotion feature V11 = F4(V5) + V5.
[0100] A2. Concatenate the first convolutional emotion feature and the second convolutional emotion feature to obtain the first concatenated emotion feature, and process the concatenated first concatenated emotion feature according to the preset convolutional network formula to obtain the fourth convolutional emotion feature.
[0101] As Figure 3 shown in the figure, the first concatenated emotion feature V9 = [V7, V8] = [F1(V1) + V1, F2(V4) + V4], and the fourth convolutional emotion feature V10 = F3(V9) + V9 = F3([F1(V1) + V1, F2(V4) + V4]) + [F1(V1) + V1, F2(V4) + V4].
[0102] A3. Concatenate the third convolutional emotion feature and the fourth convolutional emotion feature to obtain the second concatenated emotion feature, and process the second concatenated emotion feature according to the preset convolutional network formula to obtain the first fused emotion feature.
[0103] As Figure 3 shown in the figure, when the second concatenated emotion feature is represented by V12 and the first fused emotion feature is represented by V13,
[0104] V12 = [V10, V11] = [F3([F1(V1) + V1, F2(V4) + V4]) + [F1(V1) + V1, F2(V4) + V4], F4(V5) + V5];
[0105] V13 = F5(V12) + V12 = F5([F3([F1(V1) + V1, F2(V4) + V4]) + [F1(V1) + V1, F2(V4) + V4], F4(V5) + V5]) + [F3([F1(V1) + V1, F2(V4) + V4]) + [F1(V1) + V1, F2(V4) + V4], F4(V5) + V5].
[0106] S150: Extract features from the pulse wave of the target person based on the pulse wave emotion recognition model to obtain pulse wave emotion features, and fuse the pulse wave emotion features with the first fused emotion features to obtain the second fused emotion features.
[0107] The pulse wave belongs to a physiological signal and is closely related to the true expression of emotions. It can significantly distinguish true emotions from false surface emotions. The pulse wave emotion recognition model in the embodiments of this application fuses the pulse wave and all the aforementioned non - physiological signal features, thereby improving the accuracy of recognizing masked emotions.
[0108] The aforementioned image emotion recognition model, speech emotion recognition model, and text emotion recognition model can all be trained independently. After training these three models, the pulse wave emotion recognition model is trained. When training the pulse wave emotion recognition model, multiple pulse wave samples, the emotion category annotation information of each pulse wave sample, the face image emotion features extracted by the image emotion recognition model from the face sample images, the speech emotion features extracted by the speech emotion recognition model from the speech sample information and mouth sample images, and the text emotion features extracted by the text emotion recognition model from the text sample information can be used as the input of the pulse wave emotion recognition model. The pulse wave emotion recognition model first learns according to multiple pulse wave samples and the emotion category annotation information of each pulse wave sample, extracts the pulse wave emotion features of the pulse wave samples, then fuses the pulse wave emotion features with the other three non - physiological emotion features, and performs emotion recognition based on the fused emotion features. Through the above - mentioned phased training, the purpose of simultaneously recognizing surface emotions and masked emotions can be achieved.
[0109] The method of fusing the pulse wave emotion features with the first fused emotion features includes: processing the pulse wave emotion features according to a preset convolutional network formula to obtain the fifth convolutional emotion feature; splicing the fifth convolutional emotion feature with the first fused emotion features to obtain the second fused emotion features.
[0110] Such as Figure 3As shown, when the emotional feature of the pulse wave is represented by V6, the fifth convolutional emotional feature V14 = F6(V6) + V6, and the second fused emotional feature V15 = [V13, V14] = [F5([F3([F1(V1) + V1, F2(V4) + V4]) + [F1(V1) + V1, F2(V4) + V4], F4(V5) + V5]) + [F3([F1(V1) + V1, F2(V4) + V4]) + [F1(V1) + V1, F2(V4) + V4], F4(V5) + V5], V14 = F6(V6) + V6].
[0111] It should be added that there is a close connection among the face image, voice information, mouth image, text information, and pulse wave of the above-mentioned target person. They are the face image, voice information, mouth image, text information corresponding to the voice information, and the pulse wave of the target person collected by the escort robot when the target person is speaking within a period of time. Among them, one voice information may correspond to multiple face images and mouth images.
[0112] S160: Respectively obtain the image emotion recognition result obtained by the image emotion recognition model for recognizing the emotion features of the face image, the voice emotion recognition result obtained by the voice emotion recognition model for recognizing the emotion features of the voice, the text emotion recognition result obtained by the text emotion recognition model for recognizing the emotion features of the text, and the pulse wave emotion recognition result obtained by the pulse wave emotion recognition model for recognizing the second fused emotion feature.
[0113] Among them, the image emotion recognition result, voice emotion recognition result, text emotion recognition result, and pulse wave emotion recognition result all include the probabilities of each emotion category.
[0114] After each emotion recognition model extracts the corresponding emotion features, a classifier can also be used to classify and recognize the emotion features to obtain the probability of each emotion category. Among them, the classifier can adopt a softmax classifier or other classifiers.
[0115] It should be added that when the above-mentioned each emotion recognition model performs feature extraction, it can adopt networks such as CNN (Convolutional Neural Networks) or transformer. Among them, the text emotion recognition model can also use networks such as RNN (Recurrent Neural Network) or LSTM (Long Short-Term Memory) for feature extraction.
[0116] S170: Determine the target emotion of the target person according to the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result.
[0117] The methods for determining the target emotion of the target person according to the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result include but are not limited to the following two:
[0118] The first one: First, calculate the average probability of the same emotion category in the four emotion recognition results respectively, and then select the emotion category with the largest average probability as the target emotion.
[0119] The second one: Determine the target emotion w of the target person according to the preset probability processing formula;
[0120] The preset probability processing formula includes:
[0121]
[0122] Among them,
[0123]
[0124]
[0125]
[0126]
[0127] λ represents the weight for adjusting the emotion recognition result based on the non - physiological signal model. The non - physiological signal model includes the image emotion recognition model, the speech emotion recognition model, and the text emotion recognition model. represents the probability of the i - th emotion in the pulse wave emotion recognition result, n represents the total number of emotion categories, P image represents the image emotion recognition result, represents the probability of the first emotion to the n - th emotion in the image emotion recognition result, P voice represents the speech emotion recognition result, represents the probability of the first emotion to the n - th emotion in the speech emotion recognition result, P text represents the text emotion recognition result, represents the probability of the first emotion to the n - th emotion in the text emotion recognition result.
[0128] When λ = 0, the independent recognition result based on the non - physiological signal has no impact on the final result. Since the pulse wave emotion recognition model has already incorporated the features of the non - physiological signal, the value of λ in the embodiments of the present application can be less than 0.3.
[0129] The preset probability processing formula in the second method combines the importance of each emotion recognition result for the target emotion, so that compared with the first method, the accuracy of the target emotion can be further improved.
[0130] The emotion recognition method for the escort robot provided by the embodiments of the present application can not only perform emotion recognition on the emotion features after the fusion of non-physiological signal features (including facial image emotion features, speech emotion features, text emotion features) and physiological signal features (i.e., pulse wave emotion features) based on the pulse wave emotion recognition model to obtain the pulse wave emotion recognition result, but also can comprehensively determine the final target emotion according to the image emotion recognition result obtained by performing emotion recognition on the facial image emotion features by the image emotion recognition model, the speech emotion recognition result obtained by performing emotion recognition on the speech emotion features by the speech emotion recognition model, the text emotion recognition result obtained by performing emotion recognition on the text emotion features by the text emotion recognition model, and the pulse wave emotion recognition result. Therefore, compared with performing emotion recognition only through a single non-physiological signal feature, the embodiments of the present application can realize emotion recognition by fusing multiple signals of images, speech, content, and pulse waves, so that not only the emotions shown on the surface of the target person can be recognized, but also the emotions deliberately concealed by the target person can be recognized, and thus the accuracy of emotion recognition can be improved.
[0131] Corresponding to the above method embodiments, the embodiments of the present application provide an emotion recognition device for an escort robot, as Figure 4 shown, the device includes:
[0132] A first extraction unit 210, configured to extract features from the facial image of the target person based on the image emotion recognition model to obtain facial image emotion features;
[0133] A second extraction unit 220, configured to extract features from the speech information and mouth image of the target person based on the speech emotion recognition model to obtain speech emotion features;
[0134] A third extraction unit 230, configured to extract features from the text information corresponding to the speech information based on the text emotion recognition model to obtain text emotion features;
[0135] A first fusion unit 240, configured to fuse the facial image emotion features, the speech emotion features, and the text emotion features to obtain a first fused emotion feature;
[0136] A fourth extraction unit 250, configured to extract features from the pulse wave of the target person based on the pulse wave emotion recognition model to obtain pulse wave emotion features;
[0137] A second fusion unit 260, configured to fuse the pulse wave emotion feature and the first fusion emotion feature to obtain a second fusion emotion feature;
[0138] An acquisition unit 270, configured to respectively acquire an image emotion recognition result obtained by the image emotion recognition model performing emotion recognition on the face image emotion feature, a voice emotion recognition result obtained by the voice emotion recognition model performing emotion recognition on the voice emotion feature, a text emotion recognition result obtained by the text emotion recognition model performing emotion recognition on the text emotion feature, and a pulse wave emotion recognition result obtained by the pulse wave emotion recognition model performing emotion recognition on the second fusion emotion feature, wherein the image emotion recognition result, the voice emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result all include probabilities of each emotion category;
[0139] A determination unit 280, configured to determine the target emotion of the target person according to the image emotion recognition result, the voice emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result.
[0140] In one implementation, the first fusion unit 240 includes:
[0141] A first calculation module, configured to respectively process the face image emotion feature, the voice emotion feature, and the text emotion feature according to a preset convolutional network formula to obtain a first convolutional emotion feature corresponding to the face image emotion feature, a second convolutional emotion feature corresponding to the voice emotion feature, and a third convolutional emotion feature corresponding to the text emotion feature;
[0142] A first splicing module, configured to splice the first convolutional emotion feature and the second convolutional emotion feature to obtain a first spliced emotion feature;
[0143] A second calculation module, configured to process the spliced first spliced emotion feature according to the preset convolutional network formula to obtain a fourth convolutional emotion feature;
[0144] A second splicing module, configured to splice the third convolutional emotion feature and the fourth convolutional emotion feature to obtain a second spliced emotion feature;
[0145] A third calculation module, configured to process the second spliced emotion feature according to the preset convolutional network formula to obtain the first fusion emotion feature;
[0146] Among them, the preset convolutional network formula includes: Y = F(X) + X, where Y represents the calculation result of the preset convolutional network formula, X represents the emotional feature to be calculated, and F(X) represents a function determined according to the weight layer and the rectified linear unit (Relu) function in the convolutional network.
[0147] In one implementation, the second fusion unit 260 includes:
[0148] A fourth calculation module, configured to process the pulse wave emotional feature according to the preset convolutional network formula to obtain a fifth convolutional emotional feature;
[0149] A third splicing module, configured to splice the fifth convolutional emotional feature and the first fusion emotional feature to obtain the second fusion emotional feature.
[0150] In one implementation, the second extraction unit 220 includes:
[0151] An extraction module, configured to respectively extract features from the voice information and the mouth image based on the voice emotion recognition model to obtain a voice sub-emotional feature corresponding to the voice information and a mouth image emotional feature corresponding to the mouth image;
[0152] A convolutional module, configured to perform convolutional processing on the spliced voice sub-emotional feature and the mouth image emotional feature based on the voice emotion recognition model to obtain the voice emotional feature.
[0153] In one implementation, the determination unit 280 is configured to determine the target emotion w of the target person according to a preset probability processing formula;
[0154] The preset probability processing formula includes:
[0155]
[0156] Among them,
[0157]
[0158]
[0159]
[0160]
[0161] λ represents the weight for adjusting the emotion recognition result based on the non-physiological signal model. The non-physiological signal model includes the image emotion recognition model, the voice emotion recognition model, and the text emotion recognition model. represents the probability of the i-th emotion in the pulse wave emotion recognition result, where n represents the total number of emotion categories, and P image represents the image emotion recognition result, and the represents the probability of the first emotion to the n-th emotion in the image emotion recognition result, and P voice represents the speech emotion recognition result, and the represents the probability of the first emotion to the n-th emotion in the speech emotion recognition result, and P text represents the text emotion recognition result, and the represents the probability of the first emotion to the n-th emotion in the text emotion recognition result.
[0162] The emotion recognition device for the companion robot provided by the embodiment of the present application can not only perform emotion recognition on the emotion features after fusing non-physiological signal features (including face image emotion features, speech emotion features, text emotion features) and physiological signal features (i.e., pulse wave emotion features) based on the pulse wave emotion recognition model to obtain the pulse wave emotion recognition result, but also can, according to the image emotion recognition result obtained by performing emotion recognition on the face image emotion features by the image emotion recognition model, the speech emotion recognition result obtained by performing emotion recognition on the speech emotion features by the speech emotion recognition model, the text emotion recognition result obtained by performing emotion recognition on the text emotion features by the text emotion recognition model, and the pulse wave emotion recognition result, comprehensively determine the final target emotion. Therefore, compared with performing emotion recognition only through a single non-physiological signal feature, the embodiment of the present application can realize emotion recognition by fusing multiple signals of images, speech, content, and pulse waves, so that not only the emotion shown on the surface of the target person can be recognized, but also the emotion deliberately concealed by the target person can be recognized, thereby improving the accuracy of emotion recognition.
[0163] Based on the above method embodiment, another embodiment of the present application provides a storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor implements the method as described above.
[0164] Based on the above method embodiment, another embodiment of the present application provides an electronic device, including:
[0165] One or more processors;
[0166] A storage device for storing one or more programs,
[0167] wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0168] The above device embodiments correspond to the method embodiments and have the same technical effects as the method embodiments. For specific descriptions, please refer to the method embodiments. The device embodiments are obtained based on the method embodiments. For specific descriptions, please refer to the method embodiment section and will not be elaborated here. Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing this application.
[0169] Those of ordinary skill in the art can understand that the modules in the device in the embodiments can be distributed in the devices in the embodiments according to the descriptions in the embodiments, or can be correspondingly changed and located in one or more devices different from this embodiment. The modules in the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for emotion recognition of an escort robot, characterized in that, The method includes: Performing feature extraction on the face image of the target person based on an image emotion recognition model to obtain face image emotion features; Performing feature extraction on the speech information and mouth image of the target person based on a speech emotion recognition model to obtain speech emotion features; Performing feature extraction on the text information corresponding to the speech information based on a text emotion recognition model to obtain text emotion features; Fusing the face image emotion features, the speech emotion features, and the text emotion features to obtain first fused emotion features; Performing feature extraction on the pulse wave of the target person based on a pulse wave emotion recognition model to obtain pulse wave emotion features, and fusing the pulse wave emotion features with the first fused emotion features to obtain second fused emotion features; Respectively obtaining an image emotion recognition result obtained by the image emotion recognition model for emotion recognition of the face image emotion features, a speech emotion recognition result obtained by the speech emotion recognition model for emotion recognition of the speech emotion features, a text emotion recognition result obtained by the text emotion recognition model for emotion recognition of the text emotion features, and a pulse wave emotion recognition result obtained by the pulse wave emotion recognition model for emotion recognition of the second fused emotion features, where the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result all include probabilities of each emotion category; Determining the target emotion of the target person according to the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result.
2. The method according to claim 1, wherein The step of fusing the face image emotion features, the speech emotion features, and the text emotion features to obtain first fused emotion features includes: Processing the face image emotion features, the speech emotion features, and the text emotion features respectively according to a preset convolutional network formula to obtain a first convolutional emotion feature corresponding to the face image emotion features, a second convolutional emotion feature corresponding to the speech emotion features, and a third convolutional emotion feature corresponding to the text emotion features; Concatenating the first convolutional emotion feature and the second convolutional emotion feature to obtain a first concatenated emotion feature, and processing the concatenated first concatenated emotion feature according to the preset convolutional network formula to obtain a fourth convolutional emotion feature; Concatenating the third convolutional emotion feature and the fourth convolutional emotion feature to obtain a second concatenated emotion feature, and processing the second concatenated emotion feature according to the preset convolutional network formula to obtain the first fused emotion features; Wherein, the preset convolutional network formula includes: Y = F(X) + X, where Y represents the calculation result of the preset convolutional network formula, X represents the emotion feature to be calculated, and F(X) represents a function determined according to the weight layer and the rectified linear unit (Relu) function in the convolutional network.
3. The method according to claim 2, characterized in that The step of fusing the pulse wave emotion features with the first fused emotion features to obtain second fused emotion features includes: Process the pulse wave emotion features according to the preset convolutional network formula to obtain fifth convolutional emotion features; Concatenate the fifth convolutional emotion features and the first fused emotion features to obtain the second fused emotion features.
4. The method according to claim 1, wherein The feature extraction of the speech information and mouth image of the target person based on the speech emotion recognition model to obtain speech emotion features includes: Based on the speech emotion recognition model, respectively extract features from the speech information and the mouth image to obtain speech sub-emotion features corresponding to the speech information and mouth image emotion features corresponding to the mouth image; Based on the speech emotion recognition model, perform convolutional processing on the concatenated speech sub-emotion features and the mouth image emotion features to obtain the speech emotion features.
5. The method according to any one of claims 1-4, characterized in that, The determination of the target emotion of the target person according to the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result includes: Determine the target emotion w of the target person according to the preset probability processing formula; The preset probability processing formula includes: Wherein, The λ represents the weight for adjusting the emotion recognition result based on the non-physiological signal model, and the non-physiological signal model includes the image emotion recognition model, the speech emotion recognition model, and the text emotion recognition model. The represents the probability of the i-th emotion in the pulse wave emotion recognition result, the n represents the total number of emotion categories, and the P image represents the image emotion recognition result, and the represents the probabilities of the first emotion to the n-th emotion in the image emotion recognition result, and the P voice represents the speech emotion recognition result, and the represents the probabilities of the first emotion to the n-th emotion in the speech emotion recognition result, and the P text represents the text emotion recognition result, and the represents the probabilities of the first emotion to the n-th emotion in the text emotion recognition result.
6. An emotional recognition device for an escort robot, characterized in that, The device includes: A first extraction unit for extracting features from the face image of the target person based on the image emotion recognition model to obtain face image emotion features; A second extraction unit for extracting features from the speech information and mouth image of the target person based on the speech emotion recognition model to obtain speech emotion features; A third extraction unit for extracting features from the text information corresponding to the speech information based on the text emotion recognition model to obtain text emotion features; A first fusion unit for fusing the face image emotion features, the speech emotion features, and the text emotion features to obtain a first fused emotion feature; A fourth extraction unit for extracting features from the pulse wave of the target person based on the pulse wave emotion recognition model to obtain pulse wave emotion features; A second fusion unit for fusing the pulse wave emotion features and the first fused emotion features to obtain a second fused emotion feature; An acquisition unit for respectively acquiring an image emotion recognition result obtained by the image emotion recognition model for emotion recognition of the face image emotion features, a speech emotion recognition result obtained by the speech emotion recognition model for emotion recognition of the speech emotion features, a text emotion recognition result obtained by the text emotion recognition model for emotion recognition of the text emotion features, and a pulse wave emotion recognition result obtained by the pulse wave emotion recognition model for emotion recognition of the second fused emotion features, wherein the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result all include probabilities of each emotion category; A determination unit for determining the target emotion of the target person according to the image emotion recognition result, the speech emotion recognition result, the text emotion recognition result, and the pulse wave emotion recognition result.
7. The device according to claim 6, characterized in that, The first fusion unit includes: A first computing module, configured to process the facial image emotion feature, the speech emotion feature, and the text emotion feature respectively according to a preset convolutional network formula, to obtain a first convolutional emotion feature corresponding to the facial image emotion feature, a second convolutional emotion feature corresponding to the speech emotion feature, and a third convolutional emotion feature corresponding to the text emotion feature; A first splicing module, configured to splice the first convolutional emotion feature and the second convolutional emotion feature to obtain a first spliced emotion feature; A second computing module, configured to process the spliced first spliced emotion feature according to the preset convolutional network formula to obtain a fourth convolutional emotion feature; A second splicing module, configured to splice the third convolutional emotion feature and the fourth convolutional emotion feature to obtain a second spliced emotion feature; A third computing module, configured to process the second spliced emotion feature according to the preset convolutional network formula to obtain the first fused emotion feature; Wherein, the preset convolutional network formula includes: Y = F(X) + X, where Y represents the calculation result of the preset convolutional network formula, X represents the emotion feature to be calculated, and F(X) represents a function determined according to the weight layer and the rectified linear unit (Relu) function in the convolutional network.
8. The device according to claim 7, characterized in that, The second fusion unit includes: A fourth computing module, configured to process the pulse wave emotion feature according to the preset convolutional network formula to obtain a fifth convolutional emotion feature; A third splicing module, configured to splice the fifth convolutional emotion feature and the first fused emotion feature to obtain the second fused emotion feature.
9. A storage medium, characterized in that, It stores executable instructions, and when the instructions are executed by a processor, the processor implements the method according to any one of claims 1-5.
10. An electronic device, characterized in that, Including: One or more processors; A storage device, configured to store one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.
Citation Information
Patent Citations
Emotion recognition method and device, electronic equipment and storage medium
CN112926525A
Expression recognition method and related device
WO2020182121A1