Multi-modal user emotion recognition method and device, equipment and medium

By employing a multimodal user emotion recognition method that combines physiological, speech, and image data, and utilizing multiple models to predict user emotions, this approach solves the problem of traditional audio equipment being unable to recognize emotions in noisy environments. It achieves accurate emotion recognition and personalized music recommendations in noisy environments, thereby enhancing the user experience.

CN121963791APending Publication Date: 2026-05-01SHENZHEN AIRSMART TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN AIRSMART TECH CO LTD
Filing Date
2025-11-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional audio equipment relies on a single voice modality, which cannot accurately identify user emotions in noisy environments, resulting in an inability to recommend suitable music content and affecting user experience.

Method used

A multimodal user emotion recognition method is adopted. By acquiring physiological data, voice data, and image data, physiological features, voice features, semantic features, and facial features are extracted. Emotion prediction is performed using a gradient boosting tree model, a hybrid model of convolutional neural network-long short-term memory network, a BERT fine-tuning model, and a mobile network model. Finally, the comprehensive probability value is calculated based on the feature weight relationship to determine the user's emotion.

Benefits of technology

Even in noisy environments, it can accurately identify user emotions and recommend suitable music content, significantly improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963791A_ABST
    Figure CN121963791A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-mode user emotion recognition method, device and equipment and a medium, the method is applied to sound equipment, and the method comprises the steps that physiological data, voice data and image data of a target user are acquired; extracting physiological features from the physiological data; extracting voice features and semantic features from the voice data; extracting facial features from the image data; and determining the emotion of the target user according to the physiological features, the voice features, the semantic features and the facial features. Therefore, according to the technical scheme of the invention, the accurate judgment of the emotion of the target user is realized by integrating the four types of modal data of the physiological features, the voice features, the semantic features and the facial features. The four types of features respectively map the emotional state of the user from four dimensions of physiological reaction, acoustic expression, semantic connotation and visual expression to form a multi-dimensional complementary emotion recognition system. The mechanism can recommend appropriate music content to the user, the auditory demands of the user under different emotions are met, and the use experience of the user is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio equipment, and more particularly to a multimodal user emotion recognition method, device, equipment, and medium. Background Technology

[0002] With the development of the internet, audio equipment has been widely used in various scenarios such as conference rooms, living rooms, classrooms, and stages. Users can use audio equipment to amplify sound, enhance stage atmosphere, and perform other functions. As its popularity increases, users' demands for audio equipment experiences are constantly evolving, with "accurately recognizing user emotions and providing tailored services" becoming a key requirement. Audio equipment can sense the user's emotional state in real time and automatically recommend music that matches their current mood, thus accurately matching the auditory needs under different emotional states.

[0003] Currently, traditional emotion recognition solutions for audio equipment mostly rely on a single voice modality. Specifically, traditional audio equipment collects the user's voice data through a microphone. Then, it extracts voice features and semantic features from the voice data. Finally, it combines the voice features and semantic features to determine the user's emotion.

[0004] However, single-modal voice recognition has significant limitations. When audio equipment is in a noisy environment, it cannot accurately capture the user's voice data, leading to an inability to accurately identify the user's emotions. This problem directly prevents the audio equipment from recommending suitable music content, failing to meet the auditory needs of different emotions and severely impacting the user experience. Summary of the Invention

[0005] This application provides a multimodal user emotion recognition method, apparatus, device, and medium, aiming to solve the technical problem that it is difficult to accurately recognize user emotions using a single voice modality.

[0006] In a first aspect, embodiments of this application provide a multimodal user emotion recognition method, the method being applied to an audio device, comprising:

[0007] Acquire physiological, voice, and image data of the target user;

[0008] Extract the physiological characteristics of the target user from the physiological data;

[0009] Extract the voice features and semantic features of the target user from the voice data;

[0010] Extract the facial features of the target user from the image data;

[0011] The target user's emotion is determined based on the physiological characteristics, the voice characteristics, the semantic characteristics, and the facial characteristics.

[0012] In one embodiment, determining the target user's emotion based on the physiological characteristics, the voice characteristics, the semantic characteristics, and the facial characteristics includes:

[0013] The physiological characteristics are input into a preset gradient boosting tree model to obtain the first emotion probability vector of the target user. The first emotion probability vector includes probability values ​​corresponding to multiple preset emotion categories. The preset gradient boosting tree model is used to predict the emotion of the target user based on the physiological characteristics.

[0014] The speech features are input into a preset convolutional neural network-long short-term memory network hybrid model to obtain the second emotion probability vector of the target user. The second emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset convolutional neural network-long short-term memory network hybrid model is used to predict the emotion of the target user based on the speech features.

[0015] The semantic features are input into a preset BERT fine-tuning model to obtain the third emotion probability vector of the target user. The third emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset BERT fine-tuning model is used to predict the emotion of the target user based on the semantic features.

[0016] The facial features are input into a preset mobile network model to obtain the fourth emotion probability vector of the target user. The fourth emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset mobile network model is used to predict the emotion of the target user based on the facial features.

[0017] The target user's emotion is determined based on the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0018] In one embodiment, determining the target user's emotion based on the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector includes:

[0019] Determine the weighting relationships among the physiological features, the voice features, the semantic features, and the facial features;

[0020] Based on the weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset category emotions is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector;

[0021] The emotion category with the highest overall probability value among the various preset emotion categories is selected as the emotion of the target user.

[0022] In one embodiment, determining the weighted relationship among the physiological features, the voice features, the semantic features, and the facial features includes:

[0023] Obtain the scene type of the environment in which the target user is located;

[0024] According to the scene type, a preset scene type weight relationship table is queried to obtain the weight relationship. The preset scene type weight relationship table includes at least one scene type, and each scene type corresponds to a weight relationship. The weight relationship is the weight relationship between the user's physiological features, voice features, semantic features and facial features.

[0025] In one embodiment, after determining the weighting relationship among the physiological features, the voice features, the semantic features, and the facial features, the method further includes:

[0026] Based on the speech data, the signal-to-noise ratio is calculated;

[0027] Based on the signal-to-noise ratio, the weighting relationship between the physiological features, the speech features, the semantic features, and the facial features is adjusted;

[0028] The step of calculating the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories based on the weight relationship, according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector, includes:

[0029] Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0030] In one embodiment, after determining the weighting relationship among the physiological features, the voice features, the semantic features, and the facial features, the method further includes:

[0031] Calculate the image sharpness based on the image data;

[0032] Based on the image clarity, the weighting relationship between the physiological features, the voice features, the semantic features, and the facial features is adjusted;

[0033] The step of calculating the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories based on the weight relationship, according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector, includes:

[0034] Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0035] In one embodiment, after determining the weighting relationship among the physiological features, the voice features, the semantic features, and the facial features, the method further includes:

[0036] Based on the physiological data, the data missing rate is calculated;

[0037] Based on the data missing rate, the weight relationship between the physiological features, the voice features, the semantic features, and the facial features is adjusted;

[0038] The step of calculating the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories based on the weight relationship, according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector, includes:

[0039] Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0040] Secondly, embodiments of this application also provide a multimodal user emotion recognition device, which includes a unit for performing the above-described method.

[0041] Thirdly, embodiments of this application also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.

[0042] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the above-described method.

[0043] This application provides a multimodal user emotion recognition method, apparatus, device, and medium. The method, applied to an audio device, includes: acquiring physiological data, voice data, and image data of a target user; extracting physiological features of the target user from the physiological data; extracting voice and semantic features of the target user from the voice data; extracting facial features of the target user from the image data; and determining the target user's emotion based on the physiological features, voice features, semantic features, and facial features. Thus, the technical solution of this application acquires physiological data, voice data, and image data of the target user; then, extracts physiological features of the target user from the physiological data; extracts voice and semantic features of the target user from the voice data; and extracts facial features of the target user from the image data. Finally, the emotion of the target user is determined based on the physiological features, voice features, semantic features, and facial features. Therefore, the technical solution of this application achieves accurate determination of the target user's emotion by integrating four types of modal data: physiological features, voice features, semantic features, and facial features. These four types of features map the user's emotional state from four dimensions: physiological response, acoustic performance, semantic connotation, and visual expression, forming a multidimensional and complementary emotion recognition system. This application's technical solution employs a multimodal emotion recognition method. Even in noisy environments, it can effectively overcome the limitations of a single voice modality by leveraging non-voice modalities, thus completely resolving the technical problem in existing technologies where a single voice modality struggles to accurately identify user emotions. This mechanism can recommend suitable music content to users, satisfying their auditory needs under different emotional states and significantly improving the user experience. Attached Figure Description

[0044] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0047] Figure 1 A flowchart illustrating a multimodal user emotion recognition method provided in an embodiment of this application;

[0048] Figure 2 A schematic block diagram of a multimodal user emotion recognition device provided in this application embodiment;

[0049] Figure 3 A computer device provided in an embodiment of this application. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0051] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0052] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0053] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0054] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0055] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0056] To address the technical problem that existing technologies cannot accurately identify user emotions using a single voice modality, this application provides a multimodal user emotion recognition device that can accurately identify user emotions.

[0057] Figure 1 This is a flowchart illustrating a multimodal user emotion recognition method provided in an embodiment of this application. In one embodiment, the method is applied to an audio device, and the method includes steps S101-S105.

[0058] S101. Acquire the target user's physiological data, voice data, and image data.

[0059] Physiological data include, but are not limited to, heart rate, skin conductance level, skin conductance response, body surface temperature, and respiratory rate.

[0060] Voice data includes the voice signal of the target user.

[0061] The image data includes facial image data of the target user.

[0062] S102. Extract the physiological characteristics of the target user from the physiological data.

[0063] Physiological characteristics include, but are not limited to, heart rate characteristics, skin conductance characteristics, body surface temperature characteristics, and respiratory characteristics. Heart rate characteristics include, but are not limited to, mean heart rate, standard deviation of intervals between adjacent heartbeats, and standard deviation of all heartbeat intervals. Skin conductance characteristics include, but are not limited to, peak skin conductance response and baseline skin conductance levels. Body surface temperature characteristics include, but are not limited to, mean temperature and rate of temperature change. Respiratory characteristics include, but are not limited to, respiratory rate and respiratory depth.

[0064] S103. Extract the voice features and semantic features of the target user from the voice data.

[0065] Speech features include, but are not limited to, speech rate, intonation, volume, and speech spectrum.

[0066] Semantic features include, but are not limited to, the proportion of negative words, sentence length, and the frequency of modal particles.

[0067] S104. Extract the facial features of the target user from the image data.

[0068] Facial features include, but are not limited to, the degree of muscle contraction when frowning or smiling.

[0069] S105. Determine the target user's emotion based on physiological characteristics, voice characteristics, semantic characteristics, and facial characteristics.

[0070] This application's technical solution integrates four modalities of data—physiological features, voice features, semantic features, and facial features—to achieve accurate determination of the target user's emotions. These four types of features map the user's emotional state from four dimensions: physiological response, acoustic performance, semantic connotation, and visual expression, respectively, forming a multi-dimensional and complementary emotion recognition system.

[0071] This application provides a multimodal user emotion recognition method. The method is applied to an audio device and includes: acquiring physiological data, voice data, and image data of a target user; extracting physiological features of the target user from the physiological data; extracting voice features and semantic features of the target user from the voice data; extracting facial features of the target user from the image data; and determining the target user's emotion based on the physiological features, voice features, semantic features, and facial features. Therefore, this application's technical solution acquires physiological data, voice data, and image data of the target user; then, extracts physiological features of the target user from the physiological data; extracts voice features and semantic features of the target user from the voice data; and extracts facial features of the target user from the image data. Finally, it determines the target user's emotion based on the physiological features, voice features, semantic features, and facial features. Thus, this application's technical solution achieves accurate determination of the target user's emotion by integrating four types of modal data: physiological features, voice features, semantic features, and facial features. These four types of features map the user's emotional state from four dimensions: physiological response, acoustic performance, semantic connotation, and visual expression, forming a multidimensional and complementary emotion recognition system. This application's technical solution employs a multimodal emotion recognition method. Even in noisy environments, it can effectively overcome the limitations of a single voice modality by leveraging non-voice modalities, thus completely resolving the technical problem in existing technologies where a single voice modality struggles to accurately identify user emotions. This mechanism can recommend suitable music content to users, satisfying their auditory needs under different emotional states and significantly improving the user experience.

[0072] In one embodiment, S105 specifically includes the following steps: S1051-S1055.

[0073] S1051. Input the physiological characteristics into the preset gradient boosting tree model to obtain the first emotion probability vector of the target user.

[0074] The first emotion probability vector includes probability values ​​corresponding to various preset emotion categories. A preset gradient boosting tree model is used to predict the target user's emotions based on physiological characteristics. The various preset emotion categories include, but are not limited to, "happy," "neutral," "angry," "sad," "fearful," and "surprised."

[0075] It should be noted that the preset gradient boosting tree model was trained by the applicant based on practical experience. The construction process of the preset gradient boosting tree model is as follows:

[0076] First, training data was acquired. This training data was collected and extracted by the applicant. Specifically, physiological data from at least one user under different emotional states was collected. Next, physiological features were extracted from the collected physiological data and used as training data. Then, using a multi-class GBDT (Gross Gradient Boosting Tree) model as the base model, the training data was used to train, validate, and test this base model. Multiple decision trees were iteratively generated to gradually correct the prediction errors of the preceding model, ultimately resulting in a pre-defined gradient boosting tree model. This pre-defined gradient boosting tree model was used to generate probability distributions for multiple pre-defined emotional categories. For example, the generated probability distributions for multiple pre-defined emotional categories were: {"Happy": 0.02,"Neutral": 0.05,"Anger": 0.82,"Sadness": 0.03,"Fear": 0.06,"Surprise": 0.02}.

[0077] S1052. Input the speech features into a preset convolutional neural network-long short-term memory network hybrid model to obtain the second emotion probability vector of the target user.

[0078] The second emotion probability vector includes probability values ​​corresponding to various preset emotion categories. A preset convolutional neural network-long short-term memory network hybrid model is used to predict the target user's emotion based on speech features;

[0079] It should be noted that the pre-defined convolutional neural network-long short-term memory network hybrid model was trained by the applicant based on practical experience. The construction process of the pre-defined convolutional neural network-long short-term memory network hybrid model is as follows:

[0080] First, training data was acquired. This training data was collected and extracted by the applicant. Specifically, voice data from at least one user under different emotions was collected. Next, voice features were extracted from the voice data and used as training data. Then, a basic hybrid model combining a convolutional neural network (CNN) model and a long short-term memory (LSTM) network model was constructed. This hybrid model was trained, validated, and tested using the training data, ultimately resulting in a pre-defined CNN-LSTM hybrid model. It should be noted that the loss function of the pre-defined CNN-LSTM hybrid model is the category cross-entropy loss function. The pre-defined CNN-LSTM hybrid model is used to generate probability distributions for multiple pre-defined emotion categories. For example, the generated probability distributions for multiple pre-defined emotion categories are: {"Happy": 0.02, "Neutral": 0.01, "Angry": 0.87, "Sad": 0.03, "Fear": 0.05, "Surprised": 0.02}.

[0081] S1053. Input the semantic features into the preset BERT fine-tuning model to obtain the third emotion probability vector of the target user.

[0082] The third emotion probability vector includes probability values ​​corresponding to various preset emotion categories. A preset BERT fine-tuned model is used to predict the target user's emotion based on semantic features.

[0083] It should be noted that the preset BERT fine-tuning model was trained by the applicant based on practical experience. The direct Chinese translation of BERT fine-tuning model is "Transformer-based bidirectional encoder representation model." The construction process of the preset BERT fine-tuning model is as follows:

[0084] First, training data is acquired. This training data was collected and extracted by the applicant. Specifically, voice data from at least one user under different emotional states is collected. Next, the collected voice data is converted into text information. Then, the text information is cleaned, removing meaningless characters and controlling the length of sentences. Then, semantic features are extracted from the text information. Finally, based on the Chinese pre-trained model bert-base-chinese, the training data is used to train, validate, and test the Chinese pre-trained model bert-base-chinese to generate a preset BERT fine-tuning model. It should be noted that in this embodiment, the loss function of the Chinese pre-trained model bert-base-chinese is set to the category cross-entropy loss function. The preset BERT fine-tuning model is used to generate probability distributions for multiple preset emotion categories. For example, the generated probability distributions for multiple preset emotion categories are: {"Happy":0.03,"Neutral":0.02,"Angry":0.01,"Sad":0.02,"Fear":0.01,"Surprised":0.91}.

[0085] S1054. Input facial features into a preset mobile network model to obtain the fourth emotion probability vector of the target user.

[0086] The fourth emotion probability vector includes probability values ​​corresponding to various preset emotion categories. A preset mobile network model is used to predict the target user's emotion based on facial features. Specifically, the preset mobile network model is the preset MobileNet model.

[0087] It should be noted that the preset mobile network model was trained by the applicant based on practical experience. The construction process of the preset mobile network model is as follows:

[0088] First, training data was acquired, which was collected and extracted by the applicant. Specifically, image data of at least one user under different emotions was collected. Next, facial features were extracted from the image data as training data. Then, MobileNet was used as the base model, with the loss function set to the class cross-entropy loss function and the activation function set to ReLU6. Finally, the base model was trained, validated, and tested using the training data to obtain a preset mobile network model. This preset mobile network model is used to generate probability distributions for multiple preset emotion categories. For example, the generated probability distributions for multiple preset emotion categories are: {"Happy": 0.05, "Neutral": 0.02, "Angry": 0.01, "Sad": 0.01, "Fear": 0.03, "Surprised": 0.88}.

[0089] S1055. Determine the target user's emotion based on the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0090] In one embodiment, S1055 specifically includes the following steps: S10551-S10553.

[0091] S10551. Determine the weighting relationships among the physiological characteristics, voice characteristics, semantic characteristics, and facial characteristics of the target user.

[0092] In one embodiment, S10551 specifically includes the following steps: ab.

[0093] a. Obtain the scenario type of the target user's environment.

[0094] Scene types include, but are not limited to, conference rooms, living rooms, bedrooms, stages, parks, and squares. It should be noted that in this embodiment, the scene type can be obtained by the audio equipment itself sensing the surrounding environment, or by the user manually inputting the scene type of the target user's environment.

[0095] b. Query the preset scene type weight relationship table according to the scene type to obtain the weight relationship.

[0096] The preset scene type weight relationship table includes at least one scene type. Each scene type corresponds to a weight relationship. The weight relationship is the weighting relationship between a user's physiological features, voice features, semantic features, and facial features. For example, in a meeting scene, the weight relationship is: {physiological features: 0.8, voice features: 0.6, semantic features: 0.7, facial features: 0.7}. In a park scene, the weight relationship is: {physiological features: 0.8, voice features: 0.3, semantic features: 0.5, facial features: 0.7}. Clearly, the noisier the environment, the smaller the weight of voice and semantic features.

[0097] S10552. Based on the weight relationship, calculate the comprehensive probability value of each preset category emotion in multiple preset emotion categories according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector and the fourth emotion probability vector.

[0098] For example, the probability vector for the first emotion is: {"Happy":0.02,"Neutral":0.05,"Anger":0.82,"Sadness":0.03,"Fear":0.06,"Surprise":0.02}; the probability vector for the second emotion is: {"Happy":0.02,"Neutral":0.01,"Anger":0.87,"Sadness":0.03,"Fear":0.05,"Surprise":0.02}; and the probability vector for the third emotion is: {"Happy":0.03,"Neutral":0.02,"Anger":0.01,"Sadness":0.02,"Fear" The fourth emotion probability vector is: {"Happy": 0.05, "Neutral": 0.02, "Angry": 0.01, "Sad": 0.01, "Fear": 0.03, "Surprised": 0.88}; the weight relationship is: {physiological features: 0.8, voice features: 0.6, semantic features: 0.7, facial features: 0.7}. Therefore, the comprehensive probability value for the emotion "Happy" among the various preset emotion categories is: 0.02*0.8 + 0.02*0.6 + 0.03*0.7 + 0.05*0.7 = 0.084. The comprehensive probability value for the emotion "Neutral" is: 0.05*0.8 + 0.01*0.6 + 0.02*0.7 + 0.02*0.7 = 0.078. The calculation method for the overall probability value of other categories of emotions is the same as that for the overall probability value of the emotion "happiness", and will not be repeated here.

[0099] S10553. Select the preset category emotion with the highest comprehensive probability value from multiple preset emotion categories as the emotion of the target user.

[0100] In this embodiment, the emotion with the highest overall probability value among multiple preset emotion categories is selected as the emotion of the target user.

[0101] In one embodiment, after S10551 and before S10552, S1055 further includes S10554-S10555.

[0102] S10554. The signal-to-noise ratio is calculated based on the speech data.

[0103] In this application example, a general signal-to-noise ratio (SNR) calculation method can be used to calculate the SNR based on speech data. Further details will not be elaborated upon here.

[0104] S10555. Based on the signal-to-noise ratio, adjust the weight relationship between physiological features, speech features, semantic features and facial features.

[0105] It should be noted that a higher signal-to-noise ratio (SNR) indicates higher quality voice data, and thus greater reliability of the voice and semantic features extracted from it. Therefore, the weighting of voice and semantic features can be appropriately increased to improve the accuracy of user emotion recognition. Conversely, a lower SNR also indicates a lower SNR.

[0106] The above S10552 also specifically includes the following steps: A.

[0107] A. Based on the adjusted weighting relationship, calculate the comprehensive probability value of each preset emotion category among multiple preset emotion categories according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0108] In this embodiment, based on the adjusted weighting relationship, the calculated comprehensive probability value corresponding to each preset category of emotion is more accurate, thereby accurately identifying the user's emotions.

[0109] In one embodiment, after S10551 and before S10552, S1055 further includes S10556-S10557.

[0110] S10556. Calculate image sharpness based on image data.

[0111] S10557. Adjust the weight relationship between physiological features, voice features, semantic features and facial features based on image clarity.

[0112] It should be noted that S10556-S10557 will be explained in detail below.

[0113] In this embodiment, image sharpness is calculated, and the weighting relationship is adjusted based on the image sharpness. Higher image sharpness indicates more detailed facial features extracted from the image data, leading to more accurate user emotion recognition based on facial features. Therefore, the weighting ratio of facial features can be appropriately increased to improve the accuracy of user emotion recognition. Conversely, lower image sharpness also reduces the accuracy of user emotion recognition.

[0114] The above S10552 also specifically includes the following step: B.

[0115] B. Based on the adjusted weighting relationship, calculate the comprehensive probability value of each preset emotion category among multiple preset emotion categories according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0116] It should be noted that step B is the same as or similar to step A. This application will not elaborate further on this point.

[0117] In one embodiment, after S10551 and before S10552, S1055 further includes S10558-S10559.

[0118] S10558. Calculate the data missing rate based on physiological data.

[0119] It should be noted that the embodiments of this application can use a common method for calculating the missing data rate. This will not be elaborated upon further here.

[0120] S10559. Adjust the weight relationship between physiological features, voice features, semantic features and facial features based on the data missing rate.

[0121] In this embodiment, the data missing rate is calculated, and the weighting relationship is adjusted based on the data missing rate. A higher data missing rate indicates that the physiological features extracted from the physiological data are less accurate, making it difficult to accurately identify user emotions based on these features. Therefore, the weighting ratio of physiological features can be appropriately reduced to improve the accuracy of identifying user emotions. Conversely, a lower weighting ratio also indicates a lower weighting.

[0122] The above S10552 also specifically includes the following steps: C.

[0123] C. Based on the adjusted weighting relationship, calculate the comprehensive probability value of each preset emotion category among multiple preset emotion categories according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0124] It should be noted that step C is the same as or similar to step A. This application will not elaborate further on this point.

[0125] See Figure 2 , Figure 2 This is a schematic block diagram of a multimodal user emotion recognition device provided in an embodiment of this application. Corresponding to the above-described multimodal user emotion recognition method, this application also provides a multimodal user emotion recognition device. This multimodal user emotion recognition device includes a unit for executing the above-described multimodal user emotion recognition method, and can be configured in terminals such as desktop computers, tablet computers, and laptops. Specifically, the device is applied to audio equipment, and the multimodal user emotion recognition device includes:

[0126] The acquisition unit 201 is used to acquire physiological data, voice data and image data of the target user;

[0127] The first extraction unit 202 is used to extract the physiological characteristics of the target user from the physiological data;

[0128] The second extraction unit 203 is used to extract the voice features and semantic features of the target user from the voice data;

[0129] The third extraction unit 204 is used to extract the facial features of the target user from the image data;

[0130] The determining unit 205 is used to determine the emotion of the target user based on the physiological characteristics, the voice characteristics, the semantic characteristics and the facial characteristics.

[0131] In one embodiment, the determining unit 205 is specifically used for:

[0132] The physiological characteristics are input into a preset gradient boosting tree model to obtain the first emotion probability vector of the target user. The first emotion probability vector includes probability values ​​corresponding to multiple preset emotion categories. The preset gradient boosting tree model is used to predict the emotion of the target user based on the physiological characteristics.

[0133] The speech features are input into a preset convolutional neural network-long short-term memory network hybrid model to obtain the second emotion probability vector of the target user. The second emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset convolutional neural network-long short-term memory network hybrid model is used to predict the emotion of the target user based on the speech features.

[0134] The semantic features are input into a preset BERT fine-tuning model to obtain the third emotion probability vector of the target user. The third emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset BERT fine-tuning model is used to predict the emotion of the target user based on the semantic features.

[0135] The facial features are input into a preset mobile network model to obtain the fourth emotion probability vector of the target user. The fourth emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset mobile network model is used to predict the emotion of the target user based on the facial features.

[0136] The target user's emotion is determined based on the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0137] In one embodiment, the determining unit 205 is further specifically used for:

[0138] Determine the weighting relationships among the physiological features, the voice features, the semantic features, and the facial features;

[0139] Based on the weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset category emotions is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector;

[0140] The emotion category with the highest overall probability value among the various preset emotion categories is selected as the emotion of the target user.

[0141] In one embodiment, the determining unit 205 is further specifically used for:

[0142] Obtain the scene type of the environment in which the target user is located;

[0143] The preset scene type weight relationship table is queried according to the scene type to obtain the weight relationship. The preset scene type weight relationship table includes at least one scene type, and each scene type corresponds to a weight relationship. The weight relationship is the weight relationship between the user's physiological features, voice features, semantic features and facial features.

[0144] In one embodiment, the device further includes:

[0145] The calculation unit 206 is used to calculate the signal-to-noise ratio based on the speech data;

[0146] The adjustment unit 207 is used to adjust the weight relationship between the physiological features, the speech features, the semantic features and the facial features based on the signal-to-noise ratio;

[0147] The step of calculating the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories based on the weight relationship, according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector, includes:

[0148] Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0149] In one embodiment, the computing unit 206 is further configured to calculate the image sharpness based on the image data;

[0150] The adjustment unit 207 is also used to adjust the weight relationship between the physiological features, the voice features, the semantic features and the facial features based on the image clarity;

[0151] The determining unit 205 is further specifically used for:

[0152] Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0153] In one embodiment, the computing unit 206 is further configured to: calculate the data missing rate based on the physiological data;

[0154] The adjustment unit 207 is further configured to: adjust the weight relationship between the physiological features, the voice features, the semantic features and the facial features based on the data missing rate;

[0155] The determining unit 205 is further specifically used for:

[0156] Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

[0157] like Figure 3 As shown, this application provides a computer device including a processor 31, a communication interface 32, a memory 33, and a communication bus 34. The processor 31, the communication interface 32, and the memory 33 communicate with each other through the communication bus 34. The memory 33 is used to store computer programs.

[0158] In one embodiment of this application, when the processor 31 executes the program stored in the memory 33, it implements the control method for multimodal user emotion recognition provided in any of the foregoing method embodiments.

[0159] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program may be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0160] Therefore, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the multimodal user emotion recognition method provided in any of the foregoing method embodiments.

[0161] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk, or any other physical storage medium capable of storing program code. The computer-readable storage medium can be non-volatile or volatile.

[0162] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0163] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0164] The steps in the methods of this application embodiment can be adjusted, merged, or deleted according to actual needs. The units in the apparatus of this application embodiment can be merged, divided, or deleted according to actual needs. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0165] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0166] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0167] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Since these modifications and variations fall within the scope of the claims and their equivalents, this application also intends to include these modifications and variations.

[0168] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A multimodal user emotion recognition method, characterized in that, The method is applied to audio equipment, and the method includes: Acquire physiological, voice, and image data of the target user; Extract the physiological characteristics of the target user from the physiological data; Extract the voice features and semantic features of the target user from the voice data; Extract the facial features of the target user from the image data; The target user's emotion is determined based on the physiological characteristics, the voice characteristics, the semantic characteristics, and the facial characteristics.

2. The method according to claim 1, characterized in that, Determining the target user's emotion based on the physiological characteristics, voice characteristics, semantic characteristics, and facial characteristics includes: The physiological characteristics are input into a preset gradient boosting tree model to obtain the first emotion probability vector of the target user. The first emotion probability vector includes probability values ​​corresponding to multiple preset emotion categories. The preset gradient boosting tree model is used to predict the emotion of the target user based on the physiological characteristics. The speech features are input into a preset convolutional neural network-long short-term memory network hybrid model to obtain the second emotion probability vector of the target user. The second emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset convolutional neural network-long short-term memory network hybrid model is used to predict the emotion of the target user based on the speech features. The semantic features are input into a preset BERT fine-tuning model to obtain the third emotion probability vector of the target user. The third emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset BERT fine-tuning model is used to predict the emotion of the target user based on the semantic features. The facial features are input into a preset mobile network model to obtain the fourth emotion probability vector of the target user. The fourth emotion probability vector includes the probability values ​​corresponding to the various preset emotion categories. The preset mobile network model is used to predict the emotion of the target user based on the facial features. The target user's emotion is determined based on the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

3. The method according to claim 2, characterized in that, Determining the target user's emotion based on the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector includes: Determine the weighting relationships among the physiological features, the voice features, the semantic features, and the facial features; Based on the weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset category emotions is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector; The emotion category with the highest overall probability value among the various preset emotion categories is selected as the emotion of the target user.

4. The method according to claim 3, characterized in that, Determining the weighted relationship among the physiological features, the voice features, the semantic features, and the facial features includes: Obtain the scene type of the environment in which the target user is located; The preset scene type weight relationship table is queried according to the scene type to obtain the weight relationship. The preset scene type weight relationship table includes at least one scene type, and each scene type corresponds to a weight relationship. The weight relationship is the weight relationship between the user's physiological features, voice features, semantic features and facial features.

5. The method according to claim 3, characterized in that, After determining the weighted relationships among the physiological features, the voice features, the semantic features, and the facial features, the method further includes: Based on the speech data, the signal-to-noise ratio is calculated; Based on the signal-to-noise ratio, the weighting relationship between the physiological features, the speech features, the semantic features, and the facial features is adjusted; The step of calculating the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories based on the weight relationship, according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector, includes: Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

6. The method according to claim 3, characterized in that, After determining the weighted relationships among the physiological features, the voice features, the semantic features, and the facial features, the method further includes: Calculate the image sharpness based on the image data; Based on the image clarity, the weighting relationship between the physiological features, the voice features, the semantic features, and the facial features is adjusted; The step of calculating the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories based on the weight relationship, according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector, includes: Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

7. The method according to claim 3, characterized in that, After determining the weighted relationships among the physiological features, the voice features, the semantic features, and the facial features, the method further includes: Based on the physiological data, the data missing rate is calculated; Based on the data missing rate, the weight relationship between the physiological features, the voice features, the semantic features, and the facial features is adjusted; The step of calculating the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories based on the weight relationship, according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector, includes: Based on the adjusted weighting relationship, the comprehensive probability value corresponding to each preset category emotion among the multiple preset emotion categories is calculated according to the first emotion probability vector, the second emotion probability vector, the third emotion probability vector, and the fourth emotion probability vector.

8. A multimodal user emotion recognition device, characterized in that, The method is applied to audio equipment, including: The acquisition unit is used to acquire physiological data, voice data, and image data of the target user. The first extraction unit is used to extract the physiological characteristics of the target user from the physiological data; The second extraction unit is used to extract the voice features and semantic features of the target user from the voice data; The third extraction unit is used to extract the facial features of the target user from the image data; The determining unit is configured to determine the emotion of the target user based on the physiological characteristics, the voice characteristics, the semantic characteristics, and the facial characteristics.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the method as described in any one of claims 1 to 7.