A multi-modal fusion dynamic weighting recognition method and system based on facial emotion recognition and speech recognition
By employing a multimodal fusion dynamic weighting method combining facial emotion recognition and speech recognition, the accuracy problem of single-modal emotion recognition in complex environments is solved, enabling more accurate and comprehensive reflection of users' emotional states and timely intervention in their emotional states.
Patent Information
- Application Number
- CN202510733363.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing single-modal emotion recognition technologies suffer from decreased accuracy under conditions such as changes in lighting, facial occlusion, and environmental noise, making it difficult to accurately identify users' emotions in complex scenarios. In particular, they perform poorly in recognizing general or neutral emotions and cannot meet the needs of specific scenarios.
A multimodal fusion dynamic weighted recognition method based on facial emotion recognition and speech recognition is adopted. By generating an emotion probability distribution set, dynamically adjusting the weights, and comprehensively processing multimodal data, emotion recognition results are generated. Text information arbitration is introduced when necessary, and physiological and environmental monitoring are combined to improve accuracy.
It improves the accuracy and robustness of emotion recognition, accurately reflects users' true emotions in complex environments, provides timely intervention measures for negative emotions, and enhances user experience.
Smart Images

Figure CN120541788B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to a multi-modal fusion dynamic weighting recognition method and system based on facial emotion recognition and speech recognition. BACKGROUND
[0002] With the continuous development of technology, in order to meet the actual needs of more subdivided scenes, provide more personalized services for users, people need to accurately identify the emotions of users.
[0003] In the existing emotion recognition technology, facial expression image recognition or video recognition is mainly used, or the user's emotion is recognized by collecting the user's speech information. However, single facial emotion recognition or speech recognition has many shortcomings.
[0004] For example, relying only on facial expression recognition, the accuracy of recognition will be greatly reduced under conditions such as changes in lighting conditions, face occlusion, and in complex scenes (for example, scenes with multiple people) will cause large deviations in emotion recognition; and using speech recognition alone is easily disturbed by environmental noise and is difficult to accurately capture the complex emotional state of the speaker.
[0005] Therefore, the existing emotion recognition technology can only recognize emotions with obvious tendencies (such as obvious happiness, obvious sadness, etc.), and the recognition effect is poor for general emotions or near-neutral emotions. In some specific scenarios, such as nursing scenarios for the elderly, security management scenarios for bank business handling, etc., the existing recognition technology and recognition effect cannot meet the actual needs. SUMMARY
[0006] In order to overcome the above technical problems existing in the prior art, the present application provides a multi-modal fusion dynamic weighting recognition method and system based on facial emotion recognition and speech recognition, which fuses facial emotion recognition and speech recognition two modal information, uses dynamic weighting method to comprehensively process multi-modal data, which can effectively make up for the shortcomings of single mode, improve the accuracy and robustness of emotion recognition. Multi-modal fusion can obtain emotion features from different dimensions, reduce the influence of environmental factors on the recognition result, and more comprehensively reflect the real emotions of users.
[0007] To achieve the above object, the embodiment of the present application provides a multi-modal fusion dynamic weighting recognition method based on facial emotion recognition and speech recognition, which comprises: acquiring facial emotion recognition information and corresponding speech recognition information of a user based on a preset acquisition frequency; generating a corresponding first emotion probability distribution set based on the facial emotion recognition information, and generating a corresponding second emotion probability distribution set based on the speech recognition information; generating a first inclined emotion based on the first emotion probability distribution set, and generating a second inclined emotion based on the second emotion probability distribution set; judging whether the first inclined emotion and the second inclined emotion are consistent; if yes, determining a first weight of the first emotion probability distribution set and a second weight of the second emotion probability distribution set, and generating an emotion recognition result based on the first emotion probability distribution set, the first weight, the second emotion probability distribution set and the second weight.
[0008] Preferably, the generating of the emotion recognition result based on the first emotion probability distribution set, the first weight, the second emotion probability distribution set and the second weight comprises: determining a first emotion value of each emotion in the first emotion probability distribution set based on the first emotion probability distribution set and the first weight, and determining a second emotion value of each emotion in the second emotion probability distribution set based on the second emotion probability distribution set and the second weight; performing weighted average processing on the first emotion value and the second emotion value to obtain an intersection emotion value; determining an emotion proportion of each emotion based on the intersection emotion value; acquiring a preset sliding time window, and acquiring an emotion proportion at each sampling time within the sliding time window; determining emotion change data based on the emotion proportion at each sampling time; and determining an emotion recognition result of the current sliding time window based on the emotion change data.
[0009] Preferably, the method further comprises: in the case that the first inclined emotion and the second inclined emotion are inconsistent, acquiring a text information set interacted with the user; generating a third emotion probability distribution set based on the text information set; arbitrating the first inclined emotion or the second inclined emotion based on the third emotion probability distribution set to generate an arbitration result; and generating an emotion recognition result based on the arbitration result.
[0010] Preferably, the method further comprises: after the emotion recognition result is generated, judging whether the emotion recognition result is a negative emotion; if yes, acquiring an emotion source of the negative emotion; determining a corresponding interaction scheme based on the emotion source; determining a response device and response information based on the interaction scheme, and controlling the response device to perform a corresponding interaction action or outputting the response information.
[0011] Preferably, the acquiring the emotional source of the negative emotion comprises: generating inquiry information corresponding to the emotional recognition result; acquiring feedback information of the user for the inquiry information; determining an adverse factor causing the negative emotion based on the feedback information; performing type analysis on the adverse factor to generate a corresponding factor type; if the factor type is a physiological factor, determining a corresponding physiological part, and generating a corresponding emotional source based on the physiological part; and if the factor type is a psychological factor, generating an emotional source corresponding to the adverse factor.
[0012] Preferably, the method further comprises: after the inquiry information is generated and output, determining a preset waiting time; determining whether the user feeds back the inquiry information within the preset waiting time; if the inquiry information is not fed back, generating active care information; and performing a care interaction operation corresponding to the active care information, the care interaction operation comprising at least one of active concern inquiry, active approach and active interaction action.
[0013] Preferably, the user wears at least one physiological monitoring device, and the determining the corresponding physiological part comprises: acquiring physiological monitoring information of the at least one physiological monitoring device; determining a specific physiological monitoring device corresponding to the feedback information; determining whether specific physiological monitoring information corresponding to the specific physiological monitoring device is abnormal; if yes, determining a corresponding physiological part based on the specific physiological monitoring information; otherwise, acquiring abnormal monitoring information in the remaining physiological monitoring information, and determining a corresponding physiological part based on the abnormal monitoring information.
[0014] Preferably, the method further comprises: after the response device is controlled to perform the corresponding interaction action or output the response information, acquiring again an emotional recognition result for the user; determining whether the user is still in a negative emotion based on the acquired again emotional recognition result; and if yes, generating and feeding back corresponding alarm information.
[0015] Preferably, the method further comprises: before the facial emotional recognition information and the voice recognition information are acquired, acquiring environmental monitoring information of an environment in which the user is located; determining whether there is another person in a space in which the user is located based on the environmental monitoring information; if there is no other person, determining that the user is in a solitary environment; and acquiring facial emotional recognition information and corresponding voice recognition information of the user based on a preset acquisition frequency.
[0016] Correspondingly, the application further provides a multi-modal fusion dynamic weighting recognition device based on facial emotion recognition and speech recognition, the device comprising: an information acquisition unit configured to acquire facial emotion recognition information and corresponding speech recognition information of a user based on a preset acquisition frequency; a preliminary recognition unit configured to generate a corresponding first emotion probability distribution set based on the facial emotion recognition information and generate a corresponding second emotion probability distribution set based on the speech recognition information; an emotion generation unit configured to generate a first tendency emotion based on the first emotion probability distribution set and generate a second tendency emotion based on the second emotion probability distribution set; a judgment unit configured to judge whether the first tendency emotion and the second tendency emotion are consistent; and a result output unit configured to determine a first weight of the first emotion probability distribution set and a second weight of the second emotion probability distribution set in the case where the first tendency emotion and the second tendency emotion are consistent, and generate an emotion recognition result based on the first emotion probability distribution set, the first weight, the second emotion probability distribution set and the second weight.
[0017] In another aspect, the application further provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method provided by the embodiments of the application.
[0018] Through the technical solutions provided by the application, the application has at least the following technical effects:
[0019] By fusing facial emotion recognition and speech recognition two modal information, and using dynamic weighting method to comprehensively process multi-modal data, the deficiencies of single mode can be effectively made up, and the accuracy and robustness of emotion recognition can be improved. Multi-modal fusion can obtain emotion features from different dimensions, reduce the influence of environmental factors on the recognition result, and more comprehensively reflect the real emotions of the user.
[0020] Other features and advantages of the embodiments of the application will be described in detail in the following specific implementation part. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings are included to provide a further understanding of the embodiments of the application, and constitute a part of the specification, and are used together with the following specific implementation part to explain the embodiments of the application, but do not constitute a limitation on the embodiments of the application. In the drawings:
[0022] Figure 1 is a flowchart of a multi-modal fusion dynamic weighting recognition method based on facial emotion recognition and speech recognition provided by the embodiments of the application;
[0023] Figure 2 is a functional module schematic block diagram of a multi-modal fusion dynamic weighting recognition system based on facial emotion recognition and speech recognition provided by the embodiments of the application. Detailed Implementation
[0024] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0025] In this invention, the terms "system" and "network" are used interchangeably. "Multiple" refers to two or more; therefore, in this invention, "multiple" can also be understood as "at least two." "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, it should be understood that in the description of this invention, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.
[0026] Existing emotion recognition technologies often rely on a single modality, such as facial expression recognition or speech recognition alone, making it difficult to comprehensively and accurately capture a user's true emotions. Moreover, a single modality is susceptible to environmental interference, such as light affecting facial recognition and noise affecting speech recognition, and it cannot integrate multi-dimensional emotional features, leading to biased recognition results that cannot reliably reflect the user's overall emotional state.
[0027] Please see Figure 1 This invention provides a multimodal fusion dynamic weighted recognition method based on facial emotion recognition and speech recognition. The method includes: acquiring a user's facial emotion recognition information and corresponding speech recognition information based on a preset acquisition frequency; generating a first emotion probability distribution set based on the facial emotion recognition information and a second emotion probability distribution set based on the speech recognition information; generating a first tendency emotion based on the first emotion probability distribution set and a second tendency emotion based on the second emotion probability distribution set; determining whether the first tendency emotion and the second tendency emotion are consistent; if so, determining a first weight of the first emotion probability distribution set and a second weight of the second emotion probability distribution set; and generating an emotion recognition result based on the first emotion probability distribution set, the first weight, the second emotion probability distribution set, and the second weight.
[0028] In a possible embodiment, first, facial emotion recognition information (such as facial expression images, facial muscle movement data, etc.) and corresponding speech recognition information (such as speech audio, tone of speech, speech speed, keywords, etc.) of a user are continuously acquired at a preset acquisition frequency. Then, the acquired facial emotion recognition information is analyzed and processed to generate a corresponding first emotion probability distribution set, for example, facial expression images are recognized by using an existing emotion recognition algorithm to generate facial emotion probabilities, for example, facial expression images are acquired once every 100 ms, so that 10 facial emotion probabilities are generated per second, thereby generating the first emotion probability distribution set, which contains multiple emotion types identified from facial information and corresponding preliminary probabilities or intensity values; the speech recognition information is processed to generate a second emotion probability distribution set, which contains multiple emotion types identified from speech information and corresponding preliminary probabilities or intensity values. Then, the first emotion probability distribution set and the second emotion probability distribution set are analyzed respectively, and a first tendency emotion (i.e., the most likely dominant emotion under the facial modality) and a second tendency emotion (i.e., the most likely dominant emotion under the speech modality) are determined by means of statistics, weighting, etc., for example, the first tendency emotion can be determined according to the proportion, change trend, etc. of each facial emotion probability in the first emotion probability distribution set. In the embodiment of the present application, the tendency emotion can be defined as a dominant emotion type with a probability exceeding a preset threshold (such as 40%).
[0029] Since in people's daily life, a single information source cannot accurately determine the user's emotion, for example, a person may show a smiling face but actually have a low mood and express negative language, or may express positive language but actually have a lost facial expression, therefore, after the first tendency emotion corresponding to the facial expression and the second tendency emotion corresponding to the speech are generated respectively, it is determined whether the two tendency emotions are consistent, if consistent, it means that the two modalities reach a consensus on the dominant emotion, at this time, a first weight of the first emotion probability distribution set and a second weight of the second emotion probability distribution set are determined, wherein the first weight and the second weight can be preset or dynamically adjusted according to historical data, modality reliability, etc., finally, based on the two emotion probability distribution sets and their weights, a final emotion recognition result is generated by means of weighted fusion, etc., which comprehensively reflects the user's emotional state.
[0030] In another possible embodiment, the weights of the facial emotion and the speech emotion can be dynamically determined according to the confidence scores of the facial emotion and the speech emotion, so as to improve the accuracy of the weights, for example, in the embodiment of the present application, the formula of the first weight is , and the formula of the second weight is , wherein C f is the facial emotion confidence, C vVoice emotion confidence. In the implementation process, the fixed weight determined by the technician can be calculated in advance, but when the condition for triggering weight adjustment is detected, such as detecting high environmental noise level, large face cover, etc., the weight is automatically adjusted dynamically:
[0031] In an embodiment, when the environmental noise decibel value exceeds the preset threshold (such as 60dB), the voice modal weight is reduced And the face modal weight is increased When the face cover area exceeds the preset threshold (such as the cover area exceeds 30%), the face modal weight is reduced And the voice modal weight is increased For example, in the embodiment of the present application, the two weights are normalized, such as agreeing that the sum of the two weights is equal to 1, that is, the dynamic adjustment of the two weights affects each other to overcome the interference of environmental factors on multi-modal emotion recognition and improve the accuracy of multi-modal emotion recognition.
[0032] In 1000 test data, the emotion recognition accuracy of the present application is 92%, which is significantly improved compared with single face recognition (78%) and voice recognition (75%). The robustness of the dynamic weighting algorithm in noisy environment (signal-to-noise ratio <10dB) is 30% higher than that of fixed weight.
[0033] By fusing facial emotion recognition and voice recognition two modal information, using dynamic weighting method to comprehensively process multi-modal data, the shortcomings of single modal can be effectively made up, and the accuracy and robustness of emotion recognition can be improved. Multi-modal fusion can obtain emotion features from different dimensions, reduce the influence of environmental factors on recognition results, and more comprehensively reflect the real emotions of users.
[0034] In the prior art, emotion recognition is often based on single information source for analysis, and due to the diversity of people's emotional expression, the existing single recognition method has certain deviation.
[0035] In the embodiment of the present application, the generating of the emotion recognition result based on the first set of emotion probability distributions, the first weight, the second set of emotion probability distributions and the second weight comprises: determining a first emotion value of each emotion in the first set of emotion probability distributions based on the first set of emotion probability distributions and the first weight; determining a second emotion value of each emotion in the second set of emotion probability distributions based on the second set of emotion probability distributions and the second weight; performing weighted average processing on the first emotion value and the second emotion value to obtain an intersection emotion value; determining an emotion proportion of each emotion based on the intersection emotion value; obtaining a pre-set sliding time window; obtaining an emotion proportion at each sampling time within the sliding time window; determining emotion change data based on the emotion proportion at each sampling time; and determining an emotion recognition result of the current sliding time window based on the emotion change data.
[0036] In a possible embodiment, after determining the first weight of the first set of emotion probability distributions, for each emotion in the first set of emotion probability distributions, the first emotion value of the emotion is obtained by multiplying the probability or intensity value of the emotion in the first set of emotion probability distributions by the first weight. Similarly, for each emotion in the second set of emotion probability distributions, the second emotion value of the emotion is obtained by multiplying the probability or intensity value of the emotion in the second set of emotion probability distributions by the second weight. Then, weighted average processing is performed on the two sets of emotion values. Specifically, the weighted average value is taken for the emotion types that exist in both the first emotion value and the second emotion value, and the calculation formula is: wherein E f is the facial emotion value, and E v is the voice emotion value. That is, the emotion types that exist in both the two sets of emotion values are found, and the weighted average calculation is performed on the emotion types according to the corresponding weights, to obtain the intersection emotion value. By using the weighted average calculation method, the multi-modal data can be further fused, and the real emotion can be further refined and highlighted, so that the real emotion of the user can be quickly and accurately obtained. Finally, based on the intersection emotion values, the final emotion recognition result is generated through further analysis (for example, determining the dominant emotion, comprehensive evaluation of emotion intensity, etc.), so that the result only contains the emotion information recognized by both modalities.
[0037] By performing the intersection processing of the weighted average of the first emotion value and the second emotion value, the extraction efficiency and recognition accuracy of the real emotion can be effectively improved, the interference of invalid or conflicting emotion information is avoided, the generated emotion recognition result is more reliable and targeted, and the emotion that is commonly embodied by the user under the two modalities can be more accurately reflected.
[0038] The prior art usually only focuses on the emotion state at a certain time when generating the emotion recognition result, ignores the change trend of the emotion in the time dimension, cannot comprehensively analyze the dynamic change process of the user emotion, and is difficult to accurately judge the development trend and stability of the user emotion.
[0039] Therefore, after obtaining the intersection emotion value, the emotion proportion of each emotion at the current moment is further calculated, specifically, the proportion of each emotion value in the total sum of all emotion values. Then, a preset sliding time window is obtained, which can be set to the last 5 minutes or 10 minutes, etc., for analyzing the emotion data in a period of time, and the emotion change trend is updated every 30 seconds in the sliding time window. The slope of the emotion proportion is analyzed by linear regression to determine the emotion stability. Then, in the preset emotion analysis period, the emotion proportion data at each sampling moment is collected. After that, the emotion proportions at different moments are analyzed, and the change data of the emotion in time is calculated, including the increase and decrease amplitude of the emotion proportion, the conversion frequency between different emotions, etc., to reflect the dynamic change of the emotion. For example, the slope K of the emotion proportion at each sampling moment in the sliding time window is analyzed by linear regression, wherein when |k|≤0.05 / minute, it is determined that the emotion state is stable; and when |k|>0.05 / minute, it is determined that the emotion state fluctuates and needs to be focused on. Finally, based on the emotion change data, the overall emotion trend in the sliding time window is determined, and the emotion recognition result of the current sliding time window is determined, for example, whether the user is in a stable emotion state or a large fluctuation emotion state in the time period, and the change of the dominant emotion, etc.
[0040] By introducing the sliding time window and the analysis of the emotion change data, the emotion evolution process of the user in a period of time can be captured, the dynamic change characteristics of the emotion are considered, the emotion recognition result can not only reflect the current emotion state, but also reflect the change trend of the emotion, and the comprehensiveness and accuracy of the user emotion analysis are improved.
[0041] In actual application process, the first tendency emotion and the second tendency emotion are often inconsistent for a long time, which will lead to a monitoring vacuum period of the user and may lead to missed attention to the user. In order to ensure more timely and accurate emotion recognition, the first tendency emotion and the second tendency emotion need to be optimized.
[0042] In the embodiment of the present application, the method further comprises: in the case that the first tendency emotion and the second tendency emotion are inconsistent, obtaining a set of text information interacting with the user; generating a third emotion probability distribution set based on the set of text information; arbitrating the first tendency emotion or the second tendency emotion based on the third emotion probability distribution set to generate an arbitration result; and determining the arbitration result as the first tendency emotion and the second tendency emotion.
[0043] Specifically, a text information set for interaction with the user can be further acquired, for example, the text information set is a chat record between the user, a third emotion probability distribution set is generated based on the real-time chat record of the user, such as the sentiment polarity of the text keyword can be analyzed by an NLP model (such as BERT), when the first tendency emotion and the second tendency emotion are inconsistent, a third emotion set is introduced for arbitration, and a final emotion recognition result is generated through three-level fusion, that is wherein , , , E f is a facial emotion value, E v is a voice emotion value, and E t is a text emotion value.
[0044] The semantic information of the text mode is supplemented to avoid misjudgment caused by single mode deviation; the weight is adjusted in real time according to the confidence of each mode to improve the robustness in complex scenes; the visual (face), auditory (voice) and semantic (text) information are combined to more comprehensively capture the user emotion.
[0045] The existing emotion recognition technology often lacks a subsequent coping mechanism after identifying a negative emotion, cannot take effective intervention measures for the negative emotion, cannot help the user to relieve or solve the problem causing the negative emotion in time, and is difficult to realize effective interaction and emotion regulation with the user.
[0046] In the embodiment of the present application, the method further comprises: after generating the emotion recognition result, judging whether the emotion recognition result is a negative emotion; if so, acquiring an emotion source of the negative emotion; determining a corresponding interaction scheme based on the emotion source; determining a response device and response information based on the interaction scheme, and controlling the response device to perform a corresponding interaction action or outputting the response information.
[0047] In a possible embodiment, after generating the emotion recognition result through the foregoing steps, it is first determined whether the result is a negative emotion, including anger, sadness, anxiety, etc. If it is a negative emotion, the source of the negative emotion is further analyzed and obtained. The source can be determined by asking the user, analyzing historical data, combining environmental information, etc. After determining the source of the emotion, a corresponding interaction scheme is determined according to a preset rule or model, for example, if the source of the emotion is work pressure, the interaction scheme can be to provide relaxation suggestions, play soothing music, etc. If the source of the emotion is a physiological factor (such as a headache), the smart home device is controlled to adjust the indoor light to a soft mode, and white noise is played through the sound box. If it is a psychological factor (such as anxiety), a meditation guide video is pushed, and the smart watch is linked to guide breathing training. Then, according to the interaction scheme, the response device (such as a mobile phone, a sound box, etc.) and the corresponding response information (such as a voice prompt, a text message, an action instruction, etc.) that need to be executed are determined, and finally the response device is controlled to execute the corresponding interaction action or output the response information, thereby realizing intervention on the negative emotion of the user.
[0048] By identifying the negative emotion, the source of the emotion can be actively obtained, and the corresponding interaction scheme is determined according to the source, the response device is controlled to execute the interaction action or output the response information, thereby realizing timely intervention and effective response to the negative emotion of the user, improving the user experience, and helping the user to relieve the negative emotion.
[0049] In the process of obtaining the source of the negative emotion, the prior art can lack a systematic and comprehensive analysis method, which cannot accurately distinguish between negative emotions caused by physiological factors and psychological factors, or cannot take corresponding processing methods for different factor types, resulting in inaccurate judgment of the source of the emotion and affecting the effectiveness of subsequent intervention measures.
[0050] In the embodiment of the present application, the source of the negative emotion is obtained, including: generating inquiry information corresponding to the emotion recognition result; obtaining feedback information of the user for the inquiry information; determining an adverse factor causing the negative emotion based on the feedback information; performing type analysis on the adverse factor to generate a corresponding factor type; if the factor type is a physiological factor, determining a corresponding physiological part based on the physiological part to generate a corresponding source of emotion; if the factor type is a psychological factor, generating a source of emotion corresponding to the adverse factor.
[0051] In a possible embodiment, when a negative emotion is identified, first, inquiry information corresponding to the emotion identification result is generated, for example, inquiry information of "Are you now uncomfortable because of discomfort of a certain part of your body, or because of emotional reasons?" is generated and output to the user. Then, feedback information of the user to the inquiry information is obtained, for example, the user answers "recently, I have a severe headache and my mood is not good". Then, the feedback information is analyzed to determine adverse factors causing the negative emotion, for example, "headache" and "mood problem" in the above example. The adverse factors are analyzed by type to determine whether they belong to physiological factors or psychological factors. If the factor type is a physiological factor, the corresponding physiological part is further determined, for example, the physiological part is determined to be the head according to "headache", and the emotion source is generated based on the physiological part, for example, "discomfort of the head causes negative emotion"; if it is a psychological factor, the corresponding emotion source is directly generated according to the adverse factor, for example, "work pressure causes negative emotion", which provides an accurate basis for subsequent intervention.
[0052] By generating inquiry information to obtain user feedback and analyzing the feedback information by factor type, physiological factors and psychological factors are accurately distinguished, the corresponding emotion source is determined for different types, the judgment of the emotion source is more accurate, and a strong basis is provided for subsequent development of a targeted interaction scheme.
[0053] When inquiring about the source of the negative emotion of the user, the prior art may not consider the case that the user does not feedback in time, lacks an active care mechanism, and thus cannot obtain the emotion source in time, affects the intervention efficiency of the negative emotion of the user, may make the user feel neglected, and reduces the user experience.
[0054] In the embodiment of the application, the method further comprises: after the inquiry information is generated and output, a preset waiting time is determined; it is judged whether the user feeds back the inquiry information within the preset waiting time; if the inquiry information is not fed back, active care information is generated; and a care interaction operation corresponding to the active care information is performed, the care interaction operation including at least one of active concern inquiry, active approach and active interaction action.
[0055] In a possible embodiment, after the query information is generated and output to the user, a preset waiting time (such as 1 minute, 3 minutes, etc.) is determined for waiting for the feedback of the user. During the preset waiting time, it is continuously monitored whether the user has fed back the relevant information. If no feedback of the user is received within the time, active care information such as "I see that you don't seem to be in a good mood. Is there anything I can help you with?" is generated, and corresponding care interactive operations are performed. The care interactive operations include at least one of active care inquiries (such as the active care information described above), active approaches (such as the robot moving close to the user to show a caring posture), and active interactive actions (such as playing a relaxing video, providing a decompression game, etc.), which actively communicate with the user to try to obtain the source of the emotion or alleviate the negative emotion of the user.
[0056] By generating the care information and performing the care interactive operations when the user does not feed back the query information within the preset waiting time, active attention is paid to the user, the participation and satisfaction of the user are improved, and the source of the emotion or the negative emotion of the user is helped to be alleviated in time.
[0057] When the user wears the physiological monitoring device, the prior art can not effectively combine the feedback information and the physiological monitoring information of the user to determine the physiological part, can only rely on a single information source, and thus the determined physiological part is inaccurate, and the judgment and processing of the physiological factor causing the negative emotion are affected.
[0058] In the embodiment of the present application, the user wears at least one physiological monitoring device on the body, and the corresponding physiological part is determined, including: obtaining physiological monitoring information of the at least one physiological monitoring device; determining a specific physiological monitoring device corresponding to the feedback information; judging whether specific physiological monitoring information corresponding to the specific physiological monitoring device is abnormal; if yes, determining the corresponding physiological part based on the specific physiological monitoring information; otherwise, obtaining abnormal monitoring information in the remaining physiological monitoring information, and determining the corresponding physiological part based on the abnormal monitoring information.
[0059] In a possible embodiment, assuming that the user wears at least one physiological monitoring device (such as a heart rate monitor, a sphygmomanometer, a body temperature sensor, and the like), physiological monitoring information (such as heart rate, blood pressure, body temperature, and the like) of the physiological monitoring device is first acquired. Then, according to the feedback information of the user (such as the user mentioning "chest tightness"), a specific physiological monitoring device corresponding to the feedback information (such as a device for monitoring physiological indicators of the chest) is determined. Then, it is determined whether specific physiological monitoring information corresponding to the specific physiological monitoring device is abnormal (such as abnormally high heart rate, blood pressure beyond the normal range, and the like). If there is an abnormality, a corresponding physiological part is determined based on the specific physiological monitoring information (such as determining the physiological part as the chest according to the abnormal physiological indicators of the chest part). If there is no abnormality, further analysis is performed on the remaining physiological monitoring information to find whether there is abnormal monitoring information, and a corresponding physiological part is determined based on the found abnormal monitoring information (such as the user feeding back chest tightness but the physiological indicators of the chest part are normal, and it is found that the head temperature monitoring data is abnormal, and the physiological part is determined as the head), so as to accurately determine the physiological factor source related to the negative emotion.
[0060] By combining the user feedback information and the monitoring information of the physiological monitoring device, the physiological part related to the negative emotion can be more accurately determined, the subjectivity of only relying on the user feedback or the one-sidedness of only relying on the physiological monitoring can be avoided, and the accuracy of physiological factor analysis in the emotion source can be improved.
[0061] After performing the interactive action of the response device or outputting the response information, the prior art can not continuously monitor the subsequent state of the user emotion, cannot determine whether the intervention measure is effective, and cannot take further measures in time if the user is still in a negative emotion, which can cause delay in problem handling.
[0062] In the embodiment of the present application, the method further includes: after controlling the response device to perform the corresponding interactive action or output the response information, the emotion recognition result for the user is acquired again; it is determined whether the user is still in a negative emotion based on the emotion recognition result acquired again; if yes, corresponding alarm information is generated and fed back.
[0063] In a possible embodiment, after the control response device performs the corresponding interactive action or outputs the response information, the emotional recognition result for the user is acquired again according to a preset acquisition frequency and recognition method. Then, the emotional recognition result acquired again is analyzed to determine whether the user is still in a negative emotion. If it is determined that the user is still in a negative emotion (for example, the continuous multiple recognition results are all negative emotions, or the intensity of the negative emotion does not decrease significantly), corresponding alarm information (for example, sending an alarm notification to a designated contact, triggering an alarm in the system, etc.) is generated, so that relevant personnel or the system take further intervention measures, such as manual intervention communication, adjustment of the intervention scheme, etc., to ensure that the negative emotion of the user is effectively handled.
[0064] By acquiring the emotional recognition result again after the intervention, determining whether the user is still in a negative emotion, and generating alarm information if the user is still in a negative emotion, continuous attention and effective intervention on the emotional state of the user are realized, and when the intervention measure is invalid, a higher-level response measure can be taken in time to ensure the emotional health of the user.
[0065] When acquiring the facial and voice information of the user, the existing emotional recognition technology may not consider whether there are other people in the environment where the user is. In a non-private environment, the facial expression and voice of the user may be affected by others, resulting in inaccurate emotional recognition results and failing to truly reflect the real emotional state of the user when alone.
[0066] In the embodiment of the present application, the method further comprises: acquiring environmental monitoring information of the environment where the user is before acquiring the facial emotional recognition information and the voice recognition information; determining whether there are other people in the space where the user is based on the environmental monitoring information; if there are no other people, determining that the user is in a private environment; and acquiring facial emotional recognition information and corresponding voice recognition information of the user based on a preset acquisition frequency.
[0067] In a possible embodiment, first, environmental monitoring information of the environment where the user is is acquired (for example, the number of people in the environment is counted by a camera combined with face recognition technology, human body movement is detected using a millimeter wave radar to distinguish the user from other people, a microphone array locates a sound source, and whether the user is the only speaker is identified based on voice features). Then, it is determined whether there are other people in the space where the user is based on the environmental monitoring information (for example, it is detected that there is only the user in the picture, or no other human body signal is detected by the sensor). If it is determined that there are no other people, it is determined that the user is in a private environment. At this time, since the user may express his / her emotion more truly when alone, facial emotional recognition information and corresponding voice recognition information of the user are acquired according to a preset acquisition frequency, which are used for subsequent emotional recognition processing, to ensure that the acquired emotional information is not disturbed by others and can better reflect the real emotional state of the user.
[0068] By judging whether other people exist in the space where the user is before acquiring facial and voice information, if in a solitary environment, information collection is carried out, so as to avoid the influence of other people on the user's emotional expression, ensure that the collected emotional information is more real and reliable, and improve the accuracy of emotion recognition.
[0069] In the embodiments of the present application, the facial, voice and physiological data involved comply with the following security policies:
[0070] Transmission encryption: the AES-256 encryption algorithm is used in the transmission process to ensure link security;
[0071] Storage anonymization: when physiological data is stored, the user's name, ID number and other identity identifiers are removed, and only device ID associated data is retained;
[0072] User control: the user can initiate a data deletion request at any time through the terminal device, and the system completes the historical data cleaning within 72 hours.
[0073] The embodiments of the present application provide a kind of multi-modal fusion dynamic weighted recognition device based on facial emotion recognition and speech recognition, which will be described below in conjunction with the accompanying drawings.
[0074] Please refer to Figure 2 , based on the same inventive concept, the present application also provides a kind of multi-modal fusion dynamic weighted recognition device based on facial emotion recognition and speech recognition, the device includes: information acquisition unit, for obtaining the facial emotion recognition information and corresponding speech recognition information of user based on preset acquisition frequency;Preliminary identification unit, for generating corresponding first emotion probability distribution set based on the facial emotion recognition information, and generating corresponding second emotion probability distribution set based on the speech recognition information;Emotion generation unit, for generating first tendency emotion based on the first emotion probability distribution set, and generating second tendency emotion based on the second emotion probability distribution set;Judgment unit, for judging whether the first tendency emotion and the second tendency emotion are consistent;Result output unit, for determining the first weight of the first emotion probability distribution set and the second weight of the second emotion probability distribution set in the case where the first tendency emotion and the second tendency emotion are consistent, and generating emotion recognition result based on the first emotion probability distribution set, the first weight, the second emotion probability distribution set and the second weight.
[0075] Further, the embodiments of the present application also provide a computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement the method of the embodiments of the present application.
[0076] The optional implementation of the embodiments of the present application is described in detail above in combination with the drawings, but the embodiments of the present application are not limited to the specific details in the above implementation. Within the technical concept of the embodiments of the present application, the technical solutions of the embodiments of the present application can be variously and simply modified, and all the simple modifications belong to the protection scope of the embodiments of the present application.
[0077] In addition, it should be noted that each specific technical feature described in the above specific implementation can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, the various possible combinations are not described again in the embodiments of the present application.
[0078] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiment methods can be completed by programs instructing related hardware. The programs are stored in a storage medium and include a plurality of instructions for enabling a single-chip microcomputer, a chip or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk and various program code storage media.
[0079] In addition, the various different implementations of the embodiments of the present application can also be combined in any manner, as long as they do not contradict the idea of the embodiments of the present application, and they should also be considered as disclosed by the embodiments of the present application.
Claims
1. A multi-modal fusion dynamic weighting recognition method based on facial emotion recognition and speech recognition, characterized in that, The method comprises: obtaining facial emotion recognition information and corresponding speech recognition information of a user based on a preset acquisition frequency; generating a corresponding first emotion probability distribution set based on the facial emotion recognition information and a corresponding second emotion probability distribution set based on the speech recognition information; generating a first inclined emotion based on the first emotion probability distribution set and a second inclined emotion based on the second emotion probability distribution set; determining whether the first inclined emotion and the second inclined emotion are consistent; if so, determining a first weight of the first emotion probability distribution set and a second weight of the second emotion probability distribution set, and generating an emotion recognition result based on the first emotion probability distribution set, the first weight, the second emotion probability distribution set and the second weight; the emotion recognition result based on the first emotion probability distribution set, the first weight, the second emotion probability distribution set and the second weight comprises: determining a first emotion value of each emotion in the first emotion probability distribution set based on the first emotion probability distribution set and the first weight, and determining a second emotion value of each emotion in the second emotion probability distribution set based on the second emotion probability distribution set and the second weight; performing weighted average processing on the first emotion value and the second emotion value to obtain an intersection emotion value; determining an emotion proportion of each emotion based on the intersection emotion value; obtaining a preset sliding time window, and obtaining an emotion proportion at each sampling time within the sliding time window; determining emotion change data based on the emotion proportion at each sampling time; determining an emotion recognition result of the current sliding time window based on the emotion change data; in the case that the first inclined emotion and the second inclined emotion are inconsistent, obtaining a set of text information interacting with the user; generating a third emotion probability distribution set based on the set of text information; arbitrating the first inclined emotion or the second inclined emotion based on the third emotion probability distribution set to generate an arbitration result; generating an emotion recognition result based on the arbitration result.
2. The method of claim 1, wherein, The method further comprises: after generating the emotion recognition result, determining whether the emotion recognition result is a negative emotion; if so, obtaining an emotion source of the negative emotion; determining a corresponding interaction scheme based on the emotion source; determining a response device and response information based on the interaction scheme, and controlling the response device to perform a corresponding interaction action or output the response information.
3. The method of claim 2, wherein, The method further comprises: generating inquiry information corresponding to the emotion recognition result; obtaining feedback information of the user for the inquiry information; determining an adverse factor leading to the negative emotion based on the feedback information; performing type analysis on the adverse factor to generate a corresponding factor type; if the factor type is a physiological factor, determining a corresponding physiological part and generating a corresponding emotion source based on the physiological part; if the factor type is a psychological factor, generating an emotion source corresponding to the adverse factor.
4. The method of claim 3, wherein, The method further comprises: After the inquiry information is generated and output, a preset waiting time is determined; It is judged whether the user feeds back the inquiry information within the preset waiting time; If the inquiry information is not fed back, active care information is generated; A care interaction operation corresponding to the active care information is executed, and the care interaction operation includes at least one of active concern inquiry, active approach, and active interaction action.
5. The method of claim 3, wherein, The user wears at least one physiological monitoring device on the body, and the corresponding physiological part is determined, including: Obtaining physiological monitoring information of the at least one physiological monitoring device; Determining a specific physiological monitoring device corresponding to the feedback information; Judging whether the specific physiological monitoring information corresponding to the specific physiological monitoring device is abnormal; If yes, determining the corresponding physiological part based on the specific physiological monitoring information; Otherwise, obtaining abnormal monitoring information in the rest of the physiological monitoring information, and determining the corresponding physiological part based on the abnormal monitoring information.
6. The method of claim 2, wherein, The method further includes: After controlling the response device to execute the corresponding interaction action or output the response information, the emotional recognition result for the user is obtained again; Based on the emotional recognition result obtained again, it is judged whether the user is still in a negative emotion; If yes, corresponding alarm information is generated and fed back.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Before obtaining the facial emotional recognition information and the voice recognition information, obtaining environmental monitoring information of an environment in which the user is located; Based on the environmental monitoring information, it is judged whether there are other people in the space where the user is located; If there are no other people, it is determined that the user is in a solitary environment; Based on a preset collection frequency, facial emotional recognition information and corresponding voice recognition information of the user are obtained.
8. A multi-modal fusion dynamic weighted recognition device based on facial emotion recognition and speech recognition, characterized in that, The device includes: An information acquisition unit configured to obtain facial emotional recognition information and corresponding voice recognition information of a user based on a preset collection frequency; A preliminary recognition unit configured to generate a first emotional probability distribution set based on the facial emotional recognition information and generate a second emotional probability distribution set based on the voice recognition information; An emotion generation unit configured to generate a first inclined emotion based on the first emotional probability distribution set and generate a second inclined emotion based on the second emotional probability distribution set; A judgment unit configured to judge whether the first inclined emotion and the second inclined emotion are consistent; A result output unit configured to, in the case where the first inclined emotion and the second inclined emotion are consistent, determine a first weight of the first emotional probability distribution set and a second weight of the second emotional probability distribution set, and generate an emotional recognition result based on the first emotional probability distribution set, the first weight, the second emotional probability distribution set, and the second weight; The generating of the emotional recognition result based on the first emotional probability distribution set, the first weight, the second emotional probability distribution set, and the second weight includes: determining a first emotion value of each emotion in the first emotion probability distribution set based on the first emotion probability distribution set and the first weight, and determining a second emotion value of each emotion in the second emotion probability distribution set based on the second emotion probability distribution set and the second weight; performing weighted average processing on the first emotion value and the second emotion value to obtain an intersection emotion value; determining an emotion proportion of each emotion based on the intersection emotion value; obtaining a preset sliding time window, and obtaining an emotion proportion of each sampling time within the sliding time window; determining emotion change data based on the emotion proportion of each sampling time; determining an emotion recognition result of the current sliding time window based on the emotion change data; in a case where the first tendency emotion and the second tendency emotion are inconsistent, obtaining a text information set interacted with the user; generating a third emotion probability distribution set based on the text information set; arbitrating the first tendency emotion or the second tendency emotion based on the third emotion probability distribution set to generate an arbitration result; generating an emotion recognition result based on the arbitration result.
Citation Information
Patent Citations
Negative emotion-based equipment function detection method and device, equipment and storage medium
CN117558298A
Man-machine interaction method and system based on multiple modes
CN119806335A