Multi-modal data fusion emotion recognition method and system

Through non-contact video data collection and artificial intelligence models, a multimodal data fusion emotion recognition method is constructed, which solves the problems of insufficient single-modal data and the inconvenience of contact equipment, and realizes efficient and convenient emotion recognition and accurate judgment.

CN120824031AActive Publication Date: 2025-10-21ANHUI HUATU INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510875345.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-21
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing emotion recognition technology relies on single-modal data, resulting in insufficient feature dimensions and difficulty in capturing the dynamic complexity of emotions. In addition, contact-based physiological indicator collection equipment is inconvenient to use, which limits the convenience and multi-scenario application of emotion recognition.

Method used

By collecting video data contactlessly, using artificial intelligence models to build a mapping relationship between feature data and physiological indicators, extracting video feature data and predicting physiological indicators, and combining indicator calibration functions to eliminate individual differences and environmental noise influences, emotional state judgment can be achieved.

Benefits of technology

It improves the convenience and efficiency of emotion recognition, reduces equipment costs, supports rapid emotion recognition in multiple scenarios, and improves the accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120824031A_ABST
    Figure CN120824031A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion recognition method and system based on multi-modal data fusion, relates to the technical field of emotion monitoring, and solves the technical problems that the data acquisition and processing process is tedious and the timeliness and flexibility of emotion recognition are insufficient due to dependence on data acquisition equipment during emotion monitoring in the prior art. According to the method, the artificial intelligence model is trained by using a large amount of test data of a tester, and the mapping relation between the feature data and the physiological indexes is constructed through the artificial intelligence model; during emotion recognition, feature data are extracted from the video data, the feature data are input into the artificial intelligence model to predict a physiological index, and then a core index is calculated according to the physiological index to judge the emotion state of a target user; according to the method and the system, the physiological indexes can be predicted through the video data, a large number of contact acquisition devices do not need to be worn, the convenience and the efficiency of emotion recognition are improved, and meanwhile, the possibility of quickly performing emotion recognition in multiple scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of emotion monitoring, and specifically relates to an emotion recognition method and system based on multimodal data fusion. Background Art

[0002] As a comprehensive reflection of an individual's psychological state, accurate emotion recognition relies on a cross-dimensional analysis of physiological signals, behavioral characteristics, and environmental data. By integrating visual features such as facial expressions and body posture from video images with physiological indicators such as heart rate and skin conductance, a more comprehensive emotion representation model can be constructed, providing objective evidence for scenarios such as psychological counseling, medical diagnosis, and intelligent cockpits.

[0003] Traditional emotion recognition methods rely on either single-modal data, such as facial expressions or a single physiological indicator. This results in insufficient feature dimensionality and makes it difficult to capture the dynamic complexity of emotions. Furthermore, existing models often employ independent feature extraction and fusion strategies when processing spatiotemporal feature correlations. This fails to fully exploit spatiotemporal dependencies in video data (such as the temporal patterns of muscle vibrations in consecutive video frames), limiting emotion recognition accuracy. Efficiently integrating multimodal spatiotemporal features while eliminating the influence of individual differences and environmental noise has become a key challenge facing current emotion recognition technology.

[0004] Existing technologies often rely on contact-based data collection devices to capture physiological indicators. This not only requires users to wear complex equipment, resulting in inconvenience and a poor user experience, but also poses high costs, making it difficult to adopt in everyday situations. Furthermore, existing solutions often struggle to achieve real-time and efficient emotion recognition when rapidly switching between scenarios due to their high device dependency and cumbersome data collection processes. This limits their application in scenarios requiring rapid response, such as mental health screening and intelligent interaction.

[0005] The present invention provides a multimodal data fusion emotion recognition method and system for solving the above technical problems. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes an emotion recognition method and system based on multimodal data fusion, which uses a large amount of test data from testers to train an artificial intelligence model, and constructs a mapping relationship between feature data and physiological indicators through the artificial intelligence model; when performing emotion recognition, by collecting video data of the target user, feature data is extracted from the video data, and the feature data is input into the artificial intelligence model to predict physiological indicators, and then core indicators are calculated based on the physiological indicators to judge the emotional state of the target user; the present invention can be used to predict physiological indicators through video data, and the target user does not need to wear a large amount of contact collection equipment, thereby improving the convenience and efficiency of emotion recognition, and at the same time increasing the possibility of rapid emotion recognition in multiple scenarios.

[0007] To achieve the above objectives, a first aspect of the present invention provides an emotion recognition method using multimodal data fusion, comprising: S100: Collecting video data of a target user and extracting feature data from the video data; wherein the feature data includes spatial features and temporal features; S200: Inputting the characteristic data into a plurality of physiological indicators predicted by an indicator mapping model; wherein the indicator mapping model is constructed based on an artificial intelligence model; S300: obtaining a plurality of core indicators by weighted summation of a plurality of physiological indicators, and judging the emotional state of the target user based on the plurality of core indicators.

[0008] Preferably, the core indicators include impulsiveness index, stress tolerance index, anxiety index, doubt index, harmony index, self-confidence index, energy index, self-discipline index, inhibition index and neuroticism index.

[0009] Preferably, obtaining an indicator mapping model based on artificial intelligence model training includes the following steps: S211: Configure the data collection module in the standard collection environment and have the tester sit in the standard collection environment with a standard posture; S212: Setting standardized tasks to induce different emotional states of the test subjects; standardized tasks include the Trier social stress test and the Stroop color word test; S213: While inducing different emotional states of the test subject, collecting test data of the test subject through the data collection module; wherein the test data includes video data and physiological indicators; S214: Preprocess the test data of the tester to obtain standard training data; train an artificial intelligence model based on the standard training data, and mark the trained artificial intelligence model as an indicator mapping model; wherein the artificial intelligence model is CNN-LSTM.

[0010] Preferably, when the indicator mapping model is trained based on the test data of the target user, the predicted physiological indicators are not calibrated.

[0011] Preferably, when the indicator mapping model is trained based on test data of several different testers, the predicted several physiological indicators are calibrated, including: S221: Retrieving an indicator calibration function of the target user; wherein the indicator calibration function is used to calibrate the deviation between the predicted value and the actual value of the physiological indicator; S222: Calibrate the multiple physiological indicators output by the indicator mapping model through the indicator calibration function to obtain the calibrated multiple physiological indicators.

[0012] Preferably, the indicator calibration function is constructed by testing, including: S231: When collecting the video data of the target user for the first time, simultaneously collecting several physiological indicators of the target user through the contact collection device; S232: Predicting several physiological indicators of the target user based on the video data and the indicator mapping model; and obtaining an indicator calibration function by fitting the predicted results and the measured results of the physiological indicators.

[0013] Preferably, several core indicators are obtained by weighted summation of several physiological indicators, including: S231: Take physiological indicators that affect core indicators as key indicators and associate several key indicators with core indicators; S232: Determine the weight coefficients of the key indicators; perform weighted summation based on the key indicators and the corresponding weight coefficients to obtain the core indicators; wherein the method for determining the weight coefficients includes hierarchical analysis method, principal component analysis method or random forest method.

[0014] Preferably, weighted summation is performed based on key indicators and corresponding weight coefficients, including: Mark the key indicator as , the corresponding weight is marked as ; Calculate core indicators through formula ;in, is the number of key indicators.

[0015] A second aspect of the present invention provides an emotion recognition system for multimodal data fusion, comprising an emotion recognition module and a data acquisition module connected thereto; Data acquisition module: used to collect test data from the tester through data acquisition equipment and integrate the test data into standard training data; and to collect video data of the target user through non-contact acquisition equipment; wherein the test data includes video data and physiological indicators; Emotion recognition module: used to train the artificial intelligence model using standard training data to obtain an emotion mapping model; the artificial intelligence model is CNN-LSTM; and It is used to identify the feature data of video data through the emotion mapping model and predict the physiological indicators of the target user; obtain several core indicators based on the weighted sum of several physiological indicators, and judge the emotional state of the target user based on the several core indicators.

[0016] Preferably, the non-contact collection device includes: Multi-channel cameras: used to collect video data; Light source: used to adjust the ambient lighting during video data acquisition; Depth camera: used to collect three-dimensional head motion data.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention uses a large amount of test data from testers to train an artificial intelligence model, and constructs a mapping relationship between feature data and physiological indicators through the artificial intelligence model; when performing emotion recognition, by collecting video data of the target user, feature data is extracted from the video data, and the feature data is input into the artificial intelligence model to predict physiological indicators, and then core indicators are calculated based on the physiological indicators to judge the emotional state of the target user; the present invention can realize the prediction of physiological indicators through video data, and the target user does not need to wear a large amount of contact collection equipment, thereby improving the convenience and efficiency of emotion recognition, and at the same time increasing the possibility of rapid emotion recognition in multiple scenarios.

[0018] 2. In the present invention, when the indicator mapping model is trained based on the test data of multiple testers, after the physiological indicators of the target user are predicted by using the indicator mapping model, the indicator calibration function of the target user is called to calibrate the physiological indicators; the indicator calibration function can be constructed by testing the target user, and after the construction is completed, it can be used for the subsequent physiological indicator calibration of the target user; the present invention calibrates the physiological indicators through the indicator calibration function to eliminate the influence of individual differences, environmental changes, etc. on the physiological indicators, improve the prediction accuracy of the physiological indicators, and thus improve the accuracy of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1Schematic diagram of the method steps of a multimodal data fusion emotion recognition method in the present invention; Figure 2 Schematic diagram of the construction process of the indicator calibration function in the present invention; Figure 3 This is a schematic diagram of the system principle of a multimodal data fusion emotion recognition system in the present invention. DETAILED DESCRIPTION

[0021] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0022] Traditional emotion recognition methods often rely on single-modal data, such as facial expressions or a single physiological indicator. This results in insufficient feature dimensionality, making it difficult to capture the dynamic complexity of emotions, and thus limiting emotion recognition accuracy. Furthermore, some models use multiple physiological indicators to identify emotions, but accurate measurement of these indicators requires multiple contact-based measurement devices, which complicates the acquisition and processing of these indicators, impacting the efficiency of emotion recognition.

[0023] The present invention provides an emotion recognition method and system based on multimodal data fusion, which uses an artificial intelligence model to establish a mapping relationship between feature data collected by a non-contact measurement device and physiological indicators collected by a contact measurement device. When emotion recognition is required, feature data is collected by a non-contact collection device, and the trained artificial intelligence model can be used to predict the physiological indicators corresponding to the feature data. Rapid emotion recognition can be completed without using multiple types of contact collection devices, thereby improving the efficiency of emotion recognition and reducing the cost of emotion recognition.

[0024] See also Figure 1 The first embodiment of the present invention provides an emotion recognition method based on multimodal data fusion, comprising: S100: Collecting video data of the target user and extracting feature data from the video data; S200: Inputting the characteristic data into a plurality of physiological indicators predicted by an indicator mapping model; wherein the indicator mapping model is constructed based on an artificial intelligence model; S300: obtaining a plurality of core indicators by weighted summation of a plurality of physiological indicators, and judging the emotional state of the target user based on the plurality of core indicators.

[0025] S100: Collect video data of a target user and extract feature data from the video data.

[0026] The target user is the person being tested for emotion recognition and monitoring. Video data is collected in a set environment with the target user in a set posture. The duration of a single video capture ranges from 20 to 120 seconds. To balance data processing efficiency and emotion recognition accuracy, a video data duration of 60 seconds is recommended.

[0027] The emotion recognition method of multimodal data fusion of the present invention is mainly based on non-contact measurement technology. It obtains feature data by extracting them from the video data of the target user, matches the corresponding physiological indicators according to the feature data, and then calculates the core indicators for evaluating the target user's emotions based on the physiological indicators.

[0028] Since the present invention mainly uses non-contact measurement technology to assess emotional state, the equipment used is lighter and the detection process is more convenient. It is mainly suitable for regular monitoring that requires long-term monitoring, such as patients with depression and patients with mental illness.

[0029] Matching the corresponding physiological indicators according to the characteristic data is mainly achieved by using the mapping relationship between the characteristic data and physiological indicators that has been constructed in advance based on the artificial intelligence model.

[0030] In S100, feature data includes spatial features and temporal features. Spatial features reflect static or local physical properties, capturing the spatial distribution, shape, texture, and motion patterns of pixels or regions in an image, and are used to correlate immediate physiological states or anatomical features. Temporal features reflect dynamic properties that change over time, capturing the temporal evolution of spatial features. They are used to analyze the periodicity, trend, or mutation of physiological signals and reveal the persistence or volatility of emotions. The following describes a combination of spatial and temporal features.

[0031] 1. Corrugator muscle related features: 1) Corrugator supercilii vibration frequency: This parameter is associated with the power of the frontal beta wave and reflects the stress index. This parameter is the dominant frequency of the time series signal of the pixel displacement in the glabella area calculated by STFT. 2) Corrugator supercilii vibration amplitude: This parameter is associated with the electromyographic amplitude and reflects the anxiety index. This parameter is the mean Euclidean distance of pixel displacements in the glabellar region calculated using the optical flow method. 3) Zygomatic muscle vibration frequency: This parameter is correlated with the high-frequency power of HRV and reflects the pleasure index. This parameter is the main vibration frequency of the pixel at the midpoint of the line connecting the mouth corner and the zygomatic bone calculated by short-time Fourier transform; 4) Orbicularis oculi muscle contraction duration: This parameter is correlated with skin conductance and reflects the anxiety index. This parameter is the duration of continuous contraction of periocular pixels calculated by dynamic time warping (DTW). 5) Neck muscle vibration entropy: This parameter is associated with sympathetic nerve activity; this parameter is the information entropy of the vibration frequency of pixels in the neck area calculated by entropy value.

[0032] 2. Muscle synergy characteristics: 1) Symmetry of left and right facial vibration: This parameter is related to the symmetry of the left and right EEG. It is the ratio of the absolute difference between the left and right orbicularis oculi muscle vibration amplitudes and the mean value, calculated by symmetric difference. 2) Facial muscle phase consistency: This parameter is related to emotional stability; it is the phase correlation coefficient of the vibration signals of the corrugator supercilii and zygomaticus muscles calculated by the cross-correlation function.

[0033] 3. Blood volume pulse wave (BVP) characteristics

[0034] 1) Facial RGB fluctuation standard deviation: This parameter is related to heart rate and reflects the stress index. This parameter is the standard deviation of the RGB values ​​of the forehead area within 10 seconds, mainly reflecting the changes in skin color caused by blood flow pulsation; 2) Skin redness: This parameter is associated with blood oxygen saturation and reflects the energy index. This parameter is obtained through color space conversion, specifically (RG) / (R+G+B), and reflects blood oxygenation. 3) BVP period: This parameter is associated with heart rate variability; it is the blood volume pulse wave period separated from the RGB signal by independent component analysis.

[0035] 4. Facial temperature characteristics

[0036] 1) Facial temperature gradient: This parameter is associated with cross-neuronal activity and is the difference in pixel brightness between the forehead and cheek regions identified by the camera; 2) Temperature change rate: This parameter is related to vitality and affects the stress index. This parameter is the rate of change of facial brightness per unit time, calculated by inter-frame difference.

[0037] 5. Head posture characteristics: 1) Head translation amplitude: This parameter is related to fatigue and concentration. It is the root mean square displacement of the brow center point on the x / y / z axis obtained through MediaPipe posture estimation. 2) Head rotation angle: This parameter is associated with attention and is the absolute value of the pitch / yaw / roll angle calculated using Euler angles. 3) Head shaking frequency: This parameter is associated with alertness and affects the inhibition index; this parameter is the main frequency of head translation movement obtained through FFT spectrum analysis.

[0038] 6. Facial action features: 1) AU1 (brow lift) intensity: This parameter is related to emotions such as surprise and worry. It is the ratio of the vertical displacement of the medial brow peak obtained by the FACS-AU detection algorithm to the baseline value; 2) AU4 (brow down) duration: This parameter is related to emotions such as stress and anxiety. It is calculated by multiplying the frame rate by the number of frames in which the eyebrow distance is reduced, as obtained through dynamic feature tracking technology. 3) AU12 (mouth corner lift) amplitude: This parameter is related to emotions such as happiness and confidence. This parameter is the vertical displacement of the mouth corner point obtained through facial key point detection technology; 4) AU25 (mouth opening) frequency: This parameter is related to excitement and the desire to express oneself; it is the number of mouth opening movements per unit time.

[0039] 7. Eye features: 1) Blink frequency: This parameter is related to fatigue and concentration; it is the number of blinks per minute obtained through the eye state detection algorithm (upper eyelid coverage of >50% of the pupil is considered a blink); 2) Pupil diameter change rate; this parameter is associated with alertness and affects the anxiety index; this parameter is the change rate of the pupil long axis length in adjacent frames obtained by the ellipse fitting algorithm.

[0040] 8. Respiratory movement characteristics: 1) Shoulder vibration frequency: This parameter is related to the respiratory rate. This parameter is the main frequency of the vertical displacement of the clavicle midpoint obtained by the optical flow method and bandpass filtering method. 2) Chest rise and fall amplitude: This parameter is related to the depth of breathing and affects the energy index. This parameter is the peak vertical displacement of pixels in the chest area obtained by the motion tracking algorithm.

[0041] 9. Characteristics of Breathing-Emotion Coupling: 1) Abnormal breathing rate: This parameter is related to emotions such as anxiety and panic; it is the proportion of abnormal breathing cycles (such as breath holding and rapid breathing); 2) Emotion-respiration phase difference: This parameter is related to emotional stability; it is the time difference between the peak of emotional fluctuation and the peak of respiration obtained by cross-correlation analysis.

[0042] In S200, obtaining an indicator mapping model based on artificial intelligence model training includes the following steps: S211: Configure the data collection module in the standard collection environment and have the tester sit in the standard collection environment with a standard posture; S212: Setting standardized tasks to induce different emotional states of the test subjects; standardized tasks include the Trier social stress test and the Stroop color word test; S213: While inducing different emotional states of the test subject, collecting test data of the test subject through the data collection module; wherein the test data includes video data and physiological indicators; S214: Preprocess the test data of the tester to obtain standard training data; train an artificial intelligence model based on the standard training data, and mark the trained artificial intelligence model as an indicator mapping model; wherein the artificial intelligence model includes CNN-LSTM or CNN-GRU.

[0043] In S211, the data acquisition module mainly includes: 1. Non-contact acquisition equipment: multiple cameras (high-definition, high-frame rate cameras), which can be equipped with infrared fill light modules to adapt to low-light environments; at the same time, it can also be equipped with a depth camera for three-dimensional head motion tracking; 2. Contact acquisition equipment: including EEG caps, heart rate belts, electromyography patches, etc., to collect EEG, HRV, EMG, skin conductance and other signals.

[0044] Standard collection environment: 1. Lighting: Brightness 200-1000 lux, color temperature 3800K, change rate ≤ 1 lux / s, and avoid direct sunlight on the face; 2. Background: Should be a solid color with no clutter to reduce visual interference, such as a gray curtain; Standard Posture: The test subject must sit still with their face facing the multiple cameras, their head and neck centered in the frame, and their limbs in a quasi-static state (limb movement ≤ 2.5 cm).

[0045] In S212, standardized tasks refer to carefully designed task paradigms with unified processes and quantitative standards that induce specific emotional states (such as stress, anxiety, and concentration) in test subjects to facilitate scientific research or technical verification. The principles and specific testing procedures of testing methods such as the Trier Social Stress Test and the Stroop Color Word Test are already disclosed in existing protocols and will not be described in detail here.

[0046] In step S214, when acquiring standard training data, the contact and non-contact acquisition devices simultaneously collect data from the test subject. When training the artificial intelligence model using the standard training data, the video data collected by the non-contact acquisition device serves as the model input, and the physiological indicators collected by the contact acquisition device serve as the model output.

[0047] For example, the training process of an AI model is described using the CNN-LSTM model as an example: 1) Construct the network structure of the CNN-LSTM model, including the CNN layer, LSTM layer and regression layer.

[0048] The CNN layer can use the spatial features of the input video data, the LSTM layer is used to extract the temporal features of the input video data, and the regression layer is used for regression analysis to obtain physiological indicators.

[0049] 2) Time-synchronize the video frames of the standard training data with the physiological indicators, with a timestamp error of less than 10ms. Segment the video data into 10-second segments (300 frames total), and extract the mean, peak, and frequency domain features of the corresponding physiological indicators during the synchronization period. The frame rate of the video data acquisition should be ≥25 FPS, with a typical value of 30 FPS.

[0050] 3) Pre-train the CNN to independently train it to extract the spatial features of the subject's limb vibrations. After pre-training, the weights of the CNN convolutional layer are retained and the classification head is removed.

[0051] 4) When training the CNN-LSTM fusion model, CNN first identifies the spatial features of the video clips, then extracts the temporal features through LSTM, and predicts physiological indicators through the regression layer based on the spatial and temporal features.

[0052] In S200 , when the indicator mapping model is trained based on test data of several different testers, after the characteristic data is input into the indicator mapping model to obtain corresponding physiological indicators, the physiological coordinates need to be calibrated and corrected.

[0053] When training the indicator mapping model, the feature data in the standard training data comes from the tester, and the physiological indicators will also have certain deviations depending on the tester. When the fusion model is trained using the tester's feature data and physiological indicators during the test, it can be applied to most test situations. However, individual differences in the target users are likely to cause the physiological indicators predicted by the indicator mapping model to have certain deviations from the target user's actual physiological indicators. Moreover, environmental interference will also cause deviations in the prediction results. Since the present invention requires the target user to sit still in a standard posture under certain environmental conditions to collect video data, if the environmental conditions are inconsistent with the acquisition conditions of the standard training data, or the target user is not in a standard posture, etc., it will also affect the accuracy of the predicted physiological indicators.

[0054] It should be noted that the core process of inputting spatial features extracted by CNN into LSTM for temporal analysis is as follows: CNN is first used to extract spatial features of the facial region from video frames. The feature maps of consecutive frames are flattened into a temporal sequence, which is then input into LSTM to capture dynamic dependencies between frames. Finally, an attention mechanism is used to aggregate temporal features for physiological indicator prediction. This architecture significantly improves emotion recognition accuracy by jointly modeling spatial and temporal features.

[0055] In order to improve the accuracy of core indicators, the physiological indicators predicted by the indicator mapping model are calibrated.

[0056] In a preferred embodiment, several physiological indicators are calibrated, including: S221: Retrieving an indicator calibration function of the target user; wherein the indicator calibration function is used to calibrate the deviation between the predicted value and the actual value of the physiological indicator; S222: Calibrate the multiple physiological indicators output by the indicator mapping model through the indicator calibration function to obtain the calibrated multiple physiological indicators.

[0057] Emotion recognition requires predicting several physiological indicators of the target user using an indicator mapping model. Some indicators are unaffected by individual differences or environmental variations, or the effects are minimal, so calibration is not necessary and no indicator calibration function needs to be constructed. However, some indicators are affected by individual differences or environmental variations, so an indicator calibration function needs to be constructed to correct them. Therefore, the same target user may correspond to multiple indicator calibration functions for their physiological indicators. Once fitted, these calibration functions are stored in a database for easy access.

[0058] See also Figure 2 , the indicator calibration function is constructed by testing, including: S231: When collecting the video data of the target user for the first time, simultaneously collecting several physiological indicators of the target user through the contact collection device; S232: Predicting several physiological indicators of the target user based on the video data and the indicator mapping model; and obtaining an indicator calibration function by fitting the predicted results and the measured results of the physiological indicators.

[0059] When the contact data collection device first captures the target user's video data, it simultaneously measures the target user's physiological indicators. These measured indicators serve as the true values. Feature data is extracted from the video data, and the indicator mapping model is used to identify the physiological indicators corresponding to the feature data. These indicators serve as the predicted values. The indicator calibration function can be linear or nonlinear, depending on the mapping relationship between the predicted and true values.

[0060] It's worth noting that the indicator calibration function is constructed during the initial acquisition of the target user's video data. This function can be directly called upon during subsequent emotion monitoring of the target user, without requiring repeated construction. Because the emotion recognition subject in this invention is the user who requires long-term monitoring, constructing the indicator calibration function during the initial acquisition simplifies the subsequent calibration process and improves calibration efficiency.

[0061] In other preferred embodiments, if the indicator mapping model is trained based on standard training data obtained during target user testing, the physiological indicator calibration step can be omitted. The physiological indicators predicted by the indicator mapping model based on the target user's video data can be used as standard values, and no calibration is required when the environmental conditions are consistent with the test environment.

[0062] In S300, several core indicators are determined based on several physiological indicators, including: S311: Take physiological indicators that affect core indicators as key indicators and associate several key indicators with core indicators; S312: Determine the weight coefficients of the key indicators; perform weighted summation based on the key indicators and the corresponding weight coefficients to obtain the core indicators; wherein the method for determining the weight coefficients includes hierarchical analysis method, principal component analysis method or random forest method.

[0063] Core indicators include but are not limited to the following: 1. Impulsiveness Index: The impulsivity index represents an individual's level of emotional excitement and activity. This may manifest as hyperactivity, restlessness, irritability, increased movement and speech, and a degree of aggressive intent. These emotions and behaviors can stem from internal feelings or external influences. A scale is defined as: below 20 is considered low, 20-50 is considered normal, and above 50 is considered high.

[0064] 2. Pressure Index

[0065] The stress index represents the intensity of the impact of stressors on an individual's psychosomatic stability. This impact may be related to an imbalance between an individual's internal and external environmental conditions and their coping abilities, potentially leading to a range of psychological and physiological reactions, such as increased heart rate, blood pressure, shortness of breath, and increased hormone secretion. The index is categorized as: below 20 (low), 20-40 (normal), and above 40 (high).

[0066] 3. Anxiety Index

[0067] The anxiety index represents an individual's level of concern regarding major life-threatening issues, future prospects, and unpredictable and difficult-to-handle events. It can encompass a variety of emotions, including anxiety, worry, sadness, tension, uneasiness, and irritability. A scale of 15 or less is considered low, 5-40 is considered normal, and 40 or above is considered high.

[0068] 4. Doubt Index / Irritation Index

[0069] The Irritability Index (IRI) represents the intensity of an individual's emotional response. This index is derived by combining the scores of the three aforementioned indices. High values ​​in the aforementioned indices indicate a high IRI. This index serves as a comprehensive abnormality warning parameter. High values ​​indicate a significant increase in destructive and aggressive intentions, emotional instability, and possible internal stress or intense anxiety, making individuals easily irritated and prone to uncontrolled behavior. The IRI is graded as follows: below 20 is considered low, 20-50 is considered normal, and above 50 is considered high.

[0070] 5. Harmony Index

[0071] The Harmony Index represents the state of balance between the brain and body. When the brain and body are in harmonious interaction, individuals typically exhibit positive traits such as intelligence, quick thinking, strong learning abilities, high energy, and efficiency. However, when the brain and body become unbalanced, disrupted, or even conflicting, various physiological, psychological, and behavioral disorders can develop. The index is categorized as follows: below 40 is considered low, 40-80 is considered normal, and above 80 is considered high.

[0072] 6. Confidence Index

[0073] The self-confidence index represents an individual's confidence and certainty in their own abilities. It is a conscious and psychological state that actively and effectively expresses self-worth, self-respect, and self-understanding. The self-confidence index influences individual psychology and behavior in various areas, including learning, competition, employment, and achievement. A scale is defined as: below 50 is considered low, 50-80 is considered normal, and above 80 is considered high.

[0074] 7. Energy Index

[0075] The energy index represents an individual's psychological energy level. Psychological energy is divided into three dimensions: vitality, drive, and ability. The first dimension is vitality, which refers to an individual's emotional energy. Vibrant people display high energy, positive attitudes, proactive behavior, and contagious emotions. The second dimension is drive, which refers to the energy of will. Driven individuals, guided by consciousness, purpose, and planning, exhibit directionality and persistence in their behavior. The third dimension is ability, which refers to an individual's cognitive energy. Capable individuals are able to adopt effective coping strategies to solve problems. A score below 50 is considered low, between 50 and 80 is considered normal, and above 80 is considered high.

[0076] 8. Self-discipline index

[0077] The Self-Discipline Index (SDI) reflects an individual's ability to adapt their mental and physical state to environmental demands. Through self-regulation, individuals can introspect and control their own emotional state, while also understanding the emotions of others. Based on this, they assess the correctness and objectivity of their own behavior, ensuring that their actions achieve the desired results as closely as possible. A grading system is used: below 50 is considered low, 50-80 is considered normal, and above 80 is considered high.

[0078] 9. Inhibition Index

[0079] The suppression index represents the degree to which an individual's mental and physical state is suppressed and restricted by external or internal forces. An individual's mental and physical energy needs to be maintained within a relatively stable range. Excessive suppression, or a lack of necessary restraint, can lead to a variety of emotional and physical abnormalities. The index is categorized as follows: below 15 is low, 15-25 is normal, and above 25 is high.

[0080] 10. Neuroticism Index

[0081] The Neuroticism Index is used to characterize an individual's emotional stability. It's important to note that this index tends to explain normal behavior rather than mental illness. The scale is: below 10 is low, 10-50 is normal, and above 50 is high.

[0082] The analytic hierarchy process, principal component analysis, and random forest method are commonly used weight determination methods in existing solutions. The process of determining the weight coefficients will not be described in detail here. It should be noted that the data used to determine the weight coefficients are various physiological indicators collected by contact data collection equipment.

[0083] For example, taking the anxiety index as an example, assuming that its related physiological indicators include frown frequency, heart rate, eye gaze stability, breathing rate, and head movement frequency, the standardized index values ​​are: , , , , , the weight coefficients solved by the random forest method are: , , , , . Substitute the above data into the anxiety index An anxiety index can be calculated.

[0084] In 300, the emotional state of the target user is judged based on several core indicators, that is, the emotional risk corresponding to the target user is comprehensively judged based on the core indicators. Please refer to Table 1 and Table 2 below: Table 1 Single evaluation content of core indicators

[0085] Table 2 Comprehensive evaluation content based on core indicators

[0086] See also Figure 3 , a second aspect of the present invention provides an emotion recognition system for multimodal data fusion, including an emotion recognition module and a data acquisition module connected thereto; Data acquisition module: used to collect test data from the tester through data acquisition equipment and integrate the test data into standard training data; and to collect video data of the target user through non-contact acquisition equipment; wherein the test data includes video data and physiological indicators; Emotion recognition module: used to train the artificial intelligence model using standard training data to obtain an emotion mapping model; the artificial intelligence model is CNN-LSTM; and It is used to identify the feature data of video data through the emotion mapping model and predict the physiological indicators of the target user; obtain several core indicators based on the weighted sum of several physiological indicators, and judge the emotional state of the target user based on the several core indicators.

[0087] The data acquisition module mainly collects the required data through data acquisition equipment, and sends the data to the emotion recognition module or stores it after necessary processing. The emotion recognition module is responsible for the construction, training and updating of the indicator mapping model.

[0088] The non-contact acquisition equipment includes: multi-channel cameras: used to collect video data; light source: used to adjust the ambient lighting during the video data acquisition process; depth camera: used to collect three-dimensional head movement data.

[0089] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A multimodal data fusion emotion recognition method, characterized in that: include: Collecting video data of a target user and extracting feature data from the video data; wherein the feature data includes spatial features and temporal features; Inputting the characteristic data into an indicator mapping model to predict several physiological indicators; wherein the indicator mapping model is constructed based on an artificial intelligence model; A plurality of core indicators are obtained by weighted summation of the plurality of physiological indicators, and the emotional state of the target user is judged based on the plurality of core indicators.

2. The emotion recognition method based on multimodal data fusion according to claim 1, characterized in that: The core indicators include impulse index, pressure index, anxiety index, doubt index, harmony index, self-confidence index, energy index, self-discipline index, inhibition index and neuroticism index.

3. The emotion recognition method based on multimodal data fusion according to claim 1, characterized in that: The indicator mapping model obtained based on artificial intelligence model training includes the following steps: Configure the data collection module in a standard collection environment and have the tester sit in the standard collection environment with a standard posture; Setting standardized tasks to induce different emotional states of the test subject using the standardized tasks; wherein the standardized tasks include the Trier social stress test and the Stroop color word test; While inducing different emotional states of the test subject, the test data of the test subject is collected through the data collection module; wherein the test data includes video data and physiological indicators; The test data of the tester is preprocessed to obtain standard training data; an artificial intelligence model is trained based on the standard training data, and the trained artificial intelligence model is marked as an indicator mapping model; wherein the artificial intelligence model is a CNN-LSTM.

4. The emotion recognition method based on multimodal data fusion according to claim 3, characterized in that: When the indicator mapping model is trained based on the test data of the target user, the predicted physiological indicators are not calibrated.

5. The emotion recognition method based on multimodal data fusion according to claim 3, characterized in that: When the indicator mapping model is trained based on test data of several different testers, the predicted physiological indicators are calibrated, including: Retrieving an indicator calibration function for the target user; wherein the indicator calibration function is used to calibrate the deviation between the predicted value and the actual value of the physiological indicator; The multiple physiological indices output by the index mapping model are calibrated using the index calibration function to obtain the calibrated multiple physiological indices.

6. The emotion recognition method based on multimodal data fusion according to claim 5, characterized in that: The indicator calibration function is constructed by testing, including: When collecting the target user's video data for the first time, simultaneously collecting several physiological indicators of the target user through a contact collection device; Based on the video data and the indicator mapping model, several physiological indicators of the target user are predicted; and an indicator calibration function is obtained by fitting the predicted results and the measured results of the physiological indicators.

7. The emotion recognition method based on multimodal data fusion according to claim 1, characterized in that: Several core indicators are obtained by weighted summation of several physiological indicators, including: Taking physiological indicators that affect the core indicators as key indicators, and associating several of the key indicators with the core indicators; Determine the weight coefficient of the key indicator; perform weighted summation based on the key indicator and the corresponding weight coefficient to obtain the core indicator; wherein the method for determining the weight coefficient includes hierarchical analysis method, principal component analysis method or random forest method.

8. The emotion recognition method based on multimodal data fusion according to claim 7, characterized in that: Performing a weighted sum based on the key indicators and the corresponding weight coefficients includes: Mark the key indicator as , the corresponding weight is marked as ; Calculate core indicators through formula ;in, is the number of key indicators.

9. A multimodal data fusion emotion recognition system, used to implement the multimodal data fusion emotion recognition method according to any one of claims 1 to 8, characterized in that: It includes an emotion recognition module and a data acquisition module connected thereto; Data collection module: used to collect the test data of the tester through the data collection equipment and integrate the test data into standard training data; and, collecting video data of the target user through a non-contact collection device; wherein the test data includes video data and physiological indicators; Emotion recognition module: used to train the artificial intelligence model using standard training data to obtain an emotion mapping model; the artificial intelligence model is CNN-LSTM; and It is used to identify the characteristic data of the video data through the emotion mapping model and predict the physiological indicators of the target user; obtain several core indicators based on the weighted sum of several physiological indicators, and judge the emotional state of the target user based on the several core indicators.

10. The multimodal data fusion emotion recognition system according to claim 9, characterized in that: The non-contact collection device includes: Multi-channel cameras: used to collect video data; Light source: used to adjust the ambient lighting during video data acquisition; Depth camera: used to collect three-dimensional head motion data.

Citation Information

Patent Citations

  • Non-contact and contact cooperative real-time emotion intelligent monitoring system

    CN110598607A

  • Non-contact anxiety recognition method and device based on face video

    CN113326781A

  • Student psychological health group screening method and system

    CN114792553A

  • MD patient emotion fluctuation monitoring and affective disorder state evaluation method and system

    CN115517681A

  • Music intervention method and system for driver emotion adjustment

    CN117442843A

Cited By

  • AI mental health monitoring method and device based on traditional Chinese medicine data and electronic equipment

    CN121533737A