A multi-modal data fusion emotion recognition method and system

By collecting video data non-contactly and using artificial intelligence models to predict physiological indicators, the problem of insufficient recognition of single-modal data is solved, and the convenience and efficiency of emotion recognition are improved, making it suitable for multi-scenario applications.

CN120824031BActive Publication Date: 2026-03-31ANHUI HUATU INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing emotion recognition technologies rely on single-modal data, resulting in insufficient feature dimensions and difficulty in capturing the dynamic complexity of emotions. Furthermore, traditional methods require contact-based data collection devices, which are inconvenient to use and costly, making it difficult to achieve real-time and efficient recognition in multiple scenarios.

Method used

By collecting video data in a non-contact manner, using artificial intelligence models to construct a mapping relationship between feature data and physiological indicators, feature data in video data is extracted and physiological indicators are predicted. Combined with indicator calibration functions, the influence of individual differences and environmental noise is eliminated, thereby achieving emotion recognition.

Benefits of technology

It improves the convenience and efficiency of emotion recognition, reduces equipment costs, supports rapid emotion recognition in multiple scenarios, and is suitable for long-term monitoring such as the psychological state assessment of patients with depression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120824031B_ABST
    Figure CN120824031B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal data fusion emotion recognition method and system, it is related to emotional monitoring technical field, it solves the technical problem that current technology relies on acquisition data acquisition equipment when emotional monitoring, leading to data acquisition and processing process is complicated, leading to the timeliness and flexibility of emotion recognition are insufficient;The application trains artificial intelligence model using a large amount of test data of tester, and the mapping relationship between feature data and physiological indicators is constructed by artificial intelligence model;In emotion recognition, feature data is extracted from video data, the feature data is input into artificial intelligence model to predict physiological indicators, and then the core indicators are calculated according to physiological indicators to judge the emotional state of target user;Using the application can realize predicting physiological indicators through video data, without wearing a large number of contact type acquisition equipment, improve the convenience and efficiency of emotion recognition, and at the same time improve the possibility of quickly performing emotion recognition in multiple scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of emotion monitoring, specifically a method and system for emotion recognition based on multimodal data fusion. Background Technology

[0002] Emotions, as a comprehensive reflection of an individual's psychological state, require accurate identification through cross-dimensional analysis of physiological signals, behavioral characteristics, and environmental data. By integrating visual features such as facial expressions and body postures from video images, as well as physiological indicators such as heart rate and skin conductance, a more comprehensive emotion representation model can be constructed, providing objective evidence for scenarios such as psychological counseling, medical diagnosis, and intelligent cockpits.

[0003] In existing technologies, traditional emotion recognition methods either rely on single-modal data, such as facial expressions or single physiological indicators, resulting in insufficient feature dimensions and difficulty in capturing the dynamic complexity of emotions. Furthermore, existing models often employ independent feature extraction and fusion strategies when processing spatiotemporal feature correlations, failing to fully utilize the spatiotemporal dependencies in video data (such as the temporal patterns of muscle vibrations in consecutive video frames), thus limiting the accuracy of emotion recognition. How to efficiently fuse multimodal spatiotemporal features and eliminate the influence of individual differences and environmental noise has become a key problem that current emotion recognition technology urgently needs to solve.

[0004] In existing technologies, emotion recognition often relies on contact-based data collection devices to obtain physiological indicators. This not only requires users to wear complex devices, leading to inconvenience and a poor user experience, but also results in high device costs, making it difficult to popularize in everyday scenarios. Furthermore, existing solutions often struggle to achieve real-time and efficient emotion recognition when rapidly switching between scenarios due to strong device dependence and cumbersome data collection processes, limiting their application in scenarios requiring rapid response, such as mental health screening and intelligent interaction.

[0005] This invention provides a multimodal data fusion-based emotion recognition method and system to solve the aforementioned technical problems. Summary of the Invention

[0006] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention proposes a multimodal data fusion-based emotion recognition method and system. It utilizes a large amount of test data from test subjects to train an artificial intelligence model, and then constructs a mapping relationship between feature data and physiological indicators through the artificial intelligence model. During emotion recognition, video data of the target user is collected, feature data is extracted from the video data, and this feature data is input into the artificial intelligence model to predict physiological indicators. Furthermore, core indicators are calculated based on these physiological indicators to determine the target user's emotional state. This invention enables the prediction of physiological indicators through video data, eliminating the need for the target user to wear numerous contact-based data collection devices, thus improving the convenience and efficiency of emotion recognition, and simultaneously increasing the possibility of rapid emotion recognition in multiple scenarios.

[0007] To achieve the above objectives, a first aspect of the present invention provides a multimodal data fusion-based emotion recognition method, comprising:

[0008] S100: Collects video data from the target user and extracts feature data from the video data; the feature data includes spatial features and temporal features.

[0009] S200: Several physiological indicators predicted by inputting feature data into the indicator mapping model; wherein, the indicator mapping model is constructed based on an artificial intelligence model;

[0010] S300: Several core indicators are obtained by weighted summation of several physiological indicators, and the emotional state of the target user is judged based on these core indicators.

[0011] The preferred core indicators include the impulsiveness index, stress index, anxiety index, doubt index, harmony index, confidence index, energy index, self-discipline index, inhibition index, and neuroticism index.

[0012] Preferably, the process of obtaining the index mapping model based on artificial intelligence model training includes the following steps:

[0013] S211: Configure the data acquisition module in the standard acquisition environment and have the tester sit quietly in the standard acquisition environment in a standard posture.

[0014] S212: Set up standardized tasks to induce different emotional states in test subjects; among them, standardized tasks include the Trier social stress test and the Stroop color word test.

[0015] S213: While inducing different emotional states in the test subjects, the test data of the test subjects is collected through the data acquisition module; the test data includes video data and physiological indicators.

[0016] S214: Preprocess the test data of the testers to obtain standard training data; train the artificial intelligence model based on the standard training data, and label the trained artificial intelligence model as an index mapping model; wherein, the artificial intelligence model is CNN-LSTM.

[0017] Preferably, when the indicator mapping model is trained based on the test data of the target user, the predicted physiological indicators are not calibrated.

[0018] Preferably, when the indicator mapping model is trained based on test data from several different testers, several predicted physiological indicators are calibrated, including:

[0019] S221: Retrieve the target user's indicator calibration function; whereby the indicator calibration function is used to calibrate the deviation between the predicted and actual values ​​of physiological indicators;

[0020] S222: The physiological indicators output by the indicator mapping model are calibrated using the indicator calibration function to obtain the calibrated physiological indicators.

[0021] Preferably, the indicator calibration function is constructed through testing, including:

[0022] S231: When collecting video data of the target user for the first time, several physiological indicators of the target user are also collected simultaneously through a contact acquisition device.

[0023] S232: Predict several physiological indicators of the target user based on video data and an indicator mapping model; obtain the indicator calibration function by fitting the prediction and measurement results of the physiological indicators.

[0024] Preferably, several core indicators are obtained by weighted summation of several physiological indicators, including:

[0025] S231: Use physiological indicators that affect core indicators as key indicators, and associate several key indicators with core indicators.

[0026] S232: Determine the weight coefficients of key indicators; perform a weighted summation based on the key indicators and their corresponding weight coefficients to obtain the core indicators; the methods for determining the weight coefficients include the analytic hierarchy process (AHP), principal component analysis (PCA), or random forest method.

[0027] Preferably, a weighted summation is performed based on key indicators and their corresponding weight coefficients, including:

[0028] Mark the key indicators as The corresponding weight is marked as ;

[0029] Calculate core indicators using formulas ;in, The number of key indicators.

[0030] A second aspect of the present invention provides a multimodal data fusion emotion recognition system, including an emotion recognition module and a data acquisition module connected thereto;

[0031] Data acquisition module: used to collect test data from test subjects through data acquisition devices, integrate the test data into standard training data; and to collect video data from target users through non-contact acquisition devices; among which, test data includes video data and physiological indicators;

[0032] Emotion recognition module: used to train an artificial intelligence model using standard training data to obtain an emotion mapping model; wherein the artificial intelligence model is CNN-LSTM; and,

[0033] This is used to identify feature data from video data through an emotion mapping model, predict the physiological indicators of target users, and obtain several core indicators by weighted summation of several physiological indicators. Based on these core indicators, the emotional state of the target user is determined.

[0034] Preferred non-contact data acquisition devices include:

[0035] Multiple cameras: used to collect video data;

[0036] Light source: Used to adjust the ambient lighting during video data acquisition;

[0037] Depth camera: Used to collect three-dimensional motion data of the head.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] 1. This invention utilizes a large amount of test data from test subjects to train an artificial intelligence model, and constructs a mapping relationship between feature data and physiological indicators through the artificial intelligence model. When performing emotion recognition, video data of the target user is collected, feature data is extracted from the video data, and the feature data is input into the artificial intelligence model to predict physiological indicators. Then, core indicators are calculated based on the physiological indicators to determine the emotional state of the target user. Using this invention, physiological indicators can be predicted through video data. The target user does not need to wear a large number of contact collection devices, which improves the convenience and efficiency of emotion recognition, and also increases the possibility of rapid emotion recognition in multiple scenarios.

[0040] 2. In this invention, when the indicator mapping model is trained based on test data from multiple testers, after predicting the physiological indicators of the target user using the indicator mapping model, the indicator calibration function of the target user is called to calibrate the physiological indicators. The indicator calibration function can be constructed by testing the target user, and after construction, it can be used for subsequent physiological indicator calibration of the target user. This invention calibrates physiological indicators through the indicator calibration function to eliminate the influence of individual differences, environmental changes, etc. on physiological indicators, improve the prediction accuracy of physiological indicators, and thus improve the accuracy of emotion recognition. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a schematic diagram illustrating the method steps of a multimodal data fusion emotion recognition method according to the present invention;

[0043] Figure 2 This is a schematic diagram illustrating the construction process of the index calibration function in this invention;

[0044] Figure 3 This is a schematic diagram illustrating the system principle of a multimodal data fusion emotion recognition system according to the present invention. Detailed Implementation

[0045] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] In existing technologies, traditional emotion recognition methods often rely on single-modal data, such as facial expressions or single physiological indicators, resulting in insufficient feature dimensions and difficulty in capturing the dynamic complexity of emotions, thus limiting the accuracy of emotion recognition. Furthermore, some models use multiple physiological indicators to identify emotions, but accurate measurement of these indicators requires various contact-based measurement devices, making the acquisition and data processing processes complex and impacting the efficiency of emotion recognition.

[0047] This invention provides a multimodal data fusion-based emotion recognition method and system. It utilizes an artificial intelligence model to establish a mapping relationship between feature data collected by non-contact measurement devices and physiological indicators collected by contact measurement devices. When emotion recognition is required, feature data is collected through non-contact acquisition devices, and the trained artificial intelligence model can predict the corresponding physiological indicators. This eliminates the need to use multiple types of contact acquisition devices to achieve rapid emotion recognition, improving the efficiency of emotion recognition while reducing its cost.

[0048] Please see Figure 1 The first aspect of this invention provides a multimodal data fusion-based emotion recognition method, comprising:

[0049] S100: Collects video data from the target user and extracts feature data from the video data;

[0050] S200: Several physiological indicators predicted by inputting feature data into the indicator mapping model; wherein, the indicator mapping model is constructed based on an artificial intelligence model;

[0051] S300: Several core indicators are obtained by weighted summation of several physiological indicators, and the emotional state of the target user is judged based on these core indicators.

[0052] S100: Collects video data from the target user and extracts feature data from the video data.

[0053] The target users are those who need emotion recognition and monitoring. Video data refers to data collected in a set environment when the target user is in a set posture. The duration of a single video data collection ranges from 20 to 120 seconds. To balance data processing efficiency and emotion recognition accuracy, a video data duration of 60 seconds can be used.

[0054] The multimodal data fusion emotion recognition method of the present invention is mainly based on non-contact measurement technology. It extracts feature data from the target user's video data, matches corresponding physiological indicators based on the feature data, and then calculates core indicators for evaluating the target user's emotions based on the physiological indicators.

[0055] Since this invention mainly uses non-contact measurement technology to assess emotional state, the equipment used is lighter and the detection process is more convenient. It is mainly suitable for periodic monitoring that requires long-term monitoring, such as for patients with depression or mental illness.

[0056] Matching corresponding physiological indicators based on feature data is mainly achieved by utilizing the mapping relationship between feature data and physiological indicators that has been pre-built based on an artificial intelligence model.

[0057] In S100, feature data includes spatial features and temporal features. Spatial features reflect static or local physical properties, capturing the spatial distribution, shape, texture, and motion patterns of pixels or regions in an image, used to correlate immediate physiological states or anatomical features. Temporal features reflect dynamic properties that change over time, capturing the temporal evolution patterns of spatial features, used to analyze the periodicity, trend, or abruptness of physiological signals, revealing the persistence or volatility of emotions. The following provides a combined description of spatial and temporal features.

[0058] 1. Characteristics related to the corrugator supercilii muscle:

[0059] 1) Corrugator supercilii vibration frequency: This parameter is related to the power of the prefrontal cortex β wave and reflects the pressure index; this parameter is the main frequency of the timing signal of pixel displacement in the glabella region calculated by STFT;

[0060] 2) Corrugator supercilii vibration amplitude: This parameter is related to the electromyography amplitude and reflects the anxiety index; this parameter is the mean Euclidean distance of the pixel displacement in the glabella region calculated by optical flow method;

[0061] 3) Zygomatic muscle vibration frequency: This parameter is related to HRV high-frequency power and reflects the pleasure index; this parameter is the dominant vibration frequency of the pixel at the midpoint of the line connecting the corner of the mouth and the cheekbone, calculated by short-time Fourier transform.

[0062] 4) Duration of orbicularis oculi muscle contraction: This parameter is related to skin conductance and reflects the anxiety index; this parameter is the duration of continuous contraction of periocular pixels calculated by dynamic time warping (DTW);

[0063] 5) Neck muscle vibration entropy: This parameter is related to sympathetic nerve activity; this parameter is the information entropy of the vibration frequency of pixels in the neck region, calculated by entropy value.

[0064] 2. Muscle synergy characteristics:

[0065] 1) Symmetry of left and right facial vibrations: This parameter is related to the symmetry of left and right EEG; this parameter is the ratio of the absolute difference to the mean of the left and right orbicularis oculi muscle vibration amplitudes calculated by symmetric difference.

[0066] 2) Facial muscle phase consistency: This parameter is related to emotional stability; it is the phase correlation coefficient between the vibration signals of the corrugator supercilii and zygomaticus major muscles, calculated using a cross-correlation function.

[0067] 3. Blood volume pulse wave (BVP) characteristics

[0068] 1) Facial RGB fluctuation standard deviation: This parameter is related to heart rate and reflects the stress index; this parameter is the standard deviation of the statistically obtained RGB values ​​of the forehead area within 10 seconds, mainly reflecting the skin color changes caused by blood flow pulsation;

[0069] 2) Skin redness: This parameter is related to blood oxygen saturation and reflects the energy index; this parameter is obtained through color space conversion, specifically (RG) / (R+G+B), and reflects blood oxygenation.

[0070] 3) BVP cycle: This parameter is associated with heart rate variability; it is the blood volume pulse wave cycle separated from the RGB signal by independent component analysis.

[0071] 4. Facial temperature characteristics

[0072] 1) Facial temperature gradient: This parameter is related to the activity of the cross nerve; this parameter is the pixel brightness difference between the forehead and cheek areas identified by the camera;

[0073] 2) Temperature change rate: This parameter is related to vitality and affects the stress index; this parameter is the rate of change of facial brightness per unit time calculated by inter-frame difference.

[0074] 5. Head posture characteristics:

[0075] 1) Head translation amplitude: This parameter is related to fatigue and concentration; this parameter is the root mean square displacement of the center of the eyebrows on the x / y / z axes obtained by MediaPipe pose estimation;

[0076] 2) Head rotation angle: This parameter is related to attention; this parameter is the absolute value of pitch angle / yaw angle / roll angle calculated by Euler angles;

[0077] 3) Head movement frequency: This parameter is related to alertness and affects the inhibition index; this parameter is the dominant frequency of head translational movement obtained through FFT spectrum analysis.

[0078] 6. Facial movement characteristics:

[0079] 1) AU1 (brow lift) intensity: This parameter is related to emotions such as surprise and worry; this parameter is the ratio of the vertical displacement of the inner brow peak obtained by the FACS-AU detection algorithm to the baseline value;

[0080] 2) AU4 (brow drooping) duration: This parameter is related to emotions such as stress and anxiety; this parameter is the number of frames per second of eyebrow spacing reduction obtained through dynamic feature tracking technology × frame rate;

[0081] 3) AU12 (corner of the mouth lift) amplitude: This parameter is related to emotions such as pleasure and confidence; this parameter is the vertical displacement of the corner of the mouth point obtained through facial key point detection technology;

[0082] 4) AU25 (mouth opening) frequency: This parameter is related to excitement and the desire to express; this parameter is the number of mouth opening actions per unit time.

[0083] 7. Eye features:

[0084] 1) Blinking frequency: This parameter is related to fatigue and concentration; it is the number of blinks per minute obtained by an eye state detection algorithm (a blink is considered when the upper eyelid covers more than 50% of the pupil).

[0085] 2) Pupil diameter change rate; this parameter is related to alertness and affects the anxiety index; this parameter is the change rate of the major axis length of the pupil in adjacent frames obtained by ellipse fitting algorithm.

[0086] 8. Characteristics of respiratory movements:

[0087] 1) Shoulder vibration frequency: This parameter is related to the respiratory rate; this parameter is the dominant frequency of the vertical displacement of the midpoint of the clavicle obtained by optical flow method and bandpass filtering method;

[0088] 2) Chest fluctuation amplitude: This parameter is related to breathing depth and affects the energy index; this parameter is the peak value of vertical displacement of chest region pixels obtained by motion tracking algorithm.

[0089] 9. Characteristics of respiratory-emotion coupling:

[0090] 1) Abnormal breathing rate: This parameter is related to emotions such as anxiety and panic; this parameter is the percentage of abnormal breathing cycles (such as breath-holding, rapid breathing);

[0091] 2) Emotion-breath phase difference: This parameter is related to emotional stability; it is the time difference between the peak of emotional fluctuation and the peak of breathing obtained from cross-correlation analysis.

[0092] In S200, the process of training an indicator mapping model based on an artificial intelligence model includes the following steps:

[0093] S211: Configure the data acquisition module in the standard acquisition environment and have the tester sit quietly in the standard acquisition environment in a standard posture.

[0094] S212: Set up standardized tasks to induce different emotional states in test subjects; among them, standardized tasks include the Trier social stress test and the Stroop color word test.

[0095] S213: While inducing different emotional states in the test subjects, the test data of the test subjects is collected through the data acquisition module; the test data includes video data and physiological indicators.

[0096] S214: Preprocess the test data of the testers to obtain standard training data; train the artificial intelligence model based on the standard training data, and label the trained artificial intelligence model as an index mapping model; wherein, the artificial intelligence model includes CNN-LSTM or CNN-GRU.

[0097] In S211, the data acquisition module mainly includes: 1. Non-contact acquisition devices: multi-channel cameras (high-definition, high-frame-rate cameras), which can be equipped with infrared supplementary lighting modules to adapt to low-light environments; at the same time, a depth camera can also be equipped for three-dimensional head motion tracking; 2. Contact acquisition devices: including EEG caps, heart rate belts, electromyography patches, etc., to acquire EEG, HRV, EMG, skin conductance and other signals.

[0098] Standard acquisition environment: 1. Lighting: Brightness 200-1000 lux, color temperature 3800K, change rate ≤1Lux / s, and strong light should be avoided from shining directly on the face; 2. Background: It should be a solid color without clutter to reduce visual interference, such as a gray screen;

[0099] Standard posture: The test subject needs to sit still with their face facing the multiple cameras, their head and neck in the center of the frame, and their limbs in a quasi-static state (limb movement ≤2.5cm).

[0100] In S212, a standardized task refers to a carefully designed task paradigm with a unified process and quantifiable standards to induce specific emotional states (such as stress, anxiety, focus, etc.) in test subjects for scientific research or technological verification. The principles and specific testing procedures of testing methods such as the Trier Social Stress Test and the Stroop Color Word Test have been disclosed in existing schemes and will not be described in detail here.

[0101] In S214, when acquiring standard training data, both contact and non-contact acquisition devices simultaneously collect data from the test subject. When training the artificial intelligence model using the standard training data, the video data collected by the non-contact acquisition device serves as the model input, while the physiological indicators collected by the contact acquisition device serve as the model output.

[0102] For example, the training process of an artificial intelligence model is described using a CNN-LSTM model:

[0103] 1) Construct the network structure of the CNN-LSTM model, including CNN layers, LSTM layers and regression layers.

[0104] CNN layers can utilize the spatial features of the input video data, while LSTM layers are used to extract the temporal features of the input video data. Regression layers are used for regression analysis to obtain physiological indicators.

[0105] 2) Synchronize the video frames and physiological indicators in the standard training data with time, ensuring a timestamp error of <10ms; divide the video data into 10-second video segments (300 frames in total), and take the mean / peak / frequency domain characteristics of the corresponding physiological indicators during the synchronization period. The frame rate of the video data acquisition should be ≥25FPS, with a typical value of 30FPS.

[0106] 3) Pre-train the CNN to independently train it to extract the spatial features of the test subject's limb vibrations; after pre-training, retain the weights of the CNN convolutional layers and remove the classification head.

[0107] 4) When training the fusion model CNN-LSTM, CNN first identifies the spatial features of video segments, then extracts temporal features through LSTM, and predicts physiological indicators through regression layers based on the spatial and temporal features.

[0108] In S200, when the index mapping model is trained based on test data from several different testers, after the feature data is input into the index mapping model to obtain several corresponding physiological indicators, the physiological coordinates need to be calibrated and corrected.

[0109] When training the metric mapping model, the feature data in the standard training data comes from the testers, and physiological indicators will also have certain deviations depending on the testers. Using the testers' feature data and physiological indicators during testing to train the fusion model is applicable to most testing situations. However, individual differences among target users may cause the physiological indicators predicted by the metric mapping model to deviate from the target user's actual physiological indicators. Furthermore, environmental interference can also lead to prediction biases. Since this invention requires the target user to sit still in a standard posture under specific environmental conditions to collect video data, if the environmental conditions are inconsistent with the acquisition conditions of the standard training data, or if the target user is not in a standard posture, the accuracy of the predicted physiological indicators will also be affected.

[0110] It should be noted that the core process of using CNN-extracted spatial features as input to LSTM for temporal analysis is as follows: First, CNN is used to extract spatial features of facial regions from video frames, flattening the feature maps of consecutive frames into a temporal sequence. This sequence is then input into LSTM to capture dynamic dependencies between frames. Finally, an attention mechanism is used to aggregate the temporal features for physiological indicator prediction. This architecture significantly improves the accuracy of emotion recognition through joint modeling of spatiotemporal features.

[0111] To improve the accuracy of core indicators, the physiological indicators predicted by the indicator mapping model are calibrated.

[0112] In a preferred embodiment, several physiological indicators are calibrated, including:

[0113] S221: Retrieve the target user's indicator calibration function; whereby the indicator calibration function is used to calibrate the deviation between the predicted and actual values ​​of physiological indicators;

[0114] S222: The physiological indicators output by the indicator mapping model are calibrated using the indicator calibration function to obtain the calibrated physiological indicators.

[0115] In emotion recognition, it is necessary to predict several physiological indicators of the target user using an indicator mapping model. Some indicators are not affected by individual differences or environmental changes, or the impact is very small; therefore, these physiological indicators do not require calibration, and an indicator calibration function does not need to be constructed. Other indicators are affected by individual differences or environmental changes, and therefore, an indicator calibration function needs to be constructed to correct them. Thus, the same target user may correspond to multiple indicator calibration functions for different physiological indicators. The fitted indicator calibration functions are stored in a database for easy retrieval at any time.

[0116] Please see Figure 2 The indicator calibration function is constructed through testing and includes:

[0117] S231: When collecting video data of the target user for the first time, several physiological indicators of the target user are also collected simultaneously through a contact acquisition device.

[0118] S232: Predict several physiological indicators of the target user based on video data and an indicator mapping model; obtain the indicator calibration function by fitting the prediction and measurement results of the physiological indicators.

[0119] Contact-based data acquisition devices need to simultaneously measure the target user's physiological indicators during the initial acquisition of video data; these measured physiological indicators serve as the true values. Feature data is extracted from the video data, and an indicator mapping model is used to identify the corresponding physiological indicators, which serve as the predicted values. The indicator calibration function can be linear or non-linear, depending on the mapping relationship between the predicted and true values.

[0120] It is worth noting that the indicator calibration function is constructed synchronously when the target user's video data is collected for the first time. It can be directly called when the target user's emotion is monitored subsequently, without needing to be constructed repeatedly. Since the emotion recognition subject in this invention is a user who needs to be monitored over a long period, constructing the indicator calibration function during the first data collection simplifies the subsequent calibration process and improves calibration efficiency.

[0121] In some other preferred embodiments, if the indicator mapping model is trained based on standard training data obtained during target user testing, the calibration step for physiological indicators can be omitted. The physiological indicators predicted by the indicator mapping model based on the target user video data can be used as standard values, and no correction is needed when the environmental conditions are consistent with the testing environment.

[0122] In S300, several core indicators are determined based on several physiological indicators, including:

[0123] S311: Use physiological indicators that affect core indicators as key indicators, and associate several key indicators with core indicators.

[0124] S312: Determine the weight coefficients of key indicators; perform a weighted summation based on the key indicators and their corresponding weight coefficients to obtain the core indicators; the methods for determining the weight coefficients include the analytic hierarchy process (AHP), principal component analysis (PCA), or random forest method.

[0125] Key metrics include, but are not limited to, the following:

[0126] 1. Impulsiveness Index:

[0127] The Impulsivity Index characterizes an individual's level of emotional excitement and activity. It may manifest as heightened excitement, restlessness, irritability, significantly increased actions and speech, accompanied by a degree of aggression. These emotions and behaviors can originate from internal feelings or external influences. Grading: Below 20 is low, 20-50 is normal, and above 50 is high.

[0128] 2. Pressure Index

[0129] The stress index characterizes the intensity of the impact of stressors on an individual's mental and physical stability. This impact may be related to an imbalance between an individual's internal and external environmental conditions and their own coping abilities, and may result in a series of psychophysiological reactions, such as increased heart rate, elevated blood pressure, rapid breathing, and increased hormone secretion. Grading: Below 20 is low, 20-40 is normal, and above 40 is high.

[0130] 3. Anxiety Index

[0131] Anxiety index represents an individual's level of concern regarding major issues such as life-threatening safety, future prospects, and unpredictable or difficult-to-cope-with events; it may contain various emotional components such as anxiety, worry, sorrow, tension, unease, and irritability. Rating: below 15 is low, 5-40 is normal, and above 40 is high.

[0132] 4. Doubt Index / Irritability Index

[0133] The Irritability Index characterizes the intensity of an individual's emotional response. This index is derived from a comprehensive score of the three indices mentioned above; a high score in the former correlates with a high score in the latter. This index serves as a comprehensive alarm parameter for abnormal behavior. A high response value may indicate significantly increased destructive and aggressive tendencies, emotional instability, a state of high stress or intense anxiety, and susceptibility to irritation and uncontrolled behavior. Grading: Below 20 is low, 20-50 is normal, and above 50 is high.

[0134] 5. Harmony Index

[0135] The Harmony Index is used to characterize the balance between the brain and body. When the brain and body functions are in a harmonious interaction, individuals typically exhibit positive traits such as intelligence, quick thinking, strong learning ability, abundant energy, and high efficiency. Conversely, when the functions of the brain and body become unbalanced, disordered, or even antagonistic, various physiological, psychological, and behavioral disorders can occur. Grading: Below 40 is low, 40-80 is normal, and above 80 is high.

[0136] 6. Confidence Index

[0137] The self-confidence index represents an individual's level of trust and certainty in their own abilities. It is a conscious characteristic and psychological state that positively and effectively expresses self-worth, self-respect, and self-understanding. The self-confidence index influences individual psychology and behavior in many aspects, including learning, competition, employment, and achievement. Level classification: below 50 is low, 50-80 is normal, and above 80 is high.

[0138] 7. Energy Index

[0139] The energy index represents an individual's psychological energy level. Psychological energy is further divided into three dimensions: vitality, motivation, and ability. First, vitality refers to an individual's emotional energy. Vibrant individuals exhibit high energy levels, positive attitudes, proactive behavior, and infectious emotions. Second, motivation refers to willpower. Individuals with motivational characteristics demonstrate directionality and persistence in their behavior under the guidance of consciousness, purpose, and planning. Third, ability refers to an individual's cognitive energy; capable individuals can adopt effective coping strategies to solve problems. Below 50 is low, 50-80 is normal, and above 80 is high.

[0140] 8. Self-discipline index

[0141] The self-discipline index represents an individual's ability to change their mental and physical state to adapt to environmental demands. Through self-regulation, individuals introspect and control their own emotional state, while also being aware of the emotional states of others; based on this, they assess the correctness and objectivity of their behavior, ensuring that their actions achieve the desired effect as much as possible. Grading: Below 50 is low, 50-80 is normal, and above 80 is high.

[0142] 9. Inhibition Index

[0143] The inhibition index characterizes the degree to which an individual's psychosomatic state is suppressed and restricted by external or internal forces. An individual's psychosomatic energy needs to be maintained within a relatively stable range; excessive suppression or a lack of necessary limiting mechanisms can lead to various emotional and physical abnormalities. Grading: Below 15 is low, 15-25 is normal, and above 25 is high.

[0144] 10. Neuroticism Index

[0145] The Neuroticism Index is used to characterize an individual's emotional stability. It's important to note that this index tends to interpret normal behavior rather than mental illness. The scale is as follows: below 10 is low, 10-50 is normal, and above 50 is high.

[0146] The analytic hierarchy process (AHP), principal component analysis (PCA), and random forest method are commonly used methods for determining weights in existing schemes; the process of determining the weight coefficients will not be described in detail here. It should be noted that the data used to determine the weight coefficients are various physiological indicators collected through contact-based data acquisition devices.

[0147] For example, taking the anxiety index as an example, assuming its associated physiological indicators include frowning frequency, heart rate, eye fixation stability, respiratory rate, and head movement frequency, the standardized indicator values ​​are as follows: , , , , The weight coefficients for solving the problem using the random forest method are as follows: , , , , Substituting the above data into the anxiety index Anxiety index can be calculated.

[0148] In the 300, the emotional state of the target user is judged based on several core indicators, that is, the emotional risk of the target user is judged by comprehensively judging the core indicators. Please refer to Table 1 and Table 2 below:

[0149] Table 1. Single Evaluation Content of Core Indicators

[0150]

[0151] Table 2 Comprehensive Evaluation Content Based on Core Indicators

[0152]

[0153] Please see Figure 3 A second aspect of the present invention provides an emotion recognition system based on multimodal data fusion, including an emotion recognition module and a data acquisition module connected thereto;

[0154] Data acquisition module: used to collect test data from test subjects through data acquisition devices, integrate the test data into standard training data; and to collect video data from target users through non-contact acquisition devices; among which, test data includes video data and physiological indicators;

[0155] Emotion recognition module: used to train an artificial intelligence model using standard training data to obtain an emotion mapping model; wherein the artificial intelligence model is CNN-LSTM; and,

[0156] This is used to identify feature data from video data through an emotion mapping model, predict the physiological indicators of target users, and obtain several core indicators by weighted summation of several physiological indicators. Based on these core indicators, the emotional state of the target user is determined.

[0157] The data acquisition module primarily collects the necessary data through data acquisition devices, processes the data as required, and then sends it to the emotion recognition module or stores it. The emotion recognition module is responsible for building, training, and updating the indicator mapping model.

[0158] Non-contact acquisition devices include: multi-channel cameras for acquiring video data; light sources for adjusting ambient lighting during video data acquisition; and depth cameras for acquiring three-dimensional head motion data.

[0159] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A multi-modal data fusion based emotion recognition method, characterized in that, The method comprises the following steps: Collecting video data of a target user, and extracting feature data from the video data; wherein the feature data comprises spatial features and time sequence features; Inputting the feature data into an index mapping model to predict a plurality of physiological indexes; wherein the index mapping model is constructed based on an artificial intelligence model; Summing a plurality of core indexes by weighting the plurality of physiological indexes, and determining an emotional state of the target user based on the plurality of core indexes; Training the index mapping model based on the artificial intelligence model comprises the following steps: Configuring a data collection module in a standard collection environment, and letting a tester sit in a standard posture in the standard collection environment; Setting a standardized task to induce different emotional states of the tester; wherein the standardized task comprises a Trier social stress test and a Stroop color word test; Collecting test data of the tester by the data collection module while inducing different emotional states of the tester; wherein the test data comprises video data and physiological indexes; Preprocessing the test data of the tester to obtain standard training data, training an artificial intelligence model based on the standard training data, and marking the trained artificial intelligence model as an index mapping model; wherein the artificial intelligence model is a CNN-LSTM; When the index mapping model is trained based on test data of a plurality of different testers, calibrating the predicted physiological indexes, comprising: Calling an index calibration function of the target user; wherein the index calibration function is used to calibrate the deviation between the predicted value and the true value of the physiological index; Calibrating the plurality of physiological indexes output by the index mapping model through the index calibration function to obtain calibrated physiological indexes; The index calibration function is constructed by a test method, comprising: When collecting video data of the target user for the first time, simultaneously collecting a plurality of physiological indexes of the target user by a contact collection device; Predicting a plurality of physiological indexes of the target user based on the video data and the index mapping model; and fitting the index calibration function according to the prediction result and the measurement result of the physiological indexes.

2. The multi-modal data fusion based emotion recognition method as claimed in claim 1, wherein, The core indexes comprise impulsive index, pressure index, anxiety index, doubt index, harmony index, self-confidence index, energy index, self-discipline index, inhibition index and neuroticism index. 3.The multi-modal data fusion based emotion recognition method of claim 1, wherein, When the index mapping model is trained based on the test data of the target user, the plurality of predicted physiological indexes are not calibrated.

4. The multi-modal data fusion based emotion recognition method as claimed in claim 1, wherein, Summing a plurality of core indexes by weighting a plurality of physiological indexes, comprising: Taking the physiological indexes affecting the core indexes as key indexes, associating the plurality of key indexes with the core indexes; Determining a weight coefficient of the key indexes; summing the core indexes by weighting the key indexes and the corresponding weight coefficients; wherein the determination method of the weight coefficient comprises an analytic hierarchy process, a principal component analysis or a random forest method.

5. The multi-modal data fusion based emotion recognition method as claimed in claim 4, wherein, Summing the core indexes by weighting the key indexes and the corresponding weight coefficients, comprising: Key indicators are marked as Corresponding weights are marked as ; Core indicators are calculated by formula ; wherein is the number of key indicators.

6. A multi-modal data fusion emotion recognition system for implementing the multi-modal data fusion emotion recognition method of any one of claims 1 to 5, characterized in that, The system comprises an emotion recognition module and a data collection module connected thereto. The data collection module is configured to collect test data of the tester through a data collection device, and integrate the test data into standard training data. The video data of the target user is collected through a non-contact collection device. The emotion recognition module is configured to train an artificial intelligence model through the standard training data to obtain an emotion mapping model. The emotion mapping model is a CNN-LSTM model. The emotion recognition module is configured to identify feature data of the video data through the emotion mapping model, predict physiological indicators of the target user, and determine the emotional state of the target user based on the core indicators. The index mapping model is trained based on the artificial intelligence model, including the following steps: The data collection module is configured in a standard collection environment, and the tester is asked to sit in a standard posture in the standard collection environment. A standardized task is set to induce different emotional states of the tester. The standardized task includes the Trier Social Stress Test and the Stroop Color-Word Test. The test data of the tester is collected through the data collection module while the tester is in different emotional states. The test data of the tester is preprocessed to obtain standard training data. The artificial intelligence model is trained based on the standard training data, and the trained artificial intelligence model is marked as an index mapping model. The index mapping model is trained based on the test data of several different testers. The predicted physiological indicators are calibrated, including: The index calibration function is used to calibrate the deviation between the predicted value and the true value of the physiological indicators.

7. The multi-modal data fusion based emotion recognition system as claimed in claim 6, wherein, The index calibration function is used to calibrate the physiological indicators output by the index mapping model to obtain calibrated physiological indicators. The index calibration function is constructed through testing, including: The physiological indicators of the target user are collected through a contact collection device while collecting the video data of the target user for the first time. The physiological indicators of the target user are predicted based on the video data and the index mapping model. The index calibration function is obtained by fitting the prediction results and the measurement results of the physiological indicators. The non-contact collection device includes: A multi-channel camera is configured to collect video data. A light source is configured to adjust the ambient light during video data collection. A depth camera is configured to collect head three-dimensional motion data.

Citation Information

Patent Citations

  • Non-contact anxiety recognition method and device based on face video

    CN113326781A

  • Student psychological health group screening method and system

    CN114792553A

  • Pulse wave calibration method based on millimeter wave radar

    CN119344693A