Emotion recognition and intervention system based on facial micro-expression and physiological signal fusion
By constructing a multimodal data fusion and real-time emotion intervention system, the limitations of single-modal perception and lack of real-time intervention in existing technologies are solved, achieving high-precision emotion recognition and proactive psychological adjustment, which is applicable to mental health management and high-stress work scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-03-27
AI Technical Summary
Existing emotion recognition technologies face limitations in complex environments, including single-modal perception, unstable recognition accuracy, poor dynamic adaptability, and a lack of real-time closed-loop intervention mechanisms. These limitations make it difficult to meet the application needs of mental health management, human-computer interaction optimization, and high-pressure work scenarios.
The system comprises a multimodal data synchronous acquisition subsystem, a dynamic feature fusion and emotion decoding subsystem, a personalized emotion baseline modeling subsystem, a cognitive load and emotion fluctuation correlation analysis module, an adaptive feedback intervention generation subsystem, and a wearable multi-channel physiological regulation terminal, enabling efficient fusion of facial micro-expressions and physiological signals and real-time emotion intervention.
It significantly improves the accuracy and stability of emotion recognition, realizes the leap from judging instantaneous state to predicting long-term psychological trends, forms a technological paradigm for proactively regulating psychological state, and provides a system-level solution with clinical translational potential for the field of mental health.
Smart Images

Figure CN121456675B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence and health monitoring, and particularly relates to an emotion recognition and intervention system based on fusion of facial micro-expression and physiological signals. BACKGROUND
[0002] Emotion recognition and intelligent intervention technology, as a core direction in the field of human-computer interaction and mental health management, has made significant progress in recent years under the cross-promotion of artificial intelligence, biosensing and cognitive science. This technology aims to achieve objective and real-time monitoring of emotional state by analyzing individual external behavior and internal physiological activity, and accordingly provides corresponding feedback or intervention measures. Its application has covered multiple important scenarios such as auxiliary diagnosis of psychological diseases, driver fatigue warning, and educational affective computing, and has become one of the key enabling technologies to improve the quality of human life and work efficiency. Among them, non-invasive emotion perception methods have attracted widespread attention due to their good user experience and deployability.
[0003] Among them, the emotion recognition method based on multi-modal information fusion has gradually become the mainstream technology path. This method integrates facial expressions, speech intonation, heart rate variability, skin galvanic response and other physiological and behavioral signals to break through the limitations of single modal information and obtain more comprehensive and stable emotion recognition results. In particular, facial micro-expression is considered an important clue to reveal hidden emotions due to its short duration, small amplitude and strong emotion specificity; and the synchronous physiological signals such as electroencephalogram and electrocardiogram can reflect the autonomic nervous response under emotional stimulation, providing deep evidence for recognition. Therefore, how to effectively fuse these two types of heterogeneous but complementary information sources to build a high-precision and low-delay emotion recognition model has become a key research direction.
[0004] Although existing technologies attempt to combine facial micro-expression analysis and physiological signal monitoring, they still face multiple challenges in practical applications: multi-source signals have significant differences in time scale and spatial resolution, making feature alignment difficult; micro-expression capture is easily affected by light, occlusion and individual differences, resulting in unstable recognition accuracy; physiological signals are inherently sensitive, but are also vulnerable to external interference and lack a clear mapping relationship with specific emotion categories; more importantly, most systems only focus on the emotion recognition link and fail to achieve closed-loop control from perception to active intervention, which cannot meet the application needs of real-time psychological regulation. Especially in critical scenarios such as high-pressure work and remote medical care, the above defects may lead to emotional misjudgment or response lag, seriously affecting the practicality and reliability of the system. Therefore, there is an urgent need for an integrated system architecture that can efficiently fuse facial micro-expression and physiological signals and has real-time emotion intervention capability. SUMMARY
[0005] The present application aims to provide an emotion recognition and intervention system based on facial micro-expression and physiological signal fusion, to solve the problems of single modal perception limitation, unstable recognition accuracy, poor dynamic adaptation ability and lack of real-time closed-loop intervention mechanism of existing emotion recognition technology in complex environment. Current emotion recognition technology mainly relies on single information source in visual or physiological signal, leading to significant decline in recognition accuracy under light change, face occlusion, individual difference or low signal-to-noise ratio scene; at the same time, most systems only realize passive monitoring, fail to build a complete technical closed loop from accurate recognition to active adjustment, and are difficult to meet the actual application needs in psychological health management, human-computer interaction optimization and high pressure operation scene.
[0006] The technical scheme of the present application comprises a multimodal data synchronous acquisition subsystem, a dynamic feature fusion and emotion decoding subsystem, a personalized emotion baseline modeling subsystem, a cognitive load and emotion fluctuation correlation analysis module, an adaptive feedback intervention generation subsystem, a wearable multi-channel physiological regulation terminal, and a cross-period emotion evolution tracking engine. The multimodal data synchronous acquisition subsystem is composed of a high-frame-rate infrared imaging unit and a multi-channel biological signal sensing array. The former captures the subtle movements of facial muscles, especially the displacement changes in the orbicularis oculi muscle, zygomatic major muscle, and corrugator muscle region, at a sampling frequency of not less than 120 frames per second. The latter synchronously acquires skin conductance response, heart rate variability, forehead temperature gradient, and electroencephalogram alpha band power spectral density. The dynamic feature fusion and emotion decoding subsystem receives the original time series data stream. First, it models the optical flow field of micro-expression video sequences and extracts the action unit intensity time series. At the same time, it performs wavelet denoising and frequency band energy extraction on physiological signals. Then, it inputs the two types of features into a dual-path time series neural network. The visual path uses a 3D convolutional long short-term memory network to process local action unit dynamics, and the physiological path uses an attention-enhanced bidirectional gated recurrent unit to mine autonomic nervous system response patterns. The outputs of the two paths are aligned in context and complemented in information through a cross-attention mechanism with learnable weights in the intermediate layer. Finally, an eight-dimensional discrete emotion category probability distribution and continuous dimension arousal and valence estimates are output. The personalized emotion baseline modeling subsystem performs a 72-hour non-interference baseline acquisition during the initial use stage of the user, establishes a model of micro-expression activity frequency distribution and resting physiological parameter combination for the individual in the daily state, and calculates the standard deviation multiple of the deviation from the baseline in subsequent detection as a preliminary criterion for abnormal emotional activation. The cognitive load and emotion fluctuation correlation analysis module continuously monitors the trend of the θ wave and β wave energy ratio of the prefrontal lobe region electroencephalogram signal, combines the task completion delay rate and pupil diameter fluctuation amplitude to construct a cognitive resource consumption index, and performs Granger causality analysis with the emotion decoding results to identify secondary anxiety or frustration emotions caused by cognitive overload. The adaptive feedback intervention generation subsystem generates customized instruction sequences according to the emotion type, intensity level, and cause classification by calling the preset intervention strategy library. When acute anxiety is detected, a breathing guidance program is triggered. When persistent depression is identified, positive memory cues are prompted. When cognitive overload is determined, the terminal suggests a phased task interruption and plays specific frequency binaural rhythm audio. The wearable multi-channel physiological regulation terminal integrates a miniature bone conduction speaker, a distributed thermal stimulation patch, and a transcutaneous vagus nerve electrical stimulation electrode to accurately apply auditory, somatosensory, and neuroelectrophysiological triple regulation signals according to the intervention instructions. The thermal stimulation patch is distributed along both sides of the neck to activate the parasympathetic nervous reflex pathway by raising the skin temperature to the 36.5℃ to 38.0℃ interval at a rate of 0.5℃ per second.The cross-period emotion evolution tracking engine aggregates emotion recognition results by day, constructs individual emotion fluctuation heat maps, infers potential psychological state transition rules using a hidden Markov model, and predicts the probability of high-risk emotional events occurring within the next 24 hours based on long-term trends, allowing for the early deployment of preventive intervention plans.
[0007] Further, a modal confidence evaluation mechanism is introduced in the dynamic feature fusion process, which adjusts the input weight of the dual-channel in real time based on the current signal quality index; when the signal-to-noise ratio of the facial image is lower than the preset threshold, the contribution weight of the visual channel is automatically reduced to not more than 30% of the total fusion weight, while the dominance of the physiological signal channel is enhanced; the confidence evaluation includes but is not limited to pupil visibility, facial contour integrity, skin conductance baseline drift, and EEG artifact proportion, ensuring that the system can still maintain reliable output under the condition of degradation of some sensory channels.
[0008] Further, the training process of the dual-channel time series neural network adopts an adversarial domain adaptation strategy, and is jointly optimized on a multi-center database containing samples of different ages, genders, and races, forcing the network to learn shared emotion representation features across populations, and weakening the interference of individual appearance differences on micro-expression decoding; virtual occlusion enhancement technology is introduced in the training stage to simulate random mask, glasses or hand occlusion, improving the robustness of the model under real occlusion conditions.
[0009] Further, the strategy library update mechanism of the adaptive feedback intervention generation subsystem is implemented based on a reinforcement learning framework, the system records the change in emotional state before and after each intervention as a reward signal, evaluates the long-term effectiveness of different intervention methods in various situations through a Q-learning algorithm, and dynamically adjusts the priority ranking of strategies; the user's subjective rating of the intervention experience is also included as an auxiliary reward item in the update process, forming a human-machine collaborative optimization closed loop.
[0010] Further, the electrical stimulation parameters of the wearable multi-channel physiological regulation terminal follow an individualized setting protocol, the initial stimulation intensity is determined by a vagus nerve threshold test program, which increases the current output at a step rate of 5 milliamperes per minute until the user reports a slight throat constriction, and takes 60% of this threshold as the starting working intensity; the stimulation waveform uses an asymmetric rectangular pulse, with a frequency locked at 25 Hz and a duty cycle controlled at 20% to balance neural activation efficiency and wearing comfort.
[0011] Further, the cross-period emotion evolution tracking engine has a group mode analysis function, allowing the data of multiple authorized users to be aggregated and modeled after desensitization, identifying the emotional transmission path and group stress response pattern at the organizational level, and providing macro situational awareness support for enterprise team management, emergency command and dispatch, and other scenarios.
[0012] Further, the infrared imaging unit of the multi-modal data synchronous acquisition subsystem is equipped with a narrowband filter, with a center wavelength set at 850 nanometers, effectively suppressing environmental visible light interference and improving micro-expression texture contrast in dark light environments; the lens group has automatic focusing and anti-shake functions, ensuring that key facial regions are not lost when the head moves slightly.
[0013] Further, the personalized emotion baseline modeling subsystem sets a baseline update trigger condition, when the system detects that the mean value of the user's basic physiological parameters deviates from the initial baseline by more than 2 standard deviations for 7 consecutive days, a new baseline collection program is started, avoiding misjudgment caused by changes in living habits or evolution of health status.
[0014] Compared with the prior art, the advantages and positive effects of the present application are:
[0015] The system fundamentally overcomes the inherent defects of single-modal perception in complex application scenarios by constructing a spatio-temporal alignment fusion architecture of facial micro-expression and multi-dimensional physiological signals, significantly improving the accuracy and stability of emotion recognition in real environments; the dynamic weighted feature fusion mechanism endows the system with continuous operation capability under the condition of partial sensor failure or signal degradation, enhancing the engineering practicability; the combination of personalized baseline modeling and cross-period evolution tracking realizes the capability leap from instantaneous state judgment to long-term psychological trend prediction, enabling emotion management to shift from passive response to active prevention; the causal association analysis module of cognitive load and emotional fluctuation reveals the deep psychological mechanism, providing a scientific basis for precise intervention; the multi-channel wearable regulation terminal integrates three physical intervention modalities, forming a synergistic physiological regulation network, which exhibits stronger neural regulation efficacy compared to single feedback mode; the adaptive intervention strategy library relies on reinforcement learning to realize continuous evolution, ensuring that the system continuously optimizes service quality as the use time increases; the overall technical solution forms a complete closed loop of "perception-decoding-decision-intervention-feedback", not only realizing high-precision emotion recognition, but also establishing a technical paradigm for actively adjusting psychological state, providing a system-level solution with clinical transformation potential for the field of mental health, which has wide application prospects in high-risk scenarios such as intelligent driving, remote medical care, and special operation personnel monitoring. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is the overall technical scheme architecture schematic diagram of the emotion recognition and intervention system based on facial micro-expression and physiological signal fusion proposed by the present application;
[0017] Figure 2 is the core principle framework schematic diagram of the dynamic feature fusion and emotion decoding subsystem in the present application. DETAILED DESCRIPTION
[0018] Please refer to Figure 1 andFigure 2 The present application provides a facial micro-expression and physiological signal fusion-based emotion recognition and intervention system, which has a technical architecture composed of a multi-modal data synchronous acquisition subsystem, a dynamic feature fusion and emotion decoding subsystem, a personalized emotion baseline modeling subsystem, a cognitive load and emotion fluctuation correlation analysis module, an adaptive feedback intervention generation subsystem, a wearable multi-channel physiological regulation terminal, and a cross-period emotion evolution tracking engine. Each subsystem works collaboratively according to the preset data flow and timing logic, forming a complete technical chain from raw sensory signal input to closed-loop psychological regulation output. The system is deployed in a hybrid architecture of edge computing and cloud computing, where data acquisition, preliminary processing, and intervention execution tasks with high real-time requirements are completed on local edge devices, while model training, long-term trend analysis, and group-level aggregated modeling are performed on cloud platforms with high-performance computing resources.
[0019] The overall technical process of the system begins with the user wearing a wearable multi-channel physiological regulation terminal and starting the system initialization program. The multi-modal data synchronous acquisition subsystem then enters a continuous working state, synchronously acquiring the user's facial image sequence and multi-dimensional physiological signals. The dynamic feature fusion and emotion decoding subsystem receives the raw time series data stream from the previous subsystem, performs high-order feature extraction and cross-modal fusion, and outputs the current emotional state estimation result. This result is sent to the personalized emotion baseline modeling subsystem for individualized deviation assessment, and to the cognitive load and emotion fluctuation correlation analysis module to determine the type of emotional triggers. Based on the comprehensive determination result, the adaptive feedback intervention generation subsystem calls the strategy library to generate customized intervention instructions, and implements physical regulation through the wearable terminal. All historical recognition records and intervention response data are aggregated into the cross-period emotion evolution tracking engine to build individual emotion dynamic models and predict future risk events. The entire system forms a full-closed-loop structure of "perception-decoding-decision-intervention-feedback", achieving active, accurate, and continuous management of psychological state.
[0020] The multi-modal data synchronous acquisition subsystem, as the information input front end of the system, undertakes the task of capturing multi-source physiological signals with high quality, low delay, and time-space alignment. The subsystem is composed of two core sensing units: a high-frame-rate infrared imaging unit and a multi-channel biological signal sensing array. Both units are synchronized by a hardware-level timestamp synchronization mechanism to ensure that all collected data are time-aligned within a millisecond accuracy, avoiding feature misalignment caused by asynchronous sampling. The synchronization mechanism uses the IEEE 1588 Precision Time Protocol to establish a unified clock source within the main control chip, and all sensors use this as a reference for data packaging and transmission.
[0021] The high-frame-rate infrared imaging unit is installed in the front of the head-mounted frame, directly facing the center area of the user's face, with an optical axis angle of less than ± 5 degrees to the extension line of the nose bridge, ensuring coverage of the key motion unit activity areas such as the eye, cheekbone, and inter-brow. The core imaging device of the unit is a global shutter CMOS sensor with an effective pixel of not less than 1920x1080, supporting progressive scan mode to eliminate motion blur caused by rolling shutter effect. The sampling frequency is set to 120 frames per second, meeting the Nyquist sampling theorem requirement for capturing micro-expression events with a duration of less than 500 milliseconds. To improve the imaging quality in low-light environments, the unit is equipped with a narrow-band filter with a center wavelength of 850 nanometers, which is in the near-infrared region and can effectively penetrate the shallow tissue of the skin, enhance the contrast of subcutaneous capillary network and muscle texture, and suppress visible light background noise interference. The lens group adopts a fixed-focus wide-angle design with a focal length of 4 mm and a field of view angle of 85 degrees, which can maintain the complete facial profile in the frame even with slight head movement. The auto-focus algorithm is based on the Laplacian variance evaluation function to calculate the image sharpness in real time, and when the variance value drops by more than 15% of the initial stable value, the focus motor adjusts the lens position; the anti-shake function detects the device posture changes through a three-axis micro-gyroscope, and combines with the electronic image stabilization algorithm to crop and translate the video frames, compensating for the picture shift caused by head shaking.
[0022] The raw video stream output by the infrared imaging unit is in 16-bit grayscale format, and each frame contains facial surface temperature distribution and local deformation information. The system performs difference operation on consecutive video frames to generate inter-frame optical flow map for subsequent motion unit intensity quantization. To ensure privacy and security, all image data is destroyed immediately after feature extraction on the local device, without any form of long-term storage or network upload.
[0023] The multi-channel biosignal sensing array is integrated in the ear support, the inner side of the forehead band, and the neck patch, realizing non-invasive, continuous, and multi-point physiological monitoring. The array includes four independent sensing channels: skin galvanic response channel, heart rate variability channel, forehead temperature gradient channel, and electroencephalogram alpha band power spectral density channel. Each channel sensor is manufactured using flexible printed circuit board technology, with good curved surface adhesion and anti-motion artifact capability.
[0024] The skin conductance response channel uses a silver / silver chloride dry electrode pair, placed on the left ear lobe and the right wrist inner side, respectively, with a constant bias voltage of 0.5 volts applied. The skin conductance level between the two electrodes is measured. After being amplified by 1000 times by an instrument amplifier, the signal passes through a 0.05 hertz high-pass filter to remove the direct current drift, and then passes through a 50 hertz notch filter to suppress power frequency interference, and finally digitized at a sampling rate of 256 hertz. The heart rate variability channel uses the photoplethysmography method, and a red light emitting diode and photodetector combination is arranged behind the left ear lobe to detect the periodic changes in blood perfusion. After the original pulse wave signal is band-pass filtered (0.5 hertz to 5 hertz), the R-R interval sequence is extracted by a peak detection algorithm, which serves as the basis for calculating heart rate variability. The forehead temperature gradient channel is arranged with three micro thermistors in a triangular distribution 2 centimeters above the center of the eyebrows with a spacing of 1.5 centimeters, with a sampling resolution of 0.01 degrees Celsius, to capture the small temperature differences caused by local blood flow changes. The electroencephalogram alpha band power spectral density channel is set with silver paste printed electrodes on both sides of the temples, with a reference electrode at the root of the nose and a ground electrode below the occipital protuberance, forming a monopolar lead configuration. The signal is connected to a low-noise preamplifier through an impedance matching circuit, with a gain setting of 5000 and a band-pass filtering range limited to 8 hertz to 13 hertz, specifically extracting the alpha rhythm component related to the relaxation state. All physiological signals are preliminarily conditioned at the analog front end, then synchronously sampled by a 24-bit resolution analog-to-digital converter, and after being labeled with precise time tags, they are transmitted to the subsequent processing unit.
[0025] The dynamic feature fusion and emotion decoding subsystem is the core intelligent hub of the system, responsible for converting raw multi-modal signals into emotion state descriptions with semantic meaning. The subsystem first preprocesses and extracts features from two types of heterogeneous data streams, then implements deep feature fusion through a dual-path time series neural network architecture, and finally outputs an eight-dimensional discrete emotion category probability distribution and a two-dimensional continuous emotion space coordinate.
[0026] For micro-expression video sequences, the system performs optical flow field modeling. Based on the original video at 120 frames per second, the Farnebäck dense optical flow algorithm is used to calculate the pixel-level displacement vector field between adjacent frames. The algorithm builds a local space-time neighborhood model around each pixel point by constructing a quadratic polynomial approximation of the image intensity function, and solves the displacement parameters in the least squares sense. The resulting optical flow map is presented as a two-dimensional vector field, where the vector direction represents the motion direction and the modulus length represents the motion speed. The system focuses on the key regions defined by the Facial Action Coding System: left and right orbicularis oculi regions (corresponding to blinking and squinting actions), zygomaticus major region (smiling), corrugator region (frowning). Within these regions, the system calculates the absolute value of the vertical component of the average optical flow vector as a proxy indicator of action unit intensity. For example, in the left orbicularis oculi region, if the absolute value of the vertical component exceeds 2 standard deviations from the resting level for 3 consecutive frames and the duration is between 200 milliseconds and 500 milliseconds, it is determined that an AU43 (closed eye) event has occurred. All detected action unit events are arranged in chronological order to form an action unit intensity time series, which serves as the input feature of the visual channel.
[0027] For multi-channel physiological signals, the system implements wavelet denoising and sub-band energy extraction. The Daubechies wavelet basis (db4) is used to perform 5-level discrete wavelet decomposition on the original signal, decomposing the signal into detail coefficients and approximation coefficients at different scales. For electrodermal response signals, the 3rd to 5th layer detail coefficients are retained to reconstruct the signal, effectively filtering out slow drift caused by respiratory motion and high-frequency electromagnetic noise; for heart rate variability R-R interval sequences, the Morlet continuous wavelet transform is used to extract low-frequency components of 0.04 Hz to 0.15 Hz and high-frequency components of 0.15 Hz to 0.4 Hz, respectively, reflecting sympathetic and parasympathetic nerve activity; for electroencephalogram signals, the average power spectral density in the 8 Hz to 13 Hz frequency band is directly calculated as a quantitative indicator of alpha band activity level. All processed physiological signals are normalized to the [0, 1] interval to form fixed-length time window feature vectors, with a sliding window width of 10 seconds and a step size of 2 seconds.
[0028] The two types of features are input into a dual-path temporal neural network. The visual path uses a 3D convolutional long short-term memory network structure. The 3D convolutional layer is stacked with 3 layers, with kernel sizes of (3, 3, 3), (3, 3, 3), and (1, 1, 1), respectively, with a step size of 1, and zero padding to ensure that the spatial dimension remains unchanged. The time dimension convolution kernel captures short-term dynamic patterns. Each layer is followed by batch normalization and ReLU activation functions. The spatio-temporal feature maps output by the 3D convolutional layer are flattened and fed into two layers of stacked long short-term memory networks with a hidden layer dimension of 128, connected bidirectionally to utilize past and future context information simultaneously. The final output is a 256-dimensional visual semantic embedding vector.
[0029] The physiological channel adopts an attention-enhanced bidirectional gated recurrent unit network. The original physiological feature vector sequence passes through two layers of bidirectional gated recurrent units in turn, each layer containing 128 hidden units. The Bahdanau attention mechanism is introduced at the output end of the second layer to calculate the alignment weight of the current time output with the entire sequence, and the context vector is obtained by weighted summation. The attention weight reflects the importance difference of physiological signals at different time points to the current emotional state judgment. For example, in the initial stage of stress response, the skin conductance response may obtain a higher weight; while in the recovery stage, the heart rate variability index weight rises. The final output is a 256-dimensional physiological semantic embedding vector.
[0030] The outputs of the two channels are fused in the middle layer through a cross-attention mechanism with learnable weights to realize information complementation. Let the visual semantic embedding sequence be V∈ ^(T×D), and the physiological semantic embedding sequence be P∈ ^(T×D), where T is the number of time steps, and D is the feature dimension. The cross-attention calculation is as follows:
[0031]
[0032] where, 、 、 are the projection matrices of query, key, and value respectively, is the dimension of the key vector. This mechanism allows the visual features to focus on the most relevant physiological state segments, and vice versa, achieving cross-modal context alignment. The fused feature vector concatenates the original double-channel output to form a 768-dimensional joint feature representation. This representation is compressed to 128 dimensions through a fully connected layer, and then branched into two output heads: one Softmax classifier outputs the probability distribution of eight emotions: anger, disgust, fear, happiness, sadness, surprise, contempt, and neutral; the other regression head outputs the arousal and valence estimates in the [-1, 1] interval, forming the coordinates of the continuous emotional space.
[0033] To cope with the problem of partial modality signal degradation in real-world environment, a modality confidence evaluation mechanism is introduced. This mechanism dynamically adjusts the input weight of the dual-path based on real-time signal quality indicators. For the visual channel, the evaluation indicators include the pupil detection success rate, the facial contour integrity score, and the image signal-to-noise ratio. Pupil detection uses Hough circle transformation combined with template matching. If the pupil centers of both eyes cannot be located for 5 consecutive frames, it is determined that the occlusion is severe. The facial contour integrity is quantified by calculating the root mean square of the fitting residual of the active shape model. A residual exceeding 3 pixels is considered distorted. The image signal-to-noise ratio is estimated by the ratio of local variance to global variance. When any of the above indicators deteriorates below the threshold, the system starts a confidence decay function, which linearly down-regulates the contribution weight of the visual path in fusion, with the minimum down-regulation to 30% of the total weight. For the physiological channel, the evaluation indicators include the skin electricity baseline drift rate (more than 0.5 microsiemens per minute is considered unstable) and the proportion of EEG artifacts (the energy proportion of eye movement and electromyographic components is detected by independent component analysis, and more than 20% is considered severely contaminated). When the quality of physiological signals decreases, the weight of the physiological channel is reduced accordingly, and the dominance of the visual channel is enhanced. The fusion weight adjustment process is controlled by a small gating network, which takes various quality indicators as input and outputs normalized dual-path weight coefficients, ensuring that the system can still maintain reliable output under single modality degradation conditions.
[0034] The personalized emotional baseline modeling subsystem aims to establish the psychophysiological reference system of the individual in the steady state condition, providing individualized criteria for subsequent anomaly detection. The system starts a 72-hour non-interference baseline acquisition phase after the user first registers. During this period, the system runs in the background but does not make any active prompts or interventions, and only records the user's multi-modal data in natural states such as daily work, rest, and entertainment. The acquisition content includes the micro-expression activity frequency statistics (such as the number of eye blinks per hour, the proportion of smile duration) extracted every minute, the combination of resting physiological parameters (average skin electricity level, heart rate variability low / high frequency ratio, forehead average temperature, alpha band power mean), and environmental light intensity, background noise level, and other covariates.
[0035] The baseline data is stored in the local encrypted database according to the timestamp. After 72 hours, the system starts the baseline modeling algorithm. For discrete micro-expression features, kernel density estimation method is used to construct probability distribution model to determine the normal fluctuation range of each action unit frequency. For continuous physiological parameters, the mean and standard deviation of each parameter in the wakeful period are calculated to form a multivariate Gaussian distribution hypothesis. All parameters are modeled according to the circadian rhythm, with separate baseline models established for daytime (8:00-20:00) and night (20:00-8:00) to reflect the natural changes of physiological rhythm. The final individual baseline model contains a set of parameterized distribution functions and their covariance matrices, which fully characterize the user's multi-dimensional feature space distribution in the baseline state.
[0036] In the daily usage phase, the system computes the degree of deviation of current detection values from the baseline model in real-time. For single variable, the Z-score normalization formula is adopted:
[0037]
[0038] where, is the current observation, is the baseline mean, is the baseline standard deviation. For multi-variable joint deviation, the Mahalanobis distance metric is adopted:
[0039]
[0040] where, is the baseline covariance matrix. When the Mahalanobis distance exceeds the preset threshold of 3.0, it is determined that there is a significant emotional activation event, triggering the deep emotional decoding process. This mechanism effectively distinguishes real emotional fluctuations caused by external stimuli from static deviations caused by individual inherent physiological differences, greatly reducing the false positive rate.
[0041] To further improve the long-term effectiveness of the baseline model, the system sets the baseline update trigger condition. When the mean of the basic physiological parameters (such as average skin electricity level, resting heart rate) calculated daily for 7 consecutive days continuously deviates more than 2 standard deviations from the initial baseline, and this trend is confirmed by Mann-Kendall monotonicity test (p<0.05), the system determines that the user's basic physiological state has changed structurally. The reasons for the change may include long-term work and rest adjustment, chronic disease development, or drug treatment influence. At this time, the system pushes a notification to the user, suggesting to start a new 72-hour baseline collection procedure. The newly collected data is used to reconstruct the baseline model, and the old model is archived for reference. This updating mechanism ensures that the baseline always reflects the user's current true state, avoiding the performance degradation of the system caused by model aging.
[0042] The cognitive load and emotional fluctuation correlation analysis module focuses on analyzing the cognitive driving factors behind emotional responses, especially identifying secondary negative emotions caused by cognitive resource overload. The module continuously monitors the relative energy ratio of theta waves (4-8 Hz) and beta waves (13-30 Hz) in the prefrontal region of the brain. Studies have shown that an increase in the θ / β ratio is closely related to attention wandering and cognitive fatigue. The system uses an autoregressive model (AR model, order 16) to estimate the power of each frequency band in real-time, and calculates the θ / β energy ratio within a sliding window (30 seconds).
[0043] Meanwhile, the module integrates task performance data as an external validation indicator of cognitive load. When the user is performing a specific numerical task, the system records the task completion delay rate (the ratio of actual time consumption to standard time consumption) and the operation error rate. By detecting the change in pupil diameter through the eye tracking unit, the standard deviation of the pupil dilation amplitude during the task is calculated, which is positively correlated with the degree of cognitive effort. The above four indicators - theta / beta ratio, task delay rate, error rate, and pupil fluctuation amplitude - are normalized and weighted to generate the cognitive resource consumption index CRI:
[0044]
[0045] where the weight to According to the user's historical data through linear regression fitting, the sensitivity of the index to the real cognitive load is maximized.
[0046] The module performs Granger causality analysis on the CRI time series and the negative emotion intensity series in the emotion decoding results such as anxiety, frustration, and restlessness. The analysis window length is 15 minutes, and the sampling interval is 1 minute. If the CRI sequence has a significant Granger causality relationship with the negative emotion sequence (F-test p<0.01), and the lag order is 1 to 3, it is determined that the current emotional fluctuation is mainly driven by cognitive overload. This determination result directly affects the selection of intervention strategy: for primary anxiety caused by non-cognitive factors, the system prefers to use breathing guidance; while for secondary anxiety caused by cognitive overload, it focuses on task structure adjustment and attention reset training.
[0047] The adaptive feedback intervention generation subsystem selects the optimal intervention scheme from the preset strategy library according to the emotion type, intensity level, and cause classification, and generates a command sequence. The strategy library stores multiple intervention paradigms, each associated with a specific emotional situation label. The system uses a combination of rule engines and reinforcement learning to determine strategy invocation.
[0048] When the dynamic feature fusion and emotion decoding subsystem outputs an eight-dimensional emotion probability distribution with an anxiety category probability exceeding 0.7 and a wakefulness estimate higher than 0.6, the system determines an acute anxiety state. At this time, the breathing guidance program is triggered: the wearable terminal's miniature bone conduction speaker plays pre-recorded voice guidance, guiding the user to perform the 4-7-8 breathing method (inhale for 4 seconds, hold for 7 seconds, exhale for 8 seconds) for 6 cycles. The voice rhythm is precisely controlled by the background metronome, with an error of less than 50 milliseconds.
[0049] When the sadness category probability exceeds 0.65 for 5 consecutive minutes and the valence estimate is lower than -0.5, the system determines that the user is in a persistent low mood. The positive memory cue prompt is activated: the terminal generates a soft tactile signal in the left arm inner side through a vibration motor, while the bone conduction speaker plays positive sound clips related to the user's personal experiences (such as family laughter, applause at the moment of success), which are uploaded and labeled with emotional attributes by the user during the initial setup phase.
[0050] When the cognitive load and emotional fluctuation correlation analysis module confirms that emotions are triggered by cognitive overload, the system suggests a phased task interruption. The terminal screen displays a countdown interface, forcing the user to pause the current work for 5 minutes, and plays binaural beat audio with a frequency of 10 Hz. The binaural beat inputs 100 Hz and 110 Hz pure tones to the left and right ears respectively, and the brain perceives a 10 Hz beat difference frequency, which is in the alpha band, helping to induce a relaxed state and attention recovery.
[0051] The update mechanism of the strategy library is implemented based on the reinforcement learning framework. The system regards each intervention as a decision event, records the emotional state before the intervention , the selected intervention action A_t, and the state after the intervention . The reward function is defined as the improvement of emotional state:
[0052]
[0053] wherein, is a balance coefficient, taking a value of 0.5, emphasizing the importance of reduced arousal in anxiety relief. In addition, the user can give a subjective experience score of 1 to 5 after the intervention by touching the surface of the terminal, and this score is linearly mapped to be added as an auxiliary reward to the total reward. The system uses learning algorithm to update the action value function:
[0054]
[0055] wherein, is the learning rate, taking a value of 0.1; is the discount factor, taking a value of 0.9. After long-term interaction, the system learns the long-term effectiveness of various intervention methods in different situations, dynamically adjusts the priority ranking of strategies, and forms a human-machine collaborative optimization closed loop.
[0056] The wearable multi-channel physiological regulation terminal is the execution mechanism for the system to realize closed-loop intervention, integrating three physical regulation modalities of hearing, somatosensory, and neuroelectrophysiology. The terminal main body is a lightweight head-mounted device, with a total weight of not more than 85 grams, and the shell material is biocompatible polycarbonate. The internal circuit is waterproof and sweatproof, with a protection level of IPX4.
[0057] The miniature bone conduction speaker is embedded in the inner side of the left and right ear rear support, closely attached to the mastoid bone. It uses a piezoelectric ceramic transducer, with a vibration frequency response range of 20 Hz to 10,000 Hz, and a sound pressure level of up to 85 decibels, ensuring clear transmission of voice instructions and audio stimuli in noisy environments while avoiding sound leakage to others.
[0058] The distributed thermal stimulation patch has four patches, symmetrically adhered to the skin surface of the upper end of the trapezius muscle on both sides of the neck, covering the carotid sinus area. The patch has a built-in miniature Peltier thermal element, driven by a constant temperature control circuit. The warming program is set as follows: from the ambient temperature to the target interval of 36.5°C to 38.0°C at a constant rate of 0.5°C per second, maintain for 120 seconds and then naturally cool down. This temperature interval can effectively activate the carotid body chemoreceptor, reflexively enhance vagus nerve tension, promote the increase of heart rate variability HF component, and achieve rapid sedation effect.
[0059] The transcutaneous vagus nerve electrical stimulation electrode is made of medical-grade stainless steel and is located in the mastoid area behind the ear, sharing the contact surface with the bone conduction speaker. The stimulation parameters follow an individualized setting protocol. The initial intensity is determined by the vagus nerve threshold test program: the system increases the current output at a step rate of 5 mA / min, with a starting value of 1 mA, and the user wears the test device and reports the subjective feeling; when the first slight laryngeal contraction or swallowing impulse is perceived, the current value I_th is recorded; during formal use, the working intensity is set to 0.6 x I_th, ensuring effective activation while avoiding discomfort. The stimulation waveform is an asymmetric rectangular pulse with an anode phase width of 400 microseconds and a cathode phase width of 100 microseconds, with a net charge of zero to prevent tissue polarization. The frequency is locked at 25 Hz, the duty cycle is 20%, the pulse train lasts for 5 seconds, and the interval is 25 seconds, forming a periodic stimulation pattern. Numerous clinical studies have confirmed that this parameter combination has a significant effect on improving heart rate variability and reducing cortisol levels.
[0060] The three intervention modalities can be used alone or in combination. For example, for severe anxiety, the system can simultaneously start the respiratory guidance (auditory), neck heating (somatosensory), and vagus nerve electrical stimulation (neuroelectrophysiology), forming a multi-channel coordinated regulation network, which has a stronger physiological effect than any single mode.
[0061] The cross-period emotional evolution tracking engine is responsible for long-term mental health situation awareness and forward-looking risk warning. The engine automatically performs data aggregation tasks every morning, arranging the emotional recognition results (eight-dimensional probability distribution and two-dimensional coordinates) of every minute in the past 24 hours along the time axis, generating a 24 x 60 resolution emotional fluctuation heat map. The horizontal axis represents time (hours), the vertical axis represents emotional dimensions (eight discrete emotions superimposed on the arousal-valence continuous space), and the color depth represents the residence time or intensity integral of that emotional state.
[0062] On the basis of the heat map, the engine uses a hidden Markov model to infer the law of potential psychological state transition. The model sets four hidden states: calm, mild stress, high alert, and emotional breakdown. The observation symbol is the discretized emotional classification result (8 categories). Through the Baum-Welch algorithm, the state transition probability matrix and the emission probability matrix are learned. After training, the model can perform Viterbi decoding on the historical data of any time period to restore the most likely hidden state sequence.
[0063] Based on the learned state transition dynamics, the engine builds a 24-hour risk prediction model. Using survival analysis methods, the cumulative risk function of entering the "emotional breakdown" state from the current state is calculated. When the predicted risk exceeds 15%, the system deploys preventive intervention plans in advance. The plan includes adjusting the schedule to avoid high-load tasks, increasing positive social interaction reminders, and starting low-intensity background conditioning (such as playing soothing music). For organization users using the group mode, the engine aggregates the data of multiple members after desensitization processing (removing identity markers while retaining behavior patterns) to identify the emotional contagion path. For example, through Granger causal network analysis, it is found that the emotional fluctuations of a manager often lead team members by 2 hours, indicating that he is the source of emotional contagion, and targeted stress management support can be provided.
[0064] This embodiment builds an emotional recognition and intervention system with high robustness, strong adaptability, and closed-loop regulation capability through the deep coupling and collaborative operation of the above-mentioned subsystems. The multi-modal synchronous acquisition and dynamic weighted fusion mechanism fundamentally overcomes the limitations of single-source perception in complex environments, enabling the system to maintain an identification accuracy of over 90% in challenging scenarios such as changes in lighting and facial occlusion. The combination of personalized baseline modeling and cross-period tracking enables a leap in the ability to monitor instantaneous states and depict long-term psychological trajectories. The cognitive load analysis module reveals the deep cognitive roots of emotional fluctuations, providing a scientific basis for precise intervention. The three-channel wearable regulation terminal directly acts on the autonomic nervous system through physical signals, achieving rapid and effective physiological regulation. The reinforcement learning-driven strategy evolution mechanism ensures that the system continues to optimize as the use time increases. The overall scheme forms a complete "perception-decoding-decision-intervention-feedback" technical closed loop, not only improving the technical performance of emotional recognition, but also establishing an active maintenance of mental health engineering paradigm, providing a practical system-level solution for intelligent mental health services.
Claims
1. An emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals, characterized in that, include: The multimodal data synchronous acquisition subsystem is used to synchronously capture the user's facial micro-expression video stream and multi-channel physiological signals to obtain a spatiotemporally aligned raw time-series data stream; The dynamic feature fusion and emotion decoding subsystem is used to receive the original time-series data stream, extract the intensity time series of micro-expression action units and the frequency-segmented energy features of physiological signals, and input the two types of features into a dual-pathway time-series neural network containing visual and physiological pathways. The personalized emotion baseline modeling subsystem is used to establish a model of the frequency distribution of micro-expression activities and the combination of resting physiological parameters of an individual in daily life during the initial user phase, and to use this as a reference to calculate the degree of deviation from the baseline in subsequent detection as a preliminary criterion for abnormal emotional activation. The cognitive load and emotion fluctuation correlation analysis module is used to continuously monitor the changing trend of the energy ratio of theta waves to beta waves in the prefrontal cortex region of EEG signals. It constructs a cognitive resource consumption index by combining task performance data and pupil diameter fluctuation amplitude, and performs causal analysis on the cognitive resource consumption index and emotion decoding results to identify secondary emotions caused by cognitive overload. The adaptive feedback intervention generation subsystem is used to generate customized instruction sequences by calling a preset intervention strategy library based on emotion type, intensity level and trigger classification. A wearable multi-channel physiological regulation terminal is used to receive the customized instruction sequence and apply auditory, somatosensory and neuroelectrophysiological triple regulation signals to perform emotional intervention; A cross-period emotion evolution tracking engine is used to aggregate emotion recognition results on a daily basis and construct an individual emotion fluctuation heatmap; The cognitive load and emotion fluctuation correlation analysis module integrates task performance data as an external validation indicator of cognitive load. When a user is performing a specific digital task, it records the task completion delay rate and operation error rate. Eye-tracking units detect pupil diameter changes and calculate the standard deviation of pupil dilation amplitude during the task. This indicator is positively correlated with cognitive effort. The four indicators—θ / β ratio, task delay rate, error rate, and pupil fluctuation amplitude—are normalized and weighted, generating the Cognitive Resource Consumption Index (CRI). Among them, weight to Determined based on user historical data through linear regression fitting.
2. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 1, characterized in that, The multimodal data synchronous acquisition subsystem includes: a high frame rate infrared imaging unit for capturing subtle facial muscle movements at a high sampling frequency; a multi-channel biosignal sensing array for synchronously acquiring skin conductance response, heart rate variability, forehead temperature gradient, and EEG alpha band power spectral density; and a hardware-level timestamp synchronization mechanism to ensure that the facial micro-expression video stream and the multi-channel physiological signals are time-aligned with millisecond-level precision.
3. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 1, characterized in that, The dynamic feature fusion and emotion decoding subsystem comprises a visual pathway that processes local action unit dynamics and a physiological pathway that mines autonomic nervous system response patterns. The outputs of the two pathways are context-aligned and information-complementary through a cross-attention mechanism with learnable weights to output discrete emotion category probability distributions and arousal and valence estimates on continuous dimensions. It also includes a modal confidence assessment mechanism to adjust the input weights of the visual and physiological pathways in the fusion process in real time based on the current signal quality indicators. When the signal-to-noise ratio of the facial image is lower than a preset threshold, the contribution weight of the visual pathway is automatically reduced and the dominance of the physiological pathway is increased. The signal quality indicators include pupil visibility, facial contour integrity, skin conductance baseline drift, and EEG artifact ratio.
4. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 1, characterized in that, The training process of the dual-path temporal neural network adopts an adversarial domain adaptation strategy, performs joint optimization on a multi-center database containing samples of different ages, genders, and races, and introduces virtual occlusion enhancement technology during the training phase to improve the robustness of the model under real-world occlusion conditions.
5. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 1, characterized in that, The personalized emotion baseline modeling subsystem is also equipped with baseline update triggering conditions. When the system detects that the mean deviation of the user's basic physiological parameters exceeds the standard deviation of the initial baseline preset multiple for several consecutive days, a new round of baseline acquisition program is started to reconstruct the individual's micro-expression activity frequency distribution and resting physiological parameter combination model in daily life.
6. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 1, characterized in that, The strategy library update mechanism of the adaptive feedback intervention generation subsystem is based on a reinforcement learning framework. The system records the change in emotional state before and after each intervention as a reward signal, and combines it with the user's subjective rating of the intervention experience. The Q-learning algorithm is used to evaluate the long-term effectiveness of different intervention methods in various situations, so as to dynamically adjust the strategy priority ranking.
7. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 1, characterized in that, The wearable multi-channel physiological regulation terminal integrates a miniature bone conduction speaker, a distributed thermal stimulation patch, and a transcutaneous vagus nerve electrical stimulation electrode. The distributed thermal stimulation patch is distributed along both sides of the neck to activate the parasympathetic reflex pathway. The stimulation parameters of the transcutaneous vagus nerve electrical stimulation electrode follow an individualized setting protocol, and the initial stimulation intensity is determined by a vagus nerve threshold test procedure.
8. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 7, characterized in that, The vagus nerve threshold testing procedure is used to increase the current output at a step rate until the user reports a specific physiological sensation, and takes a preset percentage of the threshold as the starting working intensity. The stimulation waveform uses an asymmetric rectangular pulse, and the frequency and duty cycle are set to balance nerve activation efficiency and wearing comfort.
9. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 1, characterized in that, The cross-period emotion evolution tracking engine also has a group pattern analysis function, which allows data from multiple authorized users to be aggregated and modeled after anonymization in order to identify the emotion propagation path and group stress response patterns at the organizational level.
10. The emotion recognition and intervention system based on the fusion of facial micro-expressions and physiological signals according to claim 2, characterized in that, The high frame rate infrared imaging unit is equipped with a narrow band filter to suppress ambient visible light interference and features autofocus and image stabilization to ensure that it can effectively capture micro-expression textures in key facial areas even in low-light environments and under conditions of slight head movement.
Citation Information
Patent Citations
Emotion prediction and disease derivation method and system based on multi-modal fusion
CN121117918A
System for real-time analysis of emotional feedback during motivational presentations
DE202025104706U1
Cited By
An office worker emotion recognition method and system based on visible light and voice signals
CN122451571A