Emotion detection system based on facial recognition

Through the emotion detection system based on facial recognition, combined with time series analysis and multimodal dynamic perception modules, the weight fusion decision is dynamically adjusted, which solves the accuracy and real-time problems of emotion detection in existing technologies and realizes efficient emotion monitoring and personalized recommendations.

CN120673488APending Publication Date: 2025-09-19NORTHEAST FORESTRY UNIV
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510820823.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies lack temporal dynamic modeling and multimodal collaboration capabilities in emotion detection, and are unable to cope with complex interference in dynamic scenarios, resulting in low emotion recognition rates, low efficiency and strong subjectivity. They are unable to adapt to the differentiated needs of different scenarios, and the contradiction between computing resources and real-time performance is difficult to resolve.

Method used

An emotion detection system based on facial recognition is adopted. The feature vectors of micro-expression image sequences are extracted through the time series analysis module. The emotion classification probability and voice signal analysis are performed in combination with the multimodal dynamic perception module. The weight fusion decision is dynamically adjusted. Edge computing is used to optimize resources to achieve real-time and accurate emotion detection.

Benefits of technology

It significantly improves the accuracy and anti-interference ability of emotion detection, adapts to the needs of different scenarios, reduces computing load, improves real-time performance and resource utilization efficiency, reduces the misjudgment rate, and meets the real-time emotion monitoring needs of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673488A_ABST
    Figure CN120673488A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion detection system based on facial recognition, and relates to the technical field of computer vision and emotion calculation. A video stream time sequence analysis module is used for extracting a facial micro-expression image sequence of continuous frames, a time sequence feature vector containing a micro-expression intensity gradient, an illumination robustness coefficient and a facial action unit cooperation feature is generated, and a multi-mode dynamic sensing module is combined to carry out real-time analysis on an emotion classification probability, voice emotion parameters and physiological signals. And the fusion decision module performs dynamic weighted fusion on the multi-modal data based on the scene adaptive weight, and finally generates a comprehensive emotion score. Through multi-modal time sequence modeling and a dynamic weight optimization mechanism, the accuracy and environmental adaptability of emotion recognition are remarkably improved, and real-time perception and accurate decision making of customer emotion are realized in a target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and emotion computing, and specifically provides an emotion detection system based on facial recognition. Background Art

[0002] In target scenarios, emotion detection technology plays an important role in improving user experience and optimizing marketing strategies. However, existing technologies have significant flaws and are in urgent need of improvement:

[0003] Traditional methods such as manual observation, static expression library matching, and single-modal data analysis lack the ability to analyze temporal sequences and are unable to capture micro-expression features lasting less than 500ms. This results in low recognition rates, inefficiency, and high subjectivity for rapidly evolving emotions. In real-world scenarios like retail, lighting non-uniformity (luminance variations > 10 lux / s, sudden color temperature changes > 500K) can easily distort facial feature extraction. Existing technologies often misjudge emotions due to these environmental disturbances. Existing fusion strategies often use fixed weight allocations based on static experience, lacking the ability to dynamically adjust weights based on real-time emotional states or scene types. This results in delayed responses to gradual or sudden emotional fluctuations. The conflict between computing resources and real-time performance is another bottleneck. Traditional solutions rely on cloud-based transmission and processing of high-resolution video streams, but network latency cannot meet the real-time requirements of edge devices. This can easily lead to ineffective recommendation strategies and compromise conversion rates. Furthermore, low-computing edge devices struggle to support real-time facial feature extraction from high-resolution videos, resulting in reduced frame rates and missed key emotional nodes. High-frequency detection can also lead to computational overload when operating in low-power mode.

[0004] Further analysis reveals that existing technologies lack robustness for collaborative analysis of multimodal data and adaptation to dynamic environments. For example, in customer service scenarios, user movement can be misinterpreted as emotionally aggression. Existing systems, lacking motion compensation and noise filtering mechanisms, are unable to distinguish between physiological movements, environmental interference, and true emotional fluctuations. The lack of motion compensation mechanisms also makes it easy for existing systems to misinterpret head deflections as emotional aggression. Furthermore, fixed time windows are unable to adapt to the needs of diverse scenarios. For example, in retail scenarios, the average user dwell time is short (<30 seconds), and fixed long windows can easily filter out transient signals of interest. Customer service scenarios, on the other hand, may require more granular emotion tracking. Traditional solutions fail to consider the differentiated parameter requirements of target scenarios. These issues severely restrict the large-scale application of emotion detection technology in target scenarios. There is an urgent need for an emotion detection system based on facial recognition to improve the accuracy and adaptability of emotion recognition. Summary of the Invention

[0005] In view of the above shortcomings of the prior art, the present invention provides an emotion detection system based on facial recognition.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0007] Emotion detection system based on facial recognition, including:

[0008] A timing analysis module is used to receive real-time video streams, extract N consecutive frames of facial micro-expression image sequences, and generate a timing feature vector containing micro-expression intensity gradients, illumination robustness coefficients, and facial action unit collaborative features;

[0009] The multimodal dynamic perception module is used to perform the following processing:

[0010] According to the time series feature vector, the emotion classification probability within the sliding time window is calculated. When the variance of the probability of the positive emotion category exceeds the dynamic threshold V1 and the speech signal-to-noise ratio ≥ S1, the speech emotion compensation analysis is activated.

[0011] Calculating the cumulative amount of emotional fluctuations based on the difference in micro-expression intensity between adjacent frames, and triggering enhanced physiological signal detection when the cumulative amount of emotional fluctuations exceeds a dynamic threshold;

[0012] Extract the difference parameters of speech emotion pitch and speech speed between the current frame and the previous three frames to generate the speech emotion dynamic fluctuation coefficient, which is used to quantify the real-time changes of speech emotion;

[0013] The fusion decision module is used to perform the following operations:

[0014] When the emotion classification probability increases monotonically and the voice emotion fluctuation coefficient is lower than the preset stability threshold, the emotion classification weight is increased. W1;

[0015] When the accumulated amount of emotional fluctuation exceeds the dynamic threshold and the variance of the emotion classification probability is lower than the dynamic threshold V1, the physiological signal weight is increased. W2;

[0016] The emotion classification probability, the accumulated amount of emotion fluctuation and the physiological signal strength are weighted and fused according to the dynamic weight to generate a comprehensive emotion score.

[0017] Beneficial effects

[0018] Compared with the known public technology, the technical solution provided by the present invention has the following beneficial effects:

[0019] 1. This invention significantly improves the accuracy and anti-interference capabilities of emotion detection through multimodal temporal characterization and dynamic weight optimization. First, the system incorporates micro-expression intensity gradients, illumination robustness coefficients, and facial action unit collaborative features as temporal features. When the emotion consistency index of adjacent frames is less than 80%, a cross-frame emotion feature matching system is activated to dynamically correct the emotion label. By selecting stable illumination keyframes as reference templates (updated every 10 seconds), combined with SIFT feature point matching to correct the detection coordinates, this method effectively addresses feature distortion caused by sudden illumination changes or head rotation. This method addresses the problem of emotion misjudgment caused by image distortion or transient expression changes in traditional methods. For suspicious emotion fluctuations in the visual modality, a TCN is used to predict the emotion trend over the next 2 seconds. For the speech modality, a Kalman filter smoothing parameter is used to effectively distinguish between environmental noise and true emotion fluctuations. Compared to the high misjudgment rate of traditional single-modal systems, this solution reduces the misjudgment rate through multimodal fusion.

[0020] 2. The present invention uses dynamic weight allocation strategy to improve the weight W1, W2 is calculated in real time based on the duration t of the increasing probability of emotion classification and the cumulative fluctuation C, avoiding the lag of fixed weights and enabling scenario-adaptive fusion of multimodal data. For example, when the probability of positive emotion categories continues to rise but the emotional fluctuation of voice is low, the emotion classification weight is prioritized to capture user interest trends. Conversely, when emotional fluctuations are intense but voice features are not significant, physiological signal analysis is strengthened to identify sudden emotional changes. This dynamic logic enables the system to adapt to the differentiated needs of retail and customer service scenarios, improving the accuracy of emotion detection.

[0021] 3. The present invention optimizes the real-time resources of edge computing servers, significantly reducing the computational load while ensuring analysis accuracy. For example, when user interaction is interrupted for > 10 seconds, the video sampling rate is reduced from 30fps to 15fps, and non-critical voice threads are shut down, freeing up 42% of GPU resources. When the voice signal-to-noise ratio falls below 15dB, non-critical voice analysis threads are automatically shut down, retaining only the core visual modality detection task, and the response time is shortened from 500ms in the cloud-based solution to <200ms. Compared with traditional cloud-based solutions, this system has a shorter response time and lower memory usage, providing highly real-time, low-resource intelligent decision support for target scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0023] Figure 1 This is a system module connection diagram of the present invention;

[0024] Figure 2 This is a structural diagram of the multimodal dynamic perception module of the present invention;

[0025] Figure 3 This is a structural diagram of the fusion decision module of the present invention;

[0026] Figure 4 This is a deployment diagram of the edge computing server of the present invention. DETAILED DESCRIPTION

[0027] It is easy to understand that according to the technical solution of the present invention, without changing the essential spirit of the present invention, a person skilled in the art can propose a variety of interchangeable structural modes and implementation modes. Therefore, the following specific embodiments and drawings are only exemplary descriptions of the technical solution of the present invention and should not be regarded as the entire invention or as a limitation or restriction of the technical solution of the present invention.

[0028] Application Overview:

[0029] In existing technologies, emotion detection mostly relies on single-modal data analysis using static expression libraries, manual observation, or voice analysis. It lacks temporal dynamic modeling and multimodal collaboration capabilities, making it difficult to cope with complex interference in dynamic scenarios. Traditional methods are prone to emotional feature distortion when there are sudden changes in lighting or rapid user deflection due to the lack of quantitative analysis of micro-expression intensity gradients and lighting robustness coefficients. At the same time, the isolated processing of voice emotion compensation and physiological signal analysis makes it difficult to distinguish between environmental noise and real emotional fluctuations. Existing systems typically use fixed-weight fusion strategies, which are unable to flexibly adjust detection logic based on dynamic changes in emotional features, resulting in delayed responses to gradual interest fading or sudden emotional fluctuations. Furthermore, traditional solutions rely on cloud computing resources. In scenarios where the network of mobile terminals or edge devices is unstable, the real-time transmission of high-resolution video streams is prone to delays, making it impossible to meet the requirements of millisecond-level intelligent decision-making.

[0030] To solve the above problems, the inventors discovered that emotional states have multimodal dynamic correlation characteristics in the target scene, such as the coordinated changes in micro-expression intensity and voice emotion, and the temporal correlation between heart rate variability and user interaction behavior. The present invention establishes a coupling model of multimodal temporal coupling modeling and edge intelligent optimization, and based on this model, proposes a comprehensive solution for dynamic weight allocation and real-time resource optimization. The study found that when the probability variance of the positive emotion category exceeds the threshold, voice emotion compensation needs to be combined with Kalman filtering to eliminate instantaneous noise interference; and when the accumulated amount of emotional fluctuations surges, the increase in the frequency of physiological signal detection can effectively capture sudden emotions such as anxiety or excitement. Further experimental verification combines the sliding window dynamic adjustment mechanism with the cross-frame emotional feature matching system to form a closed-loop processing flow from data acquisition to intelligent decision-making.

[0031] Specifically, the system first uses the timing analysis module to collect video streams in real time and extract facial micro-expression image sequences, generating a timing feature vector containing micro-expression intensity gradients, illumination robustness coefficients, and facial action unit collaborative features. When the emotional consistency index of adjacent frames is lower than 80%, the cross-frame emotional feature matching system corrects the emotional label through HOG feature point matching, and the error is controlled within 5% of the original facial area. Figure 2 As shown, the multimodal dynamic perception module triggers speech emotion compensation or physiological signal enhancement detection based on the emotion classification probability variance and the accumulated emotion fluctuation in the sliding window. For example, in an environment with noise > 65dB, speech emotion compensation analysis is called first to suppress background interference; when the accumulated emotion fluctuation is detected to exceed the dynamic threshold, the heart rate variability detection frequency is increased from 1Hz to 3Hz and the detection is continued for 5 seconds; if the emotion fluctuation returns to normal within 5 seconds, the detection frequency is restored to 1Hz. The fusion decision module dynamically adjusts the weight according to the monotonicity of the emotion classification probability, the speech emotion fluctuation coefficient and the physiological signal strength. =0.3, =0.2, a comprehensive emotion score is generated. For example, when the probability of the positive emotion category continues to rise and the voice fluctuation is low, the emotion classification weight is increased to 0.5, driving personalized advertising push; when the emotion fluctuates violently but the voice characteristics are not significant, the physiological signal weight is increased to 0.4, triggering customer service intervention. The target response module executes hierarchical target decision instructions through N consecutive frame score exceeding judgment: Level 1 response: push customized coupons to the user's mobile terminal through the edge computing server, and display recommended products on the electronic screen; Level 2 response: initiate high-priority customer service intervention, and simultaneously upload user emotion data to the intelligent platform to optimize subsequent marketing strategies. Resource optimization strategy: When the system detects that the user interaction is interrupted, it suspends physiological signal detection to release computing resources; when the lighting is stable, the video stream sampling rate is reduced from 30fps to 15fps, and the non-critical voice analysis thread is closed to ensure the real-time detection of micro-expression enhancement.

[0032] Existing solutions rely on single-modal static analysis and lack the ability to adapt to dynamic environments. In specific scenarios with sudden changes in lighting or frequent user interactions, traditional methods have a high misjudgment rate because they do not introduce the lighting robustness coefficient γ and the cross-frame emotional feature matching system. This solution innovatively integrates time series feature quantization with multimodal data collaborative analysis, and significantly improves the robustness of emotion detection through dynamic weight allocation and closed-loop resource optimization mechanisms. Different from traditional fixed threshold models, this solution can intelligently adjust detection strategies based on real-time environmental interference and eliminate data noise. This method provides real-time, high-precision emotion monitoring for scenarios such as retail and customer service, and is particularly suitable for user experience optimization needs in shopping guides and online customer service scenarios with high customer traffic.

[0033] After introducing the basic concept of the present invention, embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0034] Example

[0035] The emotion detection system based on facial recognition in this embodiment is as follows: Figure 1 As shown, including:

[0036] The timing analysis module is used to receive real-time video streams, extract N consecutive frames of facial micro-expression image sequences (N≥5), and generate a timing feature vector containing micro-expression intensity gradients, illumination robustness coefficients, and facial action unit collaborative features.

[0037] The time series analysis module collects video streams in real time from cameras deployed in locations such as retail terminals and customer service centers. The system selects five consecutive frames from the video stream in chronological order, identifies and extracts facial micro-expression regions, and generates a dynamic image sequence. Using MediaPipe face tracking, the system locks onto the facial region in real time, ensuring that the ROI coordinate deviation between consecutive frames is less than 2 pixels to avoid feature misalignment caused by head movement. The system then captures the facial region of interest (ROI) to generate a five-frame micro-expression sequence. These facial region images are then arranged in sequence to form a sequence, capturing dynamic gamma changes in the face over time.

[0038] The time series feature vector also contains:

[0039] Micro-expression intensity gradient: quantifies the rate of change in the intensity of facial muscle movements between adjacent frames (0.2~1.0 scale), the formula is:

[0040]

[0041] Lighting robustness coefficient: Corrects lighting interference through the adaptive histogram equalization system. The formula is:

[0042]

[0043] Using MediaPipe's facial landmark detection, the system locates the core areas around the eyes and mouth corners, captures the facial focus areas, and generates a five-frame sequence of micro-expressions. The illumination intensity change rate (IVR) measures the dynamic variation in illumination intensity across the facial area between consecutive frames. In real-world scenarios, lighting conditions can vary significantly due to flickering spotlights, fluctuations in natural light, or shadows caused by user movement. By accurately calculating the IVR, the system can capture and quantify these disturbances in real time, providing a stable and reliable data foundation for subsequent sentiment analysis.

[0044] The facial visibility gradient reflects the dynamic trend of the occlusion of the user's facial area. In the target application scenario, the user's movement, the passing of other customers, or environmental objects may cause the facial area captured by the camera to be partially or completely occluded. The facial occlusion area is detected using a binary mask and the change rate of the occluded pixel ratio is calculated:

[0045]

[0046] in is the ratio of the occluded area to the facial area in the tth frame.

[0047] The facial region overlap between adjacent frames is a metric that measures the degree of overlap between facial regions in two adjacent frames. Under normal circumstances, facial regions between adjacent frames should have a high degree of overlap. However, when a user makes a sudden and drastic change in expression or when there are environmental disturbances, the continuity of emotional features between adjacent frames can be significantly reduced. When the cosine similarity of the temporal feature vectors of adjacent frames is less than 0.8, the dynamic emotion compensation system is activated. This system dynamically calibrates and predicts emotion values ​​by analyzing historical emotion data with local features such as pupil changes and mouth curvature in the current frame, taking into account the user's browsing behavior, such as product type and dwell time, to ensure the stability of the emotion quantification results. This mechanism enables the system to maintain high-precision emotion matching capabilities in complex scenarios, providing a reliable basis for personalized recommendations, improving target conversion rates, and improving user experience.

[0048] Phase modeling of periodic physical occlusions is a key technology for ensuring the continuity of sentiment analysis. In certain scenarios, users may experience brief facial occlusions due to habitual behaviors such as adjusting their glasses, looking down at their phone, or regular changes in ambient light, resulting in intermittent loss of emotional features. The system uses a sliding time window (window length = 2 minutes) to count the time intervals between occlusion events. The dominant period is estimated using a periodogram method. Periodic parameters such as the frequency, duration, and interval pattern of these occlusions are recorded and analyzed in real time. The system records the frequency, duration, and interval pattern of occlusions in real time and uses a phase prediction model to predict the timing and extent of the next occlusion. For example, if a user's face is partially obscured by looking down at their phone every 30 seconds, the system will proactively adjust its sentiment analysis strategy: When occlusion approaches, it will prioritize extracting key local features such as eye micro-expressions or forehead muscle dynamics for rapid emotion quantification, or interpolate and compensate based on historical emotion data to maintain the consistency of sentiment analysis. Furthermore, the system dynamically labels data from occluded periods and weights them in subsequent analysis to prevent outliers from influencing the overall sentiment trend. Through this mechanism, the system can maintain highly robust emotion detection capabilities in complex environments, thereby providing stable and reliable data support for the system to dynamically adjust advertising content, optimize product display strategies, or trigger real-time recommendation services for targeted customer service interactions, significantly improving user immersion and the accuracy of intelligent decision-making.

[0049] By generating time series feature vectors and conducting comprehensive analysis and utilization, the time series analysis module can more comprehensively and accurately analyze and process the continuous evolution of user emotions in the time dimension, providing a solid foundation for subsequent operations such as target recognition and behavior analysis, and helping to improve the performance and reliability of the entire facial recognition and recommendation system.

[0050] The cross-frame feature matching system performs the following operations:

[0051] Select unobstructed key frames in the video stream as reference templates;

[0052] SIFT feature point matching is used for low overlap frames to calculate the affine transformation matrix and then correct the detection coordinates; the corrected coordinate error satisfies:

[0053]

[0054] The corrected coordinate error is controlled within 5% of the original facial area, and the key frame reference template is updated every 10 seconds.

[0055] In the video stream, select an unobstructed keyframe as a reference template. A keyframe is typically a representative frame in the video, containing relatively complete and clear facial information. Using this keyframe as a benchmark, subsequent frames can be compared against it.

[0056] For frames with low facial overlap between adjacent frames, SIFT feature point matching is used. SIFT is a system for extracting feature points from images. These feature points are scale-invariant and rotation-invariant, enabling stable detection across different images. By matching SIFT feature points between low-overlap frames and a reference template, the correspondence between the low-overlap frames and the reference template is found. An affine transformation is a linear transformation between two-dimensional coordinates. It can describe image translation, rotation, scaling, and other transformations, maintaining image flatness and parallelism. Using the matched feature points, the affine transformation matrix of the low-overlap frame relative to the reference template is calculated. Correcting the detection coordinates involves using the calculated affine transformation matrix to modify the coordinates of the detection frame, allowing it to more accurately locate the target. Correcting the coordinates of the facial detection frame in low-overlap frames allows the detection frame to more accurately encompass the facial region.

[0057] After correcting the detection frame coordinates, the error in the corrected coordinates must be within a certain range. The specific criterion is that the difference between the corrected facial area and the original area must not exceed 5%, ensuring the accuracy and stability of the correction. To adapt to dynamic changes in video content, the keyframe reference template is updated every 10 seconds. This ensures that the reference template remains relevant to the current video content, improving the accuracy and adaptability of the system throughout the entire video stream processing process.

[0058] The γ channel brightness component of each frame in the video stream is corrected by the adaptive histogram equalization system. The correction intensity is calculated by the following formula:

[0059]

[0060] in, is the grayscale mean of the current frame. When the grayscale mean of the adjacent frames is detected, When the difference exceeds 50, segmented gamma correction is enabled. This system corrects the brightness of each frame in the video stream to improve the visual effect of the image and make the brightness distribution of the image more uniform.

[0061] Process the luminance component of the gamma channel of each frame in the video stream. In image processing, different color spaces have different channels. The gamma channel is usually related to brightness. Selecting this channel for correction can directly adjust the brightness of the image.

[0062] The adaptive histogram equalization system adjusts the histogram according to the local characteristics of the image, making the brightness distribution of the image more uniform, thereby improving the clarity and detail visibility of the image. When the difference exceeds 50, segmented gamma correction is enabled.

[0063] Compared to traditional technologies, existing emotion detection systems based on facial recognition often rely on static single-frame feature extraction or simple time-series averaging strategies, making them incapable of handling complex emotional interference in dynamic, specific scenarios. Traditional methods typically employ fixed-threshold emotion classification models, which fail to quantify the degree of emotional variation between adjacent frames. This leads to frequent misjudgments of emotional states when the user's expression changes rapidly or when lighting interferes. In terms of emotion continuity, relying solely on single-frame local feature analysis fails to capture the dynamic gradients of emotional evolution. This makes it particularly difficult to predict the discontinuity patterns of emotional features and optimize the detection logic in advance, especially in scenarios with periodic behavioral interference. Furthermore, traditional technologies often use linear smoothing systems to correct low-correlation emotion frames, resulting in significant error accumulation, which impacts the accuracy of subsequent emotion matching and recommendation decisions. This proposed solution constructs a time-series emotion vector that incorporates emotion mutation rate, interference dynamic gradient, and periodic phase parameters. Combining this with a dynamic emotion compensation system and a context-aware calibration mechanism, it effectively addresses data distortion issues in dynamic scenarios due to sudden expression changes, environmental interference, and low-correlation emotion sequences, providing robust support for precise recommendation and real-time decision-making.

[0064] Through the above technical solutions, the present application realizes highly robust analysis and multi-dimensional fusion detection of dynamic features of user emotions. The time series analysis module provides an interference-resistant time series data basis for the multimodal emotion perception module by capturing the emotion mutation rate, interference dynamic gradient, correlation of emotion features of adjacent frames and phase parameters of periodic interference in real time; the real-time monitoring of emotion correlation of adjacent frames and the dynamic emotion compensation system ensure the continuity of emotion quantification and avoid the offset of emotion values ​​caused by drastic fluctuations in expression or sudden changes in the environment. Context-aware lighting adaptation technology can more finely adjust the brightness and contrast of local areas of facial images to adapt to scenes with large differences in lighting between adjacent frames, avoiding the problem of misjudgment of emotional features due to brightness distortion. Through multi-dimensional data fusion and dynamic calibration mechanism, the system can maintain high accuracy of emotion detection in complex interactions, providing reliable underlying technical support for real-time personalized recommendations and user behavior decision analysis.

[0065] Based on the temporal emotion feature vector, the Emotional Stability Index (ESI) fluctuation trend within the sliding time window is calculated. When the ESI variance of the emotion mutation frequency exceeds the dynamic threshold V1, context-aware emotion compensation analysis is activated. This mechanism dynamically corrects emotion quantification deviations caused by transient expression anomalies or environmental interference by fusing user behavior data with local facial features, ensuring that the decision-making basis of the intelligent recommendation system is highly robust and adaptable to different scenarios. The ESI variance is calculated as follows:

[0066]

[0067] in for Always integrate sentiment scores, is the window length (seconds).

[0068] The time size of the sliding time window is dynamically adjusted according to the target interaction scenario type:

[0069] Scenario types include retail scenarios and customer service scenarios; the sliding time window for retail scenarios is set to 30 seconds, and the sliding time window for customer service scenarios is set to 10 seconds; when the frequency of emotional mutation is detected exceeding 5 times / second or the ambient light fluctuation intensity exceeds 50 lux / second, the time size of the sliding time window is shortened to 50% of the original value.

[0070] ESI change trend analysis includes:

[0071] (1) Detect the degree of deviation between the moving average of the emotion value and the instantaneous value. When the degree of deviation exceeds 25%, it is marked as a suspicious emotion interference segment;

[0072] (2) For the marked suspicious emotional interference clips, the long short-term memory (TCN) network is enabled to predict the ESI change curve for the next 2 seconds, and the prediction model is optimized in combination with the user behavior context.

[0073] The ESI value is the standard deviation of the emotional quantification value per unit time (0% to 100% scale), and is a core indicator for measuring the intensity and stability of user emotional fluctuations. During the interaction process, user emotions may fluctuate significantly due to the attractiveness of the content, interactive feedback, or external interference. The higher the ESI value, the more unstable the emotional state and the lower the reliability of the intelligent decision-making; conversely, it indicates that the emotional state is stable and the confidence of the recommendation system is higher. It is generally believed that: ESI value ≤ 15%: user emotions are stable, suitable for pushing high-value products or targeted advertisements; 15% < ESI value ≤ 30%: emotions fluctuate slightly, and the recommendation strategy needs to be adjusted based on real-time behavioral data; ESI value > 30%: emotions fluctuate severely, and it is necessary to trigger a dynamic compensation system or proactive customer service intervention to stabilize the interactive experience.

[0074] The system calculates the ESI value trend within a sliding time window. The time size of the sliding time window is dynamically adjusted according to the target scenario type. In this embodiment, the parameter settings for different scenarios are as follows:

[0075] Target scenario Sliding time window (seconds) Window time (seconds) when the frequency of emotional changes is greater than 5 times / second or the light fluctuation is greater than 50 lux / second Retail scenarios 30 15 Customer Service Scenario 10 5

[0076] During the calculation process, the system detects the degree of deviation between the moving average of the emotion value and the instantaneous value. If the deviation exceeds 25%, it is flagged as a suspected emotional interference segment. For these suspected emotional interference segments, the system uses a long short-term memory (TCN) network to predict the ESI change curve for the next two seconds. This model is optimized based on the user's behavioral context to more accurately determine the user's emotional evolution trend. Furthermore, ambient light fluctuations and the frequency of sudden changes in facial expressions also influence the adjustment of the sliding time window. When the light intensity change rate exceeds 50 lux / second or the frequency of sudden changes in emotion exceeds 5 times / second, the sliding time window is shortened to 50% of its original value, enhancing the system's response to dynamic interference and ensuring the real-time and reliable emotional quantification results. Through this mechanism, the system can quickly identify and correct abnormal emotional data in complex interactive scenarios, providing a high-confidence analytical basis for personalized service decisions.

[0077] In the target scene interaction test, the setting of dynamic threshold V1 is based on the following data:

[0078] Retail scenario: The mean ESI variance in the unstable mood state is 0.15±0.04, and in the stable mood state is 0.06±0.02. V1=0.12 is used to distinguish stable mood from interference / real fluctuations. The recall rate of real mood anomalies reaches 95% after verification by the confusion matrix.

[0079] Customer service scenario: Due to frequent content switching, the ESI variance of normal interaction can reach 0.22±0.08, and V1=0.25 can filter out 80% of environmental interference.

[0080] The setting of the dynamic threshold V1 is not fixed, but is dynamically optimized through the following mechanisms:

[0081] Sliding window association: When the window is shortened to 50% of its original value, the sample size per unit time decreases. To avoid deviation in variance calculation due to sample sparsity, the threshold V1 is simultaneously reduced to 70% of its original value to avoid variance calculation deviation caused by window reduction.

[0082] Environmental interference compensation: If the light change rate is greater than 50 lux / second, the dynamic threshold V1 is temporarily increased to 1.3 times to offset the disturbance of environmental interference on the emotion quantification system.

[0083] When the emotion classification probability exceeds the dynamic threshold V1, it indicates that there is a significant abnormality in the emotional state, but it is necessary to distinguish whether it is a real emotional fluctuation or a detection error caused by user behavior interference. For example, in a customer service scenario, the user's frequent switching of attention may cause the temporary loss of facial features, causing the ESI value to fluctuate violently. At this time, the setting of the dynamic threshold V1 must meet the following principles: Scenario differentiation: The dynamic threshold V1 of the retail scenario should be lower than that of the customer service scenario, because the shopping scenario requires more sensitive capture of user emotions to optimize the recommendation strategy; Dynamic anti-interference: When the light change rate is >50 lux / second or the frequency of expression mutation is >5 times / second, the dynamic threshold V1 is temporarily increased by 20%~30% to reduce false triggers caused by environmental or behavioral interference, while ensuring the accurate capture of key emotional signals.

[0084] In the multimodal dynamic perception module, an exponentially weighted moving average (EWMA) is used to calculate the baseline trend of sentiment values ​​and dynamically compare it with the instantaneous sentiment values. When the instantaneous value deviates from the moving average by more than 20%, for example, when a user's sentiment value suddenly rises from a stable range to a fluctuating range in a retail scenario, the system marks the period as a "suspicious emotional interference segment." For such segments, a long short-term memory (TCN) network is used to predict the ESI trend over the next two seconds. The TCN network converts behavioral features such as the difference between the current frame and the average ambient light intensity over the previous 30 seconds and the normalized value of the user's click frequency into a time series vector. This vector is then concatenated with the sentiment sequence (dimension T×1) to form a T×3 input tensor. This is then passed through a bidirectional TCN layer to capture spatiotemporal dependencies. The bidirectional recurrent structure captures the temporal dependencies of sentiment evolution. For example, if the predicted ESI value continuously rises over the next three seconds and the slope exceeds a dynamic threshold, increasing by >5% per second, the sentiment credibility is considered to have decreased, triggering dynamic compensation logic to interpolate historical sentiment baselines or integrate behavioral data to re-quantize the sentiment value.

[0085] Compared to traditional technologies, existing emotion detection systems often rely on static thresholds based on single-frame emotion values, failing to capture the dynamic evolution of emotion fluctuations. Traditional methods analyze emotion averages using fixed time windows. In customer service scenarios, due to frequent content switching and rapid changes in user expressions, this can easily misclassify normal emotional responses as abnormal. Addressing sudden changes in lighting or behavioral disturbances relies solely on simple smoothing, resulting in detection failures in strong light conditions or when users briefly obscure their faces. Furthermore, traditional solutions lack modeling of the correlation between user behavior and emotional characteristics. For example, a fluctuation in emotion caused by a user legitimately looking down at their phone can be misclassified as loss of interest. This solution, through a dynamic sliding window adjustment mechanism, a TCN time series prediction model, and a scene-adaptive threshold optimization system, addresses the core issues of traditional technologies, such as sensitivity to environmental interference, high false alarm rates, and poor cross-scenario adaptability. Specifically, dynamic window adjustment optimizes the window duration in real time based on the target scenario type, balancing detection sensitivity and interference immunity. TCN time series prediction integrates emotion sequences, behavioral context, and environmental data to accurately predict emotion evolution trends. Adaptive threshold optimization dynamically adjusts the threshold based on scene characteristics to avoid misclassifications caused by environmental interference or transient behavior. Through the above innovations, the system can accurately distinguish between true emotional intentions and instantaneous interference signals in complex intelligent interactions, provide highly reliable decision-making basis for real-time recommendation systems, and significantly improve user experience and target conversion efficiency.

[0086] Through the above technical solutions, this application achieves high-precision dynamic tracking of user emotional states and scenario-adaptive decision-making. The temporal emotion evolution analysis module quantifies the emotion mutation rate and the dynamic gradient of interference, providing an interference-resistant temporal data foundation for the multimodal dynamic perception module. The dynamic adjustment mechanism of the sliding window works in synergy with the TCN temporal prediction model to accurately identify abnormal fluctuation trends in emotion values ​​and distinguish between instantaneous interference and true emotion evolution.

[0087] The multimodal dynamic perception module further verifies the credibility of emotional signatures through contextual behavioral compensation analysis and micro-expression intensity gradient fluctuation detection, preventing misjudgment of single signals. The fusion decision module, based on a dynamic weighting strategy, integrates the Emotional Stability Index (ESI) trend, behavioral relevance, and micro-expression fluctuation coefficient into a comprehensive emotional confidence score, driving personalized recommendation strategies for target application modules.

[0088] Under the coordination of a cloud-edge collaborative computing architecture, each module ensures stable operation and millisecond-level responsiveness in complex scenarios through real-time resource optimization, providing technical support for precision marketing and user experience optimization. For example, if a user's emotional index (ESI) is detected to be continuously rising (ESI ≤ 10%) while browsing a product, the system can trigger real-time limited-time discount push or related product recommendations, significantly improving conversion rates. If the emotional index fluctuates abnormally (ESI > 30%), a dynamic compensation system is activated or customer service intervention is called to prevent user churn.

[0089] Through full-link multimodal fusion and scenario-adaptive mechanisms, this solution achieves closed-loop optimization of emotion detection and intelligent decision-making in scenarios such as retail, advertising, and customer service, and builds a highly robust emotional computing engine for the target scenario ecosystem.

[0090] The multimodal dynamic perception module is also used to perform the following processing: based on the Euler angle difference of the head posture in adjacent frames, the cumulative deviation of the user behavior is calculated. When the cumulative amount exceeds the preset angle threshold, the micro-expression enhancement detection mechanism is triggered, focusing on extracting eye dynamics and changes in the curvature of the mouth corners to eliminate the impact of posture interference on emotion quantification.

[0091] The multimodal dynamic perception module is also used to perform the following processing: extracting the difference in facial action unit intensity between the current frame and the previous three frames, generating AU dynamic fluctuation coefficients and physiological signal intensity to capture emotional fluctuation characteristics.

[0092] The AU intensity difference is normalized to eliminate the instantaneous facial muscle jitter interference caused by cargo handling. The calculation formula is:

[0093]

[0094] in: Indicates the AU intensity value of the current frame, is the AU intensity value of the previous frame;

[0095] and The maximum and minimum values ​​of AU intensity in the historical data of the current operation scene;

[0096] ε is a very small constant to prevent the denominator from being zero, and is set to 1×10⁻ 6 .

[0097] Extracting the difference in facial action unit intensity between the current frame and the three previous frames is a key step in generating the AU dynamic fluctuation coefficient. The system analyzes and calculates these intensity differences to generate the AU dynamic fluctuation coefficient. This coefficient quantifies the intensity of micro-expression changes within a 0.5-second time window (4 frames = 133ms at 30fps) on a 0-1 scale, enabling sensitive identification of abnormal facial fluctuations caused by anger or excitement.

[0098] When the accumulated emotional fluctuations C>0.5, the system triggers the physiological signal enhancement detection mechanism, increasing the heart rate variability detection frequency from the basic 1Hz to 3Hz.

[0099]

[0100] in: Indicates the heart rate variability value at the current moment. is the heart rate variability value of the previous frame. This mechanism can capture the stress response of the autonomic nervous system within 0.33 seconds and effectively identify sudden emotional changes.

[0101] For example, assuming the AU intensity value of the current frame is 0.8, and the AU intensity values ​​of the previous three frames are 0.6, 0.7, and 0.75 respectively, the difference between the adjacent frames is:

[0102]

[0103]

[0104]

[0105] Normalize the data:

[0106]

[0107]

[0108]

[0109] Generate dynamic volatility coefficient:

[0110]

[0111] Triggering HRV enhanced detection, emotional fluctuation accumulation C=0.6 (>0.5),

[0112] The heart rate variability detection frequency has been increased from 1Hz to 3Hz, capturing physiological signals in real time:

[0113] Time point HRV value (ms) Interval time t-1 85 — t 120 0.33 seconds

[0114] Calculate physiological intensity signal:

[0115]

[0116] history

[0117] After detecting the coordinated signals of the user's facial action units continuously rising and the heart rate variability mutation, the system starts the multimodal fusion decision mechanism. By analyzing the coupling characteristics of the continuous rising trend of AU and the HRV mutation, the system confirms that this pattern conforms to the typical physiological representation of anger; at the same time, the cumulative amount of emotional fluctuation C=0.6 exceeds the threshold of 0.5, effectively eliminating the possibility of environmental interference. As described in claim 8, the system dynamically increases the weight of physiological signals to =0.4, generating a comprehensive sentiment score S=0.72, and immediately triggering the secondary response mechanism: notifying high-priority customer service to intervene, initiating high-priority customer service intervention and reporting to the target platform, forming a complete closed-loop decision chain.

[0118] Compared to traditional emotion detection methods, which often rely on manual labeling or static threshold determination to analyze user emotional characteristics, simply judging a user's emotional state based on fixed rules is highly subjective and susceptible to the annotator's experience bias or the limitations of a single scenario. Furthermore, traditional technologies lack the ability to capture instantaneous micro-expressions, making it difficult to quantify the dynamic evolution of emotions. They also cannot effectively distinguish between genuine emotional expressions and facial feature anomalies caused by external interference, among other technical issues.

[0119] Through the above technical solution, this application effectively eliminates the transient jitter interference of facial micro-expressions caused by external interference such as sudden changes in ambient lighting and habitual user movements. The multimodal dynamic perception module accurately extracts the intensity difference of micro-expression action units and generates a dynamic fluctuation coefficient. A scene-adaptive normalization system then unifies the data to a comparable standard scale (0,1), making the subsequent quantitative analysis of the user's emotional state more accurate and reliable.

[0120] Head posture compensation analysis includes:

[0121] When the variance of the emotional stability index (ESI) exceeds the dynamic threshold V1, the system ignores the data generated by reasonable user behavior in the cumulative amount of posture deviation and only retains the abnormal deviation;

[0122] The Euler angle difference is smoothed by the Kalman filter to eliminate the instantaneous jitter noise caused by environmental interference.

[0123] Micro-expression enhanced detection uses the following processing:

[0124] When the accumulated amount of posture deviation exceeds the angle threshold, the AU detection frequency is increased from the first preset value to the second preset value.

[0125] In the interactive scenarios of this embodiment, dynamic analysis of changes in the user's head posture and facial micro-expressions is a core challenge in emotion detection. Because habitual user movements or environmental interference can cause reasonable deviations in head posture, and the boundary between true emotional fluctuations and behavioral interference is blurred, traditional methods struggle to effectively distinguish. To address this, this solution utilizes the collaborative processing of multimodal dynamic perception modules, combined with head posture compensation and enhanced micro-expression detection mechanisms, to achieve precise extraction of emotional features and interference filtering.

[0126] The system calculates the cumulative posture deviation by using the Euler angle differences between adjacent frames. Specifically, the system accumulates the Euler angle differences for N consecutive frames and determines an abnormal posture deviation when the cumulative difference exceeds a preset angle threshold. Posture deviation is the normalized cumulative Euler angle deviation (scaled from 0 to 1, with 1 indicating the maximum deviation).

[0127] Euler angles are parameters used to describe the angle of rotation of an object in three-dimensional space. Angle thresholds are dynamically set based on the operating scenario, such as a 30° angle threshold for retail scenarios and a 45° angle threshold for customer service scenarios. The thresholds are determined by the distribution of posture deviations between fatigued and non-fatigued states in the measured data. The thresholds are determined by the distribution of head postures in stable and abnormally fluctuating emotional states in the measured data. For example, in a retail scenario, if a user frowns frequently due to price sensitivity, and the system detects that the cumulative pitch angle difference for five consecutive frames reaches 35° (exceeding the angle threshold of 30°), micro-expression enhancement detection is triggered to capture abnormal intensity increases or decreases in emotion-related facial action units.

[0128] When the variance of the emotional stability index (ESI) exceeds the dynamic threshold V1, the system ignores the data generated by reasonable user behavior in the cumulative amount of posture deviation and only retains the abnormal deviation;

[0129] At the same time, the Euler angle difference of the head posture is smoothed by the Kalman filter to eliminate the jitter noise caused by environmental interference or instantaneous movement.

[0130] The Kalman filter is a recursive algorithm that eliminates instantaneous jitter caused by sensor noise through a weighted fusion of predictions and measurements. Smoothing refers to the fact that due to factors such as measurement error and sensor noise, the Euler angle difference may experience instantaneous jitter, which can affect the subsequent accurate judgment of the object's posture. The Kalman filter processes the Euler angle difference and corrects the current measurement based on its dynamic characteristics and noise statistics, making the output Euler angle difference smoother and reducing the impact of instantaneous jitter. Eliminating instantaneous jitter interference means that the Kalman filter's prediction and update process can effectively suppress noise and interference, extracting true Euler angle change information from noisy data. This resulting Euler angle difference more accurately reflects the actual changes in the object's posture, improving system stability and reliability.

[0131] When the cumulative deviation of user behavior exceeds the preset angle threshold, the system increases the detection frequency of micro-expression action units from the base frequency to the enhanced frequency to capture transient emotional fluctuations. AUs are defined based on the Facial Movement Coding System (FACS) and are used to quantify user expression characteristics. The detection frequency switching logic is as follows: the base frequency is used for routine monitoring, and the enhanced frequency is dynamically triggered based on the angle threshold. For example, in a product browsing scenario, if a user frequently looks down at their phone due to price sensitivity, the system will increase the AU detection frequency from 1Hz to 5Hz, capturing the instantaneous fluctuations of AU4 frowning and AU12 raising the corners of the mouth at a high frequency, thereby identifying emotional characteristics related to decision hesitation or sudden increase in interest.

[0132] Compared with traditional technologies, existing emotion detection systems mostly rely on static thresholds or single-modal data analysis, making it difficult to cope with complex interference in dynamic scenarios. Traditional methods usually use a fixed time window to analyze the emotion mean, which cannot distinguish between head posture deviations caused by reasonable user behavior and real emotional fluctuations, resulting in a high misjudgment rate; the detection of micro-expressions mostly relies on fixed-frequency sampling, which cannot capture transient facial changes when user behavior deviates drastically, especially in high-paced scenarios such as customer service. It is easy to miss key emotional signals. In addition, traditional solutions lack an active suppression mechanism for environmental noise. The instantaneous jitter in the Euler angle difference will interfere with the calculation of posture deviation, further reducing the credibility of emotion analysis. This solution achieves high-precision emotion analysis and low-interference robustness in complex interactions, providing a reliable technology engine for smart retail, digital marketing and other fields.

[0133] Through the above technical solutions, this application realizes multi-dimensional collaborative perception and dynamic optimization decision-making of the user's emotional state. The time series emotion evolution analysis module provides an interference-resistant time series data base for the multimodal dynamic perception module by quantifying the emotion mutation rate and interference dynamic gradient in real time; the head posture compensation analysis eliminates reasonable behavioral noise through Kalman filtering, retains the posture deviation characteristics driven by real emotions, and dynamically associates with the variance of the emotional stability index to achieve accurate extraction of emotional signals. The micro-expression enhancement detection module captures AU intensity fluctuations through dynamic frequency boosting (1Hz→5Hz), enhances the transient expression analysis capability when the behavioral deviation exceeds the threshold, effectively identifies key emotional features such as sudden interest increase and decision-making hesitation, and significantly improves the detection robustness and real-time performance in complex target scenarios. Figure 3 As shown, the fusion decision module is used to perform the following operations:

[0134] When the ESI value trend is monotonically increasing and the AU dynamic fluctuation coefficient is lower than the preset stability threshold, the weight of the ESI value is increased. ;

[0135] When the cumulative amount of posture deviation exceeds the angle threshold and the ESI variance is lower than the dynamic threshold V1, the AU intensity weight is increased. ;

[0136] The comprehensive sentiment confidence score is generated by dynamic weighted summation:

[0137]

[0138] Satisfy the weight normalization constraint:

[0139]

[0140] in, W1 and The calculation method of the dynamic weight increase of W2 satisfies:

[0141] (When ESI increases monotonically and < Stability Threshold)

[0142]

[0143] (When the deviation > and )

[0144]

[0145] in: =0.2, =0.3 is the preset basic weight coefficient;

[0146] The duration of the monotonically increasing ESI value, in seconds;

[0147] The maximum allowed monotonous duration of the scene;

[0148] and is a scene adaptive parameter.

[0149] In this embodiment, and Meet: Retail scenario: =0.5, =1.2; Customer service scenario: =0.8, =0.9.

[0150] Dynamic weight increase W1, W2 satisfies the weight normalization constraint: ESI weight ( ) + behavioral deviation weight ( ) + AU intensity weight ( ) = 1.

[0151] When the trend of the emotional stability index is monotonically increasing and the AU dynamic fluctuation coefficient is less than the preset stability threshold, the system will improve. Weight. This mechanism is used to capture the potential decision-making anxiety state of users whose emotional stability gradually declines but whose micro-expressions are relatively stable. In logistics operation scenarios, by analyzing a series of facial area images of N consecutive frames (N≥5), the changing trend of the ESI value within the sliding time window is calculated. When the ESI value shows a monotonically increasing trend, it means that the intensity of the user's emotional fluctuations is gradually increasing, which is an important indicator of the accumulation of decision-making pressure or the decline of interest. For example, in a retail scenario, if the ESI value continues to rise when a user browses high-priced goods, it indicates that their emotional stability has decreased due to price sensitivity or functional concerns, and their decision-making hesitation has deepened, which may be accompanied by longer page dwell time but fewer clicks.

[0152] The AU dynamic fluctuation coefficient is generated by extracting the difference in facial action unit intensity between the current frame and the three previous frames. The dynamic stability threshold should be set based on the range of normal facial expression fluctuations in different logistics scenarios. When the AU dynamic fluctuation coefficient is below the dynamic stability threshold, it indicates that the user's facial expression is relatively stable, with no significant changes in facial features due to external interference or non-emotional movements.

[0153] In practical applications, the preset stability threshold needs to be calibrated based on extensive user behavior data and the characteristics of different target interaction scenarios. For example, in retail scenarios, the preset stability threshold can be set to 0.3; in customer service scenarios, the preset stability threshold can be set to 0.2. This setting can more accurately trigger the weighting of ESI values ​​in different scenarios, thereby improving the accuracy and reliability of facial recognition emotion detection and the target application system.

[0154] Trigger AU weight The conditions for W2 improvement are: the cumulative deviation of user behavior > the angle threshold and the variance of the emotional stability index (ESI) < the dynamic threshold V1. This mechanism is used to capture scenarios where users experience significant behavioral deviations but subtle emotional fluctuations, thereby enhancing the ability to analyze transient micro-expression signals. In this case, the dynamic threshold V1, serving as the "emotional fluctuation credibility" threshold, must be higher than the mean ESI variance under normal operation to avoid misjudgments due to minor emotional fluctuations. When user behavior deviates significantly, AU intensity analysis is prioritized over the mean ESI to capture potential transient signals of interest.

[0155] When the dynamic threshold V1 and the angle threshold cover complex actions, a joint constraint model of the angle threshold and V1 is established. When the angle threshold is > 45°, the dynamic threshold V1 condition is forced to be ignored and the AU weight is directly triggered. W2 is increased; when the dynamic threshold V1 exceeds 1.5 times the historical extreme value, emotion detection is suspended and a system alarm is triggered to ensure data reliability.

[0156] The fusion decision module will dynamically adjust the weights of ESI value, posture deviation and AU intensity according to different conditions. When the change trend of ESI value is monotonically increasing and the AU dynamic fluctuation coefficient is lower than the preset stability threshold, the system will increase the weight of ESI value. When the cumulative amount of attitude deviation exceeds the angle threshold and the ESI variance is lower than the dynamic threshold V1, the AU intensity weight will be increased. W2.

[0157] In this embodiment, the parameter settings for different operation scenarios are as follows:

[0158] Scenario Type α β <![CDATA[w1]]> <![CDATA[w2]]> Retail scenarios 0.5 1.2 0.2 0.3 Customer Service Scenario 0.8 0.9 0.2 0.3

[0159] The comprehensive emotional confidence score is to dynamically adjust the weight of each indicator according to the real-time scene characteristics, so as to accurately quantify the user's emotional state. W1, After W2 is dynamically adjusted, the weights are redistributed and normalized proportionally to ensure that the sum always equals 1. This weight adjustment mechanism flexibly allocates attention based on real-time data, prioritizing key emotional signals. Through normalization and threshold constraints, standardized constraints, and interference filtering, recognition errors caused by sudden changes in lighting or natural head movements are effectively eliminated. Differentiated parameter configurations for retail and customer service scenarios enable the recognition model to precisely adapt to the application needs of different scenarios. The system specifically enhances sensitivity to customer satisfaction indicators such as pupil dilation frequency and zygomatic muscle activation, improving the accuracy of consumer sentiment prediction through multi-dimensional feature fusion.

[0160] Compared with traditional technologies, traditional methods rely on static expression thresholds for emotion judgment. For example, a pleasure judgment is triggered only by the upward angle of the mouth reaching a fixed value. In retail scenarios, this is prone to misidentification due to customers' natural posture adjustments. The discretization of micro-expression features and the sampling frequency detection of AU12 activation values ​​make it difficult to capture the instantaneous characteristics of rapid emotional transitions. This is especially true in customer service scenarios, where a user's sudden turn to answer a call may interrupt the continuity of the true emotional signal. In addition, traditional solutions lack the ability to perceive specific scenarios and cannot dynamically adjust emotion analysis strategies based on business types, resulting in limited cross-domain generalization capabilities. Through dynamic emotion weight optimization, adaptive matching of target scenario parameters, and a multi-source data fusion mechanism, this system integrates facial expressions, voice intonation, and environmental context, overcoming the key bottlenecks of traditional technologies: one-sided feature capture, high sensitivity to environmental interference, and weak target scenario transferability.

[0161] Through the above technical solutions, this system achieves multi-dimensional dynamic analysis and precise quantitative modeling of users' emotional states in target scenarios. The dynamic illumination compensation module provides a robust data foundation for multimodal emotion perception by analyzing ambient light intensity fluctuations and facial occlusion attenuation coefficients in real time. The fusion decision engine utilizes a dynamic emotion weight optimization mechanism to strengthen the weight of smile features when the intensity of joyful emotions continues to rise, and prioritizes the temporal analysis of micro-expressions during rapid expression transitions, achieving targeted enhancement of core target emotional signals. The comprehensive emotion score integrates facial action units, voice emotional fundamental frequency, and environmental semantic cues through a dynamic feature fusion framework to provide a three-dimensional measure of user emotional value. The adaptive enhancement of joyful weighting enhances the ability to capture relevant features in consumer decisions, while the contextual adjustment of anger weighting accurately identifies negative emotional inflection points during high-stress interactions. The emotional stability coefficient serves as an auxiliary dimension to balance the overall assessment system. The comprehensive emotion index quantifies customer satisfaction and consumption propensity, providing data support for intelligent decision-making such as precision marketing and service optimization. Through multi-source feature coupling analysis, the system can penetrate and identify the true emotional motivations behind users' superficial expressions, significantly improving the system practicality and target conversion efficiency in customer portrait construction, demand forecasting, and experience management scenarios.

[0162] The edge response module automatically activates the hierarchical marketing strategy when the comprehensive sentiment index exceeds the preset target value threshold for N consecutive time units.

[0163] The graded interactive responses include:

[0164] Level 1 response: triggers real-time discount pop-up windows and voice recommendations on smart terminal screens;

[0165] Secondary response: Write customer value tags through the CRM system API and synchronize them to the central data lake to build dynamic user portraits.

[0166] The hierarchical response setting is based on the quantitative analysis of the persistence of scenario-based emotions, minimizing false triggering while ensuring the capture rate of key signals, meeting the real-time requirements of the target scenario. The hierarchical corresponding judgment thresholds have been optimized through multi-scenario experiments:

[0167] The retail scenario selects Level 1 response based on the median user dwell time of 28 seconds and 90% of the effective emotional fragment duration is distributed between 3 and 5 frames;

[0168] The customer service scenario selects the second-level response based on the 92% probability that high-conflict emotion HRV>100ms+sudden increase in voice fundamental frequency lasts for ≥5 frames during the conversation, ensuring that no alarms are missed.

[0169] The CRM (Customer Relationship Management) system is the core hub of an enterprise's customer relationship management, deeply integrated with smart devices through standardized interfaces. The marketing middleware API, serving as the core interaction channel, is typically deployed within an enterprise's private cloud architecture, supporting data integration with terminal devices such as smart cameras and interactive displays. This interface connects to the emotion computing engine, acquiring real-time psychological characteristic data such as user micro-expression intensity gradients and emotion fluctuation trajectories. Through the system cluster, it generates insight reports such as consumption propensity predictions and service satisfaction assessments. Operations teams can precisely formulate personalized marketing strategies based on multidimensional emotion data, while marketing departments can optimize store traffic flow layouts using emotion heat maps. Furthermore, the CRM interface serves as an intelligent recommendation system, providing real-time emotional fuel and driving the value loop from emotion to behavior to conversion. In terms of customer loyalty management, the system integrates emotion decay curve modeling with a repurchase probability prediction system to achieve a strategic upgrade from immediate emotion perception to long-term customer value management.

[0170] When the comprehensive sentiment index exceeds the preset value threshold for N consecutive time units, a tiered marketing strategy is automatically activated. This tiered response strategy includes primary emotional guidance and secondary value conversion. When the primary response is triggered, the smart interactive terminal's real-time sentiment adaptation strategy is triggered, dynamic discount packages are pushed through the digital screen, and the AI ​​virtual shopping guide's empathetic sales pitch is activated. When the secondary response is triggered, the marketing middleware API writes real-time sentiment value tags to the customer data platform (CDP), simultaneously activating the central intelligent engine's consumer behavior prediction model to generate precise marketing instructions.

[0171] Compared to traditional technologies, existing emotion recognition systems often use a single, static criterion to trigger marketing actions, making it difficult to implement precise interventions based on the dynamic evolution of emotion intensity. Traditional approaches often set fixed response thresholds, which can lead to missed marketing opportunities due to slow response during the nascent stages of emotion, or inadequate strategies to guide consumer decisions during periods of intense emotional volatility. Interactions are limited to basic pop-up ads or SMS push notifications, lacking intelligent integration with customer relationship management (CRM) systems, preventing the triggering of personalized benefits at critical moments. Furthermore, traditional solutions lack deep integration with customer data platforms (CDPs), making it difficult to embed emotion insights into user profiles in real time, leading to a disconnect between marketing actions and real-world needs. This system, by integrating an emotion intensity gradient response mechanism with a marketing middleware API, overcomes the industry pain points of traditional technologies, including delayed opportunity capture, inefficient value conversion, and data silos. It establishes a closed-loop ecosystem from emotion perception to real-world value realization. Specifically in luxury retail scenarios, the system intelligently connects its micro-expression golden 3-second response rule with VIP service systems, enabling the precise conversion of high-net-worth customers' emotional capital into target value.

[0172] Through the aforementioned technical architecture, this system achieves dynamic gradient response and closed-loop transformation of user emotional value in target scenarios. The decision-making center tracks the emotional index over N consecutive time series units, accurately capturing the evolving trajectory of consumer decision-making tendencies. When the index crosses the emotional value threshold for three consecutive units, a primary response is triggered, establishing an emotional connection through immersive interaction. When the emotional intensity continues to rise for five consecutive units, a secondary response is activated, writing real-time emotional value tags to the customer data platform (CDP) via the marketing middleware API. This triggers the intelligent recommendation engine to generate precise marketing instructions, guiding consumer decisions at the behavioral level. The deep coupling of the marketing middleware API and the CDP enables emotional insights to be directly connected to enterprise-level intelligent systems, achieving an end-to-end closed-loop value chain from micro-expression recognition to marketing action execution. Within this multimodal collaborative architecture, the emotional computing engine provides a denoised emotional feature base. The decision-making center dynamically matches marketing strategy levels based on the emotional intensity gradient. The marketing middleware API serves as the value conversion hub, translating emotional data into actionable language, building a full-link intelligent marketing ecosystem encompassing "perception-analysis-monetization." In luxury retail verification scenarios, the system uses the golden 3-second response rule for micro-expressions to link with the VIP service system, achieving precise temporal and spatial matching of high-net-worth customers' emotional peaks and exclusive rights. While improving customer experience, it converts emotional capital into actual customer unit price growth, significantly improving the conversion rate and customer lifetime value of high-end consumption scenarios.

[0173] like Figure 4 As shown in the figure, the multimodal emotion computing engine and decision center are deployed on the edge intelligent terminal, and intelligent allocation of computing resources is achieved through the following mechanisms:

[0174] When the customer behavior recognition module detects an attention interruption, it pauses the micro-expression intensity gradient analysis and restarts emotion tracking after the gaze returns to the correct position.

[0175] When the ambient light fluctuation is stable, the frequency of emotion feature extraction of the video stream is switched from 30fps to 8-15fps adaptive sampling mode.

[0176] In the target scenario implementation plan, flexible system resource scheduling directly impacts user experience and operating costs. In a customer service scenario, when the visual sensor detects that the user's gaze has continuously strayed from the interactive interface, the affective computing engine automatically pauses eye micro-expression analysis. The passive relaxation of facial muscles at this time can lead to emotional misjudgment. Pausing detection avoids inefficient computing and reduces privacy concerns. When the infrared sensor detects a return of gaze, the system resumes full-dimensional emotion tracking using a gradual restart strategy.

[0177] When deployed in luxury stores, the system automatically activates a dynamic sampling optimization algorithm when the ambient light sensor detects stable lighting conditions. By analyzing the confidence curve of historical emotion data, it intelligently adjusts the video processing frequency to the minimum necessary level of 8fps while continuously monitoring emotional characteristics such as key zygomatic muscle activation. Field tests have shown that this strategy can reduce GPU resource utilization by 42%, increasing the concurrent processing capacity of a single device to real-time analysis of 20 customer emotion streams.

[0178] The resource optimization control module achieves a dynamic balance between computing accuracy and resource consumption through the collaboration of a scenario-based rule engine and a deep learning prediction model: it automatically increases the sampling rate to 30fps full-frequency capture during high-value customer interaction windows, and switches to a lightweight detection mode during periods of environmental interference. This "pulse-based" resource allocation strategy not only ensures the accuracy of emotion analysis at core target moments, but also significantly extends the battery life of smart terminals, reducing the system's average power consumption by 37% during an 8-hour business period. It is particularly suitable for the continuous operation needs of high-end retail scenarios.

[0179] The resource optimization controller generates sampling rate control instructions based on lighting stability (change rate <5 lux / second). For example, when lighting is stable, the video stream sampling rate can be reduced from 30fps to 15fps. The sampling rate control module receives these instructions and adjusts the camera acquisition frequency to reduce the computational load and ensure efficient operation of the edge computing unit.

[0180] The multi-source behavioral perception center captures customer interaction dynamics in real time, identifying high-value behaviors or attention interruptions. Data is transmitted to the affective computing engine via the data center, activating intelligent resource scheduling strategies. Micro-expression tracking is suspended during periods of customer distraction to avoid misjudgments caused by natural facial relaxation. The ambient light sensing module continuously monitors store lighting fluctuations, quantifying the threshold for how light changes affect emotion feature extraction, and setting a high interference threshold of 100 lux / second. Behavioral and light sensing data are synchronized via the data center to the resource optimizer and affective feature analysis module.

[0181] As an intelligent decision-making hub, the data middle platform aggregates multiple data streams from visual sensors, voice collectors, environmental monitors, etc., and realizes multi-module collaboration: it receives the golden gaze signal from the behavioral perception center, triggers the emotional computing engine to increase the weight of eye micro-expression analysis, synchronizes the illumination gradient data of the ambient light sensor, dynamically adjusts the parameters of the facial feature enhancement algorithm, and transmits the optimized emotional feature stream to the consumption prediction model to generate real-time marketing decision instructions.

[0182] The intelligent central platform monitors the system's operating status globally, setting target scenario-specific parameters such as emotional value thresholds and industry adaptation modes. It also configures emotional fluctuation tolerance ranges for differentiated scenarios, such as setting a pleasure variance threshold of V1 = 0.08 for luxury retail and V1 = 0.15 for online customer service. It also receives graded response signals from the real-time marketing engine.

[0183] Through intelligent collaboration across multiple modules, the system achieves flexible scheduling of computing resources. When detecting distracted customer attention or drastically fluctuating ambient light, it prioritizes the stable extraction of core emotional features to avoid resource congestion. Based on a multi-dimensional data fusion architecture, it dynamically assigns parsing weights for facial action units, voice emotional fundamental frequency, and environmental semantics to accurately quantify customer emotional value. The scenario adaptation engine dynamically configures parameter matrices to achieve intelligent adaptation across cross-industry scenarios. The customer emotion database and deep transfer learning model form a closed-loop data evolution loop, continuously optimizing the accuracy of the emotion-consumption correlation model and forming an intelligent iterative "perception-decision-verification" mechanism.

[0184] Empirical evidence demonstrates that the emotional intelligence solution based on this system architecture reduces emotion misjudgment in complex environments through multimodal emotion fusion computing and scenario-based parameter optimization. Furthermore, leveraging the dynamic resource scheduling of edge intelligent terminals, it ensures real-time response (<200ms latency) in high-concurrency scenarios. In some implementations, after implementing this system, an international beauty brand saw its customer emotion recognition accuracy increase to 91%. Through its emotion-driven "Golden 5 Seconds" marketing strategy, it also boosted its trial conversion rate by 34% and its average order value by 22%, validating the effectiveness of the technical solution.

[0185] In an edge server deployment architecture, the system dynamically optimizes resources by monitoring ambient lighting changes and user interaction status in real time. If a user interaction interruption lasts for more than 10 seconds, the system automatically suspends physiological signal detection and disables the galvanic skin response analysis module. Once the user re-enters the screen or the voice energy level returns to above 50dB, the system restarts the physiological signal detection mechanism after a 5-second delay.

[0186] When the system detects that the user interaction interruption lasts for more than 10 seconds:

[0187]

[0188] in To indicate continuous frame loss (30fps x 10 seconds = 300 frames), the system automatically pauses physiological signal detection and releases resources. The measured GPU resource release rate is:

[0189]

[0190] The restart mechanism uses a delay control algorithm:

[0191]

[0192] When the system is interrupted for 15 seconds, the system will restart after a delay of 7.5 seconds, effectively avoiding instantaneous interference.

[0193] In terms of lighting environment optimization, the system continuously collects illumination data through the ambient light sensor, and calculates the illumination change rate. When the light level is lower than 5 lux / second, the system starts the resource optimization process: the real-time video stream sampling rate is reduced from 30fps to 15fps, and non-critical voice analysis threads such as prosody analysis and voiceprint verification are shut down.

[0194] Calculate the illuminance gradient in real time using the ambient light sensor:

[0195]

[0196]

[0197] Optimize the video sampling rate from 30fps to 15fps and close non-critical voice threads.

[0198] Memory release calculation:

[0199]

[0200] CPU usage dropped by 28%.

[0201] This adaptive resource management mechanism enables the system to maintain optimal operating status in highly dynamic environments such as retail stores. When customers linger at the counter for a long time, the system provides accurate sentiment analysis by detecting 30fps video streams. When customers temporarily leave or the lighting in the store remains stable, it automatically switches to low-power mode, controlling the average daily energy consumption of edge devices to within 45W, achieving an intelligent balance between analysis accuracy and resource consumption.

[0202] The facial emotion intelligence system built based on this solution significantly improves the accuracy and timeliness of emotion value mining in target scenarios through the triple mechanisms of dynamic feature fusion, scenario adaptive optimization, and flexible resource scheduling. At the same time, it effectively avoids the risk of misjudgment caused by environmental interference and natural behavior, providing a reliable technical foundation for applications such as intelligent marketing and experience management, and helping enterprises achieve a strategic upgrade from traffic operation to emotional capital operation.

[0203] Specifically, the emotional intelligence system built based on this solution has the following core advantages:

[0204] Dynamic Weighting: This mechanism dynamically adjusts the weighting of multi-dimensional emotional indicators by tracking emotion intensity trends, psychological fluctuations, and transient micro-expression characteristics in real time. This prioritizes high-value signals, ensuring the sensitivity and relevance of emotional judgments to decision-making, and overcoming the issues of missed opportunities or false triggers caused by traditional static thresholds.

[0205] Target Scenario Adaptive Engine: Targeting the differentiated needs of retail shopping guides and online customer service, the engine intelligently configures micro-expression response thresholds from 0.6 to 0.4, emotion sampling windows from 500ms to 300ms, and environmental interference tolerance for sudden changes in illumination from 100lux to 50lux. This engine not only meets the real-time demands of fast-paced consumer decision-making, but also filters out non-emotional distractions like swaying curtains in fitting rooms and customers browsing through documents, achieving precise adaptation across multiple industry scenarios.

[0206] Intelligent elastic resource scheduling: Relying on edge computing nodes to monitor customer behavior, concentration, gaze duration > 3 seconds, and ambient light stability, illumination gradient < 10 lux / second in real time, non-core analysis is suspended when attention is interrupted and the head turns for > 2 seconds. When the lighting is stable, an adaptive frequency reduction strategy is activated from 30fps to dynamic 8-15fps. While ensuring the accuracy of core emotional feature extraction, the concurrent processing capacity of a single device is increased by 2.3 times.

[0207] A multi-dimensional anti-interference verification system integrates ambient light gradient patterns, a spatiotemporal feature matrix for facial occlusion compensation, decoupling of facial muscles to eliminate head sway based on 3D facial reconstruction, and enhanced micro-expression analysis using a high-frequency AU25 / 12 capture system to create a closed loop for emotional authenticity verification. This effectively distinguishes genuine consumer intent from environmental interference, keeping the target misjudgment rate below 3%.

[0208] These technological advantages directly address the three major pain points of traditional solutions in dynamic scenarios: mismarketing due to environmental sensitivity, fragmented experience caused by poor cross-industry adaptability, and computing power bottlenecks in high-concurrency scenarios. By building a three-in-one architecture combining "emotional feature enhancement, intelligent scenario adaptation, and dynamic resource allocation," this platform provides a low-latency, high-precision emotional computing foundation for high-value scenarios like retail and finance, helping enterprises achieve the value transition from traffic operations to emotional capital operations.

[0209] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. Emotion detection system based on facial recognition, characterized by: include: The timing analysis module is used to receive the real-time video stream of the target scene, extract the facial region image sequence of N consecutive frames, perform physiological signal detection, and generate a timing feature vector containing micro-expression intensity gradient, illumination robustness coefficient, and facial action unit collaborative features; The multimodal dynamic perception module is used to perform the following processing: Calculate the emotion classification probability and its changing trend within the sliding time window according to the time series feature vector, and activate speech emotion compensation analysis when the variance of the positive emotion category probability in the emotion classification probability exceeds the dynamic threshold V1 and the speech signal-to-noise ratio ≥ S1; Calculating micro-expression intensity differences based on the facial region image sequence of adjacent frames and obtaining an accumulated amount of emotional fluctuations, and triggering physiological signal enhancement detection to generate physiological signal intensity when the accumulated amount of emotional fluctuations exceeds a dynamic threshold; Extracting the facial region image sequence of the current frame and the three previous frames to calculate the difference of the voice emotion parameter, and generating the voice emotion dynamic fluctuation coefficient; The fusion decision module is used to perform the following operations: When the emotion classification probability increases monotonically and the speech emotion parameter is lower than the preset stability threshold, the emotion classification weight is increased. W1; When the accumulated amount of emotional fluctuation exceeds the dynamic threshold and the probability variance of the positive emotion category is lower than the dynamic threshold V1, the physiological signal weight is increased. W2; described W1 and The weight increase of W2 is dynamically calculated by a preset algorithm. The emotion classification probability, the accumulated amount of emotion fluctuations and the physiological signal strength are weighted and fused according to the dynamic weight to generate a comprehensive emotion score and trigger a hierarchical decision instruction.

2. The emotion detection system based on facial recognition according to claim 1, characterized in that The time series feature vector also includes: The facial region emotion consistency index of adjacent frames. When the consistency index is lower than 80%, the cross-frame emotion feature matching system is enabled to correct the emotion label. The phase parameter of periodic occlusion is used to predict the location of the next occlusion area.

3. The emotion detection system based on facial recognition according to claim 2, characterized in that The cross-frame emotion feature matching system performs the following operations: Selecting a key frame with stable lighting in the real-time video stream as an emotion reference template; In the real-time video stream, facial key point matching is performed on low-consistency frames to calculate an affine transformation matrix to correct emotion analysis coordinates; The coordinate error of the corrected emotion analysis is controlled within 5% of the original facial area, and the emotion reference template is updated every 10 seconds.

4. The emotion detection system based on facial recognition according to claim 1, characterized in that The time size of the sliding time window is dynamically adjusted according to the target scenario: The target scenarios include a retail scenario and a customer service scenario; the sliding time window of the retail scenario is set to 15 seconds, and the sliding time window of the customer service scenario is set to 5 seconds.

5. The emotion detection system based on facial recognition according to claim 1, characterized in that The specific steps of analyzing the emotion classification probability and its changing trend are as follows: Detecting the degree of deviation between the moving average value of the positive emotion category probability and the instantaneous value, and marking it as a visually suspicious fluctuation segment when the degree of deviation is greater than 25%; The marked suspicious emotion fluctuation segment is used to enable a temporal convolutional network to predict the emotion probability change curve in the next 2 seconds.

6. The emotion detection system based on facial recognition according to claim 1, characterized in that The specific steps of the speech emotion compensation analysis are as follows: Detecting the probability fluctuation amplitude of the moving average value and the instantaneous value of the probability of the positive emotion category, and marking it as a suspicious voice fluctuation segment when the fluctuation amplitude is greater than 25%; When the probability variance of the positive emotion category exceeds the dynamic threshold V1, the abnormal fluctuation data of the speech parameters caused by the environmental noise is ignored; The speech emotion parameters are smoothed by a Kalman filter to eliminate instantaneous interference.

7. The emotion detection system based on facial recognition according to claim 1, characterized in that The triggering of physiological signal enhancement detection and generating physiological signal intensity adopts the following processing: The physiological signal includes heart rate variability. When the accumulated amount of emotional fluctuation exceeds the dynamic threshold, the physiological signal enhanced detection increases the heart rate variability detection frequency from 1 Hz to 3 Hz. The physiological signal difference refers to the heart rate variability value at the current moment Heart rate variability value compared to the previous frame The difference is used to quantify the dynamic change intensity of the physiological signal between adjacent frames. The physiological signal difference generated by the physiological signal enhancement detection is normalized to eliminate the instantaneous noise interference caused by body movement. The calculation formula is: ; in: Indicates the heart rate variability value at the current moment. is the heart rate variability value of the previous frame; and The maximum and minimum values ​​of HRV intensity in the historical data of the current scene; ε is a very small constant that prevents the denominator from being zero.

8. The emotion detection system based on facial recognition according to claim 4, characterized in that: described W1 and The preset algorithm calculation method of the dynamic weight increase of W2 is: ; ; in: and is the preset basic weight coefficient; t is the duration of monotonically increasing emotion classification probability, in seconds; C is the cumulative amount of emotional fluctuations; , Automatically configured according to the target scenario type.

9. The emotion detection system based on facial recognition according to claim 1, characterized in that: The triggering hierarchical decision instruction includes: When a first-level response is triggered, personalized advertising push and coupon distribution will be initiated; When a secondary response is triggered, high-priority customer service intervention is initiated and the incident is simultaneously reported to the target intelligent platform.

10. The emotion detection system based on facial recognition according to claim 1, characterized in that The multimodal dynamic perception module and the fusion decision module are deployed on the edge computing server, and: When the interaction is interrupted for more than 10 seconds, the physiological signal detection is suspended and restarted after the interaction is resumed; When the light intensity change rate is less than 5 lux / second, the sampling rate of the real-time video stream is reduced from 30 fps to 15 fps, and a non-critical voice analysis thread is closed.

Citation Information

Cited By

  • Immersive emotional state multi-dimensional monitoring and feedback system based on VR

    CN120959745A

  • Emotion recognition method and device in video scene, equipment and medium

    CN121421538A

  • Multi-modal data video inspection concentration analysis and early warning method and system

    CN121459404A

  • Micro-expression analysis method based on facial key point recognition

    CN121640549A

  • Video editing processing method based on artificial intelligence

    CN121728310A