Visual Emotion Recognition via Physiological Signal Denoising
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing emotion recognition technologies based on facial expressions or physiological signals from images or videos are prone to misjudgment and have limited coverage in emotion expression analysis.
Innovation Solution
A visual perception-based emotion recognition method that combines face video with facial recognition algorithms, using physiological parameters like heart rate and breathing waveforms, and heart rate variability (HRV) for deep learning modeling after signal denoising and stabilization, and employs a valence-arousal-dominance (VAD) model for specific emotional representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If facial expression analysis is used for emotion recognition, then the method is simple and widely applicable, but the recognition accuracy is low and prone to misjudgment
Solution Approach 1:
The patent combines facial expression analysis with physiological signal analysis (heart rate, breathing, skin conductivity) to create a multi-modal emotion recognition system. This merging of multiple detection methods compensates for the limitations of each individual method, improving overall accuracy while maintaining ease of implementation through integrated processing.
Solution Approach 2:
The system uses composite feature extraction by combining multiple types of data (facial features, physiological signals, temporal dynamics) to form a comprehensive emotion representation. This composite approach allows the system to leverage the strengths of different signal types while mitigating their individual weaknesses.
2Adaptability or versatility
If physiological signals are extracted from face videos for emotion recognition, then the coverage of emotion expression is improved, but the signals are contaminated by motion and light changes leading to lack of rigor
Solution Approach 1:
The patent extracts and isolates specific physiological signal components from the complex face video data by focusing on controlled facial regions and using signal processing techniques to separate physiological variations from motion and illumination artifacts. This extraction approach improves signal reliability while maintaining comprehensive emotion coverage.
Solution Approach 2:
The system introduces intermediate processing steps including signal filtering, normalization, and artifact removal algorithms that act as mediators between the raw contaminated signals and the final emotion classification. These intermediary processes clean the physiological signals while preserving the emotional information.
3Device complexity
If a general categorical emotion model is used, then the classification is simplified, but the coverage of emotion expression is narrow and results in rough classification
Solution Approach 1:
The patent implements a dynamic emotion recognition model that adapts to different emotional states and intensities rather than using fixed categorical classifications. The system continuously adjusts its classification thresholds and weightings based on the detected physiological patterns, enabling it to capture a broader range of emotional expressions while maintaining manageable model complexity.
Data Source
AI summary
The present disclosure relates to the field of emotion recognition technologies, and specifically, to a visual perception-based emotion recognition method. The method in the present disclosure includes the following steps: step 1: acquiring a face video by using a camera, performing face detection, and locating a region of interest (ROI); step 2: extracting color channel information of the ROI for factorization and dimension reduction; step 3: denoising and stabilizing signals and extracting physiological parameters including heart rate and breathing waveforms, and heart rate variability (HRV); step 4: extracting pre-training nonlinear features of the physiological parameters and performing modeling and regression of a twin network-based support vector machine (SVM) model; and step 5: reading a valence-arousal-dominance (VAD) model by using regressed three-dimensional information to obtain a specific emotional representation.


