Emotional State Estimation via Multi-Sensor Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current emotional and psychological state estimation systems lack accuracy and reliability, particularly in real-time applications and environments where human behavior and environmental factors are complex, such as high-stress professions and interrogations.
Innovation Solution
A system comprising a group of sensors and a processing unit that acquires and processes multiple data types, including video, audio, physiological, and environmental parameters, using predefined reference models like NLP and color psychology to estimate emotional states with high accuracy and adaptability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple data types and sensors are used to improve estimation accuracy, then measurement precision improves, but device complexity increases
Solution Approach 1:
The system segments the emotional state estimation task into multiple independent sensing modules (video sensors for facial expressions, audio sensors for voice tone, physiological sensors for heart rate and skin conductance, environmental sensors for temperature and humidity). Each sensor type captures a specific aspect of emotional state, and the processing unit integrates these segmented data streams to achieve comprehensive and accurate estimation without requiring a single overly complex sensing mechanism.
Solution Approach 2:
The processing unit serves multiple functions simultaneously: it processes video data for facial expression analysis, analyzes audio data for voice tone detection, integrates physiological sensor data, incorporates environmental factors, and applies reference mathematical models for emotional state classification. This multi-functional approach consolidates what would otherwise require separate processing systems, improving accuracy while managing complexity through a unified processing architecture.
2Reliability
If comprehensive sensors and processing are used to improve reliability, then reliability improves, but use of energy increases
Solution Approach 1:
The system implements adaptive sampling and processing where not all sensors operate at full capacity simultaneously. The processing unit selectively activates specific sensing and processing pathways based on the current context and detected emotional state patterns. For example, during low-stress periods, full physiological monitoring may be reduced while maintaining video and audio monitoring, thereby maintaining reliability for detecting significant changes while reducing overall energy consumption during routine monitoring phases.
3Productivity
If real-time processing is implemented, then productivity improves, but use of energy increases
Solution Approach 1:
The system employs periodic processing cycles where data from all sensors is collected over a defined time window and then processed together in batches rather than continuously processing every incoming data point in real-time. This periodic approach maintains the ability to detect and respond to emotional state changes while significantly reducing the instantaneous processing load and associated energy consumption compared to continuous real-time processing of all sensor streams.
Data Source
Figure 1~2

AI summary
The present invention concerns a human emotional/behavioural/psychological state estimation system (1) comprising a group of sensors and devices (11) and a processing unit (12). The group of sensors and devices (11) includes : a video-capture device (111); a skeletal and gesture recognition and tracking device (112); a microphone (113); a proximity sensor (114); a floor pressure sensor (115); user interface means (117); and one or more environmental sensors (118). The processing unit (12) is configured to : acquire or receive a video stream captured by the video-capture device (111) and data items provided by the skeletal and gesture recognition and tracking device (112), the microphone (113), the proximity sensor (114), the floor pressure sensor (115), the environmental sensor (s) (118), and also data items indicative of interactions of a person under analysis with the user interface means (117); detect one or more facial expressions and a position of eye pupils, a body shape and features of voice and breath of the person under analysis; and estimate an emotional/behavioural/'psychological state of said person on the basis of the acquired/received data items, of the detected facial expression (s), position of the eye pupils, body shape and features of the voice and the breath of said person, and of one or more predefined reference mathematical models modelling human emotional/behavioural/psychological states.