Human-Object Interaction Emotion Recognition With Time-Space Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing emotion recognition methods are inaccurate due to the subjectivity of data sources like facial images and voice, and unreliable acquisition modes like contact sensors, limiting their application range.
Innovation Solution
A method for emotion recognition based on human-object time-space interaction behavior, using video data to construct a deep learning-based feature extraction and fusion model, integrating semantic-level information to enhance accuracy and interpretability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If facial images and voice data are used for emotion recognition, then the acquisition process is simple and non-contact, but the recognition accuracy deteriorates due to subjectivity and conformity psychology
Solution Approach 1:
The patent introduces human-object interaction behavior as an intermediary data source between the subject and the emotion recognition system. Instead of directly observing subjective facial expressions or voices, the system observes objective interaction behaviors (e.g., object manipulation, spatial relationships, temporal patterns) that indirectly reflect emotional states, thereby eliminating the subjectivity and conformity effects while maintaining non-contact acquisition simplicity
Solution Approach 2:
The patent replaces the traditional mechanical/optical observation of facial features and vocal patterns with a behavioral interaction observation system. By substituting direct facial/voice analysis with analysis of human-object interaction dynamics (position, movement, temporal-spatial relationships), the system achieves more objective emotion inference without requiring complex contact sensors
2Reliability
If physiological signals are used for emotion recognition, then the objectivity and reliability improve, but the application range deteriorates due to contact sensor requirements and user discomfort
Solution Approach 1:
The patent uses human-object interaction behavior as an intermediary that bridges objective measurement and broad applicability. By observing objective interaction behaviors through video analysis rather than direct physiological contact, the system achieves both reliability (objective data) and versatility (non-contact, comfortable for users, applicable in diverse settings)
Solution Approach 2:
The patent replaces contact-based physiological sensing mechanisms with non-contact video-based behavioral analysis. This substitution eliminates the need for wearables or sensors that cause user discomfort, thereby expanding application range while maintaining objectivity through objective observation of interaction behaviors
3Reliability
If contact sensors are used to acquire physiological signals, then the emotion recognition reliability improves, but the ease of operation deteriorates due to user discomfort and narrowed application scenarios
Solution Approach 1:
The patent introduces video-based behavioral observation as an intermediary that provides reliable emotion data without requiring direct contact with the user. By analyzing interaction behaviors captured by cameras, the system achieves both reliability (objective behavioral data) and ease of operation (user comfort, no sensors on body)
Solution Approach 2:
The patent substitutes contact-based physiological sensing with non-contact video analysis of behavior. This replacement eliminates physical sensors on the user's body, maintaining signal reliability through objective behavioral observation while significantly improving user comfort and expanding operational ease
Data Source
AI summary
An emotion recognition method includes the following steps: acquiring video data of a human-object interaction behavior process; performing data labeling on the positions of a person and an object and the interaction behaviors and emotions expressed by the person; constructing a feature extraction model based on deep learning, extracting features of interaction between the person and the object in a time-space dimension, and detecting the position and category of the human-object interaction behavior; mapping the detected interaction behavior category into a vector form through a word vector model; and finally, constructing a fusion model based on deep learning, fusing the interaction behavior vector and the time-space interaction behavior features, and identifying the emotion expressed by the interaction person.

