Meeting User State Estimation Using Voice and Image Cues
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Remote meetings via communication networks face challenges in accurately grasping the emotions and states of participants due to limited visual and auditory information, making it difficult to understand user reactions and adapt interactions effectively.
Innovation Solution
An information processing device and method that estimates user states and change reasons based on image and voice data, generating graphs and reasons for state changes to be displayed on user terminals, using learning models and algorithms to analyze user states such as interest, understanding, and fatigue.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If remote meetings are conducted via communication networks, then convenience and accessibility are improved, but the ability to accurately grasp user emotions and states deteriorates
Solution Approach 1:
The system segments user state detection into multiple independent analysis components: facial expression analysis, voice tone analysis, and keyword analysis. Each component processes specific aspects of user behavior separately and contributes to the overall user state assessment, enabling comprehensive detection despite remote communication limitations
Solution Approach 2:
The system introduces an intermediary information processing layer that analyzes intermediate signals (facial expressions, voice characteristics, selected keywords) between the user and the meeting participant. This intermediary layer extracts meaningful state information from indirect cues available in remote communication, bridging the gap between limited visual/auditory data and accurate emotion detection
2Quantity of substance
If limited visual and auditory information is used in remote meetings, then network bandwidth and device requirements are reduced, but the ability to understand user reactions deteriorates
Solution Approach 1:
The system extracts specific diagnostic features from transmitted visual and auditory data: facial expression features from video feeds, voice tone features from audio feeds, and semantic features from transcribed speech. By extracting only the most relevant features rather than transmitting and analyzing all raw data, the system maintains user state detection accuracy while minimizing data transmission requirements
Solution Approach 2:
The system transforms raw visual and auditory data into standardized state parameters (user state scores, emotion categories, reaction types). This parameter transformation enables efficient transmission and processing of user reaction information, converting complex sensory data into compact, actionable metrics that preserve essential information about user states
Data Source
AI summary
A graph showing changes over time in score indicating a user state of a user participating in a meeting and a user state change reason are estimated and displayed on a terminal of another user participating the meeting. A user state score indicating a user state of any one of a degree of interest, a degree of understanding, or a degree of fatigue of a user participating in a meeting via a communication network is estimated on the basis of at least one of image data or voice data of the user, a user state output score to be output to a user terminal of the user participating in the meeting is calculated on the basis of the estimated user state score, and a graph indicating changes in calculated user state output score and user state change reason are displayed on user terminal of another user participating in meeting.


