Active Listening Support System Using Multimodal Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current communication support systems focus on improving speaking skills but neglect the emotional aspect of listening, leading to a lack of trust and hiding of important facts in conversations, especially in face-to-face interactions.
Innovation Solution
A system that analyzes conversational behaviors using image and speech recognition to identify positive listening behaviors, providing real-time feedback to enhance emotional engagement and improve communication quality by dynamically switching roles and using multimodal interfaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If communication support systems focus on improving speaking skills through content analysis, then speaking skills are improved, but the emotional aspect of listening is neglected leading to lack of trust
Solution Approach 1:
The system transitions from analyzing only the content dimension of communication to incorporating the emotional dimension by detecting facial expressions, gestures, and tone of voice. This dimensional expansion allows simultaneous improvement of content understanding and emotional connection, resolving the contradiction between factual accuracy and trust building.
Solution Approach 2:
The system introduces multimodal sensors (cameras, microphones) as intermediaries to capture non-verbal emotional cues that mediate between the speaker's intended message and the listener's understanding. These intermediaries provide additional information channels that foster trust without compromising content accuracy.
2Reliability
If systems analyze conversational behaviors using multiple sensors, then emotional engagement is enhanced, but device complexity increases
Solution Approach 1:
The system employs multimodal sensors that serve multiple functions: cameras capture both visual content and facial expressions, microphones record both speech content and tone variations. This multi-functionality allows emotional engagement enhancement without proportionally increasing device complexity, as the same hardware serves dual purposes.
Solution Approach 2:
The system merges content analysis and emotional analysis into a unified processing framework that simultaneously handles factual and emotional dimensions of communication. This consolidation reduces overall system complexity compared to having separate independent systems for each function.
3Productivity
If real-time feedback is provided during conversations, then listening skills are improved, but privacy concerns arise due to continuous monitoring
Solution Approach 1:
The system extracts only the necessary emotional and behavioral cues needed for feedback generation, rather than continuously monitoring and storing all conversation data. This selective extraction provides real-time listening feedback while minimizing privacy intrusion by processing only essential information.
Solution Approach 2:
The system implements feedback mechanisms that provide constructive listening guidance without requiring continuous surveillance. Feedback is generated based on analyzed cues and delivered in a manner that improves listening skills while respecting participant privacy by not exposing raw monitored data.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The disclosure regards a system, method and program for assisting a person in improving its conversation skills. The system comprises a perception module configured to perceive a conversation between the person and at least one other person by generating information representing the conversation based on acquired sensor information, and a processing module configured to classify the assisted person and the at least one other person into at least one listener and a speaker based on the generated information. The processing module is further configured to determine which potential behavior of the person classified as the listener correlates with a positive quality of the conversation by evaluating the current conversation situation based on the generated information, to determine which actual behavior the person classified as the listener is performing based on the generated information, and to determine at least one additional behavior for execution by the person classified as the listener that improves the quality of the conversation based on the generated information. The system comprises an interface configured to output a recommendation for an improved communication behavior to the person classified as the listener generated based on the determined at least one additional behavior.