Video Conference Participant Prioritization for Non-Verbal Cue Visibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conference solutions fail to address the unique issues of missing important visual cues and latency in switching between active speakers, leading to missed non-verbal cues and engagement indicators among participants.
Innovation Solution
Prioritize video presentation of participants based on non-verbal cues such as emotions, gestures, and connections, using machine learning algorithms to analyze video and audio feeds, and adjust display positioning and visibility on the graphical user interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If video conference systems prioritize display based on active speaker detection, then the current speaker is highlighted, but important non-verbal cues from other participants are missed and latency occurs in switching between speakers
Solution Approach 1:
The system performs preliminary analysis of non-verbal cues (facial expressions, gestures, body language) from all participants simultaneously, preparing priority rankings in advance. This allows the system to switch between speakers more quickly by having pre-computed priority information ready, reducing the latency associated with real-time speaker switching while maintaining accurate detection of important visual cues.
2Loss of information
If video conference systems display all participants equally, then all visual cues are visible, but the interface becomes cluttered and important cues are lost in the noise
Solution Approach 1:
The system applies different display qualities and priorities to different participant videos based on their current state and relevance. Videos of participants displaying important non-verbal cues are enhanced and positioned prominently, while less relevant videos are diminished or minimized. This selective quality adjustment preserves important visual information while maintaining a clean, manageable interface layout.
Solution Approach 2:
Instead of processing and displaying all participant videos with equal detail, the system focuses computational resources on analyzing and displaying only the most relevant non-verbal cues from participants. This partial action approach highlights key visual information (such as surprised expressions or emphatic gestures) while reducing overall interface complexity by not treating all videos uniformly.
3Reliability
If video conference systems use traditional active speaker switching, then audio clarity is maintained, but non-verbal communication and participant engagement are lost
Solution Approach 1:
The video conferencing system performs multiple functions simultaneously: it continues to maintain audio clarity through traditional speaker switching while also analyzing, prioritizing, and highlighting non-verbal visual cues from all participants. The system universally applies non-verbal cue detection across all video feeds, enabling it to identify and emphasize important visual information (facial expressions, gestures, body language) alongside maintaining reliable audio communication, thus preserving both audio reliability and non-verbal communication.
Data Source
AI summary
Present principles are directed in part to prioritizing different participant videos during a video conference between remotely-located participants. For example, presentation of first video of a first video conference participant can be prioritized over second video of a second video conference participant based on one or more criteria not necessarily having to do with the first video conference participant currently speaking as part of the video conference. Therefore, in various particular non-limiting implementations, the one or more criteria can include the first video conference participant non-verbally expressing a particular emotion, making a particular non-verbal gesture, and/or having a connection outside the video conference to a viewer of the first video.


