Video Conferencing Gesture Feedback System
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems lack effective non-obtrusive feedback mechanisms for presenters to timely identify participants' intentions to intervene, leading to difficulties in managing interactions during online workshops, presentations, or meetings.
Innovation Solution
A video conferencing system that uses an image sensor to capture participants' body movements or facial expressions, processing this information with computer-vision techniques to recognize specific gestures or emotions, and then transmitting ambient graphics to the presenter's display without obscuring the main content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Difficulty of detecting and measuring
If numerous live videos of participants are displayed on the presenter's screen to enable timely identification of intervention intentions, then the ability to detect participant feedback is improved, but the original video material becomes obscured and harder to view
Solution Approach 1:
The system segments participant feedback detection from the main presentation content by using separate ambient display zones. Participant videos are segmented into small thumbnail indicators positioned in corners or edges of the screen, while the main presentation content occupies the central and primary viewing areas, ensuring both are visible without mutual obstruction.
Solution Approach 2:
The system transitions from a two-dimensional content display to a multi-dimensional spatial arrangement by utilizing peripheral and corner regions of the display screen. Participant feedback indicators are placed in spatial dimensions that do not compete with the main content area, creating a layered display hierarchy where important content remains primary while feedback information is accessible in secondary spatial zones.
2Reliability
If participants manually unmute microphones to intervene during presentations, then clear audio feedback is achieved, but the natural flow of presentation is interrupted
Solution Approach 1:
The system performs preliminary detection of participant intervention intentions through computer vision analysis of body movements and facial expressions before the participant actually speaks or unmutes. This allows the system to prepare and display feedback indicators in advance, so when the participant does intervene, the transition is smoother and requires less manual action, reducing interruption to the presentation flow.
Solution Approach 2:
The system implements continuous visual feedback to the presenter about participant engagement states through ambient graphics that change based on detected body movements and facial expressions. This feedback loop allows the presenter to sense participant readiness to intervene without requiring participants to manually unmute, maintaining presentation flow while ensuring clear communication of intervention intentions.
3Productivity
If the presenter constantly observes live videos of all participants to identify intervention intentions, then timely response to participants is improved, but the presenter's attention is diverted from the main presentation content
Solution Approach 1:
The system extracts the task of monitoring participant feedback from the presenter by implementing automated computer vision analysis. The system continuously analyzes participant videos for body movements and facial expressions, extracting meaningful feedback signals and presenting them to the presenter in simplified ambient graphic forms, eliminating the need for the presenter to manually monitor all participant videos.
Solution Approach 2:
The system enables self-service monitoring where the automated analysis and interpretation of participant feedback occurs without requiring presenter intervention or attention. The ambient graphics automatically update based on real-time participant behavior analysis, allowing the presenter to maintain focus on presentation content while the system independently manages the monitoring and feedback presentation tasks.
Data Source
AI summary
The present disclosure relates to a video conferencing system which provides non-obtrusive feedback to a presenter during a video conference to improve interactivity among participants of video conference. An image sensor, such as a camera, captures a sequence of images (i.e., a video) of the participant while a participant of the video conference is moving their port or a part of their body and sends the images to video conferencing server. The video conferencing server processes the images to recognize a type of gesture performed by the participant and selects an ambient graphic that corresponds to the recognized gesture. The video conferencing server sends the ambient graphic to a client device associated with the presenter. The client device associated with the presenter renders or displays the ambient graphic on a display screen of the client device without obscuring information displayed on the display screen of the client device.


