Video Stream Selection for Non-Verbal Activity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video conferencing systems often require users to monitor multiple video feeds to detect non-verbal activities, which can be distracting and inefficient, especially in large groups or on devices with limited display space, and may unnecessarily allocate network bandwidth.
Innovation Solution
A communication system that implements a 'Storied Experience View' by selecting and displaying video streams based on detected non-verbal events, using a central server to manage video streams and allocate bandwidth efficiently, while ensuring that active speakers' audio remains uninterrupted.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all video streams are displayed to show complete non-verbal activities of all participants, then user awareness of group dynamics is improved, but device display space is overwhelmed and user attention is lost
Solution Approach 1:
The system extracts and displays only the most relevant video streams containing significant non-verbal activities (such as reactions, gestures, or expressions) while filtering out less important feeds. This selective extraction allows the system to maintain awareness of group dynamics without overwhelming the display space with all participant feeds simultaneously.
Solution Approach 2:
Different regions or time periods of the video conference are treated with different display priorities. The system dynamically adjusts which video streams receive display priority based on the significance of non-verbal activities detected in local time windows, ensuring that important moments are captured without continuously occupying all display space.
2Loss of information
If multiple video feeds are continuously monitored to detect all non-verbal activities, then detection completeness is improved, but network bandwidth is wasted
Solution Approach 1:
The system performs preliminary analysis of video streams to identify potential non-verbal activities before full display or transmission. By pre-detecting significant events using automated video analysis, the system can selectively transmit or display only those streams containing meaningful non-verbal cues, avoiding continuous monitoring and transmission of all feeds.
Solution Approach 2:
The video conference system automatically analyzes its own video streams to detect non-verbal activities and makes autonomous decisions about which feeds to display or transmit. This self-service capability eliminates the need for manual monitoring of all feeds while optimizing bandwidth usage based on detected event significance.
3Ease of operation
If automated detection of non-verbal movements is implemented, then user engagement is improved, but system complexity increases
Solution Approach 1:
The system replaces manual user monitoring and selection of video feeds with automated computer vision algorithms that detect non-verbal movements and reactions. This substitution of mechanical/manual operations with automated image processing and pattern recognition enhances user engagement through intelligent, dynamic feed selection while managing system complexity through specialized algorithms.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In a communication event between a first user and one or more second users via a communication network, a plurality of video streams is received (S502) via the network. Each of the streams carries a moving image of at least one respective user. The moving image of a first of the video streams is displayed (S504) at a user device of the first user for a first time interval. In the moving image of a second of the video streams that is not displayed at the user device in the first time interval, a human feature of the respective user is identified (S506, S508). A movement of the identified human feature during the first time interval that matches one of a plurality of expected movements is detected (S510, S512). In response to the detected movement, at least the moving image of the second video stream is displayed (S516, S518, S520) at the user device for a second time interval.