Policy-Based Video Stream Mapping for Stable Multi-Window Displays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing systems face challenges in managing multi-window displays, leading to clutter and high computational requirements when displaying all participants, especially in large meetings, causing confusion and distraction due to frequent changes in the arrangement of windows as speakers change.
Innovation Solution
Implementing a policy-based system that maps source video streams to destination streams, minimizing the number of windows that change their contents by determining the most desirable streams based on policies such as recent speakers, motion, or geographic proximity, and using a media switch or event-aware stream router to manage these mappings, ensuring a stable and reduced number of window changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all participants are displayed on the screen at the same time, then complete information about all participants is provided, but the display becomes cluttered and computational requirements become prohibitive
Solution Approach 1:
The display is segmented into multiple functional windows rather than showing all participants in a single view. The screen is divided into a primary window for the most recent speaker and secondary windows for other relevant participants, organizing information into manageable segments that reduce clutter while maintaining completeness
Solution Approach 2:
The system extracts and prioritizes only the most relevant participants for display based on speaking activity and policy criteria. Instead of showing all participants, it selectively extracts those who have recently spoken or are most relevant to the current context, reducing display complexity while preserving essential information
2Device complexity
If a small subset of participants is displayed based on recent speakers, then computational requirements are reduced and display clutter is minimized, but the arrangement of windows changes frequently causing confusion and distraction
Solution Approach 1:
The system pre-establishes a stable mapping between destination video streams and window positions before participants change. By determining the mapping in advance and maintaining it consistently, the display composition remains stable even as speaker priorities change, preventing frequent rearrangements that cause confusion
Solution Approach 2:
The system changes the selection criteria parameters (such as time thresholds for 'recent speakers' or policy weights) rather than changing the window arrangement itself. This allows the content to adapt to new speakers while maintaining the same stable spatial composition, reducing visual disruption
3Adaptability or versatility
If the arrangement of windows changes dynamically when speakers change, then the display reflects current speaking activity, but frequent shuffling causes confusion and distraction
Solution Approach 1:
The display adapts by changing content within stable segmented windows rather than rearranging the windows themselves. Each window maintains its position and function, while the content (video stream) within each segment adapts to reflect current speaking activity, providing adaptability without compositional instability
Solution Approach 2:
The system uses an intermediary mapping layer between speaker priority and window arrangement. Instead of directly mapping speakers to window positions (which causes shuffling), it uses a stable destination stream mapping as an intermediary that decouples content changes from positional changes, maintaining stability while enabling adaptability
Data Source
AI summary
Techniques for dynamically mapping source video streams of sources to the requested destination video streams based on a policy are provided. The source video streams that are mapped to the destination video streams are changed based on events that cause changes in the mapping based on the policy. The mappings may be managed by a media switch remote from the end device or by an event aware stream router associated with the end device. The mappings are used to display images of participants associated with the source video streams where position changes in images displayed are minimized when events occur.


