Automatic Video Layout Generation for Telepresence MCU
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional multi-point and multi-stream videoconferencing systems require manual management by human operators to dynamically arrange video streams, leading to potential errors and high costs due to the need for specialized training and the difficulty in determining which video stream includes the current speaker.
Innovation Solution
A continuous presence telepresence MCU that automatically generates video stream layouts by using a processor with a stream attribute module to assign attributes to outgoing streams, determining the current speaker, and a layout manager to dynamically adjust the layout based on these attributes and endpoint configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If manual management by human operators is used to dynamically arrange video streams, then the layout can be adjusted dynamically, but human errors and costs due to specialized training increase
Solution Approach 1:
The system performs automatic layout arrangement without human intervention. The MCU automatically receives video streams, identifies current speakers, determines camera positions, and generates appropriate layouts based on endpoint configurations, eliminating the need for human operators to manually manage video stream arrangements
Solution Approach 2:
The patent replaces the mechanical system of manual human operation with an automated electronic system. The MCU uses processor-based algorithms to automatically analyze video stream attributes, detect current speakers, and dynamically generate layouts, substituting human manual management with automated computational processes
2Adaptability or versatility
If manual management by human operators is used to dynamically arrange video streams, then the layout can be adjusted dynamically, but the cost of specialized training and human operators increases
Solution Approach 1:
The system performs automatic layout arrangement without human intervention. The MCU automatically receives video streams, identifies current speakers, determines camera positions, and generates appropriate layouts based on endpoint configurations, eliminating the need for human operators to manually manage video stream arrangements
Solution Approach 2:
The patent replaces the mechanical system of manual human operation with an automated electronic system. The MCU uses processor-based algorithms to automatically analyze video stream attributes, detect current speakers, and dynamically generate layouts, substituting human manual management with automated computational processes
3Area of stationary object
If multiple video streams are received from each endpoint, then more comprehensive coverage is achieved, but the difficulty in determining which video stream includes the current speaker increases
Solution Approach 1:
The system uses feedback from the speaker locator module to identify which video stream contains the current speaker. The speaker locator detects the speaker's location, and this information feeds back to the layout manager, which uses it to determine the appropriate video stream from multiple streams for prominent display
Solution Approach 2:
The patent introduces an intermediary speaker locator module that bridges the gap between multiple video streams and the current speaker identification. This intermediary component analyzes audio-visual data to determine which video stream contains the current speaker, making the detection process easier and more reliable
4Device complexity
If static layout arrangement is used, then system complexity is reduced, but the system cannot adapt to dynamic conferencing needs
Solution Approach 1:
The system implements dynamic layout arrangement where the layout is continuously adjusted based on real-time conferencing conditions. The layout manager receives video streams with attributes, identifies current speakers, and dynamically generates layouts that adapt to changing speaker positions and endpoint configurations, making the system flexible and responsive
Solution Approach 2:
The patent changes the parameter of layout arrangement from static to dynamic. The system automatically modifies layout parameters such as video stream positioning, scaling, and visibility based on real-time detection of current speakers and endpoint configurations, enabling adaptive layout management
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A videoconference multipoint control unit, MCU, (106) automatically generates display layouts for videoconference endpoints (101-103). Display layouts are generated based on attributes associated with video streams (315, 316) received from the endpoints (102, 103) and display configuration information (329) of the endpoints (101-103). An endpoint (101-103) can include one or more attributes in each outgoing stream. Attributes can be assigned based on video streams' role, content, camera source, etc. Display layouts can be regenerated if one or more attributes change. A mixer (303) can generate video streams to be displayed at the endpoints (101-103) based on the display layout.