Automatic Video Layout Generation for Telepresence MCU

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional multi-point and multi-stream videoconferencing systems require manual management by human operators to dynamically arrange video streams, leading to potential errors and high costs due to the need for specialized training and the difficulty in determining which video stream includes the current speaker.

Innovation Solution

A continuous presence telepresence MCU that automatically generates video stream layouts by using a processor with a stream attribute module to assign attributes to outgoing streams, determining the current speaker, and a layout manager to dynamically adjust the layout based on these attributes and endpoint configurations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual management by human operators is used to dynamically arrange video streams, then the layout can be adjusted dynamically, but human errors and costs due to specialized training increase

Engineering Contradiction:
Improvedynamic layout arrangementVSAvoidhuman error
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs automatic layout arrangement without human intervention. The MCU automatically receives video streams, identifies current speakers, determines camera positions, and generates appropriate layouts based on endpoint configurations, eliminating the need for human operators to manually manage video stream arrangements

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual human operation with an automated electronic system. The MCU uses processor-based algorithms to automatically analyze video stream attributes, detect current speakers, and dynamically generate layouts, substituting human manual management with automated computational processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If manual management by human operators is used to dynamically arrange video streams, then the layout can be adjusted dynamically, but the cost of specialized training and human operators increases

Engineering Contradiction:
Improvedynamic layout arrangementVSAvoidoperational cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system performs automatic layout arrangement without human intervention. The MCU automatically receives video streams, identifies current speakers, determines camera positions, and generates appropriate layouts based on endpoint configurations, eliminating the need for human operators to manually manage video stream arrangements

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual human operation with an automated electronic system. The MCU uses processor-based algorithms to automatically analyze video stream attributes, detect current speakers, and dynamically generate layouts, substituting human manual management with automated computational processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Area of stationary object

If multiple video streams are received from each endpoint, then more comprehensive coverage is achieved, but the difficulty in determining which video stream includes the current speaker increases

Engineering Contradiction:
Improvevideo stream coverageVSAvoidspeaker identification
Core Design Contradiction:
Area of stationary objectVSDifficulty of detecting and measuring

Solution Approach 1:

The system uses feedback from the speaker locator module to identify which video stream contains the current speaker. The speaker locator detects the speaker's location, and this information feeds back to the layout manager, which uses it to determine the appropriate video stream from multiple streams for prominent display

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary speaker locator module that bridges the gap between multiple video streams and the current speaker identification. This intermediary component analyzes audio-visual data to determine which video stream contains the current speaker, making the detection process easier and more reliable

Inventive Principle:
Principle #24Intermediary (Mediator)

4Device complexity

If static layout arrangement is used, then system complexity is reduced, but the system cannot adapt to dynamic conferencing needs

Engineering Contradiction:
Improvelayout managementVSAvoiddynamic arrangement capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system implements dynamic layout arrangement where the layout is continuously adjusted based on real-time conferencing conditions. The layout manager receives video streams with attributes, identifies current speakers, and dynamically generates layouts that adapt to changing speaker positions and endpoint configurations, making the system flexible and responsive

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of layout arrangement from static to dynamic. The system automatically modifies layout parameters such as video stream positioning, scaling, and visibility based on real-time detection of current speakers and endpoint configurations, enabling adaptive layout management

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2487903B1Automatic video layouts for multi-stream multi-site telepresence conferencing system
Publication Date: 2020.01.22 POLYCOM INC
  • EP2487903B1 patent drawingFigure 1
  • EP2487903B1 patent drawingFigure 2
  • EP2487903B1 patent drawingFigure 3

AI summary

A videoconference multipoint control unit, MCU, (106) automatically generates display layouts for videoconference endpoints (101-103). Display layouts are generated based on attributes associated with video streams (315, 316) received from the endpoints (102, 103) and display configuration information (329) of the endpoints (101-103). An endpoint (101-103) can include one or more attributes in each outgoing stream. Attributes can be assigned based on video streams' role, content, camera source, etc. Display layouts can be regenerated if one or more attributes change. A mixer (303) can generate video streams to be displayed at the endpoints (101-103) based on the display layout.