Adaptive Video Layout Composition for Multipoint Conferences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing systems produce disorganized and disengaging visual layouts, failing to emulate the natural interaction and physical presence of in-person meetings, despite attempts to create dynamic compositions.

Innovation Solution

A method employing Pan Zoom Tilt (PZT) processes and face detection, combined with weighted presence rulesets and composition planes, to adaptively recompose video streams, focusing on detected faces and participant counts, and adjusting layout based on context and historical data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If simple rules are used to compose video layouts (e.g., placing pictures of last two speakers side-by-side), then the layout composition process is simple and fast, but the visual layout becomes disorganized and disengaging

Engineering Contradiction:
Improvelayout composition simplicityVSAvoidvisual engagement and intuitiveness
Core Design Contradiction:
Ease of manufactureVSEase of operation

Solution Approach 1:

The patent implements dynamic layout composition by continuously monitoring audio activity, face detection, and participant presence to automatically adjust video stream arrangements in real-time. The system transitions from static predefined layouts to dynamic adaptive layouts that respond to conference dynamics, resolving the contradiction between simple composition rules and engaging visual presentations.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms through audio activity detection, face detection algorithms, and participant presence monitoring that continuously inform layout composition decisions. This feedback loop enables the system to adapt layouts based on actual conference conditions, maintaining both simplicity in operation and high visual engagement.

Inventive Principle:
Principle #23Feedback

2Device complexity

If audio-only rules are used to recalculate display order, then the system is simple to implement, but the layout fails to emulate physical presence and feels disorganized

Engineering Contradiction:
Improvesystem implementation complexityVSAvoidphysical presence emulation capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent integrates multiple functions into a unified layout composition system that simultaneously processes audio activity, video face detection, participant presence, and contextual information. This multi-functional approach enables the system to emulate physical presence effectively while maintaining manageable implementation complexity through coordinated operation of these functions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system creates a composite video stream that combines multiple video feeds, audio streams, and visual elements (such as shared screens and participant indicators) into a unified presentation. This composite approach enriches the visual experience and better emulates physical presence compared to audio-only control methods.

Inventive Principle:
Principle #40Composite materials

3Ease of operation

If continuous presence conferences mix multiple video signals into a single composite stream, then the visual layout can be dynamic and engaging, but the system complexity increases significantly

Engineering Contradiction:
Improvevisual engagement and intuitivenessVSAvoidsignal processing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the complex video conferencing system into distinct functional modules: audio activity detection module, face detection module, participant presence detection module, and layout composition module. This segmentation reduces overall system complexity by allowing each module to perform its specific function independently while contributing to the unified layout composition.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10972702B2Intelligent adaptive and corrective layout composition
Publication Date: 2021.04.06 PEXIP
  • US10972702B2 patent drawing
  • US10972702B2 patent drawing
  • US10972702B2 patent drawing

AI summary

The present invention creates compositions of pictures in multipoint conferences that emulate natural interaction and existing aesthetic sensibilities learned from visual media by a combination of correcting and adapting the composition of the picture content and the layout, preferably in the MCN of the conference, where real-time conference data is available, in addition to statistics and knowledge of historical conference data. Further, cross checking incoming imagery against a ruleset where compositional deltas are identified is done, and these corrective transformations are applied, and the resulting corrections and remixes are applied to the layout. More advanced transformations to the final composition based on presence and context define a layout. The ruleset could be both static and dynamic, or a combination, and the final recomposition of the layout may be a result of both corrective and adaptive transformations.