Adaptive Video Layout Composition for Multipoint Conferences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems produce disorganized and disengaging visual layouts, failing to emulate the natural interaction and physical presence of in-person meetings, despite attempts to create dynamic compositions.
Innovation Solution
A method employing Pan Zoom Tilt (PZT) processes and face detection, combined with weighted presence rulesets and composition planes, to adaptively recompose video streams, focusing on detected faces and participant counts, and adjusting layout based on context and historical data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If simple rules are used to compose video layouts (e.g., placing pictures of last two speakers side-by-side), then the layout composition process is simple and fast, but the visual layout becomes disorganized and disengaging
Solution Approach 1:
The patent implements dynamic layout composition by continuously monitoring audio activity, face detection, and participant presence to automatically adjust video stream arrangements in real-time. The system transitions from static predefined layouts to dynamic adaptive layouts that respond to conference dynamics, resolving the contradiction between simple composition rules and engaging visual presentations.
Solution Approach 2:
The system incorporates feedback mechanisms through audio activity detection, face detection algorithms, and participant presence monitoring that continuously inform layout composition decisions. This feedback loop enables the system to adapt layouts based on actual conference conditions, maintaining both simplicity in operation and high visual engagement.
2Device complexity
If audio-only rules are used to recalculate display order, then the system is simple to implement, but the layout fails to emulate physical presence and feels disorganized
Solution Approach 1:
The patent integrates multiple functions into a unified layout composition system that simultaneously processes audio activity, video face detection, participant presence, and contextual information. This multi-functional approach enables the system to emulate physical presence effectively while maintaining manageable implementation complexity through coordinated operation of these functions.
Solution Approach 2:
The system creates a composite video stream that combines multiple video feeds, audio streams, and visual elements (such as shared screens and participant indicators) into a unified presentation. This composite approach enriches the visual experience and better emulates physical presence compared to audio-only control methods.
3Ease of operation
If continuous presence conferences mix multiple video signals into a single composite stream, then the visual layout can be dynamic and engaging, but the system complexity increases significantly
Solution Approach 1:
The patent segments the complex video conferencing system into distinct functional modules: audio activity detection module, face detection module, participant presence detection module, and layout composition module. This segmentation reduces overall system complexity by allowing each module to perform its specific function independently while contributing to the unified layout composition.
Data Source
AI summary
The present invention creates compositions of pictures in multipoint conferences that emulate natural interaction and existing aesthetic sensibilities learned from visual media by a combination of correcting and adapting the composition of the picture content and the layout, preferably in the MCN of the conference, where real-time conference data is available, in addition to statistics and knowledge of historical conference data. Further, cross checking incoming imagery against a ruleset where compositional deltas are identified is done, and these corrective transformations are applied, and the resulting corrections and remixes are applied to the layout. More advanced transformations to the final composition based on presence and context define a layout. The ruleset could be both static and dynamic, or a combination, and the final recomposition of the layout may be a result of both corrective and adaptive transformations.


