Multiplane Video Layering for Immersive 3D Videoconferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current videoconferencing systems lack the ability to create immersive, three-dimensional-like experiences on two-dimensional screens, limiting user interaction and engagement, as they do not allow participants to move their videos within a virtual environment or customize the background and foreground layers dynamically.
Innovation Solution
A videoconferencing system that employs a multiplane camera view by transmitting multiple video streams in different layers, enabling participants to adjust and interact with these layers, including full-body segmentation and gesture recognition, to create a dynamic and immersive experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional single-plane video streaming is used, then system complexity is low, but the three-dimensional effect and user immersion are insufficient
Solution Approach 1:
The video environment is segmented into multiple independent layers (foreground layer, background layer, participant video layer) that can be transmitted and processed separately. Each layer contains specific visual elements that can be independently controlled and customized by users, enabling three-dimensional spatial arrangement without requiring complete system redesign.
Solution Approach 2:
The system transitions from traditional two-dimensional video display to multi-dimensional spatial arrangement by adding depth perception through layered composition. Users can adjust the Z-axis positioning of different video layers, creating a three-dimensional virtual environment that enhances immersion while maintaining compatibility with standard 2D display devices.
2Adaptability or versatility
If multiple video layers are transmitted, then user customization and interaction are enhanced, but network bandwidth and processing requirements increase
Solution Approach 1:
Instead of transmitting complete high-resolution video for all layers simultaneously, the system transmits partial video data at reduced resolutions for background and foreground layers, while maintaining higher quality for participant videos. Users can dynamically adjust the detail level of each layer based on their needs, reducing overall bandwidth consumption while preserving essential customization capabilities.
3Ease of operation
If full-body segmentation and gesture recognition are implemented, then natural interaction is improved, but computational requirements and processing time increase
Solution Approach 1:
The system automatically performs full-body segmentation and gesture recognition without requiring manual user input or configuration. AI algorithms continuously analyze video feeds to identify body parts and gestures, automatically adjusting layer compositions and interactions based on detected user actions, thereby simplifying operation while managing processing through efficient automated pipelines.
Data Source
AI summary
A user interface display is provided to a participant who is viewing an event via a communication system that provides videoconferencing. The communication system provides a composite video stream that includes a plurality of different video layers, each video layer providing a different portion of the composite video stream. The participant has a participant computer for allowing the participant to receive the composite video stream for display on the user interface display. A plurality of participants view the event via user interface displays of their respective participant computers. The layers include a participant layer that displays video streams of the participants, a foreground layer, an event layer that includes video of the event, an audience layer, a background layer, and an immersive layer.


