Composite Video Stream Rendering for Immersive Multi-Party Conferences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing technologies limit the interactive and immersive experience by presenting remote participants in separate, pre-defined graphical windows, lacking the ability to create the illusion of them being physically present in the same visual scene.
Innovation Solution
A method that captures video streams from multiple cameras, identifies and segments human subjects using computer vision, and overlays these segments onto a composite video stream to create the appearance of participants being in the same scene, allowing for interactive and immersive participation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If separate pre-defined graphical windows are used to display remote and local camera streams, then the video conferencing system is simple to implement, but the interactive and immersive experience is limited
Solution Approach 1:
The patent merges multiple separate video streams (local and remote camera streams) into a single composite video stream. This is achieved by capturing video from multiple cameras, identifying human subjects in each stream, and compositing them together so that remote participants appear to be physically present in the same visual scene as local participants, thereby creating an immersive experience while maintaining implementation simplicity
Solution Approach 2:
The system introduces an intermediary processing stage that receives separate video streams, processes them through computer vision algorithms to identify and segment human subjects, and then composites these segments into a unified scene. This intermediary layer enables the transformation from separate windows to a shared visual environment without fundamentally redesigning the entire video conferencing architecture
2Adaptability or versatility
If computer vision algorithms are used to identify and segment human subjects, then the composite video stream creates the illusion of physical presence, but the processing complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the composite video rendering process into distinct stages: capturing individual video streams, identifying human subjects within each stream, segmenting the human subjects from their backgrounds, and finally compositing these segmented elements into a unified scene. This segmentation approach enables complex visual effects while managing processing complexity through modular implementation
Solution Approach 2:
The system applies partial action by focusing computer vision processing only on identifying and segmenting human subjects rather than processing entire video frames. This selective approach applies computational resources only where needed (on human subject regions) rather than uniformly across all video data, reducing overall processing complexity while achieving the desired composite effect
3Adaptability or versatility
If multiple video streams are processed and combined in real-time, then the immersive experience is enhanced, but the computational resources required increase
Solution Approach 1:
The patent implements partial action by applying intensive computer vision processing only to portions of video data that contain human subjects rather than processing all video frames uniformly. The system identifies human subjects and applies segmentation and compositing operations only to these identified regions, leaving other portions of the video stream to be processed with minimal overhead, thereby reducing overall computational resource consumption while maintaining real-time immersive rendering
Data Source
AI summary
According to a disclosed example, a first video stream is captured via a first camera associated with a first communication device engaged in a multi-party video conference. The first video stream includes a plurality of two-dimensional image frames. A subset of pixels corresponding to a first human subject is identified within each image frame of the first video stream. A second video stream is captured via a second camera associated with a second communication device engaged in the multi-party video conference. A composite video stream formed by at least a portion of the second video stream and the subset of pixels of the first video stream is rendered, and the composite video stream is output for display at one or more of the first and/or second communication devices. The composite video stream may provide the appearance of remotely located participants being physically present within the same visual scene.


