Composite Video Stream Rendering for Immersive Multi-Party Conferences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing technologies limit the interactive and immersive experience by presenting remote participants in separate, pre-defined graphical windows, lacking the ability to create the illusion of them being physically present in the same visual scene.

Innovation Solution

A method that captures video streams from multiple cameras, identifies and segments human subjects using computer vision, and overlays these segments onto a composite video stream to create the appearance of participants being in the same scene, allowing for interactive and immersive participation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If separate pre-defined graphical windows are used to display remote and local camera streams, then the video conferencing system is simple to implement, but the interactive and immersive experience is limited

Engineering Contradiction:
Improveease of implementationVSAvoidinteractive and immersive experience
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent merges multiple separate video streams (local and remote camera streams) into a single composite video stream. This is achieved by capturing video from multiple cameras, identifying human subjects in each stream, and compositing them together so that remote participants appear to be physically present in the same visual scene as local participants, thereby creating an immersive experience while maintaining implementation simplicity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an intermediary processing stage that receives separate video streams, processes them through computer vision algorithms to identify and segment human subjects, and then composites these segments into a unified scene. This intermediary layer enables the transformation from separate windows to a shared visual environment without fundamentally redesigning the entire video conferencing architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If computer vision algorithms are used to identify and segment human subjects, then the composite video stream creates the illusion of physical presence, but the processing complexity increases

Engineering Contradiction:
Improvecomposite video rendering capabilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the composite video rendering process into distinct stages: capturing individual video streams, identifying human subjects within each stream, segmenting the human subjects from their backgrounds, and finally compositing these segmented elements into a unified scene. This segmentation approach enables complex visual effects while managing processing complexity through modular implementation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial action by focusing computer vision processing only on identifying and segmenting human subjects rather than processing entire video frames. This selective approach applies computational resources only where needed (on human subject regions) rather than uniformly across all video data, reducing overall processing complexity while achieving the desired composite effect

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If multiple video streams are processed and combined in real-time, then the immersive experience is enhanced, but the computational resources required increase

Engineering Contradiction:
Improvereal-time composite renderingVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent implements partial action by applying intensive computer vision processing only to portions of video data that contain human subjects rather than processing all video frames uniformly. The system identifies human subjects and applies segmentation and compositing operations only to these identified regions, leaving other portions of the video stream to be processed with minimal overhead, thereby reducing overall computational resource consumption while maintaining real-time immersive rendering

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10609332B1Video conferencing supporting a composite video stream
Publication Date: 2020.03.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10609332B1 patent drawing
  • US10609332B1 patent drawing
  • US10609332B1 patent drawing

AI summary

According to a disclosed example, a first video stream is captured via a first camera associated with a first communication device engaged in a multi-party video conference. The first video stream includes a plurality of two-dimensional image frames. A subset of pixels corresponding to a first human subject is identified within each image frame of the first video stream. A second video stream is captured via a second camera associated with a second communication device engaged in the multi-party video conference. A composite video stream formed by at least a portion of the second video stream and the subset of pixels of the first video stream is rendered, and the composite video stream is output for display at one or more of the first and/or second communication devices. The composite video stream may provide the appearance of remotely located participants being physically present within the same visual scene.