Multi-Camera Video Compositing for Consistent Participant Visibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video conferencing systems face challenges in maintaining consistent image capture and display quality when participants move within the environment or when multiple participants are present, leading to obstructed views and limited face and expression discernment.

Innovation Solution

A video capture system employing multiple RGB and depth cameras positioned behind a display screen to capture images, selecting a foreground and background camera for each participant, and generating composite images that maintain participant visibility while minimizing obstruction, using a controller to process and transmit these images to remote systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a single camera is used to capture video conferencing participants, then the device complexity is low, but the reliability of capturing consistent views of all participants deteriorates when participants move or are in close proximity

Engineering Contradiction:
Improvereliability of capturing consistent viewsVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system divides the video capture function into multiple specialized cameras: foreground cameras for capturing close-up participant views and background cameras for capturing environmental context. This segmentation allows each camera type to optimize for its specific function, improving overall capture reliability without requiring a single complex camera system

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system merges multiple camera feeds (foreground and background) into a single composite video stream through image compositing. This combining approach integrates the strengths of different camera perspectives to deliver consistent participant views while maintaining system manageability through unified output

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If multiple cameras are deployed to capture all participants reliably, then the reliability of video capture improves, but the device complexity increases

Engineering Contradiction:
Improveconsistency of participant visibilityVSAvoidcamera system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system assigns different quality characteristics to different camera types: foreground cameras use higher resolution for detailed participant features while background cameras use lower resolution for contextual information. This local quality differentiation improves overall capture reliability while reducing total system complexity by optimizing each component for its specific role

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system extracts and processes foreground and background elements separately through image compositing techniques. By separating participant extraction from environmental capture, the system achieves reliable participant visibility while simplifying the processing pipeline compared to using a single high-complexity camera system

Inventive Principle:
Principle #2Taking out (Extraction)

3Ease of operation

If participants are allowed to move freely in the environment, then the ease of operation improves, but the quality of face and expression discernment deteriorates due to obstructions

Engineering Contradiction:
Improveparticipant movement freedomVSAvoidface and expression discernment
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts the composite image generation based on real-time participant positions detected by depth cameras. As participants move, the system automatically recalculates foreground-background compositing parameters to maintain optimal face visibility, allowing free movement while preserving measurement precision for facial features

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses depth camera feedback to continuously monitor participant positions and adjust the compositing algorithm accordingly. This feedback loop ensures that even when participants move freely or are in close proximity, the composite image maintains optimal face and expression discernment by dynamically optimizing camera selection and composition parameters

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3721617B1Video capture systems and methods
Publication Date: 2026.01.28 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3721617B1 patent drawingFigure 1
  • EP3721617B1 patent drawingFigure 2
  • EP3721617B1 patent drawingFigure 3

AI summary

Techniques for video capture including determining a position of a subject in relation to multiple cameras; selecting a foreground camera from the cameras based on at least the determined position; obtaining an RGB image captured by the foreground camera; segmenting the RGB image to identify a foreground portion corresponding to the subject, with a total height of the foreground portion being a first percentage of a total height of the RGB image; generating a foreground image from the foreground portion; producing a composite image, including compositing the foreground image and a background image to produce a portion of the composite image, with a total height of the foreground image in the composite image being a second percentage of a total height of the composite image and the second percentage being substantially less than the first percentage; and causing the composite image to be displayed on a remote system.