Multi-Camera Video Compositing for Consistent Participant Visibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing systems face challenges in maintaining consistent image capture and display quality when participants move within the environment or when multiple participants are present, leading to obstructed views and limited face and expression discernment.
Innovation Solution
A video capture system employing multiple RGB and depth cameras positioned behind a display screen to capture images, selecting a foreground and background camera for each participant, and generating composite images that maintain participant visibility while minimizing obstruction, using a controller to process and transmit these images to remote systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a single camera is used to capture video conferencing participants, then the device complexity is low, but the reliability of capturing consistent views of all participants deteriorates when participants move or are in close proximity
Solution Approach 1:
The system divides the video capture function into multiple specialized cameras: foreground cameras for capturing close-up participant views and background cameras for capturing environmental context. This segmentation allows each camera type to optimize for its specific function, improving overall capture reliability without requiring a single complex camera system
Solution Approach 2:
The system merges multiple camera feeds (foreground and background) into a single composite video stream through image compositing. This combining approach integrates the strengths of different camera perspectives to deliver consistent participant views while maintaining system manageability through unified output
2Reliability
If multiple cameras are deployed to capture all participants reliably, then the reliability of video capture improves, but the device complexity increases
Solution Approach 1:
The system assigns different quality characteristics to different camera types: foreground cameras use higher resolution for detailed participant features while background cameras use lower resolution for contextual information. This local quality differentiation improves overall capture reliability while reducing total system complexity by optimizing each component for its specific role
Solution Approach 2:
The system extracts and processes foreground and background elements separately through image compositing techniques. By separating participant extraction from environmental capture, the system achieves reliable participant visibility while simplifying the processing pipeline compared to using a single high-complexity camera system
3Ease of operation
If participants are allowed to move freely in the environment, then the ease of operation improves, but the quality of face and expression discernment deteriorates due to obstructions
Solution Approach 1:
The system dynamically adjusts the composite image generation based on real-time participant positions detected by depth cameras. As participants move, the system automatically recalculates foreground-background compositing parameters to maintain optimal face visibility, allowing free movement while preserving measurement precision for facial features
Solution Approach 2:
The system uses depth camera feedback to continuously monitor participant positions and adjust the compositing algorithm accordingly. This feedback loop ensures that even when participants move freely or are in close proximity, the composite image maintains optimal face and expression discernment by dynamically optimizing camera selection and composition parameters
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Techniques for video capture including determining a position of a subject in relation to multiple cameras; selecting a foreground camera from the cameras based on at least the determined position; obtaining an RGB image captured by the foreground camera; segmenting the RGB image to identify a foreground portion corresponding to the subject, with a total height of the foreground portion being a first percentage of a total height of the RGB image; generating a foreground image from the foreground portion; producing a composite image, including compositing the foreground image and a background image to produce a portion of the composite image, with a total height of the foreground image in the composite image being a second percentage of a total height of the composite image and the second percentage being substantially less than the first percentage; and causing the composite image to be displayed on a remote system.