Multi-Camera Eye Contact Synthesis for Video Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional video conferencing systems face an 'eye contact problem' due to cameras being mounted above or below the display, leading to an impression that the user is not looking directly at the other person, as the camera is not positioned to simulate eye contact with the image on the screen.
Innovation Solution
A multi-camera setup with cameras placed on the upper and lower edges of the display, capable of synthesizing a view from a virtual camera location within the display panel, allowing for real-time processing and transmission of images to simulate eye contact by warping and combining raw images to generate a synthesized view from the perspective of a virtual camera.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a single camera is mounted above or below the display, then the device structure is simple, but the eye contact problem occurs because the camera cannot simulate the perspective of looking directly at the other person
Solution Approach 1:
The patent uses multiple physical cameras to capture images from different positions, then synthesizes a virtual camera view that copies the perspective of a camera positioned within the display. This virtual camera copy creates the eye contact effect without requiring an actual physical camera in the display location.
Solution Approach 2:
The patent transitions from a single-dimensional camera position (above or below display) to a multi-dimensional virtual camera position (within the display plane). By synthesizing views from multiple camera angles and combining them computationally, the system creates a virtual perspective that exists in a different spatial dimension than any single physical camera.
2Ease of operation
If multiple cameras are placed on upper and lower edges of the display to synthesize a virtual camera view, then the eye contact problem is solved, but the device complexity increases
Solution Approach 1:
The patent divides the viewing task across multiple cameras positioned at different locations (upper and lower edges of the display). Each camera captures a portion of the scene from its specific position, and the image processor segments and combines these individual views to create the complete synthesized perspective.
Solution Approach 2:
The image processor acts as an intermediary that receives raw images from multiple physical cameras, processes them through warping and combining operations, and generates the final synthesized view. This intermediary computational process bridges the gap between the physical camera array and the desired virtual camera perspective.
3Reliability
If cameras are positioned to capture images for synthesizing a virtual view, then vertical occlusions are handled better, but the image processing complexity increases
Solution Approach 1:
The patent employs dynamic image processing that adapts to different scene configurations and occlusion patterns. The warping and combining operations are performed in real-time based on the relative positions of objects in the scene, allowing the system to dynamically handle vertical occlusions that would be problematic for fixed camera positions.
Data Source
AI summary
A video transmitting system includes: a display configured to display an image in a first direction; cameras including: a first camera adjacent a first edge of the display; and a second camera adjacent a second edge of the display, at least a portion of the display being proximal a convex hull that includes the first camera and the second camera, the first camera and the second camera having substantially overlapping fields of view encompassing the first direction; and an image processor to: receive a position of a virtual camera relative to the cameras substantially within the convex hull and substantially on the display, the virtual camera having a field of view encompassing the first direction; receive raw images captured by the cameras at substantially the same time; and generate processed image data from the raw images for synthesizing a view in accordance with the position of the virtual camera.


