3D Video Conferencing via Server-Side Image Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conferencing systems face challenges in providing a realistic three-dimensional imaging experience while managing bandwidth effectively, as they require expensive equipment and result in excessive data transmission and jittery displays due to the need for multiple cameras and projectors.
Innovation Solution
A method that synthesizes image data from multiple cameras to deliver a three-dimensional rendering based on the user's position, using a server network to manage media streams and delay audio accordingly, allowing for a 3D experience on a standard personal computer with reduced bandwidth consumption by selecting and transmitting only the necessary video stream.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple cameras and projectors are used to provide three-dimensional imaging, then the realism of the video conferencing experience is improved, but the bandwidth consumption and equipment cost increase significantly
Solution Approach 1:
The patent creates a virtual copy of the three-dimensional imaging experience by synthesizing depth information from two-dimensional video feeds using software algorithms. Instead of requiring multiple physical cameras and projectors, the system generates a virtual 3D model that can be viewed from different angles, effectively copying the visual experience without the physical infrastructure.
Solution Approach 2:
The patent replaces the mechanical system of multiple physical cameras and projectors with a software-based image synthesis system. The depth acquisition module and view synthesis module use computational algorithms to generate three-dimensional effects from standard video feeds, substituting hardware complexity with software processing.
2Reliability
If multiple cameras and projectors are deployed for three-dimensional imaging, then the visual quality is improved, but the device complexity and cost increase
Solution Approach 1:
The patent makes a single standard video camera perform multiple functions by using it as both a depth reference and a visual source. The system processes the video feed to extract depth information and simultaneously uses it for rendering, allowing one device to fulfill the role of multiple specialized components.
Solution Approach 2:
The system creates virtual copies of the imaging functionality through software synthesis. Instead of requiring multiple physical devices, the image synthesis module generates virtual views from a single or limited number of camera feeds, copying the visual output that would otherwise require multiple physical projectors.
3Productivity
If real-time three-dimensional rendering is provided, then the user experience is improved, but the processing time and audio-video synchronization become challenging
Solution Approach 1:
The patent performs preliminary depth extraction and scene reconstruction from video feeds before the actual view synthesis is needed. By pre-processing the video data to create depth maps and three-dimensional scene representations, the system reduces the computational burden during real-time rendering, allowing faster response when users change viewing angles.
Solution Approach 2:
The system implements feedback mechanisms where the synthesized view is continuously compared with the original video feeds, and adjustments are made to maintain lip synchronization and visual consistency. The audio-video synchronization is maintained through feedback loops that adjust rendering timing based on detected discrepancies.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method is provided in one example embodiment and includes receiving data indicative of a personal position of an end user and receiving image data associated with an object. The image data can be captured by a first camera at a first angle and a second camera at a second angle. The method also includes synthesizing the image data in order to deliver a three-dimensional rendering of the object at a selected angle, which is based on the data indicative of the personal position of the end user. In more specific embodiments, the synthesizing is executed by a server configured to be coupled to a network. Video analytics can be used to determine the personal position of the end user. In other embodiments, the method includes determining an approximate time interval for the synthesizing of the image data and then delaying audio data based on the time interval.