3D Video Conferencing via Server-Side Image Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conferencing systems face challenges in providing a realistic three-dimensional imaging experience while managing bandwidth effectively, as they require expensive equipment and result in excessive data transmission and jittery displays due to the need for multiple cameras and projectors.

Innovation Solution

A method that synthesizes image data from multiple cameras to deliver a three-dimensional rendering based on the user's position, using a server network to manage media streams and delay audio accordingly, allowing for a 3D experience on a standard personal computer with reduced bandwidth consumption by selecting and transmitting only the necessary video stream.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple cameras and projectors are used to provide three-dimensional imaging, then the realism of the video conferencing experience is improved, but the bandwidth consumption and equipment cost increase significantly

Engineering Contradiction:
Improverealism of video conferencing experienceVSAvoidbandwidth consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent creates a virtual copy of the three-dimensional imaging experience by synthesizing depth information from two-dimensional video feeds using software algorithms. Instead of requiring multiple physical cameras and projectors, the system generates a virtual 3D model that can be viewed from different angles, effectively copying the visual experience without the physical infrastructure.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical system of multiple physical cameras and projectors with a software-based image synthesis system. The depth acquisition module and view synthesis module use computational algorithms to generate three-dimensional effects from standard video feeds, substituting hardware complexity with software processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If multiple cameras and projectors are deployed for three-dimensional imaging, then the visual quality is improved, but the device complexity and cost increase

Engineering Contradiction:
Improvevisual qualityVSAvoidequipment complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent makes a single standard video camera perform multiple functions by using it as both a depth reference and a visual source. The system processes the video feed to extract depth information and simultaneously uses it for rendering, allowing one device to fulfill the role of multiple specialized components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system creates virtual copies of the imaging functionality through software synthesis. Instead of requiring multiple physical devices, the image synthesis module generates virtual views from a single or limited number of camera feeds, copying the visual output that would otherwise require multiple physical projectors.

Inventive Principle:
Principle #26Copying

3Productivity

If real-time three-dimensional rendering is provided, then the user experience is improved, but the processing time and audio-video synchronization become challenging

Engineering Contradiction:
Improveuser experience qualityVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary depth extraction and scene reconstruction from video feeds before the actual view synthesis is needed. By pre-processing the video data to create depth maps and three-dimensional scene representations, the system reduces the computational burden during real-time rendering, allowing faster response when users change viewing angles.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms where the synthesized view is continuously compared with the original video feeds, and adjustments are made to maintain lip synchronization and visual consistency. The audio-video synchronization is maintained through feedback loops that adjust rendering timing based on detected discrepancies.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP2406951B1System and method for providing three dimensional imaging in a network environment
Publication Date: 2018.07.11 CISCO TECHNOLOGY INC
  • EP2406951B1 patent drawingFigure 1
  • EP2406951B1 patent drawingFigure 2
  • EP2406951B1 patent drawingFigure 3

AI summary

A method is provided in one example embodiment and includes receiving data indicative of a personal position of an end user and receiving image data associated with an object. The image data can be captured by a first camera at a first angle and a second camera at a second angle. The method also includes synthesizing the image data in order to deliver a three-dimensional rendering of the object at a selected angle, which is based on the data indicative of the personal position of the end user. In more specific embodiments, the synthesizing is executed by a server configured to be coupled to a network. Video analytics can be used to determine the personal position of the end user. In other embodiments, the method includes determining an approximate time interval for the synthesizing of the image data and then delaying audio data based on the time interval.