Videotelephony Parallax on Monoscopic Displays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Videotelephony systems using traditional monoscopic displays lack the head-motion parallax effect, resulting in unnatural and less optimal user experiences due to the absence of parallax and stereoscopic effects.
Innovation Solution
The system employs at least one sender-side device with multiple cameras capturing video streams from different perspectives and a receiver-side device that uses either image-based rendering or model-based methods to generate output images based on the viewer's viewpoint, creating a head-motion parallax effect by determining pixel values through weighted averages or ray-casting techniques, even on traditional monoscopic displays.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional monoscopic displays are used for videotelephony, then device complexity is reduced and ease of manufacture is improved, but the user experience deteriorates due to lack of head-motion parallax effect
Solution Approach 1:
The patent applies dimensionality change by transitioning from traditional 2D monoscopic video to multi-perspective 3D video capture and rendering. Multiple cameras are arranged in three-dimensional space to capture video from different viewpoints, and the system renders these perspectives dynamically based on detected viewer head position, creating a head-motion parallax effect that adds depth and realism to the videotelephony experience.
Solution Approach 2:
The patent replaces the mechanical limitation of fixed single-camera videotelephony with a computational system that uses multiple cameras and algorithms to synthesize dynamic perspectives. Instead of physically moving a single camera, the system uses image-based rendering and model-based methods to generate views that correspond to different head positions, substituting mechanical movement with computational generation.
2Ease of operation
If multiple cameras are used to capture video from different perspectives, then head-motion parallax effect is achieved, but device complexity increases
Solution Approach 1:
The patent applies universality by designing a system where multiple cameras serve multiple functions: they simultaneously capture video from different perspectives for both standard video conferencing and immersive 3D experiences. The same camera array supports traditional monoscopic displays as well as advanced head-motion parallax effects, making the system versatile and adaptable to different display capabilities.
Solution Approach 2:
The patent uses copying by creating virtual representations of the physical camera setup through computational models. Image-based rendering methods generate synthetic video perspectives by warping and blending images from physical cameras, while model-based methods create virtual 3D models of the scene that can be viewed from any angle. These computational copies allow the system to simulate perspectives without adding physical cameras for every possible viewpoint.
3Ease of operation
If image-based rendering or model-based methods are used to generate output images, then head-motion parallax effect is achieved on monoscopic displays, but processing complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-processing video content during capture and initial rendering stages. The system pre-aligns and calibrates multiple camera feeds, pre-generates perspective transformations, and pre-computes blending weights for image-based rendering. This preliminary processing reduces the computational burden during real-time playback, allowing the system to generate head-motion parallax effects with lower latency and processing requirements.
Solution Approach 2:
The patent uses intermediaries by introducing computational layers between the physical cameras and the final display output. Image-based rendering acts as an intermediary that warps and blends camera images to simulate different viewpoints. Model-based methods introduce 3D scene reconstruction as an intermediary step that creates virtual models from which any perspective can be generated. These intermediary processing stages decouple the physical camera setup from the final visual output, enabling flexible perspective generation.
Data Source
AI summary
In one embodiment, a computing system may receive, from a second computing system, video streams of a scene, the video streams including at least a first image and a second image that are simultaneously captured by a first camera and a second camera of the second computing system, respectively. The system may determine, using a sensor system, a viewpoint of a viewer with respect to a display region of a monoscopic display associated with the first computing system. The system may generate an output image of the scene by blending, according to blending proportions computed using the viewpoint of the viewer, corresponding portions of the first image and the second image. The system may display the output image in the display region of the monoscopic display.


