Realistic Point of View Video System Using Multi-Camera Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video display systems fail to provide a realistic point of view as the viewer moves, as they are limited by occlusion distances and distortions when converting two-dimensional images to three-dimensional, restricting the viewer's movement and immersion.
Innovation Solution
A system comprising sensors arranged around a scene, a video server, and a rendering device that dynamically adjusts the video presentation based on the viewer's position, using position detection and content requestors to generate a composite view by interpolating between captured views, ensuring a realistic and interactive experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple cameras are used to capture different views of a scene, then the realism and viewer mobility are improved, but the device complexity increases
Solution Approach 1:
The system divides the scene into multiple views captured by separate cameras, each responsible for a specific angular range. This segmentation allows the system to provide realistic imagery from multiple perspectives without requiring a single complex omnidirectional camera, thereby improving viewer mobility while managing system complexity through modular camera units.
Solution Approach 2:
The video server performs multiple functions: it receives feeds from multiple cameras, detects viewer position, selects appropriate camera feeds, and generates composite views. This multi-functionality consolidates what would otherwise require separate systems into a single platform, improving adaptability while controlling overall system complexity.
2Device complexity
If a single two-dimensional image is converted to three-dimensional, then the manufacturing complexity is reduced, but the image quality and realism deteriorate due to distortion and occlusion
Solution Approach 1:
Instead of attempting to create a 3D effect from a single 2D image, the system segments the scene into multiple 2D views captured from different angles. These segmented views are then combined to create a composite 3D-like experience, avoiding the severe distortion and occlusion problems inherent in single-image 3D conversion while maintaining acceptable system complexity.
Solution Approach 2:
The video server acts as an intermediary that processes multiple camera feeds and generates composite views. Rather than directly converting a single 2D image to 3D, the system uses the video server to interpolate and combine multiple intermediate 2D views, thereby achieving superior image quality and realism while keeping the overall system complexity manageable.
3Ease of operation
If the viewer moves away from the screen, then the viewing freedom is improved, but the image realism deteriorates due to occlusion distance limitations
Solution Approach 1:
The system transitions from a single fixed 2D view to a multi-dimensional approach by capturing and combining views from multiple cameras positioned at different angles and distances. This dimensional expansion allows the system to maintain image realism across multiple viewing distances and positions, thereby improving viewing freedom without sacrificing image quality.
Solution Approach 2:
The system dynamically changes parameters such as which camera feeds are selected and how they are interpolated, based on the viewer's detected position. When the viewer moves away from the screen, the system adjusts by selecting and combining feeds from cameras that provide appropriate perspective, thereby maintaining image realism across different viewing distances and enhancing viewing freedom.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Method and apparatus for obtaining and providing realistic point of view video are described. In one innovative aspect, a device for providing video is provided. The system includes a view capture circuit configured to obtain multiple views of a scene, each view having a capture position. The system includes a receiver configured to receive a request for the scene, the request including a viewing position. The system includes a view selector configured to identify one or more views of the scene based on a comparison of the viewing position and the capture position of each view. The system includes a view generator configured to generate an output view based on the identified views and the viewing position.