Multiview Tiled Volumetric Video Signaling for Selective View Delivery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for compressing and signaling multiview tiled volumetric video are inefficient, requiring specialized hardware and consuming large bandwidth due to the large data size of point clouds and meshes, which are often uncompressed.
Innovation Solution
The method involves formatting multiview video into subpictures, assigning view numbers to images based on camera viewpoints, mapping these to subpicture identifiers, and combining images into video frames for efficient compression and transmission, with higher resolution views closer to the user's angle and lower resolution for farther views.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If point clouds and meshes are transmitted uncompressed, then the quality of volumetric video content is maintained, but the bandwidth consumption increases significantly
Solution Approach 1:
The patent segments the volumetric video content into multiple views, where each view is further divided into tiles. This segmentation allows selective transmission of only the necessary views and tiles based on user viewpoint, reducing the total data quantity transmitted while maintaining quality for the visible regions.
Solution Approach 2:
The patent applies different quality levels to different regions of the content based on their importance. Views closer to the user's current viewpoint are transmitted at higher resolution, while distant views use lower resolution. This local quality differentiation reduces overall bandwidth consumption while preserving the quality of critical regions.
2Productivity
If specialized hardware is used for decoding volumetric video, then the decoding performance is improved, but the device complexity increases
Solution Approach 1:
The patent designs the encoding and signaling system to be compatible with standard video decoding hardware. By using conventional video coding techniques and standard signaling mechanisms, the system can leverage existing universal video decoders in smartphones and HMDs, eliminating the need for specialized volumetric video decoding hardware.
Solution Approach 2:
The patent extracts and transmits only the essential viewing information (specific views and tiles) rather than the complete volumetric dataset. This extraction approach allows standard decoders to process the reduced data stream efficiently without requiring specialized hardware designed for full volumetric video decoding.
3Measurement precision
If all views are transmitted at high resolution, then the immersive experience quality is maintained, but the data size increases significantly
Solution Approach 1:
The patent applies differential resolution across different views based on their relevance to the user's viewpoint. Views that are closer to the center of the user's field of view are transmitted at high resolution, while peripheral views use lower resolution. This local quality approach maintains immersive experience quality for critical regions while reducing overall data size.
Solution Approach 2:
The patent divides each view into multiple tiles and selectively transmits only the tiles that fall within or near the user's current viewpoint. This segmentation and selective transmission strategy ensures high quality for visible regions while eliminating unnecessary data from occluded or distant regions, significantly reducing data size.
4Adaptability or versatility
If complete volumetric data is transmitted, then the versatility of content access is maintained, but the transmission time increases
Solution Approach 1:
The patent performs preliminary encoding and organization of volumetric content into views and tiles before transmission. This preliminary structuring allows the receiving device to quickly extract and display only the necessary portions based on user viewpoint, eliminating the need to wait for complete data transmission while maintaining versatile content access.
Solution Approach 2:
The patent extracts and transmits only the essential viewing information (specific views and tiles) rather than the complete volumetric dataset. This extraction approach reduces transmission time significantly while the structured signaling maintains adaptability, allowing users to access different views by requesting additional tiles as needed.
Data Source
AI summary
An apparatus includes a communication interface configured to receive a bitstream for a compressed video and a processor operably coupled to the communication interface. The processor is configured to decode the bitstream for the compressed video. The processor is also configured to identify a mapping of view numbers of a plurality of images and a plurality of subpicture identifiers, each of the plurality of subpicture identifiers associated with a defined location in a video frame, wherein the mapping is signaled in the bitstream, and wherein each one of the view numbers is assigned to one image of the plurality of images based on a corresponding one of a plurality of camera viewpoints of a scene. The processor is also configured to instruct a display of at least one image based on at least one of the plurality of images.


