Packed Video Frame Layouts for Mobile Volumetric Video Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mobile electronic devices struggle to provide a high-quality immersive video experience due to the large number of bitstreams and decoder instantiations required for volumetric video data, which are typically not equipped to synchronize multiple video decoder instantiations.
Innovation Solution
A frame packing technique is employed to consolidate video components into packed video frames, reducing the number of bitstreams and decoder instantiations by encoding multiple types of video data into a single packed frame, accompanied by supplemental enhancement information to facilitate parallel decoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple separate bitstreams are used to represent different video components (geometry, texture, occupancy), then the volumetric video quality and immersion experience are improved, but the number of decoder instantiations required increases significantly, making mobile devices unable to handle the processing load
Solution Approach 1:
The patent combines multiple separate video component bitstreams (geometry, texture, occupancy) into a single packed video frame bitstream. Different video components are packed together in a unified frame structure, allowing them to be decoded by a single decoder instantiation rather than requiring multiple separate decoders. This merging approach maintains the quality of volumetric video while reducing the computational complexity burden on mobile devices.
Solution Approach 2:
The packed video frame structure serves multiple functions simultaneously: it contains geometry information, texture information, and occupancy data all in one unified structure. A single decoder is designed to handle this multi-functional bitstream, extracting and processing different video components from the same packed frame, thereby eliminating the need for multiple specialized decoders while preserving comprehensive video quality.
2Loss of information
If multiple separate bitstreams are used for different video components, then complete volumetric information is preserved, but the synchronization of multiple decoders becomes intractable on mobile devices
Solution Approach 1:
The patent merges multiple video component streams into a single packed video frame structure where geometry, texture, and occupancy data are integrated. This unified structure ensures that all volumetric information is preserved while being delivered through a single synchronized bitstream, eliminating the synchronization problems that arise when multiple separate decoders must coordinate their operations.
3Manufacturing precision
If video components are encoded separately into individual bitstreams, then each component can be optimized independently, but the overall processing efficiency and resource utilization on mobile devices deteriorates
Solution Approach 1:
The patent combines multiple video components into a single packed video frame bitstream that can be processed efficiently by mobile devices. While each component (geometry, texture, occupancy) maintains its own encoding optimizations, they are all delivered through one unified bitstream that requires only a single decoder instantiation, dramatically improving processing efficiency and resource utilization on mobile platforms.
Data Source
AI summary
Methods, apparatus, systems and articles of manufacture to generate packed video frames are disclosed. A video encoding system disclosed herein includes a configuration determiner to create a packed video frame layout that includes regions into which video components are to be placed. The system also includes a frame generator to form packed video frames that include the video components placed into different regions. The encoding system further includes a frame information generator that generates packed video frame information that identifies characteristics of the packed video frame including (i) the identities of regions included in the packed video frame layout, (ii) types of video components included in the regions, or iii) information identifying the locations and dimensions of the regions. A video encoder of the encoding system encodes the frames and includes the packed video frame information to signal the inclusion of the packed video frames in the encoded bitstream.


