Truncated Square Pyramid Mapping for 360-Degree Video Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual reality video systems face challenges in efficiently storing and transmitting high-quality 360-degree video data due to the large amount of data required, which can exceed the capabilities of storage and transmission devices, and existing methods often result in noticeable quality gaps when reducing data resolution.
Innovation Solution
Mapping 360-degree video data onto a truncated square pyramid shape, allowing for a reduced data size while maintaining high resolution at the viewer's front view and decreasing resolution towards the back view, with a frame packing structure that stores data in a rectangular format for easier storage and transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 360-degree video data is stored at full resolution across all directions, then video quality is maintained, but data size becomes excessively large for storage and transmission
Solution Approach 1:
The patent applies local quality by assigning different resolutions to different spatial regions of the 360-degree video. The front view (base plane) maintains full resolution while the back view (top plane) and intermediate views (side planes) use progressively lower resolutions. This resolves the contradiction by preserving video quality where viewers most need it (front view) while reducing data size in less critical regions (back and side views).
Solution Approach 2:
The patent uses an asymmetric truncated square pyramid geometry instead of a symmetric cube or sphere representation. The base plane (front view) is larger and maintains higher resolution, while the top plane (back view) is smaller with lower resolution. This asymmetric structure directly addresses the contradiction by creating an imbalance between quality and data size that favors quality in the most important viewing direction while accepting lower quality elsewhere to reduce overall data size.
2Productivity
If video data is compressed to reduce data size, then storage and transmission efficiency improves, but quality gaps and fidelity loss occur
Solution Approach 1:
The patent implements local quality by applying different compression levels to different regions of the video data. The base plane (front view) undergoes minimal compression to preserve high fidelity, while the top and side planes apply progressively stronger compression. This resolves the contradiction by maintaining video fidelity where it matters most (front view for storage efficiency) while accepting controlled fidelity loss in less important regions to achieve overall compression goals.
3Ease of manufacture
If uniform resolution is applied to all views, then simplicity of processing is maintained, but inefficiency in data usage occurs
Solution Approach 1:
The patent applies local quality by implementing a hierarchical resolution structure where the base plane uses full resolution, side planes use intermediate resolution, and the top plane uses reduced resolution. This resolves the contradiction by maintaining processing simplicity through a structured approach while dramatically improving data efficiency compared to uniform full-resolution processing of all views.
4Shape
If non-rectangular data blocks are used to represent spherical video, then accurate spherical geometry is preserved, but storage and transmission complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the spherical video data into six separate planar faces (one base, one top, and four side planes) that form a truncated square pyramid. Each face can be independently processed and stored as a rectangular data block, resolving the contradiction by maintaining spherical representation accuracy through the polyhedral structure while eliminating storage complexity by using standard rectangular blocks for each segment.
Solution Approach 2:
The patent uses dimensionality change by transforming the three-dimensional spherical video data into a two-dimensional truncated square pyramid representation with six planar faces. This resolves the contradiction by preserving spherical geometry accuracy through the 3D polyhedral structure while enabling efficient 2D rectangular storage and transmission of each individual face, effectively bridging the gap between spherical accuracy and rectangular efficiency.
Data Source
Figure 1
Figure 2A~2B
Figure 3
AI summary
Techniques and systems are described for mapping 360-degree video data to a truncated square pyramid shape. A 360-degree video frame can include 360-degrees worth of pixel data, and thus be spherical in shape. By mapping the spherical video data to the planes provided by a truncated square pyramid, the total size of the 360-degree video frame can be reduced. The planes of the truncated square pyramid can be oriented such that the base of the truncated square pyramid represents a front view and the top of the truncated square pyramid represents a back view. In this way, the front view can be captured at full resolution, the back view can be captured at reduced resolution, and the left, right, up, and bottom views can be captured at decreasing resolutions. Frame packing structures can also be defined for 360-degree video data that has been mapped to a truncated square pyramid shape.