V-PCC Component Tracks for Independent Point Cloud Region Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding systems face challenges in efficiently compressing and decompressing three-dimensional (3D) point clouds, which require large amounts of data to realistically reconstruct objects and scenes in a 3D space.
Innovation Solution
The proposed system involves a video decoding device that processes a media container file containing a video-based point cloud compression (V-PCC) bitstream. The device identifies multiple V-PCC component tracks, each corresponding to an encoded version of a V-PCC component, and determines which tracks belong to the same playout track group for coordinated playback. Additionally, the system allows for independent decoding of tile groups, enabling access to specific regions within a 3D space without processing the entire point cloud.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple V-PCC component tracks are encoded and stored separately to enable flexible access to different regions of the point cloud, then accessibility and versatility are improved, but device complexity and data management complexity increase
Solution Approach 1:
The patent divides the point cloud data into multiple V-PCC component tracks, where each track represents a separate encoded version of point cloud components (geometry, attributes, occupancy). This segmentation enables independent access to different regions and components of the point cloud, improving versatility while maintaining manageable complexity through standardized track structures and grouping mechanisms.
2Speed
If tile groups are decoded independently to enable random access to specific regions, then access speed and efficiency are improved, but processing overhead and system complexity increase
Solution Approach 1:
The patent implements independent tile group decoding by dividing the point cloud into discrete tile groups that can be decoded separately. Each tile group contains self-contained information allowing random access to specific regions without processing the entire point cloud, thereby improving access speed while managing complexity through modular decoding architecture.
Solution Approach 2:
The patent performs preliminary encoding and organization of tile groups during the compression phase, preparing the data structure in advance to enable efficient random access. This preliminary action includes creating independent tile group units with proper synchronization and metadata, reducing the processing burden during playback or access operations.
3Adaptability or versatility
If alternative encoded versions of V-PCC components are provided with different attributes, then adaptability and quality options are improved, but data volume and storage requirements increase
Solution Approach 1:
The patent creates multiple V-PCC component tracks where each track represents an alternative encoded version of the same point cloud component with different attributes (e.g., different resolutions, compression ratios, or encoding parameters). This universal track structure allows a single set of tracks to serve multiple quality and access requirements, providing adaptability without proportionally increasing total data volume through efficient redundancy management.
Data Source
AI summary
Systems, methods, and instrumentalities are disclosed that relate to the processing of a media container file associated with 3D video data. The media container file may indicate that certain video-based point cloud compression (V-PCC) component tracks may be played together as a playout group. These V-PCC component tracks may represent respective encoded versions of one or more V-PCC components, and a video decoding device may play the tracks together in response to determining that the tracks belong to the same playout track group. The video decoding device may also determine from the media container file that certain PCC component tracks include tile groups that correspond to different objects in a point cloud or different parts of a same object in the point cloud. The video decoding device may decode these tile groups independently from each other so that a subset of the objects or parts of the point cloud may be accessed without also accessing the rest of the objects or parts.


