Partial Decoding of 3D Point Cloud Bitstreams via 2D Tile Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The transmission of uncompressed 3D point clouds requires significant bandwidth, and existing technologies lack efficient methods for encoding and decoding 3D video content, necessitating specialized hardware that is often expensive and not widely available.
Innovation Solution
Converting 3D point clouds into 2D representations using existing 2D video codecs for compression and reconstruction, allowing for partial decoding and rendering of point clouds, reducing bandwidth requirements and eliminating the need for specialized hardware.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D point clouds are transmitted uncompressed, then the quality and detail of the 3D data is preserved, but the bandwidth requirement becomes excessively large
Solution Approach 1:
The patent transforms 3D point cloud data into 2D video frame representations, changing the dimensional domain from three-dimensional spatial data to two-dimensional image data. This allows the use ofๆ็ 2D video compression codecs to reduce bandwidth requirements while maintaining acceptable visual quality through the projection and reconstruction process.
Solution Approach 2:
The patent divides the 3D point cloud scene into multiple independently encoded tiles, each representing a portion of the 3D space. This segmentation allows selective decoding of only the tiles corresponding to objects of interest, reducing the overall data transmission requirement while preserving the ability to reconstruct high-quality 3D data where needed.
2Productivity
If specialized hardware is used to compress 3D point clouds, then compression efficiency is improved, but device complexity and cost increase
Solution Approach 1:
The patent creates a virtual copy of the 3D point cloud data by projecting it onto 2D video frames. This copy can then be processed using existing 2D video compression infrastructure, eliminating the need for specialized 3D compression hardware while maintaining compression efficiency through the use of well-optimized 2D codecs.
Solution Approach 2:
The patent makes existing 2D video compression hardware and software universally applicable to 3D point cloud data by representing the 3D data as 2D projections. This allows standard video codecs and processing pipelines to handle both 2D and 3D content, eliminating the need for dedicated 3D compression hardware.
3Loss of information
If the entire 3D point cloud bitstream is decoded, then complete 3D scene information is obtained, but processing time and computational resources increase
Solution Approach 1:
The patent segments the 3D point cloud into multiple independently decodable tiles, each associated with specific 3D space regions or objects. Users can selectively decode only the tiles corresponding to their current view or objects of interest, significantly reducing processing time and computational resources while maintaining complete information for the regions that are decoded.
Solution Approach 2:
The patent enables partial decoding of only the necessary portions of the 3D point cloud data based on user view, object interest, or application requirements. This partial action approach provides sufficient information for the intended application without the excessive computational cost of decoding the entire 3D scene.
Data Source
AI summary
A decoding device for point cloud decoding includes a communication interface and a processor. The communication interface configured to receive a bitstream. The processor is configured to identify, from the bitstream, messages and one or more sub-bitstreams representing a three-dimensional (3D) point cloud. The processor is configured to identify a query label indicating an object of the 3D point cloud for decoding. In response to determining that the query label corresponds to a label, the processor is configured to identify a 3D scene object associated with the query label. The processor is configured to identify a 2D tile that correspond to the 3D scene object. The processor is configured to determine to decode, based on an identification of the 2D tile from the sub-bitstream, a portion of the sub-bitstream corresponding to the 2D tile to generate a portion of a video frame representing a portion of the 3D scene object.


