3D Scene Encoding via Patch Segmentation and Projection Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for encoding volumetric videos face challenges with bitrate issues due to the large amount of data required to represent 3D scenes, leading to storage space, transmission, and decoding performance problems.
Innovation Solution
A method of encoding a 3D scene in a stream by obtaining patches with de-projection data, color pictures, and geometry data, generating color and depth images, and encoding these images along with patch data items in the stream, allowing for efficient decoding of the 3D scene.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If volumetric video is encoded by projecting 3D scenes onto projection maps and packing them in color and depth images, then the immersive quality and 6DoF experience are improved, but the bitrate and data size increase significantly
Solution Approach 1:
The patent divides the 3D scene into multiple patches, each representing a specific region of the scene. Instead of encoding the entire scene as a single volumetric dataset, the system segments the scene into manageable patches that can be independently encoded and transmitted. This segmentation allows for selective transmission of only the patches visible from the current viewpoint, significantly reducing the bitrate while maintaining immersive quality.
Solution Approach 2:
The patent transitions from encoding complete volumetric data in three dimensions to encoding 2D projection maps that represent 3D scenes. By projecting the 3D scene onto 2D surfaces and encoding these projections, the system reduces the data dimensionality while preserving the essential visual information needed for 6DoF rendering, thereby reducing bitrate requirements.
2Reliability
If all patches are encoded and transmitted in the bit stream, then complete scene coverage is achieved, but storage space and transmission bandwidth are consumed
Solution Approach 1:
The patent applies local quality by encoding and transmitting only the patches that are locally relevant to the current viewpoint, rather than uniformly encoding all patches in the scene. The system determines which patches are visible from the user's current position and transmits only those, ensuring complete coverage of the visible scene while avoiding transmission of occluded or irrelevant patches, thus optimizing storage space.
3Measurement precision
If high-resolution color and depth images are used for each patch, then rendering quality is improved, but decoding performance and processing time deteriorate
Solution Approach 1:
The patent applies partial action by decoding and rendering only the patches that are currently visible from the user's viewpoint, rather than decoding all patches in the scene. This allows the system to maintain high rendering quality for the visible portions while significantly reducing the decoding workload and improving processing performance, as only a subset of the total patches needs to be processed at any given time.
Data Source
AI summary
Methods and devices are provided to encode and decode a data stream carrying data representative of a three-dimensional scene, the data stream comprising color pictures packed in a color image; depth pictures packed in a depth image; and a set of patch data items comprising de-projection data; data for retrieving a color picture in the color image and geometry data. Two types of geometry data are possible. The first type of data describes how to retrieve a depth picture in the depth image. The second type of data comprises an identifier of a parametric function and a list of parameter values for the identified parametric function.


