3D Volumetric Content Encoding Using 2D Video Atlases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Three-dimensional (3D) volumetric content files are large and costly to store and transmit, limiting their use due to high storage requirements and time-consuming data transfer, especially for real-time or on-demand applications.

Innovation Solution

The encoding of 3D volumetric content involves generating a 2D atlas by removing redundant information from multiple views and simplifying mesh-based representations through iterative edge or vertex removal, with distortion analysis ensuring minimal impact on geometry representation, allowing for efficient compression and transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If 3D volumetric content is captured using multiple cameras at different angles and locations, then the quality and completeness of the volumetric content is improved, but the file size and storage requirements increase significantly

Engineering Contradiction:
Improvequality of volumetric contentVSAvoidfile size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent divides the 3D volumetric content into multiple 2D video views captured from different camera angles and locations. Each view is processed and encoded separately, allowing for efficient compression and selective transmission based on viewer position and network conditions, thereby reducing overall storage requirements while maintaining content quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and encodes only the essential attribute information (color, texture, intensity, reflectivity) and geometry information (depth values) from the multi-camera captures. By separating and selectively encoding different types of data, the system reduces redundant information storage while preserving the necessary quality for realistic 3D reconstruction.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If all attribute information and geometry information from multiple camera views is stored in full resolution, then the accuracy of 3D reconstruction is improved, but the transmission time and storage cost increase

Engineering Contradiction:
Improveaccuracy of 3D reconstructionVSAvoidtransmission time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies different encoding precision levels to different regions of the 3D volumetric content based on their importance. High-priority regions (such as foreground objects or areas with significant geometric features) are encoded with higher precision, while less important background regions use lower precision encoding. This selective approach maintains reconstruction accuracy for critical areas while reducing overall data transmission time and storage requirements.

Inventive Principle:
Principle #3Local quality

3Quantity of substance

If redundant information is removed from multiple camera views to reduce file size, then storage efficiency is improved, but the complexity of processing and encoding increases

Engineering Contradiction:
Improvestorage efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent merges information from multiple camera views by identifying and eliminating redundant data. The system combines attribute information and geometry information from different views, using techniques such as view synthesis and depth map fusion, to create a consolidated representation that reduces storage requirements while managing processing complexity through efficient algorithms.

Inventive Principle:
Principle #5Merging (Combining)

4Quantity of substance

If mesh-based representations are simplified by removing edges or vertices, then the compression ratio is improved, but the geometric accuracy of the 3D content deteriorates

Engineering Contradiction:
Improvecompression ratioVSAvoidgeometric accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent applies partial simplification to mesh-based representations by selectively removing only certain edges or vertices that have minimal impact on overall geometric accuracy. The system uses distortion analysis to identify and remove only the least critical geometric details, achieving compression while maintaining sufficient geometric fidelity for most viewing conditions and applications.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11948338B13D volumetric content encoding using 2D videos and simplified 3D meshes
Publication Date: 2024.04.02 APPLE INC
  • US11948338B1 patent drawing
  • US11948338B1 patent drawing
  • US11948338B1 patent drawing

AI summary

An encoder encodes three-dimensional (3D) volumetric content, such as immersive media, using video encoded attribute patch images packed into a 2D atlas to communicate the attribute values for the 3D volumetric content. The encoder also uses mesh-encoded sub-meshes to communicate geometry information for portions of the 3D object or scene corresponding to the attribute patch images packed into the 2D atlas. The encoder applies decimation operations to the sub-meshes to simplify the sub-meshes before mesh encoding the sub-meshes. A distortion analysis is performed to bound the level to which the sub-meshes are simplified at the encoder. Mesh simplification at the encoder reduces the number of vertices and edges included in the sub-meshes which simplifies rendering at a decoder receiving the encoded 3D volumetric content.