Volumetric Video Encoding via 3D to 2D Projection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face inefficiencies in compressing dynamic 3D scenes due to the ill-defined problem of identifying correspondences for motion-compensation in 3D-space, where geometry and attributes can change significantly between frames.

Innovation Solution

The proposed solution involves projecting 3D scenes onto 2D planes, which can then be encoded using standard 2D video compression technologies. This approach allows for efficient temporal compression and improved 6DOF capabilities by utilizing existing video encoding hardware.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If 3D motion compensation is used for volumetric video coding, then temporal compression can be achieved, but the complexity of identifying correspondences between frames increases significantly due to geometry and attribute changes

Engineering Contradiction:
Improvetemporal compression efficiencyVSAvoidcorrespondence identification complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent projects 3D volumetric data onto 2D planes, transforming the problem from 3D space to 2D space. This dimensionality reduction allows standard 2D video coding techniques to be applied, simplifying correspondence identification while maintaining temporal compression capabilities. The projection approach converts complex 3D motion compensation into 2D plane-based operations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of manufacture

If traditional 2D video compression is used for volumetric content, then implementation is simple, but coding efficiency and scene coverage are insufficient

Engineering Contradiction:
Improveimplementation simplicityVSAvoidcoding efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent encodes volumetric data by projecting it onto multiple 2D planes rather than using a single 2D representation. This multi-plane projection approach maintains the simplicity of 2D compression techniques while improving coding efficiency and scene coverage by capturing volumetric information across multiple viewing angles and depths.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The volumetric scene is divided into multiple 2D projection planes, allowing different regions of the 3D space to be encoded separately. This segmentation enables better utilization of 2D video compression tools on each plane while collectively representing the full volumetric content, thus improving overall coding efficiency.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If full 3D mesh data is transmitted, then complete geometric information is provided, but the data size and transmission bandwidth requirements increase

Engineering Contradiction:
Improvegeometric information completenessVSAvoiddata size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential geometric information needed for reconstruction by projecting 3D mesh data onto 2D planes. Instead of transmitting complete 3D mesh data including all vertices, edges, and faces, the method extracts projected 2D representations that contain sufficient information for visual reconstruction, significantly reducing data size while maintaining geometric fidelity.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12205333B2Method, an apparatus and a computer program product for volumetric video encoding and decoding
Publication Date: 2025.01.21 NOKIA TECHNOLOGIES OY
  • US12205333B2 patent drawing
  • US12205333B2 patent drawing
  • US12205333B2 patent drawing

AI summary

A method and technical equipment for encoding, where the method comprises at least receiving a video presentation frame, where the video presentation represents a three-dimensional data in the form of mesh data (810); separating from the mesh data information on vertices, wherein the information comprises at least connectivity data defining connections between vertices (820);determining parameters relating to said connectivity data (830); encoding the parameters to a first bitstream as a video component (840); and storing the encoded first bitstream for transmission to a rendering apparatus (850). In addition to encoding, also decoding is disclosed.