Volumetric Video Projection for 2D Coding Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current volumetric video coding technologies face inefficiencies in compressing dynamic 3D scenes due to poor spatial and temporal coding performance, particularly in identifying correspondences for motion-compensation in 3D-space, where geometry and attributes change, leading to inefficient compression of volumetric data.

Innovation Solution

The method involves projecting volumetric video data onto simple geometric surfaces such as spheres, cylinders, or planes, unfolding these onto 2D planes, and applying standard 2D video coding techniques to encode texture and depth images, with relevant projection geometry information transmitted alongside for decoding and 3D reconstruction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If volumetric video data is compressed using traditional 3D-space motion-compensation methods, then compression is attempted, but coding efficiency is poor due to difficulty in identifying correspondences when geometry and attributes change

Engineering Contradiction:
Improvecompression efficiencyVSAvoidcoding complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent projects 3D volumetric video data onto 2D projection surfaces (such as spherical, cylindrical, or planar surfaces), transforming the problem from 3D space to 2D space. This dimensionality reduction allows standard 2D video coding tools to be applied effectively, improving compression efficiency while maintaining manageable complexity. The projection approach converts complex 3D motion compensation into simpler 2D frame-based coding.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If dense point clouds or voxel arrays with tens or hundreds of millions of points are used to represent the reconstructed 3D scene, then full scene coverage and 6DOF capabilities are achieved, but storage and transmission requirements become prohibitively large

Engineering Contradiction:
Improvescene coverage completenessVSAvoiddata volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Instead of storing and transmitting the full dense point cloud or voxel array, the patent creates 2D projected representations (copies) of the 3D scene on projection surfaces. These 2D projections contain sufficient information to reconstruct the full 3D scene with 6DOF capabilities, dramatically reducing data volume while maintaining scene coverage completeness. The projection process creates a compact representation that preserves essential geometric and visual information.

Inventive Principle:
Principle #26Copying

3Productivity

If 2D-video based approaches (multiview+depth) are used for compressing volumetric data, then compression efficiency improves, but full scene coverage and 6DOF capabilities are not achieved

Engineering Contradiction:
Improvecompression efficiencyVSAvoid6DOF capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent uses projection surfaces (spherical, cylindrical, or planar) to map 3D volumetric data onto 2D surfaces, enabling the application of efficient 2D video coding tools. Unlike traditional multiview approaches that use multiple discrete camera views, the projection surface method provides continuous coverage across the entire surface, enabling full 6DOF capabilities while maintaining high compression efficiency through standard 2D coding techniques.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11202086B2Apparatus, a method and a computer program for volumetric video
Publication Date: 2021.12.14 NOKIA TECHNOLOGIES OY
  • US11202086B2 patent drawing
  • US11202086B2 patent drawing
  • US11202086B2 patent drawing

AI summary

Video encoding may comprise obtaining a volumetric content containing visual information of three-dimensional objects; generating at least one patch by projecting the visual information of three-dimensional objects of the volumetric content to at least one projection plane. Video decoding may comprise obtaining neighboring pixels of a location on the 2D image based on said geometry information; determining a difference of values of the neighboring pixels on the 2D image; comparing the difference with a value range to determine a number of 3D points to be interpolated; projecting back the 2D image to create the volumetric content; wherein the projection comprises interpolating the number of 3D points on the basis of the values of the neighboring pixels.