Volumetric Video Encoding and Decoding Through 2D Projection Views

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing volumetric video coding technologies suffer from poor spatial and temporal compression performance, particularly in dynamic 3D scenes where geometry and attributes change, leading to inefficient compression.

Innovation Solution

A system for capturing, encoding, decoding, and reconstructing three-dimensional scenes using multiple cameras and microphones to create a scene model, projecting 3D data onto 2D planes, and employing standard 2D video coding tools for efficient temporal compression, with projection surfaces optimized for individual objects to improve coverage and utilize standard video encoding hardware.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If volumetric video data is represented using traditional formats (triangle meshes, point clouds, or voxel arrays), then the data can be stored and processed, but the spatial and temporal coding performance deteriorates

Engineering Contradiction:
Improvecoding performanceVSAvoiddata representation complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent transforms 3D volumetric video data into 2D projection views, projecting three-dimensional scene geometry and attributes onto two-dimensional planes. This dimensionality reduction enables the use of efficient 2D video coding standards while preserving the ability to reconstruct 3D scenes from multiple viewpoints, thereby improving coding performance without excessive complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent divides the volumetric scene into multiple discrete projection views or layers, each representing a different perspective or depth plane. By segmenting the 3D data into manageable 2D components, the system can apply standard video coding techniques to each segment independently, improving overall compression efficiency while maintaining spatial and temporal coherence

Inventive Principle:
Principle #1Segmentation

2Productivity

If standard 2D video coding tools are used for volumetric data, then coding efficiency improves, but the 3D spatial information may be lost or degraded

Engineering Contradiction:
Improvecoding efficiencyVSAvoidspatial accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent applies different coding strategies to different regions or layers of the volumetric data. By tailoring the projection and coding approach to local spatial characteristics, the system maintains high spatial accuracy in critical regions while achieving efficient compression overall, preventing uniform degradation of 3D spatial information

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent combines multiple 2D projection views or layers into a composite volumetric representation. By synthesizing information from multiple coded 2D perspectives, the system reconstructs accurate 3D spatial structure, preventing loss of spatial information while benefiting from efficient 2D coding applied to each component

Inventive Principle:
Principle #40Composite materials

Data Source

PatentEP3669330B1Encoding and decoding of volumetric video
Publication Date: 2025.07.02 NOKIA TECHNOLOGIES OY
  • EP3669330B1 patent drawingFigure 1
  • EP3669330B1 patent drawingFigure 2a~2b
  • EP3669330B1 patent drawingFigure 3a~3b

AI summary

There are provided methods, apparatuses, systems and computer program products for coding volumetric video, where a first texture picture is coded, the first texture picture comprising a first projection of texture data of a first source volume of a digital scene model, the scene model comprising a number of further source volumes, the first projection being from the first source volume to a first projection surface, a first geometry picture is coded, the first geometry picture representing a mapping of the first projection surface to the first source volume, and first projection geometry information of the first projection is coded, the first projection geometry information comprising information of position of the first projection surface in the scene model.