3D Scene Encoding via Patch Segmentation and Projection Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for encoding volumetric videos face challenges with bitrate issues due to the large amount of data required to represent 3D scenes, leading to storage space, transmission, and decoding performance problems.

Innovation Solution

A method of encoding a 3D scene in a stream by obtaining patches with de-projection data, color pictures, and geometry data, generating color and depth images, and encoding these images along with patch data items in the stream, allowing for efficient decoding of the 3D scene.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If volumetric video is encoded by projecting 3D scenes onto projection maps and packing them in color and depth images, then the immersive quality and 6DoF experience are improved, but the bitrate and data size increase significantly

Engineering Contradiction:
Improveimmersive qualityVSAvoidbitrate
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent divides the 3D scene into multiple patches, each representing a specific region of the scene. Instead of encoding the entire scene as a single volumetric dataset, the system segments the scene into manageable patches that can be independently encoded and transmitted. This segmentation allows for selective transmission of only the patches visible from the current viewpoint, significantly reducing the bitrate while maintaining immersive quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from encoding complete volumetric data in three dimensions to encoding 2D projection maps that represent 3D scenes. By projecting the 3D scene onto 2D surfaces and encoding these projections, the system reduces the data dimensionality while preserving the essential visual information needed for 6DoF rendering, thereby reducing bitrate requirements.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If all patches are encoded and transmitted in the bit stream, then complete scene coverage is achieved, but storage space and transmission bandwidth are consumed

Engineering Contradiction:
Improvescene coverageVSAvoidstorage space
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by encoding and transmitting only the patches that are locally relevant to the current viewpoint, rather than uniformly encoding all patches in the scene. The system determines which patches are visible from the user's current position and transmits only those, ensuring complete coverage of the visible scene while avoiding transmission of occluded or irrelevant patches, thus optimizing storage space.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If high-resolution color and depth images are used for each patch, then rendering quality is improved, but decoding performance and processing time deteriorate

Engineering Contradiction:
Improverendering qualityVSAvoiddecoding performance
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies partial action by decoding and rendering only the patches that are currently visible from the user's viewpoint, rather than decoding all patches in the scene. This allows the system to maintain high rendering quality for the visible portions while significantly reducing the decoding workload and improving processing performance, as only a subset of the total patches needs to be processed at any given time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250184463A1Method and apparatus for encoding and decoding three-dimensional scenes in and from a data stream
Publication Date: 2025.06.05 INTERDIGITAL VC HOLDINGS INC
  • US20250184463A1 patent drawing
  • US20250184463A1 patent drawing
  • US20250184463A1 patent drawing

AI summary

Methods and devices are provided to encode and decode a data stream carrying data representative of a three-dimensional scene, the data stream comprising color pictures packed in a color image; depth pictures packed in a depth image; and a set of patch data items comprising de-projection data; data for retrieving a color picture in the color image and geometry data. Two types of geometry data are possible. The first type of data describes how to retrieve a depth picture in the depth image. The second type of data comprises an identifier of a parametric function and a list of parameter values for the identified parametric function.