Unified 3D Coding Framework for Objects and Scenes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current 3D coding methods, such as V-PCC and MIV, separately encode 3D objects and scenes, lacking a unified approach that can represent both camera models effectively, leading to inefficient compression and limited navigation and depth representation.

Innovation Solution

A flexible camera model is used to capture 3D objects or scenes, transmitted as metadata, and combined in a 2D atlas for further compression with conventional video encoders, allowing both 3D objects and scenes to be encoded in a single bitstream by signaling a general camera concept.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate coding methods (V-PCC and MIV) are used for 3D objects and scenes, then each method can be optimized for its specific camera model, but the device complexity increases and a unified bitstream cannot be generated

Engineering Contradiction:
Improvecoding optimization for specific camera modelVSAvoidseparate coding systems
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges V-PCC and MIV coding methods into a unified 3D coding framework by defining a common syntax structure that can represent both orthographic (V-PCC) and perspective (MIV) camera models. The bitstream includes unified syntax elements such as camera model type indicators, projection parameters, and depth representation methods that work for both coding approaches, eliminating the need for separate coding systems while maintaining optimization for each camera model type.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If V-PCC coding is used for 3D objects with orthographic cameras, then object representation is optimized, but scene elements and 360 degree views cannot be efficiently coded

Engineering Contradiction:
Improve3D object representationVSAvoidscene coding capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The unified coding framework provides universality by enabling a single system to handle multiple coding scenarios: V-PCC for isolated 3D objects with orthographic cameras, MIV for 360-degree scenes with perspective cameras, and hybrid cases combining both. The syntax structure adapts to different camera models and content types through conditional syntax elements and parameter settings, making the system multi-functional without requiring separate coding paths.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If MIV coding is used for 3D scenes with perspective cameras, then scene capture is efficient, but navigation is limited to camera positions and depth representation is non-uniform

Engineering Contradiction:
Improvescene encoding efficiencyVSAvoidnavigation flexibility
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The unified framework enables parameter changes between different depth representation methods. For V-PCC, uniform depth scaling is applied to maintain consistent precision across all depths. For MIV, the system can switch between reversed depth and other depth encoding schemes. The syntax includes parameters for depth near/far planes, scaling factors, and representation type indicators that allow flexible adjustment of depth encoding parameters based on the specific coding scenario, improving both navigation flexibility and encoding efficiency.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If different syntax elements are used for V-PCC and MIV camera models, then each camera model can be precisely represented, but a unified bitstream cannot be generated and separate encoding is required

Engineering Contradiction:
Improvecamera model representationVSAvoidmultiple syntax systems
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The unified syntax structure uses segmentation to organize camera model parameters into distinct sections: a camera model type indicator field that identifies whether V-PCC or MIV is being used, followed by conditional syntax elements specific to each model type. This segmented approach allows precise representation of each camera model while maintaining a unified bitstream structure, as the parser can selectively process only the relevant syntax elements based on the model type indicator.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11196977B2Unified coding of 3D objects and scenes
Publication Date: 2021.12.07 SONY GROUP CORP
  • US11196977B2 patent drawing
  • US11196977B2 patent drawing
  • US11196977B2 patent drawing

AI summary

Methods for unified coding of 3D objects and 3D scenes are described herein. A flexible camera model is used to capture either parts of a 3D object or multiple views of a 3D scene. The flexible camera model is transmitted as metadata, and the captured 3D elements (objects and scenes) are combined in a 2D atlas image that is able to be further compressed with 2D video encoders. Described herein is a unification of two implementations for 3D coding: V-PCC and MIV. Projections of the scene/object of interest are used to map the 3D information into 2D, and then subsequently use video encoders. A general camera concept is able to represent both models via signaling to allow coding of both 3D objects and 3D scenes. With the flexible signaling of the camera model, the encoder is able to compress several different types of content into a single bitstream.