Common Bitstream for 3D Video Texture and Depth Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding technologies face challenges in efficiently encoding and decoding three-dimensional (3D) video content, particularly in integrating texture and depth information within a common bitstream, which is essential for rendering and synthesizing virtual views in 3D applications.

Innovation Solution

The proposed solution involves encoding both texture and depth information within a common bitstream, where texture and depth components are signaled together, allowing for prediction dependencies between them, and utilizing camera parameters for view synthesis prediction, enabling efficient encoding and decoding of 3D video content by reusing prediction data and synthesizing virtual views.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If texture and depth information are encoded in separate bitstreams, then decoding complexity is reduced, but bitstream efficiency and flexibility for 3D rendering are worsened

Engineering Contradiction:
Improvedecoding complexityVSAvoidbitstream efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent combines texture and depth information into a single common bitstream, allowing both data types to be transmitted together. This merging approach improves bitstream efficiency by eliminating redundant signaling and enables flexible extraction of either texture or depth components for various 3D rendering applications without requiring separate transmission channels.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If prediction dependencies between texture and depth components are established, then coding efficiency is improved, but decoding complexity increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoiddecoding complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent establishes prediction dependencies where depth components are predicted from texture components of the same view or from depth components of other views. By pre-defining these prediction relationships in the encoding stage, the system achieves better coding efficiency through redundancy removal while the decoder follows predetermined rules to reconstruct the data, managing complexity through structured prediction rather than arbitrary processing.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple view components are encoded with full independence, then decoding flexibility is improved, but bitstream size and transmission overhead increase

Engineering Contradiction:
Improvedecoding flexibilityVSAvoidbitstream size
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent creates a universal bitstream structure that can serve multiple purposes: it can be decoded for 2D display using only texture components, for 3D stereoscopic display using texture and depth together, or for virtual view synthesis using depth information for view conversion. This multi-functional encoding approach reduces bitstream size by sharing common data while maintaining the flexibility to support various decoding scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP2684372B1Coding multiview video plus depth content
Publication Date: 2018.08.01 QUALCOMM INC
  • EP2684372B1 patent drawingFigure 1
  • EP2684372B1 patent drawingFigure 2
  • EP2684372B1 patent drawingFigure 3A~3B

AI summary

This disclosure describes techniques for coding 3D video block units. In one example, a video encoder is configured to receive one or more texture components from at least a portion of an image representing a view of three dimensional video data, receive a depth map component for at least the portion of the image, code a block unit indicative of pixels of the one or more texture components for a portion of the image and the depth map component. The coding comprises receiving texture data for a temporal instance of a view of video data, receiving depth data corresponding to the texture data for the temporal instance of the view of video data, and encapsulating the texture data and the depth data in a view component for the temporal instance of the view, such that the texture data and the depth data are encapsulated within a common bitstream.