Common Bitstream for 3D Video Texture and Depth Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding technologies face challenges in efficiently encoding and decoding three-dimensional (3D) video content, particularly in integrating texture and depth information within a common bitstream, which is essential for rendering and synthesizing virtual views in 3D applications.
Innovation Solution
The proposed solution involves encoding both texture and depth information within a common bitstream, where texture and depth components are signaled together, allowing for prediction dependencies between them, and utilizing camera parameters for view synthesis prediction, enabling efficient encoding and decoding of 3D video content by reusing prediction data and synthesizing virtual views.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If texture and depth information are encoded in separate bitstreams, then decoding complexity is reduced, but bitstream efficiency and flexibility for 3D rendering are worsened
Solution Approach 1:
The patent combines texture and depth information into a single common bitstream, allowing both data types to be transmitted together. This merging approach improves bitstream efficiency by eliminating redundant signaling and enables flexible extraction of either texture or depth components for various 3D rendering applications without requiring separate transmission channels.
2Productivity
If prediction dependencies between texture and depth components are established, then coding efficiency is improved, but decoding complexity increases
Solution Approach 1:
The patent establishes prediction dependencies where depth components are predicted from texture components of the same view or from depth components of other views. By pre-defining these prediction relationships in the encoding stage, the system achieves better coding efficiency through redundancy removal while the decoder follows predetermined rules to reconstruct the data, managing complexity through structured prediction rather than arbitrary processing.
3Adaptability or versatility
If multiple view components are encoded with full independence, then decoding flexibility is improved, but bitstream size and transmission overhead increase
Solution Approach 1:
The patent creates a universal bitstream structure that can serve multiple purposes: it can be decoded for 2D display using only texture components, for 3D stereoscopic display using texture and depth together, or for virtual view synthesis using depth information for view conversion. This multi-functional encoding approach reduces bitstream size by sharing common data while maintaining the flexibility to support various decoding scenarios.
Data Source
Figure 1
Figure 2
Figure 3A~3B
AI summary
This disclosure describes techniques for coding 3D video block units. In one example, a video encoder is configured to receive one or more texture components from at least a portion of an image representing a view of three dimensional video data, receive a depth map component for at least the portion of the image, code a block unit indicative of pixels of the one or more texture components for a portion of the image and the depth map component. The coding comprises receiving texture data for a temporal instance of a view of video data, receiving depth data corresponding to the texture data for the temporal instance of the view of video data, and encapsulating the texture data and the depth data in a view component for the temporal instance of the view, such that the texture data and the depth data are encapsulated within a common bitstream.