3D Video Coding Using Depth-Based View Synthesis Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding technologies face challenges in efficiently encoding and decoding video data, particularly in 3D video applications, due to the complexity of prediction schemes and the need for increased bandwidth management, which is exacerbated by the requirement for depth-enhanced video coding.
Innovation Solution
The proposed solution involves a method for encoding and decoding video information that omits motion vector information from signaling and employs depth information for prediction, allowing for the use of view synthesis prediction and disparity compensation, thereby optimizing video coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If multiple prediction schemes (temporal, spatial, view synthesis) are used to reduce redundant information, then bandwidth requirement is reduced, but device complexity and processing speed are adversely affected
Solution Approach 1:
The patent applies dynamics by making the prediction scheme selectable and adaptable. The encoder can choose between temporal prediction, spatial prediction, and view synthesis prediction based on content characteristics and bandwidth requirements. This dynamic selection allows the system to optimize between compression efficiency and processing complexity in real-time, rather than being locked into a single complex prediction architecture.
Solution Approach 2:
The patent segments the prediction process into distinct, independently selectable schemes (temporal prediction, spatial prediction, view synthesis prediction). Each scheme can be applied to different picture types or regions, allowing the system to use simpler prediction methods where sufficient and more advanced methods only where needed, thereby reducing overall device complexity while maintaining bandwidth efficiency.
2Loss of energy
If multiple prediction schemes are implemented to improve compression efficiency, then bandwidth requirement is reduced, but processing speed is adversely affected
Solution Approach 1:
The system dynamically selects prediction schemes based on picture type indicators and content characteristics. For I pictures and P pictures, simpler temporal or spatial prediction can be used, while B pictures or specific regions may utilize view synthesis prediction. This dynamic adaptation ensures that processing speed is maintained for common cases while achieving high compression efficiency when needed.
Solution Approach 2:
The patent applies partial action by selectively applying complex view synthesis prediction only to specific picture types (B pictures) or specific regions where the compression benefit justifies the processing cost. For other pictures, simpler prediction methods are used, ensuring that processing speed is not adversely affected in cases where full complexity is unnecessary.
3Manufacturing precision
If depth-enhanced video coding is used to improve 3D video quality, then video coding efficiency is improved, but bandwidth requirement and device complexity are adversely affected
Solution Approach 1:
The patent segments depth-enhanced video coding into separate, independently processable components: depth map generation, view synthesis prediction using depth information, and residual coding. This segmentation allows the complex depth-enhanced coding to be broken down into manageable stages that can be processed independently, reducing overall device complexity while maintaining coding efficiency.
Solution Approach 2:
The system dynamically determines whether to apply depth-enhanced video coding based on picture type and content characteristics. Depth maps and view synthesis prediction are selectively applied only where they provide significant quality improvement, rather than being applied universally. This dynamic approach maintains video coding efficiency while avoiding unnecessary complexity in cases where depth enhancement provides minimal benefit.
Data Source
Figure 1a~1c
Figure 2a~2b
Figure 3a~3b
AI summary
There are disclosed various methods, apparatuses and computer program products for video encoding. The type of prediction used for a reference picture index may be signaled in the video bit-stream. The omission of motion vectors from the video bit-stream for a certain image element may be signaled; signaling may indicate to the decoder that motion vectors used in prediction are to be construed at the decoder. The construction of motion vectors may take place by using disparity information that has been obtained from depth information of the picture being used as a reference.