Additional View Block Packing for 3DoF+ Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing coding schemes for immersive media, particularly 3DoF+ video, face challenges in efficiently reducing redundancy and computational complexity for multi-view image or video data, especially in virtual reality applications, where high redundancy exists among views captured from slightly different positions, leading to increased bandwidth and computational demands.
Innovation Solution
A method involving block rearrangement and transformation of additional views to create a packed additional view, which is then split into parts and transformed to reduce size, followed by encoding with metadata describing the process, using lossless compression for metadata and lossy compression for video data, to enhance coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple views are captured from slightly different positions to enable 3DoF+ video, then immersive experience is improved, but redundancy among views increases
Solution Approach 1:
The patent combines multiple additional views into a single packed additional view by identifying and merging common image data regions across different views. This merging process eliminates redundancy while preserving the immersive experience by maintaining all unique visual information from multiple camera positions.
2Manufacturing precision
If all views are encoded separately to maintain quality, then image quality is preserved, but bitrate increases
Solution Approach 1:
The patent merges multiple additional views into a single packed additional view that contains all unique visual information from the original views. This consolidated representation significantly reduces the bitrate required to encode the same visual content while maintaining image quality, as common regions are encoded only once rather than separately in each view.
3Ease of manufacture
If conventional video codecs are used directly on multi-view data, then encoding simplicity is maintained, but coding efficiency is insufficient
Solution Approach 1:
The patent performs preliminary packing of multiple additional views into a single consolidated view before encoding. This preliminary action reorganizes the multi-view data into a format that conventional video codecs can process efficiently, eliminating the need for complex multi-view encoding algorithms while significantly improving coding efficiency through reduced redundancy.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An encoder, decoder, encoding method and decoding method for 3DoF+ video are disclosed. The encoding method comprises receiving (110) multi-view image or video data comprising a basic view and at least a first additional view of a scene. The method proceeds by identifying (220) pixels in the first additional view that need to be encoded because they contain scene-content that is not visible in the basic view. The first additional view is divided (230) into a plurality of first blocks of pixels. First blocks containing at least one of the identified pixels are retained (240); and first blocks that contain none of the identified pixels are discarded. The retained blocks are rearranged (250) so that they are contiguous in at least one dimension. A packed additional view is generated (260) from the rearranged first retained blocks and encoded (264).