Padded Reference Images for Multi-View Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern video coding systems are inefficient in recognizing and coding motion effects in multi-view image data, as they assume flat image content, failing to exploit spatial discontinuities and motion patterns present in multi-view imaging systems.
Innovation Solution
The generation of padded reference images in a cube map format, where image data from adjacent views is replicated and placed adjacent to each other, increases the likelihood of prediction matches during coding, enhancing the efficiency of video coding by maintaining image continuity across spatial boundaries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional video coding protocols are used for multi-view image data, then the coding system operates on flat image data assumptions, but it fails to recognize motion effects and spatial discontinuities in multi-view data
Solution Approach 1:
The patent extends conventional 2D flat image coding to 3D multi-view image data by introducing a depth dimension. Cube map projections are used to represent multi-view data in a three-dimensional space, allowing the coding system to handle spatial discontinuities and motion effects that occur across different views. This dimensional extension enables the coding system to recognize and exploit motion patterns that would be invisible in traditional flat 2D coding.
2Ease of operation
If image data is treated as continuous two-dimensional field, then conventional coding protocols can operate, but spatial discontinuities in multi-view data are not recognized
Solution Approach 1:
The patent segments multi-view image data into distinct cube map faces, each representing a different view direction. By dividing the continuous 3D space into discrete cubic segments, the system can apply conventional 2D coding protocols to each face while maintaining awareness of spatial relationships between views. This segmentation allows motion detection to work effectively within each view while the overall system handles multi-view spatial discontinuities.
3Quantity of substance
If object motion is small in free space, then physical movement is minimal, but it represents large spatial movements in image data
Solution Approach 1:
The patent uses cube map projections to create multiple copies of the same 3D scene from different viewpoint directions. Each cube face contains a projected view of the environment, allowing the coding system to find matching blocks across different views even when objects have moved only slightly in physical space. This copying approach enables differential coding to work effectively by comparing corresponding regions across multiple view copies.
Data Source
Figure 1~2
Figure 3(a)~3(c)
Figure 4~5
AI summary
Techniques are disclosed for coding and decoding video captured as cube map images. According to these techniques, padded reference images are generated for use during predicting input data, A reference image is stored in a cube map format, A padded reference image is generated from the reference image in which image data of a first view contained in reference image is replicated and placed adjacent to a second view contained in the cube map image. When coding a pixel block of an input image, a prediction search may be performed between the input pixel block and content of the padded reference image. When the prediction search identifies a match, the pixel block may be coded with respect to matching data from the padded reference image. Presence of replicated data in the padded reference image is expected to increase the likelihood that adequate prediction matches will be identified for input pixel block data, which will increase overall efficiency of the video coding.