Feature-Based Multi-View Coding With 3D Feature Changes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding technologies face challenges in efficiently compressing multi-view video data, particularly in capturing and encoding the spatial and temporal variations across different views, leading to increased bandwidth and storage requirements.
Innovation Solution
The proposed solution involves feature-based multi-view representation and coding, where 3D feature information is used to decode key pictures from a multi-view bitstream, utilizing pre-determined 3D feature models to determine feature changes and reconstruct images across different views, enabling efficient encoding and decoding of multi-view video data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional video coding technologies are used to encode multi-view video data, then compression can be achieved, but bandwidth and storage requirements increase due to inability to efficiently capture spatial and temporal variations across different views
Solution Approach 1:
The patent segments multi-view video data into key pictures and non-key pictures, and further divides non-key pictures into foreground objects and background regions. This segmentation allows different compression strategies to be applied to different parts of the data, improving overall compression efficiency while reducing bandwidth and storage requirements.
Solution Approach 2:
The patent changes the representation parameters by encoding non-key pictures using feature information (such as 3D position, velocity, acceleration) instead of traditional pixel-based methods. This parameter transformation enables more efficient compression by capturing essential spatial and temporal variations with fewer bits, directly addressing the bandwidth and storage requirements issue.
2Productivity
If feature-based multi-view representation and coding is used, then compression efficiency improves and bandwidth/storage requirements reduce, but complexity of encoding and decoding processes increases
Solution Approach 1:
The patent performs preliminary actions by pre-determining 3D feature models and extracting feature information from key pictures before encoding non-key pictures. This preliminary processing simplifies the subsequent encoding of non-key pictures, as only feature changes need to be encoded rather than complete picture data, thereby reducing overall encoding complexity despite the advanced feature-based approach.
Solution Approach 2:
The patent introduces feature information (such as 3D position, velocity, acceleration) as an intermediary between the original video data and the compressed representation. This intermediary layer captures the essential spatial and temporal variations across views, enabling efficient compression while managing complexity by working with simplified feature representations rather than full-resolution multi-view data.
3Measurement precision
If all pictures are encoded as key pictures using traditional methods, then reconstruction accuracy is maintained, but bandwidth and storage requirements increase significantly
Solution Approach 1:
The patent applies partial action by encoding only the essential feature information for non-key pictures rather than complete picture data. By capturing critical spatial and temporal variations through feature representation (3D position, velocity, acceleration), the system maintains adequate reconstruction accuracy for most applications while dramatically reducing bandwidth and storage requirements compared to encoding all pictures as full-resolution key pictures.
Data Source
AI summary
Aspects of the disclosure provide a method, an apparatus, and non-transitory computer-readable storage medium for video decoding. The apparatus includes processing circuitry configured to decode at least one first key picture of pictures from a multi-view bitstream. The pictures correspond to different views. The at least one first key picture corresponds to at least one first view of the different views. The processing circuitry determines first feature information of content in the at least one first key picture. The processing circuitry decode, based on the multi-view bitstream, a first feature change to the first feature information. The first feature change indicates a content change between a key picture in the at least one first key picture and a first picture. The processing circuitry reconstructs the first picture based on the decoded first feature change, the first feature information, and the key picture in the at least one first key picture.


