Layered Segmentation Decoding for Efficient Video Bitstreams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image and video coding technologies, including hybrid codecs and machine learning applications, face challenges in efficiently decoding data with scalability and adaptability to content characteristics, leading to suboptimal bitstream efficiency and processing time.
Innovation Solution
A method and apparatus for decoding data using a cascaded structure of segmentation information processing layers, which process sets of segmentation information elements in multiple layers, enabling efficient parsing and upsampling, and allowing for parallel processing on GPUs/NPUs, with trainable convolution kernels for improved decoding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional hybrid coding methods are used with separate optimization of transformation, quantization, and entropy coding, then each component can be optimized independently, but the overall decoding efficiency and adaptability to content characteristics are suboptimal
Solution Approach 1:
The patent merges traditionally separate coding components (transformation, quantization, entropy coding) into a unified neural network framework where segmentation information is processed through multiple cascaded layers that jointly optimize all components. This integration allows the system to achieve both independent component optimization and improved overall decoding efficiency simultaneously.
Solution Approach 2:
The patent applies segmentation by dividing the decoding process into multiple processing layers that handle different aspects of segmentation information separately. Each layer processes specific segmentation elements independently before combining results, enabling both modular optimization and efficient parallel processing.
2Adaptability or versatility
If machine learning is applied to determine or optimize prediction parameters, then adaptability to content characteristics improves, but the amount of side information needed in the bitstream increases
Solution Approach 1:
The patent extracts only the essential segmentation information needed for adaptive decoding while discarding redundant data. By selectively transmitting only critical segmentation elements through the bitstream, the system maintains high adaptability to content characteristics while minimizing the quantity of side information required.
Solution Approach 2:
The patent changes the representation parameters of segmentation information by encoding it in a compressed hierarchical format. This parameter transformation reduces the bitstream requirements while preserving the adaptability needed for content-specific optimization, allowing the decoder to reconstruct detailed segmentation maps from compact representations.
3Productivity
If segmentation information is processed through multiple cascaded layers with upsampling, then decoding adaptability and efficiency improve, but processing complexity increases
Solution Approach 1:
The patent resolves complexity issues by transitioning to parallel processing across multiple dimensions using GPU/NPU architectures. Instead of sequential processing through cascaded layers, the system processes segmentation information simultaneously across multiple processing units, maintaining high decoding efficiency while reducing the time dimension of complexity.
Solution Approach 2:
The patent uses copying by replicating processing layers across multiple hardware instances in parallel. Each layer is copied to dedicated processing units that operate simultaneously, reducing overall processing time and distributing complexity across multiple identical modules rather than requiring a single complex sequential processor.
Data Source
AI summary
The present disclosure relates to methods and apparatuses for decoding data for (still or video processing into a bitstream). Two or more sets of segmentation information elements are obtained from the bitstream. Then, each of the two or more sets of segmentation information elements are inputted respectively into two or more segmentation information processing layers out of a plurality of cascaded layers. In each of the two or more segmentation information processing layers, the respective sets of segmentation information are processed. The decoded data for picture or video processing are obtained based on the segmentation information processed by the plurality of cascaded layers. Accordingly, the data may be decoded from the bitstream in an efficient manner in the layered structure.


