Cascaded Segmentation Decoding for Scalable Bitstream Parsing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image and video coding technologies, including hybrid codecs and machine learning applications, face challenges in efficiently decoding data with scalability and adaptability to varying content characteristics, leading to suboptimal bitstream efficiency and processing times.
Innovation Solution
A method and apparatus for decoding data using a cascaded structure of segmentation information processing layers, which process sets of segmentation information elements in multiple layers, enabling efficient parsing and upsampling, and allowing for parallel processing on GPUs/NPUs, with trainable convolution kernels for improved decoding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional hybrid coding methods are used with separate optimization of transformation, quantization, and entropy coding, then each component can be independently optimized, but the overall decoding efficiency and adaptability to varying content characteristics deteriorate
Solution Approach 1:
The patent combines multiple coding operations (transformation, quantization, entropy coding) into a unified neural network framework where operations are jointly optimized rather than separately optimized. The cascaded processing layers integrate these functions to improve overall decoding efficiency while maintaining adaptability to content characteristics.
Solution Approach 2:
The patent introduces dynamic adaptability through neural network-based processing that can adjust to varying content characteristics. The system dynamically optimizes decoding parameters and operations based on the specific input data, enabling efficient processing across different content types while maintaining high decoding efficiency.
2Productivity
If machine learning is applied to determine or optimize prediction parameters, then coding efficiency improves, but the amount of side information that needs to be transmitted deteriorates
Solution Approach 1:
The patent extracts and processes segmentation information through cascaded neural network layers that identify and process only the most relevant features. This extraction approach enables efficient coding by transmitting only essential side information while maintaining high coding efficiency through intelligent feature selection and compression.
Solution Approach 2:
The patent transforms segmentation information through multiple processing layers that change parameters and representations of the data. This transformation process converts detailed segmentation data into compressed representations that maintain coding efficiency while reducing the quantity of side information that must be transmitted.
3Productivity
If segmentation information is processed through multiple cascaded layers with upsampling, then decoding adaptability and efficiency improve, but the processing complexity and bitstream requirements deteriorate
Solution Approach 1:
The patent divides the decoding process into multiple cascaded processing layers, each handling specific aspects of segmentation information. This segmentation of the processing pipeline enables efficient parallel processing and optimization of individual layers while maintaining overall decoding efficiency, despite increased processing complexity.
Solution Approach 2:
The patent introduces upsampling operations that change the dimensional representation of segmentation information across processing layers. This dimensional transformation enables more detailed processing and improved decoding efficiency by operating at multiple resolution levels, though it increases processing complexity and bitstream requirements.
Data Source
AI summary
The present disclosure relates to methods and apparatuses for decoding data for (still or video processing into a bitstream). Two or more sets of segmentation information elements are obtained from the bitstream. Then, each of the two or more sets of segmentation information elements are inputted respectively into two or more segmentation information processing layers out of a plurality of cascaded layers. In each of the two or more segmentation information processing layers, the respective sets of segmentation information are processed. The decoded data for picture or video processing are obtained based on the segmentation information processed by the plurality of cascaded layers. Accordingly, the data may be decoded from the bitstream in an efficient manner in the layered structure.


