Feature Map 4:2:0 Packing for Efficient VVC Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video compression technologies, such as VVC, face challenges in efficiently encoding and decoding feature maps from convolutional neural networks (CNNs), particularly in handling tensors with varying spatial dimensions and channel counts, which affects processing power and memory usage in both cloud and edge device implementations.
Innovation Solution
A method and apparatus for encoding and decoding feature maps by arranging samples in different two-dimensional arrays, allowing for efficient packing and quantization, and using VVC tools to optimize bitstream generation and decoding, thereby reducing computational overhead and improving task performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If feature maps are encoded using conventional video compression technologies, then encoding can be performed with standard tools, but computational overhead and memory usage increase
Solution Approach 1:
The patent segments the feature map encoding process into distinct stages: packing feature maps into tensors, quantizing tensors to reduce precision, and encoding using VVC tools. This segmentation allows each stage to be optimized independently, reducing overall computational overhead while maintaining encoding simplicity.
Solution Approach 2:
The patent changes the parameter precision by quantizing tensors from high precision (e.g., 32-bit floating point) to lower precision (e.g., 8-bit integers). This parameter change significantly reduces computational overhead and memory usage while preserving sufficient accuracy for video compression applications.
2Reliability
If feature maps are decoded with high precision, then task performance improves, but memory consumption and processing time increase
Solution Approach 1:
The patent applies quantization to change the precision parameter of decoded feature maps from high precision to optimized lower precision levels. This reduces memory consumption and processing time while maintaining task performance through intelligent precision management that preserves necessary accuracy.
Solution Approach 2:
The patent applies partial precision reduction through selective quantization, maintaining higher precision where needed for task performance while reducing precision in less critical areas. This partial action approach balances memory consumption with task performance requirements.
3Adaptability or versatility
If conventional encoding methods are used for feature maps, then standard video compression tools can be applied, but efficiency in handling varying spatial dimensions and channel counts decreases
Solution Approach 1:
The patent creates a universal encoding framework that handles feature maps with varying spatial dimensions and channel counts through a standardized packing and quantization process. This multi-functional approach maintains adaptability to different CNN architectures while improving encoding efficiency through consistent processing steps.
Solution Approach 2:
The patent transforms feature maps into tensors by adding a batch dimension and organizing data in a standardized multi-dimensional format. This dimensionality change enables efficient handling of varying spatial dimensions and channel counts using uniform encoding operations across all feature map configurations.
Data Source
AI summary
A method of decoding feature maps from encoded data. A plurality of samples is decoded from the encoded data. The feature maps are determined based on one image from at least a first group of samples arranged in a first two-dimensional array and a second group of samples arranged in a second two-dimensional array, where the second two-dimensional array is different from the first two-dimensional array.


