Temporal Context Encoding for Multi-Scale Feature Map Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies fail to effectively encode and decode neural networks-based multi-scale feature maps due to their high-dimensional nature, leading to inefficient processing and transmission of data associated with the spatial redundancy of natural images, and thus, may require new encoding and decoding techniques.
Innovation Solution
An encoding method that includes generating a reduced feature map by dimension reduction, channel-wise scaling, extracting prior information, and estimating a probability distribution based on temporal context, and a decoding method that involves receiving an input stream, acquiring decoded prior information, and reconstructing the feature map by reconstructing the context.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multilayer feature maps are used for machine vision and image processing, then accuracy and performance in tasks such as object recognition, tracking, and segmentation are improved, but the high-dimensional nature of the data makes efficient processing and transmission challenging
Solution Approach 1:
The patent segments the high-dimensional multilayer feature map into multiple reduced feature maps with lower dimensions. Each reduced feature map retains essential feature information while having reduced spatial or channel dimensions, making the data more manageable for processing and transmission while preserving the accuracy needed for machine vision tasks.
Solution Approach 2:
The patent transforms the high-dimensional feature map by reducing its dimensions through pooling operations. This dimensionality reduction converts the complex high-dimensional data structure into a more compact form that is easier to process and transmit, while the temporal context extraction preserves the essential information needed for maintaining accuracy in object recognition and other vision tasks.
2Productivity
If image compression techniques are applied to multilayer feature maps, then data transmission efficiency may be improved, but the techniques are ineffective because multilayer feature maps have different characteristics and structures from general images
Solution Approach 1:
The patent changes the parameters and structure of the feature maps by applying pooling operations to generate reduced feature maps with different dimensional characteristics. This transformation adapts the data structure to be more suitable for compression techniques, bridging the gap between the unique characteristics of multilayer feature maps and the requirements of efficient data transmission.
Solution Approach 2:
The patent introduces reduced feature maps as an intermediary representation between the original multilayer feature maps and the compression process. These reduced feature maps serve as a bridge that preserves essential information while having a structure more amenable to efficient encoding and transmission, thus enabling effective compression techniques to be applied.
3Loss of information
If high-dimensional multilayer feature maps are processed directly, then complete feature information is preserved, but processing and transmission become inefficient
Solution Approach 1:
The patent segments the high-dimensional feature map into multiple reduced feature maps through pooling operations. This segmentation reduces the computational burden of processing while distributing the feature information across multiple lower-dimensional representations, thereby maintaining feature completeness without requiring direct processing of the entire high-dimensional structure.
Solution Approach 2:
The patent performs preliminary dimensionality reduction by generating reduced feature maps before further processing or transmission. This preliminary action of creating compact representations upfront reduces the time required for subsequent processing operations while preserving the essential feature information needed for accurate machine vision tasks.
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
An encoding method includes generating a current reduced feature map by reducing a dimension of a current multilayer feature map, generating a scaled current reduced feature map by performing channel-wise scaling on the current reduced feature map, extracting prior information based on the scaled current reduced feature map, acquiring a scaled previous reduced feature map, extracting a temporal context based on the prior information and the scaled previous reduced feature map, estimating a probability distribution of the current reduced feature map based on the temporal context, and encoding the current reduced feature map based on the probability distribution.