Feature Map Encoding in Cascaded Neural Layers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image and video compression technologies face inefficiencies in encoding data across multiple processing layers, particularly in hybrid codecs and machine learning applications, where scalability and adaptability to content are limited, leading to suboptimal bitstream generation and decoding processes.
Innovation Solution
The method involves processing data through cascaded layers with varying resolutions, selecting a layer different from the one generating the lowest resolution feature map, and inserting information related to this selected layer into the bitstream, which includes downsampling using techniques like average or max pooling, and convolutional operations to enhance adaptivity and reduce data complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If only the layer generating the lowest resolution feature map is used for encoding, then the bitstream size is minimized, but the encoding efficiency and scalability are reduced
Solution Approach 1:
The patent segments the feature map data from multiple layers with different resolutions and selectively encodes only the necessary portions. By dividing the encoding task across layers and selectively transmitting data from layers other than the lowest resolution layer, the system achieves better encoding efficiency while controlling bitstream size through selective encoding rather than encoding all layers uniformly.
Solution Approach 2:
The patent applies partial action by encoding only the necessary information from intermediate layers rather than all layers. The method selects which layers to encode based on the specific needs of the application, allowing efficient encoding when needed while avoiding unnecessary data transmission when not needed, thus optimizing the balance between encoding efficiency and bitstream size.
2Productivity
If data from multiple layers with different resolutions is encoded, then encoding efficiency and scalability are improved, but the bitstream complexity increases
Solution Approach 1:
The patent introduces dynamic adaptability by allowing the encoder to select which layers to encode based on content characteristics and application requirements. The system can dynamically adjust the encoding strategy by choosing to encode data from intermediate layers when scalability is needed, while avoiding encoding from layers with lowest resolution when simplicity is preferred, thus managing bitstream complexity adaptively.
Solution Approach 2:
The patent applies local quality by differentiating the encoding treatment of different layers based on their specific characteristics. Rather than uniformly encoding all layers, the system identifies which layers provide the most valuable information for the specific application and encodes only those portions, reducing unnecessary complexity while maintaining encoding efficiency where beneficial.
3Loss of information
If feature maps from all layers are transmitted, then complete information is preserved, but the data volume and processing complexity increase
Solution Approach 1:
The patent extracts only the necessary information from intermediate layers rather than transmitting all feature map data from all layers. By identifying and extracting only the essential information needed for the application, the system preserves critical information while significantly reducing data volume, achieving a balance between information completeness and data efficiency.
Data Source
AI summary
The present disclosure relates to methods and apparatuses for encoding data for (still or video processing into a bitstream). In particular, the data are processed by a network which includes a plurality of cascaded layers. In the processing, feature maps are generated by the layers. The feature maps processed (output) by at least two different layers have different resolutions. In the processing, a layer is selected, out of the cascaded layers, which is different from the layer generating the feature map of the lowest resolution (e.g. latent space). The bitstream includes information related to the selected layer. With this approach, scalable processing which may operate on different resolutions is provided so that the bitstream may convey information relating to such different resolutions. Accordingly, the data may be efficiently coded within the bitstream, depending on the resolution which may vary depending on the content of the picture data coded.


