Transform-Based Feature Map Encoding for Deep Learning Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep-neural networks face challenges in efficiently compressing and processing large amounts of data from feature maps, particularly in industrial applications where communication between machines is involved, requiring improved compression efficiency and reduced data processing.
Innovation Solution
A method for encoding and decoding transform-based feature maps by extracting optimal basis vectors, acquiring transform coefficients, and generating bitstreams that include these components, allowing for efficient transmission and reconstruction of feature maps, with the option to use fixed common basis vectors or transform coefficients agreed upon by encoding and decoding apparatuses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If deep-neural networks are applied to industrial machinery with large amounts of feature map data, then the network can perform machine vision tasks, but the data processing burden and communication requirements increase significantly
Solution Approach 1:
The patent extracts only the essential feature map information (transform coefficients and basis vectors) from the complete feature map data for transmission and storage. By selecting and transmitting only the critical components needed to reconstruct the feature map, the system reduces the quantity of transmitted data while maintaining the ability to perform machine vision tasks.
Solution Approach 2:
The feature map data is segmented into distinct components: transform coefficients and basis vectors. This segmentation allows the system to process, transmit, and store these components separately, reducing the overall data burden while preserving the functional integrity needed for machine vision applications.
2Measurement precision
If complete feature map data is transmitted and processed, then accurate machine vision results can be achieved, but communication bandwidth and processing time increase
Solution Approach 1:
The patent extracts and transmits only the essential transform coefficients and basis vectors required to reconstruct the feature map, eliminating the need to transmit the complete feature map data. This extraction approach maintains measurement precision for machine vision tasks while significantly reducing data processing time and communication bandwidth requirements.
3Quantity of substance
If transform-based encoding is applied to feature maps, then compression efficiency improves, but the complexity of encoding and decoding processes increases
Solution Approach 1:
The patent performs transform encoding (such as DCT or DST) on the feature map data during the encoding phase to convert spatial domain data into frequency domain coefficients. This preliminary transformation consolidates the data in a form that is more compressible and can be efficiently transmitted, reducing the compressed data size while the decoding side simply needs to perform the inverse transform.
Data Source
AI summary
Disclosed herein are a method, an apparatus and a storage medium for encoding/decoding using a transform-based feature map. An optimal basis vector is extracted from one or more feature maps, and a transform coefficient is acquired through a transform using the basis vector. The basis vector and the transform coefficient may be transmitted through a bitstream. In an embodiment, one or more feature maps are reconstructed using the basis vector and the transform coefficient, which are decoded from the bitstream.


