Transform-Based Feature Map Encoding for Deep Learning Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep-neural networks face challenges in efficiently compressing and processing large amounts of data from feature maps, particularly in industrial applications where communication between machines is involved, requiring improved compression efficiency and reduced data processing.

Innovation Solution

A method for encoding and decoding transform-based feature maps by extracting optimal basis vectors, acquiring transform coefficients, and generating bitstreams that include these components, allowing for efficient transmission and reconstruction of feature maps, with the option to use fixed common basis vectors or transform coefficients agreed upon by encoding and decoding apparatuses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If deep-neural networks are applied to industrial machinery with large amounts of feature map data, then the network can perform machine vision tasks, but the data processing burden and communication requirements increase significantly

Engineering Contradiction:
Improvemachine vision task performanceVSAvoidfeature map data volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential feature map information (transform coefficients and basis vectors) from the complete feature map data for transmission and storage. By selecting and transmitting only the critical components needed to reconstruct the feature map, the system reduces the quantity of transmitted data while maintaining the ability to perform machine vision tasks.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The feature map data is segmented into distinct components: transform coefficients and basis vectors. This segmentation allows the system to process, transmit, and store these components separately, reducing the overall data burden while preserving the functional integrity needed for machine vision applications.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If complete feature map data is transmitted and processed, then accurate machine vision results can be achieved, but communication bandwidth and processing time increase

Engineering Contradiction:
Improvemachine vision accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and transmits only the essential transform coefficients and basis vectors required to reconstruct the feature map, eliminating the need to transmit the complete feature map data. This extraction approach maintains measurement precision for machine vision tasks while significantly reducing data processing time and communication bandwidth requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

3Quantity of substance

If transform-based encoding is applied to feature maps, then compression efficiency improves, but the complexity of encoding and decoding processes increases

Engineering Contradiction:
Improvecompressed data sizeVSAvoidencoding/decoding system complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent performs transform encoding (such as DCT or DST) on the feature map data during the encoding phase to convert spatial domain data into frequency domain coefficients. This preliminary transformation consolidates the data in a form that is more compressible and can be efficiently transmitted, reducing the compressed data size while the decoding side simply needs to perform the inverse transform.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240078710A1Method, apparatus and storage medium for encoding/decoding using transform-based feature map
Publication Date: 2024.03.07 ELECTRONICS & TELECOMM RES INST
  • US20240078710A1 patent drawing
  • US20240078710A1 patent drawing
  • US20240078710A1 patent drawing

AI summary

Disclosed herein are a method, an apparatus and a storage medium for encoding/decoding using a transform-based feature map. An optimal basis vector is extracted from one or more feature maps, and a transform coefficient is acquired through a transform using the basis vector. The basis vector and the transform coefficient may be transmitted through a bitstream. In an embodiment, one or more feature maps are reconstructed using the basis vector and the transform coefficient, which are decoded from the bitstream.