Emulated Non-Uniform Quantization for Neural Network Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep neural networks require significant computational and storage resources, making them challenging for deployment on mobile and embedded devices, and existing quantization methods are inefficient in reducing storage and computational costs without impacting accuracy.

Innovation Solution

A system and method for training and using a non-uniformly quantized neural network by applying emulated non-uniformly quantized transformations using uniformly distributed random noise, allowing for reduced storage and computational requirements while maintaining accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If quantization is applied to reduce storage and computational cost, then resource requirements are reduced, but accuracy of the neural network output deteriorates

Engineering Contradiction:
Improvestorage requirementsVSAvoidaccuracy of neural network output
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent applies different quantization strategies to different parts of the neural network. Specifically, it uses non-uniform quantization where the quantization step size varies depending on the magnitude of the weight values, allowing finer precision for important weights and coarser precision for less important ones, thus maintaining accuracy while reducing overall storage requirements

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the quantization parameters dynamically based on the distribution of weight values. It computes statistics such as mean and standard deviation of weight values and uses these to determine appropriate quantization levels and step sizes, adapting the quantization process to the specific characteristics of each layer and weight distribution

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If uniform quantization is used to simplify the quantization process, then ease of implementation is improved, but accuracy deteriorates due to loss of precision in representing weight values

Engineering Contradiction:
Improveease of quantization implementationVSAvoidprecision of weight value representation
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

Instead of applying a single uniform quantization step size across all weight values, the patent divides the weight value range into multiple intervals with different step sizes. Smaller step sizes are applied to weight values near the mean (where precision is most important), while larger step sizes are applied to extreme values, thus improving precision where needed while maintaining implementation simplicity

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent computes quantization parameters such as step size and number of levels based on the statistical properties of the weight values (mean, standard deviation, min, max). This adaptive approach allows the quantization scheme to match the actual distribution of weights, improving precision without requiring complex manual tuning

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If 32-bit floating point representation is used for weight values, then accuracy of neural network output is maintained, but storage requirements and computational cost increase

Engineering Contradiction:
Improveaccuracy of neural network outputVSAvoidstorage requirements
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies non-uniform quantization where the precision allocated to each weight value depends on its importance. Weight values near the mean (which have greater impact on network output) are represented with finer precision (smaller quantization step), while extreme values are represented with coarser precision, thus maintaining output accuracy while reducing overall storage requirements compared to uniform 32-bit floating point

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Instead of applying full 32-bit precision to all weight values, the patent applies partial precision selectively - using higher precision only where necessary (for weights with greater impact on output) and lower precision elsewhere, achieving a balance between accuracy and storage efficiency

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11972347B2System and method for emulating quantization noise for a neural network
Publication Date: 2024.04.30 RAMOT AT TEL AVIV UNIVERSITY LTD
  • US11972347B2 patent drawing
  • US11972347B2 patent drawing
  • US11972347B2 patent drawing

AI summary

A system for training a quantized neural network dataset, comprising at least one hardware processor adapted to: receive input data comprising a plurality of training input value sets and a plurality of target value sets; in each of a plurality of training iterations: for each layer, comprising a plurality of weight values, of one or more of a plurality of layers of a neural network: compute a set of transformed values by applying to a plurality of layer values one or more emulated non-uniformly quantized transformations by adding to each of the plurality of layer values one or more uniformly distributed random noise values; and compute a plurality of output values; compute a plurality of training output values; and update one or more of the plurality of weight values to decrease a value of a loss function; and output the updated plurality of weight values of the plurality of layers.