DNN Quantization Parameter Optimization via Back-Propagation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep Neural Networks (DNNs) face challenges in efficiently representing values in hardware due to the trade-off between accuracy and resource efficiency, as using fewer bits for representation reduces accuracy but increases efficiency, and existing methods lack a systematic approach to identify optimal fixed point number formats for DNN implementation.
Innovation Solution
The method involves determining quantization parameters for DNNs by simulating quantization of values to fixed point number formats using quantization blocks, calculating a cost metric combining error and size metrics, back-propagating gradients to adjust these parameters, and optimizing them to balance accuracy and resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If fewer bits are used to represent values in DNN hardware, then resource efficiency is improved, but accuracy deteriorates
Solution Approach 1:
The patent applies parameter changes by systematically varying the bit-width (quantization parameter) of fixed-point number formats to find the optimal balance between resource efficiency and accuracy. The method adjusts parameters such as integer bits and fractional bits in Q-format representations, and applies scaling factors and offset values to transform floating-point weights into fixed-point formats that maintain accuracy while reducing hardware resource requirements.
2Ease of manufacture
If existing methods are used to represent values in hardware, then implementation is simpler, but optimal fixed point number formats cannot be identified
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing optimal fixed-point number formats in lookup tables before hardware implementation. The method pre-determines the best Q-format parameters (integer bits, fractional bits, scaling factors) for different DNN layers and operating conditions, so that during actual hardware operation, the hardware can directly use these pre-optimized formats without complex real-time calculations, thus achieving both simplicity and optimality.
3Measurement precision
If floating point number formats are used, then accuracy is maintained, but hardware resource consumption increases
Solution Approach 1:
The patent applies mechanics substitution by replacing floating-point number format operations with fixed-point number format operations in hardware. The method transforms the mathematical representation from floating-point (which requires complex exponent and mantissa handling) to fixed-point (which uses simpler integer arithmetic with Q-format notation), thereby substituting a resource-intensive mechanical system with a more efficient one while maintaining computational accuracy through proper scaling and quantization.
Data Source
AI summary
Methods and systems for identifying quantisation parameters for a Deep Neural Network (DNN). The method includes determining an output of a model of the DNN in response to training data, the model of the DNN comprising one or more quantisation blocks configured to transform a set of values input to a layer of the DNN prior to processing the set of values in accordance with the layer, the transformation of the set of values simulating quantisation of the set of values to a fixed point number format defined by one or more quantisation parameters; determining a cost metric of the DNN based on the determined output and a size of the DNN based on the quantisation parameters; back-propagating a derivative of the cost metric to one or more of the quantisation parameters to generate a gradient of the cost metric for each of the one or more quantisation parameters; and adjusting one or more of the quantisation parameters based on the gradients.


