Quantization Threshold Fusion for Deep Neural Network Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep neural networks used in image processing on low-bit-width hardware platforms face reduced accuracy due to the limitations of saturated and unsaturated mapping methods, which are not suitable for all activation output layers, leading to a loss of core features.
Innovation Solution
An image processing method that calculates first and second quantization threshold values using saturated and unsaturated mapping methods respectively, performs weighted calculations to obtain an optimal threshold value, and quantifies the network model using this value, enabling effective retention of features across most activation output layers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If saturated mapping or unsaturated mapping method is used to obtain quantization threshold, then the quantization process can be completed, but the accuracy of image processing is greatly reduced due to loss of core features in activation output layers
Solution Approach 1:
The patent combines saturated mapping and unsaturated mapping methods into a unified quantization threshold calculation formula. The formula integrates both mapping approaches with adjustable parameters (α and β) to balance their respective advantages, thereby retaining core features while completing the quantization process for low-bit-width hardware deployment
Solution Approach 2:
The patent introduces adjustable parameters α and β in the quantization threshold calculation formula. By changing these parameters, the system can adapt the quantization process to different activation output layers and hardware platforms, optimizing the balance between quantization completion and feature retention to maintain image processing accuracy
2Device complexity
If a single mapping method is used for all activation output layers, then the quantization process is simplified, but it is not suitable for some activation output layers that cannot retain core features
Solution Approach 1:
The patent develops a universal quantization threshold calculation formula that can be applied to different activation output layers across various deep neural network architectures. The formula incorporates both saturated and unsaturated mapping methods with adjustable parameters, making it adaptable to different hardware platforms and network configurations without requiring layer-specific custom methods
Data Source
AI summary
An image processing method is provided. Since in the present invention, the first quantization threshold obtained by the saturated mapping method and the second quantization threshold obtained by the unsaturated mapping method are weighted, it is equivalent to fusing two quantization threshold values. The obtained optimal quantization threshold can be applied to most activations, to more effectively retain the effective information of the activations and use in subsequent image processing, thereby improving the accuracy of inference computations of the quantified deep neural network on low-bit-width hardware platforms. An image processing device and apparatus have the same beneficial effects as the above image processing method.

