Neural Network Quantization Device with Optimized Division Position
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy of neural network calculations decreases when using quantized variables, leading to reduced recognition performance, as the transition from floating-point to fixed-point numbers compromises both computation efficiency and accuracy.
Innovation Solution
An information processing device optimizes division positions for quantization to minimize quantization errors, allowing for the conversion of floating-point data into fixed-point data while maintaining calculation accuracy, by setting thresholds that reduce differences between pre- and post-quantization values, thereby improving recognition rates and computation efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If quantization is applied to convert floating-point data to fixed-point data for neural network calculations, then computation efficiency is improved, but calculation accuracy decreases
Solution Approach 1:
The patent changes the parameter of division positions in quantization from fixed to dynamically optimized values. By adjusting the division positions based on the actual data distribution and neural network requirements, the quantization process achieves better precision while maintaining computational efficiency. This parameter optimization allows the system to find the best trade-off point between speed and accuracy for each specific calculation scenario.
Solution Approach 2:
The patent introduces dynamic adjustment of quantization division positions during the neural network training and inference process. Instead of using static quantization parameters, the system adaptively modifies division positions based on data characteristics and performance requirements, enabling the quantization scheme to evolve and optimize itself for different computational contexts and accuracy demands.
2Use of energy by moving object
If quantization is applied to reduce computational resources and power consumption, then energy efficiency is improved, but recognition performance decreases
Solution Approach 1:
The patent optimizes quantization parameters including division positions and precision levels based on the specific requirements of different neural network layers and tasks. By dynamically adjusting these parameters, the system achieves significant energy savings while maintaining recognition performance above acceptable thresholds, effectively decoupling energy consumption from performance degradation.
Solution Approach 2:
The patent applies different quantization strategies and precision levels to different parts of the neural network based on their importance and sensitivity. Critical layers maintain higher precision with finer division positions, while less critical layers use coarser quantization, creating a localized quality distribution that optimizes the overall energy-performance trade-off across the entire system.
Data Source
AI summary
An information processing device that executes calculation of a neural network, includes a memory; and a processor coupled to the memory and the processor configured to: set a division position for quantization of a variable to be used for the calculation so that a quantization error based on a difference between the variable before the quantization and the variable after the quantization is reduced; and quantize the variable based on the division position set.


