Neural Network Quantization Parameter Determination Method
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Neural networks require high-precision data representation, which leads to large storage space and processing bandwidth demands, increasing costs and power consumption.
Innovation Solution
A method for adjusting data bit width through quantization, involving obtaining a data bit width, performing quantization, comparing quantization errors, and adjusting the bit width based on these errors to convert high-precision data into low-precision fixed-point data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-precision data representation is used in neural networks, then measurement precision is improved, but volume of stationary object increases
Solution Approach 1:
The patent changes the precision parameter of data from high-precision floating-point format to low-precision fixed-point format through quantization. By adjusting the bit width parameter during quantization, the system transforms data representation while maintaining acceptable precision levels, thereby reducing storage space requirements.
Solution Approach 2:
The patent creates a quantized copy of the original high-precision data. This quantized version serves as an approximation that occupies less storage space while preserving the essential information needed for neural network operations, effectively replacing the need to store full-precision data.
2Measurement precision
If high-precision data representation is used in neural networks, then measurement precision is improved, but use of energy by stationary object increases
Solution Approach 1:
The patent changes the data precision parameter to reduce power consumption. By using lower precision fixed-point arithmetic instead of high-precision floating-point operations, the energy required for computation is significantly reduced, making the system more energy-efficient.
Solution Approach 2:
The patent uses quantized data copies that require less computational resources to process. These lower-precision representations reduce the power consumption associated with data processing while maintaining sufficient accuracy for neural network tasks.
3Volume of stationary object
If data bit width is reduced through quantization, then volume of stationary object decreases, but measurement precision deteriorates
Solution Approach 1:
The patent employs dynamic quantization where the bit width and precision levels are adjusted based on the specific requirements of different neural network layers and operations. This dynamic adjustment allows the system to use lower precision where acceptable and maintain higher precision where needed, optimizing the trade-off between storage space and precision.
Solution Approach 2:
The patent applies different quantization precision levels to different parts of the neural network based on their specific requirements. Critical layers that require high precision maintain higher bit widths, while less critical layers use lower precision, creating a locally optimized precision distribution across the network.
4Use of energy by stationary object
If data bit width is reduced through quantization, then use of energy by stationary object decreases, but measurement precision deteriorates
Solution Approach 1:
The patent uses dynamic precision adjustment to balance energy consumption and precision requirements. By adaptively selecting appropriate precision levels for different computational tasks, the system minimizes energy consumption while ensuring that precision requirements are met for critical operations.
Solution Approach 2:
The patent applies differentiated precision strategies to different neural network components. Energy-intensive operations that can tolerate lower precision use reduced bit widths, while operations requiring high precision maintain higher precision levels, optimizing the energy-precision trade-off locally across the system.
Data Source
Figure 1~2
Figure 3~4
Figure 5A~5B
AI summary
The technical solution involves a board card including a storage component, an interface apparatus, a control component, and an artificial intelligence chip. The artificial intelligence chip is connected to the storage component, the control component, and the interface apparatus, respectively; the storage component is used to store data; the interface apparatus is used to implement data transfer between the artificial intelligence chip and an external device; and the control component is used to monitor a state of the artificial intelligence chip. The board card is used to perform an artificial intelligence operation.