Neural network model quantification method and electronic device
By calculating the quantization error of the baseline set of quantization sensitivity of each layer of the neural network model and the candidate bit width strategy, the bit width strategy is optimized, and the problem of maintaining the accuracy of the neural network model under limited computing resources is solved, achieving efficient and accurate quantization.
Patent Information
- Application Number
- CN202411918992.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to maintain the accuracy and efficiency of the model under limited computing resources when quantifying neural network models, especially when deployed on devices with limited computing power and energy resources.
By calculating the quantized sensitivity baseline sets for each layer and calculating the quantization error of candidate bit width strategies based on these baseline sets, the optimization algorithm is used to select the best bit width strategy to ensure that the neural network model remains efficient and accurate after quantization.
This method can improve the search performance of hybrid precision quantization algorithm without adding additional calculations, reduce model size and calculation complexity, while maintaining the accuracy of the neural network, and achieving efficient and accurate quantization.
Smart Images

Figure CN119990201A_ABST