Neural network model quantification method and electronic device

By calculating the quantization error of the baseline set of quantization sensitivity of each layer of the neural network model and the candidate bit width strategy, the bit width strategy is optimized, and the problem of maintaining the accuracy of the neural network model under limited computing resources is solved, achieving efficient and accurate quantization.

CN119990201APending Publication Date: 2025-05-13SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411918992.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to maintain the accuracy and efficiency of the model under limited computing resources when quantifying neural network models, especially when deployed on devices with limited computing power and energy resources.

Method used

By calculating the quantized sensitivity baseline sets for each layer and calculating the quantization error of candidate bit width strategies based on these baseline sets, the optimization algorithm is used to select the best bit width strategy to ensure that the neural network model remains efficient and accurate after quantization.

Benefits of technology

This method can improve the search performance of hybrid precision quantization algorithm without adding additional calculations, reduce model size and calculation complexity, while maintaining the accuracy of the neural network, and achieving efficient and accurate quantization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990201A_ABST
    Figure CN119990201A_ABST
Patent Text Reader

Abstract

The invention discloses a quantification method of a neural network model and an electronic device. The quantization method comprises: for each of a plurality of layers of a neural network model, calculating a quantization sensitivity baseline set, where the quantization sensitivity baseline set comprises a plurality of quantization sensitivity baselines, where each of the plurality of quantization sensitivity baselines is based on a different respective bit width set, respectively; calculating a quantization error of a candidate bit width policy based on the quantization sensitivity baseline set, where the candidate bit width policy allocates a target bit width for each of the plurality of layers; selecting a bit width for each of the plurality of layers based on a quantization error; and processing the multimedia data using a neural network model based on the selected bit width.
Need to check novelty before this filing date? Find Prior Art