Lightweight deep learning model quantization method
By predicting and updating the quantization scale based on the parameter distribution in previous learning steps, the method addresses the challenge of performance fluctuations in deep learning models on resource-limited mobile devices, achieving efficient and stable learning.
Patent Information
- Application Number
- PCT/KR2023/018255
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-22
AI Technical Summary
Conventional deep learning model quantization methods struggle to adapt to hardware accelerator-based designs in mobile devices, leading to performance fluctuations due to the lack of adaptive quantization technology that considers resource-limited environments.
A method for predicting and updating the quantization scale in the next learning step based on the parameter distribution in the previous learning step, involving the storage and accumulation of maximum deep learning operation values, calculation of their average, and subsequent model quantization.
This approach enables fast and high-performance learning for hardware accelerator models in mobile devices by providing a stable and adaptive quantization method that minimizes performance degradation.
Smart Images

Figure KR2023018255_22052025_PF_FP_ABST
Abstract
Description
A lightweight deep learning model quantization method
[0001] The present invention relates to a method for quantizing a deep learning model, and more particularly, to a method for updating a quantization scale for a hardware accelerator-based deep learning model installed in a mobile device.
[0002] In conventional computer vision, the majority of quantization-based object classification and object detection learning models are software-based models. The core of the quantization technology, real-time parameter value comparison for scale adjustment, is not easy to directly apply to a hardware accelerator-based design environment because it requires an additional process of re-verifying the output parameters for each layer.
[0003] Furthermore, application to mobile devices requires designs that take into account the small learning model size and associated storage memory constraints. However, design technology that satisfies these requirements is lacking. Consequently, the lack of adaptive quantization technology that can accommodate these requirements leads to performance fluctuations depending on the model structure or training data type.
[0004] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide a lightweight deep learning model quantization method that predicts and updates the quantization scale in the next learning step based on the parameter distribution in the previous learning step, as a quantization method applicable to a hardware accelerator model in a mobile device, which is a resource-limited environment.
[0005] A deep learning model quantization method according to one embodiment of the present invention for achieving the above object includes: a step of storing and accumulating maximum values of deep learning operations during training of a deep learning model; a step of calculating an average of the accumulated maximum values; a step of updating a quantization scale with the calculated average; and a step of quantizing a deep learning model based on the updated quantization scale.
[0006] The accumulation step may be to store the maximum values for each batch and accumulate them.
[0007] The calculation step may be to calculate the average of the maximum values stored per epoch.
[0008] The deep learning model quantization method according to the present invention may further include a step of performing learning of the next epoch on the quantized deep learning model.
[0009] Deep learning operations can be convolution operations.
[0010] The update step may update the quantization scale without knowing the distribution of the output feature map by the convolution operation.
[0011] Quantization can be symmetric quantization.
[0012] The accumulation step may be to store the maximum absolute value of the deep learning operation value.
[0013] Deep learning models can be deployed on mobile devices.
[0014] According to another aspect of the present invention, a deep learning operation device is provided, characterized by including: an operation unit that stores and accumulates maximum values of deep learning operations during training of a deep learning model, calculates an average of the accumulated maximum values, updates a quantization scale with the calculated average, and quantizes a deep learning model based on the updated quantization scale; and a memory that provides storage space required for the operation unit.
[0015] According to another aspect of the present invention, a method for quantizing a deep learning model is provided, comprising: a step of updating a quantization scale with an average of maximum values of deep learning operations during training of a deep learning model; a step of quantizing a deep learning model based on the updated quantization scale; and a step of training the quantized deep learning model.
[0016] According to another aspect of the present invention, a deep learning operation device is provided, characterized by including: an operation unit that updates a quantization scale as an average of maximum values of deep learning operations during training of a deep learning model, quantizes the deep learning model based on the updated quantization scale, and trains the quantized deep learning model; and a memory that provides storage space required for the operation unit.
[0017] As described above, according to embodiments of the present invention, by predicting and updating the quantization scale in the next learning step based on the parameter distribution in the previous learning step, fast and high-performance learning is possible in quantization learning of a hardware accelerator model in a mobile device, which is a resource-constrained environment.
[0018] Figure 1. Example of problems that arise in hardware structures when applying quantization learning models in software.
[0019] Figure 2. Example of a learning process based on quantization scale update technology.
[0020] Figure 3. Learning method based on quantization scale update
[0021] Figure 4-5. Comparison of maximum and average distributions when applying the current epoch learning scale based on the previous distribution values.
[0022] Figure 6-9. Example of the scale update process for each epoch and batch unit during the convolution operation.
[0023] Figure 10. Mobile deep learning computing device
[0024] Hereinafter, the present invention will be described in more detail with reference to the drawings.
[0025] Quantization technology is being widely utilized for real-time operation in various mobile device object recognition fields using cameras. Lightweighting is a key challenge, particularly when implementing learning models in memory-constrained environments such as NPU hardware accelerator design.
[0026] However, most quantization models implemented on hardware are implemented after quantization has already been completed after learning has been completed, and there are several problems in implementing a quantization learning model suitable for the hardware structure by directly applying the quantization learning method in software.
[0027] Figure 1 schematically illustrates the process generally followed during quantization learning in software and the problems that arise when the method is applied as is to a hardware accelerator.
[0028] When quantizing the convolution operation part in a CNN model, the quantization scale of the input feature map before convolution and the quantization scale of the output feature map after convolution have different values due to changes in the distribution value caused by the intermediate operation process. In the case of the scale of the output feature map, a real-time value analysis and comparison process is required to quantize the parameters based on the accurate value distribution.
[0029] However, when analyzing output feature map values in real time on hardware, it's necessary to verify all output values after the convolution operation. This comparison process requires the feature map values to be stored in memory again. This process incurs an unconditional additional cycle consumption, which can be a critical issue in hardware accelerator design.
[0030] Therefore, in order to construct a learning model that can be implemented in a hardware accelerator, the scale value for the output feature map must be known in advance. Fig. 2 is a diagram illustrating a model update method based on a quantization scale calculation method applicable to an embodiment of the present invention. A symmetric quantization structure is applied as the quantization structure. In the case of an asymmetric quantization structure, negative and positive values can be separately checked to minimize value loss occurring during quantization, but a separate calculation is added for the intermediate value that serves as a reference, making it unsuitable for hardware design.
[0031] In order to minimize the overall performance reduction of the model due to quantization, when quantization and calculation are performed on the convolution process, which involves the most calculations, the maximum absolute value of the values output after convolution is stored in a buffer.
[0032] The maximum absolute value of the convolution output is updated and maintained during learning within the batch size, and the maximum value is stored as a cumulative sum in a separate buffer at the end of each batch. The value stored as a cumulative sum is calculated as an average value at the end of each epoch of dataset learning (by dividing the cumulative sum by the number of batches) and applied as the quantization scale value for the next epoch. Similarly, the filter value is updated with a quantization scale on an epoch basis.
[0033] FIG. 3 is a diagram illustrating a quantization scale update method according to one embodiment of the present invention.
[0034] As shown, during training of a deep learning model, the maximum absolute value of the convolution output is selected and stored for each batch. Since this is performed batch-by-batch, the maximum values for each batch are accumulated. Once training is complete for an epoch, the accumulated maximum values are divided by the number of batches to calculate the average maximum value.
[0035] The quantization scale is updated with the next calculated maximum value average, and the weights and activation functions of the deep learning model are quantized based on the updated quantization scale, and learning of the next epoch is performed on the quantized deep learning model.
[0036] In this way, in an embodiment of the present invention, quantization scale update is possible without understanding the distribution of the output feature map by convolution.
[0037] FIG. 4 and FIG. 5 show a comparison of the distribution of actual values and the distribution difference of quantized scale values when performing quantized scale update according to the method presented in the embodiment of the present invention. It can be confirmed that the scale value based on the average of the accumulated maximum values for each batch (Mean: red below) has a value suitable for the parameter distribution in the current epoch learning even though the feature map distribution was not identified in real time. On the other hand, there is a problem that the scale updated with only the simple maximum value (Prev. Xpoch Abs Max: blue above)) may be updated based on an excessively large value, which may cause a problem of performance degradation during learning.
[0038] This quantization scale update, based on the average of the maximum values per batch, cannot be performed in real time, as is possible with software. However, it allows for the determination of average values across multiple image data sets and a rough understanding of distributional changes based on overall learning tendencies. Furthermore, this eliminates the need for real-time distribution parameter determination, making hardware accelerator implementation feasible.
[0039] Figure 6-9 illustrates a more detailed process for a quantization scale update method according to an embodiment of the present invention, with respect to the convolution operation process. In limited memory environments such as mobile devices, the batch size, i.e. the number of image data that can be learned at one time, is necessarily small.
[0040] This means that the number of batch learning operations per unit epoch is large, and the scale update technique, which is calculated based on the average of the maximum values for each batch, can derive an average for relatively more sample values in the limited environment, so it can be a stable method with a low influence from the bias toward high values in a specific batch.
[0041] FIG. 10 is a diagram illustrating the configuration of a mobile deep learning computing device according to another embodiment of the present invention. As illustrated, the mobile deep learning computing device according to an embodiment of the present invention comprises a communication interface (110), a deep learning computing unit (120), and a memory (130).
[0042] The communication interface (110) communicates with an external host system to receive data sets and parameters of a pre-trained deep learning model. The deep learning operator (120) quantizes and trains the loaded deep learning model according to the method presented in FIG. 3 described above. The memory (130) provides the storage space necessary for the deep learning operator (120) to perform calculations.
[0043] So far, we have described in detail a preferred embodiment of a lightweight deep learning model quantization method.
[0044] In the above embodiment, a method for quantization learning of a hardware accelerator model in a mobile device, which is an environment with limited resources, is presented by predicting and updating the average of the absolute values of the maximum values of deep learning operations per batch in the previous epoch at a quantization scale.
[0045] This enables fast and high-performance learning through a universal quantization scaling technique applicable to limited environment-based mobile devices as a quantization model applicable to hardware accelerator models.
[0046] Meanwhile, it goes without saying that the technical idea of the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.
[0047] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. A step of storing and accumulating the maximum values of deep learning operations during training of a deep learning model; A step of calculating the average of the accumulated maximum values; A step of updating the quantization scale with the computed average; and A method for quantizing a deep learning model, characterized by including a step of quantizing a deep learning model based on an updated quantization scale.
2. In claim 1, The cumulative stage is, A deep learning model quantization method characterized by storing and accumulating maximum values by batch.
3. In claim 2, The calculation steps are: A deep learning model quantization method characterized by calculating the average of the maximum values stored for each epoch.
4. In claim 3, A method for quantizing a deep learning model, characterized in that it further includes a step of performing learning of the next epoch for the quantized deep learning model.
5. In claim 1, Deep learning operations are, A deep learning model quantization method characterized by a convolution operation.
6. In claim 5, The update steps are: A deep learning model quantization method characterized by updating the quantization scale without understanding the distribution of the output feature map by the convolution operation.
7. In claim 1, Quantization is, A deep learning model quantization method characterized by symmetric quantization.
8. In claim 7, The cumulative stage is, A deep learning model quantization method characterized by storing the maximum value of the absolute value of deep learning operation values.
9. In claim 1, The deep learning model is, A method for quantizing a deep learning model, characterized by being installed on a mobile device.
10. An operator that stores and accumulates the maximum values of deep learning operations during training of a deep learning model, calculates the average of the accumulated maximum values, updates the quantization scale with the calculated average, and quantizes the deep learning model based on the updated quantization scale; and A deep learning computing device characterized by including a memory that provides storage space required for the computing device.
11. A step of updating the quantization scale with the average of the maximum values of deep learning operations during training of the deep learning model; A step of quantizing a deep learning model based on the updated quantization scale; and A method for quantizing a deep learning model, characterized by comprising a step of training a quantized deep learning model.
12. An operator that updates the quantization scale as the average of the maximum values of deep learning operations during training of a deep learning model, quantizes the deep learning model based on the updated quantization scale, and trains the quantized deep learning model; and A deep learning computing device characterized by including a memory that provides storage space required for the computing device.
Citation Information
Patent Citations
Neural network quantification method and device and computer readable storage medium
CN111401518A
Article transfer device
KR1020220102077A
Module for LED Electric Light Board with Improved Waterproofing Property
KR102274013B1
Disposable dish for sliced raw fish
KR102309916B1
Smart hood for kitchen
KR102381708B1