Method for quantizing ultra-light network model, and low-power inference accelerator applying same

The proposed end-to-end low-bit quantization method addresses inefficiencies in conventional network model quantization by using multifactors for integer-only operations, resulting in reduced energy and area consumption, and enabling efficient operation on edge devices.

WO2025127188A1PCT designated stage expired Publication Date: 2025-06-19KOREA ELECTRONICS TECH INST

Patent Information

Application Number
PCT/KR2023/020488
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2023-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Conventional network model quantization methods for object detection models are inefficient for edge devices due to high hardware resource requirements, long latency, and large area consumption, as they rely on activation functions and real number operations.

Method used

An end-to-end low-bit quantization method that calculates and stores multifactors for simultaneous quantized convolution and batch normalization operations, using only integer operations to reduce computational overhead and power consumption.

Benefits of technology

The method significantly reduces the area and energy requirements of deep learning accelerators, enabling efficient operation on edge devices with limited battery life and allowing for the creation of ultra-small edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023020488_19062025_PF_FP_ABST
    Figure KR2023020488_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method for quantizing an ultra-light network model and a low-power inference accelerator applying same. A quantized deep learning operation method according to an embodiment of the present invention calculates and stores multi-factors to be used for performing a quantized convolution operation and a quantized batch normalization operation at one time, and performs the convolution operation and the batch normalization operation at one time by using the stored multi-factors. Accordingly, a deep learning operation accelerator can be efficiently operated in a limited battery environment, thereby ensuring a longer battery time in an edge computing environment having a high battery restriction, and remarkably reducing the area of the accelerator to be mounted on an ultra-small edge device.
Need to check novelty before this filing date? Find Prior Art

Description

An ultra-lightweight network model quantization method and a low-power inference accelerator using the method.

[0001] The present invention relates to network model quantization, and more particularly, to an end-to-end low-bit quantization method optimized for hardware and deep learning computational hardware applying the same.

[0002] Conventional object detection model quantization methods for edge devices present several challenges in creating object detection models that can be accelerated on actual edge devices. Conventional quantization methods primarily utilize activation functions, which require intensive hardware resources. Consequently, hardware processing of these activation functions requires long latency and a large area. Furthermore, they do not quantize all layers, but instead use real numbers only in the first, last, and shortcut layers.

[0003] Consequently, hardware must incorporate both integer and real-valued operators, resulting in losses in area, speed, and energy. In other words, conventional quantization methods were designed without considering hardware, making them difficult to apply to real-world edge devices.

[0004] Because of these problems, it is difficult to design an efficient object detection model accelerator, and in reality, the number of object detection model accelerator technologies is very small compared to classification model accelerator technologies.

[0005] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide a lightweight deep learning network model quantization method that can be accelerated in an actual edge device and an accelerator that can operate efficiently in a limited battery environment by applying the method.

[0006] A quantized deep learning operation method according to one embodiment of the present invention for achieving the above object includes a step of calculating and storing multifactors to be used for performing a quantized convolution operation and a quantized batch normalization operation at the same time; and a step of performing a convolution operation and a batch normalization operation at the same time using the stored multifactors.

[0007] Multifactors can be calculated using values ​​that combine quantized activation values ​​and quantized weight values.

[0008] Multifactors can be computed by further utilizing parameters for batch normalization operations.

[0009] The parameters for the batch normalization operation can be applied equally to all multifactors.

[0010] The operation step may include a first summing step of multiplying the corresponding numbers by the multifactors and summing them.

[0011] The quantized deep learning operation method according to the present invention may further include a second summing step of adding AddFactor to the summing result of the first summing step.

[0012] The saving step may be to quantize and save the calculated multifactors and AddFactor.

[0013] Multifactors and AddFactor can be integer values.

[0014] The storage and computation steps can be performed on accelerators mounted on edge devices.

[0015] According to another aspect of the present invention, a deep learning operation accelerator is provided, characterized by including: a memory storing multifactors to be used for performing a quantized convolution operation and a quantized batch normalization operation at the same time; and an operation unit performing a convolution operation and a batch normalization operation at the same time using the multifactors stored in the memory.

[0016] According to another aspect of the present invention, a quantized deep learning operation method is provided, characterized by including a step of inputting multifactors to be used for performing a quantized convolution operation and a quantized batch normalization operation at the same time; and a operation step of performing a convolution operation and a batch normalization operation at the same time using the input multifactors.

[0017] According to another aspect of the present invention, a method for storing deep learning operation parameters is provided, comprising: a step of calculating multifactors to be used for performing a quantized convolution operation and a quantized batch normalization operation at the same time; a step of storing the calculated multifactors in a memory; and a step of storing the stored multifactors, characterized in that the stored multifactors are used for performing the convolution operation and the batch normalization operation at the same time.

[0018] As described above, according to embodiments of the present invention, a deep learning computational accelerator can operate efficiently in a limited battery environment through lightweight deep learning network model quantization that can be accelerated in an actual edge device, thereby ensuring a longer battery time in an edge computing environment with significant battery constraints, and drastically reducing the area of ​​the accelerator, enabling it to be mounted on an ultra-small edge device.

[0019] Figure 1 shows the accelerator energy and area by operation type.

[0020] Figure 2 is a convolution operation method using an index combination operation.

[0021] Figures 3 to 5 illustrate a multifactor and AddFactor calculation method for performing a convolution operation together with a batch normalization operation at the same time.

[0022] Figure 6 is a configuration of deep learning accelerator hardware according to another embodiment of the present invention.

[0023] Hereinafter, the present invention will be described in more detail with reference to the drawings.

[0024] As computer vision algorithms advance, efforts to accelerate computer vision models on edge devices are increasing. In particular, object detection models are being used in diverse fields such as robot vacuum cleaners, factory automation, and smart cities.

[0025] Because edge devices must accelerate object detection models in battery-constrained environments, long battery life is a key selling point for edge devices. Therefore, accelerating object detection models at low power is crucial for long battery life.

[0026] The energy and area required to accelerate object detection models vary significantly depending on the operations used. Figure 1 shows the energy and area by operation type. Figure 1 shows the relative energy / area costs of different operations, based on an 8-bit add operation. For example, an 8-bit add operation is 30 times more energy efficient and 116 times more area efficient than a 32-bit FP add operation. In other words, lightweighting is essential for accelerating object detection models at low power levels, and excluding floating-point (FP) operations is crucial when achieving lightweighting.

[0027] Accordingly, in an embodiment of the present invention, a lightweight deep learning model capable of inference using only integer operations without real number operations and an accelerator capable of accelerating the model with low power are proposed.

[0028] To this end, in an embodiment of the present invention, a convolution operation utilizing a combination of quantization steps is used instead of a general convolution operation in order to reduce the amount of computation within the accelerator while increasing offline computation.

[0029] The 2-bit quantized activations and weights can each have up to 4 quantization steps. Then, the number of steps combining the quantization steps of the activations and weights is 16 (4*4), and this value can be calculated offline in advance. In other words, the convolution operation module in the hardware does not input the quantization step value and calculate it directly, but only receives the quantization step index as input and determines which step it is using the concat operation. After the determination is completed for all input activations and weights, the step values ​​calculated offline in advance are multiplied at once to obtain the final output.

[0030] Below, a specific convolution operation method using an index combination operation in accelerator hardware is described in detail with reference to Fig. 2. Fig. 2 is a drawing for explaining a convolution operation method using an index combination operation.

[0031] Assuming 2-bit quantization, as shown in the upper part of Fig. 2, the quantized activation [a, b, c, d] and the quantized weight [A, B, C, D] can each have 4 (= 2^2) values. In accelerator hardware, these four values ​​are not directly handled, but rather indices are assigned to them and used.

[0032] The number of cases that can be created by combining these indices is 16. That is, after combining the indices, there are a total of 16 values, and the values ​​are expressed as Step = [aA, aB, aC, ..., dC, dD] in the upper part of Fig. 2. And the number of cases where each Step exists in the convolution operation is N0, N1, N2, N3, ..., N indicated in the center of Fig. 2. 15 am.

[0033] In this case, the result of the convolution operation is as shown in the lower part of Fig. 2, and if expressed as a formula, it is as shown at the bottom (Out) of Fig. 2. As can be seen from this formula, for the convolution operation, multifactors (MultiFactor0, MultiFactor1, ..., MultiFactor 15 ) and AddFactor are required.

[0034] Since the activation and weight values ​​used to create multifactors are fixed values ​​during the learning process, the corresponding multifactors can be calculated offline in advance, stored in hardware, and then input into the calculator.

[0035] Figure 3 is a diagram illustrating a method for calculating multifactors for performing a convolution operation simultaneously with a batch normalization (BN: BatchNorm) operation. In Figure 3, ConvSum is a convolution operation formula, and BN is a batch normalization formula that utilizes the result of the convolution operation.

[0036] Step0, Step1, Step2, ... Step written on the right side of the convolution operation formula (ConvSum) 15 are values ​​that combine the quantized activation values ​​and quantized weight values ​​described above. And in the batch normalization formula (BN), G bn , M bn , B bn , V bn 1 / 2 are parameters for batch normalization operation.

[0037] The result of substituting the convolution operation formula (ConvSum) into the batch normalization formula (BN) is shown in Figure 4. This is N0, N1, N2, N3, ..., N 15 If organized as an input variable, it can be expressed as in Fig. 5, and multifactors and AddFactor can be defined by this.

[0038] In multifactors, the step values ​​of the convolution operation are applied in different combinations, but the parameters for the batch normalization operation are applied with the same values ​​to all multifactors, unlike the step values ​​of the convolution operation.

[0039] Multifactors and AddFactor are pre-calculated offline, converted to integer values ​​through uniform quantization, stored, and then used as input to the accelerator hardware.

[0040] The deep learning operation replaced with the index combination operation presented in the embodiment of the present invention can reduce MAC operations by 98.6% compared to existing operations based on 3x3x128 kernels and 2-bit quantization.

[0041] Furthermore, since convolution and batch normalization operations completely exclude real-valued operations, power and area efficiency can be maximized. Since hardware inputs are limited to integers, only integer operators are used, resulting in reduced power and area consumption.

[0042] FIG. 6 is a diagram illustrating the configuration of deep learning accelerator hardware according to another embodiment of the present invention. As illustrated, the deep learning accelerator hardware according to the embodiment of the present invention is configured to include a communication interface (110), memory (120), and a deep learning operator (130).

[0043] The communication interface (110) is configured for data communication with the host, and the memory (120) stores data received through the communication interface (110) and annual data of the deep learning operator (130), and the aforementioned multifactors and AddFactor are stored.

[0044] The deep learning operator (130) is configured to perform deep learning operations, and when performing convolution operations, it utilizes multifactors and AddFactor that are pre-calculated / quantized and stored in memory (120).

[0045] So far, we have described in detail a preferred embodiment of an ultra-lightweight network model quantization method and a low-power deep learning computing device using the method.

[0046] As shown in Figure 1, integer and real number operations differ in energy and area by tens to hundreds of times. In an embodiment of the present invention, in order to use only integer operations on hardware for low-power hardware, multifactors and AddFactor are pre-calculated offline, quantized, expressed as integers, stored, and then input to the calculator.

[0047] The above embodiment presents an accelerator hardware that operates at high speed and low power on edge devices, utilizing hardware-optimized, end-to-end low-bit (1 / 2-bit) quantization technology. This hardware-hardware co-design enables the creation of an end-to-end low-bit quantized model, from the input image to the final model output, and accelerates the model at low power and high speed.

[0048] Meanwhile, when padding activation with 0, a problem may arise where the padding error 0 is recognized as index 0. To eliminate calculation errors, quantization is possible, which unconditionally sets the activation quantization value to 0. With this quantization method, activation has up to four quantization steps, including 0.

[0049] Furthermore, it is possible to apply the method of receiving the quantization step index as an input of the operation to residual operation, average pool operation, and max pool operation in addition to convolution operation.

[0050] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.

Claims

1. A step of calculating and storing multifactors to be used for performing quantized convolution operations and quantized batch normalization operations at once; A quantized deep learning operation method, characterized by including an operation step of performing a convolution operation and a batch normalization operation at the same time using stored multifactors.

2. In claim 1, Multifactors are, A quantized deep learning operation method characterized in that it is calculated using values ​​that are combinations of quantized activation values ​​and quantized weight values.

3. In claim 2, Multifactors are, A quantized deep learning operation method characterized in that it is calculated by further utilizing parameters for batch normalization operation.

4. In claim 3, The parameters for the batch normalization operation are: A quantized deep learning computational method characterized by being equally applicable to all multifactors.

5. In claim 3, The operation steps are: A quantized deep learning operation method, characterized by including a first summing step of multiplying the corresponding numbers by each multifactor and summing them.

6. In claim 5, A quantized deep learning operation method, characterized by further including a second summing step of summing AddFactor to the summing result of the first summing step.

7. In claim 6, The saving step is, A quantized deep learning operation method characterized by quantizing and storing calculated multifactors and AddFactor.

8. In claim 7, Multifactors and AddFactor, A quantized deep learning operation method characterized by having integer values.

9. In claim 1, The storage phase and the operation phase are, A quantized deep learning operation method characterized by being performed on an accelerator mounted on an edge device.

10. Memory storing multifactors to be used for performing quantized convolution operations and quantized batch normalization operations at once; A deep learning operation accelerator, characterized by including an operation unit that performs a convolution operation and a batch normalization operation at the same time using multifactors stored in memory.

11. A step of receiving multifactors to be used for performing quantized convolution operations and quantized batch normalization operations at once; A quantized deep learning operation method, characterized by including an operation step of performing a convolution operation and a batch normalization operation at the same time using input multifactors.

12. A step of calculating multifactors to be used to perform quantized convolution operations and quantized batch normalization operations at once; A step of storing the calculated multifactors in memory; The saved multifactors are, A method for storing deep learning operation parameters, characterized in that it is used to perform convolution operation and batch normalization operation at the same time.

Citation Information

Patent Citations

  • Auto encoder device, data processing system, data processing method and program

    JP2019140680A

  • Thermo hygrostat and control method of the same

    KR1020200127508A

  • System and method for safety inspection by nature freqeuncy of building structure

    KR1020240170730A

  • Device performing AI inference through quantization and batch folding

    KR102505043B1

  • KR20210127099A

Cited By

  • Feature data processing method based on edge device and related device

    CN120763718A