Method for lightening deep learning model through neural processing unit-aware pruning
By custom pruning deep learning models to align with NPU hardware specifications, the method addresses inefficiencies in existing lightweight models, achieving reduced computation and memory usage while maintaining accuracy.
Patent Information
- Application Number
- PCT/KR2023/017212
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-01
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-08
AI Technical Summary
Existing lightweight deep learning models not optimized for NPU hardware specifications lead to inefficient space utilization and unnecessary operations, particularly in resource-constrained edge devices.
A method of custom pruning deep learning models using the NPU, specifically binding output feature maps with channel units and applying filter pruning techniques based on hardware design specifications to reduce computation and memory usage.
This approach results in a more efficient use of NPU hardware, reducing unnecessary operations and improving space efficiency, while maintaining model accuracy by increasing the proportion of significant values.
Smart Images

Figure KR2023017212_08052025_PF_FP_ABST
Abstract
Description
A method for reducing the weight of deep learning models through customized pruning of neural processing units.
[0001] The present invention relates to a technology for reducing the weight of a deep learning model, and more particularly, to a method for reducing the weight of a deep learning model by utilizing a pruning technique that takes into account the hardware design specifications of an NPU (Neural Processing Unit), which is a chipset dedicated to deep learning calculations.
[0002] NPUs are being used to implement deep learning model training and inference on edge devices such as mobile devices, IoT devices, and industrial robots. However, to enable high-speed inference while using less power and memory in constrained environments, the parameters and computational load of deep learning models must be reduced.
[0003] Deep learning models can be made lightweight through various methods, including pruning, quantization, and knowledge distillation. However, lightweight models suffer from differences in hardware computational requirements depending on NPU hardware specifications. This is because lightweight models that are not considered in hardware specifications result in inefficient space and unnecessary computation.
[0004] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide a method for reducing the weight of a hardware-centered deep learning model for hardware optimization through NPU-customized pruning.
[0005] A deep learning model pruning method according to one embodiment of the present invention for achieving the above purpose includes the steps of binding an output feature map of a deep learning model to a plurality of channel units; and, if the number of feature map data of a channel unit is less than or equal to a reference value, removing the corresponding channel unit.
[0006] The removal step may be intended to reduce the amount of computation required to process the output feature map.
[0007] The deep learning model pruning method according to the present invention may further include a step of maintaining the channel unit when the number of feature map data of the channel unit exceeds a reference value.
[0008] The deep learning model pruning method according to the present invention may further include a step of filling data into an empty channel when the number of feature map data of a channel unit exceeds a reference value but is smaller than the number of channel units.
[0009] The data that fills the empty channels can be dummy data or other feature map data.
[0010] The threshold may be a specific percentage of channel units.
[0011] The number of channel units is
[0012] It is determined by the following formula:
[0013] F = α×N'
[0014] N' = N + 1 (β > α / n)
[0015] N' = N (β ≤ α / n)
[0016] F is the output feature map
[0017] α is the channel unit,
[0018] N' is the number of channel units,
[0019] N is the number of channel units filled with feature map data,
[0020] β is the number of remaining feature map data that did not fill the channel unit,
[0021] n can be an integer greater than or equal to 2.
[0022] The channel unit size can be determined by hardware specifications.
[0023] The hardware could be an NPU (Neural Process Unit) mounted on an edge device.
[0024] According to another aspect of the present invention, a system for reducing the weight of a deep learning model is provided, comprising: a processor for binding an output feature map of a deep learning model to a plurality of channel units; and a storage unit for removing a channel unit if the number of feature map data of a channel unit is less than a threshold value; and a storage unit for providing storage space required by the processor.
[0025] According to another aspect of the present invention, a method for pruning a deep learning model is provided, characterized by including: a step of binding an output feature map of a deep learning model to a plurality of channel units; and a step of maintaining the channel unit when the number of feature map data of the channel unit exceeds a reference value.
[0026] According to another aspect of the present invention, a system for reducing the weight of a deep learning model is provided, characterized in that it includes a processor for binding an output feature map of a deep learning model to a plurality of channel units and maintaining the channel unit when the number of feature map data of the channel unit exceeds a reference value; and a storage unit for providing storage space required by the processor.
[0027] As described above, according to embodiments of the present invention, hardware optimization through hardware-centric deep learning model weight reduction is possible through lightweighting of a deep learning model through NPU-customized pruning.
[0028] In particular, according to embodiments of the present invention, it is possible to use a lightweight deep learning model optimized for hardware by reducing unnecessary hardware operations and increasing the proportion of meaningful values, and inefficient space utilization of hardware can be reduced, making it possible to utilize a deep learning model even in systems with limited memory.
[0029] Figure 1. Example of output feature map according to pruning technique.
[0030] Figure 2. Example of channel unit formation in a 20-channel feature map (channel unit unit: 4 channels)
[0031] Figure 3. Example of channel unit formation in a 20-channel feature map (channel unit unit: 8 channels)
[0032] Figure 4. Example of channel unit formation in a 20-channel feature map (channel unit unit: 16 channels)
[0033] Figure 5. Comparison of general pruning and NPU-customized pruning techniques according to one embodiment of the present invention.
[0034] Figure 6. Configuration of a lightweight system for a deep learning model according to another embodiment of the present invention.
[0035] Hereinafter, the present invention will be described in more detail with reference to the drawings.
[0036] Pruning, one of the deep learning model weight reduction techniques, is a method of removing unnecessary parameters that do not significantly affect the model's accuracy. There are several types, such as weight, channel, and filter pruning.
[0037] Weight pruning is a method that repeatedly performs the process of layer-wise pruning and retraining weights that have values lower than a threshold value. Channel pruning is a method of pruning relatively unnecessary channels in a neural network model, and filter pruning is a method of reducing the amount of computation by pruning filters with low importance in a neural network model.
[0038] When writing the output feature map of a lightweight model to DRAM (Dynamic Random Access Memory), it is written as shown in Figure 1, depending on the pruning technique. Since the weight and channel pruning techniques prune a portion of the filter, the filter size remains the same, and only the pruned area has a value of 0. However, in the case of filter pruning, the entire filter is pruned, reducing the size of the output feature map.
[0039] Accordingly, in an embodiment of the present invention, a method for reducing the weight of a deep learning model is proposed through NPU-customized pruning, taking into account the hardware design specifications of the NPU (Neural Processing Unit), a chipset dedicated to deep learning calculations, based on a filter pruning technique suitable for reducing hardware computational load.
[0040] When inferring a neural network model, the output feature map of each layer is sequentially written to DRAM. As illustrated in FIGS. 2 to 4, the output feature map is bound by channel unit according to the hardware design specifications, and the bound data is referred to as a channel unit. The NPU reads the starting addresses of each channel unit from DRAM and performs the next computational process.
[0041] The output feature map is divided into multiple channel units, and the number of channel units is different because the channel unit units are different depending on the hardware design. In Fig. 2, the channel unit unit is 4 channels and the number of channel units is 5, in Fig. 3, the channel unit unit is 8 channels and the number of channel units is 3, and in Fig. 2, the channel unit unit is 16 channels and the number of channel units is 2.
[0042] In this way, the distribution of feature maps varies depending on the channel unit, and unnecessary values are included in the remaining space of the channel unit. This can be expressed as a formula, as shown in Equation 1.
[0043] [Formula 1]
[0044] F noramal = α×N + β
[0045] In Equation 1, F represents the output feature map, α represents the channel unit unit, N represents the number of channel units filled with feature map data, and β represents the number of remaining feature map data that do not fill the channel units. In Fig. 2, α=4, N=5, β=0, in Fig. 3, α=8, N=2, β=4, and in Fig. 4, α=16, N=1, β=12.
[0046] Equation 1 means that the number of channel units to be used is N + 1, and the last channel unit to which the feature map data β belongs must include unnecessary values such as dummy values.
[0047] Meanwhile, since NPUs perform operations on a per-channel basis, dummy values are not excluded from the calculations but rather ultimately increase the amount of unnecessary computation. Therefore, in an embodiment of the present invention, an NPU-specific pruning technique utilizing a filter pruning technique tailored to each channel unit, as per Equations 2 and 3, is utilized to minimize the amount of unnecessary computation.
[0048] [Formula 2]
[0049] F npu - aware = α×N'
[0050] [Formula 3]
[0051] N' = N + 1 (β > α / 2)
[0052] N' = N (β ≤ α / 2)
[0053] In Equation 2, N' is divided into two cases according to the size of β as shown in Equation 3, and accordingly, the number of channel units N' is determined according to the size of β.
[0054] Specifically, as in (a) of Fig. 5, when β is greater than half of the channel unit, the number of channel units is maintained the same as in the existing method. If the number of feature map data of a channel unit is greater than half of the channel unit, the corresponding channel unit is maintained, and if the number of feature map data of a channel unit is less than the channel unit, dummy data is filled in the empty channel.
[0055] This is because removing the remaining feature maps may actually result in greater information loss due to the small number of dummy values. However, it is also possible to improve accuracy with the same amount of computation as existing methods by filling the remaining space with other feature maps that have meaningful values instead of dummy values.
[0056] On the other hand, if β is less than half of the channel unit unit, as in (b) of Fig. 5, the channel unit containing β is removed. If the number of feature map data of the channel unit is less than half of the channel unit unit, the corresponding channel unit is removed.
[0057] This is to reduce the amount of unnecessary computation by removing dummy values, which not only reduces space utilization because they occupy most of the channel unit, but also causes a large proportion of unnecessary computation.
[0058] Meanwhile, the standard of β, which is the pruning condition of Equation 3, can be changed to α / 2, α / 3, α / 4, etc. depending on the channel unit and pruning ratio, thereby reducing the loss of accuracy through parameter adjustment.
[0059] Figure 6 is a diagram illustrating the configuration of a deep learning model lightweight system according to an embodiment of the present invention. The deep learning model lightweight system according to an embodiment of the present invention can be implemented as a computing system comprising a communication unit (110), an output unit (120), a processor (130), an input unit (140), and a storage unit (150).
[0060] The communication unit (110) is a communication interface for connection with an external network or external device. The output unit (120) is an output means for displaying the results of computation performed by the processor (130), and the input unit (140) is a user interface for receiving user commands and transmitting them to the processor (130).
[0061] The processor (130) reduces the weight of the deep learning model according to the NPU-customized pruning technique illustrated in FIG. 5 described above. The storage unit (150) provides the storage space necessary for the processor (130) to function and operate.
[0062] So far, we have described in detail preferred embodiments of methods and systems for reducing the weight of deep learning models through NPU-customized pruning.
[0063] In an embodiment of the present invention, a method for reducing the weight of a deep learning model utilizing an NPU-specific filter pruning technique for efficient inference of a neural network deep learning model on an NPU is presented. Filter pruning is performed on a channel-unit basis, tailored to the NPU's hardware design specifications.
[0064] This leads to high space efficiency and reduced computational load of NPU hardware used in edge devices, and minimizes accuracy loss by increasing the proportion of meaningful values in the same computational load.
[0065] Meanwhile, it goes without saying that the technical idea of the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.
[0066] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. A step of binding the output feature map of the deep learning model to multiple channel units; A deep learning model pruning method, characterized in that it includes a step of removing a channel unit when the number of feature map data of a channel unit is less than a reference value.
2. In claim 1, The removal step is, A deep learning model pruning method characterized by reducing the amount of computation required for processing output feature maps.
3. In claim 1. A deep learning model pruning method, characterized in that it further includes a step of maintaining the channel unit when the number of feature map data of the channel unit exceeds a reference value.
4. In claim 3, A deep learning model pruning method, characterized in that it further includes a step of filling data in an empty channel when the number of feature map data of a channel unit exceeds a reference value but is smaller than the number of channel units.
5. In claim 4, Data to fill empty channels, A method for pruning a deep learning model characterized by having dummy data or other feature map data.
6. In claim 4, The standard is, A deep learning model pruning method characterized by a specific ratio of channel units.
7. In claim 6, The number of channel units is It is determined by the following formula: F = α×N' N' = N + 1 (β > α / n) N' = N (β ≤ α / n) F is the output feature map α is the channel unit, N' is the number of channel units, N is the number of channel units filled with feature map data, β is the number of remaining feature map data that did not fill the channel unit, A deep learning model pruning method characterized in that n is an integer greater than or equal to 2.
8. In claim 7, The channel unit is, A deep learning model pruning method characterized by being determined by hardware specifications.
9. In claim 8, The hardware is, A deep learning model pruning method characterized by an NPU (Neural Processing Unit) mounted on an edge device.
10. A processor that binds the output feature map of a deep learning model to multiple channel units and, if the number of feature map data of a channel unit is below a threshold, removes the corresponding channel unit; and A system for reducing the weight of a deep learning model, characterized by including a storage unit that provides storage space required by a processor.
11. A step of binding the output feature map of the deep learning model to multiple channel units; A deep learning model pruning method, characterized in that it includes a step of maintaining the channel unit when the number of feature map data of the channel unit exceeds a reference value.
12. A processor that binds the output feature map of a deep learning model to multiple channel units and maintains the channel unit when the number of feature map data of the channel unit exceeds a threshold; and A system for reducing the weight of a deep learning model, characterized by including a storage unit that provides storage space required by a processor.
Citation Information
Patent Citations
Presure sensor
KR1020210019713A
Method for applying using inkjet apparatus
KR1020230022111A
Method and system for channel pruning of compact neural networks
KR102165273B1