Pruning Hardware Unit for Training Neural Networks
By using configurable pruning hardware units pruning weights during the training process of neural networks, the problems of incremental pruning resource consumption and time waste are solved in traditional software, and a more efficient training process is achieved.
Patent Information
- Application Number
- CN202080100459.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-06-11
AI Technical Summary
Incremental pruning implemented by traditional software requires computing, storing and accessing all weights and their gradients when training neural networks, resulting in resource consumption and time waste.
A configurable pruning hardware unit is designed to prune weights during training of neural networks. The hardware unit receives weights, gradients, and pruning indicators provided by the training engine, selects unpruned weights for pruning, and updates the pruning indicator to reduce the amount of computation and storage resources.
By pruning weights during training, the computational amount, storage resources and time required to train the neural network can be reduced without affecting the accuracy of the neural network.
Smart Images

Figure CN115552413B_ABST
Abstract
Description
Background Art
[0001] Neural networks are used to perform artificial intelligence (AI) tasks. Neural networks are trained using a training data set, and during training, they perform artificial intelligence tasks by mapping inputs to outputs. The training process involves finding and assigning weights to neurons or nodes in the neural network that are used to accurately complete the artificial intelligence tasks. The training process is iterative, and the weights are updated based on the results of the previous iteration (e.g., each epoch) during each iteration.
[0002] The backpropagation (BP) algorithm is widely used for training neural networks. Generally, backpropagation determines and fine-tunes the weights of the neurons in the neural network based on the error magnitude calculated in the previous epoch (e.g., iteration). Backpropagation is well known in the art.
[0003] Training a neural network requires a large amount of resources and time. By carefully pruning data (e.g., weights) during the training process, it is possible to reduce the computational amount, storage resources, and time required for training the neural network without affecting the accuracy of the neural network. Incremental pruning achieves a good balance between accuracy and computational overhead. However, traditional software-implemented incremental pruning does not take advantage of incremental pruning because it requires calculating, storing, and accessing all weights and their gradients, calculating the pruning criteria for all weights of the neural network, and the sorting rules for all weights of the neural network. Summary of the Invention
[0004] Disclosed herein is a system for pruning weights during the training of a neural network. The system includes a configurable pruning hardware unit configured to: receive an input including weights, gradients associated with the weights, and a pruning indicator for each weight from a training engine of the neural network; select unpruned weights for pruning; prune the selected unpruned weights for pruning; update the pruning indicator for each of the selected pruned weights; and provide the updated pruning indicator to the training engine for the next iteration or epoch. The system can be used for incremental pruning and non-incremental pruning.
[0005] In some embodiments, the pruning hardware unit includes a weight criterion calculation module that receives an input from the training engine of the neural network. The input includes the values of the weights and the values of the gradients of the nodes of the neural network. The input also includes the values of an indicator for each weight, where the value of the indicator indicates whether the associated weight is a pruned weight or an unpruned weight. The weight criterion calculation module only outputs the weight criteria for unpruned weights. The pruning hardware unit further includes a top-k module that calculates a pruning threshold based on the output of the weight criterion calculation module. The pruning hardware unit further includes a pruning module that updates the value of each weight indicator and provides the updated value to the training engine.
[0006] In some embodiments, the pruning hardware unit includes a plurality of registers that store values for configuring the pruning hardware unit. The registers are configured by the training engine and written by an application programming interface that provides a software interface between the training engine and the pruning hardware unit. The registers include a software enable register that, together with a hardware enable signal from the training engine, enables the pruning hardware unit. A criterion register defines the criterion for pruning unpruned weights. An input selection register identifies the weights and gradients obtained from the training engine. A mode register defines the pruning mode. In some embodiments, there are an incremental pruning mode and a non-incremental pruning mode. The value in the N register refers to the number N of unpruned weights. The value in the k register refers to the number k of weights that are not pruned in the current iteration or generation.
[0007] The pruning hardware unit according to an embodiment of the present invention avoids the above-mentioned drawbacks of the conventional software-implemented incremental pruning. According to an embodiment of the present invention, by carefully pruning data (e.g., weights) during the training process, it is possible to reduce the computational amount, storage resources, and time required for training a neural network without affecting the accuracy of the neural network.
[0008] Those of ordinary skill in the art will recognize the above and other objects and advantages of the embodiments of the present invention after reading the following detailed description of the embodiments shown in the various figures. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the detailed description, serve to explain the principles of the present disclosure, in which like numerals depict corresponding elements.
[0010] Figure 1 is a block diagram of an exemplary system that can implement the embodiments described herein.
[0011] Figure 2 is a block diagram of the pruning hardware unit according to an embodiment of the present invention.
[0012] Figure 3 is a block diagram of the weight criterion calculation module according to an embodiment of the present invention.
[0013] Figure 4 is a block diagram of the pruning module according to an embodiment of the present invention.
[0014] Figure 5 is a flowchart of a method for training a neural network according to an embodiment of the present invention. DETAILED DESCRIPTION
[0015] Various embodiments of the present disclosure will now be described in detail, examples of which are shown in the accompanying drawings. Although the present disclosure is described in connection with these embodiments, it should be understood that they are not intended to limit the present disclosure to these embodiments. On the contrary, the present disclosure is intended to cover alternatives, modifications, and equivalents included within the spirit and scope of the present disclosure as defined by the appended claims. In addition, in the following detailed description of the present disclosure, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it should be understood that the present disclosure may be practiced without these specific details. In some cases, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure various aspects of the present disclosure.
[0016] Certain portions of the detailed description that follow are presented in terms of procedures, logic blocks, processing, and other symbolic representations of operations on data bits within a computer memory. These descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. In the present application, a procedure, logic block, process, etc. is conceived of as a self-consistent sequence of steps or instructions leading to a desired result. These steps operate on physical quantities. Usually, though not necessarily, these physical quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as transactions, bits, values, elements, symbols, characters, samples, pixels, etc.
[0017] However, it should be borne in mind that all such and similar terms are to be associated with the appropriate physical quantities and that these terms are merely convenient labels applied to these physical quantities. Unless explicitly stated otherwise from the following discussion, it should be understood that throughout the present disclosure, descriptions using terms such as "receive", "access", "determine", "store", "select", "indicate", "prune", "update", "set", "calculate", "multiply", "provide", etc., refer to the actions and processes of a device or computer system, or similar electronic computing device or system (such as Figure 1 the system shown). A computer system or similar electronic computing device manipulates and transforms data represented as physical (electronic) quantities within a memory, register, or other such information storage device, transmission, or display device.
[0018] Some of the elements or embodiments described herein are discussed in the general context of computer-executable instructions residing on some form of computer-readable storage medium, such as program modules, that are executed by one or more computers or other devices. By way of example, and not limitation, computer-readable storage media may include non-transitory computer storage media and communication media. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. In various embodiments, the functionality of program modules may be combined or distributed as desired.
[0019] Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory (such as SSD), or other storage technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed to retrieve that information.
[0020] Communication media may contain computer-executable instructions, data structures, and program modules, and includes any information delivery media. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media. Any combination of the above is also included within the scope of computer-readable media.
[0021] The following discussion includes terms such as "weight", "gradient", "pruning indicator", etc. Unless otherwise stated, each term has a value. For example, a weight has a value, and different weights can have different values. For simplicity, the term "weight" refers to the value of the weight, unless otherwise stated or apparent from the discussion.
[0022] Figure 1 A block diagram of an exemplary system 100 is shown on which the various embodiments described herein may be implemented. In its most basic configuration, system 100 includes a computing system 170 coupled to a pruning hardware unit 190.
[0023] Computing system 170 includes at least one processing unit 102 and at least one memory 104. Each processing unit 102 may be a general-purpose processor or a special-purpose processor (such as a neural processing unit).
[0024] The computing system 170 may also have additional features and / or functionality. For example, the system 100 may also include additional storage devices (removable and / or non-removable). Such additional storage devices are shown in Figure 1 as removable storage device 108 and non-removable storage device 120. The computing system 170 may also include one or more communication connection devices 122 that allow the computing system to communicate with other devices (such as, but not limited to, pruning hardware unit 190). The computing system 170 also includes input devices 124, such as, for example, keyboards, mice, pens, voice input devices, touch input devices, etc. The computing system 170 also includes output devices 126, such as, for example, display devices, speakers, printers, etc.
[0025] In Figure 1 the example, the memory 104 includes computer-readable instructions associated with the training engine 150, data structures (which include databases), program modules, etc. The training engine 150 trains the neural network over multiple training iterations or epochs. In each iteration or epoch, the training engine 150 trains the neural network on respective mini-batches of training data, each mini-batch of training data consisting of, for example, dozens to hundreds of training data samples. Typically, the training engine 150 determines the weights and gradients of the objective (loss) function that characterizes the accuracy of the neural network and then uses this information to adjust the weights to improve the accuracy of the neural network. The training engine 150 may use, for example, the backpropagation algorithm to determine and fine-tune the weights of the neurons (or nodes) of the neural network based on the magnitude or absolute value of the error calculated in a prior epoch or iteration.
[0026] In an embodiment of the present invention, the computing system 170 outputs weights 172 and gradients 174 from the training engine 150 to the pruning hardware unit 190. In some embodiments, a "pruning indicator 176 is associated with each weight. In some embodiments, the pruning indicator 176 is a bit value that indicates whether the associated weight is unpruned or has been previously pruned. If the weight is unpruned, the pruning indicator 176 has one value (for example, the bit is set to value 1), and if the weight has been previously pruned, the pruning indicator 176 has a different value (for example, the bit is cleared and has value 0). As described below, a pruned weight also has a zero value, i.e., when a weight is pruned, its value is set to zero. As further described below, the value of the pruning indicator 176 is determined by the pruning hardware unit 190.
[0027] In some embodiments, the computing system 170 (specifically, the training engine 150) communicates with the pruning hardware unit 190 through an application programming interface (API) 180. The application programming interface 180 can be implemented on the computing system 170 or on the pruning hardware unit 190. In some embodiments, the application programming interface 180 is a software interface; however, the present invention is not limited thereto. For example, the functions provided by the application programming interface 180 can be provided by a hardware controller.
[0028] Figure 2 is a block diagram of the pruning hardware unit 190 in an embodiment of the present invention. Generally, the pruning hardware unit 190 includes a weight criterion calculation module 210, a TOP-k module 220, and a pruning module 230 controlled by a controller 250. The pruning hardware unit 190 is configured to receive an input including weights 172, gradients 174, and pruning indicators 176 for each weight from the training engine 150; select unpruned weights for pruning; prune the selected unpruned weights for pruning; update the pruning indicators for each of the selected pruned weights; and provide the updated pruning indicators for each weight to the training engine 150.
[0029] The weight criterion calculation module 210 receives the above input from the training engine 150 and outputs only the weight criterion 212 of the unpruned weights. That is, the weight criterion calculation module 210 does not operate on or use the pruned weights. As described above, the value of the pruning indicator 176 indicates whether a weight is a pruned weight or an unpruned weight. In combination with Figure 3 the weight criterion calculation module 210 is further described.
[0030] Continuing to refer to Figure 2 , the weight criterion calculation module 210 supports the calculation of different weight criteria 212. For example, the weight criterion 212 can be the criterion used by the training algorithm of the training engine 150 to minimize the objective (loss) function. For example, the weight criterion 212 can be the magnitude (absolute value) of the unpruned weights. Generally, the weight criterion 212 is used to indicate the relative importance of the weights output by the weight criterion calculation module 210.
[0031] The top-k module 220 calculates the value of the pruning threshold based on the weight criterion 212 received from the weight criterion calculation module 210. As will be further described below, the pruning threshold is used to select unpruned weights for pruning.
[0032] The pruning module 230 updates the value of each pruning indicator 176 based on the result from the top-k module 220. The values of the unpruned weights selected for pruning are set to zero, and the values of the pruned weights remain zero during the remaining training process. Thus, for the remaining training process, the values of the pruned weights are not updated, their gradients are not calculated, and the pruned weights are no longer used to perform multiplications, thereby reducing the computational amount, storage resources, and time required for training the neural network.
[0033] The pruning module 230 outputs the updated pruning indicator 232, which is used as an input to the training engine 150 and the weight criterion calculation module 210 in the next iteration or epoch. The pruning module 230 will be further described in conjunction with Figure 4 below.
[0034] Figure 2 The controller 250 updates the values used by the training engine 150 based on the output of the pruning module 230. In one embodiment, the controller 250 is a finite state machine (FSM) controller.
[0035] In some embodiments, the pruning hardware unit 190 includes a plurality of registers 240. In one embodiment, the registers include a criterion register 241, an input selection register 242, a mode register 243, an N register 244, a k register 245, and a software enable (SW_en) register 246. The registers 240 are configured by the training engine 150 and written by the Figure 1 application programming interface 180.
[0036] When both the hardware enable (HW_en) signal 188 from the training engine 150 and the value in the software enable register 246 are valid, Figure 2 the pruning hardware unit 190 is enabled.
[0037] The value (instruction) in the criterion register 241 defines the criterion for pruning the unpruned weights. The value in the input selection register 242 identifies the criterion (e.g., weights and gradients) to be obtained from the training engine 150.
[0038] The value in the mode register 243 defines the pruning mode. In some embodiments, there are an incremental pruning mode and a non-incremental pruning mode. In the incremental pruning mode, the number (or percentage) of pruned weights increases as the training process progresses. For example, in an earlier epoch or iteration, 10% of the weights are pruned; in a later epoch or iteration, 20% of the remaining weights are pruned. In the non-incremental mode, the number of weights pruned in each epoch or iteration is constant (e.g., always 10%).
[0039] The value in the N register 244 is the number N of unpruned weights. Initially, the value of N is the total number of weights. As weights are pruned, the number of unpruned weights decreases, and the value of N is updated accordingly.
[0040] The value in the k register 245 is the number of weights that will not be pruned in the current iteration or epoch. That is, the value in the k register 245 establishes the pruning threshold used by the top-k module 220. For example, when the weights output from the weight criterion calculation module 210 are sorted according to the weight criterion 212 (e.g., in descending order), the top k weights in the sorting are not pruned, while the weights outside the top k weights in the sorting are pruned.
[0041] In operation, if the value of N in the N register 244 is zero (which means the application programming interface 180 is called for the first time), the value of N is reset to the total number of weights 172, and the value of k in the k register 245 is set to the product of N and 1 minus the sparsity value (1 – <sparsity value>). The target sparsity value defines the portion of the total weights to be pruned. In one embodiment, the sparsity value is a user-specified value (e.g., a user-defined input to the training engine 150 or the application programming interface 180).
[0042] The application programming interface 180 (or the hardware controller) determines the (e.g., incremental or non-incremental) pruning mode and sets the mode register 243 accordingly. In the incremental mode, after each iteration or epoch, the value of N in the N register 244 is updated to the current value of k in the k register 245, and then the value of k in the k register 245 is updated to the product of the current value of N and 1 minus the sparsity value.
[0043] In the non-incremental mode, after each iteration or epoch, the value of N in the N register 244 remains unchanged, so the value of k in the k register 245 (the product of the current value of N and 1 minus the sparsity value) also remains unchanged.
[0044] At the start of each iteration or epoch, the application programming interface 180 (or the hardware controller) determines the inputs required to calculate the weight criterion 212 and sets the input selection register 242 accordingly. After each iteration or epoch, the application programming interface 180 or the hardware controller needs to set (update) the value in the register 240 based on the results of the just-completed iteration or epoch, and / or to establish the next iteration or epoch.
[0045] Figure 3 is a block diagram of the weight criterion calculation module 210 in an embodiment of the present invention. In Figure 3 In an embodiment, the weight criterion calculation module 210 includes two sub-modules: a criterion calculation engine 310 and a controller 320.
[0046] When (Figure 2 When the pruning hardware unit 190 operates in the incremental mode, the controller 320 provides an enable signal 326 to the criterion calculation engine 310. The value (valid or invalid) of the enable signal 326 depends on the value of the pruning indicator 176: when the pruning indicator 176 indicates that the weight 172 is not pruned, the enable signal is valid. When the pruning hardware unit 190 operates in the non-incremental mode, the enable signal 326 is always valid.
[0047] Continuing to refer to Figure 3 , according to the values of the weight read enable signal 322 and the gradient read enable signal 324, the weights 172 and the gradients 174 are selected and input into the criterion calculation engine 310. According to the ( Figure 2 ) values in the input selection register 242, the values (valid or invalid) of the weight read enable signal 322 and the gradient read enable information 324 are set.
[0048] Figure 3 The criterion calculation engine 310 outputs the weight criterion 212 to the ( Figure 2 ) top-k module 220.
[0049] Figure 4 is a block diagram of the pruning module 230 in an embodiment of the present invention. In Figure 4 the embodiment, the pruning module 230 includes a comparator 410, a data reading controller 420, and a pruning indicator updater 430.
[0050] As described above in conjunction with Figure 2 , the comparator 410 compares the weight criterion 212 with the pruning threshold 412 generated by the top-k module 220. The comparator 410 performs the comparison for one weight at a time under the control of the ( Figure 4 ) data reading controller 420. The comparator 410 does not operate on the weights that have been previously pruned. As described above, the value of the pruning indicator 176 indicates whether the weight is a pruned weight or an unpruned weight.
[0051] Comparator 410 and pruning indicator updater 430 operate in a pipelined manner. Data read controller 420 controls whether to read a weight according to the value of pruning indicator 176 associated with the weight. Data read controller 420 uses criterion access control signal 422 to synchronize pruning indicator 176 and weight criterion 212. If pruning indicator 176 indicates that the next weight to be processed (the weight associated with this pruning indicator) has been pruned, data read controller 420 will read the value of this pruning indicator, but will not read the value of the associated weight (the value of this weight is zero). If pruning indicator 176 indicates that the next weight to be processed (the weight associated with this pruning indicator) has not been pruned, data controller 420 will read the value of the pruning indicator, and after pruning indicator updater 430 finishes processing the weight currently being processed, will also read the value of the associated weight based on criterion access control signal 422.
[0052] The result of each comparison performed by comparator 410 is forwarded to pruning indicator updater 430, which also receives pruning indicator 176 of the weight associated with the comparison result. If the comparison result indicates that an unpruned weight should be pruned, pruning indicator updater 430 resets the value of pruning indicator 176 of this weight; for example, pruning indicator updater clears the bit of the pruning indicator. Pruning indicator updater 430 outputs the updated pruning indicator 232, which is used as an input to training engine 150 and ( Figure 2 the) weight criterion calculation module 210 in the next iteration or generation.
[0053] Figure 5 is the flowchart 500 of the method for training a neural network in an embodiment of the present invention. As further described in Figures 2 - 4 it, this method is implemented using Figure 1 system 100.
[0054] In Figure 5 block 502, training engine 150 calculates weights 172, gradients 174, and pruning indicators 176.
[0055] In block 504, it is determined whether to perform pruning for the current generation or iteration. That is to say, pruning is not necessarily performed during each iteration or generation. For example, the frequency of pruning can be based on user input or on the calculation results performed by training engine 150. If pruning is to be performed, flowchart 500 proceeds to block 506; otherwise, the flowchart returns to block 502.
[0056] In block 506, as described above in conjunction with Figures 2 - 4 pruning hardware unit 190 prunes unpruned weights.
[0057] In block 508, if the last epoch of the training process has been reached, flowchart 500 ends; otherwise, the flowchart returns to block 502.
[0058] Using the pruning hardware unit as described above, embodiments of the present invention avoid the disadvantages of incremental pruning implemented by traditional software. Therefore, embodiments of the present invention reduce the computational amount, storage resources, and time required for training a neural network without affecting the accuracy of the neural network by carefully pruning data (e.g., weights) during the training process.
[0059] The process parameters and step sequences described and / or illustrated herein are merely examples and can be changed as needed. For example, although the steps described and / or illustrated herein can be shown or discussed in a specific order, these steps do not necessarily need to be executed in the order shown or discussed. The various example methods described and / or illustrated herein can also omit one or more of the steps described or illustrated herein, or include additional steps in addition to the disclosed steps.
[0060] In addition, although various embodiments have been described above using specific block diagrams, flowcharts, and examples, each block diagram component, flowchart step, operation, and / or component described and / or illustrated herein can be implemented individually and / or jointly using a variety of configurations. Moreover, any disclosure of components contained in other components should be considered an example, as many other architectures can be implemented to achieve the same functionality.
[0061] Although the subject matter of the present disclosure has been described using specific language of structural features and / or method acts, it should be understood that the subject matter defined in the present disclosure is not necessarily limited to the above specific features or acts. Instead, the above specific features and acts are disclosed as example forms for implementing the present disclosure.
[0062] Thus, embodiments of the present disclosure have been described above. Although the present disclosure has been described in specific embodiments, it should be understood that the present disclosure should not be construed as being limited by these embodiments, but rather should be interpreted in accordance with the claims.
Claims
1. An apparatus, the apparatus includes a pruning hardware unit for training a neural network, the pruning hardware unit comprises: a controller; a plurality of registers coupled to the controller for configuring the pruning hardware unit; and a plurality of modules, controlled by the controller and configured to: receive an input including weights and gradients of nodes of the neural network, the input further including values of indicators for each weight, wherein the value of the indicator for each weight indicates whether the weight is a pruned weight or an unpruned weight, and the pruned weight of the weights has a zero value; select unpruned weights for pruning; prune the selected unpruned weights for pruning; update the value of the indicator for each weight according to the selected pruned weights; and provide the updated value of the indicator for each weight to a training engine of the neural network.
2. The apparatus according to claim 1, wherein, the plurality of registers includes a register for indicating one pruning mode among a plurality of pruning modes, the pruning mode includes an incremental pruning mode, and the number of pruned weights increases as training progresses.
3. The apparatus according to claim 1, wherein, the value of the selected weight for pruning is set to zero.
4. The apparatus according to claim 1, wherein, the plurality of registers includes a register for storing the value of the number of unpruned weights and a register for storing the value of the weight for selecting for pruning, wherein the value of the weight for selecting for pruning corresponds to a small fraction of the number of unpruned weights.
5. The apparatus according to claim 4, wherein, the controller is configured to calculate the value of the weight for selecting for pruning by multiplying the value of the number of unpruned weights by a value based on a target sparsity value.
6. The apparatus according to claim 1, wherein, the plurality of registers includes a register for storing the value for selecting the input.
7. The apparatus according to claim 1, wherein, the plurality of registers includes a register for storing the criterion for pruning the weights.
8. A system for training a neural network, the system comprises: a computing system for executing a training engine of the neural network, wherein the output generated by the training engine includes values of weights and gradients of nodes of the neural network; and a pruning hardware unit, coupled to the computing system via an application programming interface, wherein the pruning hardware unit includes: a plurality of registers, coupled to a controller, for configuring the pruning hardware unit; and a plurality of modules, configured to: receive the output of the training engine; receive the value of the indicator for each weight, wherein the value of the indicator for each weight indicates whether the weight is a pruned weight or an unpruned weight, and the pruned weight of the weights has a zero value; select the value of the unpruned weight from the values of the weights; prune the unpruned value; update the value of the indicator of the weights; provide the updated value of the indicator to the training engine.
9. The system according to claim 8, wherein, The plurality of registers includes a register indicating one pruning mode among a plurality of pruning modes, wherein the pruning mode includes an incremental pruning mode, and the number of pruned weights increases as training progresses.
10. The system according to claim 8, wherein, the value of the weight selected for pruning is set to zero.
11. The system according to claim 8, wherein, the plurality of registers includes a register for storing the value of the number of unpruned weights and a memory for storing the value of the weight selected for pruning, and the value of the weight selected for pruning corresponds to a small fraction of the number of unpruned weights.
12. The system according to claim 11, wherein, the application programming interface is operable to determine the value of the weight selected for pruning by multiplying the value of the number of unpruned weights by a value based on a target sparsity value.
13. The system according to claim 8, wherein, the plurality of registers includes a register for storing the value of the input selected from the output of the training engine.
14. The system according to claim 8, wherein, the plurality of registers includes a register for storing the criterion for pruning the weights.
15. The system according to claim 8, wherein, the application programming interface writes values to the plurality of registers.
16. An apparatus, the apparatus includes a pruning hardware unit for training a neural network, the pruning hardware unit comprises: a controller; a plurality of registers, coupled to the controller, for configuring the pruning hardware unit; a weight criterion calculation module for receiving inputs from a training engine of the neural network, the inputs including weight values and gradient values of nodes of the neural network, the inputs further including values of indicators for each weight, the values of the indicators for each weight indicating whether the weight is a pruned weight or an unpruned weight, and the weight criterion calculation module only outputs weight criteria for unpruned weights; a top-k module for calculating a value of a pruning threshold based on the output of the weight criterion calculation module; and a pruning module for updating the value of the indicator for each weight; wherein the controller updates the value used by the training engine of the neural network based on the output of the pruning module.
17. The apparatus according to claim 16, wherein, the plurality of registers includes a register for indicating one pruning mode among a plurality of pruning modes, the pruning mode includes an incremental pruning mode, and the number of pruned weights increases as training progresses.
18. The apparatus according to claim 16, wherein, the plurality of registers includes a register for storing the value of the number of unpruned weights and a register for storing the value of the weight selected for pruning, the value of the weight selected for pruning corresponds to a small fraction of the number of unpruned weights, and the pruning module compares the weight criterion for only unpruned weights with the pruning threshold to select the weights for pruning and updates the value of the indicator for each weight accordingly.
19. The apparatus according to claim 16, wherein, The plurality of registers includes registers that store values for selecting inputs to the weight criterion calculation module.
20. The apparatus according to claim 16, wherein, the plurality of registers includes registers that store criteria for pruning the weights.
Citation Information
Patent Citations
Method and device for predicting spectrum occupancy state based on neural network
CN103209417A
Deep neural network compression method based on improved clustering
CN108304928A