A Sparse Accelerator for On-Chip Training
By designing a sparse accelerator with multiple computing cores, using coarse and fine-grained matching modules to filter invalid operations, the problem that existing accelerators cannot eliminate all invalid operations is solved, and efficient on-chip training on terminal devices is realized.
Patent Information
- Application Number
- CN202210094538.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-26
AI Technical Summary
Existing accelerators cannot eliminate all invalid operations in on-chip training after acceleration, resulting in the inability to implement on-chip training on the terminal device to the most efficiently.
A sparse accelerator is designed, including multiple calculation cores, each calculation core includes an input numerical buffer module, a reference numerical buffer module, a mask buffer module, a coarse-grain matching module, a processing module and an accumulation module. By dynamically adjusting the data in these modules, the coarse-grained unit and the fine-grained unit filter the invalid operations respectively to eliminate all invalid operations.
It realizes efficiently and accurately eliminating all invalid operations in the three stages of on-chip training on terminal devices, and improves the hardware utilization rate of sparse accelerators.
Smart Images

Figure CN114492753B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer and electronic information technologies, and particularly relates to a sparse accelerator applied to on-chip training. Background Art
[0002] In recent years, Convolutional Neural Networks (CNNs) have performed excellently in the fields of computer vision, speech recognition, natural language processing, etc. In order to improve the recognition accuracy, it is necessary to perform on-chip training on the CNN model. On-chip training refers to the training of fine-tuning the CNN model using the user data on the terminal device, thereby improving the accuracy of the CNN model when it is used.
[0003] The process of on-chip training includes three stages, namely the forward propagation (FP) stage, the backward propagation (BP) stage, and the weight gradient calculation (WG) stage. In the FP stage, the activation value of the previous layer is used as the input activation value of the current layer, and after convolution operation with the corresponding convolution kernel weight of the current layer, the output activation value of the current layer is obtained through the activation function. In this way, the activation values of each layer in the CNN model are obtained layer by layer from front to back. Finally, the deviation between the predicted value and the true value of the label is calculated through the loss function, and the loss is obtained. In the BP stage, using the loss obtained in the FP stage, the error value of the next layer is used as the input error value of the current layer, and after convolution operation with the convolution kernel weight of the current layer, the error value of the current layer is obtained. In this way, the error values of each layer in the CNN model are obtained layer by layer from back to front. In the WG stage, according to the requirements of the chain rule, the output activation value of the previous layer and the error value of the current layer are subjected to convolution operation to obtain the weight gradient of the current layer. In this way, the weights of each layer in the CNN model are obtained layer by layer from front to back. After multiple iterative optimizations, when the model converges, the on-chip training of the entire CNN model is completed.
[0004] It can be seen that in the training process of each stage of on-chip training, the training is completed by convolving the input data and the reference data to obtain the output data. The convolution process includes obtaining the products of multiple input values and multiple reference values, and then adding the products to obtain multiple output values in the output data. However, the computational complexity of the CNN model is large, but the resources and storage capacity of the terminal device are limited, making it difficult to support the complete on-chip training on the terminal device. By transmitting the user data to the cloud server, although the complete on-chip training can be performed on the cloud server and the terminal device, the process of transmitting the user data may leak the user privacy. Therefore, it is necessary to accelerate the on-chip training to ensure the efficient implementation of on-chip training only on the terminal device.
[0005] Existing accelerators accelerate by eliminating invalid operations caused by zero input values or zero reference values. However, invalid operations also include cases where the output value is zero when both the input value and the reference value are negative. Therefore, after the existing accelerators are accelerated, since they cannot eliminate all invalid operations in on-chip training, they cannot ensure the most efficient implementation of on-chip training on terminal devices. Summary of the Invention
[0006] This application provides a sparse accelerator for on-chip training, which is used to solve the technical problem that after the existing accelerators are accelerated, since they cannot eliminate all invalid operations in on-chip training, they cannot ensure the most efficient implementation of on-chip training on terminal devices.
[0007] To solve the above technical problems, the embodiments of this application disclose the following technical solutions:
[0008] A sparse accelerator for on-chip training, the sparse accelerator includes a plurality of computing cores, each computing core is used to accelerate a convolution operation process each time, and the computing core includes an input value buffer module, a reference value buffer module, a mask buffer module, a coarse-grained matching module, a processing module, and an accumulation module, wherein:
[0009] The coarse-grained matching module includes a plurality of coarse-grained matching units, and the processing module includes a plurality of processing units arranged in an array, and each row of processing units shares one coarse-grained matching unit; wherein:
[0010] The input value buffer module is used to allocate the input values in the current acceleration stage to each row of processing units, and the current acceleration stage is the forward propagation stage, the backward propagation stage, or the weight gradient calculation stage;
[0011] The reference value buffer module is used to allocate the reference values in the current acceleration stage to each processing unit;
[0012] The mask buffer module is used to store the corresponding masks in the current acceleration stage, and the categories of the masks are input masks, reference masks, or output masks;
[0013] The coarse-grained matching unit is used to obtain a plurality of valid mask groups from the mask buffer module, and, according to an allocation request sent by any target processing unit in the corresponding row, allocate any target valid mask group in the plurality of valid mask groups to the target processing unit, and the valid mask group is a mask group that can perform non-zero convolution operations, and the mask group includes the input mask, the reference mask, and the output mask;
[0014] The target processing unit is configured to obtain target input values from the input value buffer module according to the target valid mask group, obtain target reference values from the reference value buffer module, and obtain target output values according to the target input values and the target reference values;
[0015] The accumulation module is configured to determine the sum of the target output values output by all processing units as the accumulated output value.
[0016] In an implementable manner, the input value buffer module includes a plurality of input row buffers, and each input row buffer corresponds to one row of the processing units and is configured to store the input values allocated to the corresponding row of processing units in the current acceleration stage.
[0017] In an implementable manner, the coarse-grained matching unit includes:
[0018] An arbiter, configured to receive the allocation request sent by any target processing unit in the corresponding row and convert the allocation request into allocation information, where the allocation information includes an allocation address and allocation data;
[0019] A mask matching detector, configured to pre-group the masks in the mask buffer module to obtain a plurality of mask groups, and detect the plurality of mask groups. When performing the pre-grouping, a plurality of input masks and a plurality of reference masks share one output mask;
[0020] A matching information register bank, configured to determine, according to the detection results of the plurality of mask groups, the mask groups that can perform non-zero convolution operations in the plurality of mask groups as valid mask groups, where each valid mask group corresponds to an address information and a data information, and allocate the target data information of any one target valid mask group in the plurality of valid mask groups to the processing unit according to the allocation information;
[0021] An address storage module, configured to store the address information of the plurality of valid mask groups and match the target address information from the plurality of address information according to the allocation address;
[0022] A data storage module, configured to store the data information of the plurality of valid mask groups and select the target data information according to the target address information;
[0023] A mask matching updater, configured to update the allocation information of the remaining valid mask groups after each allocation in the matching information register bank.
[0024] In an implementable manner, the processing unit includes:
[0025] A fine-grained mask matching unit, configured to obtain an input index and a reference index according to a target valid mask group, where the input index is used to determine a target input value from a corresponding input line buffer, and the reference index is used to determine a target reference value from the reference value buffer module;
[0026] A reference value register file, configured to store the reference values;
[0027] A multiply-accumulate unit, configured to determine a product of the target reference value and the target input value as a target output value;
[0028] A partial sum register file, configured to store the target output value and output the target output value to the accumulation module.
[0029] In one implementable manner, the fine-grained mask matching unit includes:
[0030] A reference mask register file, configured to store a plurality of reference masks;
[0031] A first priority encoder, configured to determine an output mask index from a target output mask, where the output mask index is used to determine a target reference mask from the reference mask register file, and the target output mask is an output mask in the target valid mask group;
[0032] An AND operation unit, configured to perform an AND operation on the target reference mask and a target input mask to obtain an AND operation result, where the target input mask is an input mask in the target valid mask group;
[0033] A second priority encoder, configured to output an AND operation result index according to the AND operation result;
[0034] A prefix sum operation unit, configured to obtain a target reference prefix sum according to the AND operation result index and the target reference mask and determine the target reference prefix sum as the reference index, and obtain a target input prefix sum according to the AND operation result index and the target input mask and determine the target input prefix sum as the input index.
[0035] In one implementable manner, the AND operation result index is the number of bits of the first value of 1 except the first bit in the AND operation result, and the first bit is counted from the 0th bit.
[0036] In one implementable manner, the obtaining a target reference prefix sum according to the AND operation result index and the target reference mask and determining the target reference prefix sum as the reference index includes:
[0037] Determining a corresponding target reference mask bit number in the target reference mask according to the AND operation result index;
[0038] Determine the sum of all values equal to one before the target reference mask bit number as the target reference prefix sum;
[0039] Determine the target reference prefix sum as the reference index.
[0040] In one implementable manner, the step of obtaining a target input prefix sum according to the AND operation result index and the target input mask, and determining the target input prefix sum as the input index includes:
[0041] Determine the corresponding target input mask bit number in the target input mask according to the AND operation result index;
[0042] Determine the sum of all values equal to one before the target input mask bit number as the target input prefix sum;
[0043] Determine the target input prefix sum as the reference index.
[0044] In one implementable manner, the accumulation module includes an accumulator and an accumulated value buffer, where:
[0045] The accumulator is used to determine the sum of the target output values output by all processing units as the accumulated output value;
[0046] The accumulated value buffer is used to store the accumulated output value.
[0047] A sparse accelerator applied to on-chip training provided by an embodiment of the present application dynamically adjusts multiple input values in an input value buffer module, multiple reference values in a reference value buffer module, and masks in a mask buffer module at different acceleration stages, so that coarse-grained units perform a preliminary screening of invalid operations in a coarse-grained manner, and fine-grained units included in each processing unit in the processing module perform a further screening of invalid operations in a fine-grained manner, thereby eliminating all invalid operations in the three training stages of on-chip training. In addition, multiple computing cores can perform parallel acceleration processing, and multiple processing units included in the processing module in multiple computing cores can also perform parallel acceleration processing, further improving the hardware utilization rate of the sparse accelerator. In this way, a sparse accelerator applied to on-chip training in the present application can efficiently and accurately eliminate all invalid operations in the three stages of on-chip training. Description of the Drawings
[0048] Figure 1 Sparse feature diagrams for different stages of on-chip training;
[0049] Figure 2 Sparse storage format diagram used by the sparse accelerator in the embodiment of the present application;
[0050] Figure 3Schematic diagram of the utilization mechanism of three types of sparsity by the sparse accelerator in the embodiments of the present application;
[0051] Figure 4 Schematic diagram of the structure of the computing core provided by the embodiments of the present application;
[0052] Figure 5 Schematic diagram of the structure of the coarse-grained matching unit provided by the embodiments of the present application;
[0053] Figure 6 Schematic diagram of the structure of the processing unit provided by the embodiments of the present application;
[0054] Figure 7 Schematic diagram of the structure of the fine-grained mask matching unit provided by the embodiments of the present application;
[0055] Figure 8 Schematic diagram of the comprehensive experimental results provided by the embodiments of the present application and comparison with other works;
[0056] Figure 9 Schematic diagram of the actual network analysis results provided by the embodiments of the present application. Detailed implementation manners
[0057] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0058] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms "a", "an", "the", "above", "said", "this" are also intended to include, for example, the expression "one or more", unless there is a clear indication to the contrary in the context. It should also be understood that in the following embodiments of the present application, "at least one", "one or more" means one, two, or more than two, and "multiple" means two or more than two. The term " / and / " is used to describe the association relationship of associated objects and indicates that three relationships may exist; for example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0059] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0060] First, the generation of invalid operations in the three stages of on-chip training will be introduced below in conjunction with the accompanying drawings.
[0061] The invalid operations in the three stages of on-chip training are generated from three types of sparsity in on-chip training, and the three types of sparsity are respectively the sparsity in the input activation layer (InputActivation, IA), the sparsity in the weight (Weight, WE), and the sparsity in the output activation layer (Output Activation, OA).
[0062] Figure 1 Exemplarily, a schematic diagram of the sparsity characteristics in different stages of on-chip training is shown, as Figure 1 shown:
[0063] In the FP stage, the training feature maps pass through the entire network layer by layer, and according to the type of network layer, the activation values of each layer are calculated sequentially through corresponding operations. In this process, the method of weight pruning is usually used to remove the weights with values lower than the preset threshold, so as to reduce model overfitting while maintaining the accuracy of the operations. In this process, sparsity in the weights will be generated. In addition, in the FP stage, the ReLU function is also used to convert the negative values in the activation layer into 0, thus bringing sparsity to the input activation layer when passing to the next layer of neural network operations.
[0064] In the BP phase, starting from the loss calculated in the FP phase, the gradients of the activation values are calculated layer by layer, i.e., the error values. The error values are backpropagated through each layer of the network from the end of the network. When backpropagating through the ReLU layer, the gradient values are also reset to 0, resulting in the sparsity of the weights. The sparse characteristics in the output activation layer in the BP phase are obtained according to the calculations in the FP phase. The batch normalization layer (Batch Normalization, BN) is used to stretch and shift the data, changing the data distribution of the input activation layer. Further, in the presence of BN, stochastic pruning is introduced. Similar to the forward weight pruning, stochastic pruning randomly prunes the gradients after backpropagation, so that smaller gradients will also be set to 0 by stochastic pruning. In this way, even in the presence of BN, zero values can be introduced, thus generating the sparsity of the gradients in the BP phase.
[0065] In the WG phase, according to the chain rule, the gradients of the weights are calculated by using the activation values (a l ) and error values (e l ) generated in the FP phase. The gradients of the weights are the error values. Therefore, when the weights are pruned and become sparse, the gradients of the weights are also sparse.
[0066] The sparse storage format used by the sparse accelerator in the embodiments of the present application will be introduced below with reference to the accompanying drawings.
[0067] Figure 2 An exemplary schematic diagram of the sparse storage format used by the sparse accelerator in the embodiments of the present application is shown, as Figure 2 shown. Such a storage format includes: MaskArray: Indexed by the coordinates of the blocks, storing a mask array of a fixed length. The non-zero values in the block correspond to 1s at the corresponding positions in the mask, and the zero values in the block correspond to 0s at the corresponding positions in the mask; PointerArray, also indexed by the coordinates of the blocks, with the pointer pointing to the starting point of the corresponding value array; ValueArray, sequentially containing all the non-zero data in the block. At the same time, the length can be changed according to different sparsity levels. In this way, the data is stored in block granularity, and the PointerArray is decoupled from the ValueArray. Therefore, when fetching data, it is more efficient to calculate the 180-degree rotated convolutional kernel and the transpose of the matrix. Therefore, the sparse storage format is suitable for the three stages of on-chip training.
[0068] The utilization mechanisms of the three types of sparsity by the sparse accelerator in the embodiments of the present application will be introduced below with reference to the accompanying drawings.
[0069] Figure 3 An exemplary schematic diagram of the utilization mechanisms of the three types of sparsity by the sparse accelerator in the embodiments of the present application is shown, as Figure 3As shown, for the position of 0 in OA, the inner products of the corresponding IA and WE will all be skipped. For the position of 1 in OA, only the positions where both IA and WE are 1 need to perform multiplication and accumulation operations. Therefore, only Figure 3 the 1 in the dashed box in Figure 3 needs to perform multiplication and accumulation operations.
[0070] To accelerate on-chip training and eliminate all invalid operations in on-chip training, an embodiment of the present application provides a sparse accelerator applied to on-chip training. The sparse accelerator includes a plurality of computing cores, and each computing core is used to accelerate a convolution operation process each time.
[0071] Figure 4 Exemplarily shows a schematic structural diagram of the computing core provided by an embodiment of the present application, as Figure 4 shown. The computing core includes an input value buffer module 100, a reference value buffer module 200, a mask buffer module 300, a coarse-grained matching module 400, a processing module 500, and an accumulation module 600. The coarse-grained matching module 400 includes a plurality of coarse-grained matching units 410, and the processing module 500 includes a plurality of processing units 510 arranged in an array. Each row of processing units 510 shares a coarse-grained matching unit 410; where:
[0072] The input value buffer module 100 is used to allocate the input values in the current acceleration stage to each row of processing units 510. The current acceleration stage is the forward propagation stage, the backward propagation stage, or the weight gradient calculation stage.
[0073] Specifically, the input value buffer module 100 includes a plurality of input row buffers 110. Each input row buffer 110 corresponds to one row of the processing units and is used to store the input values allocated to the corresponding row of processing units in the current acceleration stage.
[0074] Specifically, when the current acceleration stage is the forward propagation stage, the input value is the input activation layer value; when the current acceleration stage is the backward propagation stage, the input value is the error value of the next layer; when the current acceleration stage is the weight gradient calculation stage, the input value is the output activation layer value of the previous layer.
[0075] The reference value buffer module 200 is used to allocate the reference values in the current acceleration stage to each processing unit.
[0076] Specifically, when the current acceleration stage is the forward propagation stage, the reference value is the weight value of the current layer; when the current acceleration stage is the backward propagation stage, the reference value is the weight value of the current layer; when the current acceleration stage is the weight gradient calculation stage, the reference value is the error value of the current layer.
[0077] The mask buffer module 300 is configured to store the corresponding masks in the current acceleration stage, and the categories of the masks are input masks, reference masks, or output masks.
[0078] Specifically, when the current acceleration stage is the forward propagation stage, the masks include an input activation layer mask, a weight mask, and a preset output activation layer mask. At this time, the values in the preset output activation layer mask are all 1. After the forward propagation stage ends, an actual output activation layer mask is obtained according to the actually calculated output activation layer data, and the actual output activation layer mask is stored in the mask buffer module 300; when the current acceleration stage is the backward propagation stage, the masks include an error value mask of the subsequent layer, a weight mask, and a current layer error value mask; when the current acceleration stage is the weight gradient calculation stage, the masks include the output activation layer mask obtained in the forward propagation stage, the current layer error value mask, and the current layer weight mask.
[0079] The coarse-grained matching unit 410 is configured to obtain multiple valid mask groups from the mask buffer module 300, and, according to an allocation request sent by any one of the target processing units 510 in the corresponding row, allocate any one of the target valid mask groups in the multiple valid mask groups to the target processing unit. The valid mask group is a mask group that can perform non-zero convolution operations, and the mask group includes the input mask, the reference mask, and the output mask;.
[0080] The target processing unit 510 is configured to obtain a target input value from the input value buffer module 100 according to the target valid mask group, and obtain a target reference value from the reference value buffer module 200, and obtain a target output value according to the target input value and the target reference value.
[0081] The accumulation module 600 is configured to determine the sum of the target output values output by all the processing units 510 as the accumulated output value.
[0082] Specifically, the accumulation module 600 includes an accumulator 610 and an accumulated value buffer 620, where:
[0083] The accumulator 610 is configured to determine the sum of the target output values output by all the processing units as the accumulated output value;
[0084] The accumulated value buffer 620 is configured to store the accumulated output value.
[0085] Specifically, when the current acceleration stage is the forward propagation stage, the output value is the output activation layer value of the current layer; when the current acceleration stage is the backward propagation stage, the output value is the error value of the current layer; when the current acceleration stage is the weight gradient calculation stage, the output value is the weight value of the current layer.
[0086] Figure 5 Exemplarily shown is a schematic structural diagram of a coarse-grained matching unit provided by an embodiment of the present application, as Figure 5 shown, the coarse-grained matching unit 410 includes:
[0087] An arbiter 411, configured to receive the allocation request sent by any target processing unit 510 corresponding to a row, and convert the allocation request into allocation information, where the allocation information includes an allocation address and allocation data;
[0088] A mask matching detector 412, configured to pre-group the masks in the mask buffer module to obtain a plurality of mask groups, and detect the plurality of mask groups. When performing the pre-grouping, a plurality of input masks and a plurality of reference masks share one output mask;
[0089] A matching information register bank 413, configured to determine, according to the detection results of the plurality of mask groups, the mask groups that can perform non-zero convolution operations in the plurality of mask groups as valid mask groups, each valid mask group corresponding to an address information and a data information, and allocate the target data information of any target valid mask group in the plurality of valid mask groups to the processing unit according to the allocation information;
[0090] An address storage module 414, configured to store the plurality of address information of the plurality of valid mask groups, and match target address information from the plurality of address information according to the allocation address;
[0091] A data storage module 415, configured to store the plurality of data information of the plurality of valid mask groups, and select the target data information according to the target address information;
[0092] A mask matching updater 416, configured to update the allocation information of the remaining valid mask groups after each allocation in the matching information register bank.
[0093] Figure 6 Exemplarily shown is a schematic structural diagram of a processing unit provided by an embodiment of the present application, as Figure 6 shown, the processing unit 510 includes:
[0094] The fine-grained mask matching unit 511 is configured to obtain an input index and a reference index according to a target valid mask group. The input index is used to determine a target input value from a corresponding input row buffer, and the target input row buffer is the input row buffer corresponding to the current processing unit. The reference index is used to determine a target reference value from the reference value buffer module;
[0095] The reference value register bank 512 is configured to store a plurality of reference values;
[0096] The multiply-accumulate unit 513 is configured to determine a product of the target reference value and the target input value as a target output value;
[0097] The partial sum register bank 514 is configured to store the target output value and output the target output value to the accumulation module.
[0098] Figure 7 An exemplary structural schematic diagram of the fine-grained mask matching unit provided by an embodiment of the present application is shown, as Figure 7 shown, the fine-grained mask matching unit 511 includes:
[0099] The reference mask register bank 5111 is configured to store the reference mask;
[0100] The first priority encoder 5112 is configured to determine an output mask index from a target output mask. The output mask index is used to determine a target reference mask from the reference mask register bank, and the target output mask is the output mask in the target valid mask group;
[0101] The AND operation unit 5113 is configured to perform an AND operation on the target reference mask and a target input mask to obtain an AND operation result. The target input mask is the input mask in the target valid mask group;
[0102] Specifically, the AND operation result index is the number of bits of the first value of 1 except the first bit in the AND operation result, and the first bit starts from the 0th bit.
[0103] The second priority encoder 5114 is configured to output an AND operation result index according to the AND operation result;
[0104] The prefix sum operation unit 5115 is configured to obtain a target reference prefix sum according to the AND operation result index and the target reference mask, and determine the target reference prefix sum as the reference index,
[0105] Specifically, obtaining a target reference prefix sum according to the AND operation result index and the target reference mask, and determining the target reference prefix sum as the reference index includes:
[0106] Determine the corresponding target reference mask bit in the target reference mask according to the AND operation result index;
[0107] Determine the sum of all values that are 1 before the target reference mask bit as the target reference prefix sum;
[0108] Determine the target reference prefix sum as the reference index.
[0109] The prefix sum calculator 5115 is further configured to obtain a target input prefix sum according to the AND operation result index and the target input mask, and determine the target input prefix sum as the input index.
[0110] Specifically, the step of obtaining a target input prefix sum according to the AND operation result index and the target input mask, and determining the target input prefix sum as the input index includes:
[0111] Determine the corresponding target input mask bit in the target input mask according to the AND operation result index;
[0112] Determine the sum of all values that are 1 before the target input mask bit as the target input prefix sum;
[0113] Determine the target input prefix sum as the reference index.
[0114] Next, a specific description is given of the specific working process when the sparse accelerator provided in the embodiments of the present application performs acceleration.
[0115] Before on-chip training, for each current layer, store the input activation layer data as input activation layer values and input activation layer masks according to the sparse storage format, store the weight data as weight values and weight masks, and preset that all output activation layer masks are 1, and transfer the above data from the external DRAM (Dynamic Random Access Memory) to the corresponding buffer area.
[0116] In the forward propagation stage, obtain output activation layer data and error data according to the input activation layer values, input activation layer masks, weight values, weight masks, and preset output activation layer masks, and store the output activation layer data as output activation layer values and output activation layer masks according to the sparse storage format, and store the error data as error values and error masks;
[0117] In the backward propagation stage, determine the error value of the current layer according to the error value of the subsequent layer, the error mask of the subsequent layer, the weight value of the current layer, the weight mask of the current layer, and the error mask of the current layer;
[0118] In the weight gradient update calculation stage, the weight value of the current layer is determined based on the output activation layer value of the previous layer, the output activation layer mask of the previous layer, the error value of the current layer, the error mask of the current layer, and the weight mask of the current layer.
[0119] To evaluate the feasibility and performance of the sparse accelerator provided by the embodiments of the present application, first, the hardware design is implemented using SystemVerilog, then synthesis is performed using Design Compiler, the process library of TSMC 28nm is selected, the operating frequency is 200 MHz, the power consumption is evaluated by PT-PX in the average mode, and the area and power consumption of the buffer (SRAM) are evaluated by the CACTI tool.
[0120] Figure 8 An exemplary illustration of the experimental synthesis results provided by the embodiments of the present application and a comparison diagram with other works are shown, as Figure 8 shown, when the sparsity of the input activation layer, weights, and output activation layer is all 90%, the sparse accelerator provided by the embodiments of the present application can achieve 42.1 TOPS and 174.0 TOPS / W in throughput and energy efficiency respectively. Both of these performances exceed the previous works. This is because the sparse accelerator provided by the embodiments of the present application makes full use of three types of sparsity and eliminates all invalid operations in the three stages of on-chip training. At the same time, the design of two-level mask matching further reuses hardware units, reduces the area and energy consumption, and at the same time enables the processing units to work independently, thereby maintaining a high utilization rate.
[0121] Figure 9 An exemplary illustration of the actual network analysis results provided by the embodiments of the present application is shown, as Figure 9 shown, in the analysis of the actual network, the embodiments of the present application designed a cycle-accurate simulator to evaluate the acceleration effect of the sparse accelerator provided by the embodiments of the present application on the CIFAR10 dataset in the ResNet-50 network. It can be seen that since the sparse accelerator provided by the embodiments of the present application utilizes three types of sparsity in the BP stage and the WG stage, the throughput in the BP stage and the WG stage has been greatly improved compared with the FP stage.
[0122] Thus, a sparse accelerator for on-chip training provided by an embodiment of the present application dynamically adjusts multiple input values in an input value buffer module, multiple reference values in a reference value buffer module, and masks in a mask buffer module during different acceleration stages, so that coarse-grained units preliminarily screen out invalid operations in a coarse-grained manner, and fine-grained units included in each processing unit in a processing module further screen out invalid operations in a fine-grained manner, thereby eliminating all invalid operations in the three training stages of on-chip training. In addition, multiple computing cores can perform parallel acceleration processing, and multiple processing units included in the processing module in multiple computing cores can also perform parallel acceleration processing, further improving the hardware utilization rate of the sparse accelerator. Thus, a sparse accelerator for on-chip training according to the present application can efficiently and accurately eliminate all invalid operations in the three stages of on-chip training.
[0123] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed by the present invention; the specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.
[0124] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope; the scope of the present invention is only limited by the appended claims.
Claims
1. A sparse accelerator applied to on-chip training, characterized in that, The sparse accelerator includes a plurality of computing cores, each computing core is used to accelerate a convolution operation process each time, and the computing core includes an input value buffer module, a reference value buffer module, a mask buffer module, a coarse-grained matching module, a processing module, and an accumulation module; The coarse-grained matching module includes a plurality of coarse-grained matching units, and the processing module includes a plurality of processing units arranged in an array, and each row of processing units shares a coarse-grained matching unit; wherein: The input value buffer module is used to allocate input values in the current acceleration stage to each row of processing units, and the current acceleration stage is the forward propagation stage, the backward propagation stage, or the weight gradient calculation stage; The reference value buffer module is used to allocate reference values in the current acceleration stage to each processing unit; The mask buffer module is used to store the corresponding mask in the current acceleration stage, and the categories of the mask are input mask, reference mask, or output mask; The coarse-grained matching unit is used to obtain a plurality of valid mask groups from the mask buffer module, and, according to an allocation request sent by any target processing unit in the corresponding row, allocate any target valid mask group in the plurality of valid mask groups to the target processing unit, and the valid mask group is a mask group for performing a non-zero convolution operation, and the mask group includes the input mask, the reference mask, and the output mask; The target processing unit is used to obtain a target input value from the input value buffer module according to the target valid mask group, and obtain a target reference value from the reference value buffer module, and obtain a target output value according to the target input value and the target reference value; The accumulation module is used to determine the sum of the target output values output by all processing units as the accumulated output value.
2. The sparse accelerator applied to on-chip training according to claim 1, wherein The input value buffer module includes a plurality of input row buffers, and each input row buffer corresponds to a row of the processing units and is used to store the input values allocated to the corresponding row of processing units in the current acceleration stage.
3. A sparse accelerator applied to on-chip training according to claim 1, characterized in that, The coarse-grained matching unit includes: An arbiter, which is used to receive the allocation request sent by any target processing unit in the corresponding row and convert the allocation request into allocation information, and the allocation information includes an allocation address and allocation data; A mask matching detector, which is used to pre-group the masks in the mask buffer module to obtain a plurality of mask groups, and detect the plurality of mask groups. When performing the pre-grouping, a plurality of input masks and a plurality of reference masks share an output mask; A matching information register bank, which is used to determine, according to the detection results of the plurality of mask groups, the mask groups for performing non-zero convolution operations in the plurality of mask groups as valid mask groups, each valid mask group corresponds to an address information and a data information, and allocate the target data information of any one target valid mask group in the plurality of valid mask groups to the processing unit according to the allocation information; An address storage module, which is used to store the plurality of address information of the plurality of valid mask groups, and match the target address information from the plurality of address information according to the allocation address; A data storage module, configured to store multiple data information of multiple valid mask groups and select the target data information according to the target address information; A mask matching updater, configured to update the allocation information of the remaining valid mask groups after each allocation in the matching information register bank.
4. A sparse accelerator applied to on-chip training according to claim 2, wherein The processing unit includes: A fine-grained mask matching unit, configured to obtain an input index and a reference index according to a target valid mask group, where the input index is used to determine a target input value from a corresponding input line buffer, and the reference index is used to determine a target reference value from the reference value buffer module; A reference value register bank, configured to store multiple reference values; A multiply-accumulate unit, configured to determine a product of the target reference value and the target input value as a target output value; A partial sum register bank, configured to store the target output value and output the target output value to the accumulation module.
5. A sparse accelerator applied to on-chip training according to claim 4, characterized in that The fine-grained mask matching unit includes: A reference mask register bank, configured to store the reference mask; A first priority encoder, configured to determine an output mask index from a target output mask, where the output mask index is used to determine a target reference mask from the reference mask register bank, and the target output mask is an output mask in the target valid mask group; An AND operation unit, configured to perform an AND operation on the target reference mask and a target input mask to obtain an AND operation result, where the target input mask is an input mask in the target valid mask group; A second priority encoder, configured to output an AND operation result index according to the AND operation result; A prefix sum operation unit, configured to obtain a target reference prefix sum according to the AND operation result index and the target reference mask, and determine the target reference prefix sum as the reference index, and obtain a target input prefix sum according to the AND operation result index and the target input mask, and determine the target input prefix sum as the input index.
6. The sparse accelerator applied to on-chip training according to claim 5, wherein The AND operation result index is the number of bits of the first value of 1 except the first bit in the AND operation result, and the first bit starts from the 0th bit.
7. A sparse accelerator for on-chip training according to claim 5, characterized in that, The step of obtaining a target reference prefix sum according to the AND operation result index and the target reference mask, and determining the target reference prefix sum as the reference index includes: Determining a corresponding target reference mask bit number in the target reference mask according to the AND operation result index; Determining the sum of all values of 1 before the target reference mask bit number as the target reference prefix sum; Determining the target reference prefix sum as the reference index.
8. A sparse accelerator applied to on-chip training according to claim 5, characterized in that The step of obtaining a target input prefix sum according to the AND operation result index and the target input mask, and determining the target input prefix sum as the input index includes: Determining a corresponding target input mask bit number in the target input mask according to the AND operation result index; Determining the sum of all values of 1 before the target input mask bit number as the target input prefix sum; Determining the target input prefix sum as the reference index.
9. A sparse accelerator applied to on-chip training according to claim 1, characterized in that, The accumulation module includes an accumulator and an accumulated value buffer, where: The accumulator is configured to determine the sum of all target output values output by the processing units as an accumulated output value; The accumulation value buffer is used to store the accumulated output value.
Citation Information
Patent Citations
A sparse convolutional neural network accelerator and an implementation method
CN109635944A
Deep learning training hardware accelerator utilizing sparsity
CN111368988A