Neural network accelerator implementation system and method based on adaptive allocation

By adopting adaptively allocated dual-data bit accelerator design in neural network accelerators, using combined bit valid terms and shift accumulation operations, the problem of difficult to simultaneously utilize activation and weighted data bit sparsity in the prior art is solved, and more efficient calculations and better energy efficiency ratios are achieved.

CN115081608BActive Publication Date: 2025-05-13SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210750313.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-05-13
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

Existing neural network accelerators are difficult to effectively utilize the bit sparse characteristics of activation data and weighted data at the same time, resulting in ineffective computing efficiency and poor energy efficiency.

Method used

By building a neural network accelerator implementation system based on adaptive allocation, the overall architecture of activation and weighted dual data bit accelerator is adopted, combining effective term generation units and calculation arrays, and using the expression of combined effective term and shift accumulation operations to optimize data flow organization and synchronization to achieve dynamic load balancing.

Benefits of technology

The number of cycles of dual-data bit serial calculations is significantly shortened, PE utilization is improved, computing efficiency is enhanced, and the energy efficiency ratio of neural network accelerators is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115081608B_ABST
    Figure CN115081608B_ABST
Patent Text Reader

Abstract

The present invention provides a neural network accelerator implementation system and method based on adaptive allocation, including: module M1: constructing the overall architecture of the activation and weight double data bit accelerator, including DRAM and data loading module, write-back module, on-chip cache module, valid item generation unit and calculation array, and the connection relationship between the modules; module M2: constructing the expression mode of activation data and weight data valid item, and constructing the activation data and weight data valid item generation unit and shift accumulation operation unit according to the expression mode; module M3: determining the data flow organization mode in the calculation array, performing data grouping and synchronization, and constructing the weight data combination bit valid item expression mode. After the activation data and weight data are detected as valid bits, the present invention reduces the number of valid items in the double data bit serial calculation through the expression method of the weight data combination bit valid item, and shortens the calculation cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep neural network accelerator design, and in particular, to a neural network accelerator implementation system and method based on adaptive allocation, and in particular, to a neural network acceleration method based on adaptive allocation of activation / weight dual data bit-level sparse. Background Art

[0002] The data in the neural network is sparse. On the one hand, lightweight operations such as pruning and quantization are widely used in current deep neural networks, resulting in a large number of elements with a value of 0 distributed in the weight data of the network. On the other hand, the existence of activation layers such as ReLU in the network structure also results in a large proportion of 0 elements in the activation data. According to these two sources of sparsity in the neural network, the sparsity of the data can be divided into weight data sparsity and activation data sparsity. The sparsity of weight data is determined after the network training is completed and exists statically during the network forward reasoning process; the sparsity of activation data is generated by the real-time operation of the activation layer during the network forward reasoning process and cannot be predicted in advance, so it manifests as dynamic sparsity. Furthermore, according to the different manifestations of sparsity in the same data, it can be divided into value sparsity and bit sparsity. Value sparsity means that the entire data is 0, and bit sparsity means that several bits in a single data are 0. Compared with value sparsity, bit sparsity is a deeper and more fine-grained mining of data sparsity.

[0003] Traditional neural network accelerators cannot make good use of the bit sparsity characteristics of activation data and weight data at the same time. The computing unit of traditional neural network accelerators is often a multiplication and addition operation of fixed-point data, ignoring the large number of bits with a value of 0 in a single data. When utilizing the sparsity of data, the main exploration is to utilize the sparsity of values, and ignore zero values ​​during data transmission or data calculation by compression, skipping, etc. In the prior art, the activation data is compressed on the chip in real time through the valid data detection module and stored in the queue of valid data. During calculation, the activation data is sent to the computing unit, while the weight data needs to be indexed according to the location information of the activation data. The acquisition of activation data and the acquisition of weight data are serial behaviors, the computing efficiency is limited, and the weights need to be indexed through complex modules. In the prior art, through the design of direct indexing and stride indexing, the efficient operation of indexing is explored, but the circuit of its index module is increased. In the overall power consumption evaluation, the power consumption of the index module accounts for 34.83% of the total power consumption, affecting the energy efficiency ratio of the accelerator. The bit-serial calculation design proposed in the prior art inputs the activation data bits serially and the weight data bits in parallel, and uses a bit-serial calculation method on the activation data to skip the high-order continuous 0 bits, but only uses the sparsity of the high-order continuous 0 bits in the activation data. In the prior art, a bit-sparse design method and calculation architecture are used for both activation data and weight data. However, when the valid bit independent index calculation method is used for both input data, unnecessary invalid calculations will be introduced into the overly long serial convolution (shift accumulation) calculation process, affecting the calculation performance. At the same time, operations such as encoding are more complicated to implement, which increases the complexity of accelerator design.

[0004] In order to solve the problem of low efficiency of cumulative shift serial calculation caused by independent valid term representation methods when utilizing dual data bit sparsity of activation and weight at the same time, a weight data representation method of combined bit valid terms is proposed to improve computing performance. At the same time, in order to solve the problem of performance loss of computing units caused by the coarse-grained synchronization method that takes the worst case performance when the bit sparsity of activation data is inconsistent, a dynamic load balancing technology of computing data prefetching and adaptive distribution is proposed to further optimize the performance of neural network accelerators based on bit-level sparsity.

[0005] Patent document CN113191493A (application number: CN202110461762.1) discloses a convolutional neural network accelerator based on FPGA parallelism adaptation, including: a read command generator, a data distributor, an operation cluster group, an addition tree group, an output cache group and an output arbitrator. The accelerator configures the parallelism to multiple or single output activation parallelism depending on the structure of the convolution layer. The data distributor can analyze the data consistency of the on-chip cache and broadcast the repeated off-chip data to the corresponding cache at the same time. Multiple operation units in the operation cluster are responsible for convolution operations on different input channels; according to the convolution operation structure, the operation cluster is configured for operations on different input channels or different output activations. The output cache group contains a data routing module. For operation clusters that operate on the same output activation, the data in the corresponding output cache is connected end to end, otherwise it is output as an independent output.

[0006] Both the input and weight data in the neural network calculation process have significant bit-sparse characteristics. Most existing neural network accelerator designs either fail to make good use of this characteristic to accelerate network calculations, or only use the bit-sparse characteristics of the activation / weight data alone, and do not use the bit-sparse characteristics of both input and weight data as a support point for accelerated calculations. Although support is provided for dual data sparsity, when it simultaneously exploits the bit-sparse characteristics of activation data and weight data, it does not provide a good solution to the problem of long calculation cycles caused by bit-by-bit serial calculations of dual data. Summary of the invention

[0007] In view of the defects in the prior art, the object of the present invention is to provide a neural network accelerator implementation system and method based on adaptive allocation.

[0008] The neural network accelerator implementation system based on adaptive allocation provided by the present invention includes:

[0009] Module M1: Constructs the overall architecture of the activation and weight dual-data bit accelerator, including DRAM and data loading modules, write-back modules, on-chip cache modules, valid item generation units and computing arrays, as well as the connection relationship between each module;

[0010] Module M2: construct an expression of activation data and weight data valid items, and construct an activation data and weight data valid item generation unit and a shift accumulation operation unit according to the expression;

[0011] Module M3: Determine the data flow organization method in the computing array, perform data grouping and synchronization, and construct the weighted data combination bit valid item expression method.

[0012] Preferably, the data interaction between the on-chip cache module and the DRAM is completed through LOAD and STORE instructions, and the activation data and the weight data are converted from the on-chip cache to a valid item expression form that can be recognized by the computing array through the valid item generation unit;

[0013] The computing array includes a VPU vector computing unit, and the activation data vector and the weight data vector are deployed to the VPU. The VPU includes a basic computing unit PE, and the PE obtains valid items of activation data and weight data from the data pool of the VPU to complete the shift accumulation calculation.

[0014] Preferably, each bit of the activation data is represented by a sign field s, a value field v and a position field e, and non-negative weight data is represented in the same format; for negative weights, the negative number is represented in the form of an absolute value plus a sign bit, and any number of bits of the weight data are combined, and the v field and the e field are modified to the combined expression;

[0015] The valid item detection unit only outputs non-zero bits in the data to the calculation array. The multiplication and addition operations of the activation data and the weight data are completed through shift-accumulation operations. When the weight data valid items are represented by single bits, the e fields of the weight data and the activation data valid items are added as shift factors, and 1 is shifted to obtain the calculation result. When the weight data valid items are represented by combined bit valid items, the weight data valid items are the operands and the activation data valid items are the shift factors. The shift-accumulation operations of m activation data valid items and n weight data valid items are completed in a serial manner, consuming m*n cycles.

[0016] Preferably, the activation data vector and the weight data vector are converted into valid items and grouped and stored in the valid item data pool of each VPU respectively. The PE takes out the activation data from the valid item data pool in the group, and selects the weight data according to the channel information of the activation data to complete the shift accumulation operation. The calculated result is accumulated to the output through the addition tree.

[0017] Preferably, a valid item of activation data represents:

[0018] A i =(-1) s ×v×2 e

[0019] Among them, A i is the value of a valid item, i is the sequence number; s is the symbol field; v is the value field; e is the position field;

[0020] Valid entries for weight data are:

[0021] The activation data A0 consists of n valid items, and the weight data W0 consists of l valid items. The bit-serial calculation requires n×l clock cycles to complete, and the expression is:

[0022]

[0023] By combining the expressions of bit valid items, the number of valid items of weight data is reduced to half, and the number of clock cycles required for bit serial calculation is also reduced to half accordingly. The expression is:

[0024]

[0025] Where N is the total number of activation data items; L is the total number of weight data items; m is the number of weight vectors of the data, m = l / 2;

[0026] Two-bit serial computing data stream:

[0027] When the valid term of the activation data is represented by a single bit, and the valid term of the weight is designed with a combined bit valid term, the calculation process is expressed as:

[0028]

[0029] Among them, the v field of the weight data is used as the operand of the shift operation, the sum of the e fields of the activation data and the weight data is used as the shift factor, the shift calculation result of each cycle is stored in the accumulation unit, and finally the output is obtained. When the weight data is represented in the form of a single-bit valid item, the operand of the shift operation becomes 1.

[0030] The method for implementing a neural network accelerator based on adaptive allocation provided by the present invention includes:

[0031] Step 1: Construct the overall architecture of the activation and weight dual data bit accelerator, including DRAM and data loading modules, write-back modules, on-chip cache modules, valid item generation units and computing arrays, as well as the connection relationship between each module;

[0032] Step 2: construct an expression of activation data and weight data valid items, and construct an activation data and weight data valid item generation unit and a shift accumulation operation unit according to the expression;

[0033] Step 3: Determine the data flow organization method in the computing array, perform data grouping and synchronization, and construct a weighted data combination bit valid item expression method.

[0034] Preferably, the data interaction between the on-chip cache module and the DRAM is completed through LOAD and STORE instructions, and the activation data and the weight data are converted from the on-chip cache to a valid item expression form that can be recognized by the computing array through the valid item generation unit;

[0035] The computing array includes a VPU vector computing unit, and the activation data vector and the weight data vector are deployed to the VPU. The VPU includes a basic computing unit PE, and the PE obtains valid items of activation data and weight data from the data pool of the VPU to complete the shift accumulation calculation.

[0036] Preferably, each bit of the activation data is represented by a sign field s, a value field v and a position field e, and non-negative weight data is represented in the same format; for negative weights, the negative number is represented in the form of an absolute value plus a sign bit, and any number of bits of the weight data are combined, and the v field and the e field are modified to the combined expression;

[0037] The valid item detection unit only outputs non-zero bits in the data to the calculation array. The multiplication and addition operations of the activation data and the weight data are completed through shift-accumulation operations. When the weight data valid items are represented by single bits, the e fields of the weight data and the activation data valid items are added as shift factors, and 1 is shifted to obtain the calculation result. When the weight data valid items are represented by combined bit valid items, the weight data valid items are the operands and the activation data valid items are the shift factors. The shift-accumulation operations of m activation data valid items and n weight data valid items are completed in a serial manner, consuming m*n cycles.

[0038] Preferably, the activation data vector and the weight data vector are converted into valid items and grouped and stored in the valid item data pool of each VPU respectively. The PE takes out the activation data from the valid item data pool in the group, and selects the weight data according to the channel information of the activation data to complete the shift accumulation operation. The calculated result is accumulated to the output through the addition tree.

[0039] Preferably, a valid item of activation data represents:

[0040] A i =(-1) s ×v×2 e

[0041] Among them, A i is the value of a valid item, i is the sequence number; s is the symbol field; v is the value field; e is the position field;

[0042] Valid entries for weight data are:

[0043] The activation data A0 consists of n valid items, and the weight data W0 consists of l valid items. The bit-serial calculation requires n×l clock cycles to complete, and the expression is:

[0044]

[0045] By combining the expressions of bit valid items, the number of valid items of weight data is reduced to half, and the number of clock cycles required for bit serial calculation is also reduced to half accordingly. The expression is:

[0046]

[0047] Where N is the total number of activation data items; L is the total number of weight data items; m is the number of weight vectors of the data, m = l / 2;

[0048] Two-bit serial computing data stream:

[0049] When the valid term of the activation data is represented by a single bit, and the valid term of the weight is designed with a combined bit valid term, the calculation process is expressed as:

[0050]

[0051] Among them, the v field of the weight data is used as the operand of the shift operation, the sum of the e fields of the activation data and the weight data is used as the shift factor, the shift calculation result of each cycle is stored in the accumulation unit, and finally the output is obtained. When the weight data is represented in the form of a single-bit valid item, the operand of the shift operation becomes 1.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] (1) After valid bit detection of activation data and weight data, the method of representing valid items of weight data combination bits is used to reduce the number of valid items in double data bit serial calculation, thereby shortening the calculation cycle;

[0054] (2) To solve the problem of different computational loads caused by different numbers of valid items in activation data and weight data allocated to different PEs, the activation data grouping method is used to put multiple groups of activation data and weight data into a shared data pool of multiple PEs, adaptively and dynamically allocate the computational load of PEs, thereby improving PE utilization.

[0055] (3) By pre-fetching activation data, the corresponding weight data is selected in advance according to the channel information of the activation data, filling the time difference between obtaining the activation data and obtaining the weight data, thereby improving the calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0057] Figure 1a It is a valid item representation for activation data; Figure 1b It is a valid term representation of negative numbers in weight data;

[0058] Figure 2 is the weight data combination bit valid item representation, (a) is positive value representation, (b) is negative value representation;

[0059] Figure 3 is a double data bit sparse computing unit data stream;

[0060] Figure 4 Group valid items for activation data and weight data;

[0061] Figure 5a It is a dual data bit sparse accelerator architecture diagram; Figure 5b It is a structural diagram of a dual data bit sparse accelerator VPU;

[0062] Figure 6 The figure is a speedup ratio diagram of the main convolutional layer in Resnet18 in the present invention relative to the traditional accelerator VTA;

[0063] Figure 7 This is a graph showing the speedup of a typical neural network on a dual-bit sparse accelerator;

[0064] Figure 8 This is a comparison chart of the acceleration area efficiency of the double-data-bit sparse accelerator. DETAILED DESCRIPTION

[0065] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0066] Example:

[0067] Based on the neural network accelerator based on the dynamic activation data bit sparsity, the present invention proposes a neural network accelerator design method based on the activation and weight double data bit sparsity. After the activation data and weight data are detected for valid bits, the weight data combination bit valid item representation method is used to reduce the number of valid items in the double data bit serial calculation, thereby shortening the calculation cycle. On the other hand, in order to solve the problem that the number of valid items in the activation data and weight data allocated to different PEs is different, resulting in different calculation loads, the activation data grouping method is used to put multiple groups of activation data and weight data into a shared data pool of multiple PEs, and the calculation load of PE is adaptively and dynamically allocated, thereby improving the PE utilization rate. At the same time, by pre-fetching the activation data method, the corresponding weight data is selected in advance according to the channel information of the activation data, and the time difference between obtaining the activation data and obtaining the weight data is filled, thereby improving the calculation efficiency. In the present invention, different convolutional layers on the Resnet18 neural network are mapped and deployed. Compared with the traditional neural network accelerator VTA, the calculation speed of different network layers is increased by 2.5-12 times. In the complete neural network calculation evaluation, the present invention achieved a 3.87-fold acceleration ratio in Resnet18, a 4.75-fold acceleration ratio in Resnet50, and a 2.57-fold acceleration ratio in Vgg16, with significant improvements in both calculation speed and efficiency.

[0068] Valid entries for activation data represent:

[0069] In order to utilize the bit sparseness in the activation data, the activation data A is represented as the sum of a set of valid items. Each valid item consists of a field s (sign bit), a field v (bit value), and a field e (bit position). The value of a valid item can be represented as follows:

[0070] A i =(-1) s ×v×2 e

[0071] Taking an 8-bit data 00011010 as an example, its valid item representation is as follows Figure 1a As shown in the figure, the bits in the field v that are 0 do not contribute to the value of the data and are invalid items and can be omitted. At the same time, the activation data in the neural network is often generated by the activation layer of the ReLU activation function or the Sigmoid activation function, and is usually expressed as a positive data value. For the activation data with a positive sign, its sign bit s can be ignored and only the position information of the generated bit can be used to represent it. The 8-bit data in Figure 1 can be represented by the position information items of the three e fields 1, 3, and 4.

[0072] Valid entries for weight data are:

[0073] After the weight data is also expressed in the form of the sum of valid items, the calculation with the activation data becomes a serial calculation of bitwise accumulation and shifting as shown in the following formula. The activation data A0 consists of n valid items, and the weight data W0 consists of l valid items. The bitwise serial calculation requires n×l clock cycles to complete, and the expression is:

[0074]

[0075] In view of the performance loss caused by the serial calculation method of the above formula, the present invention proposes a method for representing the combined bit valid items of weight data, generating data valid items according to a group of adjacent data bits c = 2 / 4 / 8 data, thereby reducing the number of weight data valid items. Taking c = 2 as an example, the combined bit valid item design method is as follows Figure 2 Compared with the single-bit valid item representation, the sign bit field s of the combined bit valid item representation method has not changed, the bit value field v represents the grouped 2-bit data, and the position field e represents the position of the lower bit of the two bits after the combination. By using the combined bit valid item representation method, the number of valid items of the weight data is reduced to half, and the number of clock cycles required for bit serial calculation is also reduced to half accordingly, as shown in the following formula:

[0076]

[0077] There are both negative and positive numbers in the weight data. If the weight data is decomposed into valid items based on the complement form, negative numbers often appear as more valid items, which cannot make good use of the bit sparsity of the data. The present invention converts negative numbers into the expression of their absolute values, and records the sign of the entire data as the sign of the generated valid item. In the process of converting the complement of a negative number to a positive number, it mainly includes two steps: bitwise inversion and addition:

[0078] -W0=|W0|=~W0+1,Sign=1

[0079] Taking the 8-bit complement representation of -1 as 11111111 as an example, after converting it into an absolute value representation, the entire data can be represented by one valid item, such as Figure 1b As shown. For positive numbers in weight data, there is no need to perform absolute value conversion.

[0080] Two-bit serial computing data stream:

[0081] The calculation unit PE designed in the present invention needs to process the calculation of the activation valid item and the weight valid item. When the valid item of the activation data is represented by a single bit, and the valid item of the weight is designed by a combination bit valid item, the calculation process can be expressed as:

[0082]

[0083] The v field of the weight data is used as the operand of the shift operation, the sum of the e fields of the activation data and the weight data is used as the shift factor, and the shift calculation result of each cycle is stored in the accumulation unit, and finally the output is obtained, such as Figure 3 When the weight data is represented in the form of a single-bit valid term, the operand of the shift operation becomes 1.

[0084] Valid item grouping with two-bit sparseness:

[0085] The activation data and weight data allocated to each PE may contain different numbers of valid items, and the computational loads of each PE are different. The computational resources of the array cannot be fully utilized, resulting in reduced computational efficiency. The present invention groups the activation data and weight data. For an activation vector containing m data and a weight vector containing m data, they can be divided into p groups. The number of activation data and weight data contained in each group is q=m / p. In the group calculation, the valid items of the activation data are placed in the data pool for balanced calculation by the computing PE array in the entire group. Correspondingly, the valid items of the weight data are stored separately. When the computing PE obtains the activation valid item data from the data pool of the activation valid item, it will act on the check signal before the computing PE through the activation channel information contained in the valid item data to complete the calculation of the corresponding activation valid items and weight valid items. For each group, the data flow of its calculation Figure 4 As shown. The activation data and weight data required for calculations between different valid item groups are different. The final calculation results within the group need to be accumulated between different groups to generate the final result of the vector dot product. The weight data under different valid item groups generates a weight cycle signal based on the judgment of the calculation process to act on the pointer control module of the activated valid item, thereby realizing the calculation control of the activation data and realizing the synchronization of the calculation completion signals between different groups.

[0086] Dual-bit sparse neural network accelerator architecture design and hardware implementation:

[0087] The architecture of the neural network accelerator based on double data bit sparsity of activation and weight is as follows Figure 5a As shown, it includes a data loading module Load, a computing path, a data write back module Store, and activation, weight, and output data on-chip caches. In the computing path, it includes an activation data valid item detection unit, a weight valid item detection unit and its corresponding cache, and also includes a computing array. The activation data and weight data are loaded from DRAM to the weight on-chip cache and the activation on-chip cache respectively through the data loading module, and then the corresponding valid items are generated by the activation data valid item generation unit and the weight data valid item generation unit respectively, and distributed to the computing array. For the computing array, it is further divided into the vector computing unit VPU level. The structure of the VPU is shown in the figure below. Figure 5bAs shown, each VPU adopts a valid item grouping design method to group the valid items generated by the activation data vector and the weight data vector, and completes the dot product and accumulation of the vectors by mixing the valid items and calculating the corresponding valid items of the activation data and the weight data to obtain the convolution result.

[0088] Evaluation Methodology:

[0089] The present invention selects the FPGA platform of DIGILENT ZYBO-z7 to realize the design of a neural network accelerator based on double-bit sparse activation and weight. The FPGA development board integrates the Xilinx ZYNQ-7020 chip, which includes programmable logic and an ARMCortex-A9 dual-core processor. In the accelerator design of the ZYBO-z7 development board, the accelerator module includes 1 computing core, 16 vector computing units VPU, that is, 256 computing units PE. In addition, the present invention completes the data loading LOAD module and the data writing back STORE module based on the AXI bus protocol of FPGA to realize the support for data loading instructions and data writing back instructions.

[0090] When deploying a neural network on an acceleration platform, we first transplant the Linux operating system onto the ARM Cortex-A9 processor, and control the entire hardware backend through bus communication between the system and the programming logic. In terms of performance evaluation, we supplement the development based on the open source TVM neural network compilation platform and software frameworks such as VTA Runtime, as well as the sparse neural network accelerator architecture proposed in this invention, generate the instruction stream of the bit sparse accelerator, and control the bit sparse neural network accelerator through the control of the ARM core system and the generation of accelerator instructions.

[0091] Resource Utilization and Area:

[0092] The implementation of the neural network accelerator based on double data bit sparsity of activation and weight on FPGA has a comprehensive clock running frequency of 100MHz. The resource utilization of the neural network accelerator can be evaluated through the FPGA synthesis tool. Table 1 shows the resource utilization of the double data bit sparsity neural network accelerator (activation grouping q = 4, weight valid item 2bit combination) compared with the traditional accelerator to implement VTA.

[0093] Compared with VTA, since the present invention replaces the multiplication and accumulation operations of the traditional neural network accelerator with shift-accumulation operations, the number of DSPs used in the FPGA resources is greatly reduced, and correspondingly the usage of LUTs and FFs is significantly increased compared with VTA.

[0094] Table 1 Comparison of FPGA resource utilization

[0095]

[0096] In order to further analyze the area overhead of each module of the bit-sparse neural network accelerator, the main component modules in the accelerator implementation are synthesized by the Design Compiler tool (activation group q=4), and the comprehensive area results of the key modules are shown in Table 2. As can be seen from the table, for the double-data-bit sparse accelerator design, although the combination bit will cause the area of ​​a single computing unit to increase, under the same computing power, the number of computing units required is reduced, so with the increase of weight combination bits, the computing unit area gradually decreases. At the same time, due to the storage of weight valid items in the double-data-bit sparse, the area overhead of the weight valid item data pool is large. When the weight combination valid item parameter c=8, the accelerator degenerates into a single-data-bit sparse accelerator with sparse activation data bits, and there is no weight data valid item pool. The added other key modules have little effect on the area of ​​the accelerator, and the main part is the area of ​​the computing core, weight valid item pool and on-chip cache. Compared with the MAC unit used in VTA and other traditional accelerators, the shift-accumulate unit used in the invention has a significant area reduction, but more shift-accumulate modules are required in the design implementation. The area of ​​modules such as valid item detection does not incur significant additional area overhead.

[0097] Table 2. Area comparison of main components (um 2 )

[0098]

[0099] Accelerator performance evaluation:

[0100] By mapping and deploying the main convolutional layers of ResNet18 on the double data bit sparse accelerator design, the acceleration performance of its convolution calculation is evaluated. In order to control the design variables, combined with the analysis of data sparse utilization and area ratio under different groupings, the design grouping of the activation data bit sparse is fixed to q=4 in the accelerator design, and the design space is explored for different numbers of weight combination bits, and the accelerator performance of the convolution layer is evaluated. The evaluation results of the double data bit sparse convolution layer are shown in Figure 2. Figure 6 Shown

[0101] The evaluation results show that, under the same activation grouping, the smaller the number of combination bits, the better the acceleration effect on the convolution layer calculation. When the number of weight combination effective items c = 8, it degenerates into a design with sparse activation data bits. In contrast, double data bit sparse has a significant improvement in the accelerator effect on the convolution layer after realizing the utilization of double data bit sparsity of activation and weight. And it has a larger acceleration ratio than the traditional neural network VTA.

[0102] In order to further verify the performance benefits of the double data bit sparse design in the neural network accelerator design, the present invention deploys neural networks (such as ResNet18, ResNet50, VGG16, etc.) on an FPGA-based software and hardware collaborative platform to complete performance testing. When executing network calculations, the main convolution calculation layers and some pooling layers in the neural network are unloaded from the calculation of the main control core and loaded into the calculation array of the accelerator to complete the acceleration of the neural network calculation. The results are shown in Figure 2. Figure 7 shown.

[0103] When the dual-data-bit sparse design is adopted, the sparsity of the weight data is utilized and the utilization rate of the sparsity in the neural network data is further improved. Compared with the design with sparse activation data bits (c=8), the acceleration effect of the neural network is more obvious. Compared with the VTA design with the same computing power, when the activation grouping q=4 and the weight combination c=2, in the ResNet18 neural network evaluation, the dual-data-bit sparse accelerator design achieves a 3.87x acceleration ratio, and in ResNet50 it reaches a 4.75x acceleration ratio, and in VGG16 it still has a 2.57x acceleration ratio.

[0104] like Figure 8 ,In order to further evaluate the performance improvement of the design of ,neural network accelerator based on double data bit sparse compared to other ,accelerator design methods of other single data bit sparse are simulated and implemented, and ,their implementation is analyzed and evaluated. ,The ratio of the performance speedup ratio and the normalized overall area - the ,acceleration area efficiency is used as the evaluation index to achieve a ,comparative analysis of the acceleration area efficiency of the accelerator.

[0105] Compared with the traditional VTA accelerator, the dual-data-bit sparse design has improved the acceleration area efficiency under different neural networks. Among them, in the evaluation of ResNet18, ResNet50, and VGG16, the neural network acceleration area efficiency was improved by 2.55 times, 3.36 times, and 1.26 times respectively under the configuration of activation group q=4 and weight combination bit c=2. Compared with the Bit-Tactical design of the prior art, the dual-data-bit sparse design proposed in the present invention has higher computing acceleration area efficiency in neural network acceleration calculation. Taking the ResNet50 network evaluation as an example, the dual-data-bit sparse design of the present invention is 32.4% higher than Bit-Tactical in acceleration area efficiency under the configuration of activation group q=4 and weight combination bit c=2. Compared with the prior art Laconic, also taking the neural network implementation of ResNet50 as an example, the dual-data-bit sparse design of the present invention is 34.9% higher in acceleration area efficiency under the configuration of activation group q=4 and weight combination bit c=1.

[0106] The core technologies of the neural network accelerator design method based on double data bit sparsity of activation and weight proposed in the present invention mainly include:

[0107] 1. In the computing array design of the neural network accelerator, the activation data and weight data are converted into valid items through valid item detection to represent non-zero bits, and the representation method of the combined bit valid items is used to reduce the number of cycles of double data bit serial calculation.

[0108] 2. The problem of low PE utilization caused by different PE computing loads in the computing array is solved by grouping activation data and weight data.

[0109] The present invention implements a neural network accelerator with double data bit sparse activation and weight on the ZYBO-Z7FPGA development platform, and deploys the neural network to the acceleration platform through the TVM compilation framework to complete the calculation acceleration. The present invention is implemented on the FPGA platform with a scale of 16VPUx16PE, and the main convolutional layers and some pooling layers in the ResNet18, ResNet50, and VGG16 networks are deployed to the acceleration platform respectively. The implementation process mainly includes the following four steps:

[0110] Step 1: Based on the open source TVM neural network compilation platform and VTA Runtime and other software frameworks, as well as the sparse neural network accelerator architecture proposed in this invention, we conduct supplementary development, convert convolution and pooling operations into instruction streams of bit sparse accelerators, and generate and deploy them through control instructions of the ARM core system;

[0111] Step 2: After receiving the instruction stream, the accelerator reads the activation data and weight data from the DRAM shown in Figure 5 into the on-chip cache according to the operations specified by the instruction stream;

[0112] Step 3: The activation data and weight data are converted into valid item representation through the valid item generation unit, and grouped in the manner of activation grouping q=4 and weight combination bit c=2. After grouping, the data is distributed to each VPU. The PE in the VPU reads the activation data and selects the corresponding weight data according to the channel information of the data to complete the shift accumulation operation. A weight data and different activation channel data are respectively shifted and accumulated, and the partial sums of each channel are accumulated to complete the convolution and pooling operations, and the calculation results are saved in the on-chip cache;

[0113] Step 4: After the calculation operation is completed, the calculation results in the output on-chip cache are written back to DRAM.

[0114] Through the above steps, the effect of accelerating the calculation by utilizing double data bit sparsity in neural network calculation can be achieved. In this implementation case, compared with the VTA design with the same computing power, ResNet18 achieved a 3.87 times acceleration ratio, ResNet50 achieved a 4.75 times acceleration ratio, and VGG16 achieved a 2.57 times acceleration ratio.

[0115] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.

[0116] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A neural network accelerator implementation system based on adaptive allocation, characterized in that: include: Module M1: constructs the overall architecture of the activation and weight dual-data-bit accelerator, including DRAM and data loading modules, write-back modules, on-chip cache modules, valid item generation units and computing arrays, as well as the connection relationship between the modules; Module M2: construct an expression of activation data and weight data valid items, and construct an activation data and weight data valid item generation unit and a shift accumulation operation unit according to the expression; Module M3: Determine the data flow organization method in the computing array, perform data grouping and synchronization, and construct the weight data combination bit valid item expression method; The activation data vector and weight data vector are converted into valid items and grouped and stored in the valid item data pool of each VPU. The PE takes the activation data from the valid item data pool in the group and selects the weight data according to the channel information of the activation data to complete the shift accumulation operation. The calculation result is accumulated to the output through the addition tree. Valid entries for activation data represent: A i =(-1) s ×v×2 e Among them, A i is the value of a valid item, i is the sequence number; s is the symbol field; v is the value field; e is the position field; Valid entries for weight data are: The activation data A0 consists of n valid items, and the weight data W0 consists of l valid items. The bit-serial calculation requires n×l clock cycles to complete, and the expression is: By combining the expressions of bit valid items, the number of valid items of weight data is reduced to half, and the number of clock cycles required for bit serial calculation is also reduced to half accordingly. The expression is: Where N is the total number of activation data items; L is the total number of weight data items; m is the number of weight vectors of the data, m = l / 2; Two-bit serial computation data stream: When the valid term of the activation data is represented by a single bit, and the valid term of the weight is designed with a combined bit valid term, the calculation process is expressed as: Among them, the v field of the weight data is used as the operand of the shift operation, the sum of the e fields of the activation data and the weight data is used as the shift factor, the shift calculation result of each cycle is stored in the accumulation unit, and finally the output is obtained. When the weight data is represented in the form of a single-bit valid item, the operand of the shift operation becomes 1.

2. The neural network accelerator implementation system based on adaptive allocation according to claim 1, characterized in that: The data interaction between the on-chip cache module and DRAM is completed through LOAD and STORE instructions. The activation data and weight data are converted from the on-chip cache to the valid item expression form that can be recognized by the computing array through the valid item generation unit; The computing array includes a VPU vector computing unit, and the activation data vector and the weight data vector are deployed to the VPU. The VPU includes a basic computing unit PE, and the PE obtains valid items of activation data and weight data from the data pool of the VPU to complete the shift accumulation calculation.

3. The neural network accelerator implementation system based on adaptive allocation according to claim 1, characterized in that: Each bit of the activation data is represented by the sign field s, the value field v and the position field e. Non-negative weight data is represented in the same format. For negative weights, the negative number is represented in the form of absolute value plus the sign bit. Any bit combination of the weight data is modified to the combined expression of the v field and the e field. The valid item detection unit only outputs non-zero bits in the data to the calculation array. The multiplication and addition operations of the activation data and the weight data are completed through shift-accumulation operations. When the weight data valid items are represented by single bits, the e fields of the weight data and the activation data valid items are added as shift factors, and 1 is shifted to obtain the calculation result. When the weight data valid items are represented by combined bit valid items, the weight data valid items are the operands and the activation data valid items are the shift factors. The shift-accumulation operations of m activation data valid items and n weight data valid items are completed in a serial manner, consuming m*n cycles.

4. A method for implementing a neural network accelerator based on adaptive allocation, characterized in that: include: Step 1: Construct the overall architecture of the activation and weight dual data bit accelerator, including DRAM and data loading modules, write-back modules, on-chip cache modules, valid item generation units and computing arrays, as well as the connection relationship between each module; Step 2: construct an expression of activation data and weight data valid items, and construct an activation data and weight data valid item generation unit and a shift accumulation operation unit according to the expression; Step 3: Determine the data flow organization method in the calculation array, perform data grouping and synchronization, and construct the weight data combination bit valid item expression method; The activation data vector and weight data vector are converted into valid items and grouped and stored in the valid item data pool of each VPU. The PE takes the activation data from the valid item data pool in the group and selects the weight data according to the channel information of the activation data to complete the shift accumulation operation. The calculation result is accumulated to the output through the addition tree. Valid entries for activation data represent: A i =(-1) s ×v×2 e Among them, A i is the value of a valid item, i is the sequence number; s is the symbol field; v is the value field; e is the position field; Valid entries for weight data are: The activation data A0 consists of n valid items, and the weight data W0 consists of l valid items. The bit-serial calculation requires n×l clock cycles to complete, and the expression is: By combining the expressions of bit valid items, the number of valid items of weight data is reduced to half, and the number of clock cycles required for bit serial calculation is also reduced to half accordingly. The expression is: Where N is the total number of activation data items; L is the total number of weight data items; m is the number of weight vectors of the data, m = l / 2; Two-bit serial computing data stream: When the valid term of the activation data is represented by a single bit, and the valid term of the weight is designed with a combined bit valid term, the calculation process is expressed as: Among them, the v field of the weight data is used as the operand of the shift operation, the sum of the e fields of the activation data and the weight data is used as the shift factor, the shift calculation result of each cycle is stored in the accumulation unit, and finally the output is obtained. When the weight data is represented in the form of a single-bit valid item, the operand of the shift operation becomes 1.

5. The method for implementing a neural network accelerator based on adaptive allocation according to claim 4, characterized in that: The data interaction between the on-chip cache module and DRAM is completed through LOAD and STORE instructions. The activation data and weight data are converted from the on-chip cache to the valid item expression form that can be recognized by the computing array through the valid item generation unit; The computing array includes a VPU vector computing unit, and the activation data vector and the weight data vector are deployed to the VPU. The VPU includes a basic computing unit PE, and the PE obtains valid items of activation data and weight data from the data pool of the VPU to complete the shift accumulation calculation.

6. The method for implementing a neural network accelerator based on adaptive allocation according to claim 4, characterized in that: Each bit of the activation data is represented by the sign field s, the value field v and the position field e. Non-negative weight data is represented in the same format. For negative weights, the negative number is represented in the form of absolute value plus the sign bit. Any bit combination of the weight data is modified to the combined expression of the v field and the e field. The valid item detection unit only outputs non-zero bits in the data to the calculation array. The multiplication and addition operations of the activation data and the weight data are completed through shift-accumulation operations. When the weight data valid items are represented by single bits, the e fields of the weight data and the activation data valid items are added as shift factors, and 1 is shifted to obtain the calculation result. When the weight data valid items are represented by combined bit valid items, the weight data valid items are the operands and the activation data valid items are the shift factors. The shift-accumulation operations of m activation data valid items and n weight data valid items are completed in a serial manner, consuming m*n cycles.

Citation Information

Patent Citations

  • Convolutional neural network accelerator based on FPGA parallelism degree self-adaption

    CN113191493A

  • A convolutional neural network accelerator based on FPGA parallelism adaptation

    CN113191493B