Neural network multiply-accumulate operation device and neural network multiply-accumulate operation method

By using original code encoding and bit-by-bit calculation, and taking advantage of the data distribution characteristics of neural network, invalid calculations are terminated early, solving the problems of computing power and energy waste and high data preprocessing costs in neural network accelerators, and achieving efficient multiplication and accumulation operations.

CN120596061APending Publication Date: 2025-09-05PEKING UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510468146.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing neural network accelerators waste computing power and energy in multiplication and accumulation operations, and have high data preprocessing costs, especially when using complement calculations, which results in high power consumption. The bitmap sparsity scheme cannot effectively utilize the sparsity of output data, resulting in a waste of computing resources.

Method used

The original code encoding rules are used to perform product operations to generate positive and negative product pairs, and bit-by-bit calculations are performed through adder trees and judges. The data distribution characteristics of neural networks are used to terminate invalid calculations in advance, and the sparsity of input and output is combined to optimize the calculation process.

Benefits of technology

It reduces the amount of data processing in neural network calculations, reduces invalid calculations, improves computing power and energy efficiency, and maintains the accuracy of calculation results, especially significantly reducing the scale of calculations in neural networks with ReLU activation functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596061A_ABST
    Figure CN120596061A_ABST
Patent Text Reader

Abstract

The invention provides a neural network multiply-accumulate operation device and a neural network multiply-accumulate operation method, and the method comprises the steps: carrying out the product operation of input data and a pre-distributed corresponding weight based on a primitive code coding rule through a preprocessing module, generating a positive and negative product pair, and transmitting the positive and negative product pair to a positive and negative product temporary storage region; the adder tree sequentially calculates the numerical value of one bit position of each pair of positive and negative product pairs according to the order from the high order to the low order of the positive and negative product pairs, and accumulates the numerical values; the operation module performs subtraction operation on the positive and negative dynamic sums, and accumulates an intermediate value calculated by a current bit and an intermediate value calculated by a previous bit; the determiner determines whether to terminate the subsequent calculation of the current bit or continue the accumulation of the next bit in advance; the neuron calculation module executes corresponding activation function operation or neuron dynamics updating operation according to the type of the neural network, an operation mode of bit-by-bit calculation from a high bit to a low bit is used, the data processing amount is reduced, and the trend of a calculation result can be predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a neural network multiplication-accumulation operation device and a neural network multiplication-accumulation operation method. Background Art

[0002] The core computing units of existing neural network accelerators primarily rely on multiply-accumulate (MAC) operations, with implementation options falling into two categories: direct implementation and bitmap sparsity. Direct implementations, using a "multiplier + adder" combination, achieve high computational power but require significant hardware resources, resulting in significant increases in area and power consumption, and ineffective computations. For example, when the multiplication-accumulation result is negative, it is directly corrected to 0, but all computational steps must still be performed, resulting in redundant computational power and energy consumption. Existing neural network accelerators typically use two's complement encoding. While widely used in modern computers, this encoding scheme suffers from a high flip rate and high power consumption. Furthermore, the sign bit of the two's complement encoding has a certain weight, which is negative. If the bitwise calculation is performed when the input and weight are multiplied, the result will be mixed, with individual bits in the result fluctuating and unable to discern a clear trend.

[0003] Bitmap sparsity constructs a sparse bitmap for the input and weight tensors of each layer. In the constructed bitmap, "1" is used to indicate the presence of non-zero data at each corresponding position in the tensor, and "0" is used to indicate that the data at that location is zero. Since the product of a multiplication operation is 0 when one of the two operands is 0, which has no effect on the final accumulated result, only inputs and weights whose corresponding positions are non-zero need to be sent to the multiplication and addition unit for calculation, thereby skipping operations where the product is destined to be 0 (including data reading operations), thereby saving computational costs. However, its optimization scope is limited to the sparsity of the input data, and non-contributing multiplication and accumulation operations in the hidden layer are still fully executed. At the same time, the construction of the sparse bitmap requires additional preprocessing, and the preprocessing cost increases sharply when the number of input channels or matrix dimensions are high. Summary of the Invention

[0004] The present invention provides a neural network multiplication-accumulation operation device and a neural network multiplication-accumulation operation method, which are used to solve the defects of traditional neural network multiplication-accumulation operation devices, such as waste of computing power and energy consumption, and high data preprocessing cost.

[0005] The present invention provides a neural network multiplication and accumulation operation device, comprising: A preprocessing module is configured to receive input data, perform a product operation on the input data and the corresponding pre-assigned weights based on the original code encoding rules to generate positive and negative product pairs, and label the positive and negative product pairs with positive and negative labels and send them to the positive and negative product temporary storage area; The positive and negative product temporary storage area is used to store the positive and negative product pairs; an adder tree connected to the positive and negative product temporary storage area, for calculating the value of a bit position of each positive and negative product pair in order from high to low bits of the positive and negative product pairs, and accumulating the values ​​of the corresponding bit positions of all positive and negative product pairs to obtain a positive and negative dynamic sum; an operation module, configured to perform a subtraction operation on the positive and negative dynamic sums to obtain an intermediate value, and accumulate the intermediate value calculated for the current bit position with the intermediate value calculated for the previous bit position; a determiner for receiving the accumulated intermediate value in real time, comparing the accumulated intermediate value with a preset positive or negative threshold, and determining whether to terminate the subsequent calculation of the current bit in advance or continue the accumulation of the next bit based on the comparison result; The neuron computing module is used to receive the final multiplication and accumulation result or the zero value of the early termination output, and perform the corresponding activation function operation or neuron dynamics update operation according to the neural network type.

[0006] According to the neural network multiplication-accumulation operation device provided by the present invention, the preprocessing module is further used to screen and delete invalid product pairs with zero amplitude, and send the retained valid product pairs to the corresponding product temporary storage area.

[0007] According to the neural network multiplication and accumulation operation device provided by the present invention, the preprocessing module is also used to skip the current time step calculation and directly enter the next time step calculation when the processing neural network type is a pulse neural network, if the current time step is 0 input.

[0008] According to the neural network multiplication and accumulation operation device provided by the present invention, the positive and negative product temporary storage area includes a positive product temporary storage area and a negative product temporary storage area, and the adder tree includes a positive adder tree connected to the positive product temporary storage area and a negative adder tree connected to the negative product temporary storage area: When the calculated neural network type does not use the ReLU activation function, the positive adder tree and the negative adder tree directly communicate with the neuron circuit module, and send the complete accumulated calculation results output by the positive adder tree and the negative adder tree to the neuron circuit module.

[0009] According to the neural network multiplication-accumulation operation device provided by the present invention, the determiner is further used for: If the calculated neural network type uses ReLU as the activation function, then if the dynamic part sum output by the negative adder tree is less than the negative threshold, the calculation process is terminated early and 0 is directly output, skipping the subsequent bit multiplication and accumulation calculation; If the sum of the positive and negative dynamic parts output by the positive and negative adder trees is between the positive and negative thresholds, the next bit calculation is performed until the positive and negative dynamic parts are greater than the positive threshold or less than the negative threshold; If the sum of the positive dynamic parts output by the positive adder tree is greater than the positive threshold, the determination is stopped and the accumulated result calculated for the current bit is directly output.

[0010] According to the neural network multiplication-accumulation operation device provided by the present invention, the determiner is further used for: If the calculated neural network type does not use ReLU as the activation function, when the sum of the positive dynamic parts output by the positive adder tree is greater than the positive threshold, the multiplication and accumulation calculation result is directly output.

[0011] According to the neural network multiplication and accumulation operation device provided by the present invention, the neuron calculation module includes: The neuron unit is used to perform neuron dynamics operations according to the accumulated results of this time step when the calculated neural network type is a spiking neural network; The data processing unit is used to send the accumulated results from the current calculation unit to the high-level accumulator or specific activation function for subsequent processing when the calculated neural network type is a convolutional neural network or a Transformer model.

[0012] The neural network multiplication and accumulation operation device provided by the present invention further includes: An address processor is used to dynamically adjust the storage address of output data according to the type of neural network being calculated.

[0013] The neural network multiplication and accumulation operation device provided by the present invention further includes: The pooling calculator is configured to receive the data to be processed in a pooling mode and perform a pooling operation on the data to be processed.

[0014] The present invention also provides a neural network multiplication and accumulation operation method, comprising: Receive input data, perform a product operation on the input data and the pre-assigned corresponding weight based on the original code encoding rule to generate a positive and negative product pair, mark the positive and negative product pairs with positive and negative labels, and send them to the positive and negative product temporary storage area; Calculate the value of a bit position of each positive and negative product pair in order from high to low, and accumulate the values ​​of the corresponding bit positions of all positive and negative product pairs to obtain a positive and negative dynamic sum; Subtracting the positive and negative dynamic sums to obtain an intermediate value, and adding the intermediate value calculated for the current bit position to the intermediate value calculated for the previous bit position; Comparing the accumulated intermediate value with a preset positive or negative threshold, and determining whether to terminate the subsequent calculation of the current bit in advance or continue the accumulation of the next bit according to the comparison result; Receive the final multiplication-accumulation result or the zero value of the early termination output, and perform the corresponding activation function operation or neuron dynamics update operation according to the neural network type.

[0015] The present invention provides a neural network multiplication-accumulation operation device and a neural network multiplication-accumulation operation method. The neural network multiplication-accumulation operation device includes a preprocessing module for receiving input data, performing a product operation on the input data and a pre-assigned corresponding weight based on the original code encoding rule to generate a positive-negative product pair, marking the positive-negative product pair with a positive-negative label and sending it to a positive-negative product temporary storage area; the positive-negative product temporary storage area is used to store the positive-negative product pairs; an adder tree is connected to the positive-negative product temporary storage area, and is used to calculate the value of a bit position of each positive-negative product pair in order from high to low bits of the positive-negative product pairs, and accumulate the values ​​of the corresponding bit positions of all positive-negative product pairs to obtain a positive-negative dynamic sum; an operation module is used to perform a subtraction operation on the positive-negative dynamic sum to obtain an intermediate value. And the intermediate value calculated by the current bit is accumulated with the intermediate value calculated by the previous bit; the judge is used to receive the accumulated intermediate value in real time, compare the accumulated intermediate value with the preset positive and negative thresholds, and judge whether to terminate the subsequent calculation of the current bit in advance or continue the accumulation of the next bit according to the comparison result; the neuron calculation module is used to receive the final multiplication and accumulation result or the zero value of the early termination output, and perform the corresponding activation function operation or neuron dynamics update operation according to the neural network type. The present invention uses an operation method of bit-by-bit calculation from high to low. According to the distribution characteristics of neural network data, each bit value that has a greater impact on the final calculation result of the current layer is considered in turn, thereby reducing the data processing amount, and at the same time providing monitoring and prediction of the trend of the calculation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 This is a functional structure diagram of a neural network multiplication-accumulation operation device provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of the multiplication and accumulation operation principle of the prior art; Figure 3This is an example diagram of a neural network multiplication and accumulation calculation based on original code provided by an embodiment of the present invention; Figure 4 This is a diagram showing the principle of multiplication and accumulation calculation of a neural network based on original code provided by an embodiment of the present invention; Figure 5 This is one of the flow charts of the neural network multiplication and accumulation method provided by an embodiment of the present invention; Figure 6 This is the second flowchart of the neural network multiplication and accumulation operation method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0019] Figure 1 The functional module diagram of the neural network multiplication and accumulation operation device provided by the embodiment of the present invention is as follows: Figure 1 As shown, the neural network multiplication and accumulation operation device provided by the embodiment of the present invention includes: A preprocessing module is configured to receive input data, perform a product operation on the input data and the corresponding pre-assigned weights based on the original code encoding rules to generate positive and negative product pairs, and label the positive and negative product pairs with positive and negative labels and send them to the positive and negative product temporary storage area; The positive and negative product temporary storage area is used to store the positive and negative product pairs; an adder tree connected to the positive and negative product temporary storage area, for calculating the value of a bit position of each positive and negative product pair in order from high to low bits of the positive and negative product pairs, and accumulating the values ​​of the corresponding bit positions of all positive and negative product pairs to obtain a positive and negative dynamic sum; an operation module, configured to perform a subtraction operation on the positive and negative dynamic sums to obtain an intermediate value, and accumulate the intermediate value calculated for the current bit position with the intermediate value calculated for the previous bit position; a determiner for receiving the accumulated intermediate value in real time, comparing the accumulated intermediate value with a preset positive or negative threshold, and determining whether to terminate the subsequent calculation of the current bit in advance or continue the accumulation of the next bit based on the comparison result; The neuron computing module is used to receive the final multiplication and accumulation result or the zero value of the early termination output, and perform the corresponding activation function operation or neuron dynamics update operation according to the neural network type.

[0020] In an embodiment of the present invention, the neural network multiplication and accumulation operation device further includes: An address processor is used to dynamically adjust the storage address of output data according to the type of neural network being calculated.

[0021] The pooling calculator is configured to receive the data to be processed in a pooling mode and perform a pooling operation on the data to be processed.

[0022] In this embodiment of the present invention, the address processor processes the output addresses according to model requirements. For example, due to the transposition operation in the Transformer, the x and y coordinates of the address of a specific result need to be swapped. The pooling calculator can specifically perform four-input pooling data calculations, including maximum pooling and average pooling. This improves the data processing efficiency of the neural network multiplication and accumulation operation device.

[0023] Traditional multiplication and accumulation calculations are the most important execution process during the operation of neural network models. Existing neural network computing accelerators usually use direct implementation or bitmap sparse methods for it. Direct implementation uses multipliers to perform multiplication operations and adders to perform product accumulation operations. Therefore, a large number of multiplication and addition units are required in the accelerator to meet the high computing power requirements in the actual use of neural networks. Bitmap sparsity is to construct a sparse bitmap for the input and weight tensors of each layer. In the constructed bitmap, "1" is used to indicate that there is non-zero data in each corresponding position of the tensor, and "0" is used to indicate that the data is zero. Since the product will be 0 when one of the two operands of the multiplication operation is 0, it has no effect on the final accumulation result. Therefore, only the inputs and weights whose corresponding positions are non-zero need to be sent to the multiplication and addition unit for calculation, so that the operation where the product is destined to be 0 can be skipped, thereby saving computing costs. However, this approach has the following problems and shortcomings: (1) Existing acceleration schemes can only optimize input data and reduce the processing of 0 input and 0 weight, but cannot observe and utilize the hidden characteristics of output data, such as Figure 2 (1) As shown in the figure. The ReLU activation function is widely used in neural network calculations. Data with negative multiplication and accumulation results will be directly corrected to 0, resulting in a large number of multiplication and accumulation operations to obtain the original results, which consumes time and energy resources but does not achieve significant results and does not affect the final output feature map.

[0024] (2) Existing neural network accelerators usually use complement calculation. Although this encoding method is widely used in modern computers, studies have shown that it has the disadvantage of a high flip rate, such as Figure 2As shown in (2), in the calculation of neural network accelerators, a high flip rate means that the capacitance of circuit nodes needs to be repeatedly charged and discharged, which will lead to higher power consumption. On the other hand, the sign bit of the complement code also has a certain weight, and it is a negative weight. If the input and weight are multiplied by bits, the results of each bit in the result will be positive and negative, and the accumulated result of the product will fluctuate, making it impossible to observe a clear trend.

[0025] (3) Large-scale input data in neural network calculations leads to heavy data processing tasks. When the number of input channels of the convolutional layer increases ( Figure 2 (1) The larger N is), or the higher the dimension of the vector / matrix is, the more products need to be accumulated after the multiplication operation at the corresponding points, that is, the data of each output point is affected by the accumulation range of a longer distance and contains more input information. Such large-scale and high-precision data accumulation not only brings high time consumption, but also puts great pressure on the prediction algorithms and hardware implementation mentioned above, because the prediction needs to be combined with all the input data and weight distribution characteristics of the current output in order to make judgments and optimizations. The longer the accumulation distance, the larger the range that needs to be considered. Moreover, in the bitmap sparse scheme, in order to construct the indicator tensor composed of "1" and "0", it is necessary to traverse the input and weight data in advance ( Figure 2 (H × W × N in (1), which brings more computing and processing tasks, whether using software or hardware.

[0026] The embodiment of the present invention adopts the original code scheme to encode the input, weight and output data, which not only has a lower data flipping rate during calculation, but its more regular weight distribution is also more suitable for characteristic analysis and judgment of the output data of each calculation; it can not only optimize the input characteristics, but also study the multiplication and accumulation operation process from the output perspective, utilize the implicit properties therein, and promptly interrupt certain calculation processes that can predict the results in advance. For example, for negative outputs adjusted to 0 by the activation function, it is necessary to consider how to judge the positive and negative characteristics of the final result so as to terminate the calculation in advance), bringing considerable computing power and energy efficiency benefits.

[0027] The neural network multiplication and accumulation operation device provided by the embodiment of the present invention includes a preprocessing module for receiving input data, performing a product operation on the input data and the pre-assigned corresponding weights based on the original code encoding rules to generate positive and negative product pairs, marking the positive and negative product pairs with positive and negative labels and sending them to the positive and negative product temporary storage area; the positive and negative product temporary storage area is used to store the positive and negative product pairs; an adder tree is connected to the positive and negative product temporary storage area, and is used to calculate the value of a bit position of each pair of positive and negative product pairs in order from high to low bits of the positive and negative product pairs, and accumulate the values ​​of the corresponding bit positions of all positive and negative product pairs to obtain a positive and negative dynamic sum; an operation module is used to perform a subtraction operation on the positive and negative dynamic sum to obtain an intermediate value, and calculate the current bit position. The intermediate value is accumulated with the intermediate value calculated by the previous bit; the determiner is used to receive the accumulated intermediate value in real time, compare the accumulated intermediate value with the preset positive and negative thresholds, and judge whether to terminate the subsequent calculation of the current bit in advance or continue the accumulation of the next bit according to the comparison result; the neuron calculation module is used to receive the final multiplication and accumulation result or the zero value of the early termination output, and perform the corresponding activation function operation or neuron dynamics update operation according to the neural network type. The present invention uses an operation method of bit-by-bit calculation from high to low, and according to the data distribution characteristics of the neural network, considers each bit value that has a greater impact on the final calculation result of the current layer in turn, thereby reducing the data processing amount, and at the same time provides monitoring and prediction of the trend of the calculation results.

[0028] Based on any of the above embodiments, the embodiments of the present invention use original code encoding rules, which are applicable to the calculation process of various neural network models and support 1-8 bit input and 1-8 bit weight. The original code-based encoding method can significantly reduce the frequency of 0 / 1 data switching during the calculation process; this algorithm constructed based on the distribution properties of neural network data can determine the calculation result at an early stage of the multiplication and accumulation calculation at a relatively low computational cost, and terminate nearly half of the calculation process early, thereby reducing useless multiplication and accumulation calculations. Specifically: Each bit of the original code and the complement code except the sign bit represents a weight related to the bit order, but the sign bit of the complement code also has a weight (- ), which will result in Figure 2 In the accumulation calculation, some bits are positive and some are negative, and the sign bit of the original code only plays a role in determining the sign bit of the product result. (1): in, is the input data, is the weight, is the symbol of the input data, is the sign of the weight, according to the sign bit and Generate positive and negative product pairs, M is the total number of weight bits, N is the total number of input data bits, is the index of the weight value bit, The index of the input data bit value.

[0030] like Figure 3 As shown in Equation (1), the multiplication of a pair of original code inputs and original code weights can be decomposed into a process of multiplying (N+M-2) amplitude bits and weighted accumulation. If the carry is temporarily not considered during the calculation process, it can be regarded as the multiplication of every two amplitude bits, as shown in Figure 4 As shown in (a), results with the same bit weight are summed (i.e., vertical addition in the vertical form without carrying; results in the same column have the same bit weight). During neural network calculations, each output point is obtained by multiplying and accumulating many such "input-weight" pairs. Therefore, another way to calculate the multiplication-accumulation result is to accumulate the bits with the same bit weight in these product pairs to obtain the final result. The advantage of this method is that each multiplication calculation only requires a single bit, and the accumulated data is much smaller than the product data obtained by directly using multi-bit multiplication.

[0031] On the other hand, in neural network models, the weight distribution usually follows a normal distribution, that is, the data near 0 is the most abundant, and the data with larger absolute values ​​is less. However, input data, because it is usually regulated by the ReLU activation function, has no negative values ​​and is mostly distributed near 0, and the larger the value, the less data is distributed. In each "input-weight" product, although the high bit values ​​are relatively small, their contribution to the final result is exponentially greater than the privileged value. Therefore, these two characteristics can be combined. Figure 4 As shown in (b), taking 1152 pairs of 8-bit input and 8-bit weight as an example, to multiply and accumulate their products, the maximum privilege value without considering the carry is , which can only come from the multiplication of the highest bit of the input and the weight (6+6), totaling 1 (second row); Next, the highest 2 bits from the input and weight are cross-multiplied (6+5 and 5+6), for a total of 2, and so on. Set it to 1 (i.e. 100%, the third row), then The percentage is 1 / 2, and so on. By multiplying these data (the fourth row), we can get the contribution of each bit product to the final cumulative result. It can be seen that although the number of high-bit data is very small, their extremely high privilege value makes them have a more important influence on the cumulative result. In fact, only the first few bits of the product need to be calculated to determine the final positive or negative of most multiplication and accumulation results. Figure 4Taking (b) as an example, if the calculation is terminated after calculating a 4-bit product, the computational effort is reduced to 20.4% of the original amount; if the calculation is terminated after calculating a 6-bit product, the computational effort is reduced to 42.9%. Experiments have shown that calculating approximately 6 bits is sufficient to determine the positive or negative value of the result while maintaining the original model accuracy, and even in this case, the computational effort can be reduced by at least 40%. Considering the possibility that small values ​​may be so numerous that their impact is even greater, two thresholds, one positive and one negative, can be set. If the accumulated result exceeds the positive threshold, the result is likely positive; otherwise, it is likely negative. The granularity of setting the positive and negative thresholds can be coarse at the network layer level or more detailed at the neuron level (i.e., each calculation has a different threshold). For neural network models using the ReLU activation function, since all negative values ​​are corrected to 0, this approach effectively discards a large number of final zero outputs by simply calculating a few bits, significantly reducing the computational effort required during the neural network calculation process.

[0032] Figure 4 (c) shows the bit accumulation of 1152 pairs of 1-bit input pulses and 8-bit weights in a spiking neural network. Since the inputs in a spiking neural network are all 1-bit pulses, the impact of the input pulses need not be considered; only the 8-bit weights need to be studied. As can be seen from the table, due to the reduced amount of information carried by the input data, the accumulated results due to the weights exhibit a greater disparity between different product bits. This phenomenon facilitates prediction of the positive or negative nature of the resulting multiplication and accumulation results. The only difference is that in a spiking neural network, a positive value indicates a pulse was emitted, while a negative value indicates no pulse was emitted.

[0033] In an embodiment of the present invention, the pre-processing module is further configured to screen and delete invalid product pairs with zero amplitude, and send the remaining valid product pairs to the corresponding product temporary storage area.

[0034] Also, when the neural network type is a pulse neural network, if the current time step has 0 input, the current time step calculation is skipped and the next time step calculation is directly entered.

[0035] Sparsity is an important property in the calculation process of neural network models, and this property can be used to speed up calculations. However, due to the different ways in which data flows, existing accelerators basically process the sparsity of weights, and a small amount of sparsity is used for input data. These practices are insufficient for utilizing the sparsity of network models. The embodiment of the present invention conducts research from the perspective of output data sparsity, and proposes an energy-saving inference algorithm for neural network multiplication and accumulation calculations based on the original code, which reduces the frequency of digital switching in the calculation and simultaneously utilizes the sparsity characteristics of three aspects. Combined with the data distribution characteristics of the neural network, it successfully compresses the computational complexity of the neural network multiplication and accumulation calculations, and reduces useless calculations while maintaining a high degree of inference accuracy.

[0036] Based on any of the above embodiments, the positive and negative product temporary storage area includes a positive product temporary storage area and a negative product temporary storage area, and the adder tree includes a positive adder tree connected to the positive product temporary storage area and a negative adder tree connected to the negative product temporary storage area: When the calculated neural network type does not use the ReLU activation function, the positive adder tree and the negative adder tree directly communicate with the neuron circuit module, and send the complete accumulated calculation results output by the positive adder tree and the negative adder tree to the neuron circuit module.

[0037] Based on any of the above embodiments, the determiner is further configured to: If the calculated neural network type uses ReLU as the activation function, then if the dynamic part sum output by the negative adder tree is less than the negative threshold, the calculation process is terminated early and 0 is directly output, skipping the subsequent bit multiplication and accumulation calculation; If the sum of the positive and negative dynamic parts output by the positive and negative adder trees is between the positive and negative thresholds, the next bit calculation is performed until the positive and negative dynamic parts are greater than the positive threshold or less than the negative threshold; If the sum of the positive dynamic parts output by the positive adder tree is greater than the positive threshold, the determination is stopped and the accumulated result calculated for the current bit is directly output.

[0038] In an embodiment of the present invention, the determiner is further configured to: If the calculated neural network type does not use ReLU as the activation function, when the sum of the positive dynamic parts output by the positive adder tree is greater than the positive threshold, the multiplication and accumulation calculation result is directly output.

[0039] In an embodiment of the present invention: for the multiplication and accumulation process of each output data calculation, two thresholds, one positive and one negative, are pre-configured before the calculation, and the calculated intermediate value S is set to 0; during the calculation, the value of a bit position of each pair of product results is calculated in sequence from high to low order of the input and weight product results (no carry operation is performed), and the corresponding values ​​of all product results are accumulated, and the obtained partial sum is accumulated to the intermediate value S; during judgment, if the network uses ReLU as the activation function and S is less than the negative threshold, the calculation process is terminated and 0 is directly output; if the network uses ReLU as the activation function and S is between the positive and negative thresholds, the next bit of the result is calculated according to the above operation; if the network uses ReLU as the activation function and S is greater than the positive threshold or the network does not use ReLU as the activation function, the judgment is stopped, and each bit after the result is directly calculated according to the above operation to obtain the output result.

[0040] Based on any of the above embodiments, the neuron computing module includes: A neuron unit is used to perform neuron dynamics operations according to the accumulated results of this time step when the calculated neural network type is a spiking neural network; The data processing unit is used to send the accumulated results from the current calculation unit to the high-level accumulator or specific activation function for subsequent processing when the calculated neural network type is a convolutional neural network or a Transformer model.

[0041] Relatively few existing neural network accelerators support multiple computing modes simultaneously, including convolutional neural networks, spiking neural networks, and transformers. By modeling the accumulation processes of different models and extracting commonalities, the present invention supports hybrid intelligence of convolutional neural networks, spiking neural networks, and transformers. After the data is input, the preprocessing module first uses positive and negative labels to separate the positive and negative product pairs, and sends them to the positive and negative product temporary storage areas, filtering out products that are 0. If it is a spiking neural network, it also determines whether the current time step is a 0 input based on similar time step labels. If so, it directly skips to the next time step calculation. A circuit in the temporary storage area calculates the cumulative value of each pair of input-weights at that bit based on the bit sequence of the currently calculated product, and sends it to the positive and negative adder tree for accumulation. The positive and negative results are then subtracted and added to the previously calculated partial sum to obtain the partial sum of the product after accumulation at the current bit. The result is then fed into a discriminator, which has pre-set positive and negative thresholds. The resulting partial sum is compared with these two thresholds and, based on the algorithm's requirements, determines whether the computation should be terminated. If so, a 0 is output directly and the next subsequent module is entered. Otherwise, the next product bit is accumulated and the above process is repeated. Furthermore, if the model being calculated does not use the ReLU activation function, the full accumulation process is performed directly to obtain the result, without the need for a discriminator. During calculation, this unit can support the multiplication of up to 144 pairs of 8-bit inputs and 8-bit weights and the accumulation of these products, effectively computing a 3×3 convolution kernel with up to 16 channels. If computation of a higher number of convolution kernels is required, more such computation units can be added. Using a top-level control module, the convolution kernels are grouped into groups of 16 channels each. These are fed into different computation units, and the results of each round are then accumulated at a high level to obtain the result for the entire convolution kernel. Appropriate subsequent operations are then performed, forming a complete accelerator architecture that utilizes the aforementioned multiplication-accumulation inference algorithm.

[0042] After obtaining the multiplication and accumulation results of the "input-weight" combination, the neural computation module enters. For spiking neural network models, different neural dynamics operations are performed based on the accumulated results of this time step, combined with previous data such as residual membrane levels, neuron thresholds, and leakage patterns. This supports a variety of neuron models, including the accumulation-and-fire (IF) model, the leakage-accumulation-and-fire (LIF) model, and the degenerate Izhikevich model. If the computation of all time steps is not completed (up to 8 time steps are supported), the residual membrane levels are temporarily stored for future use. The positive and negative thresholds in the discriminator are also updated based on the model requirements and the current residual membrane levels for subsequent use. For convolutional neural network or Transformer models, the accumulated results are directly sent out of the current computation unit to a higher-level accumulator or specific activation function for further processing.

[0043] This specialized hardware design primarily targets multiplication-accumulation processes and offers excellent configurability, supporting not only ReLU activation functions but also Sigmoid, tanh, and other activation functions, as well as a variety of spiking neural network neurons, enabling it to support a wide range of neural network models. Building on basic computational operations, it leverages the data sparsity of inputs, weights, and outputs to reduce unnecessary multiplication-accumulation operations at the fine-grained bit level, achieving higher computing power and energy efficiency.

[0044] Comparison of test results of similar neural network multiplication and accumulation calculation inference hardware solutions, as shown in Table 1, shows that the technology proposed in the embodiment of the present invention achieves higher computing power and energy efficiency, successfully reduces the scale of multiplication and accumulation operations of the neural network accelerator, and maintains a high model accuracy.

[0045] Table 1 Comparison of different design schemes

[0046] It should be noted that the neural network multiplication-accumulation device provided in the embodiments of the present invention is applicable to both ASICs and FPGAs. ASIC design and implementation have been performed using a variety of tools and have passed comprehensive testing, ensuring the feasibility and high performance of the algorithm and hardware design. The embodiments of the present invention can achieve a neuron scale of 35K and a synapse scale of 4.5M, achieving a peak energy efficiency of 61.19 TOPs / W at a frequency of 500 MHz.

[0047] The neural network multiplication and accumulation device provided by the embodiment of the present invention has a preprocessing module that uses positive and negative labels to separate positive and negative product pairs and determines whether the current time step is a zero input for the spiking neural network. The positive and negative temporary storage areas and corresponding adder trees calculate the partial sums of bits to be accumulated based on the bit sequence of the current calculation result and send them to the adder tree to calculate the current partial sum. The determiner determines whether to terminate the calculation process based on pre-stored positive and negative thresholds. The spiking neural network neuron module performs neuron dynamics operations based on the accumulation result of the current time step. The convolutional neural network / Transformer data processing module sends the current calculation unit to a higher-level accumulator or a specific activation function for subsequent processing based on the current accumulation result. The address processor sets the processing output address for the model. The pooling calculator is specifically used for pooling data calculations. This hardware design can effectively implement the neural network multiplication and accumulation method, achieving high computing power and energy efficiency benefits.

[0048] The neural network multiplication and accumulation operation method provided by the present invention is described below. The neural network multiplication and accumulation operation method described below and the neural network multiplication and accumulation operation device described above can be referenced to each other.

[0049] Figure 5 The flowchart of the neural network multiplication and accumulation method provided by the embodiment of the present invention is as follows: Figure 5 As shown, the neural network multiplication and accumulation operation method provided by the embodiment of the present invention includes: Step 501: Receive input data, perform a product operation on the input data and the corresponding pre-assigned weight based on the original code encoding rule to generate positive and negative product pairs, and label the positive and negative product pairs with positive and negative labels and send them to the positive and negative product temporary storage area; Step 502: Calculate the value of a bit position of each positive and negative product pair in order from high to low, and accumulate the values ​​of the corresponding bit positions of all positive and negative product pairs to obtain a positive and negative dynamic sum; Step 503: Subtract the positive and negative dynamic sums to obtain an intermediate value, and add the intermediate value calculated for the current bit to the intermediate value calculated for the previous bit; Step 504: Compare the accumulated intermediate value with a preset positive or negative threshold, and determine whether to terminate the subsequent calculation of the current bit in advance or continue accumulation of the next bit based on the comparison result; Step 505: Receive the final multiplication-accumulation result or the zero value of the early termination output, and perform the corresponding activation function operation or neuron dynamics update operation according to the neural network type.

[0050] The embodiments of the present invention can reduce the number of multiplication and accumulation operations required to obtain an output data, thereby improving the time and energy efficiency of neural network calculations. Figure 6As shown, upon startup, the algorithm first initializes, then reads in the inputs, weights, and corresponding neuron data (positive and negative thresholds, spiking neural network dynamics parameters, etc.) to be calculated. To conserve hardware power and minimize the occurrence of zero products, each input-weight pair is digitally labeled 0 / 1 to indicate the location of positive and negative products. Because of the use of native encoding, this labeling is achieved by simply XORing the sign bit of each input and weight pair. A value of 1 indicates dissimilarity, resulting in a negative product, while a value of 0 indicates identity, resulting in a positive product. However, if at least one of the seven bits of a pair of inputs and weights is 0, the corresponding label bit is marked as "invalid." This allows for the separation of positive and negative products during calculations, eliminating invalid data that does not contribute to the final accumulation process. The computational model, whether convolutional neural network or spiking neural network, then proceeds through different steps. The Transformer, primarily performing matrix multiplication, can use the same data flow as fully connected layers and is therefore classified as a convolutional neural network. Similarly, the Spiking-Transformer model is classified as a spiking neural network. If it is a convolutional neural network model, taking 8-bit input multiplied by 8-bit weight as an example, the result weight is first calculated as The value of bit 13 comes from the input and the bit weight in the weight is (7th bit) and weight all the results of this bit in this set of "input-weight" pairs ( ) are added to get the partial sum, and the weight of the result is calculated as (12th bit), and similar process is continued until the result weight is (10th bit) and accumulate the partial sums one by one.

[0051] If the neural network uses ReLU as the activation function, it is compared with the positive and negative thresholds respectively: if the partial sum is less than the negative threshold, the calculation process is terminated and 0 is directly output, thus skipping the multiplication and accumulation calculation of subsequent bits; if the partial sum is between the positive and negative thresholds, the above steps are repeated, the next bit of the calculation result is calculated, and the comparison is repeated until the partial sum is greater than the positive threshold or less than the negative threshold; if the partial sum is greater than the positive threshold, the judgment is stopped, and each bit after the result is directly calculated according to the above operation and accumulated to obtain the output result. If the neural network does not use ReLU as the activation function, it is treated as if the partial sum is greater than the positive threshold, and the final result is calculated directly. If it is a spiking neural network, the algorithm calculation process is similar to the above, applicable to both IF and LIF neurons. The Izhikevich model only performs accumulation and does not use positive and negative thresholds for judgment. After obtaining the final result, the spiking neural network needs to perform neuron dynamics operations of the spiking neural network and update related parameters. Each time the final result is obtained, it is necessary to observe which time step it is. If it is not the final time step, the residual membrane level of this time step should be subtracted from the positive and negative thresholds respectively. Because this inference algorithm is aimed at the multiplication and accumulation process, the spiking neural network neuron calculation needs to add the accumulated result to the residual membrane level of the previous time step before it can be compared with the neuron threshold to decide whether to emit a pulse. Therefore, the positive and negative thresholds need to subtract the residual membrane potential.

[0052] The neural network multiplication and accumulation operation method provided by the embodiment of the present invention was applied to the recognition of the CIFAR-10 dataset for testing, using the Resnet-18 model. The results show that the skip rate of calculation operations in different layers of the model varies, reaching up to 63.9%, while the maximum calculation error rate is only 2.8%, and both are small values, and the impact on the recognition results is almost negligible. The skip rate and error rate of calculation operations are both related to the set positive and negative thresholds. The larger the threshold, the more difficult it is to reach, the lower the error rate, but also the lower the skip rate. The test results on the entire dataset show that the image recognition accuracy is 94.85%, which is only 0.65% lower than the software result (95.5%), proving the excellent effect of the proposed algorithm in maintaining accuracy while minimizing the calculation scale.

[0053] The neural network multiplication and accumulation operation method provided by the embodiment of the present invention uses positive and negative labels to separate positive and negative product pairs, and determines whether the current time step is a 0 input for the pulse neural network; the positive and negative temporary storage areas and the corresponding adder trees, respectively, calculate the partial sum of the bits to be accumulated according to the bit sequence of the current calculation result and send them to the adder tree to calculate the current partial sum; according to the pre-stored positive and negative thresholds, it is judged whether the calculation process needs to be terminated; according to the accumulation result of this time step, the neuron dynamics operation is executed or according to the accumulation result of this time, the current calculation unit is sent to a higher-level accumulator or a specific activation function for subsequent processing; and the output address for model setting is designed and specifically used for pooled data calculation, which can well implement the neural network multiplication and accumulation operation method and obtain higher computing power and energy efficiency benefits.

[0054] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0055] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A neural network multiplication and accumulation operation device, characterized in that: include: A preprocessing module is configured to receive input data, perform a product operation on the input data and the corresponding pre-assigned weights based on the original code encoding rules to generate positive and negative product pairs, and label the positive and negative product pairs with positive and negative labels and send them to the positive and negative product temporary storage area; The positive and negative product temporary storage area is used to store the positive and negative product pairs; an adder tree connected to the positive and negative product temporary storage area, for calculating the value of a bit position of each positive and negative product pair in order from high to low bits of the positive and negative product pairs, and accumulating the values ​​of the corresponding bit positions of all positive and negative product pairs to obtain a positive and negative dynamic sum; an operation module, configured to perform a subtraction operation on the positive and negative dynamic sums to obtain an intermediate value, and accumulate the intermediate value calculated for the current bit position with the intermediate value calculated for the previous bit position; a determiner for receiving the accumulated intermediate value in real time, comparing the accumulated intermediate value with a preset positive or negative threshold, and determining whether to terminate the subsequent calculation of the current bit in advance or continue the accumulation of the next bit based on the comparison result; The neuron computing module is used to receive the final multiplication and accumulation result or the zero value of the early termination output, and perform the corresponding activation function operation or neuron dynamics update operation according to the neural network type.

2. The neural network multiplication-accumulation operation device according to claim 1, characterized in that: The pre-processing module is further used to screen and delete invalid product pairs with zero amplitude, and send the retained valid product pairs to the corresponding product temporary storage area.

3. The neural network multiplication-accumulation operation device according to claim 1, characterized in that: The preprocessing module is also used to skip the current time step calculation and directly enter the next time step calculation when the processing neural network type is a pulse neural network if the current time step is 0 input.

4. The neural network multiplication-accumulation operation device according to claim 1, characterized in that: The positive and negative product temporary storage area includes a positive product temporary storage area and a negative product temporary storage area, and the adder tree includes a positive adder tree connected to the positive product temporary storage area and a negative adder tree connected to the negative product temporary storage area: When the calculated neural network type does not use the ReLU activation function, the positive adder tree and the negative adder tree directly communicate with the neuron circuit module, and send the complete accumulated calculation results output by the positive adder tree and the negative adder tree to the neuron circuit module.

5. The neural network multiplication-accumulation operation device according to claim 4, characterized in that: The determiner is further configured to: If the calculated neural network type uses ReLU as the activation function, then if the dynamic part sum output by the negative adder tree is less than the negative threshold, the calculation process is terminated early and 0 is directly output, skipping the subsequent bit multiplication and accumulation calculation; If the sum of the positive and negative dynamic parts output by the positive and negative adder trees is between the positive and negative thresholds, the next bit calculation is performed until the positive and negative dynamic parts are greater than the positive threshold or less than the negative threshold; If the sum of the positive dynamic parts output by the positive adder tree is greater than the positive threshold, the determination is stopped and the accumulated result calculated for the current bit is directly output.

6. The neural network multiplication-accumulation operation device according to claim 4 or 5, characterized in that: The determiner is further configured to: If the calculated neural network type does not use ReLU as the activation function, when the sum of the positive dynamic parts output by the positive adder tree is greater than the positive threshold, the multiplication and accumulation calculation result is directly output.

7. The neural network multiplication-accumulation operation device according to claim 1, characterized in that: The neuron computing module includes: The neuron unit is used to perform neuron dynamics operations according to the accumulated results of this time step when the calculated neural network type is a spiking neural network; The data processing unit is used to send the accumulated results from the current calculation unit to the high-level accumulator or specific activation function for subsequent processing when the calculated neural network type is a convolutional neural network or a Transformer model.

8. The neural network multiplication-accumulation operation device according to claim 1, characterized in that: Also includes: An address processor is used to dynamically adjust the storage address of output data according to the type of neural network being calculated.

9. The neural network multiplication-accumulation operation device according to claim 5, characterized in that: Also includes: The pooling calculator is configured to receive the data to be processed in a pooling mode and perform a pooling operation on the data to be processed.

10. A neural network multiplication and accumulation method, characterized in that: include: Receive input data, perform a product operation on the input data and the pre-assigned corresponding weight based on the original code encoding rule to generate a positive and negative product pair, mark the positive and negative product pairs with positive and negative labels, and send them to the positive and negative product temporary storage area; Calculate the value of a bit position of each positive and negative product pair in order from high to low, and accumulate the values ​​of the corresponding bit positions of all positive and negative product pairs to obtain a positive and negative dynamic sum; Subtracting the positive and negative dynamic sums to obtain an intermediate value, and adding the intermediate value calculated for the current bit position to the intermediate value calculated for the previous bit position; Comparing the accumulated intermediate value with a preset positive or negative threshold, and determining whether to terminate the subsequent calculation of the current bit in advance or continue the accumulation of the next bit according to the comparison result; Receive the final multiplication-accumulation result or the zero value of the early termination output, and perform the corresponding activation function operation or neuron dynamics update operation according to the neural network type.

Citation Information

Cited By

  • Storage and calculation integrated device, storage and calculation method, processing device, tile module and accelerator

    CN120874919A

  • Transform model accelerator for ocean drifting buoy

    CN121683911A