An Efficient Multiplication and Accumulation Unit for CNN Based on Logarithmic Data

By converting the multiplication operations in the neural network into addition calculations in the logarithmic domain, the multiplier occupies a large amount of chip area and generates high energy consumption, and the hardware design is simplified and the computing rate is improved.

CN118819464BActive Publication Date: 2025-06-20UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410979053.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2025-06-20
Estimated Expiration
2044-07-22

AI Technical Summary

Technical Problem

Existing neural networks in computer vision and speech recognition have large chip area and high energy consumption due to the large use of multipliers.

Method used

A CNN efficient multiplication and addition unit based on logarithmic data is designed, and multiplication operation is realized by converting the neuron activation value and weight of the linear domain into a logarithmic domain and adding it using an adder.

Benefits of technology

It reduces the complexity, power consumption and area of ​​hardware design, increases the maximum clock frequency of the system, and achieves an increase in computing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118819464B_ABST
    Figure CN118819464B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of integrated circuits and relates to an efficient CNN multiply-accumulate unit based on logarithmic data. The present invention includes a linear domain to logarithmic domain conversion module, a logarithmic domain addition module, and an accumulation module. In the efficient CNN multiply-accumulate unit based on logarithmic data, the multiplication operation of multiplying the weight W and the neuron activation value X is converted into an addition calculation in the logarithmic domain. First, the neuron activation value X and the weight W in the linear domain are converted into the logarithmic domain, and then the adder is used to add the already logarithmized neuron activation value X and the weight W. Finally, the output data of the logarithmic domain addition module is shifted and accumulated to obtain the result of multiplying and accumulating the neuron activation value X and the weight W in the linear domain. The present invention can reduce the complexity of hardware design, hardware power consumption and area, as well as the maximum combinational logic delay, thereby increasing the maximum clock frequency of the system. It can process multiple groups of data in a pipeline form to achieve an improvement in the operation rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of integrated circuits, and relates to an efficient CNN multiplier-accumulator unit based on logarithmic data, which is applied to digital logic circuits for convolution operations, multiplication operations, and neural network accelerators. Background Art

[0002] Deep neural networks and deep convolutional neural networks have shown excellent performance in computer vision and speech recognition. However, implementing the current state-of-the-art deep neural networks, especially convolutional neural networks, requires a large number of multipliers, which not only occupy a large amount of chip area but also consume a huge amount of power.

[0003] To solve the above problems, the prior art can convert the multiplication between weights and neuron activation values into an exclusive OR (XOR) operation by binarizing the weights and neuron activation values to +1 and -1. By reducing the magnitudes of the weights and activation values, the number of operations can be reduced to meet the reduced requirements for computing and bandwidth. However, these methods may lead to a reduction in computational accuracy or require the introduction of a fine-tuning step to reduce errors. Summary of the Invention

[0004] In view of the above problems or deficiencies, to solve the problems of large chip area and high energy consumption caused by the extensive use of multipliers in existing neural networks for computer vision and speech recognition, the present invention provides an efficient CNN multiplier-accumulator unit based on logarithmic data, which is a hardware-friendly multiplier-accumulator (MAC) unit with a high-precision logarithmic domain representation. It replaces multiplication with addition and can also simply implement power operations.

[0005] An efficient CNN multiplier-accumulator unit based on logarithmic data, as Figure 1 shown, includes three modules: a linear domain to logarithmic domain conversion module, a logarithmic domain addition module, and an accumulation module.

[0006] The linear domain to logarithmic domain conversion module (Module 1) contains 2 linear domain to logarithmic domain conversion units, which receive signed fixed-point numbers with a binary bit width of L bits in the linear domain. The highest bit is the sign bit, the next L1 bits are the integer bits, and the last L2 bits are the fractional bits. The method of converting the neuron activation value X and the weight W in the linear domain into the logarithmic domain through 2 linear domain to logarithmic domain conversion units simplifies the multiplication-accumulation operation.

[0007] For the multiplication-accumulation algorithm with n neuron activation values X and n weights W, as shown in Equation 1:

[0008] Y = X1 * W1 + X2 * W2 + … + X n *W n (1)

[0009] The simplification process is as follows. For an integer m ∈ [1, n], let:

[0010] Y m = X m * W m (2)

[0011] There is:

[0012] log2 Y m = log2X m + log2 W m (3)

[0013] sign(Y m ) = sign(X m ) xor sign(W m ) (4)

[0014] sign(Y m ) represents the sign bit of Y m , at this time:

[0015] Y = Y1 + Y2 + … + Y n (5)

[0016] The working process of the linear domain to logarithmic domain conversion unit for the neuron activation value X m , as shown in formula 6:

[0017]

[0018] where i xm represents the integer part of the logarithmic domain data log 2X m, obtains the position of the highest bit 1 of the integer part of the input fixed-point number, and the number of bits in its integer part is denoted as i xm , if the bit after the highest bit 1 in the integer part is 1, shift i xm left by 1 bit, if it is 0, no adjustment is made. Shift the input fixed-point number X m left by L1 - i xm - 1 bits to get p xm .

[0019] Approximate lnp xm using formula 6, and the approximated [(p xm - 1) + q xm *log2 e represents the fractional part of the logarithmic domain data. Since the approximation has a large error at the boundary of the value range of p xm , the absolute value of the absolute error can reach more than 0.1 without correction. Therefore, q xm is introduced to be used as the correction for p xm .

[0020] The specific correction is: judge p xmThe size.

[0021] If p xm The first 4 bits of which are equal to ’b1011, take (p xm -1) + q xm Equal to p xm The value obtained by shifting p - 1 one bit to the right plus p xm The value obtained by shifting p - 1 two bits to the right plus p xm The value obtained by shifting p - 1 four bits to the right; if p xm The first 4 bits of which are equal to ’b1010 or ’b1001, take (p xm -1) + q xm Equal to p xm The value obtained by shifting p - 1 one bit to the right plus p xm The value obtained by shifting p - 1 two bits to the right plus p xm The value obtained by shifting p - 1 three bits to the right; if p xm The first 4 bits of which are equal to ’b0110, take q xm Equal to p xm The value of p - 1 plus p xm The value obtained by shifting p - 1 three bits to the right. Formula 7 represents the correction process:

[0022]

[0023] Perform correction on p xm Using formula 7 to obtain (p xm -1) + q xm . Store i xm Into the integer part of the fixed-point number, and store (p xm -1) + q xm After removing the integer part digits into the fractional part, retaining the sign bit of the linear domain data, to obtain X m The data converted from the linear domain to the logarithmic domain.

[0024] The working process of another linear domain to logarithmic domain conversion unit for the weight W m , as shown in formula 8:

[0025]

[0026] Where i wm Represents the integer part of the logarithmic domain data log2W m , obtain the position of the highest bit 1 in the integer part of the input fixed-point number, and record its number of bits in the integer part as i wm , if the bit after the highest bit 1 in the integer part is 1, shift i wm One bit to the left, if it is 0, no adjustment is made. Shift the input fixed-point number W m Left shift L1 - i wm -1 bits to obtain p wm .

[0027] Use formula 6 to approximate lnp wm to obtain an approximation of [(p wm -1)+q wm *log2 e, which represents the fractional part of the data in the logarithmic domain. Since the approximation has a large error at the boundary of the p wm value range, the absolute value of the absolute error can reach more than 0.1 without correction. Therefore, q wm is introduced to be used as a correction for p wm .

[0028] The specific correction is as follows: Determine the magnitude of p wm .

[0029] If the first 4 bits of p wm are equal to ’b1011, take (p wm -1)+q wm to be equal to the value of p wm -1 shifted right by 1 bit plus the value of p wm -1 shifted right by 2 bits plus the value of p wm -1 shifted right by 4 bits; if the first 4 bits of p wm are equal to ’b1010 or ’b1001, take (p wm -1)+q wm to be equal to the value of p wm -1 shifted right by 1 bit plus the value of p wm -1 shifted right by 2 bits plus the value of p wm -1 shifted right by 3 bits; if the first 4 bits of p wm are equal to ’b0110, take q wm to be equal to the value of p wm -1 plus the value of p wm -1 shifted right by 3 bits. Formula 9 represents the correction process:

[0030]

[0031] Perform correction on p wm using formula 9 to obtain (p wm -1)+q wm . Store i wm in the integer part of the fixed-point number, and store (p wm -1)+q wm in the fractional part after removing the integer part bits, while retaining the sign bit of the linear domain data, to obtain the data W m converted from the linear domain to the logarithmic domain.

[0032] At the same time, send the results output from the 2 linear domain to logarithmic domain conversion units to the logarithmic domain addition module, and perform the same operations on other X and W as those on X m and W m to obtain the data of all X and W converted from the linear domain to the logarithmic domain.

[0033] The two input ports of the logarithmic domain addition module (Module 2) are respectively connected to the outputs after the neuron activation value X and the weight W are converted to the logarithmic domain through the linear domain to logarithmic domain conversion module. Taking X m and W m as an example, according to Formula 3, the adder is used to add the already logarithmized neuron activation value X m and the weight W m as shown in Formulas 10 and 11:

[0034] j m = i xm + i wm (10)

[0035] f m = [(p xm - 1) + q xm + (p wm - 1) + q wm * log2e (11)

[0036] where j m represents the integer part after the logarithmic domain addition operation, that is, the sum of the integer parts i m of X m and W xm and i wm . f m represents the fractional part after the logarithmic domain addition operation, that is, the sum of the fractional parts (p m of X m and W xm - 1) + q xm and (p wm - 1) + q wm .

[0037] The obtained data is converted back to the linear domain according to Formula 12:

[0038]

[0039]

[0040] Using Formula 8 to approximate , the approximated (f m + g m ) ln 2 represents the fractional part after the logarithmic domain addition operation. Since the error is relatively large at the boundary of the value range of f m , the absolute value of the absolute error can reach more than 0.2 without correction. Therefore, g m is introduced to correct f m .

[0041] The correction is specifically as follows: Determine the magnitude of f m . Since the integer part of f m is always 0, only the magnitude of the fractional part needs to be determined. If the sign bit of f m is 1 and the first 3 bits of the fractional part are ’b110, take g m to be equal to the value of -f m shifted right by three bits; if the sign bit of f m is 0 and the first 2 bits of the fractional part are ’b01, take g m to be equal to the value of f m shifted right by 3 bits; if the sign bit of f m is 0 and the first 4 bits of the fractional part are ’b1010 or the first 3 bits are ’b100, take g m to be equal to the value of f m shifted right by 2 bits; if the sign bit of f m is 0 and the first 4 bits of the fractional part are ’b1011 or the first 2 bits are ’b11, take g m to be equal to the value of f m shifted right by 4 bits plus the value of f m shifted right by 2 bits. Formula 13 represents the correction process:

[0042]

[0043] Set 2 k+1 different data storage units according to the integer - bit length k of the linear - range fixed - point number. Among them, 2 k data storage units iacc0, iacc1, iacc2......iacc2 k - 1 are used to store the number of occurrences of the integer parts j1, j2......j n of Y1, Y2......Y n for each value in the range from 0 to 2 k - 1. Additionally, 2 k data storage units facc0, facc1, facc2......facc2 k - 1 are used to store the fractional parts f1 + g1, f2 + g2......f n of Y1, Y2......Y k when j1, j2......j n take different values in the range from 0 to 2 n + g n ; for the same value of j m1 = j m2 = r (m1, m2 ∈ [1, n] and m1 ≠ m2, r ∈ [1, 2 k ), add f m1 + g m1 and f m2 + g m2After addition, store it in the faccr data storage unit, f m +g m The coefficient ln 2 can cancel out with log2 e in Equation 6, so there is no need to operate on the coefficient. According to Equation 5, the stored data needs to be accumulated with the previously stored data to finally obtain the sum of the integer part and the fractional part for different values.

[0044] The accumulation module (Module 3) shifts and accumulates the output data of the logarithmic domain addition module, shifts the internally stored data to the left according to the labels of the iacc and facc memories, and the shift length is the size of the label. For example, both the integer part and the fractional part stored in iaccm and faccm need to be shifted to the left by m bits. Sum the finally obtained integer part value and fractional part value to get the result Y of the multiplication and accumulation of the neuron activation value X and the weight W in the linear domain.

[0045] Furthermore, the above-mentioned CNN high-efficiency multiply-accumulate unit based on logarithmic data is applied to the digital logic circuits of convolution operations, multiplication operations, and neural network accelerators to reduce the chip area and power consumption generated.

[0046] In summary, in the CNN high-efficiency multiply-accumulate unit based on logarithmic data of the present invention, the multiplication operation of multiplying the weight W and the neuron activation value X is converted into an addition calculation in the logarithmic domain. After testing, with the increase of the input data in the present invention, the relative error rate also increases. When the input quantity is about 5000, the relative error reaches 20%, which is sufficient for most convolution layer calculations. The present invention can reduce the complexity in hardware design, reduce the hardware power consumption and area, reduce the maximum combinational logic delay, thereby increasing the maximum clock frequency of the system, and can process multiple groups of data in a pipeline form to achieve an increase in the operation rate. Brief Description of the Drawings

[0047] Figure 1 is a structural schematic block diagram of the present invention;

[0048] Figure 2 is a data error comparison diagram for converting the linear domain to the logarithmic domain in the embodiment;

[0049] Figure 3 is a data error comparison diagram for converting the logarithmic domain to the linear domain in the embodiment;

[0050] Figure 4 is a relationship diagram between the input data quantity and the error in the embodiment;

[0051] Figure 5 is a schematic diagram of the representation method of logarithmic domain fixed-point numbers in the embodiment;

[0052] Figure 6 is a schematic diagram of the principle of the linear domain to logarithmic domain conversion module of the present invention;

[0053] Figure 7 Schematic diagram of the logarithmic domain addition module for the embodiment;

[0054] Figure 8 Schematic diagram of the accumulation module for the embodiment. Detailed implementation manners

[0055] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0056] An efficient CNN multiplication and addition unit based on logarithmic data proposed in this embodiment, as Figure 1 shown, includes three modules: Module 1 includes two linear domain to logarithmic domain conversion units, Module 2 includes a logarithmic domain addition unit, and Module 3 includes an accumulation unit. It receives the neuron activation value X represented by a 16-bit fixed-point number with a bit width L and the weight value W represented by a 16-bit fixed-point number with a bit width L, where the first bit is the sign bit, the next 10 bits are the integer bits, and the last 5 bits are the fractional bits. The multiplication and addition operations are simplified through logarithmic methods.

[0057] As Figure 5 shown, this embodiment uses a 16-bit fixed-point number to represent the logarithmic domain fixed-point number. The highest bit is the sign bit of this data in the linear domain, the second highest 4 bits are the integer bits, representing the integer part of the data in the logarithmic domain, and the last 11 bits are the fractional bits, representing the fractional part of the data in the logarithmic domain. In particular, since 0 cannot be logarithmized, 0 is represented by 16'b0111100000000000 and is specially processed in the calculation. When 0 appears as an addend in the logarithmic addition, the calculation result will be set to zero. At the same time, other 16'bx1111xxxxxxxxxxx are recognized as illegal data. Therefore, the data range that can be represented by the representation method of the logarithmic domain fixed-point number is (-2 15 , +2 15 ).

[0058] As Figure 6 shown, the two linear domain to logarithmic domain conversion units included in Module 1 respectively receive 128 Xs and 128 Ws, convert them into logarithmic domain values log2X and log2W with base 2, and send them to the logarithmic domain addition unit of Module 2. Taking X1 as an example to illustrate the working process of the linear domain to logarithmic domain conversion unit, as shown in Formula 14:

[0059]

[0060] After the 16-bit fixed-point number X1 is input, it enters the comparator. After removing the highest bit sign bit, the highest bit 1 of X1 is bitwise ORed with 15'b0 to obtain I x1 , and I X1 is right-shifted by the fractional bit number 5 in the fixed-point number X1. At this time, I x1Indicates the position of the highest bit of the integer part of X1, and obtains the number of bits i of the highest bit through the decoder x1 = log2I x1 , obtains the value of the bit after the highest bit 1 of X1. If it is 1, then i x1 is shifted left by 1 bit. If it is 0, no adjustment is made. At this time, i x1 is the integer part of log2X1.

[0061] Shift X1 to the left by 10 - i x1 - 1 bits to obtain the fractional part of log2X1 According to the value of p x1 , approximate and correct ln p x1 : If the first 4 bits of p x1 are equal to 'b1011, take (p x1 - 1) + q1 equal to the value of p x1 - 1 shifted right by 1 bit plus the value of p x1 - 1 shifted right by 2 bits plus the value of p x1 - 1 shifted right by 4 bits; if the first 4 bits of p x1 are equal to 'b1010 or 'b1001, take (p x1 - 1) + q x1 equal to the value of p x1 - 1 shifted right by 1 bit plus the value of p x1 - 1 shifted right by 2 bits plus the value of p x1 - 1 shifted right by 4 bits; if the first 4 bits of p x1 are equal to 'b0110, take q x1 equal to the value of p x1 - 1 plus the value of p x1 - 1 shifted right by 3 bits. As shown in Formula 15:

[0062]

[0063] The error of this formula is as Figure 2 shown. Fitted line1 is the curve without qx1 correction, and Fitted line2 is the curve with qx1 correction. Error1 and Error2 are the curves of the absolute values of the absolute errors of Fitted line1 and Fitted line2 relative to y = log2x. Through correction, the absolute value of the absolute error within the value range can be reduced to 0.02.

[0064] Discard the coefficient log2 e and take (p x1 - 1) + q x1 as the fractional part of log2X1, and take i1 and (p x1 - 1) + q x1Concatenate according to the representation of logarithmic domain fixed-point numbers. The highest bit of the logarithmic domain fixed-point number is the sign bit of X1 in the linear domain. Assign i1 to the integer part represented by bits [14:11] of the logarithmic domain fixed-point number, and assign the [(p x1 -1)+q x1 bits [15:5] to bits [10:0] of the logarithmic domain fixed-point number to represent the fractional part. The discarded log2 e coefficient will be canceled in the subsequent module. At this time, the logarithmic domain representation of X1 is obtained.

[0065] Perform the same operation on the weight W1 as on the neuron activation value X1. At the same time, send the results output from the 2 linear domain to logarithmic domain conversion units to the logarithmic domain addition module. Perform the same operation on other X and W as on X1 to obtain the data of all X and W converted from the linear domain to the logarithmic domain.

[0066] As Figure 7 shown, module 2 is a logarithmic domain addition unit. The two inputs are X1 and W1 after logarithmic domain conversion respectively. Exclusive OR the sign bits of the two to obtain the sign bit of the result, as shown in formula 16:

[0067] sign(Y1) = sign(X1) xor sign(W1) (16)

[0068] Add the 15-bit data bits to obtain the data bits of the result. Convert the result in the logarithmic domain to the linear domain according to formula 17:

[0069]

[0070] where j1 represents the integer part after the logarithmic domain addition operation, that is, the integer part of Y1. Use formula 8 to approximate, judge the size of f1. The integer bit of f1 is always 0, so only need to judge the size of the fractional bit: if the sign bit of f1 is 1 and the first 3 bits of the fractional bit are 'b110, take g1 equal to the value of -f1 shifted right by three bits; if the sign bit of f1 is 0 and the first 2 bits of the fractional bit are 'b01, take g1 equal to the value of f1 shifted right by 3 bits; if the sign bit of f1 is 0 and the first 4 bits of the fractional bit are 'b1010 or the first 3 bits are 'b100, take g1 equal to the value of f1 shifted right by 2 bits; if the sign bit of f1 is 0 and the first 4 bits of the fractional bit are 'b1011 or the first 2 bits are 'b11, take g1 equal to the value of f1 shifted right by 4 bits plus the value of f1 shifted right by one bit. As shown in formula 18:

[0071]

[0072] The error of this formula is as Figure 3As shown, Fitted line1 is the curve without g1 correction, and Fitted line2 is the curve with g1 correction. Error1 and Error2 are the curves of the absolute values of the absolute errors of Fitted line1 and Fitted line2 relative to y = 2 x respectively. The absolute value of the absolute error within the value range can be reduced to 0.025 through correction.

[0073] The approximately obtained (f1 + g1)ln 2 represents the fractional part after the logarithmic domain addition operation. Since the log2 e coefficient is omitted during storage in Module 1, and the ln2 coefficient is omitted in this module to cancel it out, the linear domain result of multiplying X1 and W1 is

[0074] According to the magnitude of j1, sign(Y1)*1 and sign(Y1)*(f1 + g1) are sent to different accumulators respectively. There are 32 accumulators iacc0, iacc1, iacc2......iacc15 and facco, facc1, facc2......facc15 that store the accumulated values of sign*1 and sign*f corresponding to j respectively. According to the value of j1, sign(Y1)*1 is stored in iaccj1, and sign(Y1)*(f1 + g1) is stored in iaccf1. For subsequent calculations of Y2, Y3……Y 128 for j2, j3……j 128 when there is an equal j value as before during the calculation of j, the corresponding sign*1 and sign*f are accumulated with the values already stored in the register. The data in the register waits for the next accumulation or is used as output for Module 3.

[0075] Such as Figure 8 shown, Module 3 is an accumulation module used to recombine the data in different registers in Module 2 into the output result y. Controlled by a counter, the data in the integer register and the fractional register are taken out one by one through a multiplexer. When the counter counts to h, the data in iacch and facch are taken out respectively, and shifted left by h bits to get iacch*2 h and facch*2 h , and the data is sent to the accumulator to be accumulated with the data sent to the accumulator when the counter indicates 0, 1, 2……h - 1 before. The intermediate results of integer accumulation and fractional accumulation are stored in Y i and Y f registers respectively and wait for the next accumulation. When the counter indicates 15, after the last accumulation, the data in Y i is shifted right by 11 bits and added to the data in Y f to obtain the final output Y.

[0076] After testing, as Figure 4 shown ( Figure 4 which is the relationship diagram between the input data quantity and the error of the embodiment), as the input data increases, the relative error rate also increases. When the input quantity is around 5000, the relative error reaches 20%, but it is sufficient for most convolutional layer calculations.

[0077] As can be seen from the above embodiments, in the CNN efficient multiply-accumulate unit based on logarithmic data of the present invention, the multiplication operation of multiplying the weight W and the neuron activation value X is converted into an addition calculation in the logarithmic domain. First, the neuron activation value X and the weight W in the linear domain are converted into the logarithmic domain, and then the adder is used to add the logarithmized neuron activation value X and the weight W. Finally, the output data of the logarithmic domain addition module is shifted and accumulated to obtain the result of multiplying and accumulating the neuron activation value X and the weight W in the linear domain. The present invention can reduce the complexity in hardware design, reduce the hardware power consumption and area, reduce the maximum combinational logic delay, thereby increasing the maximum clock frequency of the system, and can process multiple groups of data in a pipeline form to achieve an improvement in the operation rate.

Claims

1. A CNN efficient multiplication and addition unit based on logarithmic data, characterized by: It includes three modules: linear domain to logarithmic domain conversion module, logarithmic domain addition module and accumulation module; The linear domain to logarithmic domain conversion module includes two linear domain to logarithmic domain conversion units, which receive a signed fixed-point number input with a binary bit width of L bits in the linear domain, the highest bit is the sign bit, the next L1 bits are integer bits, and the last L2 bits are decimal bits; the neuron activation value X and the weight W in the linear domain are respectively converted into the logarithmic domain through the two linear domain to logarithmic domain conversion units; The multiplication and addition algorithm for n neuron activation values ​​X and n weights W is shown in Formula 1: Y=X1*W1+X2*W2+…+X n *W n (1) The simplified process is as follows. For an integer m∈[1,n], let: Y m =X m *W m (2) have: log2 Y m = log2 X m + log2 W m (3) sign(Y m )=sign(X m )xor sign(W m ) (4) sign(Y m ) indicates Y m The sign bit of , at this time: Y= Y1+ Y2+…+Y n (5) The working process of the linear domain to logarithmic domain conversion unit is for the neuron activation value X m , as shown in Formula 6: where i xm Represents logarithmic domain data log2 X m The integer part of the input fixed-point number is obtained by obtaining the position of the highest bit 1 in the integer part of the input fixed-point number. The number of bits in the integer part is recorded as i xm , if the integer part has a 1 after the highest 1, i xm Shift left by 1 bit. If it is 0, no adjustment will be made. m Shift left L1-i xm -1 bit gets p xm ; Using formula 6, lnp xm Approximate, the approximate [(p xm -1)+q xm ]*log2 e represents the decimal part of the logarithmic domain data, and introduces q xm Used as a pair xm amendments; P xm Correction is performed to obtain (p xm -1)+q xm , change i xm Store the integer part of the fixed-point number and (p xm -1)+q xm After removing the integer part, store the decimal part and retain the sign bit of the linear domain data to get X m Convert data from linear to logarithmic domain; Another linear domain to logarithmic domain conversion unit works for the weight W m , as shown in Formula 8: where i wm Represents logarithmic domain data log2 W m The integer part of the input fixed-point number is obtained by obtaining the position of the highest bit 1 in the integer part of the input fixed-point number. The number of bits in the integer part is recorded as i wm , if the integer part has a 1 after the highest 1, i wm Shift left by 1 bit. If it is 0, no adjustment will be made. m Shift left L1-i wm -1 bit gets p wm ; Using formula 6, lnp wm Approximate, the approximate [(p wm -1)+q wm ]*log2 e represents the decimal part of the logarithmic domain data, and introduces q wm Used as a pair wm amendments; P wm Correction is performed to obtain (p wm -1)+q wm , change i wm Store the integer part of the fixed-point number and (p wm -1)+q wm After removing the integer part, store the decimal part and retain the sign bit of the linear domain data to get W m Convert data from linear to logarithmic domain; At the same time, the output results from the two linear domain to logarithmic domain conversion units are sent to the logarithmic domain addition module, and the other X and W are ANDed with X m and W m The same operation is performed to obtain data of all X and W converted from the linear domain to the logarithmic domain; The two input ports of the logarithmic domain addition module are respectively connected to the output of the neuron activation value X and the weight W after being converted to the logarithmic domain by the linear domain to logarithmic domain conversion module, and the logarithmic neuron activation value X is added by an adder. m and weight W m Add them together, as shown in formulas 10 and 11: j m =i xm +i wm (10) f m =[(p xm -1)+q xm +(p wm -1)+q wm ]*log2 e (11) where j m represents the integer part after the logarithmic domain addition operation, that is, X m and W m The integer part i xm and i wm Add and; f m Represents the decimal part after the logarithmic domain addition operation, that is, X m and W m The fractional part (p xm -1)+q xm and (p wm -1)+q wm Add and; The obtained data is converted back to the linear domain according to Equation 12: Using formula 8 for e fm*ln2 Approximate, the approximate (f m +g m )ln2 represents the decimal part after the logarithmic domain addition operation, and introduces g m Use as f m amendments; Set 2 according to the integer bit length k of the linear domain fixed-point number k+1 Different data storage units, 2 of which k data storage units iacc0, iacc1, iacc2, ..., iacc2 k -1 is used to store Y1, Y2, ... Y n The integer part j1, j2, ... j n Between 0 and 2 k -1 is the number of values, and the other 2 k facc0, facc1, facc2, ... facc2 k -1 is used to store j1, j2, ... j n 0 to 2 k -1 when Y1, Y2, ... Y n The decimal part of f1+g1, f2+g2...f n +g n For the same value of j m1 =j m2 = r, f m1 +g m1 、f m2 +g m2 After addition, the data is stored in the faccr data storage unit, where m1,m2∈[1,n] and m1≠m2,r∈[1,2 k ]; and f m +g m The coefficient ln2 can cancel each other with log2 e in formula 6, so there is no need to operate the coefficient; according to formula 5, the stored data needs to be accumulated with the previously stored data to finally obtain the sum of the integer part and the decimal part when j is different values; The accumulation module shifts and accumulates the output data of the logarithmic domain addition module, and left-shifts the internally stored data according to the labels of the iacc and facc memories, and the length of the shift is the size of the label; then the integer part value and the decimal part value finally obtained are summed to obtain the linear domain neuron activation value X and the weight W multiplied and accumulated result Y.

2. The CNN efficient multiplication and addition unit based on logarithmic data as claimed in claim 1, characterized in that: The introduction of q xm Used as a pair xm The specific correction is: judge p xm size; If p xm The first 4 bits are equal to 'b1011, take (p xm -1)+q xm Equal to p xm -1 right shift 1 bit value plus p xm -1 right shift 2 bits plus p xm -1 right shift 4 bits; if p xm The first 4 bits are equal to 'b1010 or 'b1001, take (p xm -1)+q xm Equal to p xm -1 right shift 1 bit value plus p xm -1 right shift 2 bits plus p xm -1 right shift 3 bits; if p xm The first 4 bits are equal to 'b0110, take q xm Equal to p xm The value of -1 plus p xm -1 is the value shifted 3 bits right; Formula 7 shows the correction process: P xm Formula 7 is used to correct (p xm -1)+q xm .

3. The CNN efficient multiplication and addition unit based on logarithmic data as claimed in claim 1, characterized in that: The introduction of q wm Used as a pair wm The specific correction is: judge p wm size; If p wm The first 4 bits are equal to 'b1011, take (p wm -1)+q wm Equal to p wm -1 right shift 1 bit value plus p wm -1 right shift 2 bits plus p wm -1 right shift 4 bits; if p wm The first 4 bits are equal to 'b1010 or 'b1001, take (p wm -1)+q wm Equal to p wm -1 right shift 1 bit value plus p wm -1 right shift 2 bits plus p wm -1 right shift 3 bits; if p wm The first 4 bits are equal to 'b0110, take q wm Equal to p wm The value of -1 plus p wm -1 is the value shifted 3 bits right; Formula 9 shows the correction process: P wm Formula 9 is used to correct (p wm -1)+q wm , change i wm Store the integer part of the fixed-point number and (p wm -1)+q wm After removing the integer part, store the decimal part and retain the sign bit of the linear domain data to get W m Convert data from linear to logarithmic domain.

4. The CNN efficient multiplication and addition unit based on logarithmic data as claimed in claim 1, characterized in that: The introduction of m Use as f m The specific correction is: judge f m The size of f m The integer digit is always 0, so we only need to determine the size of the decimal digit; If f m The sign bit is 1 and the first 3 decimal places are 'b110, take g m Equals -f m Shift the value right by three bits; if f m The sign bit is 0 and the first two decimal places are 'b01, take g m Equal to f m The value shifted right by 3 bits; if f m The sign bit is 0 and the first 4 decimal places are 'b1010 or the first 3 decimal places are 'b100, select g m Equal to f m The value shifted right by 2 bits; if f m The sign bit is 0 and the first 4 decimal places are 'b1011 or the first 2 decimal places are 'b11, select g m Equal to f m Shift right 4 bits and add f m The value shifted right by 2 bits; Formula 13 shows the correction process:

5. The CNN efficient multiplication-addition unit based on logarithmic data according to any one of claims 1 to 4, characterized in that: Digital logic circuits used for convolution operations, multiplication operations, and neural network accelerators to reduce chip area and power consumption.

Citation Information

Patent Citations

  • Multiplying Hardware Circuit, System-on-Chip and Electronic device

    CN109521994A

  • Fine-grained precision-adjustable Multiplier-Accumulator

    KR102037043B1