A Precision Configurable Multiply-Accumulate Unit for Neural Network Accelerators

By designing a multiplication and accumulation unit with configurable accuracy, a multi-selectable 2bit multiplication unit, bit splicing and approximate adder is used to solve the calculation-intensive problem of neural network accelerator in the quantization model, reducing power consumption and improving hardware performance.

CN116108904BActive Publication Date: 2025-07-04SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310221906.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-09
Publication Date
2025-07-04
Estimated Expiration
2043-03-09

AI Technical Summary

Technical Problem

The existing neural network accelerators fail to effectively utilize precision configurability in the quantitative model, resulting in unsolved computing-intensive problems, increasing bandwidth and logical resource requirements, high power consumption and low hardware utilization.

Method used

Design a multiplication and accumulation unit with configurable accuracy, adopting multiple selection 2bit multiplication unit, bit splicing, shift unit and approximate adder to simplify the calculation unit logic, introduce approximate calculation to adapt to convolutional neural network layers with different bit widths, and reduce power consumption through bit-level flexibility.

Benefits of technology

It realizes that while ensuring computing accuracy, power consumption and computing complexity are reduced, accelerator flexibility and hardware performance are improved, and parameters of different quantization methods are adapted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116108904B_ABST
    Figure CN116108904B_ABST
Patent Text Reader

Abstract

The present invention discloses a precision-configurable multiply-accumulate unit designed for a neural network accelerator. First, the zero-th computing unit and the first computing unit are summed through a fifth summation module, then the second computing unit and the third computing unit are summed through a sixth summation module, and finally the parts generated by the above two summations are finally summed through a third summation module. All the summation modules in the precision-configurable multiply-accumulate unit are correspondingly turned off according to different input bit widths to adapt to the multiply-accumulate operations of each layer of convolutional neural networks with different bit widths of 2 bits, 4 bits, and 8 bits. While ensuring a certain precision of the CNN, they not only simplify the computing unit and the external complex configurable logic, but also introduce the concept of approximate computing. Approximate computing is introduced into the configurable computing unit for the first time, enabling the architecture to further reduce power consumption based on bit-level flexibility and adapt to parameters from different quantization methods of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology of a precision-configurable multiply-accumulate unit for a neural network accelerator, belonging to the field of configurable computing technology. Background Art

[0002] CNN (Convolutional Neural Network) has achieved great success in many computer vision tasks such as image recognition and object recognition. However, the continuously increasing network model size has led to significant demands for memory size, bandwidth, and computing resources. Many model compression methods have been proposed, such as pruning and quantization, to reduce the storage and computing requirements of CNN.

[0003] The quantization method can significantly reduce the size of the CNN model, alleviate the memory-intensive problem, and thus help reduce the bandwidth requirement. However, most current accelerators fail to utilize the quantized model to solve the compute-intensive problem. Most accelerators perform MAC (Multiply Accumulate) operations with fixed high precision, but many quantized MAC operations do not require such high precision. It is difficult for quantization technology to improve the throughput rate and power efficiency of the precisely fixed accelerator. Different applications have different requirements for the accelerator in various aspects, and the precisely fixed accelerator lacks the flexibility to meet these requirements.

[0004] Therefore, recently, a precision-configurable CNN accelerator has been proposed, in which the precision of the activation and the weights can be partially or fully scaled. However, with the improvement of performance, the part with lower precision requires complex configurable logic, and the reduction of precision also requires more activations and weights to perform the calculation of the precision-configurable unit. The increase in the requirements for activations and weights leads to an increase in the requirements for bandwidth and logic resources, resulting in higher power consumption and lower hardware utilization. Summary of the Invention

[0005] Technical Problem: Aiming at the above problems, the present invention discloses a precision-configurable multiply-accumulate unit for a neural network accelerator, which not only simplifies the computing unit and the external complex configurable logic, but also introduces the concept of approximate computing while ensuring a certain precision of the CNN. The approximate computing is introduced into the configurable computing unit for the first time, enabling the architecture to further reduce power consumption based on bit-level flexibility and adapt to the parameters from different quantization methods of the network.

[0006] Technical solution: An accuracy-configurable multiply-accumulate unit for a neural network accelerator according to the present invention includes four computing units, namely a zero-th computing unit, a first computing unit, a second computing unit, and a third computing unit, and a third summing module, a fifth summing module, and a sixth summing module. First, the zero-th computing unit and the first computing unit are summed through the fifth summing module, then the second computing unit and the third computing unit are summed through the sixth summing module, and finally the parts generated by the above two summations are finally summed through the third summing module. Among them, all the summing modules in the accuracy-configurable multiply-accumulate unit are turned off accordingly according to the different input bit widths to adapt to the multiply-accumulate operations of each layer of the convolutional neural network with different bit widths of 2 bits, 4 bits, and 8 bits.

[0007] The four units, namely the zero-th computing unit, the first computing unit, the second computing unit, and the third computing unit, are respectively minimum computing units. Each minimum computing unit includes a 2-bit multiplication unit based on multiplexing, bit concatenation, a shift unit, and a summing module. Among them,

[0008] The bit concatenation includes a first bit concatenation, a second bit concatenation, a third bit concatenation, a fourth bit concatenation, a fifth bit concatenation, and a sixth bit concatenation.

[0009] The shift unit includes a first shift unit, a second shift unit, a third shift unit, a fourth shift unit, a fifth shift unit, a sixth shift unit, a seventh shift unit, an eighth shift unit, and a ninth shift unit.

[0010] The summing module includes a first summing module, a second summing module, a third summing module, a fourth summing module, a fifth summing module, a sixth summing module, a seventh summing module, an eighth summing module, and a ninth summing module.

[0011] In the zero-th computing unit, the 2-bit multiplication unit based on multiplexing is composed of 4-way multiplication units as inputs. Among them, the outputs of the first-way multiplication unit and the second-way multiplication unit are respectively connected to the inputs of the first bit concatenation, the outputs of the third-way multiplication unit and the fourth-way multiplication unit are respectively connected to the inputs of the second bit concatenation, the output of the first bit concatenation is connected to the first summing module, the output of the second bit concatenation is connected to the first shift unit, the output of the first shift unit is connected to the first summing module, and the output of the first summing module is connected to the fifth summing module for summation.

[0012] In the first computing unit described above, the 2-bit multiplication unit based on multiplexing is composed of four multiplication units as inputs. Among them, the outputs of the first multiplication unit and the second multiplication unit are respectively connected to the inputs of the fifth-bit concatenation. The outputs of the third multiplication unit and the fourth multiplication unit are respectively connected to the inputs of the sixth-bit concatenation. The output of the fifth-bit concatenation is connected to the second summation module. The output of the sixth-bit concatenation is connected to the third shift unit. The output of the third shift unit is connected to the second summation module. The output of the second summation module is connected to the second shift unit. The output of the second shift unit is connected to the fifth summation module for summation.

[0013] In the second computing unit described above, the 2-bit multiplication unit based on multiplexing is composed of four multiplication units as inputs. Among them, the outputs of the first multiplication unit and the second multiplication unit are respectively connected to the inputs of the third-bit concatenation. The outputs of the third multiplication unit and the fourth multiplication unit are respectively connected to the inputs of the fourth-bit concatenation. The output of the third-bit concatenation is connected to the fourth summation module. The output of the fourth-bit concatenation is connected to the fourth shift unit. The output of the fourth shift unit is connected to the fourth summation module. The output of the fourth summation module is connected to the fifth shift unit. The output of the fifth shift unit is connected to the sixth summation module for summation.

[0014] In the third computing unit described above, the 2-bit multiplication unit based on multiplexing is composed of four multiplication units as inputs. Among them, the output of the first multiplication unit is connected to the seventh summation module. The output of the second multiplication unit is connected to the seventh shift unit. The output of the seventh shift unit is connected to the seventh summation module. The output of the seventh summation module is connected to the ninth summation module. The output of the third multiplication unit is connected to the eighth summation module. The output of the fourth multiplication unit is connected to the sixth shift unit. The output of the sixth shift unit is connected to the eighth summation module. The output of the eighth summation module is connected to the eighth shift unit. The output of the eighth shift unit is connected to the ninth summation module. The output of the ninth summation module is connected to the ninth shift unit. The output of the ninth shift unit is connected to the sixth summation module for summation.

[0015] Among the shift units described above, the first shift unit, the second shift unit, the third shift unit, the fourth shift unit, the sixth shift unit, the seventh shift unit, and the eighth shift unit are two-bit shifts. The fifth shift unit is a four-bit shift, and the ninth shift unit is a six-bit shift.

[0016] The bit concatenation directly combines the high two bits and the low two bits to generate the partial product, replacing the adder, effectively avoiding the accumulation operation. The bit concatenation stitches together the products of two of the 2-bit multiplication units based on multiplexing. These two products do not affect each other in the final product, and the bits they occupy do not overlap. The partial sum of the high bits is directly placed in the high 4 bits, and the partial sum of the low bits is directly placed in the low 4 bits.

[0017] The described shift unit is composed of 2-bit, 4-bit, and 6-bit left shifters. By using shifters to replace the function of adders, the purpose of reducing the number of adders or reducing the bit width of adders is achieved. Ultimately, less circuit area is used to improve the circuit performance.

[0018] The described summation module splits the adder into two non-overlapping calculation modules according to the high bits and low bits. The split high-bit calculation module is implemented by an exact adder, and the split low-bit calculation module is implemented by bitwise OR of an approximate adder based on an OR gate; the summation partial result of the split low-bit part is approximate. To compensate for the error caused by the approximate calculation, an AND gate is used for the highest bit of the split low-bit part as an error compensation.

[0019] Beneficial effects: Since the present invention adopts the above technical solutions, the present invention has the following advantages:

[0020] 1. The 2-bit multiplication unit based on the multiplexer selects signed multiplication, unsigned multiplication, or signed and unsigned mixed multiplication through the multiplexer, avoiding the addition of an extra sign bit and saving the area of the minimum multiplication unit.

[0021] 2. The number of multipliers, shifters, and adders required by the method based on bit concatenation is significantly less than the number of multipliers, shifters, and adders required by the existing multiply-accumulate units, and this advantage becomes more and more obvious as the number of bits of the calculation multiplier increases. It has a smaller area and lower power consumption.

[0022] 3. Due to the fault tolerance of CNN, the use of approximate adders is allowed while ensuring accuracy. Among various approximate adders, the use of a logical OR gate for logical OR operation on the low bits, LOA has a smaller area and power consumption.

[0023] 4. For different configuration modes, the input bandwidth is 8 bits, but the corresponding output bandwidth varies greatly in different configuration modes. To relieve the huge pressure on the output bandwidth, the bandwidth in different modes is reused. In the 8×8 and 4×4 input modes, the output bandwidth is reused for the 2×2 mode, and the final total output bandwidth is 64 bits, thus reducing the area and power loss caused by the bandwidth.

[0024] From the perspective of improving the flexibility of the accelerator, the present invention designs a precision-configurable multiply-accumulate unit, which adapts to various network structures while ensuring low power consumption, and the hardware performance of the accelerator will be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is the overall structural schematic diagram of a configurable MAC unit implemented by the present invention;

[0026] Figure 2 isFigure 1 The 2BM module in [description], i.e., the structural schematic diagram of a 2-bit multiplication unit;

[0027] Figure 3 It is the structural schematic diagram of a 4-bit multiplication unit using bit concatenation implemented in the present invention;

[0028] Figure 4 It is the structural schematic diagram of an approximate LOA adder unit implemented in the present invention.

[0029] Where A and B are inputs, 2BM is a 2-bit multiplier, M is bit concatenation, "<<" is a shift unit, "+" is a summation module, and "×" is a multiplier. Specific implementation manner

[0030] The technical solution of the present invention will be further specifically described below through embodiments and in conjunction with the accompanying drawings.

[0031] As Figure 1 shown, a precision-configurable multiply-accumulate unit for a neural network accelerator according to the present invention includes four computing units, namely, a zero-th computing unit Cell0, a first computing unit Cell1, a second computing unit Cell2, and a third computing unit Cell3, and three summation modules, namely, a third summation module 4.3, a fifth summation module 4.5, and a sixth summation module 4.6; first, the zero-th computing unit Cell0 and the first computing unit Cell1 are summed through the fifth summation module 4.5, then the second computing unit Cell2 and the third computing unit Cell3 are summed through the sixth summation module 4.6, and finally, the parts generated by the above two summations are finally summed through the third summation module 4.3; among them, all summation modules in the precision-configurable multiply-accumulate unit are turned off accordingly according to different input bit widths to adapt to the multiply-accumulate operations of each layer of a convolutional neural network with different bit widths of 2 bits, 4 bits, and 8 bits.

[0032] The described zero - th computing unit Cell0, first computing unit Cell1, second computing unit Cell2, and third computing unit Cell3 are respectively the minimum computing units. Each minimum computing unit includes a 2 - bit multiplication unit 1 based on multiplexing, bit concatenation, shift units, and summation modules. Among them, the bit concatenation includes the first bit concatenation 2.1, the second bit concatenation 2.2, the third bit concatenation 2.3, the fourth bit concatenation 2.4, the fifth bit concatenation 2.5, and the sixth bit concatenation 2.6. The shift units include the first shift unit 3.1, the second shift unit 3.2, the third shift unit 3.3, the fourth shift unit 3.4, the fifth shift unit 3.5, the sixth shift unit 3.6, the seventh shift unit 3.7, the eighth shift unit 3.8, and the ninth shift unit 3.9. The summation modules include the first summation module 4.1, the second summation module 4.2, the third summation module 4.3, the fourth summation module 4.4, the fifth summation module 4.5, the sixth summation module 4.6, the seventh summation module 4.7, the eighth summation module 4.8, and the ninth summation module 4.9. In the described zero - th computing unit Cell0, the 2 - bit multiplication unit 1 based on multiplexing is composed of 4 - way multiplication units as inputs. Among them, the outputs of the first - way multiplication unit and the second - way multiplication unit are respectively connected to the inputs of the first bit concatenation 2.1. The outputs of the third - way multiplication unit and the fourth - way multiplication unit are respectively connected to the inputs of the second bit concatenation 2.2. The output of the first bit concatenation 2.1 is connected to the first summation module 4.1. The output of the second bit concatenation 2.2 is connected to the first shift unit 3.1. The output of the first shift unit 3.1 is connected to the first summation module 4.1. The output of the first summation module 4.1 is connected to the fifth summation module 4.5 for summation. In the described first computing unit Cell1, the 2 - bit multiplication unit 1 based on multiplexing is composed of 4 - way multiplication units as inputs. Among them, the outputs of the first - way multiplication unit and the second - way multiplication unit are respectively connected to the inputs of the fifth bit concatenation 2.5. The outputs of the third - way multiplication unit and the fourth - way multiplication unit are respectively connected to the inputs of the sixth bit concatenation 2.6. The output of the fifth bit concatenation 2.5 is connected to the second summation module 4.2. The output of the sixth bit concatenation 2.6 is connected to the third shift unit 3.3. The output of the third shift unit 3.3 is connected to the second summation module 4.2. The output of the second summation module 4.1 is connected to the second shift unit 3.2. The output of the second shift unit 3.2 is connected to the fifth summation module 4.5 for summation.In the described zero - th computing unit Cell2, the 2 - bit multiplication unit 1 based on multiplexing is composed of 4 multiplication units as inputs. Among them, the outputs of the first multiplication unit and the second multiplication unit are respectively connected to the inputs of the third - bit concatenation 2.3. The outputs of the third multiplication unit and the fourth multiplication unit are respectively connected to the inputs of the fourth - bit concatenation 2.4. The output of the third - bit concatenation 2.3 is connected to the fourth summation module 4.4. The output of the fourth - bit concatenation 2.4 is connected to the fourth shift unit 3.4. The output of the fourth shift unit 3.4 is connected to the fourth summation module 4.4. The output of the fourth summation module 4.4 is connected to the fifth shift unit 3.5. The output of the fifth shift unit 3.5 is connected to the sixth summation module 4.6 for summation. In the described zero - th computing unit Cell3, the 2 - bit multiplication unit 1 based on multiplexing is composed of 4 multiplication units as inputs. Among them, the output of the first multiplication unit is connected to the seventh summation module 4.7. The output of the second multiplication unit is connected to the seventh shift unit 3.7. The output of the seventh shift unit 3.7 is connected to the seventh summation module 4.7. The output of the seventh summation module 4.7 is connected to the ninth summation module 4.9. The output of the third multiplication unit is connected to the eighth summation module 4.8. The output of the fourth multiplication unit is connected to the sixth shift unit 3.6. The output of the sixth shift unit 3.6 is connected to the eighth summation module 4.8. The output of the eighth summation module 4.8 is connected to the eighth shift unit 3.8. The output of the eighth shift unit 3.8 is connected to the ninth summation module 4.9. The output of the ninth summation module 4.9 is connected to the ninth shift unit 3.9. The output of the ninth shift unit 3.9 is connected to the sixth summation module 4.6 for summation. Among the shift units, the first shift unit 3.1, the second shift unit 3.2, the third shift unit 3.3, the fourth shift unit 3.4, the sixth shift unit 3.6, the seventh shift unit 3.7, and the eighth shift unit 3.8 are two - bit shifts. The fifth shift unit 3.5 is a four - bit shift, and the ninth shift unit 3.9 is a six - bit shift.

[0033] The configurable multiply - accumulate unit is applied to the LeNet5 network, where the activations and weights are quantized to 8 bits, 4 bits, and 2 bits respectively. Since the quantization of activations has a greater impact on accuracy than the quantization of weights, the precision - configurable MAC unit supports 8×8, 4×4, and 2×2 calculations, which can greatly reduce the hardware computational complexity within the allowable accuracy loss range.

[0034] The precision-configurable multiply-accumulate unit for a neural network accelerator of the present invention includes: a 2-bit multiplication unit based on multiplexing, bit splicing, a shift unit, and an approximate LOA adder unit. The precision-configurable multiply-accumulate unit arranges 16 2-bit multiplication units based on multiplexing in a certain order in space to adapt to the MAC operations of 2-bit, 4-bit, and 8-bit DNN layers. It is compatible with inputting 16 2-bit multipliers or 4 4-bit input multipliers or 1 8-bit input multiplier. The steps for implementing the multiplication of two 8-bit signed numbers A and B in the configurable MAC are as follows:

[0035] Step 1: Divide A and B into 4 2-bit numbers respectively - A[1:0] (denoted as A 0,1 ), A[3:2] (denoted as A 2,3 ), A[5:4] (denoted as A 4,5 ), A[7:6] (denoted as A 6,7 ), B[1:0] (denoted as B 0,1 ), B[3:2] (denoted as B 2,3 ), B[5:4] (denoted as B 4,5 ) and B[7:6] (denoted as B 6,7 ).

[0036] Step 2: A 0,1 ,, A 2,3 , A 4,5 and A 6,7 are respectively broadcast to each 2-bit multiplication unit in 4 cells.

[0037] Step 3: Each cell respectively receives the 4 2-bit numbers obtained by splitting B and distributes them to the 2-bit multiplication units according to a certain rule.

[0038] Step 4: After the partial sums obtained in the 4 cells are shifted correspondingly, they are accumulated to obtain the product of A and B. As Figure 2 shown, the 2-bit multiplication unit based on the multiplexer is a bit-level processing element. As the configurable minimum multiplication unit, it realizes the multiplication of 2-bit signed / unsigned numbers. It includes: a tenth shift unit 1.1, a first multiplier 1.2.1, a second multiplier 1.2.2, a tenth summation module 1.3, a first multiplexer 1.4.1, and a second multiplexer 1.4.2;

[0039] When this design performs 2-bit multiplication, the multiplexer determines the signed or unsigned operands, thus avoiding adding a sign bit. For 2-bit signed / unsigned numbers A and B, the inputs and outputs of the multiplier are in the form of signed two's complement. cond(1) represents the case of "unsigned A × unsigned B"; cond(2) represents "signed A × unsigned B"; cond(3) represents the case of "signed A × signed B". As shown in the following formula, the multiplication of a signed number and an unsigned number can reuse the results of two unsigned number multiplications. Through the multiplexer, the product of 2-bit signed / unsigned numbers can be calculated without adding an extra sign bit.

[0040] cond(1): A×B = (2A[1]+A[0])×(2B[1]+B[0]) = 4A[1]B[1]+2A[1]B[0]+2A[0]B[1]+A[0]B[0]. cond(2): A×B = (-2A[1]+A[0])×(2B[1]+B[0]) = -4A[1]B[1]-2A[1]B[0]+2A[0]B[1]+A[0]B[0] = cond(1)+([2B[1]+B[0]]<<2)×A[1].

[0041] cond(3): A×B = (-2A[1]+A[0])×(-2B[1]+B[0]) = 4A[1]B[1]-2A[1]B[0]-2A[0]B[1]+A[0]B[0].

[0042] The bit concatenation mentioned above directly combines non-interfering partial sums through bit concatenation to achieve 4 / 8-bit multiplication. As Figure 3 shown, taking 4×4 as an example, first split 4-bit A into 2-bit A[3:2] and A[1:0], and B is also split into 2-bit B[3:2] and B[1:0]. The result of multiplying the lower 2 bits A[1:0] and B[1:0] is in the lower 4 bits, and the result of multiplying the higher 2 bits A[3:2] and B[3:2] is in the higher 4 bits. These two partial products do not interfere with each other, so they can be directly concatenated by bits without the need for accumulation.

[0043] As Figure 4As shown, the approximate unit LOA includes an exact adder and an approximate adder based on an OR gate. The approximate unit splits the addition into two non-overlapping sub-adders. The MSB is implemented by a traditional exact adder, and the LSB is implemented by bitwise OR of an approximate adder based on an OR gate. Since the required approximate bit widths for adders with different bit widths are different, each adder was modeled in MATLAB, certain equal intervals were selected, and its MRED was tested using the Monte Carlo method. Considering that the MRED of the computing unit is usually required to be less than 5%, for 6 / 8 / 10 / 16-bit wide LOA adders, the approximate bit widths are 2 / 3 / 4 / 6 bits respectively.

Claims

1. A precision-configurable multiply-accumulate unit for a neural network accelerator, characterized in that: The multiply-accumulate unit includes four computing units, namely the zero-th computing unit (Cell0), the first computing unit (Cell1), the second computing unit (Cell2), and the third computing unit (Cell3), and three summation modules, namely the third summation module (4.3), the fifth summation module (4.5), and the sixth summation module (4.6). First, the zero-th computing unit (Cell0) and the first computing unit (Cell1) are summed through the fifth summation module (4.5). Then, the second computing unit (Cell2) and the third computing unit (Cell3) are summed through the sixth summation module (4.6). Finally, the partial results generated by the above two summations are finally summed through the third summation module (4.3). Among them, all summation modules in the precision-configurable multiply-accumulate unit are correspondingly turned off according to different input bit widths to adapt to the multiply-accumulate operations of each layer of the convolutional neural network with 2-bit, 4-bit, and 8-bit different bit widths. The four units, namely the zero-th computing unit (Cell0), the first computing unit (Cell1), the second computing unit (Cell2), and the third computing unit (Cell3), are respectively the minimum computing units. Each minimum computing unit includes a 2-bit multiplication unit (1) based on multiplexing, a bit concatenation unit, a shift unit, and a summation module. Among them, The bit concatenation unit includes the first bit concatenation (2.1), the second bit concatenation (2.2), the third bit concatenation (2.3), the fourth bit concatenation (2.4), the fifth bit concatenation (2.5), and the sixth bit concatenation (2.6). The shift unit includes the first shift unit (3.1), the second shift unit (3.2), the third shift unit (3.3), the fourth shift unit (3.4), the fifth shift unit (3.5), the sixth shift unit (3.6), the seventh shift unit (3.7), the eighth shift unit (3.8), and the ninth shift unit (3.9). The summation modules include the first summation module (4.1), the second summation module (4.2), the third summation module (4.3), the fourth summation module (4.4), the fifth summation module (4.5), the sixth summation module (4.6), the seventh summation module (4.7), the eighth summation module (4.8), and the ninth summation module (4.9).

2. The precision-configurable multiply-accumulate unit for a neural network accelerator according to claim 1, characterized in that: In the zero-th computing unit (Cell0), the 2-bit multiplication unit (1) based on multiplexing is composed of 4-way multiplication units as inputs. Among them, the outputs of the first-way multiplication unit and the second-way multiplication unit are respectively connected to the inputs of the first bit concatenation (2.1). The outputs of the third-way multiplication unit and the fourth-way multiplication unit are respectively connected to the inputs of the second bit concatenation (2.2). The output of the first bit concatenation (2.1) is connected to the first summation module (4.1). The output of the second bit concatenation (2.2) is connected to the first shift unit (3.1). The output of the first shift unit (3.1) is connected to the first summation module (4.1). The output of the first summation module (4.1) is connected to the fifth summation module (4.5) for summation.

3. The precision-configurable multiply-accumulate unit for a neural network accelerator according to claim 1, characterized in that: In the described first computing unit (Cell1), the multiplexer-based 2-bit multiplication unit (1) is composed of four multiplication units as inputs. Among them, the outputs of the first multiplication unit and the second multiplication unit are respectively connected to the inputs of the fifth-bit concatenation (2.5). The outputs of the third multiplication unit and the fourth multiplication unit are respectively connected to the inputs of the sixth-bit concatenation (2.6). The output of the fifth-bit concatenation (2.5) is connected to the second summation module (4.2). The output of the sixth-bit concatenation (2.6) is connected to the third shift unit (3.3). The output of the third shift unit (3.3) is connected to the second summation module (4.2). The output of the second summation module (4.2) is connected to the second shift unit (3.2). The output of the second shift unit (3.2) is connected to the fifth summation module (4.5) for summation.

4. The precision-configurable multiply-accumulate unit for a neural network accelerator according to claim 1, wherein: In the described second computing unit (Cell2), the multiplexer-based 2-bit multiplication unit (1) is composed of four multiplication units as inputs. Among them, the outputs of the first multiplication unit and the second multiplication unit are respectively connected to the inputs of the third-bit concatenation (2.3). The outputs of the third multiplication unit and the fourth multiplication unit are respectively connected to the inputs of the fourth-bit concatenation (2.4). The output of the third-bit concatenation (2.3) is connected to the fourth summation module (4.4). The output of the fourth-bit concatenation (2.4) is connected to the fourth shift unit (3.4). The output of the fourth shift unit (3.4) is connected to the fourth summation module (4.4). The output of the fourth summation module (4.4) is connected to the fifth shift unit (3.5). The output of the fifth shift unit (3.5) is connected to the sixth summation module (4.6) for summation.

5. The precision-configurable multiply-accumulate unit for a neural network accelerator according to claim 1, wherein: In the described third computing unit (Cell3), the multiplexer-based 2-bit multiplication unit (1) is composed of four multiplication units as inputs. Among them, the output of the first multiplication unit is connected to the seventh summation module (4.7). The output of the second multiplication unit is connected to the seventh shift unit (3.7). The output of the seventh shift unit (3.7) is connected to the seventh summation module (4.7). The output of the seventh summation module (4.7) is connected to the ninth summation module (4.9). The output of the third multiplication unit is connected to the eighth summation module (4.8). The output of the fourth multiplication unit is connected to the sixth shift unit (3.6). The output of the sixth shift unit (3.6) is connected to the eighth summation module (4.8). The output of the eighth summation module (4.8) is connected to the eighth shift unit (3.8). The output of the eighth shift unit (3.8) is connected to the ninth summation module (4.9). The output of the ninth summation module (4.9) is connected to the ninth shift unit (3.9). The output of the ninth shift unit (3.9) is connected to the sixth summation module (4.6) for summation.

6. The precision-configurable multiply-accumulate unit for a neural network accelerator according to claim 1, wherein: The first shift unit (3.1), the second shift unit (3.2), the third shift unit (3.3), the fourth shift unit (3.4), the sixth shift unit (3.6), the seventh shift unit (3.7), and the eighth shift unit (3.8) in the shift unit are two-bit shifts, the fifth shift unit (3.5) is a four-bit shift, and the ninth shift unit (3.9) is a six-bit shift.

7. The precision-configurable multiply-accumulate unit for a neural network accelerator according to claim 1, wherein: The partial products generated by directly merging the high two bits and the low two bits through bit concatenation replace the adder and the shifter, effectively avoiding the operations of shifting and accumulation; the bit concatenation stitches together the products of 2 of the 2-bit multiplication units (1) based on multiplexing. These 2 products do not affect each other in the final product, and the bits they occupy do not overlap. The partial sum of the high bits is directly placed in the high 4 bits, and the partial sum of the low bits is directly placed in the low 4 bits.

8. The precision-configurable multiply-accumulate unit for a neural network accelerator according to claim 1, characterized in that: The shift unit is composed of 2-bit, 4-bit, and 6-bit left shifters. By using the shifter to replace the function of the adder, the purpose of reducing the number of adders or reducing the bit width of the adder is achieved. Ultimately, less circuit area is used to improve the circuit performance.

9. The precision-configurable multiply-accumulate unit for a neural network accelerator according to claim 1, characterized in that: The summation module splits the adder into two non-overlapping calculation modules according to the high bits and the low bits. The split high-bit calculation module is implemented by an exact adder, and the split low-bit calculation module is implemented by bitwise OR of an approximate adder based on an OR gate; the summation partial result of the split low-bit part is approximate. To compensate for the error caused by the approximate calculation, an AND gate is used for the highest bit of the split low-bit part as an error compensation.

Citation Information

Patent Citations

  • High-energy-efficiency approximate multiplier with reconfigurable precision and bit width

    CN115033204A

  • Arithmetic unit compatible with asymmetric multi-precision hybrid multiply-accumulate operation

    CN115357214A