High speed low power approximate multiply accumulate operator for image convolution processing
By designing an approximate multiply-accumulate, and employing an approximate radix-8 Buss encoder and an approximate carry-add module, the imbalance in speed, power consumption, and area of existing multiply-accumulates is resolved, achieving high-speed, low-power image convolution processing while maintaining image quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2026-04-07
AI Technical Summary
Existing multiply-accumulate units are difficult to balance in terms of speed, power consumption, and area, and lack designs specifically for image convolution processing, failing to meet the precision tolerance and low power consumption requirements of image processing.
An approximate multiply-accumulate converter was designed, employing an approximate radix-8-Bose encoder, a 3:2 compressor, and an approximate carry-add module for image convolution processing. By using approximate algorithms in the partial product generation, compression, and accumulation stages, the circuit complexity and power consumption are reduced.
It achieves high-speed, low-power image convolution processing, reduces circuit resources, increases speed, and reduces power consumption and area, while maintaining good image quality within tolerance limits.
Smart Images

Figure CN115480729B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of integrated circuits, and particularly relates to a high-speed and low-power approximate multiplier accumulator for image convolution processing. BACKGROUND
[0002] In recent years, with the development of big data and artificial intelligence, Internet of Things devices need to process a large amount of pictures, and the size and bit depth of the pictures are also increasing, which continuously improves the requirements for speed and power consumption of image processing devices. Convolution is often used in image smoothing, blurring, sharpening, denoising and edge detection, and is a common operation in image processing. Convolution calculation is a multiplication and accumulation operation, and the core is a multiplier accumulator (MAC) operator. Therefore, it is of great significance to design a high-speed and low-power multiplier accumulator for image convolution processing.
[0003] Internet of Things devices make approximate calculation promising for the following reasons: 1. In image, audio and video processing, the final result is interpreted by human senses, and the fact that human senses are limited reduces the strict requirement for accuracy. 2. Redundant and noisy data need to be processed. The basic premise of approximate calculation is to replace the traditional complex and energy-wasting data processing block with a data processing block with less logic and lower complexity. This method effectively reduces chip energy consumption and area and improves speed at the cost of reducing the accuracy of processed data.
[0004] At present, related research on multiplier accumulators is not sufficient, and there is no circuit that balances speed, power consumption and area performance. In addition, there are few multiplier accumulators designed specifically for image convolution processing. Therefore, the application designs a multiplier accumulator for image convolution processing and achieves high speed and low power consumption by using approximate calculation. SUMMARY
[0005] The approximate multiplier accumulator designed in the application can comprehensively improve the speed, power consumption and area performance compared with traditional circuits. In addition, the circuit is used for image convolution processing and can meet the tolerance of image processing accuracy.
[0006] The high-speed and low-power approximate multiplier accumulator for image convolution processing processes 8-bit signed numbers and 8-bit unsigned numbers. The unsigned numbers are used to process image pixels, and the signed numbers are used to process convolution kernels. For the input signed number, a sign bit is added, and the value is the 8th bit of the signed number.
[0007] The approximate accumulator is divided into four modules, i.e. a partial product generation module, a partial product compression module, a carry addition module and an accumulation module. The partial product generation module is an approximate base-8-Booth encoder, and generates a partial product matrix with a height of 3. The partial product compression module is a 3:2 compressor. The carry addition module approximately adds 16-bit carry to 4-bit addition. The accumulator is composed of an adder and a D flip-flop, and the addition operation adopts the approximate addition of 4-bit carry.
[0008] The approximate base-8-Booth encoder is characterized in that it is composed of two two-input XOR gates, one two-input XNOR gate, one two-input OR gate, one three-input OR gate, one four-input OR gate, five two-input AND gates, four three-input AND gates, two four-input AND gates and one three-input NAND gate, and sequentially comprises a first two-input XOR gate, a second two-input XOR gate, a two-input XNOR gate, a two-input first OR gate, a three-input second OR gate, a four-input third OR gate, a two-input first AND gate, a two-input second AND gate, a two-input third AND gate, a two-input fourth AND gate, a two-input fifth AND gate, a three-input sixth AND gate, a three-input seventh AND gate, a three-input eighth AND gate, a three-input ninth AND gate, a four-input tenth AND gate, a four-input eleventh AND gate and a three-input NAND gate.
[0009] The multiplier (Y) is encoded into 4-bit groups, i.e. y 3i+2 , y 3i+1 , y 3i , y 3i-1 , as inputs, the multiplicand x i , x i-1 , x i-2 , as inputs, and the generated partial product PP ij is output.
[0010] The first input end of the first XOR gate is connected with a first input signal (y 3i+2 ), the second input end is connected with a second input signal (y 3i+1 ), and the output end is connected with the first input end of the first AND gate.
[0011] The first input end of the first XNOR gate is connected with a third input signal (y 3i ), the second input end is connected with a fourth input signal (y 3i-1 ), and the output end is connected with the second input end of the first AND gate.
[0012] The first input end of the second AND gate is connected with the output end of the first AND gate, the second input end is connected with a fifth input signal (x i ), and the output end is connected with the first input end of the second OR gate.
[0013] The first input end of the tenth AND gate is connected with the first input signal (y 3i+2 ), the second input end is connected with the inverted second input signal (y3i+1 ), the third input of which is connected to the third input signal (y 3i ), the fourth input of which is connected to the negated fourth input signal (y 3i-1 ), the output of which is connected to the first input of the third OR gate;
[0014] the first input of the eleventh AND gate is connected to the negated first input signal (y 3i+2 ), the second input of which is connected to the second input signal (y 3i+1 ), the third input of which is connected to the negated third input signal (y 3i ), the fourth input of which is connected to the fourth input signal (y 3i-1 ), the output of which is connected to the second input of the third OR gate;
[0015] the first input of the sixth AND gate is connected to the second input signal (y 3i+1 ), the second input of which is connected to the negated third input signal (y 3i ), the third input of which is connected to the negated third input signal (y 3i-1 ), the output of which is connected to the third input of the third OR gate;
[0016] the first input of the seventh AND gate is connected to the negated second input signal (y 3i+1 ), the second input of which is connected to the third input signal (y 3i ), the third input of which is connected to the fourth input signal (y 3i-1 ), the output of which is connected to the fourth input of the third OR gate;
[0017] the first input of the third AND gate is connected to the output of the third OR gate, the second input of which is connected to the sixth input signal (x i-1 ), the output of which is connected to the second input of the second OR gate;
[0018] the first input of the eighth AND gate is connected to the negated first input signal (y 3i+2 ), the second input of which is connected to the second input signal (y 3i+1 ), the third input of which is connected to the third input signal (y 3i ), the output of which is connected to the first input of the first OR gate;
[0019] the first input of the ninth AND gate is connected to the first input signal (y 3i+2 ), the second input of which is connected to the negated second input signal (y 3i+1 ), the third input of which is connected to the negated third input signal (y 3i ), the output of which is connected to the second input of the first OR gate;
[0020] the first input of the fourth AND gate is connected to the seventh input signal (xi-2 ), the second input of which is connected to the output of the first OR gate, and the output of which is connected to the third input of the second OR gate;
[0021] the first input of the NAND gate is connected to the second input signal (y 3i+1 ), the second input of which is connected to the third input signal (y 3i ), the third input of which is connected to the fourth input signal (y 3i-1 ), the output of which is connected to the first input of the fifth AND gate;
[0022] the first input of the fifth AND gate is connected to the output of the NAND gate, the second input of which is connected to the first input signal (y 3i+2 ), the output of which is connected to the second input of the second XOR gate;
[0023] the first input of the second XOR gate is connected to the output of the second OR gate, the second input of which is connected to the output of the fifth AND gate, and the output of which is connected to the output signal PP ij ;
[0024] The partial product compression module described in the application is characterized in that the 3:2 compressor is composed of 16 full adders, the first input of the first full adder is the first bit data of the first row of the partial product matrix, the second input of the first full adder is the first bit data of the second row of the partial product matrix, the third input of the first full adder is the first bit data of the third row of the partial product matrix, and so on, the first input of each full adder is the data of the first row, the second input is the data of the second row, and the third input is the data of the third row, each full adder processes a column of data of the partial product, and the two outputs of the full adder form a matrix with a height of 2.
[0025] The approximate carry adder module described in the application is characterized in that in the approximate carry adder, a binary integer is represented as A, B, and each bit of the integer A, B is represented as a i-1 ,b i-1 . In order to add two n-bit integers A and B, a generation signal g i , a propagation signal p i and a termination signal k i are defined for each bit, wherein g i is equal to a i and b i , p i is equal to a i XOR b i ; and k i is equal to a i OR NOT b i .
[0026] Using these signals, a carry output signal c iand is used to calculate the sum s of each bit i .c i and s i The recursive formula is:
[0027] When K i is equal to 1, c i is equal to 0; when g i is equal to 1, c i is equal to 1; when p i is equal to 1, c i is equal to c i-1 .
[0028] s i is equal to a i XOR b i XOR c i-1 .
[0029] For the 16-bit adder of the present application, its longest sequence is set to 4, therefore, s i is calculated using the input bits of the first 5 bits from the i-th bit. That is, S5 is calculated by a0~a5, b0~b5, S6 is calculated by a1~a6, b1~b6, and so on, and S 15 is calculated by a 10 ~a 15 , b 10 ~b 15 . There are 11 six-bit adders in total.
[0030] The accumulator module described in the present application is characterized in that: the accumulator is composed of an approximate adder and a D flip-flop, wherein the approximate adder is the approximate carry adder described above, the first input end of the approximate adder is connected with an input signal (din), the second input end is connected with the output (sum_d) of the D flip-flop, and the output end is connected with the input (dout) and the output signal (dout) of the D flip-flop; the input end of the D flip-flop is connected with the output (dout) of the approximate adder, and the output end is connected with the input (sum_d) of the approximate adder.
[0031] Compared with the prior art, the present application has the beneficial effects that:
[0032] 1. The approximate multiply-accumulator designed in the present application is dedicated to image processing, which is for 8-bit unsigned number x 8-bit signed number, wherein the unsigned number is used for processing image pixels, and the signed number is used for processing convolution kernels, thereby reducing circuit resources, improving speed, reducing power consumption, and reducing area.
[0033] 2. The approximate radix-8-Booth algorithm is used, which simplifies the circuit structure, reduces power consumption and area.
[0034] 3. The approximate 4-carry adder is used, which reduces the delay and improves the speed.
[0035] 4. Using the approximate circuit for image convolution processing, the image quality is within the tolerance. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 The schematic diagram of the approximate multiplier-accumulator.
[0037] Figure 2 The operation process of the multiplier-accumulator.
[0038] Figure 3 The value of the approximate base-8-Booth encoder.
[0039] Figure 4 The approximate base-8-Booth encoder circuit.
[0040] Figure 5 The partial product row of the base-8-Booth encoding multiplier.
[0041] Figure 6 The schematic diagram of the approximate adder.
[0042] Figure 7 The performance comparison of three multiplier-accumulators.
[0043] Figure 8 The picture after image processing using the approximate multiplier-accumulator. DETAILED DESCRIPTION
[0044] In this embodiment, a high-speed low-power approximate multiplier-accumulator for image convolution processing has a structure as shown in the figure. Figure 1 The approximate multiplier-accumulator designed in the application is used for image convolution processing. The pixels of the picture are 8-bit unsigned numbers, and the convolution kernel can be positive or negative, such as the mean convolution kernel 1 / 9 [1, 1, 1; 1, 1, 1; 1, 1, 1], the sharpening convolution kernel [-1, -1, -1; -1, 9, -1; -1, -1, -1], etc. In order to adapt to this situation, the approximate multiplier-accumulator designed in the application is for 8-bit unsigned numbers (X) x 8-bit signed numbers (Y), wherein the unsigned numbers are used to process the image pixels, and the signed numbers are used to process the convolution kernel. The implementation method is as follows: for the input signed number Y, the highest bit is used as the sign bit to obtain Y_MAC, i.e. Y_MAC = {B[7], Y[7:0]}.
[0045] The multiplier-accumulator is composed of a multiplier and an adder, and is divided into four stages of partial product generation, partial product compression, carry addition, and accumulation, as shown in the figure. Figure 2 The approximate algorithm is used in the stages of partial product generation, carry addition, and accumulation.
[0046] Partial product generation, for n-bit multiplicand X and multiplier Y, the bit multiplication of multiplicand and multiplier with AND logic gate in array multiplication, generates n x n partial products, the height of partial product matrix is n. In order to improve the speed, the Booth encoder is commonly used, base 2 r Booth encoder encodes the multiplier (Y) into a set of r+1 bits, overlapping 1 bit, the height of partial product matrix is [n / r], there are [n / r] x n partial products. With the increase of r value, the height of partial product matrix decreases, the partial product compression stage is simplified, but the complexity of Booth encoder increases. Taking base 4-Booth encoder as an example, the height of partial product matrix is reduced from n to [n / 2], and the partial products of 0, ±1 and ±2 times of multiplicand (X) are needed. Similarly, base 8-Booth encoder, the height of partial product matrix is reduced from n to [n / 3], and the partial products of 0, ±1, ±2, ±3 and ±4 times of multiplicand (X) are needed.
[0047] Since the multiplier-accumulator in this application is used for image convolution processing, the multiplicand X and the multiplier Y are 8-bit wide, and the base 8-Booth encoder is suitable in terms of the height of partial product matrix and the complexity of the encoder. In the base 8-Booth encoder, 3X is not a power of 2, which cannot be obtained by shifting X. The 3X term is usually obtained by adding the X term and the 2X term through an n-bit adder, which is called a recoded adder. The encoding adder increases the overall critical path delay, which is the main disadvantage of the base 8-Booth encoder.
[0048] In order to reduce the complexity in the base 8-Booth encoder, an approximate base 8-Booth encoder is proposed. The approximate method is to approximate the 3X term which consumes a lot of resources to 2X term and 4X term, reducing the complexity of the circuit, as shown in Figure 3 As shown in the figure, the 3X input terms (0101, 1010) and (0110, 1001) are approximated to 2X and 4X terms, and in the figure, y 3i+2 y 3i+1 y 3i y 3i-1 The multiplicand is represented by pp ij The accurate partial product is represented by R8ABE1, and the error of the approximate encoder is represented by ED.
[0049] The circuit of the approximate base 8-Booth encoder after approximation is as follows: Figure 4As shown, the circuit is described as: composed of two two-input XOR gates, a two-input XNOR gate, a two-input OR gate, a three-input OR gate, a four-input OR gate, 5 two-input AND gates, 4 three-input AND gates, 2 four-input AND gates, a three-input NAND gate, and in turn: a first two-input XOR gate, a second XOR gate, a two-input XNOR gate, a first two-input OR gate, a second three-input OR gate, a third four-input OR gate, a first two-input AND gate, a second two-input AND gate, a third two-input AND gate, a fourth two-input AND gate, a fifth two-input AND gate, a sixth three-input AND gate, a seventh three-input AND gate, an eighth three-input AND gate, a ninth three-input AND gate, a tenth four-input AND gate, an eleventh four-input AND gate, a three-input NAND gate.
[0050] The multiplier (Y) is encoded as 4-bit groups, respectively y 3i+2 , y 3i+1 , y 3i , y 3i-1 As input, the multiplicand x i , x i-1 , x i-2 As input, the generated partial product PP ij is output.
[0051] The first input end of the first XOR gate is connected to the first input signal (y 3i+2 ), the second input end is connected to the second input signal (y 3i+1 ), and the output end is connected to the first input end of the first AND gate.
[0052] The first input end of the first XNOR gate is connected to the third input signal (y 3i ), the second input end is connected to the fourth input signal (y 3i-1 ), and the output end is connected to the second input end of the first AND gate.
[0053] The first input end of the second AND gate is connected to the output end of the first AND gate, the second input end is connected to the fifth input signal (x i ), and the output end is connected to the first input end of the second OR gate.
[0054] The first input end of the tenth AND gate is connected to the first input signal (y 3i+2 ), the second input end is connected to the negated second input signal (y 3i+1 ), the third input end is connected to the third input signal (y 3i ), the fourth input end is connected to the negated fourth input signal (y 3i-1 ), and the output end is connected to the first input end of the third OR gate.
[0055] The first input end of the eleventh AND gate is connected to the negated first input signal (y 3i+2 ), the second input end is connected to the second input signal (y 3i+1), the third input of which is connected to the negated third input signal (y 3i ), the fourth input of which is connected to the fourth input signal (y 3i-1 ), the output of which is connected to the second input of the third OR gate;
[0056] the first input of the sixth AND gate is connected to the second input signal (y 3i+1 ), the second input of which is connected to the negated third input signal (y 3i ), the third input of which is connected to the negated third input signal (y 3i-1 ), the output of which is connected to the third input of the third OR gate;
[0057] the first input of the seventh AND gate is connected to the negated second input signal (y 3i+1 ), the second input of which is connected to the third input signal (y 3i ), the third input of which is connected to the fourth input signal (y 3i-1 ), the output of which is connected to the fourth input of the third OR gate;
[0058] the first input of the third AND gate is connected to the output of the third OR gate, the second input of which is connected to the sixth input signal (x i-1 ), the output of which is connected to the second input of the second OR gate;
[0059] the first input of the eighth AND gate is connected to the negated first input signal (y 3i+2 ), the second input of which is connected to the second input signal (y 3i+1 ), the third input of which is connected to the third input signal (y 3i ), the output of which is connected to the first input of the first OR gate;
[0060] the first input of the ninth AND gate is connected to the first input signal (y 3i+2 ), the second input of which is connected to the negated second input signal (y 3i+1 ), the third input of which is connected to the negated third input signal (y 3i ), the output of which is connected to the second input of the first OR gate;
[0061] the first input of the fourth AND gate is connected to the seventh input signal (x i-2 ), the second input of which is connected to the output of the first OR gate, the output of which is connected to the third input of the second OR gate;
[0062] the first input of the NAND gate is connected to the second input signal (y 3i+1 ), the second input of which is connected to the third input signal (y 3i ), the third input of which is connected to the fourth input signal (y 3i-1 ), the output of which is connected to the first input of the fifth AND gate;
[0063] The first input of the fifth AND gate is connected to the output of the NAND gate, the second input is connected to the first input signal (y 3i+2 ), and the output is connected to the second input of the second XOR gate;
[0064] The first input of the second XOR gate is connected to the output of the second OR gate, the second input is connected to the output of the fifth AND gate, and the output is connected to the output signal PP ij ;
[0065] The expression of the simplified logic PP ij is given by:
[0066] PP ij = [x i (y 3i+2 ⊙y 3i+1 )(y 3i ⊕y 3i-1 )+x i-1 ((~y 3i+1 )y 3i y 3i-1 +y 3i+1 (~y 3i )(~y 3i-1 )+(y 3i+2 )y 3i+1 (~y 3i )y 3i-1 +y 3i+2 (~y 3i+1 )y 3i (~y 3i-1 ))+x i-1 ((~y 3i+2 )y 3i+1 y 3i +y 3i+2 (~y 3i=1 )(~y 3i ))]⊕y 3i+2 (~(y 3i+1 y 3i y 3i-1 ))
[0067] For the negative partial product row, one also adds one at the LSB. Figure 5 The partial product row of the radix-8 Booth encoded multiplier is shown in Fig. 1 using dot notation, where each dot represents a partial product bit.
[0068] The partial product compression part, the height of the matrix generated by the partial product is 3, uses a 3:2 compressor to complete the partial product compression part. The 3:2 compressor is composed of 16 full adders, the first input of the first full adder is the first bit data of the first row of the partial product matrix, the second input of the first full adder is the first bit data of the second row of the partial product matrix, the third input of the first full adder is the first bit data of the third row of the partial product matrix, and so on. The first input of each full adder is the data of the first row, the second input is the data of the second row, and the third input is the data of the third row. Each full adder processes a column of partial products, and the two outputs of the full adder form a matrix with a height of 2.
[0069] Carry addition, in the exact carry addition, binary integers are represented as A, B, the highest bit of the integer A, B is represented as a i-1 ,b i-1 . In order to add two n-bit integers A and B, the generation signal g i , the propagation signal p i and the termination signal k i of each bit can be defined as follows:
[0070] g i =a i b i ,p i =a i ⊕b i , k i =~(a i +b i )
[0071] Using these signals, the carry output signal c i of each bit i is generated, and is used to calculate the sum of each bit. The recursive formula of c i and the sum of each bit are as follows.
[0072] c i =0(k i =1),c i =1(g i =1),c i =c i-1 (p i =1)
[0073] s i =a i ⊕b i ⊕c i-1
[0074] Only when the propagation signal p i is true, the carry c i depends on the carry c i-1 , otherwise ci According to g i and k i The value of is determined only when p i-1 When true, c i-1 Depends on c i-2 This means that only when p i and p i-1 When both are true, c i Only depends on c i-2 Generally speaking, only when position c i and c i-k+1 Each propagating signal p between i When true, c i Only depends on c i-k .
[0075] For the addition of two 16-bit integers, the longest sequence is 14, but the average propagation sequence is (log2n), which is 4. Therefore, the approximation method of this invention is to set the longest sequence of the propagation signal to 4. Figure 6 In this example, we want to add two 16-bit integers. Since the longest sequence of propagated signals is 4, therefore, s i It can only be calculated using the first 5 input bits starting from the i-th bit. That is, S5 is calculated using a0~a5, b0~b5, S6 is calculated using a1~a6, b1~b6, and so on. 15 By a 10 ~a 15 ,b 10 ~b 15 Calculation. A total of 11 6-bit adders are formed, such as... Figure 6 As shown, the latency of this special 20-bit adder is practically the same as that of a 6-bit adder. This approximation method reduces the accumulator latency and increases speed.
[0076] The accumulator consists of an approximate adder and a D flip-flop. The approximate adder is the aforementioned approximate carry adder. The first input of the approximate adder is connected to the input signal (din), the second input is connected to the output (sum_d) of the D flip-flop, and the output is connected to the input (dout) and output signal (dout) of the D flip-flop. The input of the D flip-flop is connected to the output (dout) of the approximate adder, and the output is connected to the input (sum_d) of the approximate adder.
[0077] In summary, the approximate multiply-accumulate operator for image convolution processing has the following advantages: the accumulator of 8-bit signed number*8-bit unsigned number is designed according to the characteristics of convolution image processing, so that the resource waste of the circuit is avoided, the delay, area and power consumption are reduced; the approximate base 8-Booth encoder is used in the partial product generation stage, so that the power consumption and area are reduced; the approximate algorithm is used in the addition stage, so that the number of carry bits is reduced, thereby the critical path and delay time are reduced and the operation speed is improved.
[0078] The hardware design of the multiply-accumulate operator is designed by using the verilog language, and the delay, area, power consumption and other performances of the circuit are obtained by using the library of scc018ug_hd_rvt_ff_v1p98_-40c in DC. Figure 7 Among the non-optimized MAC, the multiplier and the adder directly use the * and + arithmetic symbols, and the multiply-accumulate operator needs to process image convolution, so the non-optimized MAC is 16-bit*16-bit signed number. The optimized precise MAC is 8-bit signed number*8-bit unsigned number, but the approximate algorithm is not used. Figure 7 It can be seen that compared with the non-optimized MAC, the delay time is reduced by 44.29%, the power consumption is reduced by 57.03%, and the area is reduced by 71.42%; compared with the optimized precise MAC, the delay time is reduced by 6.02%, the power consumption is reduced by 4.84%, and the area is reduced by 13.68%.
[0079] The process of image convolution processing is as follows: the pixels of the picture are obtained by using matlab to process the picture, the pixels of the image are convolved by using the approximate accumulator, and then the processed pixels are imported into matlab for image quality evaluation. The picture after sharpening convolution image processing is as follows: Figure 8 Among them, the gray picture is the original picture without processing, the second picture is the picture processed by using the convolution of matlab, and the third picture is the picture processed by using the approximate multiply-accumulate operator. The peak signal-to-noise ratio (PSNR) of the two pictures after image processing is 11.0636. It can be seen that the quality of the image after using the approximate multiply-accumulate operator for image convolution processing does not decrease greatly.
Claims
1. A high-speed, low-power approximate multiply-accumulate arithmetic unit for image convolution processing, characterized in that, This high-speed, low-power approximate multiply-accumulate arithmetic unit is used to process 8-bit signed numbers × 8-bit unsigned numbers, where the 8-bit unsigned numbers are used to process image pixels and the 8-bit signed numbers are used to process convolution kernels. For the input signed number, a sign bit is added, and the value of the sign bit is the 8th bit of the signed number. This high-speed, low-power approximate multiply-accumulate arithmetic unit is divided into four modules: a partial product generation module, a partial product compression module, a carry addition module, and an accumulation module; these modules are connected sequentially. The partial product generation module is an approximate radix-8 Booth encoder that generates a partial product matrix with a height of 3. The partial product compression module is a 3:2 compressor; the 3:2 compressor consists of 16 full adders. The first input of the first full adder is the first data of the first row of the partial product matrix, the second input of the first full adder is the first data of the second row of the partial product matrix, the third input of the first full adder is the first data of the third row of the partial product matrix, and so on. The first input of each full adder is the data of the first row, the second input is the data of the second row, and the third input is the data of the third row. Each full adder processes one column of partial product data, and the two outputs of the full adder form a matrix of height 2. The carry-add module approximates a 16-bit carry-add as a 4-bit add; The accumulation module consists of an adder and a D flip-flop, wherein the addition operation adopts a base-4 approximate addition. The approximate radix 8-Bouse encoder consists of two two-input XOR gates, one two-input XNOR gate, one two-input OR gate, one three-input OR gate, one four-input OR gate, five two-input AND gates, four three-input AND gates, two four-input AND gates, and one three-input NAND gate, in the following order: two-input first XOR gate, two-input XNOR gate, two-input first OR gate, three-input second OR gate, four-input third OR gate, two-input first AND gate, second AND gate, third AND gate, fourth AND gate, fifth AND gate, three-input sixth AND gate, seventh AND gate, eighth AND gate, ninth AND gate, four-input tenth AND gate, eleventh AND gate, and three-input NAND gate; The multiplier Y is encoded into groups of 4 bits, which are y 3i+2 y 3i+1 y 3i y 3i-1 As input, the multiplicand x i x i-1 x i-2 As input, the generated partial product PP ij For output; The first input terminal of the first XOR gate is connected to the first input signal y. 3i+2 Its second input terminal is connected to the second input signal y. 3i+1 Its output is connected to the first input of the first AND gate; The first input of the first XOR gate is connected to the third input signal y. 3i Its second input terminal is connected to the fourth input signal y. 3i-1 Its output is connected to the second input of the first AND gate; The first input of the second AND gate is connected to the output of the first AND gate, and its second input is connected to the fifth input signal x. i Its output is connected to the first input of the second OR gate; The first input terminal of the tenth AND gate is connected to the first input signal y. 3i+2 Its second input terminal is connected to the inverted second input signal y. 3i+1 Its third input terminal is connected to the third input signal y. 3i Its fourth input terminal is connected to the inverted fourth input signal y. 3i-1 Its output is connected to the first input of the third OR gate; The first input terminal of the eleventh AND gate is connected to the inverted first input signal y. 3i+2 Its second input terminal is connected to the second input signal y. 3i+1 Its third input terminal is connected to the inverted third input signal y. 3i Its fourth input terminal is connected to the fourth input signal y. 3i-1 Its output is connected to the second input of the third OR gate; The first input of the sixth AND gate is connected to the second input signal y. 3i+1 Its second input terminal is connected to the inverted third input signal y. 3i Its third input terminal is connected to the inverted third input signal y. 3i-1 Its output is connected to the third input of the third OR gate; The first input of the seventh AND gate is connected to the inverted second input signal y. 3i+1 Its second input terminal is connected to the third input signal y. 3i Its third input terminal is connected to the fourth input signal y. 3i-1 Its output is connected to the fourth input of the third OR gate; The first input of the third AND gate is connected to the output of the third OR gate, and its second input is connected to the sixth input signal x. i-1 Its output is connected to the second input of the second OR gate; The first input terminal of the eighth AND gate is connected to the inverted first input signal y. 3i+2 Its second input terminal is connected to the second input signal y. 3i+1 Its third input terminal is connected to the third input signal y. 3i Its output is connected to the first input of the first OR gate; The first input terminal of the ninth AND gate is connected to the first input signal y. 3i+2 Its second input terminal is connected to the inverted second input signal y. 3i+1 Its third input terminal is connected to the inverted third input signal y. 3i Its output is connected to the second input of the first OR gate; The first input of the fourth AND gate is connected to the seventh input signal x. i-2 Its second input is connected to the output of the first OR gate, and its output is connected to the third input of the second OR gate; The first input terminal of the NAND gate is connected to the second input signal y. 3i+1 Its second input terminal is connected to the third input signal y. 3i Its third input terminal is connected to the fourth input signal y. 3i-1 Its output is connected to the first input of the fifth AND gate; The first input of the fifth AND gate is connected to the output of the NAND gate, and its second input is connected to the first input signal y. 3i+2 Its output is connected to the second input of the second XOR gate; The first input of the second XOR gate is connected to the output of the second OR gate, its second input is connected to the output of the fifth AND gate, and its output is connected to the output signal PP. ij .
2. The high-speed, low-power approximate multiply-accumulate arithmetic unit for image convolution processing according to claim 1, characterized in that: In the approximate carry addition module, binary integers are represented as A and B, and the bits of integers A and B are represented as a. i ,b i To add two n-bit integers A and B, a generation signal g is defined at each bit. i , propagation signal p i and termination signal k i , where g i equals a i With b i p i equals a i XOR b i ;k i equals a i or not b i ; Using these signals, generate the carry output signal c for each bit i. i And used to calculate the sum of each bit s i c i and s i The recursive formula is: When k i When c equals 1, i Equals 0; when g i When c equals 1, i Equals 1; when p i When c equals 1, i equals c i-1 ; s i equals a i XOR b i XOR c i−1 ; For a 16-bit adder, its longest sequence is set to 4, s i The calculation is performed using the first 5 input bits starting from the i-th bit; that is, S5 is calculated from a0 to a5, b0 to b5, S6 is calculated from a1 to a6, b1 to b6, and so on, until S... 15 By a 10 ~a 15 , b 10 ~b 15 Calculation; a total of 11 6-bit adders are formed.
3. The high-speed, low-power approximate multiply-accumulate arithmetic unit for image convolution processing according to claim 1, characterized in that: The accumulator module consists of an approximate adder and a D flip-flop. The approximate adder is the approximate carry adder as described in claim 2. The first input terminal of the approximate adder is connected to the input signal din, and its second input terminal is connected to the output of the D flip-flop. The output terminal of the approximate adder is connected to the input of the D flip-flop. The input terminal of the D flip-flop is connected to the output dout of the approximate adder, and its output terminal is connected to the input of the approximate adder.