Memristor-based 4-2 approximate compressor, control method, device and application
By using memristor-based majority gate circuit design, the logic depth and power consumption of the 4-2 approximation compressor are simplified, achieving high-efficiency error performance and computation speed, and solving the problems of high circuit complexity and high power consumption in the prior art.
Patent Information
- Application Number
- CN202510648751.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Existing 4-2 approximation compressor circuits typically have deep logic, complex circuit structures, high power consumption, and low computation speed, making it difficult to achieve efficient error performance.
Employing a memristor-based majority gate design, including a voltage comparator and three memristors, we construct a conventional majority gate circuit with inverting logic to realize the carry and pseudo-sum results of a 4-2 approximation compressor, simplifying the logic depth and reducing power consumption.
A 4-2 compressor with good error performance, low circuit complexity, high computing speed and low power consumption, was achieved. The average error was reduced and the computing efficiency was improved by the simple circuit structure and the low power consumption characteristics of memristors.
Smart Images

Figure CN120601893B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of integrated circuit technology, and more specifically, relates to a 4-2 approximate compressor based on memristors, a control method, an apparatus, and an application. Background Technology
[0002] In recent years, approximate computation, as an effective method for optimizing circuit performance under the premise of tolerating computational errors, has been widely studied and applied in application scenarios where high precision requirements are not high, such as image blurring, video compression, and forward inference of convolutional neural networks in high-speed parallel computing scenarios. In high-speed parallel computing, the problem of accumulating multiple operands is frequently encountered, such as the compression of partial products in parallel multipliers. The 4-2 compressor is currently the most commonly used compressor module, effectively reducing the number of bits in intermediate results. Therefore, researching a 4-2 approximate compressor is of great significance.
[0003] Existing 4-2 approximation compressor circuits are mostly implemented using CMOS structures. In order to obtain better error performance, they usually rely on a large number of basic logic units such as NAND gates and NOR gates to be cascaded. The logic depth is relatively deep, the overall circuit structure is relatively complex, and frequent switching of logic gates is usually required during the operation, resulting in high power consumption and low calculation speed. Summary of the Invention
[0004] In view of the above-mentioned defects or improvement needs of the prior art, the present invention provides a 4-2 approximate compressor based on memristors, a control method, an apparatus and an application, the purpose of which is to achieve a 4-2 compressor with good error performance with lower circuit complexity, higher operation speed and lower power consumption.
[0005] To achieve the above objectives, in a first aspect, the present invention provides a memristor-based 4-2 approximation compressor for approximating compression of four 1-bit numbers X1, X2, X3, and X4, comprising: two identical first majority gate circuits and second majority gate circuits;
[0006] Both the first and second majority gate circuits include: a voltage comparator, a resistor, and three identical memristors; the resistance of the resistor is the low-resistance resistance value of the memristor; the negative terminals of the three memristors are connected to one end of the resistor, and the positive terminals of the three memristors serve as the first, second, and third input terminals of the corresponding majority gate circuit, respectively; the other end of the resistor is grounded; the first input terminal of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator outputs a first level when the voltage at its first input terminal is greater than the voltage at its second input terminal, and uses the first level as the output result "1" of the corresponding majority gate circuit; it outputs a second level when the voltage at its first input terminal is less than or equal to the voltage at its second input terminal, and uses the second level as the output result "0" of the corresponding majority gate circuit; the output terminal of the voltage comparator serves as the output terminal of the corresponding majority gate circuit.
[0007] The three memristors in the first majority gate circuit are used to store the resistance values corresponding to X1, X3, and X4 respectively; the positive terminals of the three memristors in the first majority gate circuit are all connected to voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is connected to the reference voltage V. Ref1 The output of the first majority gate is used to output the carry result of the 4-2 approximate compressor.
[0008] The three memristors in the second majority gate circuit are used to store the resistance values corresponding to X1, X2, and 1 respectively, and their positive terminals are used to connect voltages 0, V2, and V2 respectively; the second input terminal of the voltage comparator in the second majority gate circuit is used to connect the reference voltage V. Ref2 The output of the second majority gate is used to output the pseudo-sum result of the 4-2 approximate compressor.
[0009] Where 0 < V1 < V set V1 / 2 < V Ref1 <2V1 / 3; 0<V2<V set V² / 3 < V Ref2 <V2 / 2;V set The threshold for a memristor to transition from a high-resistance state to a low-resistance state.
[0010] In a second aspect, the present invention provides a control method for the 4-2 approximate compressor in the first aspect described above, comprising:
[0011] X1, X3, and X4 are written into the three memristors in the first majority gate circuit, one by one. A voltage V1 is applied to the positive terminal of each of the three memristors in the first majority gate circuit. A reference voltage V is connected to the second input terminal of the voltage comparator in the first majority gate circuit. Ref1 The carry result of the 4-2 approximate compressor is obtained at the output of the first majority gate circuit.
[0012] X1, X2, and 1 are written into the three memristors in the second majority gate circuit, one by one. A voltage of 0V is applied to the positive terminal of the memristor into which X1 is written, and a voltage of V2 is applied to the positive terminals of the memristors into which X2 and 1 are written. A reference voltage V is connected to the second input terminal of the voltage comparator in the second majority gate circuit. Ref2 At the output of the second majority gate circuit, the pseudo-sum result of the 4-2 approximate compressor is obtained;
[0013] Where 0 < V1 < V set V1 / 2 < V Ref1 <2V1 / 3; 0<V2<V set V² / 3 < V Ref2 <V2 / 2;V set The threshold for a memristor to transition from a high-resistance state to a low-resistance state.
[0014] Thirdly, the present invention provides a memristor-based 4-2 approximation compressor for approximating compression of four 1-bit numbers X1, X2, X3, and X4, comprising: two identical first majority gate circuits and second majority gate circuits;
[0015] Both the first and second majority gate circuits include: a voltage comparator, a resistor, and three identical memristors; the resistance of the resistor is the low-resistance resistance value of the memristor; the negative terminals of the three memristors are connected to one end of the resistor, and the positive terminals of the three memristors serve as the first, second, and third input terminals of the corresponding majority gate circuit, respectively; the other end of the resistor is grounded; the first input terminal of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator outputs a first level when the voltage at its first input terminal is greater than the voltage at its second input terminal, and uses the first level as the output result "1" of the corresponding majority gate circuit; it outputs a second level when the voltage at its first input terminal is less than or equal to the voltage at its second input terminal, and uses the second level as the output result "0" of the corresponding majority gate circuit; the output terminal of the voltage comparator serves as the output terminal of the corresponding majority gate circuit.
[0016] The three memristors in the first majority gate circuit are used to store the resistance values corresponding to X3, X4, and 1 respectively; the positive terminals of the three memristors in the first majority gate circuit are all connected to voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is connected to the reference voltage V. Ref1 The output of the first majority gate is used to output the carry result of the 4-2 approximate compressor.
[0017] The three memristors in the second majority gate circuit are used to store the resistance values corresponding to 0, X1, and X4 respectively, and their positive terminals are used to connect voltages 0, V2, and V3 respectively; the second input terminal of the voltage comparator in the second majority gate circuit is used to connect the reference voltage V. Ref2 The output of the second majority gate is used to output the pseudo-sum result of the 4-2 approximate compressor.
[0018] Where 0 < V1 < V set V1 / 2 < V Ref1 <2V1 / 3; 0<V2<V set V² / 3 < V Ref2 <V2 / 2;V set The threshold for a memristor to transition from a high-resistance state to a low-resistance state.
[0019] Fourthly, the present invention provides a control method for the approximate compressor 4-2 in the third aspect above, comprising:
[0020] X3, X4, and 1 are written into the three memristors in the first majority gate circuit, respectively, and a voltage V1 is applied to the positive terminal of each of the three memristors in the first majority gate circuit. A reference voltage V is connected to the second input terminal of the voltage comparator in the first majority gate circuit. Ref1 The carry result of the 4-2 approximate compressor is obtained at the output of the first majority gate circuit.
[0021] Write 0, X1, and X4 into the three memristors in the second majority gate circuit, one by one. Apply a voltage of 0V to the positive terminal of the memristor containing 0, and apply a voltage V2 to the positive terminals of the memristors containing X1 and X4. Connect the reference voltage V to the second input terminal of the voltage comparator in the second majority gate circuit. Ref2 At the output of the second majority gate circuit, the pseudo-sum result of the 4-2 approximate compressor is obtained;
[0022] Where 0 < V1 < V set V1 / 2 < V Ref1 <2V1 / 3; 0<V2<V set V² / 3 < V Ref2 <V2 / 2;V set The threshold for a memristor to transition from a high-resistance state to a low-resistance state.
[0023] Fifthly, the present invention provides a 4-2 approximate compressor device based on a memristor;
[0024] The 4-2 approximate compressor device includes the 4-2 approximate compressor provided in the first aspect of the present invention and a controller for performing the control method provided in the second aspect of the present invention;
[0025] or,
[0026] The 4-2 approximate compressor device includes the 4-2 approximate compressor provided in the third aspect of the present invention and a controller for performing the control method provided in the fourth aspect of the present invention.
[0027] Sixthly, the present invention provides an approximate multiplication operation method, comprising:
[0028] Perform a logical AND operation between each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; construct a partial product compression tree with n rows and 2n-1 columns from the n×n partial products; n≥4;
[0029] The partial product compression tree is divided into an exact compression part of n1 columns, an approximate compression part of n2 columns, and a truncated part of n3 columns; the weights of the partial products in the exact compression part, the approximate compression part, and the truncated part decrease sequentially; n1 + n2 + n3 = 2n - 1; n1, n2, and n3 are all positive integers;
[0030] The partial product in the approximate compression section is approximately compressed using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half-adder, and a full-adder; the partial product in the accurate compression section is accurately compressed using a precise 4-2 compressor, a half-adder, and a full-adder.
[0031] The partial product of the truncated portion is truncated to obtain a truncated result of n3 bits all zeros;
[0032] Arrange the partially products after precise compression as the high-order bits and the partially products after approximate compression as the low-order bits to obtain the combined result. Then, add the partially products of each column in the combined result from low to high order, and add the carry-over result to the current column when there is a carry-over in the current column to obtain the final carry-over result c0 and the sum of each column. Arrange the carry-over result c0, the sum of each column and the truncated result of n3 bits all zeros in order from high to low order to form the final approximate multiplication result.
[0033] Wherein, the first approximate 4-2 compressor is the 4-2 approximate compressor provided in the first aspect of the present invention; the second approximate 4-2 compressor is the 4-2 approximate compressor provided in the fourth aspect of the present invention.
[0034] More preferably, the partial product in the approximate compression section is approximated by using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half-adder, and a full-adder, including:
[0035] By using a first approximate 4-2 compressor, a half-adder, and a full adder to perform approximate compression on the partial product in the approximate compression part, a first approximate compression array of size 4×n2 is obtained.
[0036] The partial product in the first approximate compression array is approximated by using the second approximate 4-2 compressor to obtain a second approximate compression array of size 2×n2.
[0037] When performing compression on the part to be compressed, the compression operation is performed column by column in the order from the last column to the first column of the part to be compressed; the first column of the part to be compressed has the highest weight value and the last column has the lowest weight value; the part to be compressed is an approximate compressed part or a first approximate compressed array.
[0038] The number of first approximate 4-2 compressors used in the j-th column of the approximate compression section, p j Number of full adders r j The number of half-adders q j All satisfy g j -4p j -3r j -2q j +p j +q j +r j +m j+1 =4; m j+1 =p j+1 +q j+1 +r j+1 j=1,2,......,n2-1; 4≤g h ≤n; p h ≥0; q h ≥0; r h ≥0; h=1,2,......,n2; g h This represents the number of partial products in the h-th column of the approximate compressed portion;
[0039] When there is a compression operation in the h-th column of the approximate compression part, the h-th column of the first approximate compression array includes: the pseudo sum result obtained by compressing the h-th column of the approximate compression part;
[0040] When the h-th column of the approximate compression portion contains a partial product that does not undergo compression, the h-th column of the first approximate compression array includes: the partial product in the h-th column of the approximate compression portion that does not undergo compression.
[0041] When there is a compression operation in the (j+1)th column of the approximate compression part, the jth column of the first approximate compression array includes: the carry result after compressing the (j+1)th column of the approximate compression part;
[0042] The n2th column of the second approximate compression array includes: the pseudo-sum result after compressing the n2th column of the first approximate compression array and 0; the jth column of the second approximate compression array includes: the pseudo-sum result after compressing the jth column of the first approximate compression array and the carry result after compressing the (j+1)th column of the first approximate compression array.
[0043] More preferably, when the j-th column of the first approximate compression array contains the pseudo-sum result obtained by compressing the j-th column of the approximate compression portion, the carry result after compressing the (j+1)-th column of the approximate compression portion, and the partial product of the j-th column of the approximate compression portion without compression, the pseudo-sum result obtained by compressing the j-th column of the approximate compression portion, the carry result after compressing the (j+1)-th column of the approximate compression portion, and the partial product of the j-th column of the approximate compression portion without compression are arranged sequentially along the column direction;
[0044] When the j-th column of the first approximate compression array contains only the pseudo-sum result obtained by compressing the j-th column of the approximate compression part and the carry result obtained by compressing the (j+1)-th column of the approximate compression part, the pseudo-sum result obtained by compressing the j-th column of the approximate compression part and the carry result obtained by compressing the (j+1)-th column of the approximate compression part are arranged sequentially along the column direction.
[0045] When the j-th column of the first approximate compression array contains only the carry result after compressing the (j+1)-th column of the approximate compression part and the partial product of the j-th column of the approximate compression part that is not compressed, the carry result after compressing the (j+1)-th column of the approximate compression part and the partial product of the j-th column of the approximate compression part that is not compressed are arranged sequentially along the column direction.
[0046] When the h-th column of the first approximate compression array includes only the pseudo-sum result obtained by compressing the h-th column of the approximate compression part and the partial product of the j-th column of the approximate compression part that is not compressed, the pseudo-sum result obtained by compressing the h-th column of the approximate compression part and the partial product of the j-th column of the approximate compression part that is not compressed are arranged sequentially along the column direction.
[0047] When there are multiple compression operations in the h-th column of the approximate compression part, there are multiple pseudo sums obtained after compressing the h-th column of the approximate compression part. Their arrangement order in the h-th column of the first approximate compression array is consistent with the sorting order of the compression objects of the corresponding compression operation in the h-th column of the approximate compression part.
[0048] When there are multiple partial products in the h-th column of the approximate compression section that do not undergo compression, their arrangement order in the h-th column of the first approximate compression array is consistent with their sorting order in the h-th column of the approximate compression section.
[0049] More preferably, when the first approximate 4-2 compressor is used for compression, the compression object is four partial products arranged sequentially along the column direction in a certain column of the approximate compression part, and they are respectively used as 1-bit numbers X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression.
[0050] When the second approximation 4-2 compressor is used for compression, the compression object is four partial products arranged sequentially along the column direction in a certain column of the first approximation compression array. These are then input into the second approximation 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively.
[0051] More preferably, n=8; the n-3 columns with the highest weight in the partial product compression tree are used as the exact compression part, the n-2 columns with the second highest weight in the partial product compression tree are used as the approximate compression part, and the 4 columns with the lowest weight in the partial product compression tree are used as the truncated part.
[0052] The n-digit multiplier is a7a6a5a4a3a2a1a0; the n-digit multiplicand is b7b6b5b4b3b2b1b0;
[0053] In the partially multiplicative compressed tree, rows 1 to 8, sorted in descending order of weights, are: P 07 P 06 P 05 P 04 P 03 P 02 P 01 P 00 P 17 P 16 P 15 P 14 P 13 P 12 P 01 P 00 P 27 P 26 P 25 P 24 P 23 P 22 P 21 P 20 P 37 P 36 P 35 P 34 P 33 P 32 P 31 P 30 P 47 P 46 P 45 P 44 P43 P 42 P 41 P 40 P 57 P 56 P 55 P 54 P 53 P 52 P 51 P 50 P 67 P 66 P 65 P 64 P 63 P 62 P 61 P 60 P 77 P 76 P 75 P 74 P 73 P 72 P 71 P 70 Among them, P 07 P is the partial product of a0 and b7. 06 P is the partial product of a0 and b6, and so on. 70 It is the partial product of a7 and b0;
[0054] The approximate compression part consists of 6 columns of partial products; the partial products in the first column, from top to bottom, are: P 27 P 36 P 45 P 54 P 63 P 72 The partial products in the second column, from top to bottom, are: P 17 P 26 P 35 P 44 P 53 P 62 P 71 The partial products in the third column, from top to bottom, are: P 07 P 16 P 25 P 34 P 43 P 52 P 61 P 70 The partial products in the fourth column, from top to bottom, are: P 06 P 15 P 24 P 33 P 42 P 51 P 60The partial products in the fifth column, from top to bottom, are: P 05 P 14 P 23 P 32 P 41 P 50 The partial products in the sixth column, from top to bottom, are: P 04 P 13 P 22 P 31 P 40 The first column in the approximate compressed portion has the highest weight value, and the last column has the lowest weight value.
[0055] The above-described approximate compression operation, which utilizes a first approximate 4-2 compressor, a half-adder, and a full adder to perform approximate compression on the partial product in the approximate compression portion, yields a first approximate compression array of size 4×n², including:
[0056] Using a half-adder to P 04 and P 13 Compression is performed to obtain the pseudo-sum result S. h04 Sum of carry result C h15 ; P 05 P 14 P 23 P 32 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. c05 Sum of carry result C c26 ; P 06 P 15 P 24 P 33 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. c06 Sum of carry result C c27 Using a half-adder to P 42 and P 51 Compression is performed to obtain the pseudo-sum result S. h16 Sum of carry result C h37 ; P 07 P 16 P 25 P 34 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. c07 Sum of carry result C c28 ; P 43 P 52 P 61 P 70Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. c17 Sum of carry result C c38 ; P 17 P 26 P 35 P 44 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. c08 Sum of carry result C c29 Using a full adder to P 62 and P 71 Compression is performed to obtain the result S. f18 Sum of carry result C f39 ; P 27 P 36 P 45 P 54 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. c09 Sum of carry result C c110 Using a half-adder to P 63 and P 72 Compression is performed to obtain the pseudo-sum result S. h19 Sum of carry result C h210 ; P 37 P 46 P 55 P 64 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. e010 Sum of carry result C f012 ;
[0057] Therefore, the first column in the first approximate compressed array is S. c09 S h19 C c29 C f39 The second column is S c08 S h18 C c28 C f38 The third column is S c07 S c17 C c27 C h37 The fourth column is S c06 S h16 C c26 P 60 The fifth column is S c05 C h15 P41 P 50 The sixth column is S h04 P 22 P 31 P 40 ;
[0058] The above-described approximate compression operation, which uses a second approximate 4-2 compressor to perform approximate compression on partial products in the first approximate compression array, yields a second approximate compression array of size 2×n², including:
[0059] S h04 P 22 P 31 P 40 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l4 Sum of carry result C l4 ; will S c05 C h15 P 41 P 50 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l5 Sum of carry result C l5 ; will S c06 S h16 C c26 P 60 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l6 Sum of carry result C l6 ;S c07 S c17 C c27 C h37 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l7 Sum of carry result C l7 ;S c08 S f18 C c28 C c38 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l8 Sum of carry result C l8 ; will S c09 S h19 C c29 C f39Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l9 Sum of carry result C l9 ;
[0060] This results in the first column of the second approximate compressed array including S. l9 and C l8 The second column includes S l8 and C l7 The third column includes S l7 and C l6 The fourth column includes S l6 and C l5 The fifth column includes S l5 and C l4 The sixth column includes S l4 And 0.
[0061] In a seventh aspect, the present invention provides an approximate multiplication apparatus, comprising: a partial product generation module, an exact partial compression module, an approximate partial compression module, a truncation processing module, and a carry adder module;
[0062] The partial product generation module performs a logical AND operation between each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; these n×n partial products are then used to construct a partial product compression tree with n rows and 2n-1 columns; n≥4; the partial product compression tree is divided into an exact compression part with n1 columns, an approximate compression part with n2 columns, and a truncated part with n3 columns; the weights of the partial products in the exact compression part, approximate compression part, and truncated part decrease sequentially; n1+n2+n3=2n-1; n1, n2, and n3 are all positive integers;
[0063] The precise partial compression module is used to approximate the partial product in the approximate compression section using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder; and to precisely compress the partial product in the precise compression section using a precise 4-2 compressor, a half adder, and a full adder.
[0064] The truncation module is used to truncate the partial product of the truncated part to obtain a truncated result of n3 bits all zero;
[0065] The carry adder module is used to arrange the partially products after precise compression as the high-order bits and the partially products after approximate compression as the low-order bits to obtain the combined result. The partial products of each column in the combined result are added from the low-order bits to the high-order bits. When there is a carry in the current column, the carry result is added to the current column to obtain the final carry result c0 and the sum of each column. The carry result c0, the sum of each column and the truncated result of n3 bits all zeros are arranged in order from the high-order bits to the low-order bits to form the final approximate multiplication result.
[0066] Wherein, the first approximate 4-2 compressor is the 4-2 approximate compressor provided in the first aspect of the present invention; the second approximate 4-2 compressor is the 4-2 approximate compressor provided in the fourth aspect of the present invention.
[0067] In summary, the above-described technical solutions conceived in this invention can achieve the following beneficial effects:
[0068] 1. The first aspect of this invention provides a 4-2 approximation compressor based on memristors, which approximates and compresses four 1-bit numbers X1, X2, X3, and X4 using two identical memristor-based majority gate circuits; wherein, the majority gate circuit includes a voltage comparator, a resistor, and three identical memristors; the two majority gate circuits respectively implement conventional majority gate logic and a novel majority gate logic including inversion logic; the first majority gate circuit is used to implement the conventional majority gate logic M(X1, X3, X4) to obtain the carry result of the 4-2 approximation compressor; the second majority gate circuit is used to implement the novel majority gate logic including inversion logic. The pseudo-sum result of the 4-2 approximate compressor is obtained. This invention does not rely on a large number of cascaded basic logic gates, has a shallow logic depth, a simple circuit structure, and low latency. Combined with the low power consumption characteristics of memristors, it maintains good error performance while achieving high-efficiency compression. In particular, the input dominant value is retained in the input design of the first majority gate circuit and the second majority gate circuit, which helps to reduce the average error. Based on this, this invention can achieve a 4-2 compressor with good error performance with low circuit complexity, high operation speed and low power consumption.
[0069] 2. A third aspect of the present invention provides a 4-2 approximation compressor based on memristors, which approximates and compresses four 1-bit numbers X1, X2, X3, and X4 using two identical memristor-based majority gate circuits; wherein, the majority gate circuit includes a voltage comparator, a resistor, and three identical memristors; the two majority gate circuits respectively implement conventional majority gate logic and a novel majority gate logic including inversion logic; the first majority gate circuit is used to implement the conventional majority gate logic M(X3,X4,1) to obtain the carry result of the 4-2 approximation compressor; the second majority gate circuit is used to implement the novel majority gate logic including inversion logic. The pseudo-sum result of the 4-2 approximation compressor is obtained. This invention does not rely on a large number of cascaded basic logic gates, has a shallow logic depth, simple device composition, low latency, and low power consumption. By introducing fixed values "1" and "0" at the input, a structural logic bias is effectively constructed, giving the pseudo-sum result a clear error direction, which is beneficial for targeted error guidance and neutralization in subsequent compression. Therefore, this invention not only has low hardware complexity but also exhibits better error equalization capability in a controllable error structure. Based on this, this invention can achieve a 4-2 compressor with good error performance with low circuit complexity, high operation speed, and low power consumption.
[0070] 3. The sixth aspect of this invention provides an approximate multiplication method, which utilizes the first approximate 4-2 compressor provided in the first aspect of this invention, the approximate 4-2 compressor provided in the third aspect of this invention, a half adder, and a full adder to approximate the partial product in the approximate compression part. The first approximate 4-2 compressor omits some logic inputs while maintaining the main compression path, resulting in a simple structure, low area and power consumption overhead, and is suitable for compression paths of intermediate weight columns where high accuracy is not required; its expected error is negative. The second approximate 4-2 compressor is an error-guided approximate compressor that introduces a fixed input "1" to construct an output bias, making the output error directional and adjustable; it is suitable for compression paths of approximate multipliers where error guidance or error mean control is required, and its expected error is positive. By utilizing 4-2 approximate compressors with different expected error directions, an approximate multiplication method with controllable error, reconfigurable structure, simple structure, high operation speed, and good resource utilization can be achieved.
[0071] 4. Further, the approximate multiplication method provided by this invention divides the approximate compression of the partial product in the approximate compression part into two stages. In the first stage, a first approximate 4-2 compressor, a half-adder, and a full adder are used to perform approximate compression on the partial product in the approximate compression part to obtain a first approximate compression array of size 4×n2. In the second stage, a second approximate 4-2 compressor is used to perform approximate compression on the partial product in the first approximate compression array to obtain a second approximate compression array of size 2×n2. The compression in the first stage results in a first approximate compression array with 4 rows, which allows the second stage to use only the second approximate 4-2 compressor for compression. This facilitates the adjustment and compensation of the error direction of the output result of the first stage in the second stage. Since the pseudo-sum result output by the first approximate 4-2 compressor has a positive error distribution, and the second approximate 4-2 compressor has the characteristic of structurally guiding positive errors, applying them together in the second stage can achieve complementary cancellation of positive and negative errors between columns, effectively balancing the mean error of the overall multiplication operation output, making the expected error approach zero, and further reducing the error of the multiplication operation.
[0072] 5. Furthermore, the approximate multiplication method provided by the present invention is preferably applicable to the multiplication of an 8-bit multiplier and an 8-bit multiplicand. By reasonably setting the positions of the approximate 4-2 compressor, half adder and full adder in different stages of the compression process of the approximate compression part, the positive and negative errors can be mutually canceled between columns to the greatest extent possible, thereby effectively balancing the mean error of the overall multiplication result, making the output error expected to approach zero, and further reducing the error of the multiplication operation. Attached Figure Description
[0073] Figure 1 This is a schematic diagram of the 4-2 approximate compressor provided in Embodiment 1 of the present invention.
[0074] Figure 2 This is a schematic diagram of the structure of most gate circuits provided in Embodiment 1 of the present invention.
[0075] Figure 3 This is a schematic diagram of the equivalent circuit after writing logic values to the memristor according to Embodiment 1 of the present invention.
[0076] Figure 4 This is a schematic diagram of the 4-2 approximate compressor provided in Embodiment 2 of the present invention.
[0077] Figure 5 This is a schematic diagram of the gate-level circuit of the half-adder provided in Embodiment 3 of the present invention.
[0078] Figure 6 This is a schematic diagram of the gate-level circuit of the full adder provided in Embodiment 3 of the present invention.
[0079] Figure 7 This is a schematic diagram of the gate-level circuit of the precise 4-2 compressor provided in Embodiment 3 of the present invention.
[0080] Figure 8 This is a schematic diagram of the first stage of the approximate multiplication operation provided in Embodiment 3 of the present invention.
[0081] Figure 9 This is a schematic diagram representing the devices involved in the approximate multiplication operation provided in Embodiment 3 of the present invention.
[0082] Figure 10 This is a schematic diagram of stage two of the approximate multiplication operation provided in Embodiment 3 of the present invention.
[0083] Figure 11 This is a schematic diagram of stage three and stage four of the approximate multiplication operation provided in Embodiment 3 of the present invention. Detailed Implementation
[0084] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0085] Example 1
[0086] This embodiment provides a memristor-based 4-2 approximation compressor for approximating the compression of four 1-bit numbers X1, X2, X3, and X4, including: two identical first majority gate circuits and second majority gate circuits;
[0087] Both the first and second majority gate circuits include: a voltage comparator, a resistor, and three identical memristors; the resistance of the resistor is the low-resistance resistance value of the memristor; the negative terminals of the three memristors are connected to one end of the resistor, and the positive terminals of the three memristors serve as the first, second, and third input terminals of the corresponding majority gate circuit, respectively; the other end of the resistor is grounded; the first input terminal of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator outputs a first level when the voltage at its first input terminal is greater than the voltage at its second input terminal, and uses the first level as the output result "1" of the corresponding majority gate circuit; it outputs a second level when the voltage at its first input terminal is less than or equal to the voltage at its second input terminal, and uses the second level as the output result "0" of the corresponding majority gate circuit; the output terminal of the voltage comparator serves as the output terminal of the corresponding majority gate circuit.
[0088] The three memristors in the first majority gate circuit are used to store the resistance values corresponding to X1, X3, and X4 respectively; the positive terminals of the three memristors in the first majority gate circuit are all connected to voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is connected to the reference voltage V. Ref1 The output of the first majority gate is used to output the carry result of the 4-2 approximate compressor.
[0089] The three memristors in the second majority gate circuit are used to store the resistance values corresponding to X1, X2, and 1 respectively. The positive terminal of the memristor storing the resistance value corresponding to X1 is connected to voltage 0, the positive terminal of the memristor storing the resistance value corresponding to X2 is connected to voltage V2, and the positive terminal of the memristor storing the resistance value corresponding to 1 is connected to voltage V2. The second input terminal of the voltage comparator in the second majority gate circuit is connected to the reference voltage V. Ref2 The output of the second majority gate is used to output the pseudo-sum result of the 4-2 approximate compressor.
[0090] Where 0 < V1 < V setV1 / 2 < V Ref1 <2V1 / 3; 0<V2<V set V² / 3 < V Ref2 <V2 / 2;V set The threshold for a memristor to transition from a high-resistance state to a low-resistance state.
[0091] Understandably, the correspondence between the voltage comparator's output level and its logic value can be flexibly set. For example, with the first input terminal as positive and the second input terminal as negative, the first level can be high and the second level can be low, with the high level representing the result of the majority gate logic operation as "1" and the low level representing the result of the majority gate logic operation as "0". Alternatively, with the first input terminal as negative and the second input terminal as positive, the first level can be low and the second level can be high, with the low level representing the result of the majority gate logic operation as "1" and the high level representing the result of the majority gate logic operation as "0".
[0092] In this embodiment, the voltage comparator uses the first input terminal as the positive terminal and the second input terminal as the negative terminal. The first level is high level and the second level is low level. The high level is used as the result of the majority gate logic operation "1", and the low level is used as the result of the majority gate logic operation "0".
[0093] A schematic diagram of the 4-2 approximate compressor in this embodiment is shown below. Figure 1 As shown. The 4-2 approximation compressor in this embodiment consists of a majority gate (MAJ) and a majority gate with an inverted input (MAJF); the majority gate (MAJ) is the first majority gate circuit in this embodiment, and its corresponding logical operation is M(X1,X3,X4); the majority gate with an inverted input (MAJF) is the second majority gate circuit in this embodiment, and its corresponding logical operation is... The carry result (Carry signal) is output by the first majority gate circuit, and the pseudo-sum result (Sum signal) is output by the second majority gate circuit with an inverted input.
[0094] The first majority gate circuit and the second majority gate circuit have the same structure. The majority gate circuit in this embodiment is as follows: Figure 2 As shown, the circuit includes three memristors (M1, M2, and M3), a fixed resistor, and a voltage comparator. The positive terminals of memristors M1, M2, and M3 are controlled via control terminals T1, T2, and T3 respectively. One end of the fixed resistor is controlled via control terminal T4. The other end of the fixed resistor is connected to the negative terminals of memristors M1, M2, and M3. The voltage at this common node is represented by V. Com This indicates that the control voltages for the other four control ports T1, T2, T3, and T4 are respectively V. T1 V T2 V T3 and V T4The positive terminal of the voltage comparator is connected to the common node of the memristor and the series resistor, and the negative terminal is connected to the reference voltage V. Ref The output voltage is V o .
[0095] A memristor consists of two resistive states: a high-resistance state (HRS) and a low-resistance state (LRS). By applying a port voltage of a specific direction and magnitude, the memristor can transition between these two resistive states. Specifically, by applying a voltage greater than V... set A positive voltage can switch a memristor from a high-resistance state to a low-resistance state; conversely, by applying a voltage less than V... reset A negative voltage allows a memristor to transition from a low-resistance state to a high-resistance state. V set V is the threshold voltage at which the memristor transitions from a high-resistance state to a low-resistance state. reset It is the threshold voltage at which a memristor transitions from a low-resistance state to a high-resistance state.
[0096] The high and low configuration resistance values of the memristor are R and R, respectively. H and R L R H < <R L In this embodiment, the resistance value is equal to the low-resistance value of the memristor. The high-resistance R of the memristor... H =100KΩ, low resistance R L When the resistance is 1KΩ, the resistance R = 1KΩ.
[0097] The three input variables a, b, and c of the majority gate logic are mapped to the resistance states of memristors M1, M2, and M3, respectively. The high resistance state of the memristor corresponds to the logic value "0", and the low resistance state corresponds to the logic value "1". The output of the logic is the voltage value at the comparator output port, where the output V... o = V+ indicates that the output is high, corresponding to a logic result of "1"; output V o = V- indicates that the output is low, and the corresponding logic result is "0".
[0098] Based on the above logic circuit, only the control voltages applied to control ports T1, T2, T3, and T4 need to be changed to realize M(a,b,c) and Two types of majority gate logic. The specific implementation steps are as follows:
[0099] (1) Implementation of the majority gate logic M(a,b,c):
[0100] By applying voltages V1, V1, V1, 0 to control ports T1, T2, T3, and T4 respectively, V Ref = V Ref1The result of the logic operation is obtained through the output of the voltage comparator, and the port voltage V1 is compared with the voltage comparator reference voltage V. Ref The following constraints must be met: 0 < V1 < V set V1 / 2 < V Ref1 <2V1 / 3; The value of V1 must ensure that it does not change the logic value currently written to the memristor, while V Ref1 The value of must be chosen to ensure that the voltage comparator can output the correct calculation result.
[0101] The principle behind this logic implementation is that when at least two of the input memristors M1, M2, and M3 are in a low-resistance state, the voltage V at the common node... Com Greater than the comparator reference voltage V Ref1 This results in a high-level output "1"; conversely, when at most one of the input memristors M1, M2, and M3 is in a low-impedance state, the voltage V at the common node... Com Less than the comparator reference voltage V Ref1 This outputs a low level "0". This enables the implementation of a majority gate logic M(a,b,c).
[0102] by Figure 3 Taking the equivalent circuit after writing logic values to the memristors as an example, we can further illustrate how the above method can yield correct logical operation results. Other combinations follow the same logic. In the second combination, we input a logic value "1" and two logic values "0", writing them respectively to memristors M1, M2, and M3. The memristor corresponding to the logic value "1" is in a low-resistance state, and the memristor corresponding to the logic value "0" is in a high-resistance state. Therefore, the equivalent circuit after writing the logic values is as follows: Figure 3 As shown, at this time, the voltage V of the common node is... Com for:
[0103]
[0104] That is, the voltage V of the common node Com It is less than and close to V1 / 2, while V1 / 2 < V Ref1 Therefore, voltage V Com Less than V Ref1 Therefore, the voltage comparator outputs a low level, indicating a logic value of "0", and the result of the majority gate logic operation is correct.
[0105] (2) Implementation of majority gate logic:
[0106] By applying voltages 0, V2, V2, 0, and V to control ports T1, T2, T3, and T4 respectively. Ref = V Ref2 The result of the logic operation is obtained through the output of the voltage comparator, where the port voltage V2 is compared with the comparator reference voltage V. Ref2The following constraint must be satisfied: 0 < V2 < V set V² / 3 < V Ref2 <V2 / 2, the value of V2 must ensure that it does not change the logic value currently written to the memristor, while V Ref2 The value of must be chosen to ensure that the voltage comparator can output the correct calculation result.
[0107] The principle behind this logic implementation is that when both input memristors M2 and M3 are low-impedance "1", they can pull up the voltage V of the common node. Com The voltage V of the common node Com Greater than the comparator reference voltage V Ref2 This outputs a high level "1"; however, when input memristors M2 and M3 have high impedance "0" and input memristor M1 has low impedance "1", it will pull down the voltage V of the common node. Com The voltage V of the common node Com Less than the comparator reference voltage V Ref2 This outputs a low level "0". This is used to implement majority gate logic. .
[0108] This embodiment also provides a control method for the above-mentioned 4-2 approximate compressor, including:
[0109] X1, X3, and X4 are written into the three memristors in the first majority gate circuit, one by one. A voltage V1 is applied to the positive terminal of each of the three memristors in the first majority gate circuit. A reference voltage V is connected to the second input terminal of the voltage comparator in the first majority gate circuit. Ref1 The carry result of the 4-2 approximate compressor is obtained at the output of the first majority gate circuit.
[0110] X1, X2, and 1 are written into the three memristors in the second majority gate circuit, one by one. A voltage of 0V is applied to the positive terminal of the memristor into which X1 is written, and a voltage of V2 is applied to the positive terminals of the memristors into which X2 and 1 are written. A reference voltage V is connected to the second input terminal of the voltage comparator in the second majority gate circuit. Ref2 At the output of the second majority gate circuit, the pseudo-sum result of the 4-2 approximate compressor is obtained;
[0111] Where 0 < V1 < V set V1 / 2 < V Ref1 <2V1 / 3; 0<V2<V set V² / 3 < V Ref2 <V2 / 2;V set The threshold for a memristor to transition from a high-resistance state to a low-resistance state.
[0112] This embodiment also provides a memristor-based 4-2 approximation compressor device, including the 4-2 approximation compressor and controller provided in this embodiment;
[0113] The controller is used to execute the control method of the approximate compressor described in 4-2 in this embodiment.
[0114] Example 2
[0115] like Figure 4 As shown, this embodiment provides a 4-2 approximation compressor based on memristors for approximating compression of four 1-bit numbers X1, X2, X3, and X4, including: two identical first majority gate circuits and second majority gate circuits;
[0116] Both the first and second majority gate circuits include: a voltage comparator, a resistor, and three identical memristors; the resistance of the resistor is the low-resistance resistance value of the memristor; the negative terminals of the three memristors are connected to one end of the resistor, and the positive terminals of the three memristors serve as the first, second, and third input terminals of the corresponding majority gate circuit, respectively; the other end of the resistor is grounded; the first input terminal of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator outputs a first level when the voltage at its first input terminal is greater than the voltage at its second input terminal, and uses the first level as the output result "1" of the corresponding majority gate circuit; it outputs a second level when the voltage at its first input terminal is less than or equal to the voltage at its second input terminal, and uses the second level as the output result "0" of the corresponding majority gate circuit; the output terminal of the voltage comparator serves as the output terminal of the corresponding majority gate circuit.
[0117] The three memristors in the first majority gate circuit are used to store the resistance values corresponding to X3, X4, and 1 respectively; the positive terminals of the three memristors in the first majority gate circuit are all connected to voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is connected to the reference voltage V. Ref1 The output of the first majority gate is used to output the carry result of the 4-2 approximate compressor.
[0118] The three memristors in the second majority gate circuit are used to store the resistance values corresponding to 0, X1, and X4 respectively, and their positive terminals are used to connect voltages 0, V2, and V3 respectively; the second input terminal of the voltage comparator in the second majority gate circuit is used to connect the reference voltage V. Ref2 The output of the second majority gate is used to output the pseudo-sum result of the 4-2 approximate compressor.
[0119] Where 0 < V1 < V set V1 / 2 < V Ref1 <2V1 / 3; 0<V2<V set V² / 3 < VRef2 <V2 / 2;V set The threshold for a memristor to transition from a high-resistance state to a low-resistance state.
[0120] Similar to Example 1, the 4-2 approximation compressor in this example also consists of a majority gate (MAJ) and a majority gate with an inverted input (MAJF). The majority gate (MAJ) is the first majority gate circuit in this example, and its corresponding logical operation is M(X1,X3,X4). The majority gate with an inverted input (MAJF) is the second majority gate circuit in this example, and its corresponding logical operation is... The carry result (Carry signal) is output by the first majority gate circuit, and the pseudo-sum result (Sum signal) is output by the second majority gate circuit with an inverted input.
[0121] Unlike the 4-2 approximation compressor in Example 1, which employs a fully data-driven majority logic design, the 4-2 approximation compressor provided in this example effectively constructs a structured logic bias by introducing fixed values "1" and "0" at the input. This gives the pseudo-sum result a clear error direction, which is beneficial for targeted error guidance and neutralization in subsequent compression. When used for multiplication operations, it can optimize the overall error distribution of the multiplication operation output. Therefore, this design not only has low hardware complexity but also exhibits better error equalization capabilities within a controllable error structure.
[0122] The truth tables for the 4-2 approximate compressor (compressor one) in Example 1 and the 4-2 approximate compressor (compressor two) in Example 2 are shown in Table 1.
[0123]
[0124] Most of the gate circuit structures and detailed descriptions in this embodiment are the same as in Embodiment 1, and will not be repeated here.
[0125] This embodiment also provides a control method for the above-mentioned 4-2 approximate compressor, including:
[0126] X3, X4, and 1 are written into the three memristors in the first majority gate circuit, respectively, and a voltage V1 is applied to the positive terminal of each of the three memristors in the first majority gate circuit. A reference voltage V is connected to the second input terminal of the voltage comparator in the first majority gate circuit. Ref1 The carry result of the 4-2 approximate compressor is obtained at the output of the first majority gate circuit.
[0127] Write 0, X1, and X4 into the three memristors in the second majority gate circuit, one by one. Apply a voltage of 0V to the positive terminal of the memristor containing 0, and apply a voltage V2 to the positive terminals of the memristors containing X1 and X4. Connect the reference voltage V to the second input terminal of the voltage comparator in the second majority gate circuit. Ref2 At the output of the second majority gate circuit, the pseudo-sum result of the 4-2 approximate compressor is obtained;
[0128] Where 0 < V1 < V set V1 / 2 < V Ref1 <2V1 / 3; 0<V2<V set V² / 3 < V Ref2 <V2 / 2;V set The threshold for a memristor to transition from a high-resistance state to a low-resistance state.
[0129] This embodiment also provides a memristor-based 4-2 approximation compressor device, including the 4-2 approximation compressor and controller provided in this embodiment;
[0130] The controller is used to execute the control method of the approximate compressor described in 4-2 in this embodiment.
[0131] Example 3
[0132] As a core computational module in digital systems, the performance of multipliers directly affects the overall data processing speed, energy consumption, and system throughput. Traditional multiplier designs are typically based on partial product reduction strategies, relying on numerous basic units such as full adders (FA), half adders (HA), and 4:2 compressors to compress and merge multi-bit partial products step by step. However, in traditional CMOS processes, the compression network of multipliers usually requires deep logic gate cascading and multi-stage carry propagation, leading to severe latency accumulation, significantly increased power consumption, and exponential growth in hardware resource consumption with large bit width expansion, making it difficult to meet the requirements of low-power, highly integrated computing systems.
[0133] This embodiment provides an approximate multiplication method, including:
[0134] Perform a logical AND operation between each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; construct a partial product compression tree with n rows and 2n-1 columns from the n×n partial products; n≥4;
[0135] The partial product compression tree is divided into an exact compression part of n1 columns, an approximate compression part of n2 columns, and a truncated part of n3 columns; the weights of the partial products in the exact compression part, the approximate compression part, and the truncated part decrease sequentially; n1 + n2 + n3 = 2n - 1; n1, n2, and n3 are all positive integers;
[0136] The partial product in the approximate compression section is approximately compressed using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half-adder, and a full-adder; the partial product in the accurate compression section is accurately compressed using a precise 4-2 compressor, a half-adder, and a full-adder.
[0137] The partial product of the truncated portion is truncated to obtain a truncated result of n3 bits all zeros;
[0138] Arrange the partially products after precise compression as the high-order bits and the partially products after approximate compression as the low-order bits to obtain the combined result. Then, add the partially products of each column in the combined result from low to high order, and add the carry-over result to the current column when there is a carry-over in the current column to obtain the final carry-over result c0 and the sum of each column. Arrange the carry-over result c0, the sum of each column and the truncated result of n3 bits all zeros in order from high to low order to form the final approximate multiplication result.
[0139] The first approximate 4-2 compressor is the 4-2 approximate compressor provided in Embodiment 1 of this invention; the second approximate 4-2 compressor is the 4-2 approximate compressor provided in Embodiment 2 of this invention. The first approximate 4-2 compressor omits some logic inputs while maintaining the main compression path, resulting in a simple structure and low area and power consumption overhead, making it suitable for intermediate weight column compression paths where high accuracy is not required. The second approximate 4-2 compressor is an error-guided approximate compressor. By introducing a fixed input "1" to construct the output bias, it makes the output error directional and adjustable, making it suitable for approximate multiplier compression paths where error guidance or error mean control is required. This embodiment, through the aforementioned first and second approximate 4-2 compressors, constructs a multi-stage compression network for the multiplier, achieving an approximate multiplication operation scheme with controllable error, reconfigurable structure, and good resource utilization.
[0140] The relevant technical solutions are the same as those in Embodiments 1 and 2 of this invention, and will not be repeated here.
[0141] Furthermore, in a preferred embodiment of this example, the partial product in the approximate compression portion is approximated using a first approximation 4-2 compressor, a second approximation 4-2 compressor, a half-adder, and a full-adder, including:
[0142] By using a first approximate 4-2 compressor, a half-adder, and a full adder to perform approximate compression on the partial product in the approximate compression part, a first approximate compression array of size 4×n2 is obtained.
[0143] The partial product in the first approximate compression array is approximated by using the second approximate 4-2 compressor to obtain a second approximate compression array of size 2×n2.
[0144] When performing compression on the part to be compressed, the compression operation is performed column by column in the order from the last column to the first column of the part to be compressed; the first column of the part to be compressed has the highest weight value and the last column has the lowest weight value; the part to be compressed is an approximate compressed part or a first approximate compressed array.
[0145] The number of first approximate 4-2 compressors used in the j-th column of the approximate compression section, p j Number of full adders r j The number of half-adders q j All satisfy g j -4p j -3r j -2q j +p j +q j +r j +m j+1 =4; m j+1 =p j+1 +q j+1 +r j+1 j=1,2,......,n2-1; 4≤g h ≤n; p h ≥0; q h ≥0; r h ≥0; h=1,2,......,n2; g h This represents the number of partial products in the h-th column of the approximate compressed portion;
[0146] When there is a compression operation in the h-th column of the approximate compression part, the h-th column of the first approximate compression array includes: the pseudo sum result obtained by compressing the h-th column of the approximate compression part;
[0147] When the h-th column of the approximate compression portion contains a partial product that does not undergo compression, the h-th column of the first approximate compression array includes: the partial product in the h-th column of the approximate compression portion that does not undergo compression.
[0148] When there is a compression operation in the (j+1)th column of the approximate compression part, the jth column of the first approximate compression array includes: the carry result after compressing the (j+1)th column of the approximate compression part;
[0149] The n2th column of the second approximate compression array includes: the pseudo-sum result after compressing the n2th column of the first approximate compression array and 0; the jth column of the second approximate compression array includes: the pseudo-sum result after compressing the jth column of the first approximate compression array and the carry result after compressing the (j+1)th column of the first approximate compression array.
[0150] The above design divides the approximate compression of the partial product in the approximate compression part into two stages. The compression in the first stage results in a first approximate compression array with 4 rows, which enables the second stage to use only the second approximate 4-2 compressor for compression, further reducing the error of the multiplication operation.
[0151] It should be noted that the precise 4-2 compressor can be any existing precise 4-2 compressor, such as the 4-2 compressor structure based on an optimized CMOS full adder proposed in the paper "Improved CMOS (4;2) compressor designs for parallel multipliers," and is not limited here. The half adder can be any existing half adder, such as the half adder design based on CMOS technology proposed in the paper "CMOS Half Adder Design & Simulation Using Different Foundry," and is not limited here. The full adder can be any existing full adder, such as the full adder design based on CMOS technology proposed in the paper "Design and Analysis of CMOS Full Adder," and is not limited here.
[0152] In this embodiment, the precise 4-2 compressor, half-adder, and full-adder are all designed based on majority gate circuits.
[0153] like Figure 5 As shown, the half-adder is used to process three 1-bit numbers a i b i c i Approximate compression is performed, including four identical first majority gate circuits, second majority gate circuits, third majority gate circuits, and fourth majority gate circuits; the first majority gate circuit, second majority gate circuit, third majority gate circuit, and fourth majority gate circuit are all the same as the majority gate circuits in Embodiment 1; the first majority gate circuit is used to implement the logic operation M(a i , b i The first gate (0) outputs the carry result of the half-adder (Carry signal); the second majority gate is used to implement logical operations. The third majority gate circuit is used to implement logic operations. The outputs of the second and third majority gates are both used as inputs to the fourth majority gate, which is used to implement logical operations. The output of the half-adder is the pseudo-sum result (Sum signal).
[0154] The aforementioned half-adder can directly perform addition without introducing traditional inverters and XOR gates, reducing the number of logic levels and devices, which is beneficial for improving computing speed and hardware integration.
[0155] like Figure 6 As shown, the full adder is used to process three 1-bit numbers a i b i c i Approximate compression is performed, including three identical first majority gate circuits, second majority gate circuits, and third majority gate circuits; the first majority gate circuit, second majority gate circuit, and third majority gate circuit are all the same as the majority gate circuits in Embodiment 1; the first majority gate circuit is used to implement the logical operation M(a i ,b i , c i The second majority gate circuit is used to implement logic operations, and the carry result (Carry signal) of the full adder is obtained. The outputs of the first and second majority gates are used as inputs to the fourth majority gate, and the third majority gate is used to implement logical operations. This yields the pseudo-sum result (Sum signal) of the full adder.
[0156] The above-mentioned full adder has a compact logic path and can generate carry and summation results simultaneously, reducing the delay caused by carry propagation in traditional full adders and improving the overall operation speed. It is particularly suitable for use in in-memory computing or low-power systems that require efficient addition calculations.
[0157] The above-mentioned half-adder and full-adder structures can be efficiently implemented in most logic platforms. They have shallow logic levels and low resource overhead, and are suitable for embedded designs with various compressor structures.
[0158] like Figure 7 As shown, the Accurate 4:2 Compressor consists of two cascaded full adders; the first-stage full adder takes X1, X2, and X3 as inputs and outputs the intermediate carry (Carry1) and sum (Sum1); the second-stage full adder takes X4, C... inSum1 is the input, and the outputs are the final carry-in Carry2 and the sum-in Sum2. Carry1 serves as the carry-out result (Cout signal) of the precise 4:2 compressor, while Carry2 and Sum2 serve as the carry-out result (Carry signal) and the sum-in result (Sum signal) of the precise 4:2 compressor, respectively. Through this two-stage full adder cascaded design, this precise 4:2 compressor can accurately compress four input bits and one carry-in bit, outputting two sum-in bits and one carry-in bit. Its logic function is consistent with that of a traditional CMOS-implemented 4:2 compressor. Furthermore, because it is built using majority gates and majority gates with inverted inputs, it further optimizes the logic depth and hardware area, helping to reduce latency and improve energy efficiency. It is suitable for application in high-performance approximate multipliers or partial product compression networks. It should be noted that the full adder used in the precise 4:2 compressor can be any existing full adder, preferably the above-mentioned full adder based on majority gates. This structure is functionally equivalent to the traditional CMOS precise 4:2 compressor, has precise compression function, and is implemented based on majority gate logic, with good integrability and logic consistency.
[0159] This embodiment uses 8-bit binary approximate multiplication as an example for detailed description, as follows:
[0160] like Figure 8 , Figure 10 and Figure 11 The diagram shown illustrates the working principle of the approximate multiplication operation provided in this embodiment; where, as Figure 8 This is a schematic diagram of stage one of the approximate multiplication operation; as shown below. Figure 9 A schematic diagram representing the devices involved in approximate multiplication operations; such as Figure 10 This is a schematic diagram of stage two of the approximate multiplication operation; as shown below. Figure 11 This is a schematic diagram of stages three and four of the approximate multiplication operation.
[0161] The entire multiplier operation is divided into three stages. The first two stages involve the use of an approximate compressor. To control the multiplier error within a certain range, the partial product calculation of the higher-order bits uses an accurate compressor to approximate the partial product compression. Similarly, to reduce hardware usage, the partial product of the lower-order bits is ignored. Only the partial product of the middle-order bits is approximated and compressed.
[0162] The first-stage approximation uses only the aforementioned first 4-2 approximation compressor, and the second-stage approximation uses only the aforementioned second approximation 4-2 compressor. The source of the operator in each stage of the approximation region is also marked on the figure. Hollow circles indicate ignored operators or operators filled with 0s. P indicates that the operator is the result of directly multiplying two multipliers, and S... h This represents Sum, C of a half-adder.h This indicates that S is the Carry of the half-adder. f Sum, C represents the sum of the full adders. f This indicates that S is the Carry of the full adder. c This indicates that Sum, C is an approximate compressor one. c This indicates that Carry,S is an approximate compressor one. e Sum,C represents the precision compressor. e This refers to the Carry, C of the precision compressor. o It is the C of the precision compressor out S l C l These are Sum and Carry, which are the approximate compressors used in the second stage.
[0163] Perform a logical AND operation between each bit of the 8-bit multiplier a7a6a5a4a3a2a1a0 and each bit of the 8-bit multiplicand b7b6b5b4b3b2b1b0 to obtain 64 partial products. Then, construct an 8-row, 15-column partial product compressed tree, as shown below. Figure 8 As shown. In the partially compressed tree, the first row, from the highest digit (highest weight value) to the lowest digit (lowest weight value), is P. 07 P 06 P 05 P 04 P 03 P 02 P 01 P 00 Similarly, in the last row of the partial product compressed tree, the values from the highest digit (highest weight value) to the lowest digit (lowest weight value) are P. 77 P 76 P 75 P 74 P 73 P 72 P 71 P 70 Among them, P 07 P is the partial product of a0 and b7 after a logical AND operation; 00 P is the partial product of a0 and b0 (AND logical operation); 70 P is the partial product of a7 and b0 (AND logical operation); 77 It is the partial product of a7 and b7 after multiplication (AND logical operation).
[0164] For phase one, P 04 P 13 As input to the half-adder, we obtain S h04 C h15 P 05 P 14 P23 P 32 As input to the first 4-2 approximate compressor, S is obtained. c05 C c26 P 06 P 15 P 24 P 33 As input to the first 4-2 approximate compressor, S is obtained. c06 C c27 P 42 P 51 As input to the half-adder, we obtain S h16 C h37 P 07 P 16 P 25 P 34 As input to the first 4-2 approximate compressor, S is obtained. c07 C c28 P 43 P 52 P 61 P 70 As input to the first 4-2 approximate compressor, S is obtained. c17 C c38 P 17 P 26 P 35 P 44 As input to the first 4-2 approximate compressor, S is obtained. c08 C c29 P 53 P 62 P 71 As input to a full adder, S is obtained. f18 C f39 P 27 P 36 P 45 P 54 As input to the first 4-2 approximate compressor, S is obtained. c09 C c110 P 63 P 72 As input to the half-adder, we obtain S h19 C h210 P 37 P 46 P 55 P 64 As input to a precise 4-2 compressor, S is obtained. e010 C e111 C o211 P 47 P 56 P 65As input to a full adder, S is obtained. f011 C f012 .
[0165] For phase two, S h04 P 22 P 31 P 40 As input to the second 4-2 approximate compressor, S is obtained. l4 C l4 S c05 C h15 P 41 P 50 As input to the second 4-2 approximate compressor, S is obtained. l5 C l5 S c06 S h16 C c26 P 60 As input to the second 4-2 approximate compressor, S is obtained. l6 C l6 S c07 S c17 C c27 C h37 As input to the second 4-2 approximate compressor, S is obtained. l7 C l7 S c08 S f18 C c28 C c38 As input to the second 4-2 approximate compressor, S is obtained. l8 C l8 S c09 S h19 C c29 C f39 As input to the second 4-2 approximate compressor, S is obtained. l9 C l9 S e010 C c110 C h210 P 73 As input to a precise 4-2 compressor, S is obtained. e10 C e10 C o10 S f011 C e111 C o211 P 74 As input to a precise 4-2 compressor, S is obtained. e11 C e11 C o11 C f012 P 57 P 66P 75 As input to a precise 4-2 compressor, S is obtained. e12 C e12 C o12 P 67 P 76 As input to the half-adder, we obtain S h13 C h13 .
[0166] For phase three, S e11 C e10 C o10 As input to a full adder, S is obtained. f11 C f11 S e12 C e11 C o11 As input to a full adder, S is obtained. f12 C f12 S h13 C e12 C o12 As input to a full adder, S is obtained. f13 C f13 .
[0167] For stage four, all operators are summed using a full adder to obtain the final result.
[0168] It should be noted that there are multiple ways to perform precise compression on the partial product in the precise compression section; it can be achieved using a precise 4-2 compressor, a full adder, and a half adder. The above method is only one preferred method, not the only one. In stage two of the above method, the carry bit of the precise 4-2 compressor input is defaulted to 0, which makes the approximate multiplication operation relatively fast. However, in other implementations, when using the precise 4-2 compressor to compress the column with the lowest weight in the compression section, the carry bit of the input is defaulted to 0. When using the precise 4-2 compressor to compress other columns, the carry bit is C, which is the carry when compressing the previous column using the precise 4-2 compressor. out .
[0169] It should be noted that, firstly, it is assumed that the multiplier of the multiplier is equally likely to be 0 or 1. Therefore, for the partial product P, the probability of P=1 is 0.25, and the probability of P=0 is 0.75. Based on this, the error caused by using the first approximation 4-2 compressor in stage 1 can be calculated: since the four inputs of each compressor in stage 1 are all P, the expected error of each compressor is 0.125. If the weight of the first column of the approximation region is set to unit 1, then the total expected error of stage 1, Error1, is = 0×1 + 0.125×2 + 0.125×4 + 0.25×8 + 0.125×16 + 0.125×32 = 8.75. The error of stage 2 can also be calculated, but since the inputs of columns with different weights are different in stage 2, the error of each column needs to be calculated separately, which will not be elaborated here. The expected error of stage 2 is directly given as Error2 = -8.75 = -Error1.
[0170] In traditional approximate calculation circuit design, a single compressor structure usually has a bias error in a fixed direction (e.g., the overall bias output is too high or too low). When multiple columns are superimposed in the multiplier, it is easy to cause a systematic error shift in the overall output, affecting the final calculation accuracy.
[0171] This invention proposes to design approximate compressors with different expected error directions (the first 4-2 approximate compressor has a negative expected error, and the second 4-2 approximate compressor has a positive expected error), and further to rationally combine and apply them in the multiplier column arrangement. By utilizing the complementary cancellation of positive and negative errors between columns, the mean error of the overall multiplication operation result can be effectively balanced, so that the expected output error approaches zero.
[0172] By employing a complementary error expectation configuration, not only can systematic drift caused by unidirectional accumulation be avoided, but the overall stability and robustness of the multiplier output can also be improved while maintaining low hardware resource overhead. This strategy is particularly suitable for applications that are sensitive to mean error but can tolerate local random errors, such as neural network inference, image processing, and edge computing.
[0173] Furthermore, compared to designs that use precise compressors entirely, this invention, by reasonably introducing approximate calculations, can significantly reduce the logic cascade depth and the number of devices, thereby reducing power consumption and area overhead. At the same time, it achieves controllable overall accuracy through an error complementarity mechanism, balancing energy efficiency and computational accuracy, and has high engineering application value in resource-constrained systems.
[0174] Example 4
[0175] An approximate multiplication operation device includes: a partial product generation module, an accurate partial compression module, an approximate partial compression module, a truncation processing module, and a carry adder module;
[0176] The partial product generation module performs a logical AND operation between each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; these n×n partial products are then used to construct a partial product compression tree with n rows and 2n-1 columns; n≥4; the partial product compression tree is divided into an exact compression part with n1 columns, an approximate compression part with n2 columns, and a truncated part with n3 columns; the weights of the partial products in the exact compression part, approximate compression part, and truncated part decrease sequentially; n1+n2+n3=2n-1; n1, n2, and n3 are all positive integers;
[0177] The precise partial compression module is used to approximate the partial product in the approximate compression section using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder; and to precisely compress the partial product in the precise compression section using a precise 4-2 compressor, a half adder, and a full adder.
[0178] The truncation module is used to truncate the partial product of the truncated part to obtain a truncated result of n3 bits all zero;
[0179] The carry adder module is used to arrange the partially products after precise compression as the high-order bits and the partially products after approximate compression as the low-order bits to obtain the combined result. The partial products of each column in the combined result are added from the low-order bits to the high-order bits. When there is a carry in the current column, the carry result is added to the current column to obtain the final carry result c0 and the sum of each column. The carry result c0, the sum of each column and the truncated result of n3 bits all zeros are arranged in order from the high-order bits to the low-order bits to form the final approximate multiplication result.
[0180] The first approximate 4-2 compressor is the 4-2 approximate compressor provided in Embodiment 1 of the present invention; the second approximate 4-2 compressor is the 4-2 approximate compressor provided in Embodiment 2 of the present invention.
[0181] The relevant technical solutions are the same as those in Embodiments 1, 2 and 3 of this invention, and will not be repeated here.
[0182] In summary, this invention addresses the structural optimization requirements of approximate multipliers by proposing a compressor construction method based on majority gate logic and majority gate logic with inverted inputs. By designing approximate compressors with different error characteristics and configuring them rationally according to column weights and error distribution patterns, this invention effectively achieves directional compensation and mean control of the overall multiplier error. The proposed multi-stage compression structure reduces hardware overhead and latency while ensuring the predictability and stability of output accuracy, avoiding the problem of systematic error accumulation in traditional multipliers. Leveraging the unified implementation of majority gate logic, this invention exhibits significant advantages in device utilization, energy efficiency, and scalability, making it suitable for widespread application in novel computing systems with stringent requirements for power consumption, area, and computational fault tolerance.
[0183] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A memristor-based 4-2 approximate compressor, characterized in that, Comprising: A 4-2 approximate compressor for compressing 4 1-bit numbers X1, X2, X3, X4, comprising: two identical first majority gate circuits and second majority gate circuits; Each of the first majority gate circuit and the second majority gate circuit comprises: a voltage comparator, a resistor, and three identical memristors; the resistance of the resistor is the low resistance value of the memristor; the negative poles of the three memristors are connected to one end of the resistor, and the positive poles of the three memristors are respectively used as the first input end, the second input end, and the third input end of the corresponding majority gate circuit; the other end of the resistor is used for grounding; the first input end of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator is used for outputting a first level when the voltage at the first input end is greater than the voltage at the second input end, and taking the first level as the output result "1" of the corresponding majority gate circuit, and outputting a second level when the voltage at the first input end is less than or equal to the voltage at the second input end, and taking the second level as the output result "0" of the corresponding majority gate circuit; the output end of the voltage comparator is used as the output end of the corresponding majority gate circuit; The three memristors in the first majority gate circuit are used for storing the resistance values corresponding to X1, X3 and X4 respectively one by one; the anodes of the three memristors in the first majority gate circuit are all used for connecting the voltage V1; the second input end of the voltage comparator in the first majority gate circuit is used for connecting the reference voltage V Ref1 ; and the output end of the first majority gate circuit is used for outputting the carry result of the 4-2 approximate compressor. The three memristors in the second majority gate circuit are used for storing the resistance values corresponding to X1, X2 and 1 respectively one by one, and the anodes thereof are used for being connected to the voltages 0, V2 and V2 respectively one by one; the second input end of the voltage comparator in the second majority gate circuit is used for being connected to the reference voltage V Ref2 ; and the output end of the second majority gate circuit is used for outputting the pseudo-sum result of the 4-2 approximate compressor. wherein 0 < V1 < V set ; V1 / 2 < V Ref1 < 2V1 / 3; 0 < V2 < V set ; V2 / 3 < V Ref2 < V2 / 2; V set is the threshold value for the memristor to change from a high resistance state to a low resistance state.
2. A control method of a 4-2 approximation compressor, characterized by, The 4-2 approximate compressor is the 4-2 approximate compressor of claim 1; and the control method comprises: X1, X3, X4 are written into the three memristors in the first majority gate circuit of the 4-2 approximate compressor one by one in a one-to-one correspondence, a voltage V1 is applied to the positive electrodes of the three memristors in the first majority gate circuit, and a reference voltage V Ref1 The carry result of the 4-2 approximate compressor is obtained at the output end of the first majority gate circuit. X1, X2, 1 are written into the three memristors in the second majority gate of the 4-2 approximate compressor one by one, a voltage of 0V is applied to the positive electrode of the memristor written with X1, a voltage V2 is applied to the positive electrode of the memristors written with X2 and 1, and a reference voltage V is connected to the second input end of the voltage comparator in the second majority gate Ref2 The pseudo-sum result of the 4-2 approximate compressor is obtained at the output end of the second majority gate. wherein 0 < V1 < V set ; V1 / 2 < V Ref1 < 2V1 / 3; 0 < V2 < V set ; V2 / 3 < V Ref2 < V2 / 2; V set is the threshold value for the memristor to change from a high resistance state to a low resistance state.
3. A memristor-based 4-2 approximate compressor characterized by, A 4-2 approximate compressor for compressing 4 1-bit numbers X1, X2, X3, X4, comprising: two identical first majority gate circuits and second majority gate circuits; Each of the first majority gate circuit and the second majority gate circuit comprises: a voltage comparator, a resistor, and three identical memristors; the resistance of the resistor is the low resistance value of the memristor; the negative poles of the three memristors are connected to one end of the resistor, and the positive poles of the three memristors are respectively used as the first input end, the second input end, and the third input end of the corresponding majority gate circuit; the other end of the resistor is used for grounding; the first input end of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator is used for outputting a first level when the voltage at the first input end is greater than the voltage at the second input end, and taking the first level as the output result "1" of the corresponding majority gate circuit, and outputting a second level when the voltage at the first input end is less than or equal to the voltage at the second input end, and taking the second level as the output result "0" of the corresponding majority gate circuit; the output end of the voltage comparator is used as the output end of the corresponding majority gate circuit; The three memristors in the first majority gate circuit are used for storing the resistance values corresponding to X3, X4 and 1 respectively one by one; the anodes of the three memristors in the first majority gate circuit are all used for connecting the voltage V1; the second input end of the voltage comparator in the first majority gate circuit is used for connecting the reference voltage V Ref1 ; and the output end of the first majority gate circuit is used for outputting the carry result of the 4-2 approximate compressor. The three memristors in the second majority gate circuit are used for storing the resistance values corresponding to 0, X1 and X4 respectively one by one, and the anodes thereof are used for connecting the voltages 0, V2 and V2 respectively one by one; the second input end of the voltage comparator in the second majority gate circuit is used for connecting the reference voltage V Ref2 ; and the output end of the second majority gate circuit is used for outputting the pseudo-sum result of the 4-2 approximate compressor. wherein 0 < V1 < V set ; V1 / 2 < V Ref1 < 2V1 / 3; 0 < V2 < V set ; V2 / 3 < V Ref2 < V2 / 2; V set is the threshold value for the memristor to change from a high resistance state to a low resistance state.
4. A control method of a 4-2 approximation compressor, characterized by, The 4-2 approximate compressor is the 4-2 approximate compressor of claim 3; and the control method comprises: X3, X4, 1 are written into the three memristors in the first majority gate of the 4-2 approximate compressor one by one, a voltage V1 is applied to the positive poles of the three memristors in the first majority gate, a reference voltage V Ref1 The carry result of the 4-2 approximate compressor is obtained at the output end of the first majority gate. correspondingly to 0, X1, X4, and a voltage V2 is applied to the anode of the memristor writing X1 and X4, the second input end of the voltage comparator in the second majority gate circuit is connected to a reference voltage V Ref2 The pseudo-sum result of the 4-2 approximate compressor is obtained at the output end of the second majority gate circuit. wherein 0 < V1 < V set ; V1 / 2 < V Ref1 < 2V1 / 3; 0 < V2 < V set ; V2 / 3 < V Ref2 < V2 / 2; V set is the threshold value for the memristor to change from a high resistance state to a low resistance state.
5. A memristor-based 4-2 approximate compressor device, characterized in that, The 4-2 approximate compressor device comprises the 4-2 approximate compressor of claim 1 and a controller for executing the control method of claim 2; or The 4-2 approximate compressor device comprises the 4-2 approximate compressor of claim 3 and a controller for executing the control method of claim 4. Comprising:
6. A method of approximate multiplication, characterized by, Performing an AND logic operation between each bit of an n-bit multiplier and each bit of an n-bit multiplicand to obtain n×n partial products; Constructing the n×n partial products into a partial product compression tree with n rows and 2n-1 columns; n≥4; The partial products are divided into an accurate compression part of n1 columns, an approximate compression part of n2 columns and a truncated part of n3 columns; the weight values of the partial products in the accurate compression part, the approximate compression part and the truncated part decrease in turn; n1+n2+n3=2n-1; n1, n2 and n3 are positive integers; The first approximate 4-2 compressor, the second approximate 4-2 compressor, the half adder and the full adder are used to perform approximate compression on the partial products in the approximate compression part; the accurate 4-2 compressor, the half adder and the full adder are used to perform accurate compression on the partial products in the accurate compression part; The partial products in the truncated part are subjected to truncation processing to obtain a truncated result of n3 bits of all zeros; The partial products after the accurate compression are arranged as high bits, and the partial products after the approximate compression are arranged as low bits to obtain a combination result; each column of partial products in the combination result is added in a direction from low bits to high bits, and when there is a carry from a previous column, the carry result is added to the current column to obtain a final carry result c0 and a summation result of each column; the carry result c0, the summation result of each column and the truncated result of n3 bits of all zeros are arranged in an order from high bits to low bits to form a final approximate multiplication result; The first approximate 4-2 compressor is the 4-2 approximate compressor of claim 1; and the second approximate 4-2 compressor is the 4-2 approximate compressor of claim 3.
7. The approximate multiplication method according to claim 6, wherein The approximate compression on the partial products in the approximate compression part by using the first approximate 4-2 compressor, the second approximate 4-2 compressor, the half adder and the full adder comprises: The first approximate 4-2 compressor, the half adder and the full adder are used to perform approximate compression on the partial products in the approximate compression part to obtain a first approximate compression array of 4×n2; The second approximate 4-2 compressor is used to perform approximate compression on the partial products in the first approximate compression array to obtain a second approximate compression array of 2×n2; When the compression operation is performed on the to-be-compressed part, the compression operation is performed on the to-be-compressed part in turn in an order from the last column to the first column; the weight value of the first column in the to-be-compressed part is the highest, and the weight value of the last column is the lowest; the to-be-compressed part is the approximate compression part or the first approximate compression array; The number of first approximate 4-2 compressors used in the j-th column of the approximate compression section, p j Number of full adders r j The number of half-adders q j All satisfy g j -4p j -3r j -2q j +p j +q j +r j +m j+1 =4; m j+1 =p j+1 +q j+1 +r j+1 j=1,2,......,n2-1; 4≤g h ≤n; p h ≥0; q h ≥0; r h ≥0; h=1,2,......,n2; g h This represents the number of partial products in the h-th column of the approximate compressed portion; When the hth column of the approximate compression part has a compression operation, the hth column of the first approximate compression array comprises a pseudo-sum result obtained by compressing the hth column of the approximate compression part; When the hth column of the approximate compression part contains partial products without compression operation, the hth column of the first approximate compression array comprises the partial products without compression operation in the hth column of the approximate compression part; When the j+1th column of the approximate compression part has a compression operation, the jth column of the first approximate compression array comprises a carry result after the compression of the j+1th column of the approximate compression part; The n2th column of the second approximate compression array includes a pseudo-sum result after compression of the n2th column of the first approximate compression array and 0; and the jth column of the second approximate compression array includes a pseudo-sum result after compression of the jth column of the first approximate compression array and a carry result after compression of the j+1th column of the first approximate compression array.
8. The approximate multiplication method according to claim 7, wherein When the jth column of the first approximate compression array contains a pseudo-sum result after compression of the jth column of the approximate compression part, a carry result after compression of the j+1th column of the approximate compression part, and partial products in the jth column of the approximate compression part that are not subjected to compression operations, the pseudo-sum result after compression of the jth column of the approximate compression part, the carry result after compression of the j+1th column of the approximate compression part, and the partial products in the jth column of the approximate compression part that are not subjected to compression operations are sequentially arranged in the column direction; When the jth column of the first approximate compression array contains only a pseudo-sum result after compression of the jth column of the approximate compression part and a carry result after compression of the j+1th column of the approximate compression part, the pseudo-sum result after compression of the jth column of the approximate compression part and the carry result after compression of the j+1th column of the approximate compression part are sequentially arranged in the column direction; When the jth column of the first approximate compression array contains only a carry result after compression of the j+1th column of the approximate compression part and partial products in the jth column of the approximate compression part that are not subjected to compression operations, the carry result after compression of the j+1th column of the approximate compression part and the partial products in the jth column of the approximate compression part that are not subjected to compression operations are sequentially arranged in the column direction; When the hth column of the first approximate compression array contains only a pseudo-sum result after compression of the hth column of the approximate compression part and partial products in the jth column of the approximate compression part that are not subjected to compression operations, the pseudo-sum result after compression of the hth column of the approximate compression part and the partial products in the jth column of the approximate compression part that are not subjected to compression operations are sequentially arranged in the column direction; When there are multiple compression operations in the hth column of the approximate compression part, there are multiple pseudo-sum results after compression of the hth column of the approximate compression part, and the arrangement order of the pseudo-sum results in the hth column of the first approximate compression array is consistent with the sorting order of the compression objects in the hth column of the approximate compression part; When there are multiple partial products in the hth column of the approximate compression part that are not subjected to compression operations, the arrangement order of the partial products in the hth column of the first approximate compression array is consistent with the sorting order of the partial products in the hth column of the approximate compression part; When compression is performed using the first approximate 4-2 compressor, the compression objects are four partial products sequentially arranged in the column direction in a column of the approximate compression part, which are one-to-one corresponding to 1-bit numbers X1, X2, X3, and X4, respectively, and are input into the first approximate 4-2 compressor for compression; When compressed by the second approximate 4-2 compressor, the compression object is four partial products in a column of the first approximate compression array, which are sequentially arranged along the column direction, and are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3 and X4 in one-to-one correspondence.
9. The approximate multiplication method according to claim 8, wherein n=8; the n-3 columns with the highest weights in the partial product compression tree are taken as the accurate compression part, the n-2 columns with the second highest weights in the partial product compression tree are taken as the approximate compression part, and the four columns with the lowest weights in the partial product compression tree are taken as the truncated part; The n-bit multiplier is a7a6a5a4a3a2a1a0, and the n-bit multiplicand is b7b6b5b4b3b2b1b0. The first to eighth rows in the partial product compression tree, sorted in descending order of weight values, are as follows: 07 P 06 P 05 P 04 P 03 P 02 P 01 P 00 P 17 P 16 P 15 P 14 P 13 P 12 P 01 P 00 P 27 P 26 P 25 P 24 P 23 P 22 P 21 P 20 P 37 P 36 P 35 P 34 P 33 P 32 P 31 P 30 P 47 P 46 P 45 P 44 P 43 P 42 P 41 P 40 P 57 P 56 P 55 P 54 P 53 P 52 P 51 P 50 P 67 P 66 P 65 P 64 P 63 P 62 P 61 P 60 P 77 P 76 P 75 P 74 P 73 P 72 P 71 P 70 ; where P 07 is the partial product of the multiplication of a0 and b7, P 06 is the partial product of the multiplication of a0 and b6, and so on, P 70 is the partial product of the multiplication of a7 and b0. The approximate compression part includes 6 columns of partial products; the partial products in the first column from top to bottom are in turn: P 27 P 36 P 45 P 54 P 63 P 72 ; the partial products in the second column from top to bottom are in turn: P 17 P 26 P 35 P 44 P 53 P 62 P 71 ; the partial products in the third column from top to bottom are in turn: P 07 P 16 P 25 P 34 P 43 P 52 P 61 P 70 ; the partial products in the fourth column from top to bottom are in turn: P 06 P 15 P 24 P 33 P 42 P 51 P 60 ; the partial products in the fifth column from top to bottom are in turn: P 05 P 14 P 23 P 32 P 41 P 50 ; the partial products in the sixth column from top to bottom are in turn: P 04 P 13 P 22 P 31 P 40 ; the weight value of the first column in the approximate compression part is the highest, and the weight value of the last column is the lowest; The approximate compression operation on the partial products in the approximate compression part is performed by using the first approximate 4-2 compressor, the half adder and the full adder, and a first approximate compression array with a size of 4×n2 is obtained, which includes: P 04 and P 13 are compressed by using a half adder to obtain pseudo sum result S h04 and carry result C h15 ; P 05 , P 14 , P 23 , P 32 are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, X4 respectively in one-to-one correspondence to be compressed to obtain pseudo sum result S c05 and carry result C c26 ; P 06 , P 15 , P 24 , P 33 are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, X4 respectively in one-to-one correspondence to be compressed to obtain pseudo sum result S c06 and carry result C c27 ; P 42 and P 51 are compressed by using a half adder to obtain pseudo sum result S h16 and carry result C h37 ; P 07 , P 16 , P 25 , P 34 are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, X4 respectively in one-to-one correspondence to be compressed to obtain pseudo sum result S c07 and carry result C c28 ; P 43 , P 52 , P 61 , P 70 are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, X4 respectively in one-to-one correspondence to be compressed to obtain pseudo sum result S c17 and carry result C c38 ; P 17 , P 26 , P 35 , P 44 are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, X4 respectively in one-to-one correspondence to be compressed to obtain pseudo sum result S c08 and carry result C c29 ; P 62 and P 71 are compressed by using a full adder to obtain sum result S f18 and carry result C f39 ; P 27 , P 36 , P 45 , P 54 are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, X4 respectively in one-to-one correspondence to be compressed to obtain pseudo-sum result S c09 and carry result C c110 ; P 63 and P 72 are compressed by using a half adder to obtain pseudo-sum result S h19 and carry result C h210 ; P 37 , P 46 , P 55 , P 64 are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, X4 respectively in one-to-one correspondence to be compressed to obtain pseudo-sum result S e010 and carry result C f012 ; Thus the first column in the first approximated compression array is S c09 S h19 S c29 S f39 The second column is S c08 S h18 S c28 S f38 The third column is S c07 S c17 S c27 S h37 The fourth column is S c06 S h16 S c26 S 60 The fifth column is S c05 S h15 S 41 S 50 The sixth column is S h04 S 22 S 31 S 40 ; The approximate compression operation on the partial products in the first approximate compression array is performed by using the second approximate 4-2 compressor, and a second approximate compression array with a size of 2×n2 is obtained, which includes: S h04 P 22 P 31 P 40 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l4 Sum of carry result C l4 ;S c05 C h15 P 41 P 50 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l5 Sum of carry result C l5 ;S c06 S h16 C c26 P 60 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l6 Sum of carry result C l6 ;S c07 S c17 C c27 C h37 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l7 Sum of carry result C l7 ;S c08 S f18 C c28 C c38 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the first approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l8 Sum of carry result C l8 ;S c09 S h19 C c29 C f39 Each bit is correspondingly represented as a 1-bit number X1, X2, X3, and X4, and input into the second approximate 4-2 compressor for compression, yielding the pseudo-sum result S. l9 Sum of carry result C l9 ; Thus the first column in the second approximated compression array comprises S l9 and C l8 , the second column comprises S l8 and C l7 , the third column comprises S l7 and C l6 , the fourth column comprises S l6 and C l5 , the fifth column comprises S l5 and C l4 , the sixth column comprises S l4 and 0.
10. An approximate multiplication operation device, characterized by comprising: It includes: The partial product generation module, the accurate part compression module, the approximate part compression module, the truncation processing module and the carry adder module; The partial product generation module is configured to perform AND logical operation on each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; and the n×n partial products are arranged into a partial product compression tree with n rows and 2n-1 columns. n≥4; the partial product compression tree is divided into an accurate compression part with n1 columns, an approximate compression part with n2 columns and a truncated part with n3 columns; the weight values of the partial products in the accurate compression part, the approximate compression part and the truncated part decrease in turn; n1+n2+n3=2n-1; n1, n2 and n3 are positive integers. The accurate part compression module is configured to perform approximate compression on the partial products in the approximate compression part by using the first approximate 4-2 compressor, the second approximate 4-2 compressor, the half adder and the full adder; and perform accurate compression on the partial products in the accurate compression part by using the accurate 4-2 compressor, the half adder and the full adder. The truncation processing module is configured to perform truncation processing on the partial products in the truncated part to obtain a truncated result of n3 all-zeroes. The carry adder module is configured to arrange the partial products after accurate compression as high bits and the partial products after approximate compression as low bits to obtain a combination result; perform addition operation on each column of partial products in the combination result in the direction from low bits to high bits, and when there is a carry from the previous column, add the carry result to the current column to obtain a final carry result c0 and a summation result of each column; and arrange the carry result c0, the summation result of each column and the truncated result of n3 all-zeroes in the order from high bits to low bits to form a final approximate multiplication operation result. The first approximate 4-2 compressor is the 4-2 approximate compressor in claim 1, and the second approximate 4-2 compressor is the 4-2 approximate compressor in claim 3.