4-2 approximate compressor based on memristor, control method, device and application

By designing a majority gate circuit based on memristors, simplifying the logic depth and introducing fixed input values, the problems of high power consumption and low speed of the existing 4-2 approximate compressor are solved, and an approximate compressor with low complexity, high speed and excellent error performance is realized.

CN120601893AActive Publication Date: 2025-09-05HUAZHONG UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510648751.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-05
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing 4-2 approximate compressor circuit has high power consumption and low computing speed due to its deep logic depth and frequent logic gate switching, making it difficult to achieve approximate calculation with good error performance.

Method used

A memristor-based majority gate circuit design is adopted, including two identical majority gate circuits and a majority gate circuit with inversion logic. Voltage comparators and memristors are used to achieve 4-2 approximate compression, simplify the logic depth, and introduce fixed input values ​​to construct logic bias, thereby reducing power consumption and controlling the error direction.

Benefits of technology

A 4-2 approximate compressor with low circuit complexity, low power consumption and high computing rate is implemented, with excellent error performance, suitable for high-speed parallel computing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120601893A_ABST
    Figure CN120601893A_ABST
Patent Text Reader

Abstract

The invention discloses a 4-2 approximate compressor based on memristors, a control method, a device and application. Approximate compression is performed on four 1-bit numbers X1, X2, X3 and X4 based on two same majority gate circuits based on the memristors; the two majority gate circuits are used for respectively realizing traditional majority gate logic and novel majority gate logic containing negation logic; the first majority gate circuit is used for realizing traditional majority gate logic M (X1, X3, X4) to obtain a carry result of a 4-2 approximate compressor; the second majority gate circuit is used for realizing a novel majority gate logic # imgabs0 # containing negation logic to obtain a pseudo sum result of a 4-2 approximate compressor; the 4-2 compressor does not need to depend on cascade construction of a large number of basic logic units such as NAND gates and NOR gates, the logic depth is relatively shallow, and the 4-2 compressor with relatively good error performance can be realized with relatively low circuit complexity, relatively high operation rate and relatively low power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of integrated circuit technology, and more specifically, relates to a memristor-based 4-2 approximate compressor, a control method, a device, and applications. Background Art

[0002] In recent years, approximate computing, as an effective method for optimizing circuit performance while maintaining tolerable computational errors, has been widely studied and applied to applications with low precision requirements, such as image blur processing, video compression, and forward reasoning in convolutional neural networks. In high-speed parallel computing, the accumulation of multiple operands is often a challenge, such as compressing partial products in parallel multipliers. The 4-2 compressor is the most commonly used compressor module, effectively reducing the number of intermediate result bits. Therefore, the study of a 4-2 approximate compressor is of great significance.

[0003] Existing 4-2 approximate compressor circuits are mostly implemented using CMOS structures. In order to obtain better error performance, they usually need to rely on a large number of basic logic units such as NAND gates and NOR gates to be cascaded. The logic depth is deep and the overall circuit structure is relatively complex. In addition, frequent logic gate switching is usually required during the calculation process, resulting in high power consumption and low calculation speed. Summary of the Invention

[0004] In response to the above-mentioned defects or improvement needs of the prior art, the present invention provides a 4-2 approximate compressor, control method, device and application based on memristor, the purpose of which is to realize a 4-2 compressor with better error performance with lower circuit complexity, higher computing rate and lower power consumption.

[0005] To achieve the above objectives, in a first aspect, the present invention provides a memristor-based 4-2 approximate compressor for approximate compression of four 1-bit numbers X1, X2, X3, and X4, comprising: two identical first majority gate circuits and a second majority gate circuit; The first majority gate circuit and the second majority gate circuit both include: a voltage comparator, a resistor, and three identical memristors; the resistance value of the resistor is the low-resistance resistance value of the memristor; the negative electrodes of the three memristors are connected to one end of the resistor, and the positive electrodes of the three memristors serve as the first input end, the second input end, and the third input end of the corresponding majority gate circuit respectively; the other end of the resistor is used for grounding; the first input end of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator is used to output a first level when the voltage at its first input end is greater than the voltage at its second input end, and use the first level as the output result "1" of the corresponding majority gate circuit, and to output a second level when the voltage at its first input end is less than or equal to the voltage at its second input end, and use the second level as the output result "0" of the corresponding majority gate circuit; the output end of the voltage comparator serves as the output end of the corresponding majority gate circuit; The three memristors in the first majority gate circuit are used to store the resistance values ​​corresponding to X1, X3, and X4 respectively; the positive electrodes of the three memristors in the first majority gate circuit are all used to access the voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is used to access the reference voltage V Ref1 The output terminal of the first majority gate circuit is used to output the carry result of the 4-2 approximate compressor; The three memristors in the second majority gate circuit are used to store the resistance values ​​corresponding to X1, X2, and 1 respectively, and their positive electrodes are used to connect to voltages 0, V2, and V2 respectively; the second input terminal of the voltage comparator in the second majority gate circuit is used to connect to the reference voltage V Ref2 The output terminal of the second majority gate circuit is used to output the pseudo-sum result of the 4-2 approximate compressor; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

[0006] In a second aspect, the present invention provides a control method for the 4-2 approximation compressor in the first aspect, comprising: X1, X3, and X4 are written into the three memristors in the first majority gate circuit one by one. A voltage V1 is applied to the positive electrodes of the three memristors in the first majority gate circuit. The second input terminal of the voltage comparator in the first majority gate circuit is connected to the reference voltage V Ref1 , obtaining the carry result of the 4-2 approximate compressor at the output end of the first majority gate circuit; Write X1, X2, and 1 into the three memristors in the second majority gate circuit one by one. Apply a 0V voltage to the positive electrode of the memristor written with X1, and apply a voltage V2 to the positive electrodes of the memristors written with X2 and 1. Connect the reference voltage V to the second input of the voltage comparator in the second majority gate circuit. Ref2 , a pseudo-sum result of the 4-2 approximate compressor is obtained at the output end of the second majority gate circuit; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

[0007] In a third aspect, the present invention provides a memristor-based 4-2 approximate compressor for approximate compression of four 1-bit numbers X1, X2, X3, and X4, comprising: two identical first majority gate circuits and a second majority gate circuit; The first majority gate circuit and the second majority gate circuit both include: a voltage comparator, a resistor, and three identical memristors; the resistance value of the resistor is the low-resistance resistance value of the memristor; the negative electrodes of the three memristors are connected to one end of the resistor, and the positive electrodes of the three memristors serve as the first input end, the second input end, and the third input end of the corresponding majority gate circuit respectively; the other end of the resistor is used for grounding; the first input end of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator is used to output a first level when the voltage at its first input end is greater than the voltage at its second input end, and use the first level as the output result "1" of the corresponding majority gate circuit, and to output a second level when the voltage at its first input end is less than or equal to the voltage at its second input end, and use the second level as the output result "0" of the corresponding majority gate circuit; the output end of the voltage comparator serves as the output end of the corresponding majority gate circuit; The three memristors in the first majority gate circuit are used to store the resistance values ​​corresponding to X3, X4, and 1 respectively; the positive electrodes of the three memristors in the first majority gate circuit are all used to access the voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is used to access the reference voltage V Ref1 The output terminal of the first majority gate circuit is used to output the carry result of the 4-2 approximate compressor; The three memristors in the second majority gate circuit are used to store the resistance values ​​corresponding to 0, X1, and X4 respectively, and their positive electrodes are used to connect to voltages 0, V2, and V2 respectively; the second input terminal of the voltage comparator in the second majority gate circuit is used to connect to the reference voltage V Ref2 The output terminal of the second majority gate circuit is used to output the pseudo-sum result of the 4-2 approximate compressor; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

[0008] In a fourth aspect, the present invention provides a control method for the 4-2 approximation compressor according to the third aspect, comprising: Write X3, X4, and 1 into the three memristors in the first majority gate circuit one by one. Apply voltage V1 to the positive electrodes of the three memristors in the first majority gate circuit. Connect the reference voltage V to the second input terminal of the voltage comparator in the first majority gate circuit. Ref1 , obtaining the carry result of the 4-2 approximate compressor at the output end of the first majority gate circuit; Write 0, X1, and X4 into the three memristors in the second majority gate circuit one by one. Apply a voltage of 0V to the positive electrode of the memristor written with 0, and apply a voltage of V2 to the positive electrodes of the memristors written with X1 and X4. Connect the reference voltage V to the second input of the voltage comparator in the second majority gate circuit. Ref2 , a pseudo-sum result of the 4-2 approximate compressor is obtained at the output end of the second majority gate circuit; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

[0009] In a fifth aspect, the present invention provides a 4-2 approximation compressor device based on a memristor; The 4-2 approximate compressor device comprises the 4-2 approximate compressor provided by the first aspect of the present invention and a controller for executing the control method provided by the second aspect of the present invention; or, The 4-2 approximate compressor device includes the 4-2 approximate compressor provided by the third aspect of the present invention and a controller for executing the control method provided by the fourth aspect of the present invention.

[0010] In a sixth aspect, the present invention provides an approximate multiplication method, comprising: Perform a logical AND operation on each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; construct a partial product compression tree with n rows and 2n-1 columns from the n×n partial products; n ≥ 4; The partial product compression tree is divided into an exact compression part of column n1, an approximate compression part of column n2, and a truncated part of column n3; the weights of the partial products in the exact compression part, the approximate compression part, and the truncated part decrease in sequence; n1+n2+n3=2n-1; n1, n2, and n3 are all positive integers; Approximately compressing the partial products in the approximate compression portion using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder; and accurately compressing the partial products in the accurate compression portion using an accurate 4-2 compressor, a half adder, and a full adder; The partial product of the truncated part is truncated to obtain a truncated result with n3 bits all zero; The partial products after exact compression are used as the high bits, and the partial products after approximate compression are used as the low bits to obtain a combination result; the partial products of each column in the combination result are added in the direction from low bits to high bits, and when there is a carry in the previous column, the carry result is added to the current column to obtain the final carry result c0 and the summation result of each column; the carry result c0, the summation results of each column, and the truncation result with n3 bits all zero are arranged in order from high bits to low bits to form the final approximate multiplication result; Among them, the first approximate 4-2 compressor is the 4-2 approximate compressor provided by the first aspect of the present invention; the second approximate 4-2 compressor is the 4-2 approximate compressor provided by the fourth aspect of the present invention.

[0011] Further preferably, the partial products in the approximate compression part are approximate compressed using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder, comprising: Performing an approximate compression operation on the partial products in the approximate compression part by using a first approximate 4-2 compressor, a half adder, and a full adder to obtain a first approximate compression array of size 4×n2; Using a second approximate 4-2 compressor to perform an approximate compression operation on the partial products in the first approximate compression array, thereby obtaining a second approximate compression array of size 2×n2; When performing a compression operation on the portion to be compressed, the compression operation is performed column by column in the order from the last column to the first column of the portion to be compressed; the first column in the portion to be compressed has the highest weight value, and the last column has the lowest weight value; the portion to be compressed is an approximate compression portion or a first approximate compression array; The number of first approximate 4-2 compressors p used for the jth column of the approximate compressed part j , the number of full adders r j and the number of half adders q j All satisfy g j -4p j -3r j -2qj +p j +q j +r j +m j+1 =4;m j+1 =p j+1 +q j+1 +r j+1 ;j=1,2,......,n2-1;4≤g h ≤n;p h ≥0;q h ≥0; r h ≥0; h=1,2,......,n2; g h is the number of partial products in the hth column of the approximate compressed part; When a compression operation occurs on the hth column of the approximately compressed portion, the hth column of the first approximately compressed array includes: a pseudo-sum result obtained by compressing the hth column of the approximately compressed portion; When the hth column of the approximate compressed portion includes partial products that are not subjected to the compression operation, the hth column of the first approximate compressed array includes: the partial products that are not subjected to the compression operation in the hth column of the approximate compressed portion; When a compression operation occurs on the j+1th column of the approximately compressed portion, the jth column of the first approximately compressed array includes: a carry result after compressing the j+1th column of the approximately compressed portion; The n2th column of the second approximate compressed array includes: the pseudo-sum result after compressing the n2th column of the first approximate compressed array and 0; the jth column of the second approximate compressed array includes: the pseudo-sum result after compressing the jth column of the first approximate compressed array and the carry result after compressing the j+1th column of the first approximate compressed array.

[0012] Further preferably, when the j-th column of the first approximate compressed array includes a pseudo-sum result obtained by compressing the j-th column of the approximate compressed portion, a carry result after compressing the j+1-th column of the approximate compressed portion, and a partial product of the j-th column of the approximate compressed portion that is not subjected to the compression operation, the pseudo-sum result obtained by compressing the j-th column of the approximate compressed portion, the carry result after compressing the j+1-th column of the approximate compressed portion, and the partial product of the j-th column of the approximate compressed portion that is not subjected to the compression operation are arranged sequentially along the column direction; When the j-th column of the first approximate compression array contains only the pseudo-sum result obtained by compressing the j-th column of the approximate compression part and the carry result after compressing the j+1-th column of the approximate compression part, the pseudo-sum result obtained by compressing the j-th column of the approximate compression part and the carry result after compressing the j+1-th column of the approximate compression part are arranged sequentially along the column direction; When the j-th column of the first approximate compressed array contains only the carry result after compression of the j+1-th column of the approximate compressed part and the partial product of the j-th column of the approximate compressed part that is not subjected to the compression operation, the carry result after compression of the j+1-th column of the approximate compressed part and the partial product of the j-th column of the approximate compressed part that is not subjected to the compression operation are arranged sequentially along the column direction; When the h-th column of the first approximate compressed array includes only the pseudo-sum result obtained by compressing the h-th column of the approximate compressed part and the partial product of the j-th column of the approximate compressed part without the compression operation, the pseudo-sum result obtained by compressing the h-th column of the approximate compressed part and the partial product of the j-th column of the approximate compressed part without the compression operation are arranged sequentially along the column direction; When there are multiple compression operations on the h-th column of the approximately compressed part, there are multiple pseudo-sum results obtained after compressing the h-th column of the approximately compressed part, and the arrangement order of the pseudo-sum results in the h-th column of the first approximate compression array is consistent with the sorting order of the compressed objects of the corresponding compression operations in the h-th column of the approximately compressed part; When there are multiple partial products that are not subjected to the compression operation in the h-th column of the approximate compression part, their arrangement order in the h-th column of the first approximate compression array is consistent with their sorting order in the h-th column of the approximate compression part.

[0013] Further preferably, when the first approximate 4-2 compressor is used for compression, the compression objects are four partial products arranged sequentially along the column direction in a column of the approximate compressed part, which are respectively used as 1-bit numbers X1, X2, X3, and X4 in a one-to-one correspondence and input into the first approximate 4-2 compressor for compression; When the second approximate 4-2 compressor is used for compression, the compression objects are the four partial products arranged in sequence along the column direction in a column of the first approximate compression array, which are input one-to-one as 1-bit numbers X1, X2, X3, and X4 respectively and compressed into the second approximate 4-2 compressor for compression.

[0014] Further preferably, n=8; the n-3 columns with the highest weight in the partial product compression tree are used as the exact compression part, the n-2 columns with the second highest weight in the partial product compression tree are used as the approximate compression part, and the 4 columns with the lowest weight in the partial product compression tree are used as the truncated part; The n-bit multiplier is a7a6a5a4a3a2a1a0; the n-bit multiplicand is b7b6b5b4b3b2b1b0; The first to eighth rows in the partial product compression tree, sorted in descending order of weight values, are: P 07 P 06 P 05 P 04 P 03 P02 P 01 P 00 、P 17 P 16 P 15 P 14 P 13 P 12 P 01 P 00 、P 27 P 26 P 25 P 24 P 23 P 22 P 21 P 20 、P 37 P 36 P 35 P 34 P 33 P 32 P 31 P 30 、P 47 P 46 P 45 P 44 P 43 P 42 P 41 P 40 、P 57 P 56 P 55 P 54 P 53 P 52 P 51 P 50 、P 67 P 66 P 65 P 64 P 63 P 62 P 61 P 60 、P 77 P 76 P 75 P 74 P 73 P 72 P 71 P 70 Among them, P 07 is the partial product of a0 and b7, P 06 is the partial product of a0 and b6, and so on. 70 is the partial product of a7 and b0; The approximate compression part includes 6 columns of partial products; the partial products in the first column are as follows from top to bottom: P 27 P 36 P45 P 54 P 63 P 72 ; The partial products in the second column are from top to bottom: P 17 P 26 P 35 P 44 P 53 P 62 P 71 ; The partial products in the third column are from top to bottom: P 07 P 16 P 25 P 34 P 43 P 52 P 61 P 70 ; The partial products in the fourth column from top to bottom are: P 06 P 15 P 24 P 33 P 42 P 51 P 60 ; The partial products in the fifth column from top to bottom are: P 05 P 14 P 23 P 32 P 41 P 50 ; The partial products in the sixth column from top to bottom are: P 04 P 13 P 22 P 31 P 40 ; The first column in the approximate compression part has the highest weight value, and the last column has the lowest weight value; The first approximate 4-2 compressor, the half adder, and the full adder are used to perform an approximate compression operation on the partial products in the approximate compression part, thereby obtaining a first approximate compression array of size 4×n2, including: Using half adder to P 04 and P 13 Compression is performed to obtain the pseudo-sum result S h04 and carry result C h15 ;P 05 、P 14 、P 23 、P 32 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c05 and carry result C c26 ;P 06 、P 15 、P 24 、P33 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c06 and carry result C c27 ; Use half adder to P 42 and P 51 Compression is performed to obtain the pseudo-sum result S h16 and carry result C h37 ;P 07 、P 16 、P 25 、P 34 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c07 and carry result C c28 ;P 43 、P 52 、P 61 、P 70 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c17 and carry result C c38 ;P 17 、P 26 、P 35 、P 44 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c08 and carry result C c29 ; Use full adder to P 62 and P 71 Compress and get the result S f18 and carry result C f39 ;P 27 、P 36 、P 45 、P 54 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c09 and carry result C c110 ; Use half adder to P 63 and P 72 Compression is performed to obtain the pseudo-sum result S h19 and carry result C h210 ;P 37 、P 46 、P 55 、P 64They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. e010 and carry result C f012 ; The first column of the first approximate compressed array is S c09 S h19 C c29 C f39 , the second column is S c08 S h18 C c28 C f38 , the third column is S c07 S c17 C c27 C h37 , the fourth column is S c06 S h16 C c26 P 60 , the fifth column is S c05 C h15 P 41 P 50 , the sixth column is S h04 P 22 P 31 P 40 ; The above-mentioned second approximate 4-2 compressor is used to perform an approximate compression operation on the partial products in the first approximate compressed array, thereby obtaining a second approximate compressed array of size 2×n2, including: S h04 、P 22 、P 31 、P 40 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l4 and carry result C l4 ; S c05 、C h15 、P 41 、P 50 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l5 and carry result C l5 ; S c06 、S h16 、C c26 、P 60 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l6 and carry result C l6 ; Sc07 、S c17 、C c27 、C h37 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l7 and carry result C l7 ; S c08 、S f18 、C c28 、C c38 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l8 and carry result C l8 ; S c09 、S h19 、C c29 、C f39 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l9 and carry result C l9 ; The first column of the second approximate compressed array thus obtained includes S l9 and C l8 , the second column includes S l8 and C l7 , the third column includes S l7 and C l6 , the fourth column includes S l6 and C l5 , the fifth column includes S l5 and C l4 , the sixth column includes S l4 and 0.

[0015] In a seventh aspect, the present invention provides an approximate multiplication operation device, comprising: a partial product generation module, an accurate partial compression module, an approximate partial compression module, a truncation processing module, and a carry adder module; The partial product generation module is used to perform a logical AND operation on each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; the n×n partial products are constructed into a partial product compression tree with n rows and 2n-1 columns; n ≥ 4; the partial product compression tree is divided into an exact compression part of column n1, an approximate compression part of column n2, and a truncated part of column n3; the weights of the partial products in the exact compression part, the approximate compression part, and the truncated part decrease in sequence; n1+n2+n3=2n-1; n1, n2, and n3 are all positive integers; The precise partial compression module is used to approximately compress the partial products in the approximate compression part using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder; and precisely compress the partial products in the precise compression part using a precise 4-2 compressor, a half adder, and a full adder; The truncation processing module is used to truncate the partial product of the truncation part to obtain a truncation result with n3 bits all zero; The carry adder module is used to arrange the partial products after exact compression as the high bits and the partial products after approximate compression as the low bits to obtain a combination result; the partial products of each column in the combination result are added in the direction from low bits to high bits, and when there is a carry in the previous column, the carry result is added to the current column to obtain the final carry result c0 and the summation result of each column; the carry result c0, the summation results of each column, and the truncation result of n3 bits all zero are arranged in order from high bits to low bits to form the final approximate multiplication result; Among them, the first approximate 4-2 compressor is the 4-2 approximate compressor provided by the first aspect of the present invention; the second approximate 4-2 compressor is the 4-2 approximate compressor provided by the fourth aspect of the present invention.

[0016] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects: 1. A first aspect of the present invention provides a memristor-based 4-2 approximate compressor, which performs approximate compression on four 1-bit numbers X1, X2, X3, and X4 based on two identical memristor-based majority gate circuits; wherein the majority gate circuit includes a voltage comparator, a resistor, and three identical memristors; the two majority gate circuits respectively implement traditional majority gate logic and a new majority gate logic including negation logic; the first majority gate circuit is used to implement traditional majority gate logic M(X1, X3, X4) to obtain a carry result of the 4-2 approximate compressor; the second majority gate circuit is used to implement the new majority gate logic including negation logic , and obtain the pseudo-sum result of the 4-2 approximate compressor; the present invention does not need to rely on a large number of basic logic gate cascades, has a shallow logic depth, a simple circuit structure, and a low delay. Combined with the low power consumption characteristics of the memristor itself, it maintains better error performance while achieving efficient compression, especially in the input design of the first majority gate circuit and the second majority gate circuit, the input dominant value is retained, which helps to reduce the average error; based on this, the present invention can realize a 4-2 compressor with better error performance with lower circuit complexity, higher computing speed and lower power consumption.

[0017] 2. A third aspect of the present invention provides a memristor-based 4-2 approximate compressor, which performs approximate compression on four 1-bit numbers X1, X2, X3, and X4 based on two identical memristor-based majority gate circuits; wherein the majority gate circuit includes a voltage comparator, a resistor, and three identical memristors; the two majority gate circuits respectively implement traditional majority gate logic and a new majority gate logic including inversion logic; the first majority gate circuit is used to implement traditional majority gate logic M(X3, X4, 1) to obtain a carry result of the 4-2 approximate compressor; the second majority gate circuit is used to implement the new majority gate logic including inversion logic , obtaining the pseudo-sum result of a 4-2 approximate compressor; the present invention does not rely on a large number of basic logic gate cascades, has a shallow logic depth, a simple device composition, low latency, and low power consumption; by introducing fixed values ​​of "1" and "0" at the input, a structural logic bias is effectively constructed, so that the pseudo-sum result has a clear error direction, which is conducive to the purposeful error guidance and neutralization in subsequent compression. Therefore, the present invention not only has low hardware complexity, but also exhibits better error equalization capabilities in a controllable error structure; based on this, the present invention can realize a 4-2 compressor with better error performance with lower circuit complexity, higher computing speed, and lower power consumption.

[0018] 3. The sixth aspect of the present invention provides an approximate multiplication method, which utilizes the first approximate 4-2 compressor provided by the first aspect of the present invention, the approximate 4-2 compressor provided by the third aspect of the present invention, a half adder, and a full adder to approximate compression of the partial product in the approximate compression part; wherein the first approximate 4-2 compressor omits some logic inputs while maintaining the main compression path, has a simple structure, low area and power consumption, and is suitable for intermediate weight column compression paths with low accuracy requirements, and its error expectation is negative; the second approximate 4-2 compressor is an error-guided approximate compressor, which introduces a fixed input "1" to construct an output bias, so that the output error has directionality and adjustability, and is suitable for use in error-guided or error mean-controlled approximate multiplier compression paths, and its error expectation is positive. By utilizing 4-2 approximate compressors with different error expectation directions, an approximate multiplication method with controllable error, reconfigurable structure, simple structure, high computing speed, and good resource utilization can be achieved.

[0019] 4. Furthermore, the approximate multiplication method provided by the present invention divides the approximate compression of the partial products in the approximate compression portion into two stages. In the first stage, the partial products in the approximate compression portion are approximate compressed using a first approximate 4-2 compressor, a half adder, and a full adder, resulting in a first approximate compressed array of size 4×n². In the second stage, the partial products in the first approximate compressed array are approximate compressed using a second approximate 4-2 compressor, resulting in a second approximate compressed array of size 2×n². The compression in the first stage results in the generated first approximate compressed array having four rows, which enables the second stage to use only the second approximate 4-2 compressor for compression, thereby facilitating the second stage to adjust and compensate for the error direction of the output of the first stage. Since the pseudo-sum output of the first approximate 4-2 compressor has a positive error distribution, and the second approximate 4-2 compressor has the characteristic of structurally guiding positive errors, its unified application in the second stage can achieve complementary cancellation of positive and negative errors between columns, effectively balancing the error mean of the overall multiplication output, making the error expectation approach zero, and further reducing the error of the multiplication operation.

[0020] 5. Furthermore, the approximate multiplication method provided by the present invention is preferably applicable to the multiplication operation of an 8-bit multiplier and an 8-bit multiplicand. By reasonably setting the positions of the approximate 4-2 compressor, half adder and full adder in the compression process of different stages of the approximate compression part, the positive and negative errors can be offset complementary between columns to the greatest extent possible, thereby effectively balancing the mean error of the overall multiplication operation result, making the output error expectation close to zero, and further reducing the error of the multiplication operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Schematic diagram of the 4-2 approximate compressor provided in Example 1 of the present invention.

[0022] Figure 2 This is a schematic diagram of the structure of the majority gate circuit provided in Example 1 of the present invention.

[0023] Figure 3 This is a schematic diagram of the equivalent circuit of the memristor after the logic value is written into the memristor provided in Example 1 of the present invention.

[0024] Figure 4 Schematic diagram of a 4-2 approximate compressor provided in Example 2 of the present invention.

[0025] Figure 5 This is a gate-level circuit diagram of a half adder provided in Example 3 of the present invention.

[0026] Figure 6 This is a gate-level circuit diagram of a full adder provided in Example 3 of the present invention.

[0027] Figure 7 Schematic diagram of the gate-level circuit of the precise 4-2 compressor provided in Example 3 of the present invention.

[0028] Figure 8 Schematic diagram of stage 1 of the approximate multiplication operation provided in embodiment 3 of the present invention.

[0029] Figure 9 This is a schematic diagram showing the devices involved in the approximate multiplication operation provided in Example 3 of the present invention.

[0030] Figure 10 This is a schematic diagram of stage 2 of the approximate multiplication operation provided in embodiment 3 of the present invention.

[0031] Figure 11 Schematic diagram of stages three and four of the approximate multiplication operation provided in embodiment 3 of the present invention. DETAILED DESCRIPTION

[0032] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0033] Example 1 This embodiment provides a memristor-based 4-2 approximate compressor for approximate compression of four 1-bit numbers X1, X2, X3, and X4, including: two identical first majority gate circuits and a second majority gate circuit; The first majority gate circuit and the second majority gate circuit both include: a voltage comparator, a resistor, and three identical memristors; the resistance value of the resistor is the low-resistance resistance value of the memristor; the negative electrodes of the three memristors are connected to one end of the resistor, and the positive electrodes of the three memristors serve as the first input end, the second input end, and the third input end of the corresponding majority gate circuit respectively; the other end of the resistor is used for grounding; the first input end of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator is used to output a first level when the voltage at its first input end is greater than the voltage at its second input end, and use the first level as the output result "1" of the corresponding majority gate circuit, and to output a second level when the voltage at its first input end is less than or equal to the voltage at its second input end, and use the second level as the output result "0" of the corresponding majority gate circuit; the output end of the voltage comparator serves as the output end of the corresponding majority gate circuit; The three memristors in the first majority gate circuit are used to store the resistance values ​​corresponding to X1, X3, and X4 respectively; the positive electrodes of the three memristors in the first majority gate circuit are all used to access the voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is used to access the reference voltage V Ref1 The output terminal of the first majority gate circuit is used to output the carry result of the 4-2 approximate compressor; The three memristors in the second majority gate circuit are used to store the resistance values ​​corresponding to X1, X2, and 1 respectively, and the positive electrode of the memristor storing the resistance value corresponding to X1 is used to access voltage 0, the positive electrode of the memristor storing the resistance value corresponding to X2 is used to access voltage V2, and the positive electrode of the memristor storing the resistance value corresponding to 1 is used to access voltage V2; the second input terminal of the voltage comparator in the second majority gate circuit is used to access the reference voltage V Ref2 The output terminal of the second majority gate circuit is used to output the pseudo-sum result of the 4-2 approximate compressor; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

[0034] It is understood that the correspondence between the voltage comparator output level and the logical value can be flexibly configured. For example, with the first input terminal as the positive terminal, the second input terminal as the negative terminal, the first level as a high level, and the second level as a low level, the high level is used as the result of a majority logic operation, "1", and the low level is used as the result of a majority logic operation, "0". Alternatively, with the first input terminal as the negative terminal, the second input terminal as the positive terminal, the first level as a low level, and the second level as a high level, the low level is used as the result of a majority logic operation, "1", and the high level is used as the result of a majority logic operation, "0".

[0035] The voltage comparator in this embodiment uses the first input terminal as the positive terminal, the second input terminal as the negative terminal, the first level as the high level, and the second level as the low level. The high level is used as the result of the majority gate logic operation "1", and the low level is used as the result of the majority gate logic operation "0".

[0036] The schematic diagram of the 4-2 approximate compressor in this embodiment is as follows Figure 1As shown. The 4-2 approximate compressor in this embodiment is composed of a majority gate (MAJ) and a majority gate with an inverted input (MAJF); the majority gate (MAJ) is the first majority gate circuit in this embodiment, and its corresponding logical operation is M(X1, X3, X4); the majority gate with an inverted input (MAJF) is the second majority gate circuit in this embodiment, and its corresponding logical operation is ; Among them, the carry result (Carry signal) is output by the first majority gate circuit, and the pseudo-sum result (Sum signal) is output by the second majority gate circuit with an inverted input.

[0037] The first majority gate circuit and the second majority gate circuit have the same structure. In this embodiment, the majority gate circuit is as follows: Figure 2 As shown, it includes: three memristors M1, M2 and M3, a fixed resistor and a voltage comparator; wherein, the positive electrodes of the memristors M1, M2 and M3 are controlled by the control terminals T1, T2 and T3 respectively, and one end of the fixed resistor is controlled by the control terminal T4; the other end of the fixed resistor is connected to the negative electrodes of the memristors M1, M2 and M3, and the voltage of the common node is represented by V Com Indicates that the control voltages of the other four control ports T1, T2, T3 and T4 are V T1 、V T2 、V T3 and V T4 The positive terminal of the voltage comparator is connected to the common node of the memristor and the series resistor, and the negative terminal is connected to the reference voltage V Ref , the output voltage is V o .

[0038] Memristor has two resistance states, namely high resistance state (HRS) and low resistance state (LRS). By applying a port voltage of a specific direction and magnitude, the memristor can switch between the two resistance states. set The memristor can be set from high resistance to low resistance by applying a forward voltage less than V reset The memristor can change from low resistance state to high resistance state by applying negative voltage V set V is the threshold voltage at which the memristor changes from a high-resistance state to a low-resistance state; reset is the threshold voltage at which the memristor changes from a low-resistance state to a high-resistance state.

[0039] The high and low configuration resistance values ​​of the memristor are R H and R L , R H < <R L In this embodiment, the resistance of the resistor is equal to the resistance of the low resistance of the memristor.H =100KΩ, low resistance R L =1KΩ, the resistance R=1KΩ.

[0040] The three input variables a, b, and c of the majority gate logic are mapped to the resistance states of the memristors M1, M2, and M3, respectively. The high resistance state of the memristor corresponds to the logic value "0", and the low resistance state of the memristor corresponds to the logic value "1". The output of the logic is the voltage value of the comparator output port, where the output V o = V+ indicates that the output is high level, and the corresponding logic result is "1"; the output V o = V-, indicating that the output is low level, and the corresponding logic result is "0".

[0041] Based on the above logic circuit, only the control voltage applied to the control ports T1, T2, T3 and T4 needs to be changed to realize M(a,b,c) and Two types of majority gate logic. The specific implementation steps are as follows: (1) Implementation of majority gate logic M(a,b,c): By applying voltages V1, V1, V1, 0 to the control ports T1, T2, T3 and T4 respectively, V Ref = V Ref1 The result of the logic operation is obtained through the output of the voltage comparator, and the port voltage V1 is equal to the reference voltage V Ref The following constraints must be met: 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3; the value of V1 must ensure that the logic value currently written in the memristor cannot be changed, and V Ref1 The value of must ensure that the voltage comparator can output the correct calculation result.

[0042] The principle of the logic implementation is that when at least two of the input memristors M1, M2, and M3 are in a low resistance state, the voltage V Com Greater than the comparator reference voltage V Ref1 , thus outputting a high level "1"; conversely, when at most one of the input memristors M1, M2, and M3 is in a low resistance state, the voltage V Com Less than the comparator reference voltage V Ref1 , thus outputting a low level "0". This realizes the majority gate logic M(a,b,c).

[0043] by Figure 3The equivalent circuit after writing the logical value to the memristor in the example further illustrates that the above method can obtain the correct logical operation result. The same applies to other combinations. In the second combination, a logical value "1" and two logical values ​​"0" are input and written into the memristors M1, M2 and M3 respectively. The memristor corresponding to the logical value "1" is in a low resistance state, and the memristor corresponding to the logical value "0" is in a high resistance state. Therefore, the equivalent circuit after writing the logical value is as follows: Figure 3 As shown, at this time, the voltage of the common node V Com for: That is, the voltage of the common node V Com Less than and close to V1 / 2, while V1 / 2<V Ref1 , therefore, the voltage V Com Less than V Ref1 , so the voltage comparator outputs a low level, indicating a logical value of "0", and the results of the majority gate logic operations are correct.

[0044] (2) Implementation of majority gate logic: By applying voltages 0, V2, V2, 0 to the control ports T1, T2, T3 and T4 respectively, V Ref = V Ref2 The result of the logic operation is obtained through the output of the voltage comparator, and the port voltage V2 is equal to the comparator reference voltage V Ref2 The following constraints must be met: 0<V2<V set ; V2 / 3<V Ref2 <V2 / 2, the value of V2 must ensure that the logic value currently written in the memristor cannot be changed, and V Ref2 The value of must ensure that the voltage comparator can output the correct calculation result.

[0045] The principle of logic implementation is that when the input memristors M2 and M3 are both low resistance "1", the voltage V Com , the voltage of the common node V Com Greater than the comparator reference voltage V Ref2 , thus outputting a high level "1"; and when the input memristors M2 and M3 have high resistance "0" and the input memristor M1 is low resistance "1", the voltage V of the common node will be lowered Com , the voltage of the common node V Com Less than the comparator reference voltage V Ref2 , thus outputting a low level "0". This is to achieve the majority gate logic .

[0046] This embodiment also provides a control method for the 4-2 approximate compressor, including: X1, X3, and X4 are written into the three memristors in the first majority gate circuit one by one. A voltage V1 is applied to the positive electrodes of the three memristors in the first majority gate circuit. The second input terminal of the voltage comparator in the first majority gate circuit is connected to the reference voltage V Ref1 , obtaining the carry result of the 4-2 approximate compressor at the output end of the first majority gate circuit; Write X1, X2, and 1 into the three memristors in the second majority gate circuit one by one. Apply a 0V voltage to the positive electrode of the memristor written with X1, and apply a voltage V2 to the positive electrodes of the memristors written with X2 and 1. Connect the reference voltage V to the second input of the voltage comparator in the second majority gate circuit. Ref2 , a pseudo-sum result of the 4-2 approximate compressor is obtained at the output end of the second majority gate circuit; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

[0047] This embodiment also provides a memristor-based 4-2 approximate compressor device, including the above-mentioned 4-2 approximate compressor and a controller provided in this embodiment; The controller is used to execute the control method of the above-mentioned 4-2 approximate compressor provided in this embodiment.

[0048] Example 2 like Figure 4 As shown, this embodiment provides a 4-2 approximate compressor based on a memristor, for approximate compression of four 1-bit numbers X1, X2, X3, and X4, including: two identical first majority gate circuits and a second majority gate circuit; The first majority gate circuit and the second majority gate circuit both include: a voltage comparator, a resistor, and three identical memristors; the resistance value of the resistor is the low-resistance resistance value of the memristor; the negative electrodes of the three memristors are connected to one end of the resistor, and the positive electrodes of the three memristors serve as the first input end, the second input end, and the third input end of the corresponding majority gate circuit respectively; the other end of the resistor is used for grounding; the first input end of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator is used to output a first level when the voltage at its first input end is greater than the voltage at its second input end, and use the first level as the output result "1" of the corresponding majority gate circuit, and to output a second level when the voltage at its first input end is less than or equal to the voltage at its second input end, and use the second level as the output result "0" of the corresponding majority gate circuit; the output end of the voltage comparator serves as the output end of the corresponding majority gate circuit; The three memristors in the first majority gate circuit are used to store the resistance values ​​corresponding to X3, X4, and 1 respectively; the positive electrodes of the three memristors in the first majority gate circuit are all used to access the voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is used to access the reference voltage V Ref1 The output terminal of the first majority gate circuit is used to output the carry result of the 4-2 approximate compressor; The three memristors in the second majority gate circuit are used to store the resistance values ​​corresponding to 0, X1, and X4 respectively, and their positive electrodes are used to connect to voltages 0, V2, and V2 respectively; the second input terminal of the voltage comparator in the second majority gate circuit is used to connect to the reference voltage V Ref2 The output terminal of the second majority gate circuit is used to output the pseudo-sum result of the 4-2 approximate compressor; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

[0049] Similar to Example 1, the 4-2 approximate compressor in this embodiment is also composed of a majority gate (MAJ) and a majority gate with an inverted input (MAJF); the majority gate (MAJ) is the first majority gate circuit in this embodiment, and the corresponding logic operation is M(X1, X3, X4); the majority gate with an inverted input (MAJF) is the second majority gate circuit in this embodiment, and the corresponding logic operation is ; Among them, the carry result (Carry signal) is output by the first majority gate circuit, and the pseudo-sum result (Sum signal) is output by the second majority gate circuit with an inverted input.

[0050] Unlike the 4-2 approximate compressor in Example 1, which uses a fully data-driven majority logic design, the 4-2 approximate compressor provided in this embodiment effectively constructs a structural logic bias by introducing fixed values ​​of "1" and "0" at the input. This effectively creates a structural logic bias, giving the pseudo-sum result a clear error direction. This facilitates targeted error guidance and neutralization in subsequent compression. When used for multiplication operations, it can optimize the error distribution of the overall multiplication output. Therefore, this design not only has low hardware complexity but also exhibits better error equalization capabilities within a controllable error structure.

[0051] The truth tables of the 4-2 approximate compressor (compressor 1) in Example 1 and the 4-2 approximate compressor (compressor 2) in Example 2 are shown in Table 1.

[0052] The structures and detailed descriptions of most gate circuits in this embodiment are the same as those in embodiment 1 and are not described in detail here.

[0053] This embodiment also provides a control method for the 4-2 approximate compressor, including: Write X3, X4, and 1 into the three memristors in the first majority gate circuit one by one. Apply voltage V1 to the positive electrodes of the three memristors in the first majority gate circuit. Connect the reference voltage V to the second input terminal of the voltage comparator in the first majority gate circuit. Ref1 , obtaining the carry result of the 4-2 approximate compressor at the output end of the first majority gate circuit; Write 0, X1, and X4 into the three memristors in the second majority gate circuit one by one. Apply a voltage of 0V to the positive electrode of the memristor written with 0, and apply a voltage of V2 to the positive electrodes of the memristors written with X1 and X4. Connect the reference voltage V to the second input of the voltage comparator in the second majority gate circuit. Ref2 , a pseudo-sum result of the 4-2 approximate compressor is obtained at the output end of the second majority gate circuit; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

[0054] This embodiment also provides a memristor-based 4-2 approximate compressor device, including the above-mentioned 4-2 approximate compressor and a controller provided in this embodiment; The controller is used to execute the control method of the above-mentioned 4-2 approximate compressor provided in this embodiment.

[0055] Example 3 As the core computing module in digital systems, the performance of multipliers is directly related to the overall data processing speed, energy consumption, and system throughput. Traditional multiplier designs are typically based on a partial product reduction (PPR) strategy, relying on a large number of basic units such as full adders (FAs), half adders (HAs), and 4:2 compressors to compress and merge multi-bit partial products in a step-by-step manner. However, in traditional CMOS processes, the compression network of multipliers typically requires deep logic gate cascades and multi-stage carry propagation, resulting in severe delay accumulation and a significant increase in power consumption. Furthermore, hardware resource consumption increases exponentially with larger bit widths, making it difficult to meet the requirements of low-power, highly integrated computing systems.

[0056] This embodiment provides an approximate multiplication method, including: Perform a logical AND operation on each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; construct a partial product compression tree with n rows and 2n-1 columns from the n×n partial products; n ≥ 4; The partial product compression tree is divided into an exact compression part of column n1, an approximate compression part of column n2, and a truncated part of column n3; the weights of the partial products in the exact compression part, the approximate compression part, and the truncated part decrease in sequence; n1+n2+n3=2n-1; n1, n2, and n3 are all positive integers; Approximately compressing the partial products in the approximate compression portion using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder; and accurately compressing the partial products in the accurate compression portion using an accurate 4-2 compressor, a half adder, and a full adder; The partial product of the truncated part is truncated to obtain a truncated result with n3 bits all zero; The partial products after exact compression are used as the high bits, and the partial products after approximate compression are used as the low bits to obtain a combination result; the partial products of each column in the combination result are added in the direction from low bits to high bits, and when there is a carry in the previous column, the carry result is added to the current column to obtain the final carry result c0 and the summation result of each column; the carry result c0, the summation results of each column, and the truncation result with n3 bits all zero are arranged in order from high bits to low bits to form the final approximate multiplication result; Among them, the first approximate 4-2 compressor is the 4-2 approximate compressor provided in Example 1 of the present invention; the second approximate 4-2 compressor is the 4-2 approximate compressor provided in Example 2 of the present invention. The first approximate 4-2 compressor omits some logic inputs while maintaining the main compression path. It has a simple structure and low area and power consumption overhead. It is suitable for intermediate weight column compression paths with low accuracy requirements. The second approximate 4-2 compressor is an error-guided approximate compressor. By introducing a fixed input "1" to construct an output bias, the output error has directionality and adjustability, and is suitable for use in error-guided or error mean-controlled approximate multiplier compression paths. This embodiment uses the above-mentioned first 4-2 approximate compressor and the second approximate 4-2 compressor to build a multi-stage compression network of multipliers, realizing an approximate multiplication operation scheme with controllable errors, reconfigurable structure, and good resource utilization.

[0057] The relevant technical solutions are the same as those in Embodiment 1 and Embodiment 2 of the present invention and will not be described in detail here.

[0058] Furthermore, in a preferred implementation of this embodiment, the partial products in the approximate compression part are approximate compressed using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder, including: Performing an approximate compression operation on the partial products in the approximate compression part by using a first approximate 4-2 compressor, a half adder, and a full adder to obtain a first approximate compression array of size 4×n2; Using a second approximate 4-2 compressor to perform an approximate compression operation on the partial products in the first approximate compression array, thereby obtaining a second approximate compression array of size 2×n2; When performing a compression operation on the portion to be compressed, the compression operation is performed column by column in the order from the last column to the first column of the portion to be compressed; the first column in the portion to be compressed has the highest weight value, and the last column has the lowest weight value; the portion to be compressed is an approximate compression portion or a first approximate compression array; The number of first approximate 4-2 compressors p used for the jth column of the approximate compressed part j , the number of full adders r j and the number of half adders q j All satisfy g j -4p j -3r j -2q j +p j +q j +r j +m j+1 =4;m j+1 =p j+1 +q j+1 +r j+1 ;j=1,2,......,n2-1;4≤g h ≤n;p h ≥0;q h ≥0; r h ≥0; h=1,2,......,n2; g h is the number of partial products in the hth column of the approximate compressed part; When a compression operation occurs on the hth column of the approximately compressed portion, the hth column of the first approximately compressed array includes: a pseudo-sum result obtained by compressing the hth column of the approximately compressed portion; When the hth column of the approximate compressed portion includes partial products that are not subjected to the compression operation, the hth column of the first approximate compressed array includes: the partial products that are not subjected to the compression operation in the hth column of the approximate compressed portion; When a compression operation occurs on the j+1th column of the approximately compressed portion, the jth column of the first approximately compressed array includes: a carry result after compressing the j+1th column of the approximately compressed portion; The n2th column of the second approximate compressed array includes: the pseudo-sum result after compressing the n2th column of the first approximate compressed array and 0; the jth column of the second approximate compressed array includes: the pseudo-sum result after compressing the jth column of the first approximate compressed array and the carry result after compressing the j+1th column of the first approximate compressed array.

[0059] Through the above design, the approximate compression of the partial products in the approximate compression part is divided into two stages. The compression in the first stage makes the generated first approximate compression array have 4 rows, which enables the second stage to use only the second approximate 4-2 compressor for compression, further reducing the error of the multiplication operation.

[0060] It should be noted that the precise 4-2 compressor can be any existing precise 4-2 compressor, such as the 4-2 compressor structure based on an optimized CMOS full adder proposed in the document "Improved CMOS (4;2) compressor designs for parallel multipliers," and is not limited here. The half adder can be any existing half adder, such as the CMOS-based half adder design proposed in the document "CMOS Half Adder Design & Simulation Using Different Foundry," and is not limited here. The full adder can be any existing full adder, such as the CMOS-based full adder design proposed in the document "Design and Analysis of CMOS Full Adder," and is not limited here.

[0061] In this embodiment, the precise 4-2 compressor, half adder, and full adder are all designed based on majority gate circuits.

[0062] like Figure 5 As shown, the half adder is used to add three 1-bit numbers a i 、b i 、c i Approximate compression is performed, including four identical first majority gate circuits, second majority gate circuits, third majority gate circuits and fourth majority gate circuits; the first majority gate circuit, the second majority gate circuit, the third majority gate circuit and the fourth majority gate circuit are all the same as the majority gate circuit in Example 1; the first majority gate circuit is used to implement the logical operation M(a i , b i , 0), outputs the carry result of the half adder (Carry signal); the second majority gate circuit is used to implement logical operations ; The third majority gate circuit is used to implement logical operations The outputs of the second majority gate circuit and the third majority gate circuit are both used as inputs of the fourth majority gate circuit, which is used to implement logical operations. , output the pseudo sum result of the half adder (Sum signal).

[0063] The above-mentioned half adder can directly realize the addition function without introducing traditional inverters and XOR gates, reducing the logic level and the number of devices, which is conducive to improving the computing speed and hardware integration.

[0064] like Figure 6 As shown, the full adder is used to add three 1-bit numbers a i 、b i 、c i The approximate compression is performed, comprising three identical first majority gate circuits, second majority gate circuits and third majority gate circuits; the first majority gate circuit, the second majority gate circuit and the third majority gate circuit are all the same as the majority gate circuit in Example 1; the first majority gate circuit is used to implement the logical operation M(a i ,b i , c i ), get the carry result (Carry signal) of the full adder; the second majority gate circuit is used to implement logical operations The outputs of the first majority gate circuit and the second majority gate circuit are both used as inputs of the fourth majority gate circuit, and the third majority gate circuit is used to implement logical operations , and get the pseudo-sum result of the full adder (Sum signal).

[0065] The above-mentioned full adder structure has a compact logic path and can synchronously generate carry and sum results, reducing the delay caused by carry propagation in traditional full adders and improving the overall computing speed. It is particularly suitable for application in storage and computing integrated or low-power systems that require efficient addition calculations.

[0066] The above-mentioned half adder and full adder structures can be efficiently implemented in most logic platforms, with shallow logic levels and low resource overhead, and are suitable for embedded designs of various compressor structures.

[0067] like Figure 7 As shown in the figure, the Accurate 4:2 Compressor consists of two cascaded full adders; the first stage full adder takes X1, X2, and X3 as inputs and outputs the intermediate carry Carry1 and the sum Sum1; the second stage full adder takes X4, C in, Sum1 as input, and outputs the final carry Carry2 and sum Sum2. Carry1 serves as the carry result (i.e., Cout signal) output by the precise 4-2 compressor, while Carry2 and Sum2 serve as the carry result (i.e., Carry signal) and sum result (i.e., Sum signal) of the precise 4-2 compressor, respectively. Through the two-stage series design of full adders, the precise 4:2 compressor can accurately compress four-bit input and one-bit carry input, outputting two-bit summation value and one-bit carry signal. Its logical function is consistent with that of a traditional CMOS-implemented 4:2 compressor. Furthermore, due to its construction based on majority gates and majority gates with inverted inputs, it further optimizes logic depth and hardware area, helping to reduce latency and improve energy efficiency. It is suitable for use in high-performance approximate multipliers or partial product compression networks. It should be noted that the full adder used in the precise 4-2 compressor can be any existing full adder, preferably the above-mentioned full adder based on majority gates. This structure is functionally equivalent to the traditional CMOS precise 4:2 compressor, has precise compression function, and is implemented based on majority gate logic, with good integrability and logical consistency.

[0068] This embodiment takes 8-bit binary approximate multiplication operation as an example to describe in detail, as follows: like Figure 8 、 Figure 10 and Figure 11 FIG. 1 is a schematic diagram showing the working principle of the approximate multiplication operation provided by this embodiment; wherein, Figure 8 is a schematic diagram of the first stage of the approximate multiplication operation; Figure 9 A schematic diagram showing the devices involved in the approximate multiplication operation; Figure 10 is a schematic diagram of the second stage of the approximate multiplication operation; Figure 11 Schematic diagram of stages three and four of approximate multiplication operation.

[0069] The entire multiplier operation process is divided into three stages. The first two stages involve the use of approximate compressors. To control the multiplier error within a certain range, the high-order partial products are approximated using an accurate compressor. Similarly, to reduce hardware usage, the low-order partial products are directly ignored. Only the partial products in the middle digits are approximated and compressed.

[0070] The approximate compressors used in the first stage of the approximation are all the first 4-2 approximate compressors mentioned above, and the compressors used in the second stage of the approximation are all the second approximate 4-2 compressors mentioned above. The sources of the operators in each stage of the approximation area are also marked on the figure. The hollow circles represent operators that are ignored or filled with 0. P means that the operator is the result of the direct multiplication of two multipliers, S h It represents the Sum of the half adder, C hIndicates the Carry of the half adder, S f It represents the Sum of the full adder, C f Indicates the Carry of the full adder, S c It represents the Sum of the approximate compressor, C c It represents the Carry,S of the approximate compressor. e It represents the Sum of the exact compressor, C e Indicates the Carry of the precise compressor, C o is the C of the precise compressor out , S l 、C l They are the Sum and Carry of the approximate compressor 2 used in the second stage.

[0071] Perform a logical AND operation on each bit of the 8-bit multiplier a7a6a5a4a3a2a1a0 and each bit of the 8-bit multiplicand b7b6b5b4b3b2b1b0 to obtain 64 partial products. The 64 partial products are constructed into a partial product compression tree with 8 rows and 15 columns, as shown in the following example: Figure 8 As shown. Among them, the first row in the partial product compression tree is P from high position (high weight value) to low position (low weight value). 07 P 06 P 05 P 04 P 03 P 02 P 01 P 00 , and so on, the last row in the partial product compression tree is P from high (high weight value) to low (low weight value) 77 P 76 P 75 P 74 P 73 P 72 P 71 P 70 Among them, P 07 It is the partial product of a0 and b7 after multiplication (and logical operation); P 00 It is the partial product after multiplication (and logical operation) of a0 and b0; P 70 is the partial product of a7 and b0 after multiplication (and logical operation); P 77 It is the partial product of a7 and b7 after multiplication (AND logical operation).

[0072] For stage 1, P 04 、P 13 As the input of the half adder, we get S h04 , C h15 .P 05 、P 14 、P 23、P 32 As the input of the first 4-2 approximation compressor, we get S c05 , C c26 .P 06 、P 15 、P 24 、P 33 As the input of the first 4-2 approximation compressor, we get S c06 , C c27 .P 42 、P 51 As the input of the half adder, we get S h16 , C h37 .P 07 、P 16 、P 25 、P 34 As the input of the first 4-2 approximation compressor, we get S c07 , C c28 .P 43 、P 52 、P 61 、P 70 As the input of the first 4-2 approximation compressor, we get S c17 , C c38 .P 17 、P 26 、P 35 、P 44 As the input of the first 4-2 approximation compressor, we get S c08 , C c29 .P 53 、P 62 、P 71 As the input of the full adder, we get S f18 , C f39 .P 27 、P 36 、P 45 、P 54 As the input of the first 4-2 approximation compressor, we get S c09 , C c110 .P 63 、P 72 As the input of the half adder, we get S h19 , C h210 .P 37 、P 46 、P 55 、P 64 As the input of the exact 4-2 compressor, we get S e010 , C e111 , C o211 .P 47 、P 56 、P 65 As the input of the full adder, we get Sf011 , C f012 .

[0073] For stage 2, S h04 、P 22 、P 31 、P 40 As the input of the second 4-2 approximation compressor, we get S l4 , C l4 . S c05 、C h15 、P 41 、P 50 As the input of the second 4-2 approximation compressor, we get S l5 , C l5 . S c06 、S h16 、C c26 、P 60 As the input of the second 4-2 approximation compressor, we get S l6 , C l6 . S c07 、S c17 、C c27 、C h37 As the input of the second 4-2 approximation compressor, we get S l7 , C l7 . S c08 、S f18 、C c28 、C c38 As the input of the second 4-2 approximation compressor, we get S l8 , C l8 . S c09 、S h19 、C c29 、C f39 As the input of the second 4-2 approximation compressor, we get S l9 , C l9 . S e010 、C c110 、C h210 、P 73 As the input of the exact 4-2 compressor, we get S e10 , C e10 , C o10 . S f011 、C e111 、C o211 、P 74 As the input of the exact 4-2 compressor, we get S e11 , C e11 , C o11 . C f012 、P 57 、P 66 、P 75As the input of the exact 4-2 compressor, we get S e12 , C e12 , C o12 .P 67 、P 76 As the input of the half adder, we get S h13 , C h13 .

[0074] For stage three, S e11 、C e10 、C o10 As the input of the full adder, we get S f11 、C f11 ;S e12 、C e11 、C o11 As the input of the full adder, we get S f12 、C f12 ;S h13 、C e12 、C o12 As the input of the full adder, we get S f13 、C f13 .

[0075] For stage 4, all operators are summed using full adders to obtain the final result.

[0076] It should be noted that there are many ways to accurately compress the partial products in the accurate compression part, which can be achieved by using an accurate 4-2 compressor, a full adder, and a half adder. The above method is only one of the preferred methods, but not the only one. In the second stage of the above method, the carry bit of the input of the accurate 4-2 compressor is defaulted to 0, which makes the speed of the approximate multiplication operation relatively fast; but in other implementations, when the accurate 4-2 compressor is used to compress the column with the lowest weight in the compressed part, the carry bit of the input is defaulted to 0, and when the accurate 4-2 compressor is used to compress other columns, the carry bit is C when the accurate 4-2 compressor is used to compress the previous column. out .

[0077] It's important to note that we first assume that the multipliers 0 and 1 are equally likely. Therefore, for the partial product P, the probability of P = 1 is 0.25, and the probability of P = 0 is 0.75. Based on this, we can calculate the error caused by using the first approximation 4-2 compressor in stage 1: Because all four inputs to each compressor in stage 1 are P, the expected error for each compressor is 0.125. If the weight of the first column in the approximation area is set to unity, the total expected error for stage 1 is Error1 = 0 × 1 + 0.125 × 2 + 0.125 × 4 + 0.25 × 8 + 0.125 × 16 + 0.125 × 32 = 8.75. The error for stage 2 can be calculated similarly, but since the inputs to columns with different weights are different in stage 2, the error for each column must be calculated separately, which we will not elaborate on here. The expected error for stage 2 is Error2 = -8.75 = -Error1.

[0078] In traditional approximate computing circuit design, a single compressor structure usually has a fixed-direction bias error (for example, the overall bias output is too high or too low). After multiple columns are superimposed in the multiplier, it is easy to cause a systematic error offset in the overall output, affecting the final calculation accuracy.

[0079] The present invention proposes to design approximate compressors with different expected error directions (the error expectation of the first 4-2 approximate compressor is negative, and the error expectation of the second 4-2 approximate compressor is positive), and further reasonably match and apply them in the multiplier column arrangement, and use the positive and negative errors to complement and offset each other between columns, so as to effectively balance the mean error of the overall multiplication operation result and make the output error expectation approach zero.

[0080] By configuring error expectations in a complementary manner, we can not only avoid systematic drift caused by accumulation in a single direction, but also improve the overall stability and robustness of the multiplier output while maintaining low hardware resource overhead. This strategy is particularly suitable for applications that are sensitive to mean error but can tolerate local random errors, such as neural network reasoning, image processing, and edge computing.

[0081] Furthermore, compared with the design scheme that fully uses precise compressors, the present invention can significantly reduce the logic cascade depth and the number of devices, reduce power consumption and area overhead by reasonably introducing approximate calculations. At the same time, it achieves overall controllable accuracy through the error compensation mechanism, taking into account both energy efficiency and calculation accuracy, and has high engineering application value in resource-constrained systems.

[0082] Example 4 An approximate multiplication operation device includes: a partial product generation module, an accurate partial compression module, an approximate partial compression module, a truncation processing module and a carry adder module; The partial product generation module is used to perform a logical AND operation on each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; the n×n partial products are constructed into a partial product compression tree with n rows and 2n-1 columns; n ≥ 4; the partial product compression tree is divided into an exact compression part of column n1, an approximate compression part of column n2, and a truncated part of column n3; the weights of the partial products in the exact compression part, the approximate compression part, and the truncated part decrease in sequence; n1+n2+n3=2n-1; n1, n2, and n3 are all positive integers; The precise partial compression module is used to approximately compress the partial products in the approximate compression part using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder; and precisely compress the partial products in the precise compression part using a precise 4-2 compressor, a half adder, and a full adder; The truncation processing module is used to truncate the partial product of the truncation part to obtain a truncation result with n3 bits all zero; The carry adder module is used to arrange the partial products after exact compression as the high bits and the partial products after approximate compression as the low bits to obtain a combination result; the partial products of each column in the combination result are added in the direction from low bits to high bits, and when there is a carry in the previous column, the carry result is added to the current column to obtain the final carry result c0 and the summation result of each column; the carry result c0, the summation results of each column, and the truncation result of n3 bits all zero are arranged in order from high bits to low bits to form the final approximate multiplication result; The first approximate 4-2 compressor is the 4-2 approximate compressor provided in Example 1 of the present invention; the second approximate 4-2 compressor is the 4-2 approximate compressor provided in Example 2 of the present invention.

[0083] The relevant technical solutions are the same as those in Embodiment 1, Embodiment 2 and Embodiment 3 of the present invention and are not described in detail here.

[0084] In summary, the present invention focuses on the structural optimization requirements of approximate multipliers and proposes a method for constructing a compressor based on majority gate logic and majority gate logic with inverted input. By designing approximate compressors with different error characteristics and rationally configuring them according to column weights and error distribution rules, the present invention effectively achieves directional compensation and mean control of the overall error of the multiplier. The proposed multi-level compression structure ensures the predictability and stability of the output accuracy while reducing hardware overhead and delay, and avoids the problem of systematic error accumulation in traditional multipliers. With the help of the unified implementation characteristics of majority gate logic, the present invention shows obvious advantages in device utilization, energy efficiency and scalability, and is suitable for promotion and application in new computing systems with strict requirements on power consumption, area and computational fault tolerance.

[0085] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A 4-2 approximation compressor based on a memristor, characterized in that: include: Used for approximate compression of four 1-bit numbers X1, X2, X3, and X4, comprising: two identical first majority gate circuits and a second majority gate circuit; The first majority gate circuit and the second majority gate circuit both include: a voltage comparator, a resistor, and three identical memristors; the resistance value of the resistor is the low-resistance resistance value of the memristor; the negative electrodes of the three memristors are connected to one end of the resistor, and the positive electrodes of the three memristors serve as the first input terminal, the second input terminal, and the third input terminal of the corresponding majority gate circuit respectively; the other end of the resistor is grounded; the first input terminal of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator is configured to output a first level when the voltage at its first input terminal is greater than the voltage at its second input terminal, and use the first level as the output result "1" of the corresponding majority gate circuit; and output a second level when the voltage at its first input terminal is less than or equal to the voltage at its second input terminal, and use the second level as the output result "0" of the corresponding majority gate circuit; the output terminal of the voltage comparator serves as the output terminal of the corresponding majority gate circuit; The three memristors in the first majority gate circuit are used to store the resistance values ​​corresponding to X1, X3, and X4 respectively; the positive electrodes of the three memristors in the first majority gate circuit are all used to access the voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is used to access the reference voltage V Ref1 The output terminal of the first majority gate circuit is used to output the carry result of the 4-2 approximate compressor; The three memristors in the second majority gate circuit are used to store the resistance values ​​corresponding to X1, X2, and 1 respectively, and their positive electrodes are used to connect to voltages 0, V2, and V2 respectively; the second input terminal of the voltage comparator in the second majority gate circuit is used to connect to the reference voltage V Ref2 The output terminal of the second majority gate circuit is used to output the pseudo-sum result of the 4-2 approximate compressor; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

2. The control method of the 4-2 approximation compressor according to claim 1, characterized in that: include: Write X1, X3, and X4 into the three memristors in the first majority gate circuit of the 4-2 approximate compressor one by one, apply voltage V1 to the positive electrodes of the three memristors in the first majority gate circuit, and connect the reference voltage V to the second input terminal of the voltage comparator in the first majority gate circuit. Ref1 , obtaining a carry result of the 4-2 approximate compressor at an output end of the first majority gate circuit; Write X1, X2, and 1 into the three memristors in the second majority gate circuit of the 4-2 approximate compressor one by one, apply a 0V voltage to the positive electrode of the memristor written with X1, apply a voltage V2 to the positive electrodes of the memristors written with X2 and 1, and connect the reference voltage V to the second input terminal of the voltage comparator in the second majority gate circuit. Ref2 , obtaining a pseudo-sum result of the 4-2 approximate compressor at an output end of the second majority gate circuit; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

3. A 4-2 approximation compressor based on a memristor, characterized in that: Used for approximate compression of four 1-bit numbers X1, X2, X3, and X4, comprising: two identical first majority gate circuits and a second majority gate circuit; The first majority gate circuit and the second majority gate circuit both include: a voltage comparator, a resistor, and three identical memristors; the resistance value of the resistor is the low-resistance resistance value of the memristor; the negative electrodes of the three memristors are connected to one end of the resistor, and the positive electrodes of the three memristors serve as the first input terminal, the second input terminal, and the third input terminal of the corresponding majority gate circuit respectively; the other end of the resistor is grounded; the first input terminal of the voltage comparator is connected to the common node of the three memristors and the resistor; the voltage comparator is configured to output a first level when the voltage at its first input terminal is greater than the voltage at its second input terminal, and use the first level as the output result "1" of the corresponding majority gate circuit; and output a second level when the voltage at its first input terminal is less than or equal to the voltage at its second input terminal, and use the second level as the output result "0" of the corresponding majority gate circuit; the output terminal of the voltage comparator serves as the output terminal of the corresponding majority gate circuit; The three memristors in the first majority gate circuit are used to store the resistance values ​​corresponding to X3, X4, and 1 respectively; the positive electrodes of the three memristors in the first majority gate circuit are all used to access the voltage V1; the second input terminal of the voltage comparator in the first majority gate circuit is used to access the reference voltage V Ref1 The output terminal of the first majority gate circuit is used to output the carry result of the 4-2 approximate compressor; The three memristors in the second majority gate circuit are used to store the resistance values ​​corresponding to 0, X1, and X4 respectively, and their positive electrodes are used to connect to voltages 0, V2, and V2 respectively; the second input terminal of the voltage comparator in the second majority gate circuit is used to connect to the reference voltage V Ref2 The output terminal of the second majority gate circuit is used to output the pseudo-sum result of the 4-2 approximate compressor; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

4. The 4-2 approximation compressor according to claim 3, characterized in that include: Write X3, X4, and 1 into the three memristors in the first majority gate circuit of the 4-2 approximate compressor one by one, apply voltage V1 to the positive electrodes of the three memristors in the first majority gate circuit, and connect the reference voltage V to the second input terminal of the voltage comparator in the first majority gate circuit. Ref1 , obtaining a carry result of the 4-2 approximate compressor at an output end of the first majority gate circuit; Write 0, X1, and X4 into the three memristors in the second majority gate circuit one by one. Apply a voltage of 0V to the positive electrode of the memristor written with 0, and apply a voltage of V2 to the positive electrodes of the memristors written with X1 and X4. Connect the reference voltage V to the second input terminal of the voltage comparator in the second majority gate circuit. Ref2 , obtaining a pseudo-sum result of the 4-2 approximate compressor at an output end of the second majority gate circuit; Among them, 0<V1<V set ; V1 / 2<V Ref1 <2V1 / 3;0<V2<V set ; V2 / 3<V Ref2 <V2 / 2;V set is the threshold value at which the memristor changes from a high-resistance state to a low-resistance state.

5. A 4-2 approximation compressor device based on a memristor, characterized in that: The 4-2 approximation compressor device comprises the 4-2 approximation compressor according to claim 1 and a controller for executing the control method according to claim 2; or, The 4-2 approximation compressor device includes the 4-2 approximation compressor described in claim 3 and a controller for executing the control method described in claim 4.

6. An approximate multiplication method, characterized in that: include: Perform a logical AND operation on each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; The n×n partial products are constructed into a partial product compression tree with n rows and 2n-1 columns; n≥4; The partial product compression tree is divided into an exact compression part of column n1, an approximate compression part of column n2, and a truncated part of column n3; the weights of the partial products in the exact compression part, the approximate compression part, and the truncated part decrease in sequence; n1+n2+n3=2n-1; n1, n2, and n3 are all positive integers; Approximately compressing the partial products in the approximate compression portion using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder; and accurately compressing the partial products in the accurate compression portion using an accurate 4-2 compressor, a half adder, and a full adder; The partial product of the truncated part is truncated to obtain a truncated result with n3 bits all zero; The partial products after exact compression are used as the high bits, and the partial products after approximate compression are used as the low bits to obtain a combination result; the partial products of each column in the combination result are added in the direction from low bits to high bits, and when there is a carry in the previous column, the carry result is added to the current column to obtain the final carry result c0 and the summation result of each column; the carry result c0, the summation results of each column, and the truncation result with n3 bits all zero are arranged in order from high bits to low bits to form the final approximate multiplication result; Wherein, the first approximate 4-2 compressor is the 4-2 approximate compressor described in claim 1; the second approximate 4-2 compressor is the 4-2 approximate compressor described in claim 3.

7. The approximate multiplication method according to claim 6, wherein: The method of performing approximate compression on the partial products in the approximate compression part by using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder comprises: Performing an approximate compression operation on the partial products in the approximate compression part by using a first approximate 4-2 compressor, a half adder, and a full adder to obtain a first approximate compression array of size 4×n2; Using a second approximate 4-2 compressor to perform an approximate compression operation on the partial products in the first approximate compression array, thereby obtaining a second approximate compression array of size 2×n2; When performing a compression operation on the portion to be compressed, the compression operation is performed column by column in the order from the last column to the first column of the portion to be compressed; the first column in the portion to be compressed has the highest weight value, and the last column has the lowest weight value; the portion to be compressed is an approximate compression portion or a first approximate compression array; The number of first approximate 4-2 compressors p used for the jth column of the approximate compressed part j , the number of full adders r j and the number of half adders q j All satisfy g j -4p j -3r j -2q j +p j +q j +r j +m j+1 =4;m j+1 =p j+1 +q j+1 +r j+1 ;j=1,2,......,n2-1;4≤g h ≤n;p h ≥0;q h ≥0; r h ≥0; h=1,2,......,n2; g h is the number of partial products in the hth column of the approximate compressed part; When a compression operation occurs on the hth column of the approximately compressed portion, the hth column of the first approximately compressed array includes: a pseudo-sum result obtained by compressing the hth column of the approximately compressed portion; When the hth column of the approximate compressed portion includes partial products that are not subjected to the compression operation, the hth column of the first approximate compressed array includes: the partial products that are not subjected to the compression operation in the hth column of the approximate compressed portion; When a compression operation occurs on the j+1th column of the approximately compressed portion, the jth column of the first approximately compressed array includes: a carry result after compressing the j+1th column of the approximately compressed portion; The n2th column of the second approximate compressed array includes: the pseudo-sum result after compressing the n2th column of the first approximate compressed array and 0; the jth column of the second approximate compressed array includes: the pseudo-sum result after compressing the jth column of the first approximate compressed array and the carry result after compressing the j+1th column of the first approximate compressed array.

8. The approximate multiplication method according to claim 7, wherein: When the j-th column of the first approximate compressed array includes a pseudo-sum result obtained by compressing the j-th column of the approximate compressed part, a carry result after compressing the j+1-th column of the approximate compressed part, and a partial product of the j-th column of the approximate compressed part that is not subjected to the compression operation, the pseudo-sum result obtained by compressing the j-th column of the approximate compressed part, the carry result after compressing the j+1-th column of the approximate compressed part, and the partial product of the j-th column of the approximate compressed part that is not subjected to the compression operation are arranged sequentially along the column direction; When the j-th column of the first approximate compression array contains only the pseudo-sum result obtained by compressing the j-th column of the approximate compression part and the carry result after compressing the j+1-th column of the approximate compression part, the pseudo-sum result obtained by compressing the j-th column of the approximate compression part and the carry result after compressing the j+1-th column of the approximate compression part are arranged sequentially along the column direction; When the j-th column of the first approximate compressed array contains only the carry result after compression of the j+1-th column of the approximate compressed part and the partial product of the j-th column of the approximate compressed part that is not subjected to the compression operation, the carry result after compression of the j+1-th column of the approximate compressed part and the partial product of the j-th column of the approximate compressed part that is not subjected to the compression operation are arranged sequentially along the column direction; When the h-th column of the first approximate compressed array includes only the pseudo-sum result obtained by compressing the h-th column of the approximate compressed part and the partial product of the j-th column of the approximate compressed part without the compression operation, the pseudo-sum result obtained by compressing the h-th column of the approximate compressed part and the partial product of the j-th column of the approximate compressed part without the compression operation are arranged sequentially along the column direction; When there are multiple compression operations on the h-th column of the approximately compressed part, there are multiple pseudo-sum results obtained after compressing the h-th column of the approximately compressed part, and the arrangement order of the pseudo-sum results in the h-th column of the first approximate compression array is consistent with the sorting order of the compressed objects of the corresponding compression operations in the h-th column of the approximately compressed part; When there are multiple partial products that are not subjected to the compression operation in the h-th column of the approximate compressed part, their arrangement order in the h-th column of the first approximate compressed array is consistent with their sorting order in the h-th column of the approximate compressed part; When the first approximate 4-2 compressor is used for compression, the compression object is the four partial products arranged in sequence along the column direction in a column of the approximate compressed part, which are respectively used as 1-bit numbers X1, X2, X3, and X4 in a one-to-one correspondence and input into the first approximate 4-2 compressor for compression; When the second approximate 4-2 compressor is used for compression, the compression objects are the four partial products arranged in sequence along the column direction in a column of the first approximate compression array, which are input one-to-one as 1-bit numbers X1, X2, X3, and X4 respectively and compressed into the second approximate 4-2 compressor for compression.

9. The approximate multiplication method according to claim 8, wherein: n=8; the n-3 columns with the highest weight in the partial product compression tree are used as the exact compression part, the n-2 columns with the second highest weight in the partial product compression tree are used as the approximate compression part, and the 4 columns with the lowest weight in the partial product compression tree are used as the truncated part; The n-bit multiplier is a7a6a5a4a3a2a1a0; the n-bit multiplicand is b7b6b5b4b3b2b1b0; The first to eighth rows in the partial product compression tree, sorted in descending order of weight values, are: P 07 P 06 P 05 P 04 P 03 P 02 P 01 P 00 、P 17 P 16 P 15 P 14 P 13 P 12 P 01 P 00 、P 27 P 26 P 25 P 24 P 23 P 22 P 21 P 20 、P 37 P 36 P 35 P 34 P 33 P 32 P 31 P 30 、P 47 P 46 P 45 P 44 P 43 P 42 P 41 P 40 、P 57 P 56 P 55 P 54 P 53 P 52 P 51 P 50 、P 67 P 66 P 65 P 64 P 63 P 62 P 61 P 60 、P 77 P 76 P 75 P 74 P 73 P 72 P 71 P 70 ; Among them, P 07 is the partial product of a0 and b7, P 06 is the partial product of a0 and b6, and so on. 70 is the partial product of a7 and b0; The approximate compression part includes 6 columns of partial products; the partial products in the first column are as follows from top to bottom: P 27 P 36 P 45 P 54 P 63 P 72 ; The partial products in the second column are from top to bottom: P 17 P 26 P 35 P 44 P 53 P 62 P 71 ; The partial products in the third column from top to bottom are: P 07 P 16 P 25 P 34 P 43 P 52 P 61 P 70 ; The partial products in the fourth column from top to bottom are: P 06 P 15 P 24 P 33 P 42 P 51 P 60 ; The partial products in the fifth column from top to bottom are: P 05 P 14 P 23 P 32 P 41 P 50 ; The partial products in the sixth column from top to bottom are: P 04 P 13 P 22 P 31 P 40 ; The first column in the approximate compression part has the highest weight value, and the last column has the lowest weight value; The method of performing an approximate compression operation on the partial products in the approximate compression part by using a first approximate 4-2 compressor, a half adder, and a full adder to obtain a first approximate compression array of size 4×n2 includes: Using half adder to P 04 and P 13 Compression is performed to obtain the pseudo-sum result S h04 and carry result C h15 ;P 05 、P 14 、P 23 、P 32 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c05 and carry result C c26 ;P 06 、P 15 、P 24 、P 33 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c06 and carry result C c27 ; Use half adder to P 42 and P 51 Compression is performed to obtain the pseudo-sum result S h16 and carry result C h37 ;P 07 、P 16 、P 25 、P 34 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c07 and carry result C c28 ;P 43 、P 52 、P 61 、P 70 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c17 and carry result C c38 ;P 17 、P 26 、P 35 、P 44 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c08 and carry result C c29 ; Use full adder to P 62 and P 71 Compress and get the result S f18 and carry result C f39 ;P 27 、P 36 、P 45 、P 54 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. c09 and carry result C c110 ; Use half adder to P 63 and P 72 Compression is performed to obtain the pseudo-sum result S h19 and carry result C h210 ;P 37 、P 46 、P 55 、P 64 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. e010 and carry result C f012 ; The first column of the first approximate compressed array is S c09 S h19 C c29 C f39 , the second column is S c08 S h18 C c28 C f38 , the third column is S c07 S c17 C c27 C h37 , the fourth column is S c06 S h16 C c26 P 60 , the fifth column is S c05 C h15 P 41 P 50 , the sixth column is S h04 P 22 P 31 P 40 ; The method of using a second approximate 4-2 compressor to perform an approximate compression operation on the partial products in the first approximate compressed array to obtain a second approximate compressed array of size 2×n2 includes: S h04 、P 22 、P 31 、P 40 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l4 and carry result C l4 ; S c05 、C h15 、P 41 、P 50 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l5 and carry result C l5 ; S c06 、S h16 、C c26 、P 60 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l6 and carry result C l6 ; S c07 、S c17 、C c27 、C h37 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l7 and carry result C l7 ; S c08 、S f18 、C c28 、C c38 They are input into the first approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l8 and carry result C l8 ; S c09 、S h19 、C c29 、C f39 They are input into the second approximate 4-2 compressor as 1-bit numbers X1, X2, X3, and X4 respectively for compression, and the pseudo-sum result S is obtained. l9 and carry result C l9 ; The first column of the second approximate compressed array thus obtained includes S l9 and C l8 , the second column includes S l8 and C l7 , the third column includes S l7 and C l6 , the fourth column includes S l6 and C l5 , the fifth column includes S l5 and C l4 , the sixth column includes S l4 and 0.

10. An approximate multiplication operation device, characterized in that: include: A partial product generation module, an exact partial compression module, an approximate partial compression module, a truncation processing module, and a carry adder module; The partial product generation module is used to perform a logical AND operation on each bit of the n-bit multiplier and each bit of the n-bit multiplicand to obtain n×n partial products; the n×n partial products are constructed into a partial product compression tree with n rows and 2n-1 columns; n≥4; the partial product compression tree is divided into the exact compression part of column n1, the approximate compression part of column n2, and the truncated part of column n3; the weights of the partial products in the exact compression part, the approximate compression part, and the truncated part decrease in sequence; n1+n2+n3=2n-1; n1, n2, and n3 are all positive integers; The precise partial compression module is used to approximately compress the partial products in the approximate compression part using a first approximate 4-2 compressor, a second approximate 4-2 compressor, a half adder, and a full adder; and precisely compress the partial products in the precise compression part using a precise 4-2 compressor, a half adder, and a full adder; The truncation processing module is used to truncate the partial product of the truncation part to obtain a truncation result with n3 bits all zero; The carry adder module is used to arrange the partial products after exact compression as the high bits and the partial products after approximate compression as the low bits to obtain a combination result; the partial products of each column in the combination result are added in the direction from low bits to high bits, and when there is a carry in the previous column, the carry result is added to the current column to obtain the final carry result c0 and the summation result of each column; the carry result c0, the summation results of each column, and the truncation result of n3 bits all zero are arranged in order from high bits to low bits to form the final approximate multiplication result; Wherein, the first approximate 4-2 compressor is the 4-2 approximate compressor described in claim 1; the second approximate 4-2 compressor is the 4-2 approximate compressor described in claim 3.

Citation Information

Patent Citations

  • Memristor-based hardware convolutional neural network model suitable for being realized on FPGA (Field Programmable Gate Array)

    CN117010466A

  • Approximate addition circuit control method based on memristor and approximate addition operation device

    CN119883182A

  • Control method of approximate adder circuit based on memristor and approximate adder device

    CN119883183A