Approximate floating point multiplier, chip and computing device

By introducing an approximate mantissa multiplier and an approximate 4-2 compressor that compensates each other in the floating-point multiplier, the energy efficiency and complexity problems of floating-point multiplication approximate calculation in the prior art are solved, and optimization in terms of accuracy, power consumption and area are achieved.

CN120104094APending Publication Date: 2025-06-06NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510030205.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When implementing floating point multiplication approximation calculations, it is difficult for the prior art to effectively utilize the characteristics of floating point multiplication, resulting in optimization space in terms of overhead such as accuracy, power consumption and area. The circuit of the traditional approximation 4-2 compressor is relatively complex and has low compression efficiency.

Method used

An approximate floating point multiplier is designed, using a floating point multiplier body with an approximate mantissa multiplier. By implementing a bit-by-bit "Agree" operation of the input operand and a partial product array with the logic unit and the compressor unit, the approximate 4-2 compressor and high-precision 4-2 compressor designed with neglected carry are used to reduce the total error by mutual compensation.

Benefits of technology

It realizes optimization in terms of overhead such as accuracy, power consumption and area, reduces the complexity and error rate of traditional approximate 4-2 compressors, improves the energy efficiency of floating point multiplication, and is suitable for applications with half-precision floating point operation requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104094A_ABST
    Figure CN120104094A_ABST
Patent Text Reader

Abstract

The invention discloses an approximate floating point multiplier, a chip and computing equipment, an approximate mantissa multiplier of the approximate floating point multiplier comprises an AND logic unit and a compressor unit, and the AND logic unit is used for performing AND operation on two input operands bit by bit to generate a partial product array with the size of 11 rows and 21 columns; the compressor unit is used for compressing the 11th column to the 21st column by column to obtain a final approximate mantissa, the compressor unit comprises two novel approximate 4-2 compressors ignoring carry design, and the error rate of the approximate 4-2 compressors is within an acceptable range by utilizing mutual compensation inside the compressors. The invention aims to excavate and use the characteristics of floating point multiplication to further improve the energy efficiency of floating point multiplication, realize the optimization of the approximate floating point multiplier in the overhead aspects of precision, power consumption, area and the like, and solve the problems of relatively complex circuit and low compression efficiency of the traditional approximate 4-2 compressor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of microprocessor design, and in particular to an approximate floating-point multiplier, a chip and a computing device. Background Art

[0002] With the rapid development of large-scale applications such as scientific computing, multimedia, the Internet of Things, and artificial intelligence, the total energy consumption of computer systems is growing at an alarming rate. In order to process the ever-increasing amount of information, energy efficiency has become the key to solving the continued growth in power consumption. However, as Dennard scaling has almost failed and Moore's Law has gradually come to an end, the method of significantly reducing power consumption, area and other resource overheads by reducing the size of transistors has become unsustainable. In order to cope with the increasing performance and power consumption challenges, researchers are actively exploring a new computing paradigm - approximate computing. Approximate computing reduces the logical complexity of the design and simplifies the hardware structure by sacrificing a certain degree of accuracy, thereby significantly reducing the power consumption, latency and area of ​​the hardware.

[0003] In computationally intensive applications such as neural networks and image processing, multiplication is the most typical and most frequently occurring operation. Fully accurate multiplication calculations consume a lot of power and latency. At the same time, such applications have a certain degree of fault tolerance, so approximate multipliers are increasingly attracting the attention of researchers in the field of approximate computing. Current research on approximate multiplication focuses on fixed-point design, and there is a lack of approximate research on floating-point multiplication, especially 16-bit half-precision floating-point (FP16) multiplication. The 16-bit half-precision floating-point number defined by the IEEE 754 standard consists of three parts: a 1-bit sign bit, a 5-bit offset exponent bit, and a 10-bit truncated mantissa bit. The 10-bit truncated mantissa bits plus a leading hidden bit "1" form a complete 11-bit mantissa. The multiplication of two floating-point numbers involves the multiplication of the mantissa bits, the XOR of the sign bit, the addition of the exponent bits, and the normalization and rounding steps.

[0004] Although many fixed-point approximate design methods can be applied to floating-point approximate design, the characteristics of floating-point multiplication itself have not been fully utilized, resulting in the fixed-point approximate design methods not being able to achieve the best results in floating-point multiplication. Therefore, how to further improve the energy efficiency of floating-point multiplication by utilizing the characteristics of floating-point multiplication itself has become a key technical issue that needs to be solved urgently. Summary of the invention

[0005] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, an approximate floating-point multiplier, chip and computing device are provided. The present invention aims to further improve the energy efficiency of floating-point multiplication by exploiting the characteristics of floating-point multiplication itself, achieve the optimization of the approximate floating-point multiplier in terms of accuracy, power consumption, area and other overheads, and solve the problem of relatively complex circuits and low compression efficiency of traditional approximate 4-2 compressors.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: An approximate floating-point multiplier comprises a floating-point multiplier body with an approximate mantissa multiplier, wherein the approximate mantissa multiplier comprises an AND logic unit and a compressor unit, wherein the AND logic unit is used to perform an AND operation on two operands input to the approximate floating-point multiplier bit by bit to generate a partial product array with a size of 11 rows and 21 columns, wherein each row in the partial product array is 11 valid partial products, and the compressor unit is used to compress the partial product arrays of columns 11 to 21 obtained by truncating columns 1 to 10 by column to obtain a final approximate mantissa, and the compressor unit comprises an approximate 4-2 compressor with a carry-ignoring design, wherein the approximate 4-2 compressor with a carry-ignoring design makes the error rate within an acceptable range by utilizing mutual compensation inside the compressor.

[0007] Optionally, the carry-ignoring approximate 4-2 compressor comprises an approximate 4-2 compressor AC 1 , the approximate 4-2 compressor AC 1 includes an OR gate, and the approximate 4-2 compressor AC 1 Ignore the input operand x when performing the operation 2 , the input operand x 1 As the output sum, the input operand x 3 and x 4 After passing through the OR gate, the carry bit is used as the output, and the approximate 4-2 compressor AC 1 The truth table is as follows: ① When the operand x 2 and x 1 When the operand x is 00, 4 and x 3 If the operand x is 00, the output is 00. 4 and x 3 If the operand x is 01, the output is 10. 4 and x 3 If the operand x is 10, the output is 10. 4 and x 3 If the operand x is 11, the output is 10; ②When the operand x 2 and x 1 When it is 01, if the operand x 4 and x 3 If the operand x is 00, the output is 01. 4 and x 3 If the operand x is 01, the output is 11. 4 and x 3 If the operand x is 10, the output is 11. 4 and x 3 If the operand x is 11, the output is 11; 2and x 1 When the operand x is 10, 4 and x 3 If the operand x is 00, the output is 00. 4 and x 3 If the operand x is 01, the output is 10. 4 and x 3 If the operand x is 10, the output is 10. 4 and x 3 If the operand x is 11, the output is 10; 2 and x 1 When the operand x is 11, 4 and x 3 If the operand x is 00, the output is 01. 4 and x 3 If the operand x is 01, the output is 11. 4 and x 3 If the operand x is 10, the output is 11. 4 and x 3 is 11, the output is 11; so that the approximate 4-2 compressor AC 1 Only on operand x 2 、x 1 、x 4 and x 3 In the case of 0001, 0010, 0101, and 0110, 4 "+1" errors are generated. 2 、x 1 、x 4 and x 3 In the four cases of 1000, 1011, 1100, and 1111, 4 "-1" errors are generated. The 4 "+1" errors and the 4 "-1" errors compensate each other in probability to reduce the total error.

[0008] Optionally, the carry-ignoring approximate 4-2 compressor comprises an approximate 4-2 compressor AC 2 , the approximate 4-2 compressor AC 2 The approximate 4-2 compressor AC includes an AND gate and two OR gates. 1 The operand x is input when performing the operation 1 and operand x 3 The two pass through the first OR gate to get the output sum, the input operand x 3 and operand x 4 Input AND gate, output of AND gate and operand x 2 The two pass through the second OR gate to get the output Carry, the approximate 4-2 compressor AC 2 The truth table is as follows: ① When the operand x 2and x 1 When the operand x is 00, 4 and x 3 If the operand x is 00, the output is 00. 4 and x 3 If the operand x is 01, the output is 01. 4 and x 3 If the operand x is 10, the output is 00. 4 and x 3 If the operand x is 11, the output is 11; ②When the operand x 2 and x 1 When it is 01, if the operand x 4 and x 3 If the operand x is 00, the output is 01. 4 and x 3 If the operand x is 01, the output is 01. 4 and x 3 If the operand x is 10, the output is 01. 4 and x 3 If the operand x is 11, the output is 11; 2 and x 1 When the operand x is 10, 4 and x 3 If the operand x is 00, the output is 10. 4 and x 3 If the operand x is 01, the output is 11. 4 and x 3 If the operand x is 10, the output is 10. 4 and x 3 If the operand x is 11, the output is 11; 2 and x 1 When the operand x is 11, 4 and x 3 If the operand x is 00, the output is 11. 4 and x 3 If the operand x is 01, the output is 11. 4 and x 3 If the operand x is 10, the output is 11. 4 and x 3 is 11, the output is 11; so that the approximate 4-2 compressor AC 2 Only on operand x 2 、x 1 、x 4 and x 3 In the case of 1000, 1001, 1100, and 0011, four "+1" errors are generated. 2 、x 1 、x 4 and x3 Four "-1" errors are generated in the four cases of 0010, 0101, 0110, and 1111. The four "+1" errors and the four "-1" errors compensate each other in probability to reduce the total error.

[0009] Optionally, the compressor unit further comprises a compressor which is only operand x 2 、x 1 、x 4 and x 3 A high-precision 4-2 compressor that will only generate a "-2" error when the value is "1111", the high-precision 4-2 compressor includes three AND gates, two OR gates and two XOR gates, and the operand x 4 and x 3 Input the first AND gate together, operand x 2 and x 1 Input to the second AND gate together, operand x 4 and x 3 Also input into the first XOR gate, operand x 2 and x 1 The outputs of the two XOR gates are also input into the second XOR gate, the outputs of the two XOR gates are input into the third AND gate, the outputs of the three AND gates are input into the first OR gate to obtain the output Carry, and the outputs of the two XOR gates are also input into the second OR gate to obtain the output sum.

[0010] Optionally, when the approximate mantissa multiplier first generates a partial product array of 11 rows and 21 columns by performing a bit-by-bit AND operation on the two operands input into the approximate floating-point multiplier through an AND logic unit, the highest bit of each row of the partial product array and all bits of the last row are directly generated as 1.

[0011] Optionally, the compressor unit comprises an approximately 4-2 compressor AC 1 , similar to 4-2 compressor AC 2 , a high-precision 4-2 compressor HPC, a full adder, a half adder and an accurate 4-2 compressor are used to compress the 11th to 21st columns by column to obtain the final approximate mantissa.

[0012] Optionally, the method of compressing the 11th to 21st columns by column to obtain the final approximate mantissa comprises four stages: in the first stage, in the partial product array, the 1st, 2nd, 11th rows in the column with index 10 and the compensated "1" are combined using an approximate 4-2 compressor AC 1 The remaining 8 lines use two approximate 4-2 compressors AC 2 Compression is completed; columns with indexes 11, 12, and 13 use two approximate 4-2 compressors AC respectively 2 Compression is done; the column with index 14 uses an approximate 4-2 compressor AC2 Complete compression; columns with indexes 15 to 19 are compressed using an accurate 4-2 compressor respectively; in the second stage, in the result obtained in the first stage, the column with index 10 is compressed using a full adder, the column with index 11 is compressed using two high-precision 4-2 compressors HPC, the columns with indexes 12, 13, and 14 are compressed using a high-precision 4-2 compressor HPC respectively, and the columns with indexes 15, 16, and 19 are compressed using a full adder respectively; in the third stage, in the result obtained in the second stage, the columns with indexes 12 and 14 are compressed using a high-precision 4-2 compressor HPC respectively, the columns with indexes 13, 15, and 16 are compressed using a half adder respectively, and the columns with index 17 are compressed using a full adder respectively; in the fourth stage, the results obtained in the third stage are added through a carry propagation adder to obtain the final approximate mantissa.

[0013] Optionally, the floating-point multiplier body includes a sign bit XOR module, an exponent addition module, an approximate mantissa multiplier, a normalization module and a rounding module, the sign bit XOR module is used to perform an XOR operation on the sign bits of the multiplier and the multiplicand to serve as the sign bit of the output result, the exponent addition module is used to add the exponents of the multiplier and the multiplicand, the approximate mantissa multiplier is used to multiply the mantissas of the multiplier and the multiplicand, the normalization module is used to normalize the result of the mantissa multiplication and send it to the rounding module, the rounding module is used to round the normalized result as the mantissa of the output result, and control the exponent adjustment module to adjust the exponent obtained by the exponent addition to serve as the mantissa exponent of the output result.

[0014] In addition, this embodiment further provides a chip, including a chip body and a computing unit captured in the chip body, wherein the computing unit includes the approximate floating-point multiplier.

[0015] In addition, this embodiment also provides a computing device, including a microprocessor and a memory connected to each other, wherein the microprocessor includes the approximate floating-point multiplier.

[0016] Compared with the prior art, the present invention mainly has the following advantages: the demand for half-precision floating-point operations in applications such as machine learning is increasing, but there is currently little research on half-precision approximate floating-point multipliers, and there is a huge room for optimization in terms of precision and overhead such as power consumption and area. The approximate mantissa multiplier of the approximate floating-point multiplier of the present invention includes an AND logic unit and a compressor unit, the AND logic unit is used to perform an AND operation on two input operands bit by bit to generate a partial product array with a size of 11 rows and 21 columns, the compressor unit is used to compress the 11th to 21st columns by column to obtain the final approximate mantissa, and the compressor unit includes two extremely simple approximate 4-2 compressors with ignoring carry designs, and the approximate 4-2 compressor makes the error rate within an acceptable range by utilizing the mutual compensation inside the compressor, so that the energy efficiency of the floating-point multiplication can be further improved by exploiting the characteristics of the floating-point multiplication itself, and the optimization of the approximate floating-point multiplier in terms of precision and overhead such as power consumption and area is achieved, and the problem of relatively complex circuits and low compression efficiency of traditional approximate 4-2 compressors is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the structure of an approximate floating-point multiplier in an embodiment of the present invention.

[0018] Figure 2 The AC compressor is approximately 4-2 according to the embodiment of the present invention. 1 Logic circuit diagram.

[0019] Figure 3 The AC compressor is approximately 4-2 according to the embodiment of the present invention. 1 Truth table diagram of .

[0020] Figure 4 The AC compressor is approximately 4-2 according to the embodiment of the present invention. 2 Logic circuit diagram.

[0021] Figure 5 The AC compressor is approximately 4-2 according to the embodiment of the present invention. 2 Truth table diagram of .

[0022] Figure 6 4-2 compressor in an embodiment of the present invention.

[0023] Figure 7 Schematic diagram of a partial product array directly obtained in an embodiment of the present invention.

[0024] Figure 8 Schematic diagram of the working principle of the compressor unit in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The present invention aims to arrange the targeted new efficient approximate 4-2 compressor in the partial product matrix in the best way according to the characteristics of floating-point mantissa multiplication, and use other high-precision compressors to form a good approximate mixed compression strategy, and finally combine truncation processing and tuning compensation to achieve the design goals of high-precision and low-overhead multipliers. The design includes two aspects, one is the design of two efficient approximate 4-2 compressors with mutual error compensation characteristics, and the other is the design of a floating-point multiplier that comprehensively uses multiple approximate technologies. In order to enable personnel in this technical field to better understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described in combination with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.

[0026] like Figure 1 As shown, this embodiment provides an approximate floating-point multiplier, including a floating-point multiplier body with an approximate mantissa multiplier, the approximate mantissa multiplier including an AND logic unit and a compressor unit, the AND logic unit is used to perform an AND operation on two operands input to the approximate floating-point multiplier bit by bit to generate a partial product array with a size of 11 rows and 21 columns, each row in the partial product array is 11 valid partial products, the compressor unit is used to compress the partial product array of columns 11 to 21 obtained after truncating the columns 1 to 10 by column to obtain a final approximate mantissa, the compressor unit includes an approximate 4-2 compressor with a carry-ignoring design, and the approximate 4-2 compressor with a carry-ignoring design makes the error rate within an acceptable range by utilizing mutual compensation inside the compressor.

[0027] The design concept of the approximate 4-2 compressor ignoring the carry design in this embodiment is to use the mutual compensation inside the compressor to make the error rate within an acceptable range, while realizing the extremely simple design of the compressor structure as much as possible, greatly reducing power consumption, delay and area, etc. Specifically, the approximate 4-2 compressor ignoring the carry design in this embodiment includes the approximate 4-2 compressor AC 1 ,like Figure 2 As shown, the approximate 4-2 compressor AC 1 includes an OR gate, and the approximate 4-2 compressor AC 1 Ignore the input operand x when performing the operation 2 , the input operand x 1 As the output sum, the input operand x 3 and x 4 The carry bit is output after passing through the OR gate, such as Figure 3 As shown, the approximate 4-2 compressor AC1 The truth table is as follows: ① When the operand x 2 and x 1 When the operand x is 00, 4 and x 3 If the operand x is 00, the output is 00. 4 and x 3 If the operand x is 01, the output is 10. 4 and x 3 If the operand x is 10, the output is 10. 4 and x 3 If the operand x is 11, the output is 10; ②When the operand x 2 and x 1 When it is 01, if the operand x 4 and x 3 If the operand x is 00, the output is 01. 4 and x 3 If the operand x is 01, the output is 11. 4 and x 3 If the operand x is 10, the output is 11. 4 and x 3 If the operand x is 11, the output is 11; 2 and x 1 When the operand x is 10, 4 and x 3 If the operand x is 00, the output is 00. 4 and x 3 If the operand x is 01, the output is 10. 4 and x 3 If the operand x is 10, the output is 10. 4 and x 3 If the operand x is 11, the output is 10; 2 and x 1 When the operand x is 11, 4 and x 3 If the operand x is 00, the output is 01. 4 and x 3 If the operand x is 01, the output is 11. 4 and x 3 If the operand x is 10, the output is 11. 4 and x 3 is 11, the output is 11; so that the approximate 4-2 compressor AC 1 Only on operand x 2 、x 1 、x 4 and x 3 In the case of 0001, 0010, 0101, and 0110, 4 "+1" errors are generated.2 、x 1 、x 4 and x 3 In the four cases of 1000, 1011, 1100, and 1111, 4 "-1" errors are generated. The 4 "+1" errors and the 4 "-1" errors compensate each other in probability to reduce the total error. Figure 3 The solid line box in the figure indicates that a "+1" error is generated, and the dotted line box indicates that a "-1" error is generated. Approximate 4-2 compressor AC 1 At the same time, 4 "+1" errors and 4 "-1" errors are generated, which can compensate each other in probability, thereby reducing the total error. The average error rate is 12.5%. In the designed floating-point multiplier, the average error rate will be further reduced. Figure 2 It can be seen that the approximate 4-2 compressor AC 1 The structure is extremely simple, requiring only one OR gate. Compared with the complex gate-level circuit structure of the precise compressor, it is similar to the 4-2 compressor AC 1 The cost is greatly reduced.

[0028] Although the approximate 4-2 compressor AC 1 The structure is simple and the compression efficiency is high, but the high error rate is still not negligible. Therefore, in the approximate 4-2 compressor AC 1 Based on the design, this embodiment further designs an approximate 4-2 compressor AC with higher accuracy. 2 The approximate 4-2 compressor designed by ignoring carry in this embodiment includes an approximate 4-2 compressor AC 2 ,like Figure 4 As shown, the approximate 4-2 compressor AC 2 It consists of an AND gate and two OR gates, and is similar to a 4-2 compressor AC 1 The operand x is input when performing the operation 1 and operand x 3 The two pass through the first OR gate to get the output sum, the input operand x 3 and operand x 4 Input AND gate, output of AND gate and operand x 2 The two pass through the second OR gate to get the output carry, such as Figure 5 As shown, the approximate 4-2 compressor AC 2 The truth table is as follows: ① When the operand x 2 and x 1 When the operand x is 00, 4 and x 3 If the operand x is 00, the output is 00. 4 and x 3 If the operand x is 01, the output is 01. 4 and x3 If the operand x is 10, the output is 00. 4 and x 3 If the operand x is 11, the output is 11; ②When the operand x 2 and x 1 When it is 01, if the operand x 4 and x 3 If the operand x is 00, the output is 01. 4 and x 3 If the operand x is 01, the output is 01. 4 and x 3 If the operand x is 10, the output is 01. 4 and x 3 If the operand x is 11, the output is 11; 2 and x 1 When the operand x is 10, 4 and x 3 If the operand x is 00, the output is 10. 4 and x 3 If the operand x is 01, the output is 11. 4 and x 3 If the operand x is 10, the output is 10. 4 and x 3 If the operand x is 11, the output is 11; 2 and x 1 When the operand x is 11, 4 and x 3 If the operand x is 00, the output is 11. 4 and x 3 If the operand x is 01, the output is 11. 4 and x 3 If the operand x is 10, the output is 11. 4 and x 3 is 11, the output is 11; so that the approximate 4-2 compressor AC 2 Only on operand x 2 、x 1 、x 4 and x 3 In the case of 1000, 1001, 1100, and 0011, 4 "+1" errors are generated. 2 、x 1 、x 4 and x 3 Four "-1" errors are generated in the four cases of 0010, 0101, 0110, and 1111. The four "+1" errors and the four "-1" errors compensate each other in probability to reduce the total error. Figure 5The solid line box indicates that a "+1" error is generated, and the dotted line box indicates that a "-1" error is generated. Approximate 4-2 compressor AC 2 With an approximate 4-2 compressor AC 1 Similarly, 4 "+1" errors and 4 "-1" errors are also generated at the same time, thus playing a role of mutual compensation. Compared with the approximate 4-2 compressor AC 1 , similar to a 4-2 compressor AC 2 Only one AND gate and one OR gate are added, but the probability of mutual compensation of errors is more balanced, and the average error rate is only 3.125%.

[0029] In addition to the two proposed high-efficiency approximate 4-2 compressors, other high-precision approximate compressors are required at different stages of the partial accumulation process to ensure that the accuracy loss is within an acceptable range. Therefore, a high-precision 4-2 compressor that only produces a "-2" error when the input is "1111" is a good choice for balancing accuracy and overhead. Therefore, the compressor unit of this embodiment also includes a high-precision 4-2 compressor that only produces a "-2" error when the input operand x is "1111". 2 、x 1 、x 4 and x 3 A high-precision 4-2 compressor that only produces a "-2" error when the bit is "1111" is used. Figure 6 As shown, the high-precision 4-2 compressor includes three AND gates, two OR gates and two XOR gates. 4 and x 3 The first AND gate (A in the figure) is fed together with the operand x 2 and x 1 The second AND gate (B in the figure) is fed with the operand x 4 and x 3 It is also input into the first XOR gate (represented as C in the figure), and the operand x 2 and x 1 The outputs of the two XOR gates are also input into the second XOR gate (indicated as D in the figure), the outputs of the two XOR gates are input into the third AND gate (indicated as E in the figure), the outputs of the three AND gates are input into the first OR gate (indicated as F in the figure) to obtain the output Carry, and the outputs of the two XOR gates are also input into the second OR gate (indicated as G in the figure) to obtain the output sum. The logical expression of the high-precision 4-2 compressor is: , , in, For peace, is XOR, For carry.

[0030] The design of the approximate multiplier mainly introduces approximate calculation in the mantissa multiplication stage. The approximate methods used include: truncation and tuning compensation for the low-significant bits of the partial product array, approximate mixed compression for the middle part, and precise calculation for the high-significant bits. Since there is always an implicit "1" in the highest bit of the mantissa, this feature is used. In this embodiment, the approximate mantissa multiplier first generates a partial product array of 11 rows and 21 columns by performing a bit-by-bit "AND" operation on the two operands input to the approximate floating-point multiplier through the logic unit. The highest bit of each row of the partial product array and all bits of the last row are directly generated as 1, which can be directly obtained as follows Figure 7 The black partial product shown in FIG. 4 reduces 21 AND gates in the partial product generation process, thereby achieving the purpose of reducing resource overhead.

[0031] like Figure 8 As shown, the compressor unit in this embodiment includes approximately 4-2 compressor AC 1 , similar to 4-2 compressor AC 2 , a high-precision 4-2 compressor HPC, a full adder, a half adder and an accurate 4-2 compressor are used to compress the 11th to 21st columns by column to obtain the final approximate mantissa. Figure 8 As shown, in this embodiment, the 11th to 21st columns are compressed column by column to obtain the final approximate mantissa, which includes four stages: In the first stage, the 1st, 2nd, and 11th rows in the column with index 10 in the partial product array are combined with the compensated "1" and then used in an approximate 4-2 compressor AC 1 The remaining 8 lines use two approximate 4-2 compressors AC 2 Compression is completed; columns with indexes 11, 12, and 13 use two approximate 4-2 compressors AC respectively 2 Compression is done; the column with index 14 uses an approximate 4-2 compressor AC 2 The compression is completed; the columns with indexes 15 to 19 are compressed using an exact 4-2 compressor respectively; In the second stage, in the results obtained in the first stage, the column with index 10 is compressed using a full adder, the column with index 11 is compressed using two high-precision 4-2 compressors HPC, the columns with indexes 12, 13, and 14 are compressed using a high-precision 4-2 compressor HPC respectively, and the columns with indexes 15, 16, and 19 are compressed using a full adder respectively; In the third stage, in the results obtained in the second stage, the columns with indexes 12 and 14 are compressed using a high-precision 4-2 compressor HPC, the columns with indexes 13, 15 and 16 are compressed using a half adder, and the column with index 17 is compressed using a full adder. In the fourth stage, the results obtained in the third stage are added through a carry-propagation adder to form the final approximate mantissa.

[0032] See also Figure 8 In this embodiment, truncation is used in columns 0-9 of the partial product matrix, approximate mixed compression is used in columns 10-14, and precise calculation is used in columns 15-21. In half-precision floating-point multiplication, mantissa multiplication generates 22-bit partial products during calculation, but only 10 significant digits are retained after normalization. Therefore, according to the characteristics of mantissa normalization, an approximate processing method is adopted to truncate the lower 10 columns of the partial product matrix and add "1" to the 10th column for tuning compensation. In the approximate compression part of columns 10-14, the proposed AC 1 and AC 2 Two approximate 4-2 compressors, combined with the direct generation of the partial product matrix, adopt a compression strategy that is beneficial to the balance between accuracy and overhead. In the first stage of partial accumulation, considering AC 1 The error is large and the compression efficiency is high. It is applied to both ends of the partial product of the 10th column to reduce the error rate and improve the compression efficiency. The partial products of the 10-14 columns use AC with higher accuracy. 2 Compression involves directly generating both ends of the columns of partial products ( Figure 8 The blue part in the middle) is the AC 2 The input is compressed to further reduce the error rate. In the second and third stages of partial accumulation and addition, a high-precision 4-2 compressor HPC is used, and accurate full adders and half adders are used in the rest to maintain high accuracy and gain overhead benefits. After repeatedly using the high-precision 4-2 compressor HPC, full adders, and half adders, two rows of partial products are obtained in the fourth stage, and the final two rows of partial products are added to produce the final result.

[0033] In addition, the floating-point multiplier body of the present embodiment includes a sign bit XOR module, an exponent addition module, an approximate mantissa multiplier, a normalization module and a rounding module, wherein the sign bit XOR module is used to perform an XOR operation on the sign bits of the multiplier and the multiplicand to serve as the sign bit of the output result, the exponent addition module is used to add the exponents of the multiplier and the multiplicand, the approximate mantissa multiplier is used to multiply the mantissas of the multiplier and the multiplicand, the normalization module is used to normalize the result of the mantissa multiplication and send it to the rounding module, the rounding module is used to round the normalized result as the mantissa of the output result, and control the exponent adjustment module to adjust the exponent obtained by the exponent addition to serve as the mantissa exponent of the output result.

[0034] The quasi-floating-point multiplier of this embodiment has the following characteristics: (1) The two approximate 4-2 compressors proposed in the approximate floating-point multiplier of this embodiment use the idea of ​​mutual error compensation, and have extremely high compression efficiency under the condition of acceptable precision loss. (2) The precision loss of the approximate floating-point multiplier of this embodiment is small, and the approximate design method based on the characteristics of the new compressor and mantissa multiplication can be extended to other floating-point type formats. (3) The overall power consumption and area overhead of the approximate floating-point multiplier of this embodiment after synthesis is small. Compared with the precise floating-point multiplier, the power consumption and area are reduced by about half, and the power consumption delay product is reduced by 60%. The low-overhead design makes the approximate floating-point multiplier of this embodiment suitable for the design of various general-purpose processors and hardware accelerators.

[0035] In addition, this embodiment also provides a chip, including a chip body and a computing unit captured in the chip body, wherein the computing unit includes the approximate floating-point multiplier. Since the computing unit is a basic circuit structure in the chip, the structure of the chip is not described in detail here.

[0036] In addition, this embodiment also provides a computing device, including a microprocessor and a memory connected to each other, wherein the microprocessor includes the approximate floating-point multiplier.

[0037] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. An approximate floating-point multiplier, comprising a floating-point multiplier body with an approximate mantissa multiplier, characterized in that: The approximate mantissa multiplier includes an AND logic unit and a compressor unit. The AND logic unit is used to perform an AND operation on two operands input to the approximate floating-point multiplier bit by bit to generate a partial product array with a size of 11 rows and 21 columns, each row of the partial product array is 11 valid partial products, and the compressor unit is used to compress the partial product array of columns 11 to 21 obtained by truncating columns 1 to 10 by column to obtain the final approximate mantissa. The compressor unit includes an approximate 4-2 compressor with a carry-ignoring design. The approximate 4-2 compressor with a carry-ignoring design makes the error rate within an acceptable range by utilizing mutual compensation inside the compressor.

2. The approximate floating-point multiplier according to claim 1, wherein: The approximate 4-2 compressor designed with ignore carry includes an approximate 4-2 compressor AC1, and the approximate 4-2 compressor AC1 includes an OR gate, and the approximate 4-2 compressor AC1 ignores the input operand x2 when performing operation, takes the input operand x1 as the output sum, and takes the input operands x3 and x4 as the output carry after passing through the OR gate. The truth table of the approximate 4-2 compressor AC1 is as follows: ① When the operands x2 and x1 are 00, If operands x4 and x3 are 00, the output is 00; if operands x4 and x3 are 01, the output is 10; if operands x4 and x3 are 10, the output is 10; if operands x4 and x3 are 11, the output is 10; ② When operands x2 and x1 are 01, if operands x4 and x3 are 00, the output is 01; if operands x4 and x3 are 01, the output is 11; if operands x4 and x3 are 10, the output is 11; if operands x4 and x3 are 11, the output is 11; ③When operands x2 and x1 are 10, if operands x4 and x3 are 00, the output is 00; if operands x4 and x3 are 01, the output is 10; if operands x4 and x3 are 10, the output is 10; if operands x4 and x3 are 11, the output is 10; ④When operands x2 and x1 are 11, if operands x4 and x3 are 00, the output is 01; if operands x4 and x3 are 01, the output is 11; if operands x4 and x3 are 10, the output is 11; If x4 and x3 are 11, the output is 11; so that the approximate 4-2 compressor AC1 only generates 4 "+1" errors when the operands x2, x1, x4 and x3 are 0001, 0010, 0101, and 0110, and generates 4 "-1" errors when the operands x2, x1, x4 and x3 are 1000, 1011, 1100, and 1111. The 4 "+1" errors and the 4 "-1" errors compensate each other in probability to reduce the total error.

3. The approximate floating-point multiplier according to claim 1, wherein: The approximate 4-2 compressor designed to ignore carry includes an approximate 4-2 compressor AC2, which includes an AND gate and two OR gates. When the approximate 4-2 compressor AC1 performs an operation, the input operand x1 and the operand x3 are both passed through the first OR gate to obtain the output sum, the input operand x3 and the operand x4 are input to the AND gate, and the output of the AND gate and the operand x2 are both passed through the second OR gate to obtain the output carry Carry. The approximate 4-2 compressor AC2 The truth table is as follows: ① When operands x2 and x1 are 00, if operands x4 and x3 are 00, the output is 00, if operands x4 and x3 are 01, the output is 01, if operands x4 and x3 are 10, the output is 00, and if operands x4 and x3 are 11, the output is 11; ② When operands x2 and x1 are 01, if operands x4 and x3 are 00, the output is 01, if operands x4 and x3 are 01, the output is 01, if operands x4 and x3 are 10, the output is 01, and if operands If x4 and x3 are 11, the output is 11; ③ When operands x2 and x1 are 10, if operands x4 and x3 are 00, the output is 10, if operands x4 and x3 are 01, the output is 11, if operands x4 and x3 are 10, the output is 10, if operands x4 and x3 are 11, the output is 11; ④ When operands x2 and x1 are 11, if operands x4 and x3 are 00, the output is 11, if operands x4 and x3 are 01, the output is 11, if operands x4 and x3 are 10, the output is is 11, and if the operands x4 and x3 are 11, the output is 11; so that the approximate 4-2 compressor AC2 only generates 4 "+1" errors when the operands x2, x1, x4 and x3 are 1000, 1001, 1100, and 0011, and generates 4 "-1" errors when the operands x2, x1, x4 and x3 are 0010, 0101, 0110, and 1111, and the 4 "+1" errors and the 4 "-1" errors compensate each other in probability to reduce the total error.

4. The approximate floating-point multiplier according to claim 1, wherein: The compressor unit also includes a high-precision 4-2 compressor that will only generate a "-2" error when the input operands x2, x1, x4 and x3 are "1111". The high-precision 4-2 compressor includes three AND gates, two OR gates and two XOR gates. The operands x4 and x3 are input into the first AND gate together, the operands x2 and x1 are input into the second AND gate together, the operands x4 and x3 are also input into the first XOR gate together, the operands x2 and x1 are also input into the second XOR gate together, the outputs of the two XOR gates are input into the third AND gate together, the outputs of the three AND gates are input into the first OR gate together to obtain the output Carry, and the outputs of the two XOR gates are also input into the second OR gate together to obtain the output sum.

5. The approximate floating-point multiplier according to claim 1, wherein: The approximate mantissa multiplier first performs an AND operation on the two operands input to the approximate floating-point multiplier bit by bit through an AND logic unit to generate a partial product array with a size of 11 rows and 21 columns. The highest bit of each row of the partial product array and all bits of the last row are directly generated as 1.

6. The approximate floating-point multiplier according to claim 1, wherein: The compressor unit includes an approximate 4-2 compressor AC1, an approximate 4-2 compressor AC2, a high-precision 4-2 compressor HPC, a full adder, a half adder and an accurate 4-2 compressor for compressing the 11th to 21st columns by column to obtain the final approximate mantissa.

7. The approximate floating-point multiplier according to claim 6, characterized in that: The method of compressing the 11th to 21st columns by column to obtain the final approximate mantissa includes four stages: in the first stage, in the partial product array, the 1st, 2nd, and 11th rows in the column with index 10 are combined with the compensated "1" and compressed using an approximate 4-2 compressor AC1, and the remaining 8 rows are compressed using two approximate 4-2 compressors AC2; the columns with index 11, 12, and 13 are compressed using two approximate 4-2 compressors AC2 respectively; the column with index 14 is compressed using an approximate 4-2 compressor AC2; the columns with indexes 15 to 19 are compressed using an exact 4-2 compressor respectively; in the second stage, in the result obtained in the first stage, the column with index 10 is compressed using Compression is completed using a full adder, the column with index 11 is compressed using two high-precision 4-2 compressors HPC, the columns with indexes 12, 13, and 14 are compressed using a high-precision 4-2 compressor HPC respectively, and the columns with indexes 15, 16, and 19 are compressed using a full adder respectively; in the third stage, in the results obtained in the second stage, the columns with indexes 12 and 14 are compressed using a high-precision 4-2 compressor HPC respectively, the columns with indexes 13, 15, and 16 are compressed using a half adder respectively, and the columns with index 17 are compressed using a full adder respectively; in the fourth stage, the results obtained in the third stage are added together to obtain the final approximate mantissa.

8. The approximate floating-point multiplier according to claim 1, wherein: The floating-point multiplier body comprises a sign bit XOR module, an exponent addition module, an approximate mantissa multiplier, a normalization module and a rounding module. The sign bit XOR module is used to perform an XOR operation on the sign bits of the multiplier and the multiplicand to serve as the sign bit of the output result. The exponent addition module is used to add the exponents of the multiplier and the multiplicand. The approximate mantissa multiplier is used to multiply the mantissas of the multiplier and the multiplicand. The normalization module is used to perform normalization processing on the result of mantissa multiplication and send it to the rounding module. The rounding module is used to round the normalized result as the mantissa of the output result and control the exponent adjustment module to adjust the exponent obtained by exponent addition to serve as the mantissa exponent of the output result.

9. A chip, comprising a chip body and a computing unit captured in the chip body, characterized in that: The operation unit comprises the approximate floating-point multiplier as claimed in any one of claims 1 to 8.

10. A computing device comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor comprises the approximate floating-point multiplier according to any one of claims 1 to 8.

Citation Information

Cited By

  • Approximate 4-2 compressor and approximation multiplier based on pairwise error compensation

    CN121680780A

  • An approximate 4-2 compressor and an approximate multiplier based on pair-wise error compensation

    CN121680780B