Floating-point number addition and subtraction device and method
The floating-point arithmetic unit optimizes operations by minimizing code conversions and integrating normalization logic, improving speed and efficiency while maintaining precision, addressing the limitations of existing designs.
Patent Information
- Application Number
- CN202510385272.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
Traditional floating-point number addition and subtraction devices have high logic complexity and limited computing speed, making it difficult to meet the needs of modern high-performance computing.
The mantissa alignment module is used for order alignment shifting, and multiple conversions between the original code and the complement are avoided by taking the inverse control signal. It combines the mantissa operation module for efficient operation, uses the transmission gate array submodule for efficient normalization shifting, and integrates the leading zero preprocessing function to simplify the logical path.
It significantly reduces circuit complexity and hardware resource waste, improves computing speed and efficiency, meets the needs of high precision and low power consumption, and is suitable for high-performance computing scenarios.
Smart Images

Figure CN120315673A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of integrated circuits, and particularly to a floating-point addition and subtraction device and method. Background Art
[0002] Floating-point numbers are a numerical representation widely used in the field of computing, consisting of a sign bit, an exponent, and a mantissa. Compared with fixed-point numbers, floating-point numbers have a larger value range and higher calculation accuracy, so they are widely adopted in fields such as artificial intelligence, big data processing, embedded systems, and high-performance computing. The floating-point addition and subtraction operations based on the IEEE-754 standard require a complex processing process, including steps such as exponent alignment, mantissa addition and subtraction, and result normalization, to ensure the accuracy of the calculation result and the standardization of the format.
[0003] In related technologies, floating-point adders and subtractors are usually implemented using an architecture based on the IEEE-754 standard. However, there are problems of high logical complexity and limited calculation speed in traditional designs, which makes it difficult to further improve the overall performance of floating-point adders and subtractors. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a floating-point addition and subtraction device and method, which can solve the problems of high logical complexity and limited calculation speed of the floating-point addition and subtraction architecture.
[0005] To solve the above technical problems, an embodiment of the present invention provides a floating-point addition and subtraction device for implementing the addition and subtraction operations of a first floating-point operand and a second floating-point operand. The first floating-point operand includes a first exponent and a first mantissa, and the second floating-point operand includes a second exponent and a second mantissa. The floating-point addition and subtraction device includes: a mantissa alignment module for performing an exponent alignment shift operation on the smaller mantissa and performing an inversion operation on the smaller mantissa after the exponent alignment shift operation based on an inversion control signal, and outputting a mantissa alignment result; wherein, the smaller mantissa is the mantissa of the operand with the smaller absolute value among the first floating-point operand and the second floating-point operand; a mantissa operation module connected to the mantissa alignment module for performing an operation on the larger mantissa and the mantissa alignment result, and outputting a mantissa operation result; wherein, the larger mantissa is the mantissa of the operand with the larger absolute value among the first floating-point operand and the second floating-point operand; a mantissa processing module connected to the mantissa operation module for performing normalization processing on the mantissa operation result and outputting a normalized mantissa; the mantissa processing module includes: a leading zero preprocessing sub-module for generating a prefix or an operation result based on the mantissa operation result; a transmission gate control sub-module for generating a transmission gate control signal based on the prefix or the operation result; a transmission gate array sub-module for normalizing the mantissa operation result based on the transmission gate control signal.
[0006] In addition, a floating-point addition and subtraction device provided by an embodiment of the present invention further includes: the mantissa alignment module is further configured to output a remainder; the mantissa processing module is further configured to receive the remainder and supplement the remainder into the normalized mantissa. By retaining the remainder after aligning the smaller mantissa and performing a restoration process on it during the normalization stage, the accuracy of the floating-point addition and subtraction operation device is significantly improved, ensuring the accuracy of the result, and is particularly applicable to application scenarios with strict requirements for high-precision calculations.
[0007] An embodiment of the present invention also provides a floating-point addition and subtraction method, which is applied to the floating-point addition and subtraction device described above, and includes: receiving a first floating-point operand and a second floating-point operand; wherein, the first floating-point operand includes a first exponent and a first mantissa, and the second floating-point operand includes a second exponent and a second mantissa; generating an inversion flag and a final sign bit based on the sign bits and the operator of the first floating-point operand and the second floating-point operand; performing a stage alignment shift operation on the smaller mantissa, and determining whether to perform an inversion operation on the aligned mantissa according to the inversion flag; performing an addition operation on the aligned mantissa and the larger mantissa to obtain a mantissa operation value; generating a prefix or an operation result based on the mantissa operation value, and generating a transmission gate control signal and a leading zero count value; performing a normalization shift based on the transmission gate control signal and outputting a normalized mantissa; adjusting the exponent of the operand with the larger exponent based on the leading zero count value and outputting an exponent adjustment value; performing a normalization process on the normalized mantissa, the final sign bit signal, and the exponent adjustment value, and outputting a floating-point number that meets a preset standard.
[0008] The floating-point addition and subtraction device provided by the embodiment of the present invention first, through the control mechanism of the negation control signal, only performs one's complement operation on the smaller mantissa once or does not need to perform it, avoiding multiple conversions between the original code and the complement code, significantly reducing the circuit complexity and waste of hardware resources, thereby improving the operation speed and efficiency. Subsequently, the mantissa operation module efficiently operates on the aligned smaller mantissa and larger mantissa, and ensures that the final mantissa output is always positive, simplifying the leading zero / one statistical logic to a leading zero statistical operation, further reducing the circuit scale and complexity. In the mantissa processing module, through the sharing mechanism with the leading zero preprocessing function, the prefix or result generated by the leading zero preprocessing module can be directly used as the input of the transmission gate control sub-module, thereby generating a dynamic transmission gate control signal. This design integrates some functions of the leading zero counting module into the normalization circuit, not only effectively shortening the logic path but also reducing the delay caused by too deep cascaded logic. Finally, the transmission gate array sub-module adopts an efficient mantissa normalization shift operation, significantly improving the mantissa shift efficiency. The overall solution is fully compatible with the IEEE-754 standard, and can significantly improve the addition and subtraction operation efficiency and energy efficiency performance while ensuring high precision, meeting the requirements of modern high-performance computing for high efficiency, low power consumption, and high precision. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a proportional limitation.
[0010] Figure 1 is the floating-point format defined in the IEEE-754 specification;
[0011] Figure 2 is the structural diagram of the floating-point addition and subtraction device in the related art;
[0012] Figure 3 is the structural diagram of implementing shift by a common single multiplexer in the floating-point addition and subtraction device in the related art;
[0013] Figure 4 is the structural diagram of implementing shift by a common multi-stage two-way selector in the floating-point addition and subtraction device in the related art;
[0014] Figure 5 is the structural diagram of the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0015] Figure 6 is the structural diagram of the comparison and decision module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0016] Figure 7It is the operation schematic diagram of the mantissa restorer in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0017] Figure 8 It is the structural diagram of the mantissa alignment module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0018] Figure 9a It is the operation schematic diagram of the mantissa alignment module in the floating-point addition and subtraction device provided by the embodiment of the present invention when eSHIFT = 0;
[0019] Figure 9b It is the operation schematic diagram of the mantissa alignment module in the floating-point addition and subtraction device provided by the embodiment of the present invention when eSHIFT > 0;
[0020] Figure 10 It is the structural diagram of the mantissa operation module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0021] Figure 11 It is the structural diagram of the mantissa processing module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0022] Figure 12a It is the serial cascade electrical structure of the leading zero preprocessing sub-module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0023] Figure 12b It is the parallel direct connection structural diagram of the leading zero preprocessing sub-module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0024] Figure 13 It is the structural diagram of the zero-padding signal generation sub-module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0025] Figure 14 It is the structural diagram of the strobe signal generation sub-module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0026] Figure 15 It is the structural diagram of the transmission gate array sub-module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0027] Figure 16 It is the structural diagram of the leading zero counting sub-module in the floating-point addition and subtraction device provided by the embodiment of the present invention;
[0028] Figure 17a It is the operation schematic diagram of the mantissa processing module in the floating-point addition and subtraction device provided by the embodiment of the present invention when SUM[n] = 1;
[0029] Figure 17bIt is a schematic diagram of the operation of the mantissa processing module in the floating-point addition and subtraction device provided by the embodiment of the present invention when SUM[n,n - 1]=01;
[0030] Figure 17c It is a schematic diagram of the operation of the mantissa processing module in the floating-point addition and subtraction device provided by the embodiment of the present invention when SUM[n,n - 4]=0001;
[0031] Figure 18 It is a flowchart of the floating-point addition and subtraction method provided by the embodiment of the present invention. Detailed implementation manners
[0032] Floating-point numbers are a numerical representation widely used in the field of computing. As shown in Figure 1 , they consist of a sign bit (S), an exponent (E), and a mantissa (M). Compared with fixed-point numbers, floating-point numbers have a larger value range and higher calculation accuracy, so they are widely used in fields such as artificial intelligence, big data processing, embedded systems, and high-performance computing. Based on the IEEE-754 standard, floating-point addition and subtraction operations usually require a complex processing process, including exponent alignment, mantissa addition and subtraction, result normalization, etc., to ensure the accuracy of the calculation result and the standardization of the format.
[0033] According to the definition of the IEEE-754 specification, the mantissa (also known as the significand) of a floating-point number is transmitted and stored in the form of a sign bit plus the original code. For single-precision or double-precision floating-point formats, the all-0 and all-1 encodings of the exponent field are not used as normal exponents, but are reserved for special identifiers to represent special cases such as absolute zero, subnormal numbers (the integer part of the significand is 0), infinity, or not a number (NaN). Under this specification, referring to Figure 1 , the true value of a floating-point number can be expressed as:
[0034] N = (-1) S × 2 E-k × (I + M);
[0035] where k = 2 w-1 - 1 is the exponent offset (related to the bit width W of the exponent E), I is the integer part of the significand, which takes 1 for normal numbers (the exponent field is neither all 1 nor all 0), and 0 for subnormal numbers (the exponent field is all 0 and the mantissa is not 0); M is the fractional part of the mantissa (significand).
[0036] In the addition and subtraction operations of floating-point numbers, due to the need to comply with the IEEE-754 format and take into account the performance requirements of processors / large computing power chips, existing adders and subtracters often play a crucial role. However, traditional implementation methods face significant bottlenecks in terms of logical complexity and circuit scale, and their overall structure can usually be referred to in Appendix Figure 2As shown below, the specific process is as follows:
[0037] 1. Shift decision and exponent alignment operation: The shift decision unit compares the exponent values in the two floating-point operands A and B to determine the magnitude relationship of their exponents, and calculates the absolute difference between the two exponent values as the shift bit control signal for the exponent alignment shifter; The first MUX selects the mantissa of the operand with the smaller exponent according to the comparison result of the shift decision unit and sends it to the exponent alignment shifter for exponent alignment shift. After the exponent alignment shift, the mantissa parts of the two operands are on the same exponent basis, so that addition / subtraction operations can be performed.
[0038] 2. Sign / magnitude to two's complement conversion and fixed-point addition: The mantissa after the exponent alignment shift is sent to the first two's complement obtaining module for conversion from sign / magnitude to two's complement; The second MUX sends the mantissa of the operand with the larger exponent to the second two's complement obtaining module to complete the conversion from sign / magnitude to two's complement; The operator control signal (add / sub) determines the two's complement operation in the two two's complement obtaining modules; The two's complement formatted mantissas output from the two two's complement obtaining modules are sent to the fixed-point adder for binary fixed-point addition operation of signed numbers (including addition / subtraction logic). In this process, the fixed-point adder usually needs to support effective processing of mantissas with negative sign bits.
[0039] 3. Leading 0 / 1 counting and mantissa normalization: The output result of the fixed-point adder may have leading 0s or leading 1s depending on the sign bit. The leading 0 / 1 counter counts the output and sends the counting result to the mantissa left shift operation module for normalization shift operation; After the normalization shift, the result is then sent to the mantissa sign / magnitude restoration module to be converted back from two's complement to sign / magnitude format and finally sent to the result normalization module.
[0040] 4. Exponent adjustment and result normalization: The exponent adjustment module makes corresponding increases or decreases to the larger exponent in the input operands according to the result of the leading 0 / 1 counter; The result normalization module judges and processes overflow or underflow conditions according to the IEEE-754 specification requirements, and forms the final output result that conforms to the floating-point format with the adjusted exponent and the sign / magnitude mantissa.
[0041] It should be noted that the fixed-point adder involved in the above floating-point addition and subtraction operations (such as Figure 2 ) adopts a fixed-point adder for signed numbers, which is used to support binary addition and subtraction operations where the mantissa may be negative. Since the mantissa segment in the IEEE-754 format is an unsigned sign / magnitude representation, different from the two's complement form of fixed-point numbers, it is necessary to complete the conversion from sign / magnitude to two's complement before performing fixed-point addition and subtraction calculations, and convert the two's complement back to sign / magnitude before outputting to the overflow judgment and normalization modules.
[0042] In this way, these two steps require a total of three original code - complement conversion (or vice versa) circuits at two levels (such as before alignment, before the adder input, after the adder output, etc.). Their carry operations of bitwise inversion and adding 1 will bring a huge logical hierarchy, greatly increasing the circuit delay, becoming the main bottleneck for improving the floating - point addition and subtraction operation speed, thus limiting the working frequency of the entire arithmetic circuit and affecting the computing performance.
[0043] In addition, after the addition and subtraction operations are completed, the mantissa needs to be normalized and shifted to ensure that the output result meets the IEEE - 754 standard. In the existing solutions, the normalization and shifting operations on the mantissa addition and subtraction operation results all adopt a single multiplexer (see Figure 3 ) or a cascade of multi - level 2 - to - 1 selectors (see Figure 4 ), or a multi - level multi - input selector structure between the two.
[0044] 1. Single multiplexer: The unnormalized mantissa SUM output by the adder is used as the input. Each output bit corresponds to a multiplexer (such as 16 - to - 1). The shift control signal output by the leading - zero counter determines which bit of SUM the bit should be taken from. If the value index exceeds the actual bit width, it is replaced by 0. This scheme completes the shift through multiplexing at one time, but the input fan - in of the multiplexer is large, the hardware scale is huge, the logical path is long, and it is easy to limit the circuit speed.
[0045] 2. Cascade of multi - level 2 - to - 1 selectors: The one - time large - scale shift is split into multi - level small - scale shifts. Each level is a 2 - to - 1 selector (2:1 MUX), and the left or right path is determined by each bit of the shift control signal in turn; after all are cascaded, an equivalent overall left - shift is achieved. The bits with boundary overflow are also directly replaced by 0. Although the fan - in of a single 2:1 selector is small, multi - level cascading is required, and the logical hierarchy is deep, which will still increase the signal propagation delay and the circuit scale.
[0046] Generally speaking, when normalizing and shifting the mantissa, both of these two types of schemes require a large number of selector modules. Whether it is a single multiplexer or a multi - level 2 - to - 1 selector, it may lead to a large amount of hardware resource usage and an increase in the critical path delay, thus affecting the overall performance of floating - point operations.
[0047] Although traditional floating - point adders and subtractors follow the IEEE - 754 standard and meet the basic addition and subtraction operation requirements, they still face the following problems:
[0048] 1. High logical complexity and speed limitation: The mantissa needs to be converted between the original code and the complement multiple times before and after the operation. Each conversion requires carry processing, increasing the logical hierarchy and easily becoming the bottleneck of the circuit working frequency;
[0049] 2. The normalization shift module is huge: To achieve fast normalization, existing solutions usually use multi-level selectors to complete multi-bit shift operations. As the precision or bit width of floating-point numbers increases, the circuit scale further expands, seriously affecting the operation performance and power consumption.
[0050] Therefore, how to improve the operation process and normalization shift of floating-point addition and subtraction while maintaining compatibility with the IEEE-754 standard, optimizing the logic level, and improving the processing speed has become an urgent technical problem in the current design of floating-point operation circuits. In this regard, there have been many solutions in the industry trying to optimize the mantissa encoding and operation architecture, but there are still bottlenecks in the limited improvement of computing energy efficiency and performance, making it difficult to meet the higher requirements of large computing power chips for high-speed floating-point operation units. For this reason, the present invention proposes a solution for floating-point adders and subtractors, aiming to improve the overall computing performance and efficiency on the basis of reducing logical complexity while meeting the IEEE-754 specification.
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will elaborate on each embodiment of the present invention with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present invention, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation on the specific implementation manner of the present invention. Each embodiment can be combined and cross-referenced with each other on the premise of no contradiction.
[0052] One embodiment of the present invention relates to a floating-point addition and subtraction device, which can be applied to a variety of computing devices, including but not limited to terminal devices such as mobile phones, computers, and tablets, as well as scenarios such as embedded systems, computing power chips, high-performance computing servers, and artificial intelligence processors. The embodiment of the present invention is used to implement the addition and subtraction operations of a first floating-point operand and a second floating-point operand. The first floating-point operand includes a first exponent and a first mantissa, and the second floating-point operand includes a second exponent and a second mantissa. The floating-point addition and subtraction device includes: a mantissa alignment module, configured to perform a mantissa alignment shift operation on the smaller mantissa and perform an inversion operation on the smaller mantissa after the mantissa alignment shift operation based on an inversion control signal, and output a mantissa alignment result; wherein, the smaller mantissa is the mantissa of the operand with the smaller absolute value among the first floating-point operand and the second floating-point operand; a mantissa operation module, connected to the mantissa alignment module, configured to perform an operation on the larger mantissa and the mantissa alignment result, and output a mantissa operation result; wherein, the larger mantissa is the mantissa of the operand with the larger absolute value among the first floating-point operand and the second floating-point operand; a mantissa processing module, connected to the mantissa operation module, configured to perform a normalization process on the mantissa operation result and output a normalized mantissa; the mantissa processing module includes: a leading zero preprocessing sub-module, configured to generate a prefix or an operation result based on the mantissa operation result; a transmission gate control sub-module, configured to generate a transmission gate control signal based on the prefix or the operation result; a transmission gate array sub-module, configured to normalize the mantissa operation result based on the transmission gate control signal. The floating-point addition and subtraction device provided by the embodiment of the present invention, firstly, through the control mechanism of the inversion control signal, only performs one or no complement calculation operation on the smaller mantissa, avoiding multiple conversions between the original code and the complement code, significantly reducing the circuit complexity and hardware resource waste, and thus improving the operation speed and efficiency. Subsequently, the mantissa operation module performs an efficient operation on the aligned smaller mantissa and the larger mantissa, and ensures that the final mantissa output is always positive, simplifying the leading zero / one counting logic to a leading zero counting operation, further reducing the circuit scale and complexity. In the mantissa processing module, through the sharing mechanism with the leading zero preprocessing function, the prefix or result generated by the leading zero preprocessing module can be directly used as the input of the transmission gate control sub-module to generate a dynamic transmission gate control signal. This design integrates part of the functions of the leading zero counting module into the normalization circuit, not only effectively shortening the logic path but also reducing the delay caused by too deep cascaded logic. Finally, the transmission gate array sub-module adopts an efficient mantissa normalization shift operation, significantly improving the mantissa shift efficiency. The overall solution is fully compatible with the IEEE-754 standard, capable of significantly improving the addition and subtraction operation efficiency and energy efficiency performance while ensuring high precision, meeting the requirements of modern high-performance computing for high efficiency, low power consumption, and high precision.
[0053] The following is combined withFigure 5 , the implementation details of the floating - point addition and subtraction device of the embodiments of the present invention are specifically described. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.
[0054] As Figure 5 shown, the floating - point addition and subtraction device of the embodiments of the present invention includes a comparison and decision - making module 1, a mantissa alignment module 2, a sign - processing module 3, a mantissa arithmetic module 4, a mantissa - processing module 5, an exponent adjustment module 6, and a result normalization module 7. The signals between each module are closely related, and they cooperate to achieve efficient and accurate floating - point operations. Specifically, the comparison and decision - making module 1 compares the exponents and mantissas of the operands, and generates control signals for mantissa alignment and sign processing; the mantissa alignment module 2 aligns and shifts the smaller mantissa according to the exponent difference, and completes the mantissa negation in combination with the negation control signal generated by the sign - processing module 3. The mantissa arithmetic module 4 performs unsigned addition operations on the aligned mantissas, and the results are passed to the mantissa - processing module 5 for leading - zero output and mantissa normalization processing respectively. The exponent adjustment module 6 corrects the exponent in combination with the larger exponent value and the leading - zero count value, and the result normalization module 7 combines the normalized mantissa, the corrected exponent, and the sign bit into a floating - point number that conforms to the IEEE - 754 standard for output.
[0055] The structures and functions of each module related to the embodiments of the present invention are as follows:
[0056] The comparison and decision - making module 1, that is, Figure 5 the absolute - value comparison and decision - making module in . It is connected to the mantissa alignment module 2 and the mantissa arithmetic module 4, and is used to compare the absolute values of the first floating - point operand and the second floating - point operand, and output the smaller mantissa and the absolute value of the difference between the exponents of the first floating - point operand and the second floating - point operand. The comparison and decision - making module 1 is the core part of the floating - point addition and subtraction device. It is mainly used to receive the order and mantissa data of the first floating - point operand and the second floating - point operand, determine the smaller and larger operands by comparing the absolute values of the two operands, and provide necessary input signals for subsequent alignment shifting and mantissa arithmetic operations.
[0057] The outputs of the comparison and decision - making module 1 include the smaller mantissa (LIT), the larger mantissa (BIG), the absolute value of the exponent difference (eSHIFT), the larger exponent value (eMAX), and the control signal of the mantissa size relationship (ALTB). The comparison and decision - making module 1 ensures that the subsequent modules can operate efficiently with correct data and control signals through precise exponent comparison, mantissa restoration, and mantissa size relationship judgment, thereby supporting the performance optimization and accuracy guarantee of the entire floating - point adder - subtractor.
[0058] In one embodiment, as Figure 6As shown, the comparison decision module 1 includes sub-modules such as a shift decision maker 101, a mantissa restorer 102, a mantissa comparator 103, and a MUX (multiplexer). Among them, the shift decision maker 101 completes a series of logical operations by comparing the exponent parts of the input operands, including judging whether the exponents are equal and outputting a signal eq (the value of 1 indicates that the exponents are equal, and 0 indicates that the exponents are not equal), comparing the magnitude relationship of the exponents and outputting a signal eaLTeb (the value of 1 indicates that the exponent of the first operand is less than the exponent of the second operand, and 0 indicates the opposite), and calculating the absolute value of the difference between the exponents of the two operands eSHIFT = ∣eA - eB∣. In addition, the shift decision maker also outputs the larger exponent value eMAX = Max(eA - k, eB - k), where k is the exponent offset value, which is related to the bit width of the exponent encoding in the IEEE-754 standard. The above calculation results provide the necessary inputs for the subsequent exponent alignment shift operation.
[0059] The above mantissa restorer 102 restores the hidden integer bits in the mantissa segment according to the exponent value of the input operand. As Figure 7 shown, for a normalized number (i.e., an operand whose exponent segment is neither all 1s nor all 0s), the restored mantissa integer bit is 1; for a sub-normalized number (i.e., an operand whose exponent segment is all 0s but the mantissa segment is non-zero), the restored mantissa integer bit is 0. This process ensures that the restored mantissa meets the requirements of the floating-point number format in the IEEE-754 standard and provides complete data for subsequent mantissa comparison.
[0060] The above mantissa comparator 103 compares the absolute values of the two mantissas output by the mantissa restorer 102. If eq = 1 (the exponents are equal), the mantissa comparator directly compares the magnitudes of the mantissa segments and outputs a signal ALTB; the value of 1 indicates that the mantissa of the first operand is less than the mantissa of the second operand, and 0 indicates the opposite. If eq = 0 (the exponents are not equal), the output signal ALTB is directly determined by the exponent magnitude relationship signal eaLTeb.
[0061] The above multiplexer (MUX) controls the data path according to the output signal ALTB of the comparison decision module 1, transfers the mantissa with the smaller absolute value (LIT) to the mantissa alignment module, and transfers the mantissa with the larger absolute value (BIG) to the mantissa operation module. Through this data path control, it is ensured that the exponent alignment shift and mantissa operation modules can correctly process the input data.
[0062] In the above embodiment, taking specific operations as an example, assume that the exponent eA of the first operand is 5 and the mantissa MA is 0.25; the exponent eB of the second operand is 3 and the mantissa MB is 0.75. The shift decision maker calculates eSHIFT = |eA - eB| = 2 and outputs eq = 0 (indicating that the exponents are not equal) and eaLTeb = 0 (indicating that the exponent of the first operand is larger). The mantissa restorer supplements the hidden integer bits of the mantissa and calculates that the mantissa MantA of the first operand is (1 + 0.25) = 1.25, and the mantissa MantB of the second operand is (1 + 0.75) = 1.75. It outputs ALTB = 0 (indicating that the mantissa of the first operand is larger). The multiplexer finally outputs the mantissa LIT of the smaller operand as 1.75, the mantissa BIG of the larger operand as 1.25, and the exponent difference eSHIFT as 2. Based on the collaborative work of sub-modules such as the shift decision maker 101, mantissa restorer 102, mantissa comparator 103, and MUX (multiplexer) in the above embodiment, the inputs of the comparison and decision module 1 include the exponent eA and mantissa MA of the first floating-point operand, and the exponent eB and mantissa MB of the second floating-point operand. The output signals of the module include the smaller mantissa (LIT), the larger mantissa (BIG), the absolute value of the exponent difference (eSHIFT), the larger exponent value (eMAX), and the mantissa comparison result signal (ALTB). Through the difference calculation and magnitude comparison of the exponent and mantissa, this module generates the required accurate control signals and data inputs for subsequent modules.
[0063] The mantissa alignment module 2, that is Figure 5 the exponent alignment shift and negation module in
[0064] In one embodiment, the exponent alignment shift operation of the mantissa alignment module 2 is implemented by a multiplexer (such as attached Figure 8 ). The selector dynamically selects the corresponding bits in the input mantissa LIT for output according to the value of the shift control signal eSHIFT. As attached Figure 9a , when eSHIFT = 0, all bits of LIT are directly output to LITTLE; as attached Figure 9b, when eSHIFT > 0, the high bits of LIT are shifted to the right by eSHIFT bits in sequence, and the low bits are filled with zeros. Each bit of LITTLE is determined by the corresponding shifted result. For example, when the output mantissa LITTLE[0] selects LIT[15:0] as the input of the selector, LITTLE[1] will correspondingly select LIT[16:1] as the selector input, and so on, until the output mantissa LITTLE[i] uses LIT[i + 15:i] as the selector input. Through this design, the module can flexibly adjust the alignment result of the mantissa according to the shifting requirements to ensure the accuracy and integrity of the output data.
[0065] In the above embodiment, the output of the mantissa alignment module 2 is divided into a high significant bit segment LITTLE and a low significant bit segment TAIL. LITTLE is the mantissa after alignment shifting, and its bit width is the same as that of the input mantissa LIT; TAIL is the remainder segment generated during the shifting process and is used for subsequent precision compensation processing. The bit width of TAIL is determined by the bit width of eSHIFT. For example, when the bit width of eSHIFT is 4 bits, the bit width of TAIL is 2 4 -1 = 15. The generation of TAIL uses the same multiplexer structure as LITTLE (as shown in the appendix Figure 8 ), but the selected data comes from the lower bit part of LIT. For example, when the output remainder TAIL
[14] uses LIT[14:-1] as the selector input, the output remainder TAIL
[13] will correspondingly select LIT[13:-2] as the selector input, and so on, until the lowest bit of the output remainder TAIL[0] uses LIT[0:-15] as the selector input.
[0066] In the above embodiment, if during the shifting process, the selected LIT bit positions exceed the actual data bit width boundary of the input mantissa LIT (including the high bit overrun on the left and the low bit overrun on the right), as shown in the appendix Figure 9b , the selector will automatically fill with 0 to ensure the integrity of the output data. For example, if the actual bit width of LIT is 24 bits (i.e., LIT[23:0]), for the bits outside this range (such as LIT
[24] or LIT[-1]), the selector will automatically fill with 0 to ensure the integrity of the output data.
[0067] In the above embodiment, the mantissa alignment module 2 also includes an inversion operation. As Figure 8 , after the shifting is completed, the mantissa alignment module 2 performs a two's complement inversion operation on LITTLE and TAIL according to the inversion control signal sub generated by the sign processing module 3. The inversion is implemented through an exclusive OR gate. When sub = 0, it represents an addition operation, and the exclusive OR gate directly outputs the result of the multiplexer; when sub = 1, it represents a subtraction operation, and the exclusive OR gate outputs the inverted result.
[0068] In addition, to support the operation of modified complement, the mantissa alignment module 2 also provides a carry-in flag CIN. Only when eSHIFT = 0 and sub = 1 (i.e., the exponents of the two operands are equal and a subtraction operation is performed), the CIN outputs 1 to correct the bias of the modified complement; in other cases, the CIN outputs 0.
[0069] In the above embodiment, LITTLE is directly passed to the mantissa operation module for unsigned fixed-point addition operation. And TAIL, as an optional item for precision compensation, can be passed to the mantissa processing module 5 for normalization operation. In scenarios with lower precision requirements, the output of TAIL and its related circuits can be selectively disabled, thereby reducing circuit resource consumption and power consumption.
[0070] The design of the above mantissa alignment module 2 provides support for high-precision and efficient calculation in floating-point addition and subtraction operations by separating the output of high and low bit segment data, flexibly controlling the complement operation, and the generation and selection of remainders. The overall design not only reduces the logical complexity but also improves the computational performance and resource utilization rate of the circuit, can adapt to a variety of application scenarios, and fully meets the requirements of efficient and accurate floating-point calculations.
[0071] Table 1
[0072]
[0073] The sign processing module 3, namely Figure 5 the sign bit processing module in
[0074] In one embodiment, the symbol processing module 3 generates a mantissa operation type control signal sub and a result sign bit signal sign by comprehensively analyzing the sign bits (SA and SB) of the input operands A and B, the operator add_sub (for selecting addition or subtraction operations), and the ALTB signal output by the comparison decision module 1 (for indicating the selection of the larger operand). The relationship between its input values and output signals is shown in Table 1. Specifically, the sub signal is used to indicate the type of mantissa operation: when sub = 1, it means to perform an absolute value subtraction operation (subtracting the smaller mantissa from the larger mantissa); when sub = 0, it means to perform an absolute value addition operation (adding the smaller mantissa to the larger mantissa). The sign signal represents the sign bit of the operation result, sign = 0 indicates that the result is positive, and sign = 1 indicates that the result is negative. The generated sub signal is directly sent to the mantissa alignment module 2 to control whether to perform a two's complement inversion operation on the mantissa after mantissa alignment, while the sign signal is transmitted to the result normalization module 7 and combined with the exponent and mantissa of the calculation result to form a floating-point number format output that conforms to the IEEE-754 standard.
[0075] The design of the above symbol processing module 3 greatly simplifies the overall operation logic by completing symbol judgment and mantissa inversion control before the operation. The smaller mantissa only performs inversion processing when sub = 1 (absolute value subtraction operation), while the larger mantissa always participates in the operation in its original code form, avoiding multiple conversions between the original code and the two's complement in traditional logic. This optimization effectively reduces the complexity of the operation path, significantly reduces the critical path delay, and thus improves the overall calculation speed and efficiency of the floating-point adder / subtractor. In addition, the symbol processing module 3 completes the symbol processing logic in advance by directly generating and providing the sign bit signal of the final result to the result normalization module 7, optimizing the complexity of the subsequent result normalization process, and further enhancing the calculation performance and resource utilization rate. This module design fully considers the accuracy and efficiency of symbol processing, providing strong support for the high-performance operation of the floating-point operation device.
[0076] The mantissa operation module 4, that is Figure 5 the unsigned fixed-point full adder in. It is connected to the mantissa alignment module 2 and is used to perform operations on the larger mantissa (BIG) and the mantissa alignment result (LITTLE), and output the mantissa operation result (SUM). Among them, the larger mantissa is the mantissa of the operand with the larger absolute value among the first floating-point operand and the second floating-point operand. The mantissa operation module 4 is implemented by cascading fixed-point full adders, and its internal logic architecture is as shown in the appendix Figure 10 shown, which can efficiently complete the mantissa operation and transfer the result to the subsequent module for further processing.
[0077] In one embodiment, the inputs of the mantissa operation module 4 include the larger mantissa BIG, the aligned smaller mantissa LITTLE, and the two's complement plus 1 flag signal CIN from the mantissa alignment module 2. The adder calculates BIG and LITTLE in a bitwise unsigned addition manner to generate the mantissa operation result SUM. Among them, SUM[n:0] is the mantissa operation result, and the highest bit SUM[n] represents the carry flag of the addition operation. SUM and the carry flag are directly passed to the mantissa processing module 5 for further normalization operations and result correction.
[0078] In this embodiment, when the CIN signal is valid (i.e., CIN = 1), the adder performs a two's complement plus 1 operation on the aligned and shifted smaller mantissa when the orders of the two operands are equal (eSHIFT = 0) and it is an absolute value subtraction operation (sub = 1), so as to accurately correct the mantissa deviation that may occur during the two's complement calculation process. In the case where the orders of the operands are not equal, since the influence of the two's complement plus 1 operation has been absorbed by the remainder TAIL generated in the shift operation, and the accuracy correction is completed through the remainder supplement in the subsequent normalization step, the CIN signal outputs 0 in this case to avoid imposing an additional burden on the operation logic of the adder.
[0079] The mantissa processing module 5 (as shown in the appendix Figure 5 ), is connected to the mantissa operation module 4 and is used to perform normalization processing on the mantissa operation result (SUM) and output the normalized mantissa (MantiY). Among them, the mantissa processing module includes: a leading zero preprocessing sub-module, which is used to generate a prefix or operation result (OROUT) based on the mantissa operation result; a transmission gate control sub-module, which is used to generate a transmission gate control signal based on the prefix or operation result; a transmission gate array sub-module, which is used to normalize the mantissa operation result based on the transmission gate control signal.
[0080] In the above module structure, part of the functions of the leading zero counting module are integrated into the normalization circuit through the leading zero preprocessing operation. The dynamically generated OROUT supports the generation of both the transmission gate control signal and the zero-padding signal at the same time, significantly optimizing the logic path and reducing the delay caused by the cascaded logic depth.
[0081] Moreover, based on the above mantissa processing module 5, a leading zero counting sub-module is further included, which is used to generate a leading zero count value (LZCNT) based on the prefix or operation result (OROUT) output by the leading zero preprocessing sub-module.
[0082] In one embodiment, as shown in the appendix Figure 11, the mantissa processing module 5 includes a leading zero preprocessing sub-module 501, a transmission gate control sub-module, a transmission gate array sub-module 504, and a leading zero counting sub-module 505. Among them, the transmission gate control sub-module includes a zero-padding signal generation sub-module 502 and a strobe signal generation sub-module 503. Based on the above structure, the mantissa processing module 5 receives the mantissa operation result SUM and the remainder TAIL as inputs, and the outputs include two parts of data: the leading zero count value LZCNT and the normalized mantissa MantiY. The leading zero count value LZCNT represents the number of "0"s before the first "1" counted from the high bit to the low bit in the SUM data. For example, for SUM[24:0] with a bit width of 25 bits, if the value of SUM is 0_00011100_11000010_10001101, then LZCNT = 4, indicating that there are 4 "0"s before the first "1". The normalized mantissa MantiY is the result after normalized shift processing and high significant bit truncation. Its highest significant bit is fixed at 1, and the rest are decimal bits, conforming to the IEEE-754 specification.
[0083] The leading zero preprocessing module 501 in the above structure generates a preprocessing result OROUT by performing a bitwise "OR" operation on the concatenated data {SUM, TAIL} of SUM and TAIL. The calculation formula for the i-th bit of OROUT is:
[0084] OROUT[i] = SUM
[24] || SUM
[23] ||... || TAIL[i];
[0085] For example, SUM[24:0] = 0100_1100_xxxx_xxxx_xxxx_xxxx, TAIL[14:0] = xxx_xxxx_xxxx_xxxx, then OROUT[39:0] = 0111_1111_1111_1111_1111_1111_1111_1111_1111_1111. In addition, under the bit width configuration of 25-bit SUM and 15-bit TAIL, the hardware implementation of the leading zero preprocessing sub-module 501 is as shown in the appendix Figure 12a and the appendix Figure 12b shown, including two main circuit architectures: serial cascading mode ( Figure 12a ) and parallel direct connection mode ( Figure 12b ). Figure 12a The serial cascading implementation shown uses an OR gate tree structure, with the least amount of logic units used, but the lowest bit output needs to go through more logic levels, which may affect the circuit speed; Figure 12bThe parallel direct connection implementation in uses a preprocessing circuit with parallel outputs. Although it uses more logic units, the logical level of the least significant bit output is significantly reduced, resulting in a faster circuit speed. In practical applications, the appropriate architecture can be selected according to the specific scenario, and it is not limited to the above two architectures.
[0086] The zero-padding control signal generation sub-module 502 in the above structure generates a control signal yclr based on the input OROUT, which is used to control the transmission gates in the sub-module 504 to complete the zero-padding operation on the lower bits of the normalized mantissa MantiY. The bit width of the generated yclr signal is the same as the bit width of the output mantissa MantiY after the normalized shift operation. For example, in the Figure 13 configuration shown in the appendix, the bit width of MantiY is 24 bits, so the bit width of yclr is 24 bits, that is, yclr[23:0]. The generation logic of yclr is: only process the lower 24 bits of OROUT (OROUT[23:0]). The specific operation is to invert each bit of OROUT[23:0] and then swap the high and low bit orders before outputting. Its calculation formula is: yclr[i] = ~OROUT[23 - i]. The appendix Figure 13 shows the implementation method of this generation logic and its circuit structure. Based on this circuit structure, the input-output truth table shown in Table 2 can be realized.
[0087] Table 2
[0088]
[0089] The transmission gate gating signal generation sub-module 503 in the above structure generates a transmission gate control signal sel through the input OROUT, which is used to control the switching state of the shift transmission gate array. Specifically, the sel signal determines which group of transmission gates is in the open state, and the remaining transmission gates remain closed, so as to achieve precise control of the shift path. The bit width of sel is equal to the sum of the bit widths of the concatenated mantissa SUM and the remainder TAIL. For example, in the appendix Figure 14 shown bit width configuration, SUM is 25 bits and TAIL is 15 bits, so the bit width of sel is 40 bits. The transmission gate gating signal generation sub-module 503 is implemented using an exclusive OR gate circuit. The appendix Figure 14 shows its circuit structure and specific implementation method. Based on this circuit structure, the sel signal output truth table shown in Table 3 can be realized.
[0090] Table 3
[0091]
[0092]
[0093] In the above structure, the normalized shift transmission gate array module 504 selects a specific shift path through the control signal sel, and combines the zero-padding control signal yclr to perform zero-padding processing on the lower bits after shifting, so as to implement the normalized shift operation on the spliced data {SUM, TAIL}. The finally output normalized mantissa MantiY, its most significant bit is the integer part, fixed as 1, and the remaining bits represent the decimal part, meeting the mantissa format requirements of the IEEE-754 standard. The specific circuit implementation of this module is as shown in the appendix Figure 15 As shown, through the efficient transmission gate array design, the accuracy and hardware efficiency of mantissa normalization shift are ensured.
[0094] In the above structure, the leading zero counter module 505 is used to count the number of "0"s in the preprocessing result OROUT output by the leading zero preprocessing module 501. Its implementation method is as shown in the appendix Figure 16 As shown, this circuit example combines the bitwise inversion logic and the multi-bit accumulator structure to complete the zero counting operation. Specifically, the input data OROUT first undergoes bitwise inversion processing, and then is sent to the bitwise accumulator. The accumulator accumulates the values of all bits to obtain the total number of zeros LZCNT. For example, when the input data OROUT is 0101...011, the counting result is the accumulated sum of 0+1+0+1+...+0+1+1.
[0095] As an optimization option, if the floating-point calculation accuracy requirement does not require the remainder TAIL to participate in the accuracy compensation, the logic related to TAIL can be selectively disabled to reduce hardware resources and power consumption. In the case where TAIL participates, {SUM, TAIL} participates in the leading zero counting and normalization shift processing as a whole. For example, for the left shift operation, the processing logic is {SUM, TAIL}×2^LZCNT, achieving higher calculation accuracy.
[0096] Based on the above embodiment structure, the operations shown in Figure 17a and Figure 17b can be implemented. As shown in Figure 17a , when SUM[n]=1, it means no shift is required and the mantissa is directly output. At this time, the leading zero count value LZCNT is 0; as shown in Figure 17b , when SUM[n,n-1]=01, the mantissa SUM will be shifted left by 1 bit. At this time, the leading zero count value LZCNT is 1; as shown in Figure 17c , when SUM[n,n-4]=0001, the mantissa SUM will be shifted left by 3 bits. At this time, the leading zero count value LZCNT is 3.
[0097] In the above embodiments, some functions of the leading zero counting module (i.e., the prefix or processing operation) are integrated into the normalization circuit. By directly generating OROUT through the leading zero preprocessing module and providing support for the generation of the transmission gate control signal and the zero-padding signal, not only the complexity of the independent leading zero counting circuit is reduced, but also the logical path is significantly shortened, and the delay caused by too deep cascaded logic is reduced. In addition, by dynamically generating the transmission gate control signal and combining with an efficient transmission gate array design, this module further optimizes the logical process of the mantissa normalization process, reduces the occupancy of hardware resources, and at the same time greatly improves the accuracy and operation efficiency of the mantissa calculation.
[0098] The exponent adjustment module 6 (such as Figure 5 ) is connected to the comparison and decision module 1 and the mantissa processing module 5, and is used to adjust a larger exponent based on the leading zero count value to generate an exponent adjustment value representing the normalized mantissa order; wherein, the larger exponent is the exponent of the operand with the larger absolute value among the first floating-point operand and the second floating-point operand.
[0099] In one embodiment, the exponent adjustment module 6 is used to adjust and correct the exponent according to the larger exponent value eMAX output by the comparison and decision module 1 and the leading zero count value LZCNT generated by the mantissa processing module 5 to generate an order value eOUT that matches the result mantissa MantY after normalization shift. This module completes the exponent correction calculation through the formula eOUT = eMAX - LZCNT + 1, where eMAX represents the exponent value of the larger operand, LZCNT is the number of leading zeros in the mantissa, and the correction factor “+1” is used to compensate for the influence of the normalization shift operation on the exponent.
[0100] In the above embodiments, when LZCNT = 0, the mantissa does not need to be shifted and the exponent remains unchanged, eOUT = eMAX. When LZCNT > 0, the exponent decreases according to the left shift number of the mantissa to ensure that the mantissa and the exponent match and meet the requirements of the floating-point format of the IEEE-754 standard. The output eOUT of the exponent adjustment module 6 is passed to the result normalization module 7 to form the final floating-point result together with the mantissa and the sign bit.
[0101] The result normalization module 7 (such as in the appendix Figure 5 ) is used to combine the normalized mantissa, the final sign bit signal, and the exponent adjustment value into a floating-point number that meets a preset standard.
[0102] In one embodiment, the result normalization module 7 combines the input exponent adjustment value eOUT and the normalized mantissa MantiY to complete the normalized combination of the result, while handling special cases such as subnormal numbers or overflow results. When eOUT ≤ -k (where k is the exponent offset defined by the IEEE-754 standard), the module converts the mantissa MantiY into the form of a subnormal number. The specific operations include recalculating the exponent segment of the floating-point number eY = eOUT + k, and hiding the integer bits in the mantissa, only retaining the fractional part of the mantissa to form a mantissa segment that conforms to the definition of a subnormal number. In this case, the output floating-point result is represented as a subnormal number, meeting the format requirements of the standard for subnormal numbers.
[0103] In the above embodiment, for cases that do not meet the conditions of subnormal numbers, the result normalization module 7 directly combines the mantissa MantiY and the exponent adjustment value eOUT to generate the exponent segment and the mantissa segment of the floating-point result, while adding the sign bit to the result, and finally outputs the complete floating-point number Y. The format of Y is (-1) sign ×2 eY ×MantiY, where the sign bit is provided by the sign processing module 3, and the exponent segment eY and the mantissa segment are truncated or rounded according to the IEEE-754 standard.
[0104] The above result normalization module 7 ensures the result processing in case of overflow or underflow. For example, when eOUT exceeds the maximum exponent value defined by the standard, the result is directly marked as infinity; when eOUT is too small and the mantissa cannot be expressed as a subnormal number, the result is marked as zero. This processing method effectively prevents the output of abnormal data during the operation process.
[0105] In an embodiment of the present invention, first, through the negation control signal control mechanism, only one or no negation operation is performed on the smaller mantissa, avoiding multiple conversions between the original code and the complement code, significantly reducing the circuit complexity and hardware resource waste, and thus improving the operation speed and efficiency. Subsequently, the mantissa operation module performs efficient operations on the aligned smaller mantissa and larger mantissa, and ensures that the final mantissa output is always positive, simplifying the leading zero / one counting logic to a leading zero counting operation, further reducing the circuit scale and complexity. In the mantissa processing module, through the sharing mechanism with the leading zero preprocessing function, the prefix or result generated by the leading zero preprocessing module can be directly used as the input of the transmission gate control sub-module to generate a dynamic transmission gate control signal. This design integrates part of the functions of the leading zero counting module into the normalization circuit, not only effectively shortening the logic path but also reducing the delay caused by too deep cascaded logic. Finally, the transmission gate array sub-module adopts an efficient mantissa normalization shift operation, significantly improving the mantissa shift efficiency. The overall solution is fully compatible with the IEEE-754 standard, can significantly improve the addition and subtraction operation efficiency and energy efficiency performance while ensuring high precision, and meets the requirements of modern high-performance computing for high efficiency, low power consumption, and high precision.
[0106] In addition, the examples mentioned in the above embodiments can be freely combined, and any combination method can be understood as an embodiment. The "embodiment" or "example" that appears at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments.
[0107] It is worth mentioning that each module involved in this embodiment is a logic module. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or can be implemented by a combination of multiple physical units. In addition, to highlight the innovative part of the present invention, units not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0108] Another embodiment of the present invention relates to a floating-point addition and subtraction method, and the specific steps are as Figure 18As shown, it is applied to the floating-point addition and subtraction device described above, including: receiving a first floating-point operand and a second floating-point operand; wherein, the first floating-point operand includes a first exponent and a first mantissa, and the second floating-point operand includes a second exponent and a second mantissa; generating an inversion flag and a final sign bit based on the sign bits of the first floating-point operand and the second floating-point operand and the operator; performing an order alignment shift operation on the smaller mantissa, and determining whether to perform an inversion operation on the aligned mantissa according to the inversion flag; performing an addition operation on the aligned mantissa and the larger mantissa to obtain a mantissa operation value; generating a prefix or an operation result based on the mantissa operation value, and generating a transmission gate control signal and a leading zero count value; performing a normalization shift based on the transmission gate control signal to output a normalized mantissa; adjusting the exponent of the operand with a larger exponent based on the leading zero count value to output an exponent adjustment value; performing a normalization process on the normalized mantissa, the final sign bit signal, and the exponent adjustment value to output a floating-point number that meets the preset standard.
[0109] In an alternative embodiment, the above-mentioned performing a normalization shift based on the transmission gate control signal to output a normalized mantissa specifically includes: inputting the prefix or the operation result into a transmission gate control module to generate a strobe signal and a zero-padding signal; shifting the mantissa operation value based on the strobe signal in the transmission gate control signal, and performing a zero-padding process on the normalized shifted mantissa based on the zero-padding signal in the transmission gate control signal to output a normalized mantissa.
[0110] In an alternative embodiment, the above-mentioned floating-point addition and subtraction method may further include: outputting a remainder when performing an order alignment shift operation on the smaller mantissa; after performing the normalization shift, supplementing the remainder to the normalized mantissa. Specifically, if a remainder TAIL is output when the mantissa alignment module in step 4 above performs an order alignment shift operation on the smaller mantissa. Then, after the mantissa processing module in step 5 performs the normalization shift, the remainder TAIL will be supplemented to the normalized mantissa, thereby further improving the accuracy of the floating-point addition and subtraction result.
[0111] When the above method embodiment is applied to the device embodiment of the present invention, the specific steps executed are as follows:
[0112] 1. The floating-point addition and subtraction device receives a first floating-point operand and a second floating-point operand, which respectively include a first exponent and a first mantissa, and a second exponent and a second mantissa. These operands are used as the basic inputs for subsequent operations and are parsed and compared by the comparison and decision module 1 and the sign processing module 3 of the floating-point addition and subtraction device.
[0113] 2. The comparison and decision module 1 generates the smaller mantissa LIT and the larger mantissa BIG by comparing the exponents and mantissas of the operands, calculates the absolute value eSHIFT of the exponent difference, and the comparison result ALTB for mantissa alignment and sign processing. The smaller mantissa LIT is passed to the mantissa alignment module 2 to complete the order alignment shift operation, while the larger exponent value eMAX is used for subsequent exponent adjustment operations.
[0114] 3. Based on the sign bits SA and SB of the operands, the arithmetic operator add_sub, and the comparison result ALTB generated by the comparison and decision module 1, the sign processing module 3 generates the arithmetic type control signal inversion flag sub and the final sign bit signal sign. Among them, sub is used to indicate whether the mantissa needs to be inverted, and sign represents the sign of the operation result, which is passed to the result normalization module 8 to construct the final output.
[0115] 4. The mantissa alignment module 2 performs a right shift operation on the smaller mantissa LIT according to the exponent difference eSHIFT to complete the order alignment shift operation. At the same time, it judges whether to perform two's complement inversion on the aligned mantissa according to the sub signal. The output mantissa alignment result LITTLE is used as the input for mantissa arithmetic, and the remainder segment TAIL is used for precision compensation or can be optionally disabled to optimize hardware resources and power consumption.
[0116] 5. The mantissa arithmetic module 4 performs an unsigned fixed-point addition operation on the larger mantissa BIG and the aligned smaller mantissa LITTLE to generate the mantissa arithmetic value SUM. If the CIN signal is valid (CIN = 1), the module corrects the two's complement deviation by adding 1 to ensure the accuracy and correctness of the operation.
[0117] 6. The mantissa processing module 5 receives the mantissa arithmetic result SUM, and after processing, outputs the leading zero count value LZCNT and the normalized shift result MantY. LZCNT is used to control the exponent adjustment module 6 to correct the exponent.
[0118] 7. The exponent adjustment module 6 combines the larger exponent eMAX and LZCNT to perform a decreasing correction on the exponent, generating the exponent adjustment value eOUT that matches the normalized mantissa. The corrected exponent reflects the impact of the mantissa shift on the result representation, ensuring that the exponent and the mantissa are consistent.
[0119] 8. The result normalization module 7 combines the normalized mantissa MantiY, the final sign bit sign, and the corrected exponent eOUT into a complete floating-point format for output. For special cases (such as subnormal numbers or exponent overflow), the module performs special processing to ensure that the result conforms to the IEEE-754 standard.
[0120] The final output results of the above steps 1 to 8 are output in floating-point form, including a sign bit, a mantissa, and an exponent, meeting the preset calculation accuracy and specification requirements.
[0121] Inside the mantissa processing module 5 in step 6 above, the following steps are also executed among the sub-modules:
[0122] 6.1 The leading zero preprocessing sub-module 501 receives the mantissa operation result SUM, performs a bitwise "OR" operation on SUM, and generates a preprocessing result OROUT. Among them, the calculation formula for the i-th bit of OROUT is:
[0123] OROUT[i] = SUM
[24] || SUM
[23] ||... || TAIL[i];
[0124] 6.2 The zero-padding control signal generation sub-module 502 generates a control signal yclr based on the input OROUT. For example, in the Figure 13 configuration shown, the bit width of MantiY is 24 bits. Therefore, the bit width of yclr is 24 bits, that is, yclr[23:0]. The generation logic of yclr is: only process the lower 24 bits of OROUT (OROUT[23:0]). The specific operation is to reverse the order of the high and low bits after taking the bitwise inverse of OROUT[23:0] and then output. Its calculation formula is: yclr[i] = ~OROUT[23 - i].
[0125] 6.3 When step 6.2 is executed, the strobe signal generation sub-module 503 generates a transmission gate control signal sel through the input OROUT to control the on / off state of the shift transmission gate array.
[0126] 6.4 The transmission gate array module 504 receives the strobe signal sel and the zero-padding signal yclr, selects a specific shift path through the control signal sel, and combines the zero-padding control signal yclr to perform zero-padding processing on the shifted low bits, thereby realizing the normalization shift operation of SUM.
[0127] 6.5 When step 6.4 is executed, the leading zero counter module 505 receives the prefix OR processing result OROUT and counts the number of "0"s in OROUT. Specifically, the input data OROUT first undergoes a bitwise inverse process and then is sent to a bitwise accumulator (as shown in the Figure 16 attachment), and the accumulator accumulates the values of all bits to obtain the total number of zeros LZCNT. For example, when the input data OROUT is 0101...011, the counting result is the cumulative sum of 0 + 1 + 0 + 1 +... + 0 + 1 + 1.
[0128] The step division of the above method is only for clear description. When implemented, it can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this patent; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, is within the protection scope of this patent.
[0129] It is not difficult to find that this embodiment is a method embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details mentioned in the above method embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.
[0130] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made to them in form and details without departing from the spirit and scope of the present invention.
Claims
1. A floating-point addition and subtraction device for performing addition and subtraction operations on a first floating-point operand and a second floating-point operand, where the first floating-point operand includes a first exponent and a first mantissa, and the second floating-point operand includes a second exponent and a second mantissa, characterized in that, The floating-point addition and subtraction device includes: A mantissa alignment module, which is used to perform order alignment shift operations on the smaller mantissa, and perform an inversion operation on the smaller mantissa after the order alignment shift operation based on an inversion control signal, and output a mantissa alignment result; wherein, the smaller mantissa is the mantissa of the operand with a smaller absolute value among the first floating-point operand and the second floating-point operand; A mantissa operation module, connected to the mantissa alignment module, which is used to perform an operation on the larger mantissa and the mantissa alignment result, and output a mantissa operation result; wherein, the larger mantissa is the mantissa of the operand with a larger absolute value among the first floating-point operand and the second floating-point operand; A mantissa processing module, connected to the mantissa operation module, which is used to perform normalization processing on the mantissa operation result and output a normalized mantissa; the mantissa processing module includes: a leading zero preprocessing sub-module, which is used to generate a prefix or an operation result based on the mantissa operation result; a transmission gate control sub-module, which is used to generate a transmission gate control signal based on the prefix or the operation result; a transmission gate array sub-module, which is used to normalize the mantissa operation result based on the transmission gate control signal.
2. The floating-point addition and subtraction device according to claim 1, characterized in that The device further includes: A comparison and decision module, connected to the mantissa alignment module and the mantissa operation module, which is used to compare the absolute value sizes of the first floating-point operand and the second floating-point operand, and output the smaller mantissa and the absolute value of the difference between the exponents of the first floating-point operand and the second floating-point operand; The mantissa alignment module performs a shift operation on the smaller mantissa based on the absolute value of the difference to complete the order alignment shift operation.
3. A floating-point addition and subtraction device according to claim 2, characterized in that, The device further includes: A sign processing module, connected to the comparison and decision module and the mantissa alignment module, which is used to generate the inversion control signal based on the comparison result of the comparison and decision module, the operator, the sign bit of the first floating-point operand, and the sign bit of the second floating-point operand.
4. A floating-point addition and subtraction device according to claim 3, wherein The mantissa processing module further includes a leading zero counting sub-module, which is used to generate a leading zero count value based on the prefix or the operation result; The device further includes: An exponent adjustment module, connected to the comparison and decision module and the mantissa processing module, which is used to adjust the larger exponent based on the leading zero count value to generate an exponent adjustment value representing the order of the normalized mantissa; wherein, the larger exponent is the exponent of the operand with a larger absolute value among the first floating-point operand and the second floating-point operand.
5. A floating-point addition and subtraction device according to claim 4, wherein The sign processing module is further used to generate a final sign bit signal when generating the inversion control signal; the device further includes: A result normalization module, which is used to combine the normalized mantissa, the final sign bit signal, and the exponent adjustment value into a floating-point number that meets a preset standard.
6. The floating-point addition and subtraction device according to claim 1, characterized in that, The transmission gate control signal includes: A strobe signal, which is used to control the normalization shift operation of the transmission gate array sub-module; A zero-padding signal, which is used to control the zero-padding operation of the transmission gate array sub-module; The transmission gate control sub-module includes: A strobe signal generation sub-module, which generates the strobe signal based on the prefix or the operation result; The zero-padding signal generation sub-module generates the zero-padding signal based on the prefix or the operation result.
7. A floating-point addition and subtraction device according to any one of claims 1 to 6, characterized in that The mantissa alignment module is further configured to output a remainder. The mantissa processing module is further configured to receive the remainder and supplement the remainder to the normalized mantissa.
8. A floating-point addition and subtraction method, characterized in that, Applied to the floating-point addition and subtraction device according to any one of claims 1 to 7, comprising: Receiving a first floating-point operand and a second floating-point operand; wherein, the first floating-point operand includes a first exponent and a first mantissa, and the second floating-point operand includes a second exponent and a second mantissa; Generating an inversion flag and a final sign bit based on the sign bits of the first floating-point operand and the second floating-point operand and the operator; Performing a mantissa alignment shift operation on the smaller mantissa, and determining whether to perform an inversion operation on the aligned mantissa according to the inversion flag; Performing an addition operation on the aligned mantissa and the larger mantissa to obtain a mantissa operation value; Generating a prefix or an operation result based on the mantissa operation value, and generating a transmission gate control signal and a leading zero count value; Performing a normalization shift based on the transmission gate control signal and outputting a normalized mantissa; Adjusting the exponent of the operand with the larger exponent based on the leading zero count value and outputting an exponent adjustment value; Normalizing the normalized mantissa, the final sign bit signal and the exponent adjustment value, and outputting a floating-point number that meets a preset standard.
9. A floating-point addition and subtraction method according to claim 8, characterized in that, The performing a normalization shift based on the transmission gate control signal and outputting a normalized mantissa specifically includes: Inputting the prefix or the operation result into a transmission gate control module to generate a strobe signal and a zero-padding signal; Shifting the mantissa operation value based on the strobe signal in the transmission gate control signal, and performing zero-padding processing on the normalized shifted mantissa based on the zero-padding signal in the transmission gate control signal, and outputting a normalized mantissa.
10. A floating-point addition and subtraction method according to claim 8 or 9, characterized in that Further comprising: Outputting a remainder when performing a mantissa alignment shift operation on the smaller mantissa; After performing the normalization shift, supplementing the remainder to the normalized mantissa.