Floating-point number operation method and device
By aligning exponents and performing operations directly on the original code, the method optimizes floating-point number arithmetic circuits, reducing logical depth and enhancing computational efficiency and density.
Patent Information
- Application Number
- CN202510365848.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-15
AI Technical Summary
The traditional floating-point number addition and subtraction calculation architecture requires two-level original code-complement conversion, resulting in an increase in the logic level of the circuit, limiting the computing speed and performance.
By aligning the absolute values of the order digits of floating-point numbers, unsigned fixed-point full addition operation is used to simplify the mantissa operation logic, eliminate the original code-complement conversion, and directly use the unsigned fixed-point binary adder of the original code method.
The working frequency of floating-point number addition and subtraction circuits has been improved, computing energy efficiency has been improved, the scale of logic circuits has been reduced, and the computing power density has been improved.
Smart Images

Figure CN120315670A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of data processing, and in particular, to a floating-point operation method and apparatus. Background Art
[0002] A floating-point number consists of a sign bit, an exponent, and a mantissa, which provides a much larger value range and higher calculation accuracy than a fixed-point number. Processors that support floating-point calculations and computing power chips in the field of artificial intelligence are widely used. The addition and subtraction calculation circuits of floating-point numbers often become important unit circuits that seriously affect key indicators such as computing performance, computing energy efficiency, and computing power density in processor chips and high-computing-power chips.
[0003] According to the definition of the IEEE-754 specification, the traditional floating-point format is defined as Figure 1 shown. Whether it is double precision (abbreviated as FP64), single precision (abbreviated as FP32), or half precision (abbreviated as FP16), it consists of a sign bit, an exponent segment, and a mantissa segment. Among them, S is the sign bit, E is the exponent segment, and M is the mantissa segment. The encodings of all 0s and all 1s in the exponent segment are not used as normal exponents, but are special identifiers reserved for operands to mark absolute 0, subnormal numbers (the integer part of the significant number is 0), infinity, and non-numerical data (abbreviated as NaN). The true value formula of this floating-point number is: N = (-1) s × 2 E-k × (I + M), where: k = 2 w-1 -1 is the exponent offset value, which is related to the bit width W of the exponent E, and is used to represent the actual positive and negative exponents with an unsigned number of W bits. I is the integer part of the significant number, which takes 1 when the floating-point number is a normal number (the exponent bit is neither all 1s nor all 0s), and takes 0 when it is a subnormal number (the exponent bit is all 0s and the mantissa is not 0); M is the pure decimal part of the mantissa (significant number), and the bit [x] from left to right represents 1 / 2 x , and the detailed definition of its specification can be found in the IEEE 754 technical standard specification.
[0004] The mantissas of traditional floating-point numbers are all transmitted and stored in the way of sign bit + original code. Therefore, in the traditional floating-point addition and subtraction calculation architecture, the mantissa is first complemented, and after the binary addition calculation is completed, the original code of the calculation result is obtained again. The changes between these original codes and complement codes require an operation of inverting each bit and then adding 1.
[0005] The inventors found that there are at least the following problems in the related art: This traditional floating-point addition and subtraction calculation architecture requires two levels of a total of three original code - complement code conversion (or vice versa) circuits. These conversion circuits are all implemented by taking the one's complement bit by bit and then adding 1. The carry processing of the add 1 operation will increase the logical level of the circuit and become the bottleneck of the calculation speed of the entire circuit, thereby restricting the operating frequency of the floating-point addition and subtraction operation circuit and affecting the calculation performance. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a floating-point operation method and device, which simplify the complement code calculation required in the calculation process and the logic circuit for restoring the original code from the calculation result, achieving the effects of reducing the logical level of the calculation, increasing the operating frequency of the floating-point addition and subtraction circuit, improving the calculation energy efficiency, and increasing the computing power density.
[0007] To solve the above technical problems, an embodiment of the present invention provides a floating-point operation method, including: aligning the mantissa of the smaller number with the exponent by using the absolute value of the difference between the exponent bits of two floating-point numbers to obtain a shifted remainder and a mantissa of the smaller number with the exponent aligned; wherein, the mantissa of the smaller number is the complete mantissa of the smaller floating-point number obtained according to the absolute value size relationship between the two floating-point numbers; obtaining a mantissa operation type signal and a final calculation result sign bit through logical operations according to the sign bits of the two floating-point numbers, the type of the floating-point operation operator, and the absolute value size relationship between the two floating-point numbers; the mantissa operation type signal is used to indicate whether to perform an inversion operation on the mantissa of the smaller number with the exponent aligned; if the mantissa operation type signal indicates to perform an inversion operation, then perform an unsigned fixed-point full addition operation on the mantissa of the larger number and the mantissa of the smaller number after the inversion operation, otherwise directly perform an unsigned fixed-point full addition operation on the mantissa of the smaller number with the exponent aligned and the mantissa of the larger number; wherein, the mantissa of the larger number is the complete mantissa of the larger floating-point number obtained according to the absolute value size relationship between the two floating-point numbers; perform a normalization shift operation on the result of the unsigned fixed-point full addition operation through the number of leading zeros in the result of the unsigned fixed-point full addition operation and the shifted remainder to obtain a normalized result; perform an overflow judgment and result normalization process on the normalized result by using the number of leading zeros to adjust the exponent of the larger exponent bit among the two floating-point numbers and the final calculation result sign bit.
[0008] An embodiment of the present invention further provides a floating-point arithmetic unit, including: a mantissa alignment module, configured to align the mantissa of the smaller mantissa with the exponent code by using the absolute value of the difference between the exponent bits of two floating-point numbers, and obtain a shifted remainder and a mantissa of the smaller mantissa with the exponent code aligned; wherein, the smaller mantissa is the complete mantissa of the smaller floating-point number obtained according to the absolute value size relationship between the two floating-point numbers; a sign bit processing module, configured to obtain a mantissa operation type signal and a final calculation result sign bit through logical operations according to the sign bits of the two floating-point numbers, the type of floating-point operation operator, and the absolute value size relationship between the two floating-point numbers; the mantissa operation type signal is used to indicate whether to perform an inversion operation on the mantissa of the smaller mantissa with the exponent code aligned; a full addition operation module, configured to, if the mantissa operation type signal indicates to perform an inversion operation, perform an unsigned fixed-point full addition operation on the larger mantissa and the mantissa of the smaller mantissa after the inversion operation, otherwise directly perform an unsigned fixed-point full addition operation on the mantissa of the smaller mantissa with the exponent code aligned and the larger mantissa; wherein, the larger mantissa is the complete mantissa of the larger floating-point number obtained according to the absolute value size relationship between the two floating-point numbers; a normalization module, configured to perform a normalization shift operation on the result of the unsigned fixed-point full addition operation through the number of leading zeros of the result of the unsigned fixed-point full addition operation and the shifted remainder, and obtain a normalization result; a normalization processing module, configured to perform an exponent adjustment on the larger exponent bit of the two floating-point numbers based on the number of leading zeros, and perform an overflow judgment and result normalization processing on the normalization result by using the exponent adjustment result and the final calculation result sign bit.
[0009] In the embodiment of the present invention, the small mantissa is aligned with the exponent code by using the absolute value of the difference between the exponent bits of two floating-point numbers, and the shifted remainder and the small mantissa with aligned exponent code are obtained; wherein, the small mantissa is the complete mantissa of the smaller floating-point number obtained according to the absolute value size relationship between the two floating-point numbers; according to the sign bits of the two floating-point numbers, the floating-point operation operator type, and the absolute value size relationship between the two floating-point numbers, the mantissa operation type signal and the final calculation result sign bit are obtained through logical operations; the mantissa operation type signal is used to indicate whether to perform a negation operation on the small mantissa with aligned exponent code; if the mantissa operation type signal indicates to perform a negation operation, then an unsigned fixed-point full addition operation is performed using the large mantissa and the negated small mantissa, otherwise, the small mantissa with aligned exponent code and the large mantissa are directly subjected to an unsigned fixed-point full addition operation; wherein, the large mantissa is the complete mantissa of the larger floating-point number obtained according to the absolute value size relationship between the two floating-point numbers; the result of the unsigned fixed-point full addition operation is subjected to a normalization shift operation through the number of leading zeros and the shifted remainder of the result of the unsigned fixed-point full addition operation to obtain a normalized result; based on the number of leading zeros, the exponent bit of the larger one of the two floating-point numbers is adjusted exponentially, and the overflow judgment and result normalization processing are performed on the normalized result by using the exponential adjustment result and the final calculation result sign bit. By using the size relationship information, floating-point sign bit information, and floating-point type of the floating-point numbers for optimized sign bit processing, the processing logic of the mantissa part operation in the floating-point addition and subtraction operation is simplified. The unsigned fixed-point binary adder in the original code mode is directly used, most of the operation hardware for the original code - complement code conversion is omitted, the design of the leading 0 / 1 counter is simplified, and the calculation accuracy is ensured without increasing the bit width of the adder by using the shifted remainder of the exponent alignment directly participating in the result mantissa normalization shift. These circuit optimizations significantly reduce the level of the overall logic, achieving the effects of improving the working frequency of the floating-point addition and subtraction circuit, enhancing the calculation energy efficiency, reducing the scale of the logic circuit, and increasing the computing power density.
[0010] In addition, the performing the unsigned fixed-point full addition operation using the large mantissa and the negated small mantissa includes: when the exponent bits of the two floating-point numbers are equal, performing an increment operation on the negated small mantissa, and performing an unsigned fixed-point full addition operation on the large mantissa and the incremented small mantissa; when the exponent bits of the two floating-point numbers are not equal, performing an unsigned fixed-point full addition operation on the large mantissa and the negated small mantissa. By using an unsigned fixed-point full adder to replace the signed fixed-point adder in the existing solution, the increment operation in the calculation process of converting the smaller operand from the original code to the complement code is completed by sharing its circuit, optimizing the processing logic. The unsigned fixed-point full adder performs the mantissa addition without sign participation on the small mantissa and the large mantissa after exponent alignment, and the mantissa is directly sent to the binary fixed-point adder for operation. The operation of obtaining the complement code of the mantissa of the larger operand is omitted.
[0011] In addition, the normalization shift operation on the result of the unsigned fixed-point full addition operation by using the number of leading zeros and the shift remainder of the result of the unsigned fixed-point full addition operation to obtain a normalized result includes: performing a left shift operation on the result of the unsigned fixed-point full addition operation by using the number of leading zeros; and performing precision compensation on the result of the unsigned fixed-point full addition operation after the left shift operation by using the shift remainder to obtain a normalized result. By adding the output of the remainder bit information during the alignment shift operation, it is used to compensate the precision of the mantissa addition calculation result during the mantissa normalization shift operation, saving the effective width of the adder while ensuring the calculation precision of the overall floating-point number.
[0012] In addition, the obtaining of the mantissa operation type signal and the sign bit of the final calculation result through logical operations according to the sign bits of two floating-point numbers, the type of floating-point operation operator, and the magnitude relationship of the absolute values of the two floating-point numbers includes: when the floating-point operation operator type is the addition type and the sign bits of the two floating-point numbers are of the same sign, obtaining the mantissa operation type signal indicating no inversion operation and the sign bit of the final calculation result that is the same as the sign bits of the two floating-point numbers; when the floating-point operation operator type is the subtraction type and the sign bits of the two floating-point numbers are of different signs, obtaining the mantissa operation type signal indicating no inversion operation and the sign bit of the final calculation result that is the same as the sign bit of the minuend among the two floating-point numbers; when the floating-point operation operator type is the subtraction type and the sign bits of the two floating-point numbers are of the same sign, obtaining the mantissa operation type signal indicating an inversion operation and determining the sign bit of the final calculation result by using the magnitude relationship of the absolute values of the two floating-point numbers; when the floating-point operation operator type is the addition type and the sign bits of the two floating-point numbers are of different signs, obtaining the mantissa operation type signal indicating an inversion operation and determining the sign bit of the final calculation result by using the magnitude relationship of the absolute values of the two floating-point numbers. By adding a sign bit and operator processing module, through simple logical operations, according to the sign bit information of the two operands, the output result of the mantissa segment comparator, information such as the operator type add / sub, etc., the sign bit information of the final result and the complement preprocessing control signal of the smaller operand are output.
[0013] In addition, the exponential adjustment of the larger exponent bit of two floating-point numbers based on the number of leading zeros includes: correcting the larger exponent bit of the two floating-point numbers after subtracting the exponent offset value by using the number of leading zeros, and the formed exponential adjustment result is an exponent value that matches the normalized result. Simplify the leading 0 / 1 counter to a leading 0 leading counter.
[0014] In addition, the overflow judgment and result normalization processing of the normalization result by using the exponent adjustment result and the sign bit of the final calculation result include: converting the normalization result into the form of a subnormal number according to the comparison result between the exponent adjustment result and the exponent offset value; calculating the exponent bit of the floating-point number specification representation according to the exponent adjustment result and the exponent offset value, hiding the integer bit in the normalization result to form the mantissa segment of the floating-point number specification representation, and combining the sign bit of the final calculation result to obtain the normalization processing result. The result exponent adjustment of the large order number is performed through the result of the leading zero count. The conversion calculation operation from the complement code to the original code for the mantissa result after the normalization shift operation is omitted.
[0015] In addition, before aligning the small mantissa by using the absolute value of the difference between the exponent bits of two floating-point numbers, compare the exponent bits of the two floating-point numbers; when the exponent bits of the two floating-point numbers are equal, obtain the magnitude relationship of the absolute values of the two floating-point numbers by comparing the absolute values of the mantissa bits of the two floating-point numbers, so as to simplify the internal calculation logic through the magnitude relationship of the absolute values. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a proportional limitation.
[0017] Figure 1 is a schematic diagram of the definition of a traditional floating-point format provided according to the related art;
[0018] Figure 2 is a schematic diagram of the overall structure of a conventional floating-point adder / subtractor in the related art;
[0019] Figure 3 is a flowchart of a floating-point operation method provided according to an embodiment of the present invention;
[0020] Figure 4 is an architecture diagram of a floating-point adder / subtractor using sign bit processing provided according to an embodiment of the present invention;
[0021] Figure 5 is a hardware operation data flowchart provided according to an embodiment of the present invention;
[0022] Figure 6 is a block diagram of an absolute value comparison and decision module provided according to an embodiment of the present invention;
[0023] Figure 7 is a schematic diagram of restoring the integer bit of the mantissa provided according to an embodiment of the present invention;
[0024] Figure 8 It is a schematic diagram of the hardware structure of the alignment shift and negation module provided according to an embodiment of the present invention;
[0025] Figure 9a It is a schematic diagram of a shift operation of the alignment shift and negation module provided according to an embodiment of the present invention;
[0026] Figure 9b It is a schematic diagram of a shift operation of the alignment shift and negation module provided according to an embodiment of the present invention;
[0027] Figure 10 It is a circuit structure diagram of an unsigned fixed-point full adder provided according to an embodiment of the present invention;
[0028] Figure 11a It is a schematic diagram of a result mantissa normalization shift operation provided according to an embodiment of the present invention;
[0029] Figure 11b It is a schematic diagram of a result mantissa normalization shift operation provided according to an embodiment of the present invention;
[0030] Figure 11c It is a schematic diagram of a result mantissa normalization shift operation provided according to an embodiment of the present invention;
[0031] Figure 12 It is a schematic diagram of the structure of a floating-point arithmetic unit according to another embodiment of the present invention. Specific embodiments
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will elaborate on each embodiment of the present invention with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present invention, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation to the specific implementation of the present invention. Each embodiment can be combined and cross-referenced with each other on the premise of no contradiction.
[0033] The overall structure of a conventional floating-point adder / subtractor, such as Figure 2As shown in the figure, assume that the exponent fields in two floating-point operands A and B are A exponent digits and B exponent digits respectively. Then, it is necessary to compare the values of A exponent digits and B exponent digits through a shift decision maker to determine the magnitude relationship of their exponents, and calculate the absolute difference between the two exponent values as the shift number. The shift number serves as the shift number control signal for shifting the mantissa in the alignment shifter; the first MUX selects the mantissa of the operand with the smaller exponent according to the comparison result "eA < eB" of the shift decision maker and sends it to the alignment shifter for alignment shifting; the mantissa after shift alignment is sent to the first two's complement obtaining module for conversion from the original code to the two's complement; the second MUX sends the mantissa of the operand with the larger exponent to the second two's complement obtaining module to complete the conversion from the original code to the two's complement; the operator control signal add / sub is sent to the two two's complement obtaining modules for calculation of obtaining the two's complement in combination with the sign bit; the mantissas of the operands in the two's complement format output from the two two's complement obtaining modules are sent to a fixed-point adder for signed binary fixed-point addition operation; the leading 0 / 1 counter counts the leading 0 (if the sign bit of the adder calculation result is 0) or leading 1 (if the sign bit of the adder calculation result is 1) of the output of the fixed-point adder, and sends the counting result to the mantissa left shift operation module for normalization shift operation. After its result is converted to the original code format through the mantissa original code restoration module, it is sent to the result normalization module; the exponent adjustment module uses the result of the leading 0 / 1 counter to perform adjustment calculation on the larger exponent of the input operand, and the exponent adjustment result is sent to the result normalization module; the result normalization module performs overflow judgment processing on the adjusted result exponent and the result mantissa in the original code form according to the IEEE-754 specification requirements, and finally outputs the calculation result in the normalized floating-point format. Figure 2 The fixed-point adder in the example shown uses a signed fixed-point adder to support binary addition or subtraction operations on mantissas whose sign bits may be negative. Since the numerical format of the mantissa field in the IEEE-754 format is defined as the original code rather than the two's complement, the mantissa fields in the input and output floating-point numbers of the floating-point adder / subtractor in the example are unsigned original codes. It is necessary to complete the conversion from the original code to the two's complement before performing fixed-point addition and subtraction processing, and complete the restoration calculation from the two's complement to the original code before sending it to overflow judgment and output normalization processing. This series of operations of converting from the original code to the two's complement before performing fixed-point addition and subtraction processing on the unsigned original code and completing the restoration calculation from the two's complement to the original code before sending it to overflow judgment and output normalization processing require two levels and a total of three original code - two's complement conversion (or vice versa) circuits. This increases the logical level of the circuit, thereby restricting the operating frequency of the floating-point addition and subtraction operation circuit and affecting the calculation performance.
[0034] To solve the above technical problems, an embodiment of the present invention relates to a floating-point operation method, which can be applied to electronic devices capable of performing floating-point operations, such as mobile phones, computers, chips, etc., or can also be an aggregate composed of multiple electronic devices, and can also be mounted on, for example, Figure 4 the architecture of a floating-point adder / subtractor using sign bit processing provided by an embodiment of the present invention as shown. In this embodiment, the absolute value of the difference between the exponent bits of two floating-point numbers is used to align the mantissa of the smaller mantissa to the exponent code, obtaining a shifted remainder and a mantissa of the smaller mantissa with aligned exponent code; wherein, the smaller mantissa is the complete mantissa of the smaller floating-point number obtained according to the absolute value size relationship of the two floating-point numbers; according to the sign bits of the two floating-point numbers, the floating-point operation operator type, and the absolute value size relationship of the two floating-point numbers, a mantissa operation type signal and the final calculation result sign bit are obtained through logical operations; the mantissa operation type signal is used to indicate whether to perform an inversion operation on the mantissa of the smaller mantissa with aligned exponent code; if the mantissa operation type signal indicates to perform an inversion operation, an unsigned fixed-point full addition operation is performed using the larger mantissa and the inverted smaller mantissa, otherwise the mantissa of the smaller mantissa with aligned exponent code and the larger mantissa are directly subjected to an unsigned fixed-point full addition operation; wherein, the larger mantissa is the complete mantissa of the larger floating-point number obtained according to the absolute value size relationship of the two floating-point numbers; through the number of leading zeros in the result of the unsigned fixed-point full addition operation and the shifted remainder, a normalization shift operation is performed on the result of the unsigned fixed-point full addition operation to obtain a normalized result; based on the number of leading zeros, an exponential adjustment is performed on the larger exponent bit in the two floating-point numbers, and an overflow judgment and result normalization process are performed on the normalized result using the exponential adjustment result and the final calculation result sign bit. By using the size relationship information, floating-point sign bit information, and floating-point type of floating-point numbers for optimized sign bit processing, the processing logic of the mantissa part operation in the floating-point addition and subtraction operation is simplified, an unsigned fixed-point binary adder in the original code form is directly used, most of the operation hardware for original code-complement code conversion is omitted, the design of the leading 0 / 1 counter is simplified, and the shifted remainder for alignment is directly used to participate in the normalization shift of the result mantissa, ensuring the calculation accuracy without increasing the adder bit width. These circuit optimizations significantly reduce the level of the overall logic, achieving the effects of improving the working frequency of the floating-point addition and subtraction circuit, enhancing the calculation energy efficiency, reducing the scale of the logic circuit, and increasing the computing power density. The implementation details of the floating-point operation method in this embodiment will be specifically described below. The following content is only provided for facilitating understanding of the implementation details and is not necessary for implementing this solution.
[0035] Overall, the floating-point operation method process is as Figure 3As shown, in step 301, the smaller mantissa is aligned with the exponent by using the absolute value of the difference between the exponent bits of two floating-point numbers, obtaining a shifted remainder and a mantissa of the smaller floating-point number with aligned exponent; where the smaller mantissa is the complete mantissa of the smaller floating-point number obtained according to the magnitude relationship of the absolute values of the two floating-point numbers; in step 302, according to the sign bits of the two floating-point numbers, the type of floating-point operation operator, and the magnitude relationship of the absolute values of the two floating-point numbers, a mantissa operation type signal and a final calculation result sign bit are obtained through logical operations; the mantissa operation type signal is used to indicate whether to perform a negation operation on the mantissa of the smaller floating-point number with aligned exponent; in step 303, if the mantissa operation type signal indicates to perform a negation operation, an unsigned fixed-point full addition operation is performed using the larger mantissa and the negated mantissa of the smaller floating-point number, otherwise, the mantissa of the smaller floating-point number with aligned exponent and the larger mantissa are directly subjected to an unsigned fixed-point full addition operation; where the larger mantissa is the complete mantissa of the larger floating-point number obtained according to the magnitude relationship of the absolute values of the two floating-point numbers; in step 304, a normalization shift operation is performed on the result of the unsigned fixed-point full addition operation according to the number of leading zeros in the result of the unsigned fixed-point full addition operation and the shifted remainder, obtaining a normalized result; in step 305, the exponent bit of the larger of the two floating-point numbers is adjusted based on the number of leading zeros, and an overflow judgment and result normalization process are performed on the normalized result using the exponent adjustment result and the final calculation result sign bit.
[0036] To better understand the execution principle of the floating-point number operation method of the present invention, this embodiment provides an optimized architecture of a floating-point adder / subtractor as a carrier of the floating-point number operation method. The optimized architecture of the floating-point adder / subtractor is as Figure 4 shown. The adder / subtractor in this example can better complete the addition or subtraction operation of two floating-point numbers A and B compared to Figure 2 the example shown. It should be noted that Figure 4 is a relatively complete and detailed example of a specific optimized architecture provided for easy understanding. The floating-point number operation method involved in the present invention can act on this structure, but it does not mean that the floating-point number operation method can only act on this structure. Figure 4 The structure in includes: an absolute value comparison and decision maker 1, a mantissa alignment shift and negation module 2, a sign bit processing module 3, an unsigned fixed-point full adder 4, a leading 0 counter 5, a result mantissa normalization shift module 6, an exponent adjustment module 7, and a result normalization processing module 8. These modules have corresponding functions and corresponding expansion methods, which will be discussed in detail in turn below.
[0037] As Figure 4The adder / subtractor shown can perform the addition or subtraction operations on two floating-point numbers A and B: Y = A + B or Y = A - B, where the input operands A, B, and the output result Y are all floating-point numbers that meet the definition of the IEEE-754 standard. SignA and SignB are the sign bits of A and B respectively, eA and eB are the exponent fields of A and B respectively, and MA and MB are the mantissa fields of A and B respectively. The data operation process corresponding to this architecture is as Figure 5 shown. For the input operands A and B, the operands A and B will be split and then sent to the absolute value comparison and decision module. The absolute value comparison and decision module receives the exponent bits eA / eB and mantissa bits MA / MB data of the two input operands. After internal processing, it outputs the complete mantissas MantA and MantB of the two operands with the hidden integer bits restored, as well as the comparison result ALTB representing the magnitude relationship of the absolute values of the two operands A and B. The internal processing includes: comparing the magnitudes of the exponent bits, comparing the absolute values of the mantissas when eA = eB, restoring the complete mantissas MantA / MantB, calculating the larger exponent value, and calculating the absolute value of the exponent difference. The comparison result of the magnitude relationship is used together with the operator for sign bit processing, and the sign bit of the final result and the indication that the smaller number needs to be inverted are output. The smaller mantissa output by the absolute value comparison and decision module is used together with the sign bit of the final result output by the sign bit processing and the indication that the smaller number needs to be inverted for the alignment shifter operation, so as to output the mantissa of the smaller operand and the shift remainder after the exponents are aligned through the alignment shifter operation, and decide whether to invert according to the sign bit processing result. The alignment shifter operation will obtain the shift remainder and the mantissa after alignment. The mantissa after alignment is used for unsigned addition of the mantissas with the larger mantissa output by the absolute value comparison and decision module, and then the leading 0 count of the result is performed, and the result normalization shift is performed with the shift remainder. The leading 0 count of the result can also be used for result exponent adjustment with the larger exponent obtained by the absolute value comparison and decision module. Finally, combining the results of sign bit processing, result normalization shift, and result exponent adjustment, the overflow judgment and result normalization processing are jointly completed.
[0038] By interpreting Figure 4 it can be seen that Figure 3The small mantissa in step 301 is the complete mantissa of the smaller floating-point number obtained according to the magnitude relationship of the absolute values of two floating-point numbers. The way to obtain the small mantissa can be: comparing the exponent bits of the two floating-point numbers to get the absolute value of the difference in exponent bits; then calculating the magnitude relationship of the absolute values of the two floating-point numbers. In an example, before aligning the small mantissa with the exponent code using the absolute value of the difference in the exponent bits of the two floating-point numbers, compare the exponent bits of the two floating-point numbers; when the exponent bits of the two floating-point numbers are equal, obtain the magnitude relationship of the absolute values of the two floating-point numbers by comparing the absolute values of the mantissa bits of the two floating-point numbers. Because when the large exponent and the small exponent are equal, the absolute values of the mantissa bits of the two floating-point numbers can be compared through the mantissa segment comparison circuit to obtain the magnitude relationship of the absolute values of the two floating-point numbers.
[0039] Such as Figure 4 The absolute value comparison and decision maker 1 shown is mainly used to receive the exponent bits eA / eB and mantissa bits MA / MB data of two input operands, and after internal processing, output the complete mantissas MantA and MantB of the two operands with the hidden integer bits restored, output the comparison result ALTB representing the magnitude relationship of the absolute values of the two floating-point numbers A and B, output the larger exponent bit eMAX = Max(eA - k, eB - k) of the two floating-point numbers after subtracting the offset value k. At the same time, this module also outputs the unsigned number eSHIFT representing the absolute value of the difference in the two exponent bits |eA - eB|. For example, the internal structure of the absolute value comparison and decision maker module is as Figure 6 shown. The absolute value comparison decision maker includes: sub-module 101 - shift decision maker, sub-module 102 - mantissa restorer, sub-module 103 - mantissa comparator, and a multiplexer MUX for calculating and processing the exponent bits of the two input operands, restoring and comparing the complete mantissas of the two input operands, outputting the larger exponent value eMAX after removing the exponent offset, the control signal eSHIFT for the alignment shifter, the gating control signal ALTB for the mantissa data path, and the complete mantissas MantA and MantB of the two operands for addition calculation. In an example, the complete mantissas of the two floating-point numbers are obtained by respectively performing the processing of restoring the hidden integer bits on the mantissa bits of the two floating-point numbers.
[0040] Figure 6The sub-module 101 - shift decision maker in it is used to complete the comparison and calculation operations of the exponent bits eA and eB of two operands, including: comparing whether eA is equal to eB, if equal, output the signal eq = 1, otherwise output eq = 0; comparing the magnitude relationship between eA and eB, if eA < eB, output the signal eaLTeb = 1, otherwise output eaLTeb = 0; calculating and outputting the absolute value eSHIFT = |eA - eB| of the difference between eA and eB; outputting the larger value eMAX = MAX(eA - k, eB - k) after removing the exponent offset k from eA and eB, etc.
[0041] Figure 6 The sub-module 102 - mantissa restoration module in it is used to complement the hidden mantissa integer bit information according to the eA and eB values of the exponent segment. When the exponent bits are all 0, the corresponding restored complete mantissa integer bit is 0, otherwise the integer bit is 1. Taking operand A as an example, the operation schematic diagram of this sub-module is as shown in the appendix Figure 7 as follows;
[0042] Figure 6 The sub-module 103 - mantissa comparator in it is used to directly compare the magnitudes of the mantissa segment values MA and MB. If MA is less than MB, output 1, otherwise output 0. This comparison result will only be selected and output from port ALTB when the control signal eq = 1. When eq = 0, ALTB directly selects and outputs the exponent comparison result eaLTeb of sub-module 101, and the comparison result of sub-module 103 will be ignored. This path selection function is implemented by a two-way selector.
[0043] such as Figure 4The absolute value comparison in it and the output signal ALTB of the decision maker module 1 are used to control the first MUX to send the complete mantissa LIT = Min(MantA, MantB) of the smaller operand to the alignment shift and negation module 2 for alignment shift operation; control the second MUX to send the complete mantissa BIG = Max(MantA, MantB) of the larger number to the unsigned fixed-point full adder 4 for addition operation. For example: when ALTB = 1, it means that the absolute value of operand A (including the exponent bit and mantissa segment data) is less than the absolute value of operand B. This signal controls the first MUX to send the complete mantissa MantA of the smaller absolute value number A to the alignment shifter, and controls the second MUX to send the complete mantissa MantB of the larger absolute value number B to the fixed-point negation adder; when ALTB = 0, the data selection path is exactly the opposite; the ALTB signal is also directly used by the sign bit processing module 3 to participate in the sign bit calculation to simplify the internal calculation logic; the output signal eSHIFT of the absolute value comparison and decision maker 1 is sent to the alignment shift and negation module 2 to control the operation of the alignment shifter, and its value represents the number of shift bits for the right shift operation of the alignment shifter; the eMAX output by the absolute value comparison and decision maker 1 is sent to the exponent adjustment module 7 for finally calculating the exponent code of the output result.
[0044] In step 302, according to the sign bits of the two floating-point numbers, the floating-point operation operator type, and the magnitude relationship of the absolute values of the two floating-point numbers, the mantissa operation type signal and the sign bit of the final calculation result are obtained through logical operations; the mantissa operation type signal is used to indicate whether to perform a negation operation on the small mantissa after aligning the exponent codes. The implementation method is as follows:
[0045] Figure 4 The sign bit processing module 3 in it outputs the sign bit sign and the mantissa operation type signal sub of the final calculation result through logical operations according to the sign bits SA and SB of the input operands A and B, the operation operator (add_sub), and the ALTB signal output by the absolute value comparison and decision maker 1.
[0046] In one example, step 302 further includes the following operations: when the floating-point operation operator type is the addition type and the sign bits of the two floating-point numbers have the same sign, a mantissa operation type control signal indicating no negation operation is obtained, and a final calculation result sign bit that is the same as the sign bits of the two floating-point numbers; when the floating-point operation operator type is the subtraction type and the sign bits of the two floating-point numbers have different signs, a mantissa operation type control signal indicating no negation operation is obtained, and a final calculation result sign bit that is the same as the sign bit of the minuend among the two floating-point numbers; when the floating-point operation operator type is the subtraction type and the sign bits of the two floating-point numbers have the same sign, a mantissa operation type control signal indicating a negation operation is obtained, and the final calculation result sign bit is determined using the magnitude relationship of the absolute values of the two floating-point numbers; when the floating-point operation operator type is the addition type and the sign bits of the two floating-point numbers have different signs, a mantissa operation type control signal indicating a negation operation is obtained, and the final calculation result sign bit is determined using the magnitude relationship of the absolute values of the two floating-point numbers. Specifically: sub = 1 represents subtracting the mantissa of the smaller operand from the mantissa of the larger operand; sub = 0 represents adding the mantissa of the smaller operand to the mantissa of the larger operand. The input-output logic operation truth table of the sign bit processing module 3 is as follows.
[0047]
[0048]
[0049] In one example, the output signal sub of the sign bit processing module 3 is sent to the alignment shift and negation module 2 to control the negation preprocessing for two's complement conversion of the data after the alignment right shift operation before output; another output signal sign is the sign bit of the entire addition and subtraction operation result. sign = 0 represents that the final result is positive, and sign = 1 represents that the final output result is negative. This signal is sent to the result normalization processing module 8 to be combined with the exponent bit and mantissa bit of the final calculation result into a floating-point format for output.
[0050] In step 303, if the mantissa operation type control signal indicates a negation operation, an unsigned fixed-point full addition operation is performed using the large mantissa and the small mantissa after the negation operation, otherwise, the small mantissa and the large mantissa with aligned exponents are directly subjected to an unsigned fixed-point full addition operation; where the large mantissa is the complete mantissa of the larger floating-point number obtained based on the magnitude relationship of the absolute values of the two floating-point numbers. An example implementation method is as follows:
[0051] Such as Figure 4 in the alignment shift and negation module 2, according to the value of eSHIFT, a right shift operation of the input mantissa LIT (i.e., the small mantissa) is completed, and then according to the status of the sub signal (i.e., the mantissa operation type control signal) output by the sign bit processing module 3, a negation preprocessing for two's complement calculation of the output result is performed.Figure 8 Disclosed is a hardware structure of the exponent alignment shift and inversion module 2. This example takes the exponent alignment shift range of 0 to 15 as an example. In this example, a multiplexer is used to implement the right shift operation. Each bit of the output mantissa uses an independent multiplexer to select one bit from the possible bits of the input mantissa for output, and the selection control signal is the input shift bit control signal eSHIFT[3:0]; each multiplexer is a 16-to-1 selector, and the input of the selector is 16-bit data, which comes from 16 higher-order bits on the left starting from this position in the input mantissa (including this position). For example, the selector input of the output mantissa LITTLE[0] is LIT[15:0], the selector input of the output mantissa LITTLE[1] is LIT[16:1], and so on. The selector input of the output mantissa LITTLE[i] is LIT[i + 15:i].
[0052] In one example, as a preferred design for improving calculation accuracy, such as Figure 8 The exponent alignment shift and inversion module shown can output the shift remainder TAIL generated by the right shift operation, which is used for precision compensation during the mantissa normalization shift operation of the addition result. The data width of the TAIL output is determined by the bit width of eSHIFT. Taking Figure 8 as an example, if the bit width w of eSHIFT[3:0] is 4, then the data bit width of TAIL is 2 w - 1 = 15, that is, TAIL[14:0]. The generation of TAIL uses the same multiplexing circuit to obtain it by right shifting the input mantissa LIT. Its selector also uses the eSHIFT[3:0] input for control, but the input data of its selector comes from 15 lower-order bits on the right starting from this position in the input mantissa, excluding this position. For example, the selector input of TAIL
[14] is LIT[14:-1], the selector input of TAIL
[13] is LIT[13:-2], and so on, until the lowest bit TAIL[0] uses LIT[0:-15] as the selector input.
[0053] In one example, for the exponent alignment shift and inversion module 2 shown as Figure 8 When the LIT bit selected by the selector input exceeds the actual data bit width boundary of the input mantissa LIT (including the high-order limit on the left and the low-order limit on the right), 0 is directly used to access the selector input terminal. For example, if the bit width of the smaller mantissa LIT of the actual input module is 24 bits, that is, LIT[23:0], all bits other than LIT[23:0] are directly used 0 to access the selector. That is, when x < 0 or x > 23, LIT[x] is directly replaced by 0.
[0054] In one example, step 303 may further include the following operations: when the exponent bits of two floating-point numbers are equal, perform an increment operation on the mantissa of the negated operation, and perform an unsigned fixed-point full addition operation on the large mantissa and the incremented mantissa of the small mantissa; when the exponent bits of two floating-point numbers are not equal, perform an unsigned fixed-point full addition operation on the large mantissa and the mantissa of the small mantissa after the negated operation. An example implementation method is as follows:
[0055] In one example, for Figure 8 the alignment shift and negation module shown, to simplify the two's complement calculation operation of the output data LITTLE before fixed-point addition, this module provides selectable data reverse output and the carry-in flag output CIN required in the two's complement calculation. The input signal sub comes from the sign bit operator processing module and is used to control whether to perform the two's complement operation on the output data. As shown in the appendix Figure 8 shown, the bitwise negation operation is implemented using an exclusive OR gate. When sub = 0, it means adding the original number, and the exclusive OR gate directly outputs the output of the selector; when sub = 1, it means subtracting the original number, and the exclusive OR gate outputs the negated result of the output of the selector. To simplify the calculation logic, the increment operation of the two's complement processing in this module only takes effect when eSHIFT = 0, that is, when the exponents of the two input operands are equal and no alignment shift of LIT is required, that is: when eSHIFT = 0 (when the exponent bits of two floating-point numbers are equal) and sub = 1 (subtracting the original number), CIN outputs 1, otherwise it outputs 0; the deviation in two's complement calculation in other cases is directly absorbed into the error of the remainder compensation, which also means that when the exponent bits of two floating-point numbers are equal, perform an increment operation on the mantissa of the small mantissa after the negated operation.
[0056] In one example, for Figure 8 the alignment shift and negation module shown, to keep the output remainder TAIL (i.e., the shifted remainder) and the mantissa LITTLE after alignment (i.e., the small mantissa with aligned exponents) having the same two's complement attribute, the output of the remainder TAIL adopts the same negation output control logic as LITTLE.
[0057] In one example, for Figure 8 the alignment shift and negation module shown, as one of the optimization options, when the overall floating-point calculation accuracy requirement does not require the remainder to be compensated for accuracy, as an optional item, the remainder TAIL is not output, and its corresponding generation circuit can be optimized to save circuit resources and area power consumption.
[0058] In one example, the data output by the alignment shift and negation module 2 is divided into high and low two segments of data, Figures 9a to 9bDiscloses the shift data correspondence between the input mantissa LIT and the high significant bit data segment LITTLE and the low significant bit data segment TAIL of the output; after alignment and shift, the mantissa LITTLE[n-1:0] of the output maintains the same format as the input mantissa LIT[n-1:0], and its data bit width is the same as that of the input mantissa; while the bit width of the output remainder TAIL[m-1:0] is determined by the bit width k of eSHIFT[k:0], that is, m = 2 k -1, and the bits exceeding the input bit width are filled with 0; for example: when the value of eSHIFT is 0, as Figure 9a shown, there is no shift and it is determined whether to take the bitwise inversion according to the state of the sub signal. When the bit width of eSHIFT is 4 bits, as Figure 9b shown, the corresponding maximum right shift amplitude is 0 to 15, the data bit width of TAIL is 15, that is, m = 15, and it is determined whether to take the bitwise inversion according to the state of the sub signal.
[0059] In an example, the high significant bit segment LITTLE output by the alignment shift and inversion module 2 is sent to the unsigned fixed-point full adder 4 for unsigned fixed-point addition calculation; as an optional item to improve the calculation accuracy, the low significant bit data segment TAIL output by the alignment shift and inversion module 2 can skip the adder and be directly sent to the leading zero counter 5 to participate in the leading zero counting operation and be sent to the result mantissa normalization shift module 6 to participate in the left shift operation of the calculation result mantissa.
[0060] In an example, the single-bit output signal CIN output by the alignment shift and inversion module 2 is connected to the unsigned fixed-point full adder 4 as the complement code plus 1 flag bit, and is used to correct the complement code deviation of the input LITTLE participating in the fixed-point addition calculation when the exponent bits of the two operands are exactly equal (eA = eB) and sub = 1 (subtraction operation).
[0061] Figure 4 The unsigned fixed-point full adder 4 in realizes the function of adding two complete mantissas BIG and LITTLE after alignment of the exponents, and the output result SUM[n:0] = BIG[n-1:0] + LITTLE[n-1:0], where the highest bit SUM[n] is the carry output bit; when the exponent bits of the two floating-point numbers are equal, perform a plus-one operation on the inverted small mantissa, and perform an unsigned fixed-point full addition operation on the large mantissa and the small mantissa after the plus-one operation; when the exponent bits of the two floating-point numbers are not equal, perform an unsigned fixed-point full addition operation on the large mantissa and the inverted small mantissa. In an example, the adder is implemented by cascaded fixed-point full adders, and its internal logic architecture is as attached Figure 10As shown, the CIN signal from the output of the alignment shift and negation module 2 is connected to the lowest carry input of the full adder, supporting the addition of 1 after bitwise negation for obtaining the two's complement of the mantissa of the smaller operand after alignment shift when the exponents of the two operands are equal; when the exponents of the two operands are not equal, the addition of 1 during the two's complement calculation is absorbed as an error into the remainder TAIL output from the alignment shift.
[0062] Figure 4 The leading zero counter 5 in is used to count the number of leading zeros in the output result SUM of the unsigned fixed-point full adder 4, that is, to count how many '0's there are before the first '1' from the high bit to the low bit in SUM[n:0]. For example, taking the 24-bit wide fixed-point full adder used in the architecture as an example, the output result SUM[24:0] of the unsigned fixed-point full adder 4 is 0_00011100_11000010_10001101, and its highest bit SUM
[24] is the carry output bit, then the leading zero count result LZCNT = 4, indicating that there are 4 zeros before the first 1 in SUM[24:0].
[0063] In one example, while the leading zero count result LZCNT is used by the result mantissa normalization shift module 6 to control the left shift operation of the result mantissa, it is also sent to the exponent adjustment module 7 for correcting the order (exponent) of the final output result. That is, before overflow judgment and result normalization processing, the result exponent is adjusted for large exponents through the result of leading zero counting.
[0064] In one example, step 304 performs a normalization shift operation on the result of the unsigned fixed-point full addition operation through the number of leading zeros and the shift remainder of the result of the unsigned fixed-point full addition operation to obtain a normalized result. The implementation method is as follows:
[0065] As Figure 4 The result mantissa normalization shift module 6 in Figure 4 is used to perform a normalization shift operation on the calculation result of the unsigned fixed-point full adder 4 in Figure 4 to ensure that the most significant bit in the mantissa result MantiY output after its shift is fixed as 1 and is the only integer bit of the mantissa, achieving the purpose of normalization.
[0066] In one example, the output result MantiY is fixedly intercepted starting from the highest bit after the shift, and the intercepted bit width meets the requirements defined by the IEEE-754 standard. The remaining data bits after interception can be used for rounding processing, which can be implemented according to different rounding algorithms.
[0067] In one example, step 304 may be to perform a left shift operation on the result of the unsigned fixed-point full addition operation by using the number of leading zeros; use the shift remainder to perform precision compensation on the result of the unsigned fixed-point full addition operation after the left shift operation to obtain a normalized result. In Figure 4 it is shown as:
[0068] In one example, the shift bit number of the normalization shift operation is determined by the output signal LZCNT of the leading 0 counter 5, that is, MantiY = SUM * 2 LZCNT , when LZCNT is 0, it means that no shift operation is required; when LZCNT is i, it means that the result mantissa SUM is left-shifted by i bits.
[0069] In one example, as an optimization option, the remainder TAIL output from the alignment shift and negation module 2 can be used as the low significant bits to participate in the leading 0 counting of the module together with SUM, that is, the leading 0 counting is performed on {SUM[n:0], TAIL[m - 1, 0]} as a whole, and it is sent to the result mantissa normalization shift module 6 to participate in the result mantissa normalization shift operation together with SUM, that is, Y = {SUM, TAIL} * 2 LZCNT , to improve the final calculation accuracy, the left shift operation of the result mantissa normalization shift module 6 is for example as Figure 11a 、 Figure 11b 、 Figure 11c shown, where Figure 11a represents the case where SUM[n] = 1 and LZCNT = 0, without shifting, Figure 11b represents the case where SUM[n, n - 1] = 01 and LZCNT = 1, with a left shift of 1 bit, Figure 11c represents the case where SUM[n:n - 4] = 0001 and LZCNT = 3, with a left shift of 3 bits.
[0070] In one example, in step 305, the exponential adjustment of the larger exponent bit among the two floating-point numbers based on the number of leading zeros may be: correcting the larger exponent bit among the two floating-point numbers after subtracting the exponent offset value by the number of leading zeros, and the formed exponential adjustment result is the exponent value that matches the normalized result. In Figure 4 it is shown as:
[0071] Such as Figure 4 the exponential adjustment module 7 in is used to adjust and correct eMAX output from the absolute value comparison and decision maker 1 to form an exponent value that matches the result mantissa MantY output after the result mantissa normalization shift operation, and its output result eOUT = eMAX - LZCNT + 1.
[0072] In one example, the process of determining overflow and normalizing the normalization result by using the exponent adjustment result and the sign bit of the final calculation result in step 305 may be as follows: according to the comparison result between the exponent adjustment result and the exponent offset value, convert the normalization result into the form of a subnormal number; calculate the exponent bit of the floating-point number standard representation based on the exponent adjustment result and the exponent offset value, hide the integer bit in the normalization result to form the mantissa segment of the floating-point number standard representation, and combine the sign bit of the final calculation result to obtain the normalized processing result. An example implementation method is as follows:
[0073] Such as Figure 4 The result normalization processing module 8 in it is used to complete the overflow judgment and normalization processing of the calculation result, including judging that when eOUT <= -k (k is the order offset defined by the IEEE-754 standard), converting the mantissa into the form of a subnormal number, calculating the exponent segment eY = eOUT + k of the floating-point number standard representation, hiding the integer bit in the mantissa to form the mantissa segment of the floating-point number standard representation, and finally outputting the calculation result floating-point number Y of the overall floating-point adder / subtractor from this module to ensure that the processed result meets the requirements of the IEEE-754 standard.
[0074] In this embodiment, through the optimized sign bit processing using the size relationship information, sign bit information, and floating-point type of the floating-point number, the processing logic of the mantissa part operation in the floating-point addition and subtraction operation is simplified. The unsigned fixed-point binary adder in the original code mode is directly used, most of the operation hardware for the original code - complement code conversion is omitted, the leading 0 / 1 counter design is simplified, and the mantissa normalization shift of the result is directly participated by the alignment shift remainder. Without increasing the bit width of the adder, the calculation accuracy is guaranteed. These circuit optimizations significantly reduce the level of the overall logic, achieving the effects of improving the working frequency of the floating-point addition and subtraction circuit, enhancing the calculation energy efficiency, reducing the scale of the logic circuit, and increasing the computing power density.
[0075] The step division of the above method is only for clear description. When implementing, it can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, are all within the protection scope of this application.
[0076] Another embodiment of the present invention relates to a floating-point operation device, such as Figure 12As shown in the figure, it includes: a mantissa alignment module 1201, which is used to align the mantissa of the smaller mantissa by using the absolute value of the difference between the exponent bits of two floating-point numbers to obtain the shifted remainder and the mantissa of the smaller mantissa with aligned exponent; among them, the smaller mantissa is the complete mantissa of the smaller floating-point number obtained according to the absolute value size relationship of the two floating-point numbers; a sign bit processing module 1202, which is used to obtain the mantissa operation type signal and the sign bit of the final calculation result through logical operations according to the sign bits of the two floating-point numbers, the floating-point operation operator type, and the absolute value size relationship of the two floating-point numbers; the mantissa operation type signal is used to indicate whether to perform an inversion operation on the mantissa of the smaller mantissa with aligned exponent; a full adder operation module 1203, which is used to perform an unsigned fixed-point full adder operation on the larger mantissa and the inverted smaller mantissa if the mantissa operation type signal indicates an inversion operation, otherwise directly perform an unsigned fixed-point full adder operation on the mantissa of the smaller mantissa with aligned exponent and the larger mantissa; among them, the larger mantissa is the complete mantissa of the larger floating-point number obtained according to the absolute value size relationship of the two floating-point numbers; a normalization module 1204, which is used to perform a normalization shift operation on the result of the unsigned fixed-point full adder operation through the number of leading zeros and the shifted remainder of the result of the unsigned fixed-point full adder operation to obtain the normalized result; a normalization processing module 1205, which is used to adjust the exponent of the larger exponent bit among the two floating-point numbers based on the number of leading zeros, and perform an overflow judgment and result normalization processing on the normalized result by using the exponent adjustment result and the sign bit of the final calculation result.
[0077] In an example, performing an unsigned fixed-point full adder operation on the larger mantissa and the inverted smaller mantissa includes: when the exponent bits of the two floating-point numbers are equal, performing an increment operation on the inverted smaller mantissa, and performing an unsigned fixed-point full adder operation on the larger mantissa and the incremented smaller mantissa; when the exponent bits of the two floating-point numbers are not equal, performing an unsigned fixed-point full adder operation on the larger mantissa and the inverted smaller mantissa.
[0078] In an example, performing a normalization shift operation on the result of the unsigned fixed-point full adder operation through the number of leading zeros and the shifted remainder of the result of the unsigned fixed-point full adder operation to obtain the normalized result includes: performing a left shift operation on the result of the unsigned fixed-point full adder operation by using the number of leading zeros; performing precision compensation on the result of the unsigned fixed-point full adder operation after the left shift operation by using the shifted remainder to obtain the normalized result.
[0079] In one example, according to the sign bits of two floating-point numbers, the type of floating-point operation operator, and the magnitude relationship of the absolute values of the two floating-point numbers, a mantissa operation type signal and the sign bit of the final calculation result are obtained through logical operations, including: when the type of floating-point operation operator is the addition type and the sign bits of the two floating-point numbers are the same, a mantissa operation type signal indicating no negation operation is obtained, and the sign bit of the final calculation result is the same as the sign bits of the two floating-point numbers; when the type of floating-point operation operator is the subtraction type and the sign bits of the two floating-point numbers are different, a mantissa operation type signal indicating no negation operation is obtained, and the sign bit of the final calculation result is the same as the sign bit of the minuend among the two floating-point numbers; when the type of floating-point operation operator is the subtraction type and the sign bits of the two floating-point numbers are the same, a mantissa operation type signal indicating a negation operation is obtained, and the sign bit of the final calculation result is determined using the magnitude relationship of the absolute values of the two floating-point numbers; when the type of floating-point operation operator is the addition type and the sign bits of the two floating-point numbers are different, a mantissa operation type signal indicating a negation operation is obtained, and the sign bit of the final calculation result is determined using the magnitude relationship of the absolute values of the two floating-point numbers.
[0080] In one example, the exponent of the larger exponent bit among two floating-point numbers is adjusted based on the number of leading zeros, including: correcting the larger exponent bit among the two floating-point numbers after subtracting the exponent offset value by the number of leading zeros, and the resulting exponent adjustment result is an exponent value that matches the normalized result.
[0081] In one example, overflow judgment and result normalization processing are performed on the normalized result using the exponent adjustment result and the sign bit of the final calculation result, including: converting the normalized result into the form of a subnormal number according to the comparison result between the exponent adjustment result and the exponent offset value; calculating the exponent bit of the floating-point number's canonical representation based on the exponent adjustment result and the exponent offset value, hiding the integer bit in the normalized result to form the mantissa segment of the floating-point number's canonical representation, and combining the sign bit of the final calculation result to obtain the normalized processing result.
[0082] In one example, the device further includes: an absolute value magnitude relationship fast calculation module, configured to compare the exponent bits of two floating-point numbers before aligning the mantissa with the smaller exponent by using the absolute difference between the exponent bits of the two floating-point numbers; when the exponent bits of the two floating-point numbers are equal, obtain the magnitude relationship of the absolute values of the two floating-point numbers by comparing the absolute values of the mantissa bits of the two floating-point numbers.
[0083] In one example, the complete mantissa of each of the two floating-point numbers is obtained by separately performing the processing of restoring and hiding the integer bit on the mantissa bits of each of the two floating-point numbers.
[0084] It is not difficult to find that this embodiment is a device embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details mentioned in the above method embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.
[0085] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present invention, units that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0086] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made to them in form and details without departing from the spirit and scope of the present invention.
Claims
1. A floating-point operation method, characterized in that, Including: Aligning the mantissa by using the absolute difference of the exponent bits of two floating-point numbers to obtain a shifted remainder and a mantissa with aligned exponent; wherein, the mantissa is the complete mantissa of the smaller floating-point number obtained according to the magnitude relationship of the absolute values of the two floating-point numbers; Obtaining a mantissa operation type signal and a final calculation result sign bit through logical operations according to the sign bits of the two floating-point numbers, the floating-point operation operator type, and the magnitude relationship of the absolute values of the two floating-point numbers; the mantissa operation type signal is used to indicate whether to perform a negation operation on the mantissa with aligned exponent; If the mantissa operation type signal indicates to perform a negation operation, performing an unsigned fixed-point full addition operation using the large mantissa and the negated mantissa, otherwise directly performing an unsigned fixed-point full addition operation on the mantissa with aligned exponent and the large mantissa; wherein, the large mantissa is the complete mantissa of the larger floating-point number obtained according to the magnitude relationship of the absolute values of the two floating-point numbers; Performing a normalization shift operation on the result of the unsigned fixed-point full addition operation according to the number of leading zeros in the result of the unsigned fixed-point full addition operation and the shifted remainder to obtain a normalized result; Based on the number of leading zeros, performing an exponent adjustment on the larger exponent bit of the two floating-point numbers, and performing an overflow judgment and result normalization process on the normalized result by using the exponent adjustment result and the final calculation result sign bit.
2. The floating-point operation method according to claim 1, wherein The performing an unsigned fixed-point full addition operation using the large mantissa and the negated mantissa includes: When the exponent bits of the two floating-point numbers are equal, performing an increment operation on the negated mantissa, and performing an unsigned fixed-point full addition operation on the large mantissa and the incremented mantissa; When the exponent bits of the two floating-point numbers are not equal, performing an unsigned fixed-point full addition operation on the large mantissa and the negated mantissa.
3. The floating-point operation method according to claim 1, wherein The performing a normalization shift operation on the result of the unsigned fixed-point full addition operation according to the number of leading zeros in the result of the unsigned fixed-point full addition operation and the shifted remainder to obtain a normalized result includes: Performing a left shift operation on the result of the unsigned fixed-point full addition operation by using the number of leading zeros; Performing a precision compensation on the result of the unsigned fixed-point full addition operation after the left shift operation by using the shifted remainder to obtain a normalized result.
4. The floating-point operation method according to claim 1, characterized in that, The obtaining the mantissa operation type signal and the final calculation result sign bit through logical operations according to the sign bits of the two floating-point numbers, the floating-point operation operator type, and the magnitude relationship of the absolute values of the two floating-point numbers includes: When the floating-point operation operator type is an addition type and the sign bits of the two floating-point numbers are the same, obtaining the mantissa operation type signal indicating not to perform a negation operation, and the final calculation result sign bit the same as the sign bits of the two floating-point numbers; When the floating-point operation operator type is a subtraction type and the sign bits of the two floating-point numbers are different, obtaining the mantissa operation type signal indicating not to perform a negation operation, and the final calculation result sign bit the same as the sign bit of the minuend in the two floating-point numbers; When the floating-point operation operator type is the subtraction type and the sign bits of the two floating-point numbers are of the same sign, a mantissa operation type control signal indicating a negation operation is obtained, and the sign bit of the final calculation result is determined by using the magnitude relationship of the absolute values of the two floating-point numbers; When the floating-point operation operator type is the addition type and the sign bits of the two floating-point numbers are of different signs, a mantissa operation type control signal indicating a negation operation is obtained, and the sign bit of the final calculation result is determined by using the magnitude relationship of the absolute values of the two floating-point numbers.
5. The floating-point operation method according to claim 1, wherein The exponent adjustment of the larger exponent bit of the two floating-point numbers based on the number of leading zeros includes: The larger exponent bit of the two floating-point numbers after subtracting the exponent offset value is corrected by the number of leading zeros, and the formed exponent adjustment result is an exponent value matching the normalized result.
6. The floating-point operation method according to claim 5, wherein The overflow judgment and result normalization processing of the normalized result by using the exponent adjustment result and the sign bit of the final calculation result include: According to the comparison result of the exponent adjustment result and the exponent offset value, the normalized result is converted into the form of a subnormal number; The exponent bit of the floating-point number specification representation is calculated according to the exponent adjustment result and the exponent offset value, the integer bit in the normalized result is hidden to form the mantissa segment of the floating-point number specification representation, and combined with the sign bit of the final calculation result, the normalization processing result is obtained.
7. The floating-point number operation method according to claim 1, wherein, The method further includes: Before aligning the small mantissa with the exponent code by using the absolute value of the difference between the exponent bits of the two floating-point numbers, compare the exponent bits of the two floating-point numbers; When the exponent bits of the two floating-point numbers are equal, the magnitude relationship of the two floating-point numbers is obtained by comparing the absolute values of the mantissa bits of the two floating-point numbers.
8. The floating-point operation method according to any one of claims 1-7, characterized in that The complete mantissa of each of the two floating-point numbers is obtained by respectively performing the process of restoring and hiding the integer bit on the mantissa bits of the two floating-point numbers.
9. A floating-point arithmetic unit, characterized in that, It includes: A mantissa alignment module, configured to align the small mantissa with the exponent code by using the absolute value of the difference between the exponent bits of the two floating-point numbers to obtain a shifted remainder and the small mantissa with the exponent code aligned; wherein, the small mantissa is the complete mantissa of the smaller floating-point number obtained according to the magnitude relationship of the two floating-point numbers; A sign bit processing module, configured to obtain a mantissa operation type control signal and a sign bit of the final calculation result through logical operations according to the sign bits of the two floating-point numbers, the floating-point operation operator type, and the magnitude relationship of the absolute values of the two floating-point numbers; the mantissa operation type control signal is used to indicate whether to perform a negation operation on the small mantissa with the exponent code aligned; A full addition operation module, configured to, if the mantissa operation type control signal indicates a negation operation, perform an unsigned fixed-point full addition operation by using the large mantissa and the small mantissa after the negation operation, otherwise directly perform an unsigned fixed-point full addition operation on the small mantissa with the exponent code aligned and the large mantissa; wherein, the large mantissa is the complete mantissa of the larger floating-point number obtained according to the magnitude relationship of the two floating-point numbers; A normalization module, configured to perform a normalization shift operation on the result of the unsigned fixed-point full addition operation according to the number of leading zeros of the result of the unsigned fixed-point full addition operation and the shift remainder, so as to obtain a normalization result; A normalization processing module, configured to perform an exponent adjustment on the larger exponent bit of two floating-point numbers based on the number of leading zeros, and perform an overflow judgment and a result normalization processing on the normalization result by using the exponent adjustment result and the sign bit of the final calculation result.
10. The floating-point arithmetic unit according to claim 9, wherein The performing the unsigned fixed-point full addition operation by using the large mantissa and the small mantissa after the negation operation includes: When the exponent bits of two floating-point numbers are equal, performing an increment operation on the small mantissa after the negation operation, and performing an unsigned fixed-point full addition operation on the large mantissa and the small mantissa after the increment operation; When the exponent bits of two floating-point numbers are not equal, performing an unsigned fixed-point full addition operation on the large mantissa and the small mantissa after the negation operation.
Citation Information
Cited By
Data link waveform data preprocessing method based on digital operation
CN121958767A
A data link waveform data preprocessing method based on digital operation
CN121958767B