A hardware computing device and method for floating-point division and square root calculation

CN116578344BActive Publication Date: 2026-09-01青岛本原微电子有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310378233.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2026-09-01
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

[0003]目前,浮点处理装置对诸如浮点除法、浮点平方根等计算的功能操作,主要依赖应用软件来实现,只是输出运算结果,对于运算过程中存在的数据异常操作及特殊类型未进行标注,使得运算结果不可靠,无法满足对计算速度可靠性要求较高的场合,当运算结果不可靠时需要重复运算,使得运算实时性差

Benefits of technology

[0033] To ensure accurate calculations with minimal hardware overhead, this device first integrates floating-point division and floating-point square root calculations, requiring only one device for each calculation. The device considers comprehensive functionality, including handling denormalized operands, using full-width data for rounding floating-point mantissas during division and square root operations, supporting five rounding cases, and supporting five exception flags. This ensures the highest accuracy requirements for floating-point division and square root calculations, while also providing comprehensive flags for special calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116578344B_ABST
    Figure CN116578344B_ABST
Patent Text Reader

Abstract

This invention relates to the field of floating-point arithmetic technology within microprocessors, and discloses a hardware calculation device and method for floating-point division and square root calculation. The device employs a 16-stage pipeline structure, divided into three parts: the first part is a data preprocessing section with one pipeline stage; the second part is an iteration section with 14 pipeline stages, used for mantissa division and square root iteration operations, and for obtaining exponent results; the third part is a final data processing section with one pipeline stage, used for special data processing, denormalization processing, five types of rounding, normalization, and five types of exception flag processing. The device and method disclosed in this invention offer high computational accuracy, small hardware resources, and comprehensive functionality. This device improves resource utilization by reducing the complexity of rounding mode processing and reusing resources; it ensures the highest accuracy requirements for floating-point division and square root calculations while simultaneously deriving comprehensive flags for special calculations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of floating-point arithmetic technology in microprocessors, and particularly to a hardware computing device and method for floating-point division and square root calculation. Background Technology

[0002] For DSP processors, signal processing functions are frequently called. Many current signal processing functions include calls to trigonometric functions, typically implemented by calling system library functions, which is inefficient, time-consuming, and involves numerous instructions. Adding a trigonometric function computation unit (TCU) instead of calling system library functions can significantly improve the computation speed of functions like FFT, thereby enhancing performance. The TCU provides hardware support for common trigonometric function operations and is an extension of the floating-point unit. By efficiently executing trigonometric and arithmetic operations commonly used in control system applications, it enhances the instruction set architecture of the DSP chip. Floating-point division and floating-point square root are part of trigonometric function computations.

[0003] Currently, floating-point processing devices mainly rely on application software to perform calculations such as floating-point division and floating-point square root. They only output the calculation results and do not label abnormal data operations and special types that occur during the calculation process, making the calculation results unreliable. This cannot meet the requirements of occasions with high requirements for calculation speed and reliability. When the calculation results are unreliable, the calculation needs to be repeated, resulting in poor real-time performance.

[0004] When implemented in hardware, floating-point division and floating-point square root operations have higher hardware complexity than other basic arithmetic operations (adders, subtractors, and multipliers), requiring a larger area while achieving relatively lower performance. Improving the hardware implementation of floating-point division and floating-point square root operations to reduce resource consumption while increasing functionality has become an important research direction in current arithmetic unit research. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a hardware computing device and method for floating-point division and square root calculation, which significantly enhances the real-time performance and accuracy of DSP chips in calculating signal processing-related functions, while also significantly reducing hardware resource overhead.

[0006] To achieve the above objectives, the technical solution of the present invention is as follows:

[0007] A hardware computing device for floating-point division and square root calculation adopts a 16-stage pipeline structure and is divided into three parts;

[0008] The first part is the data preprocessing section, which has one pipeline stage, including the input DFF, enable control module, input data preprocessing module, leading zero detection module 1, leading zero detection module 2, 8-bit addition module, 24-bit left shift module, sign bit processing module, special data detection module and first-stage DFF;

[0009] The second part is the iteration part, which has a total of 14 pipeline stages. It is used to process mantissa division and square root iterative operations and to obtain exponent results. The second stage includes the exponent operation module and the iteration unit.

[0010] The third part is the final data processing section, which consists of one pipeline stage. It is used for special data processing, denormalization processing, five types of rounding, normalization, and five types of exception flag processing. The third stage includes a denormalization processing module, a sign bit calculation module, a preprocessing module, a postprocessing module, a result splicing module, and the last stage DFF.

[0011] In the above scheme, the iteration unit includes a floating-point mantissa square root operation iteration unit and a floating-point mantissa division operation iteration unit. The floating-point mantissa square root operation iteration unit and the floating-point mantissa division operation iteration unit share a 29-bit adder 1, a 29-bit adder 2, a quotient operation module, a rounding bit merging module, an iteration control module, and the second to fifteenth level DFF.

[0012] In a further technical solution, the floating-point mantissa square root operation iteration unit also includes a mantissa processing module, a fast calculation module, an inversion module 1, an inversion module 2, MUX1, MUX2, MUX3, MUX4, a data processing module 1, a data processing module 2, and a rounding bit calculation module.

[0013] In a further technical solution, the floating-point mantissa division operation iteration unit also includes an operand 1 preprocessing module, an operand 2 preprocessing module, MUX5, MUX6, MUX7, a data processing module 3, a data processing module 4, and a 29-bit comparator.

[0014] A hardware method for floating-point division and square root calculation includes the following steps:

[0015] Phase 1: Data Preprocessing

[0016] The input data consists of the opcode, opcode enable, rounding mode, operand fa, and operand fb, all of which are stored in the input DFF.

[0017] The opcode and opcode validity enable stored in the DFF register are entered into the enable control module to obtain the current floating-point division operation enable signal and floating-point square root operation enable signal; in rounding mode, it directly enters the first-level DFF; operands fa and fb enter the input data preprocessing module to split the data and obtain the sign bit, exponent bit, and mantissa bit of operands fa and fb. At the same time, it judges the output of the enable control module. If all are invalid, the sign bit, exponent bit, and mantissa bit are set to 0.

[0018] The mantissa bit output by the input data preprocessing module is fed to the leading zero detection module 1 and the leading zero detection module 2 for leading zero detection to obtain the number of high-order zeros. This number is then transmitted to the 8-bit addition module and the 24-bit left shift module. The 8-bit addition module adds this data to the exponent bit output by the input data preprocessing module to obtain the normalized exponent bit. The 24-bit left shift module uses this data as the shift value to shift the number of bits output by the input data preprocessing module to obtain the normalized mantissa bit. The sign bit output by the input data preprocessing module is transmitted to the sign bit processing module to calculate the XOR operation result of the sign bits of fa and fb to obtain the sign bit result of the division. The sign bit of the square root is directly obtained from the sign bit of fa. The special data detection module performs positive 0, negative 0, positive infinity, negative infinity, QNaN, and SNaN detection based on the sign bit, exponent bit, and mantissa bit of the input data preprocessing module to obtain special data indication signals.

[0019] The floating-point division enable signal, the floating-point square root enable signal, the normalized exponent and mantissa, the sign bit, and the special data indicator signal are all registered in the first-stage DFF, waiting for the next stage of pipeline operation.

[0020] Phase Two: Iteration Phase

[0021] The calculation is then divided into three paths: sign bit calculation, exponent bit calculation, and mantissa bit calculation. The sign bit calculation is completed in the third stage, while the exponent bit calculation and mantissa bit calculation are completed in the second stage.

[0022] For exponent calculation, the exponent operation module is used to obtain the result of floating-point division: fa exponent plus fb exponent plus 63. The square root exponent operation is the result of shifting fa exponent one bit to the right, adding the least significant bit of fa exponent, and then adding 127. Since the least significant bit is processed when it is odd, the mantissa processing module is used to shift the mantissa of fa one bit to the left based on the least significant bit of fa exponent to obtain the accurate exponent.

[0023] For mantissa calculations, iterative units are used, specifically including iterative units for floating-point mantissa square root operations and floating-point mantissa division operations. Both are implemented using the same main framework. The 29-bit adder 1, 29-bit adder 2, quotient operation module, iteration control module, rounding bit merging module, and the second to fifteenth levels of DFFs reuse resources, having the same functionality but using different operands, as detailed below:

[0024] For floating-point square roots, the normalized mantissa is transmitted to the mantissa processing module, where it undergoes shifting and adjustment before entering the fast calculation module. The high 2 bits are used to generate a carry and 2 data bits, which are then transmitted to 29-bit adder 1 and data processing module 1 respectively. Simultaneously, the low 2 bits are used to generate a carry and 2 data bits, which are transmitted to 29-bit adder 2 and data processing module 2 respectively. Since this is the first time entering the loop, the 29 bits of 0 are transmitted via MUX1 to 29-bit adder 1 and rounding bit calculation module, and the 29 bits of 1 are transmitted via MUX2 to the same module. The rounding bit calculation module derives the rounding bit based on the operands and transmits it to the rounding bit merging module. After bit adder 1 completes the addition operation, it obtains the carry and Sum. The carry is transmitted to the quotient operation module to obtain the first quotient data. At the same time, the carry is transmitted to MUX4 to select one from the output of the quotient operation module and the output of the inverting module 2, and then transmits it to 29-bit adder 2. Sum is transmitted to data processing module 1 to complete the bit concatenation operation, and then transmitted to 29-bit adder 2. After bit adder 2 completes the addition operation, it obtains the carry and Sum. The carry is transmitted to the quotient operation module to obtain the first quotient data, and Sum is transmitted to data processing module 2 to complete the bit concatenation operation.

[0025] The iteration control module is activated, and the counter is 0. When one of the floating-point division operation enable signals or floating-point square root operation enable signals from the enable control module in the first-level DFF register is valid, the carry outputs of the enable control module, exponentiation operation module, data processing module 2, quotient operation module, iteration control module, and 29-bit adder 2 in the first-level DFF register are registered to the second-level DFF. This completes one cycle. Afterward, the rounding bit merging module completes the rounding bit and quotient merging, and outputs it as the mantissa. The next iteration process begins. The output of data processing module 2 in the second-level DFF register is fed into 29-bit adder 1 and rounding bit calculation module through MUX1. The carry generated by 29-bit adder 2 in the second-level DFF register is used as a selection signal to select the result of quotient operation module in the second-level DFF register or the result of inversion logic operation after quotient operation module is transmitted to inversion module 2. This result is then transmitted to 29-bit adder 1 and rounding bit calculation module through MUX2. The rounding calculation module calculates the rounding bits based on the operands and transmits them to the rounding bit merging module. The carry and 2-bit data generated by the fast calculation module are transmitted to the 29-bit adder 1 and data processing module 1, respectively. The operation is then the same as above until the second iteration is completed. When one of the floating-point division operation enable signal or floating-point square root operation enable signal from the enable control module of the previous level DFF is valid, the carry outputs of the enable control module, the exponent operation module, the data processing module 2, the quotient operation module, the iteration control module, and the 29-bit adder 2 are registered in the third level DFF. When invalid, the data remains unchanged. The operation is then the same as above until the 14th iteration. The iteration control module outputs an iteration end signal, which is valid and registered in the fifteenth level DFF. In the next cycle, the iteration control module sets the internal counter value to 0 and waits for the next calculation. The above outputs the exponent result, the mantissa result, and the iteration end signal.

[0026] For floating-point division, the normalized mantissa is processed by the operand 1 preprocessing module and the operand 2 preprocessing module to complete data concatenation. Since it is the first time entering the loop, the data is transmitted through MUX5 and MUX6 to the 29-bit adder 1 to perform addition and obtain the carry and Sum. The carry signal is transmitted to the quotient operation module to obtain the first quotient data. Sum is transmitted to the data processing module 3 to complete the bit concatenation operation, and then transmitted to the 29-bit adder 2 and the 29-bit comparator. The most significant bit of Sum is transmitted to MUX7 as a selection signal to transmit the output data of the operand 2 preprocessing module to the 29-bit adder 2.

[0027] After the 29-bit adder 2 completes the addition operation, it obtains the carry and Sum. The carry signal is transmitted to the quotient operation module to obtain the second quotient data. Sum is transmitted to the data processing module 4 to complete the bit concatenation operation, and then transmitted to the 29-bit comparator to complete the comparison operation and obtain the rounding bit. The iteration control module is activated, and the counter is 0. When one of the floating-point division operation enable signal or the floating-point square root operation enable signal from the enable control module in the first-level DFF register is valid, the outputs of the enable control module, the exponentiation operation module, the data processing module 4, the quotient operation module, the iteration control module, and the 29-bit adder 2 carry are registered in the second-level DFF. This completes one cycle. After that, the rounding bit merging module completes the merging of the rounding bit and the quotient, and outputs it as the mantissa. The next iteration process begins. The output result of the data processing module 4 registered in the second-level DFF is transmitted to the 29-bit adder 4 through the MUX5. Adder 1, and simultaneously the carry stored in the second-level DFF is transmitted to MUX6 as a selection signal, transmitting the data of the operand 2 preprocessing module to 29-bit adder 1. The operation then proceeds as described above until the second iteration is completed. When one of the floating-point division or floating-point square root operation enable signals from the enable control module of the previous-level DFF is valid, the carry outputs of the enable control module, the exponent operation module, data processing module 2, quotient operation module, iteration control module, and 29-bit adder 2 are stored in the third-level DFF. When invalid, the data remains unchanged. The operation then proceeds as described above until the 14th iteration. The iteration control module outputs an iteration end signal, which is valid and stored in the fifteenth-level DFF. In the next cycle, the iteration control module sets its internal counter value to 0, waiting for the next operation. The above outputs the exponent result, mantissa result, and iteration end signal.

[0028] Phase 3: Final Data Processing Phase

[0029] For sign bit calculation, the sign bit calculation module is used. Based on the floating-point division operation enable signal and the floating-point square root operation enable signal, special data is processed first. When there is no special data, normal data is processed to obtain the sign bit data.

[0030] When the calculated exponent is less than 0, the denormalization module is used to perform a shift operation and output the denormalized number data to the preprocessing module. Then, the preprocessing module processes the special data according to the special data indication signal and enable control signal generated in the first stage to obtain the exponent and mantissa results and the NV exception flag. If it is not a special value, the exponent and mantissa generated in the second stage and the exponent and mantissa results from the denormalization module are used to perform boundary value processing and output the exponent, mantissa, rounding bits and OF and UF exception flags of the first result processing.

[0031] The result then enters the post-processing module, which performs five rounding operations based on the mantissa obtained from the pre-processing module, the rounding bits stored in the first-level DFF register, and the sign bit generated by the sign bit calculation module. If there is a carry, the exponent is incremented by 1. The rounded result is then subjected to boundary checks to determine if there is overflow or underflow, and an exception flag, exponent, and mantissa are output. Finally, the result concatenation module merges the exponent, mantissa, and exception flag generated by the pre-processing module, as well as the exception flag, exponent, and mantissa output by the post-processing module. The exponent, mantissa, and the output of the sign bit calculation module are concatenated to obtain the final single-precision floating-point division and square root operation results, along with the five exception flags. Finally, these results, along with the control enable, are registered in the last-level DFF.

[0032] The floating-point division and square root calculation hardware device and method provided by the present invention through the above technical solution have the following beneficial effects:

[0033] To ensure accurate calculations with minimal hardware overhead, this device first integrates floating-point division and floating-point square root calculations, requiring only one device for each calculation. The device considers comprehensive functionality, including handling denormalized operands, using full-width data for rounding floating-point mantissas during division and square root operations, supporting five rounding cases, and supporting five exception flags. This ensures the highest accuracy requirements for floating-point division and square root calculations, while also providing comprehensive flags for special calculations.

[0034] Because of its comprehensive functionality, implementing it independently would incur significant hardware resource overhead. Therefore, this device employs extensive resource reuse during its implementation. Firstly, denormalized operand processing is integrated into the normalized data stream to maximize hardware resource sharing and improve resource reuse. Secondly, when performing division and square root operations on floating-point mantissas using full-width data for rounding, fewer logic operations are employed to achieve the required functionality. Thirdly, this device improves resource reuse by reducing the complexity of rounding mode processing. Finally, it improves resource reuse by reducing the complexity of exception flag handling. In summary, the advantages of this device are high computational accuracy, small hardware resource requirements, and comprehensive functionality. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0036] Figure 1 This is a schematic diagram of the overall process framework of a floating-point division and square root extraction device disclosed in an embodiment of the present invention.

[0037] Figure 2 This is a schematic diagram of the floating-point mantissa square root operation iterative device framework disclosed in the embodiments of the present invention.

[0038] Figure 3 This is a schematic diagram of the floating-point mantissa division iterative device framework disclosed in the embodiments of the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and figures and tables. The specific implementation will be illustrated using the IEEE 754-2008 standard floating-point single-precision format as an example. This example can be extended to floating-point data format operations with an m-bit exponent and n-bit mantissa.

[0040] First, we introduce floating-point data, rounding modes, and exception flag handling. The specific representation of floating-point data is shown in Table 1. Besides normalized numbers, some special values ​​have specific representations in the standard. These special values ​​include positive 0, negative 0, positive infinity (+Inf), negative infinity (-Inf), denormalized numbers, and Not-a-N (NaN). Denormalized numbers and Not-a-N are handled differently in floating-point arithmetic. General floating-point arithmetic does not support denormalized number operations, but this device does, resulting in higher arithmetic precision. Not-a-N can be represented as Quiet-NaN (QNaN) or Signaling-NaN (SNaN) based on its most significant bit. The exception flag settings differ for these two types of Not-a-N.

[0041] Table 1 Floating-point format data

[0042]

[0043]

[0044] The specific meanings of the rounding modes are shown in Table 2. There are a total of 5 rounding modes: Round to nearest neighbor to even number (RNE), Round to zero (RTZ), Round to negative infinity (RDZ), Round to positive infinity (RUP), and Round to nearest neighbor to farthest number (RMM).

[0045] Table 2 Meaning of Rounding Mode

[0046]

[0047] The exception flags are set as follows:

[0048] (1) NV anomaly indicators:

[0049] If any operand is NaN, the output is Canonical-NaN, i.e., 0x7fc00000. If any operand is SNaN, the NV exception flag needs to be set. If floating-point division results in calculations of 0 / 0 or Inf / Inf, or if floating-point square root operations result in negative operands (excluding negative 0), the output is Canonical-NaN, and the NV exception flag needs to be set.

[0050] (2) DZ anomaly indicators:

[0051] For floating-point division operations, if the calculation form is a / 0, the output result is Inf, and the DZ exception flag is set.

[0052] (3) OF Abnormal Indicators:

[0053] An overflow exception signal should be issued if and only if the target data range is unbounded and the resulting floating-point value exceeds the maximum finite number of the target format. The default result should be determined by the rounding direction attribute and the sign of intermediate results. RNE and RMM rounding use the sign of intermediate results to carry all overflows to Inf. RTZ rounding carries all overflows to the maximum finite number of the format with the sign of intermediate results. RDN rounding mode carries positive overflows to the maximum finite number of the format and negative overflows to -Inf. RUP rounding mode carries negative overflows to the least negative finite number of the format and positive overflows to +Inf.

[0054] In addition, under the default exception handling for overflow, the OF exception flag should be set and the NX exception flag signal should be emitted.

[0055] (4) UF anomaly indicators:

[0056] The IEEE 754-2008 floating-point standard offers two options: round-before detection and round-in detection. This design uses round-before detection. After rounding, the result that is non-zero (the data range is unbounded) is strictly within ±b. emin When b is 2 and emin is the minimum normalization exponent -126, b emin It is the normalized number with the smallest absolute value.

[0057] The default exception handling for underflow must always deliver a rounded result, which could be zero, a denormalized number, or ±b. emin The absolute value ratio of b emin Small numbers can be represented as b emin or -b emin 0 and denormalized numbers are used. The output of this design is represented by denormalized numbers.

[0058] Under the default underflow exception handling, if the rounding result is inaccurate, the UF flag must be set and an NX exception must be sent. If the rounding result is accurate, no flag is set and no UF exception flag is signaled.

[0059] (5) NX anomaly flags:

[0060] If the result of a rounding operation is inaccurate, the NX exception flag must be sent.

[0061] The implementation scheme of this device will be described in detail below.

[0062] This invention provides a hardware computing device for floating-point division and square root calculation, which adopts a 16-stage pipeline structure and is divided into three parts:

[0063] Part 1

[0064] The first part is the data preprocessing section, which consists of one pipeline stage, including the input DFF, enable control module, input data preprocessing module, leading zero detection module 1, leading zero detection module 2, 8-bit addition module, 24-bit left shift module, sign bit processing module, special data detection module, and the first-stage DFF.

[0065] The input DFF is an edge-sensitive flip-flop in a sequential circuit. It is enabled by the opcode and used to register the input opcode, opcode enable, rounding mode, fa and fb data when the enable is valid.

[0066] The enable control module performs enable operations based on the input opcode and the opcode's valid enable, and obtains the current floating-point division operation enable signal and floating-point square root operation enable signal.

[0067] The input data preprocessing module extracts the sign bit, exponent bit, and mantissa bit of the data based on the two input operands. When normalizing numbers, a 1 is added to the hidden mantissa bit; when denormalizing numbers, a 0 is added to the hidden mantissa bit to ensure data integrity. The module outputs the sign bit, exponent bit, and mantissa bit of fa and fb. Simultaneously, it checks the output of the enable control module; if all checks are invalid, the sign bit, exponent bit, and mantissa bit are set to 0.

[0068] Leading zero detection module 1 performs high-order zero detection on the mantissa of fa output by the input data preprocessing module and outputs the number of high-order zeros.

[0069] Leading zero detection module 2 performs high-order zero detection on the mantissa of fa output by the input data preprocessing module and outputs the number of high-order zeros.

[0070] The 8-bit addition module performs calculations based on the number of high-order zeros in fa and fb output by leading zero detection 1 and leading zero detection 2, and the 8-bit exponents of fa and fb output by the input data preprocessing module. The calculated exponent is subtracted from the number of high-order zeros and then 1 is added to obtain the normalized exponents of fa and fb.

[0071] The 24-bit left shift module uses the number of high-order zeros in the outputs of fa and fb from leading zero detection 1 and leading zero detection 2 as the shift value to shift the mantissas of the input data preprocessing module's outputs fa and fb to the left, thus obtaining the normalized mantissas of fa and fb.

[0072] The sign bit processing module determines the sign bit to be corrected based on the operation enable signal output by the enable control. The operation data is the sign bits of fa and fb output by the input data preprocessing module. The sign bits of fa and fb are XORed to obtain the division sign bit. The sign bits of the output result fa / fb and the sign bit of fa are sent to the sign bit processing module of the next stage of the pipeline cycle.

[0073] The special data detection module, based on the sign bit, exponent bit, and mantissa bit of the input data preprocessing module's output fa and fb, detects positive 0, negative 0, positive infinity, negative infinity, SNAN, and QNAN, and outputs the corresponding indicator signals for these special values.

[0074] The first-level DFF uses the floating-point division and floating-point square root operation enable signals output by the enable control module as control enable signals. When one of them is valid, it is used to register the output results of the input DFF, enable control module, 8-bit addition module, 24-bit left shift module, sign bit processing module, and special data indicator module. When all are invalid, the original data remains unchanged.

[0075] Part Two

[0076] The second part is the iteration section, which consists of 14 pipeline stages. It is used to process mantissa division and square root iteration operations and to obtain exponent results. The second stage includes the exponent operation module and the iteration unit.

[0077] The exponentiation module, based on the normalized exponents of fa and fb obtained from the 8-bit addition module in the first stage of the first-level DFF register and the control signal from the enable control module, performs floating-point division or floating-point square root exponentiation. When the floating-point number is single-precision floating-point, the division exponentiation is fa exponent plus fb exponent plus 63, and the square root exponentiation is fa exponent right-shifted by one bit plus the least significant bit of fa exponent plus 127. The module outputs the division or square root exponentiation result. The iteration unit performs calculations based on the control enable, fa exponent and mantissa, and fb exponent and mantissa. Division is performed on the mantissas of fa and fb, and square root is performed on the mantissa of fa. Finally, the division mantissa result and the square root mantissa result are output. Figure 2 and Figure 3 As shown.

[0078] like Figure 2 and Figure 3 The diagram shows the floating-point mantissa square root operation iteration unit and the floating-point mantissa division operation iteration unit. These are used for mantissa division operations of fa and fb, as well as the square root operation of the mantissa of fa. Both mantissa division and square root operations employ an iterative algorithm without restoring the remainder. The division and square root operation iteration units share two 29-bit adders, a quotient operation module, a rounding bit merging module, an iteration processing module, and the second to fifteenth level DFFs, thus reusing resources and reducing resource consumption. The square root operation iteration unit and the division operation iteration unit will be described separately below.

[0079] like Figure 2 As shown, this is the floating-point mantissa square root operation iteration unit, which mainly includes a mantissa processing module, a fast calculation module, an inversion module 1, an inversion module 2, MUX 1, MUX 2, MUX 3, MUX 4, a 29-bit adder 1, a 29-bit adder 2, a data processing module 1, a data processing module 2, a rounding bit calculation module, a quotient operation module, a rounding bit merging module, an iteration processing module, and the second to fifteenth level DFFs.

[0080] The mantissa processing module shifts the mantissa of fa by one bit to the left based on the least significant bit of the exponent of fa output by the first-stage DFF.

[0081] The fast calculation module, in each iteration, takes 4 bits of data sequentially from the most significant bit to the least significant bit, based on the output of the mantissa processing module. These are then divided into two parts: the high 2 bits and the low 2 bits, and processed in this manner 14 times. If the output value from the mantissa processing module is insufficient, it is padded with 0s. The two extracted two-bit data are then subjected to a logical AND operation to output the carry; subsequently, they are subjected to a logical XOR operation, and the output is used as the high 2 bits. This high 2 bits are then inverted with the low 2 bits of the extracted two-bit data and used as the low 2 bits. The resulting bits are then concatenated to output two two-bit data. Therefore, this module outputs both the carry and the two-bit result calculated using the high 2 bits, and the carry and the two-bit result calculated using the low 2 bits.

[0082] Invert module 1 inverts the output of the quotient operation module registered in the second to fifteenth level DFFs, and then outputs the result.

[0083] Invert module 2 inverts the output of the quotient operation module and then outputs the result.

[0084] MUX1 is a selector unit that uses whether it is the first time entering the loop as the selection signal. When the selection signal is valid, it outputs 29 bits of 0; when the selection signal is invalid, it selects the output result of the data processing module 2 registered in the second to fifteenth levels of DFF as the output.

[0085] MUX2 is a selector unit that uses whether it is the first time entering the loop as the selection signal. When the selection signal is valid, it outputs 29 bits as 1; when the selection signal is invalid, it selects the output result of MUX3 as the output.

[0086] MUX3 is a selector unit that uses the carry value of the 29-bit adder 2 stored in the second to fifteenth level DFF as the selection signal. When the selection signal is valid, the output of the inverting module 1 is selected as the output; when the selection signal is invalid, the output of the quotient operation module stored in the second to fifteenth level DFF is selected as the output.

[0087] MUX4 is a selector unit that uses the carry output from the 29-bit adder 1 as the selection signal. When the selection signal is valid, it selects the output of the inverting module 2 for output; when the selection signal is invalid, it selects the output of the quotient operation module for output.

[0088] The 29-bit adder 1 is an adder with carry. The first operand uses the output of MUX1, the second operand uses the output of MUX2, and the carry value is generated by the higher two bits of the fast calculation module's output. The adder ultimately obtains the Sum and the carry result, which is the quotient, and outputs them.

[0089] The 29-bit adder 2 is an adder with carry. The first operand uses the output of data processing module 1, the second operand uses the output of MUX4, and the carry value is generated by the lower two bits of the fast calculation module output. The adder ultimately obtains the Sum and the carry result, which is the quotient, and outputs them.

[0090] Data processing module 1 operates on the Sum output by 29-bit adder 1 and the overall two-bit output of the fast calculation module, and performs bit concatenation of Sum

[28] , Sum[25:0] and the overall two-bit output result, and finally outputs the result.

[0091] Data processing module 2 operates on the Sum output from 29-bit adder 1 and the overall two-bit output from the fast calculation module, performing bit concatenation of Sum

[28] , Sum[25:0], and the overall two-bit output result, and finally outputting the result. This module has the same function as data processing module 1, but selects different data.

[0092] The quotient operation module shifts the quotient value one bit to the left based on the carry value generated by either 29-bit adder 1 or 29-bit adder 2, uses the carry value generated by either 29-bit adder 1 or 29-bit adder 2 as the least significant bit, and finally outputs the quotient value.

[0093] The iterative control module uses a counter to increment from 0, adding 1 with each iteration. When the counter reaches 13, an iteration end signal is generated and output, and the value is set to 0. The enable control module, using the first-level DFF register, outputs floating-point division and floating-point square root operation enable signals as the control enable to start counting.

[0094] Levels 2-15 of DFF, Level 2 DFF is based on... Figure 1 The floating-point division and floating-point square root operations enable signals output by the enable control module of the first-level DFF register serve as control enable signals. When one of them is active, it is used for register operation. Figure 1 The enable control module of the first-level DFF register, Figure 1 When the carry values ​​output by the exponentiation module, data processing module 2, quotient operation module, and 29-bit adder 2, as well as the output data of the iteration control module, are all invalid, the original data remains unchanged. For the third to fifteenth level DFFs, the floating-point division and floating-point square root operation enable signals output by the enable control module registered in the previous level DFF are used as control enable signals. When one of these signals is valid, it is used to register the carry values ​​output by the enable control module, the exponentiation module, data processing module 2, quotient operation module, and 29-bit adder 2, as well as the output data of the iteration control module registered in the previous level DFF. When all of these are invalid, the original data remains unchanged.

[0095] The rounding bit calculation module operates based on the outputs of MUX1 and MUX2. First, the output of MUX2 is shifted left by one bit and padded with 1 for the least significant bit. Then, it is added to the output of MUX1 to obtain Sum. The most significant bit of MUX1 is used as the selection signal. If valid, Sum is selected for bitwise OR operation; if invalid, the output of MUX1 is selected for bitwise OR operation. Finally, the rounding bit is output.

[0096] The rounding merging module performs an OR operation between the least significant bit of the quotient value output by the quotient operation module and the rounding bit generated by the rounding calculation module, and outputs the quotient value, which is used as the mantissa output in the second stage.

[0097] like Figure 2 As shown, this is the iterative unit for the square root operation, which finally outputs the quotient result of the division of the last digit.

[0098] like Figure 3 The diagram shows the division operation iteration unit, which includes an operand 1 preprocessing module, an operand 2 preprocessing module, MUX5, MUX6, MUX7, a 29-bit adder 1, a 29-bit adder 2, a data processing module 3, a data processing module 4, a 29-bit comparator, a quotient operation module, an iteration control module, a rounding bit merging module, and second to fifteenth level DFFs. The 29-bit adder 1, 29-bit adder 2, quotient operation module, rounding bit merging module, second to fifteenth level DFFs, and iteration control module are integrated with... Figure 2 Modules with the same name can reuse resources, have the same function, but use different operands.

[0099] The function of the operand 1 preprocessing module is to... Figure 1 The first-level DFF register fb is operated on, and the operand result is output by bit concatenating 1'b0, the mantissa of fb, 1'b1, and 3'b0.

[0100] The function of the operand 2 preprocessing module is to... Figure 1 The most significant bit of the mantissa of fb in the first-level DFF register is padded with 0 and then inverted. The inverted mantissa of fb, 1'b1, and 3'b0 are then concatenated to output the first operand result. The inverted mantissa of fb and 4'b0 are then concatenated to output the second operand result.

[0101] MUX5 is a selector unit. The selection signal is based on whether it is the first time entering the loop. When the selection signal is valid, the output result of the preprocessing of operand 1 is selected as the output. When the selection signal is invalid, the output result of the data processing module 4 registered in the second to fifteenth DFF is selected as the output.

[0102] MUX6 is a selector unit that uses the carry signal from the 29-bit adder 2 output registered in the second to fifteenth DFF as the selection signal, depending on whether it is the first time entering the loop or the carry signal. The first time entering the loop has higher priority. When the selection signal is valid, the first operand result of the preprocessed output of operand 2 is selected; when the selection signal is invalid, the second operand result of the preprocessed output of operand 2 is selected.

[0103] MUX7 is a selector unit that uses the most significant bit of the Sum output from the 29-bit adder 1 as the selection signal. When the selection signal is valid, it selects the first operand result of the preprocessed output of operand 2; when the selection signal is invalid, it selects the second operand result of the preprocessed output of operand 2.

[0104] 29-bit adder 1, with the same function as... Figure 2 In contrast, the operands chosen are different. The first operand uses the output of MUX5, and the second operand uses the output of MUX6. The 29-bit adder 1 ultimately obtains the Sum and carry result, which is the quotient, and outputs it.

[0105] 29-bit adder 2, with the same function, and Figure 2 In contrast, the operands selected are different. The first operand is the data output from data processing module 3, and the second operand uses the output of MUX7. The 29-bit adder 2 ultimately obtains the Sum and carry result, which is the quotient, and outputs it.

[0106] Data processing module 3 operates based on the Sum output by 29-bit adder 1. First, it inverts Sum

[28] , then it concatenates the inverted data of Sum[27:3] and Sum

[28] with 3'b0, and finally outputs the result.

[0107] Data processing module 4 operates on the Sum and carry output by 29-bit adder 2, performing bit concatenation on Sum[27:3], carry value, and 3'b0, and finally outputting the result.

[0108] A 29-bit comparator compares the data from data processing module 3 with the data from data processing module 4. If they are equal, the rounding bit is 0; otherwise, it is 1.

[0109] The quotient operation module shifts the quotient value one bit to the left based on the carry value generated by either 29-bit adder 1 or 29-bit adder 2, uses the carry value generated by either 29-bit adder 1 or 29-bit adder 2 as the least significant bit, and finally outputs the quotient value.

[0110] The iterative control module uses a counter to increment from 0, adding 1 with each iteration. When the counter reaches 13, an iteration end signal is generated and output, and the value is set to 0. The enable control module, using the first-level DFF register, outputs floating-point division and floating-point square root operation enable signals as the control enable to start counting.

[0111] Levels 2-15 of DFF, Level 2 DFF is based on... Figure 1 The floating-point division and floating-point square root operations enable signals output by the enable control module of the first-level DFF register serve as control enable signals. When one of them is active, it is used for register operation. Figure 1 The enable control module of the first-level DFF register, Figure 1 When the carry values ​​from the exponentiation module, data processing module 4, quotient operation module, 29-bit adder 2, and iteration control module are all invalid, the original data remains unchanged. For DFFs at levels three through fifteen, the floating-point division and floating-point square root operation enable signals from the enable control module registered in the previous DFF are used as control enable signals. When one of these signals is valid, it is used to register the output data from the enable control module, the carry values ​​from the exponentiation module, data processing module 4, quotient operation module, 29-bit adder 2, and iteration control module registered in the previous DFF. When all of these are invalid, the original data remains unchanged.

[0112] The rounding merging module performs an OR operation between the least significant bit of the quotient value output by the quotient operation module and the rounding bit generated by the rounding calculation module, and outputs the quotient value, which is used as the mantissa output in the second stage.

[0113] like Figure 3 As shown, this is the division operation iteration unit, which finally outputs the quotient result of the mantissa division.

[0114] Part Three

[0115] The third part is the final data processing section, which consists of one pipeline stage. It is used for special data processing, denormalization processing, five types of rounding, normalization, and five types of exception flag processing. The third stage includes a denormalization processing module, a sign bit calculation module, a preprocessing module, a postprocessing module, a result splicing module, and the last stage DFF.

[0116] The denormalization module uses the exponent output from the second stage as a judgment signal, calculates the absolute value of the exponent as the shift value, performs a right shift operation on the mantissa result output from the second stage, and finally outputs the exponent and mantissa results represented by the denormalized number.

[0117] The sign bit calculation module outputs the sign bit of the special value based on the special data indication signal and the enable control signal generated in the first stage. If it is not a special value calculation, it selects the sign bit of the division or the square root sign bit of the sign bit of the first-level DFF register to be output based on the enable control signal of the first-level DFF register.

[0118] The preprocessing module processes the special data based on the special data indication signal and enable control signal generated in the first stage to obtain the exponent and mantissa results and the NV anomaly flag. If it is not a special value, it performs boundary value processing based on the exponent and mantissa generated in the second stage and the exponent and mantissa results from the denormalization processing module, and outputs the exponent, mantissa, rounding bits and OF and UF anomaly flags of the first result processing.

[0119] The post-processing module performs five rounding operations based on the mantissa obtained from the pre-processing module, the rounding bits stored in the first-level DFF register, and the sign bit generated by the sign bit processing module. If a carry-in exists, the exponent is incremented by 1. The rounded result is then subjected to boundary checks to determine if overflow or underflow occurs, and an exception flag, exponent, and mantissa are output.

[0120] The result concatenation module merges the exponent, mantissa, and exception flags generated by the preprocessing module and the exception flags, exponent, and mantissa output by the postprocessing module. It then concatenates the outputs of the exponent, mantissa, and sign bit calculation module to obtain the final single-precision floating-point division and square root operation results and five exception flags.

[0121] The final DFF uses the output of the iteration control module registered in the fifteenth-level DFF and the floating-point division and floating-point square root operation enable signals output by the first-level DFF enable control module as control enable signals. When the output of the iteration control module registered in the fifteenth-level DFF is valid and one of the floating-point division and floating-point square root operation enable signals output by the enable control module is valid, it is used to store the operation result and exception flag data generated by the result splicing module. When invalid, the data is retained.

[0122] In summary, special data processing only requires the first and third stages of processing, which necessitates a 2-stage pipeline. Normal data processing requires the first, second, and third stages, which necessitates a 16-stage pipeline.

[0123] A hardware method for floating-point division and square root calculation includes the following steps:

[0124] Phase 1: Data Preprocessing

[0125] The input data consists of the opcode, opcode enable, rounding mode, operand fa, and operand fb, all of which are stored in the input DFF.

[0126] The opcode and opcode validity enable stored in the DFF register are entered into the enable control module to obtain the current floating-point division operation enable signal and floating-point square root operation enable signal; in rounding mode, it directly enters the first-level DFF; operands fa and fb enter the input data preprocessing module to split the data and obtain the sign bit, exponent bit, and mantissa bit of operands fa and fb. At the same time, it judges the output of the enable control module. If all are invalid, the sign bit, exponent bit, and mantissa bit are set to 0.

[0127] The mantissa bit output by the input data preprocessing module is fed to the leading zero detection module 1 and the leading zero detection module 2 for leading zero detection to obtain the number of high-order zeros. This number is then transmitted to the 8-bit addition module and the 24-bit left shift module. The 8-bit addition module adds this data to the exponent bit output by the input data preprocessing module to obtain the normalized exponent bit. The 24-bit left shift module uses this data as the shift value to shift the number of bits output by the input data preprocessing module to obtain the normalized mantissa bit. The sign bit output by the input data preprocessing module is transmitted to the sign bit processing module to calculate the XOR operation result of the sign bits of fa and fb to obtain the sign bit result of the division. The sign bit of the square root is directly obtained from the sign bit of fa. The special data detection module performs positive 0, negative 0, positive infinity, negative infinity, QNaN, and SNaN detection based on the sign bit, exponent bit, and mantissa bit of the input data preprocessing module to obtain special data indication signals.

[0128] The floating-point division enable signal, the floating-point square root enable signal, the normalized exponent and mantissa, the sign bit, and the special data indicator signal are all registered in the first-stage DFF, waiting for the next stage of pipeline operation.

[0129] Phase Two: Iteration Phase

[0130] The calculation is then divided into three paths: sign bit calculation, exponent bit calculation, and mantissa bit calculation. The sign bit calculation is completed in the third stage, while the exponent bit calculation and mantissa bit calculation are completed in the second stage.

[0131] For exponent calculation, the exponent operation module is used to obtain the result of floating-point division: fa exponent plus fb exponent plus 63. The square root exponent operation is the result of shifting fa exponent one bit to the right, adding the least significant bit of fa exponent, and then adding 127. Since the least significant bit is processed when it is odd, the mantissa processing module is used to shift the mantissa of fa one bit to the left based on the least significant bit of fa exponent to obtain the accurate exponent.

[0132] For mantissa calculations, iterative units are used, specifically including iterative units for floating-point mantissa square root operations and floating-point mantissa division operations. Both are implemented using the same main framework. The 29-bit adder 1, 29-bit adder 2, quotient operation module, iteration control module, rounding bit merging module, and the second to fifteenth levels of DFFs reuse resources, having the same functionality but using different operands, as detailed below:

[0133] For floating-point square roots, the normalized mantissa is transmitted to the mantissa processing module, where it undergoes shifting and adjustment before entering the fast calculation module. The high 2 bits are used to generate a carry and 2 data bits, which are then transmitted to 29-bit adder 1 and data processing module 1 respectively. Simultaneously, the low 2 bits are used to generate a carry and 2 data bits, which are transmitted to 29-bit adder 2 and data processing module 2 respectively. Since this is the first time entering the loop, the 29 bits of 0 are transmitted via MUX1 to 29-bit adder 1 and rounding bit calculation module, and the 29 bits of 1 are transmitted via MUX2 to the same module. The rounding bit calculation module derives the rounding bit based on the operands and transmits it to the rounding bit merging module. After bit adder 1 completes the addition operation, it obtains the carry and Sum. The carry is transmitted to the quotient operation module to obtain the first quotient data. At the same time, the carry is transmitted to MUX4 to select one from the output of the quotient operation module and the output of the inverting module 2, and then transmits it to 29-bit adder 2. Sum is transmitted to data processing module 1 to complete the bit concatenation operation, and then transmitted to 29-bit adder 2. After bit adder 2 completes the addition operation, it obtains the carry and Sum. The carry is transmitted to the quotient operation module to obtain the first quotient data, and Sum is transmitted to data processing module 2 to complete the bit concatenation operation.

[0134] The iteration control module is activated, and the counter is 0. When one of the floating-point division operation enable signals or floating-point square root operation enable signals from the enable control module in the first-level DFF register is valid, the carry outputs of the enable control module, exponentiation operation module, data processing module 2, quotient operation module, iteration control module, and 29-bit adder 2 in the first-level DFF register are registered to the second-level DFF. This completes one cycle. Afterward, the rounding bit merging module completes the rounding bit and quotient merging, and outputs it as the mantissa. The next iteration process begins. The output of data processing module 2 in the second-level DFF register is fed into 29-bit adder 1 and rounding bit calculation module through MUX1. The carry generated by 29-bit adder 2 in the second-level DFF register is used as a selection signal to select the result of quotient operation module in the second-level DFF register or the result of inversion logic operation after quotient operation module is transmitted to inversion module 2. This result is then transmitted to 29-bit adder 1 and rounding bit calculation module through MUX2. The rounding calculation module calculates the rounding bits based on the operands and transmits them to the rounding bit merging module. The carry and 2-bit data generated by the fast calculation module are transmitted to the 29-bit adder 1 and data processing module 1, respectively. The operation is then the same as above until the second iteration is completed. When one of the floating-point division operation enable signal or floating-point square root operation enable signal from the enable control module of the previous level DFF is valid, the carry outputs of the enable control module, the exponent operation module, the data processing module 2, the quotient operation module, the iteration control module, and the 29-bit adder 2 are registered in the third level DFF. When invalid, the data remains unchanged. The operation is then the same as above until the 14th iteration. The iteration control module outputs an iteration end signal, which is valid and registered in the fifteenth level DFF. In the next cycle, the iteration control module sets the internal counter value to 0 and waits for the next calculation. The above outputs the exponent result, the mantissa result, and the iteration end signal.

[0135] For floating-point division, the normalized mantissa is processed by the operand 1 preprocessing module and the operand 2 preprocessing module to complete data concatenation. Since it is the first time entering the loop, the data is transmitted through MUX5 and MUX6 to the 29-bit adder 1 to perform addition and obtain the carry and Sum. The carry signal is transmitted to the quotient operation module to obtain the first quotient data. Sum is transmitted to the data processing module 3 to complete the bit concatenation operation, and then transmitted to the 29-bit adder 2 and the 29-bit comparator. The most significant bit of Sum is transmitted to MUX7 as a selection signal to transmit the output data of the operand 2 preprocessing module to the 29-bit adder 2.

[0136] After the 29-bit adder 2 completes the addition operation, it obtains the carry and Sum. The carry signal is transmitted to the quotient operation module to obtain the second quotient data. Sum is transmitted to the data processing module 4 to complete the bit concatenation operation, and then transmitted to the 29-bit comparator to complete the comparison operation and obtain the rounding bit. The iteration control module is activated, and the counter is 0. When one of the floating-point division operation enable signal or the floating-point square root operation enable signal from the enable control module in the first-level DFF register is valid, the outputs of the enable control module, the exponentiation operation module, the data processing module 4, the quotient operation module, the iteration control module, and the 29-bit adder 2 carry are registered in the second-level DFF. This completes one cycle. After that, the rounding bit merging module completes the merging of the rounding bit and the quotient, and outputs it as the mantissa. The next iteration process begins. The output result of the data processing module 4 registered in the second-level DFF is transmitted to the 29-bit adder 4 through the MUX5. Adder 1, and simultaneously the carry stored in the second-level DFF is transmitted to MUX6 as a selection signal, transmitting the data of the operand 2 preprocessing module to 29-bit adder 1. The operation then proceeds as described above until the second iteration is completed. When one of the floating-point division or floating-point square root operation enable signals from the enable control module of the previous-level DFF is valid, the carry outputs of the enable control module, the exponent operation module, data processing module 2, quotient operation module, iteration control module, and 29-bit adder 2 are stored in the third-level DFF. When invalid, the data remains unchanged. The operation then proceeds as described above until the 14th iteration. The iteration control module outputs an iteration end signal, which is valid and stored in the fifteenth-level DFF. In the next cycle, the iteration control module sets its internal counter value to 0, waiting for the next operation. The above outputs the exponent result, mantissa result, and iteration end signal.

[0137] Phase 3: Final Data Processing Phase

[0138] For sign bit calculation, the sign bit calculation module is used. Based on the floating-point division operation enable signal and the floating-point square root operation enable signal, special data is processed first. When there is no special data, normal data is processed to obtain the sign bit data.

[0139] When the calculated exponent is less than 0, the denormalization module is used to perform a shift operation and output the denormalized number data to the preprocessing module. Then, the preprocessing module processes the special data according to the special data indication signal and enable control signal generated in the first stage to obtain the exponent and mantissa results and the NV exception flag. If it is not a special value, the exponent and mantissa generated in the second stage and the exponent and mantissa results from the denormalization module are used to perform boundary value processing and output the exponent, mantissa, rounding bits and OF and UF exception flags of the first result processing.

[0140] The result then enters the post-processing module, which performs five rounding operations based on the mantissa obtained from the pre-processing module, the rounding bits stored in the first-level DFF register, and the sign bit generated by the sign bit calculation module. If there is a carry, the exponent is incremented by 1. The rounded result is then subjected to boundary checks to determine if there is overflow or underflow, and an exception flag, exponent, and mantissa are output. Finally, the result concatenation module merges the exponent, mantissa, and exception flag generated by the pre-processing module, as well as the exception flag, exponent, and mantissa output by the post-processing module. The exponent, mantissa, and the output of the sign bit calculation module are concatenated to obtain the final single-precision floating-point division and square root operation results, along with the five exception flags. Finally, these results, along with the control enable, are registered in the last-level DFF.

[0141] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A hardware calculation method for floating-point division and square root, characterized in that, The process includes the following: Phase 1: Data Preprocessing The input data consists of the opcode, opcode enable, rounding mode, operand fa, and operand fb, all of which are stored in the input DFF. The opcode and opcode validity enable stored in the DFF register are entered into the enable control module to obtain the current floating-point division operation enable signal and floating-point square root operation enable signal; in rounding mode, it directly enters the first-level DFF; operands fa and fb enter the input data preprocessing module to split the data and obtain the sign bit, exponent bit, and mantissa bit of operands fa and fb. At the same time, it judges the output of the enable control module. If all are invalid, the sign bit, exponent bit, and mantissa bit are set to 0. The mantissa bit output by the input data preprocessing module is fed to the leading zero detection module 1 and the leading zero detection module 2 for leading zero detection to obtain the number of high-order zeros. This number is then transmitted to the 8-bit addition module and the 24-bit left shift module. The 8-bit addition module adds this data to the exponent bit output by the input data preprocessing module to obtain the normalized exponent bit. The 24-bit left shift module uses this data as the shift value to shift the number of bits output by the input data preprocessing module to obtain the normalized mantissa bit. The sign bit output by the input data preprocessing module is transmitted to the sign bit processing module to calculate the XOR operation result of the sign bits of fa and fb to obtain the sign bit result of the division. The sign bit of the square root is directly obtained from the sign bit of fa. The special data detection module performs positive 0, negative 0, positive infinity, negative infinity, QNaN, and SNaN detection based on the sign bit, exponent bit, and mantissa bit of the input data preprocessing module to obtain special data indication signals. The floating-point division enable signal, the floating-point square root enable signal, the normalized exponent and mantissa, the sign bit, and the special data indicator signal are all registered in the first-stage DFF, waiting for the next stage of pipeline operation. Phase Two: Iteration Phase The calculation is then divided into three paths: sign bit calculation, exponent bit calculation, and mantissa bit calculation. The sign bit calculation is completed in the third stage, while the exponent bit calculation and mantissa bit calculation are completed in the second stage. For exponent calculation, the exponent operation module is used to obtain the result of floating-point division: fa exponent plus fb exponent plus 63. The square root exponent operation is the result of shifting fa exponent one bit to the right, adding the least significant bit of fa exponent, and then adding 127. Since the least significant bit is processed when it is odd, the mantissa processing module is used to shift the mantissa of fa one bit to the left based on the least significant bit of fa exponent to obtain the accurate exponent. For mantissa calculations, iterative units are used, specifically including iterative units for floating-point mantissa square root operations and floating-point mantissa division operations. Both are implemented using the same main framework. The 29-bit adder 1, 29-bit adder 2, quotient operation module, iteration control module, rounding bit merging module, and the second to fifteenth levels of DFFs reuse resources, having the same functionality but using different operands, as detailed below: For floating-point square roots, the normalized mantissa is transmitted to the mantissa processing module, where it is shifted and adjusted before entering the fast calculation module. The high 2 bits are used to generate the carry and 2 data bits, which are then transmitted to 29-bit adder 1 and data processing module 1, respectively. Simultaneously, the low 2 bits are used to generate the carry and 2 data bits, which are transmitted to 29-bit adder 2 and data processing module 2, respectively. Since this is the first time entering the loop, the 29 bits of 0 are transmitted through MUX1 to 29-bit adder 1 and rounding bit calculation module, and the 29 bits of 1 are transmitted through MUX2 to 29-bit adder 1 and rounding bit calculation module. The rounding calculation module derives the rounding bit based on the operands and transmits it to the rounding bit merging module. After the 29-bit adder 1 completes the addition operation, it obtains the carry and Sum. The carry is transmitted to the quotient operation module to obtain the first quotient data. At the same time, the carry is transmitted to MUX4 to select one from the output of the quotient operation module and the output of the inverting module 2, and then transmits it to the 29-bit adder 2. Sum is transmitted to the data processing module 1 to complete the bit concatenation operation, and then transmitted to the 29-bit adder 2. After the 29-bit adder 2 completes the addition operation, it obtains the carry and Sum. The carry is transmitted to the quotient operation module to obtain the first quotient data, and Sum is transmitted to the data processing module 2 to complete the bit concatenation operation. The iteration control module is activated, and the counter is 0. When one of the floating-point division operation enable signal or the floating-point square root operation enable signal output from the enable control module in the first-level DFF register is valid, the carry output of the enable control module, exponentiation operation module, data processing module 2, quotient operation module, iteration control module, and 29-bit adder 2 in the first-level DFF register is registered into the second-level DFF. This completes one cycle. After that, the rounding bit merging module completes the merging of the rounding bit and the quotient value, and outputs it as the mantissa. The next iteration process begins. The output of data processing module 2 in the second-level DFF register is sent to 29-bit adder 1 and rounding bit calculation module through MUX1. The carry generated by 29-bit adder 2 in the second-level DFF register is used as a selection signal to select the result of quotient operation module in the second-level DFF register or the result of inversion logic operation after quotient operation module is transmitted to inversion module 2 and output through MUX2. The rounding calculation module derives the rounding bit based on the operand and transmits it to the rounding bit merging module. The carry and 2-bit data generated by the fast calculation module are transmitted to 29-bit adder 1 and data processing module 1, respectively. The operation is then the same as above until the second iteration is completed. When one of the floating-point division operation enable signal or floating-point square root operation enable signal from the enable control module of the previous level DFF is valid, the carry outputs of the enable control module, the exponent operation module, the data processing module 2, the quotient operation module, the iteration control module, and 29-bit adder 2 are registered in the third level DFF. When invalid, the data remains unchanged. The operation is then the same as above until the 14th iteration. The iteration control module outputs an iteration end signal, which is valid and registered in the fifteenth level DFF. In the next cycle, the iteration control module sets its internal counter value to 0 and waits for the next calculation. The above outputs the exponent result, the mantissa result, and the iteration end signal. For floating-point division, the normalized mantissa is processed by the operand 1 preprocessing module and the operand 2 preprocessing module to complete data concatenation. Since it is the first time entering the loop, the data is transmitted through MUX5 and MUX6 to the 29-bit adder 1 to perform addition and obtain the carry and Sum. The carry signal is transmitted to the quotient operation module to obtain the first quotient data. Sum is transmitted to the data processing module 3 to complete the bit concatenation operation, and then transmitted to the 29-bit adder 2 and the 29-bit comparator. The most significant bit of Sum is transmitted to MUX7 as a selection signal to transmit the output data of the operand 2 preprocessing module to the 29-bit adder 2. After the 29-bit adder 2 completes the addition operation, it obtains the carry and Sum. The carry signal is transmitted to the quotient operation module to obtain the second quotient data. Sum is transmitted to the data processing module 4 to complete the bit concatenation operation, and then transmitted to the 29-bit comparator to complete the comparison operation and obtain the rounding bit. The iteration control module is activated, and the counter is 0. When one of the floating-point division operation enable signal or the floating-point square root operation enable signal from the enable control module in the first-level DFF register is valid, the outputs of the enable control module, the exponentiation operation module, the data processing module 4, the quotient operation module, the iteration control module, and the 29-bit adder 2 carry are registered in the second-level DFF. This completes one cycle. After that, the rounding bit merging module completes the merging of the rounding bit and the quotient, and outputs it as the mantissa. The next iteration process begins. The output result of the data processing module 4 registered in the second-level DFF is transmitted to the 29-bit adder 4 through the MUX5. Adder 1, and simultaneously the carry stored in the second-level DFF is transmitted to MUX6 as a selection signal, transmitting the data of the operand 2 preprocessing module to 29-bit adder 1. The operation then proceeds as described above until the second iteration is completed. When one of the floating-point division or floating-point square root operation enable signals from the enable control module of the previous-level DFF is valid, the carry outputs of the enable control module, the exponent operation module, data processing module 2, quotient operation module, iteration control module, and 29-bit adder 2 are stored in the third-level DFF. When invalid, the data remains unchanged. The operation then proceeds as described above until the 14th iteration. The iteration control module outputs an iteration end signal, which is valid and stored in the fifteenth-level DFF. In the next cycle, the iteration control module sets its internal counter value to 0, waiting for the next operation. The above outputs the exponent result, mantissa result, and iteration end signal. Phase 3: Final Data Processing Phase For sign bit calculation, the sign bit calculation module is used. Based on the floating-point division operation enable signal and the floating-point square root operation enable signal, special data is processed first. When there is no special data, normal data is processed to obtain the sign bit data. When the calculated exponent is less than 0, the denormalization module is used to perform a shift operation and output the denormalized number data to the preprocessing module. Then the preprocessing module processes the special data according to the special data indication signal and enable control signal generated in the first stage to obtain the exponent and mantissa results and the NV abnormality flag. If it is not a special value, it performs boundary value processing according to the exponent and mantissa generated in the second stage and the exponent and mantissa results from the denormalization processing module, and outputs the exponent, mantissa, rounding bits and OF and UF abnormality flags of the first result processing. The result then enters the post-processing module, which performs five rounding operations based on the mantissa obtained from the pre-processing module, the rounding bits stored in the first-level DFF register, and the sign bit generated by the sign bit calculation module. If there is a carry, the exponent is incremented by 1. The rounded result is then subjected to boundary checks to determine if there is overflow or underflow, and an exception flag, exponent, and mantissa are output. Finally, the result concatenation module merges the exponent, mantissa, and exception flag generated by the pre-processing module, as well as the exception flag, exponent, and mantissa output by the post-processing module. The exponent, mantissa, and the output of the sign bit calculation module are concatenated to obtain the final single-precision floating-point division and square root operation results, along with the five exception flags. Finally, these results, along with the control enable, are registered in the last-level DFF.

2. A hardware computing device for floating-point division and square root calculation, employing the method as described in claim 1, characterized in that, It adopts a 16-stage production line structure, divided into three parts; The first part is the data preprocessing section, which has one pipeline stage, including the input DFF, enable control module, input data preprocessing module, leading zero detection module 1, leading zero detection module 2, 8-bit addition module, 24-bit left shift module, sign bit processing module, special data detection module and first-stage DFF; The second part is the iteration section, which consists of 14 pipeline stages and is used to process mantissa division and square root iteration operations and to obtain exponent results; the second stage includes the exponent operation module and the iteration unit. The third part is the final data processing section, which consists of one pipeline level and is used for special data processing, denormalization processing, five types of rounding, normalization, and five types of exception flag processing. The third stage includes a denormalization processing module, a sign bit calculation module, a preprocessing module, a postprocessing module, a result splicing module, and the final DFF stage.

3. The floating-point division and square root calculation device according to claim 2, characterized in that, The iteration unit includes a floating-point mantissa square root operation iteration unit and a floating-point mantissa division operation iteration unit. The floating-point mantissa square root operation iteration unit and the floating-point mantissa division operation iteration unit share a 29-bit adder 1, a 29-bit adder 2, a quotient operation module, a rounding bit merging module, an iteration control module, and second to fifteenth level DFFs.

4. The floating-point division and square root calculation device according to claim 3, characterized in that, The floating-point mantissa square root operation iteration unit also includes a mantissa processing module, a fast calculation module, an inversion module 1, an inversion module 2, MUX1, MUX2, MUX3, MUX4, a data processing module 1, a data processing module 2, and a rounding bit calculation module.

5. A floating-point division and square root calculation device according to claim 3, characterized in that, The floating-point mantissa division operation iteration unit also includes an operand 1 preprocessing module, an operand 2 preprocessing module, MUX5, MUX6, MUX7, a data processing module 3, a data processing module 4, and a 29-bit comparator.

Citation Information

Patent Citations

  • Circuit for dual-mode floating-point division by square root

    CN109298848A