Single-precision floating-point Nth square root calculation architecture, method and system based on piecewise quadratic polynomial approximation

Through the single-precision floating point number N-order root calculation architecture based on segmented quadratic polynomial approximation, the problem of high hardware resource consumption is solved, and the calculation effects of high precision, low latency and low power consumption are achieved.

CN115495046BActive Publication Date: 2025-07-22NANJING UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210943023.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2025-07-22
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

The prior art has problems of large hardware resource consumption and long delay in the calculation of N-order root numbers in floating point numbers. Especially when using the segmented quadratic polynomial approximation method, the size of the lookup table becomes a performance limiting factor.

Method used

A single-precision floating-point number N-order root number calculation architecture based on segmented quadratic polynomial approximation is adopted, including log2 and exp2 segmented quadratic polynomial approximation modules. Through fine-grained segmentation and basic operation modules, hardware resource occupation is reduced, and the calculation process is coordinated through the control module.

Benefits of technology

It realizes high-precision, low latency and low power consumption, reduces resource overhead and optimizes computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115495046B_ABST
    Figure CN115495046B_ABST
Patent Text Reader

Abstract

The present invention relates to a single-precision floating-point Nth root calculation architecture, method and system based on piecewise quadratic polynomial approximation, including: a log2 piecewise quadratic polynomial approximation module that calculates the result of the logarithmic function with any true number as the base 2 through piecewise quadratic polynomial approximation; an exp2 piecewise quadratic polynomial approximation module that calculates the result of the exponential function with any exponent as the base 2 through piecewise quadratic polynomial approximation; a basic operation module includes a floating-point conversion unit, an addition unit, a lookup table unit and a multiplication unit; a control module controls the overall calculation process and outputs the result by calling each module and calculation unit. The present invention can simultaneously meet the requirements of high precision, low latency, low resource occupancy rate and low power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of hardware implementation of function calculation, and specifically to a single-precision floating-point Nth root calculation architecture based on piecewise quadratic polynomial approximation. Background Art

[0002] The Nth root calculation of floating-point numbers has wide applications in various fields: digital signal processing, three-dimensional graphics systems, and so on. Although the accuracy is very high, the software processing function is very time-consuming, which makes researchers turn to hardware implementation to obtain better performance. In the past few decades, several high-speed or high-precision methods have been proposed, and these methods can be divided into the following four types: Newton-Raphson method, digital recursive algorithm, CORDIC-based method, and polynomial approximation method.

[0003] Piecewise quadratic polynomial approximation segments the source number into subintervals and uses low-degree polynomials for approximation within the subintervals. It usually uses a look-up table (LUT) to store coefficients, so the size of the table becomes a key factor restricting performance. Lyu et al. proposed a PWL-based Nth root calculation architecture and achieved low latency in half-precision BFP.

[0004] However, this design uses adders to select appropriate segmentation intervals, which greatly consumes hardware resources. Second-order or higher-order piecewise quadratic polynomial approximation has not been explored in the field of floating-point Nth root calculation either. Summary of the Invention

[0005] Object of the Invention: To overcome the deficiencies of the above prior art, a method for hardware implementation of single-precision floating-point Nth root calculation based on piecewise quadratic polynomial approximation is provided, which uses fewer hardware resources while ensuring calculation accuracy.

[0006] Technical Solution: A single-precision floating-point Nth root calculation architecture based on piecewise quadratic polynomial approximation includes the following modules:

[0007] A basic operation module, including a floating-point conversion unit for receiving a floating-point data R and separating its exponent E and mantissa M, a look-up table unit for finding the reciprocal of the root extraction times N, an addition unit, and a multiplication unit that receives the output of the addition unit as an input and outputs an integer part and a positive decimal part;

[0008] A log2 piecewise quadratic polynomial approximation module, which receives the mantissa and calculates the logarithmic function result of the mantissa with base 2 through quadratic polynomial approximation; the logarithmic function result and the exponent are used as the input of the addition unit;

[0009] The exp2 piecewise quadratic polynomial approximation module receives the positive fractional part and calculates the exponential function result with base 2 of the positive fractional part through piecewise quadratic polynomial approximation;

[0010] The core algorithm control module calls each unit in the basic operation module, the log2 piecewise quadratic polynomial approximation module, and the exp2 piecewise quadratic polynomial approximation module to control the overall calculation process and output the result.

[0011] According to one aspect of the present application, the single-precision floating-point Nth root calculation architecture based on piecewise quadratic polynomial approximation includes:

[0012] The log2 piecewise quadratic polynomial approximation module calculates the logarithmic function result with base 2 of any true number through piecewise quadratic polynomial approximation;

[0013] The exp2 piecewise quadratic polynomial approximation module calculates the exponential function result with base 2 of any exponent through piecewise quadratic polynomial approximation;

[0014] The basic operation module includes four basic operation units, namely: a floating-point conversion unit, a lookup table unit, an addition unit, and a multiplication unit;

[0015] The control module calls each module and calculation unit to control the overall calculation process and output the result.

[0016] According to one aspect of the present application, the control module connects each unit in the basic operation module and assigns tasks to the log2 piecewise quadratic polynomial approximation module, the exp2 piecewise quadratic polynomial approximation module, and the lookup table module, and finally outputs the Nth root function operation result.

[0017] According to one aspect of the present application, the process of the computing system calculating the function is as follows:

[0018] Set the function R is represented in floating-point form as R = 1.M * 2 E , M is the mantissa, E is the exponent; N is sent to the lookup table unit to obtain the reciprocal of N; the floating-point conversion unit separates the exponent E and the mantissa M from the floating-point R, and the mantissa M is sent to the log2 piecewise quadratic polynomial approximation module;

[0019] The exponent E and the output result of the log2 piecewise quadratic polynomial approximation module are sent to the addition unit together, and the result is used as a multiplication factor and sent to the multiplication unit together with the output result obtained from the lookup table unit;

[0020] The output result of the multiplication unit is converted into an integer part and a positive decimal part. The integer part is used as the exponent of the final exponential function result, and the positive decimal part is used as the input of the exp2 piecewise quadratic polynomial approximation module; the decimal part of its output result is truncated to the upper 23 bits as the mantissa of the final exponential function result.

[0021] According to one aspect of the present application, the log2 piecewise quadratic polynomial approximation module adopts fine-grained segmentation to divide the input into three segments. The value corresponding to the first segment input is obtained through a lookup table to obtain the constant term coefficient, the linear term coefficient and the quadratic term coefficient of the corresponding quadratic polynomial. The linear term coefficient is multiplied with the values corresponding to the second and third segment inputs respectively, and the two results are added together with the constant term coefficient; the second segment input is squared and multiplied with the quadratic term coefficient. The sum of the three addition results and the multiplication output result is the output of the log2 piecewise quadratic polynomial approximation module.

[0022] According to one aspect of the present application, the exp2 piecewise quadratic polynomial approximation module divides the input into three segments. The value corresponding to the first segment input is obtained through a lookup table to obtain the constant term coefficient, the linear term coefficient and the quadratic term coefficient of the corresponding quadratic polynomial. The linear term coefficient is multiplied by the values corresponding to the second and third segment inputs respectively, and the two results are added together with the constant term coefficient; the second segment input is squared and multiplied by the quadratic term coefficient. The sum of the three addition results and the multiplication output result is the output of the exp2 piecewise quadratic polynomial approximation module.

[0023] According to one aspect of the present application, the addition unit in the basic operation module implements the signed data addition function, which includes two addition factors: the exponent E separated by the floating point conversion unit and the output result of the log2 piecewise quadratic polynomial approximation module.

[0024] According to one aspect of the present application, the multiplication unit in the basic operation module has a floating-point data multiplication function, and includes two multiplication factors: the reciprocal of the number of root operations N of the input obtained from the lookup table and the output result obtained from the addition unit.

[0025] Beneficial effects: The method for implementing the Nth square root of a single-precision floating-point number based on piecewise quadratic polynomial approximation of the present invention is the first to introduce high-order piecewise quadratic polynomial approximation into the study of calculating the Nth square root of a single-precision floating-point number. This method proposes fine-grained segmentation for the first time, greatly reducing resource overhead, and achieving low power consumption and low latency with small errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a module schematic diagram of the N-th square root calculation architecture of a single-precision floating-point number based on piecewise quadratic polynomial approximation of the present invention.

[0027] Figure 2 It is a flowchart for calculating the Nth root function by an Nth root calculation architecture.

[0028] Figure 3 They are the architecture diagrams of the log2 piecewise quadratic polynomial approximation module and the exp2 piecewise quadratic polynomial approximation module.

[0029] Figure 4 It is a graph of the precision performance indicators for calculating the Nth root function.

[0030] Figure 5 They are the graphs of the power consumption and area indicators of the present invention. Detailed implementation manners

[0031] To solve the above problems, the applicant has conducted in-depth research on the prior art. In 2005, F. de Dinechin et al. proposed a multi-part table method, which decomposes a large LUT into several small LUTs, plus a simple multi-operand addition. The extension of this method is achieved by adding small multipliers or additional tables for higher-order approximations. Piecewise linear (PWL) approximation is a first-order piecewise polynomial method that uses linear functions for approximation to quickly calculate the Nth root. The PWL method is widely used in unary functions. In 2016, Liu et al. proposed an error-flat non-uniform region linear approximation algorithm. However, there are still various technical problems with the above algorithms and further improvements are needed.

[0032] The solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0033] As Figure 1 shown, the hardware implementation method for calculating the Nth root of any single-precision floating-point number based on piecewise quadratic polynomial approximation in this example, the hardware mainly includes a core algorithm control module (control module for short), a log2 piecewise quadratic polynomial approximation module, an exp2 piecewise quadratic polynomial approximation module architecture diagram, and a basic operation module. Among them, the basic operation module further includes the following four units: a floating-point conversion unit, a lookup table unit, an addition unit, and a multiplication unit. Figure 1 shows the topological structure of this hardware part. However, it should be noted that in order to more clearly show the design points of the present invention, this figure does not specifically describe relevant details, such as the signal connection part.

[0034] The basic working principle of this application is as follows:

[0035] Generally, for the function (N is a constant and N>1, the domain is the real number R),

[0036] Assume y = 2 S, R can be represented in floating-point form as R = 1.M * 2 E (where M is the mantissa and E is the exponent), then

[0037] Since 1.M ∈ [1, 2), using the log2 piecewise quadratic polynomial approximation module, we get log21.M = P1 ∈ [0, 1), that is, E + log21.M = P2 ∈ [E, E + 1). So (where C is the integer part and D is the positive fractional part), then

[0038] Since D ∈ [0, 1), using the exp2 piecewise quadratic polynomial approximation module, we get 2 D ∈ [1, 2). If we let 2 D = 1.W (where 1 is the integer part and W is the fractional part), then y = 1.W * 2 C .

[0039] In summary, it can be seen that the input is the floating-point number R and the fixed-point number N, and the output is the floating-point number y.

[0040] Next, a detailed description of an example of the present invention will be given, and design verification will be carried out through MATLAB simulation and RTL-level description of Verilog.

[0041] The present invention is based on Figure 1 the hardware architecture shown to calculate the Nth root function. In the example, it is assumed that the range of the root extraction times N ∈ [2, 256], R belongs to the range of single-precision floating-point numbers, and the input R, N, and the output y are all 32-bit data ([a 31, a 30, …, a 2, a0]). For the input R, [a 31 represents the sign bit, [a 30, a 29, …, a 24, a 23 represents the exponent plus the exponent bias 127, [a 22, a 21, …, a 1, a0] represents the mantissa; for the input N, [a 31 represents the sign bit (since N must be greater than 0, this bit is always 0), [a 30, a 29, …, a 24, a 23 represents the exponent, [a 22, a 21, …, a 1, a0] represents the mantissa (since N is an integer, these 23 bits are always 0); for the output y, [a 31represents the sign bit, [a 30, a 29, …,a 24, a 23 represents the exponent, [a 22, a 21, …,a 1, a0] represents the mantissa.

[0042] As Figure 2 shown, the square root order N is transmitted to the lookup table unit to obtain the reciprocal of the square root order N; the floating-point conversion unit separates the exponent E and the mantissa M from the floating-point R, and the mantissa M is transmitted to the log2 piecewise quadratic polynomial approximation module; the exponent E and the output result of the log2 piecewise quadratic polynomial approximation module are transmitted to the adder unit together, and the result is used as a multiplication factor, which is transmitted to the multiplication unit together with the output result obtained from the lookup table unit; the output result of the multiplication unit is converted into an integer part and a positive decimal part, the integer part (plus the exponent bias 127) is used as the exponent of the final exponential function result, and the positive decimal part is used as the input of the exp2 piecewise quadratic polynomial approximation module; the high 23 bits of the fractional part of its output result are intercepted as the mantissa of the final exponential function result.

[0043] Regarding the design of the log2 piecewise quadratic polynomial approximation module, as Figure 3 shown. The input variable is set to 23 bits, the output variable is set to 26 bits, the input is divided into three segments M0, M1, and M2, which are set to 6 bits, 10 bits, and 7 bits respectively. The value corresponding to M0 is transmitted to the lookup table unit to obtain the constant term coefficient W0, the first-order term coefficient W1, and the second-order term coefficient W2 of the corresponding quadratic polynomial.

[0044] Among them, W1 is used as a multiplication factor and is transmitted to the multiplication unit together with M1 and M2 respectively. The output results of the two multiplication units are transmitted to the carry-save adder (CSA) together with W0.

[0045] After M1 is squared, it is transmitted to the multiplication unit together with W2, and the output result is used as a multiplication factor and is transmitted to the adder unit together with the output result of the carry-save adder unit. The output result of the adder unit intercepts the high 26 bits as the output result of the log2 piecewise quadratic polynomial approximation module.

[0046] Regarding the design of the exp2 piecewise quadratic polynomial approximation module, the Figure 3 similar or identical architecture is adopted. The input variable is set to 23 bits, the output variable is set to 26 bits, the input is divided into three segments M0, M1, and M2, which are set to 6 bits, 10 bits, and 7 bits respectively. The value corresponding to M0 is transmitted to the lookup table unit to obtain the constant term coefficient W0, the first-order term coefficient W1, and the second-order term coefficient W2 of the corresponding quadratic polynomial.

[0047] Among them, W1, as a multiplication factor, is transmitted to the multiplication unit together with M1 and M2 respectively. The output results of the two multiplication units are transmitted to the carry-save adder (CSA) together with W0; after M1 performs a squaring operation, it is transmitted to the multiplication unit together with W2, and its output result is used as a multiplication factor and transmitted to the adder unit together with the output result of the carry-save adder unit. The high 26 bits of the output result of the adder unit are intercepted as the output result of the exp2 piecewise quadratic polynomial approximation module.

[0048] According to the above example, 2,550,000 samples are extracted in MATLAB (10,000 floating-point numbers are randomly selected within a certain range of R, and the value of N ranges from 2 to 255) for precision test simulation, and specific precision performance indicators can be obtained as Figure 4 shown. The errors in the figure are all relative errors. After synthesis in the TSMC 40nm process library, the performance indicators as Figure 5 shown can be obtained.

[0049] From Figure 4 and Figure 5 it can be seen that the hardware implementation method of the Nth square root of single-precision floating-point numbers based on piecewise quadratic polynomial approximation not only has high calculation precision, but also occupies less resources and has low power consumption, providing an excellent solution to the problems that the traditional linear approximation method and the CORDIC method occupy too much resources and it is difficult to balance speed and accuracy.

Claims

1. A single-precision floating-point Nth root calculation architecture based on piecewise quadratic polynomial approximation, characterized in that Including: A basic operation module, including a floating-point conversion unit for receiving a floating-point data R and separating its exponent E and mantissa M, a lookup table unit for finding the reciprocal of the root extraction times N, an addition unit, and a multiplication unit for receiving the output of the addition unit as an input and outputting an integer part and a positive fractional part; A log2 piecewise quadratic polynomial approximation module, receiving the mantissa M, and approximately calculating the logarithmic function result of the mantissa with base 2 through a quadratic polynomial; this logarithmic function result and the exponent E serve as the input quantities of the addition unit; An exp2 piecewise quadratic polynomial approximation module, receiving the positive fractional part, and approximately calculating the exponential function result with base 2 of the positive fractional part through a quadratic polynomial; A core algorithm control module, calling each unit in the basic operation module, the log2 piecewise quadratic polynomial approximation module, and the exp2 piecewise quadratic polynomial approximation module, controlling the overall calculation process and outputting a result; The log2 piecewise quadratic polynomial approximation module and the exp2 piecewise quadratic polynomial approximation module adopt fine-grained segmentation: The input is divided into three segments; The value corresponding to the first segment input obtains the constant term coefficient, the first-order term coefficient, and the second-order term coefficient of the corresponding quadratic polynomial through a lookup table; The first-order term coefficients are respectively multiplied by the values corresponding to the second segment and the third segment inputs, and the two results are added together with the constant term coefficient; The second segment input is squared and then multiplied by the second-order term coefficient; the sum of the three is added to the output of this multiplication to obtain the output.

2. The single-precision floating-point Nth root calculation architecture based on piecewise quadratic polynomial approximation according to claim 1, characterized in that The exponent minus the exponent bias is used as one of the input quantities of the addition unit, and the integer part plus the exponent bias is used as the exponent of the final exponential function result.

3. The square root calculation architecture for single-precision floating-point numbers to the Nth power based on piecewise quadratic polynomial approximation according to claim 1, characterized in that, The calculation process is as follows: Set function R is represented in floating-point form as R = 1.M * 2 E , where M is the mantissa and E is the exponent; The root extraction times N is transmitted to the lookup table unit to obtain its reciprocal; the floating-point conversion unit separates the exponent E and the mantissa M from the floating-point R, and the mantissa M is transmitted to the log2 piecewise quadratic polynomial approximation module; The exponent E and the output result of the log2 piecewise quadratic polynomial approximation module are transmitted to the addition unit together, and its result serves as a multiplication factor, and is transmitted to the multiplication unit together with the output result obtained from the lookup table unit; The output result of the multiplication unit is converted into an integer part and a positive fractional part. The integer part serves as the exponent of the final exponential function result, and the positive fractional part serves as the input of the exp2 piecewise quadratic polynomial approximation module; the fractional part of its output result intercepts the high 23 bits as the mantissa of the final exponential function result.

4. The single-precision floating-point Nth root calculation architecture based on piecewise quadratic polynomial approximation according to claim 3, characterized in that The addition unit realizes the function of adding signed data, and includes two addition factors respectively: the exponent separated by the floating-point conversion unit, and the result calculated by the log2 linear piecewise quadratic polynomial approximation module.

5. The single-precision floating-point Nth root calculation architecture based on piecewise quadratic polynomial approximation according to claim 3, characterized in that The multiplication unit has the function of multiplying floating-point data, and includes two multiplication factors respectively: the reciprocal of the fixed-point root extraction times N found from the lookup table unit, and the output result obtained from the addition unit.

6. A method for calculating the Nth square root of a single-precision floating-point number based on piecewise quadratic polynomial approximation, characterized in that, Including the following steps: Set function R is represented in floating point as R = 1.M * 2 E , where M is the mantissa and E is the exponent; Transfer N to the lookup table unit to obtain the reciprocal of N; the floating-point conversion unit separates the exponent E and the mantissa M from the floating-point R, and the mantissa M is transferred to the log2 piecewise quadratic polynomial approximation module; The exponent E and the output result of the log2 piecewise quadratic polynomial approximation module are sent to the addition unit together, and the result is used as a multiplication factor, which is sent to the multiplication unit together with the output result obtained from the lookup table unit; The output result of the multiplication unit is converted into an integer part and a positive fractional part. The integer part is used as the exponent of the final exponential function result, and the positive fractional part is used as the input of the exp2 piecewise quadratic polynomial approximation module; the high 23 bits of the fractional part of the output result are intercepted as the mantissa of the final exponential function result; The exp2 piecewise quadratic polynomial approximation module uses fine-grained segmentation and divides the input into three segments; the values corresponding to the inputs of the first segment are used to obtain the constant term coefficient, the first-order term coefficient, and the second-order term coefficient of the corresponding quadratic polynomial through a lookup table; the first-order term coefficients are multiplied by the values corresponding to the inputs of the second and third segments respectively, and the two results are added together with the constant term coefficient; the input of the second segment is squared and then multiplied by the second-order term coefficient; the sum of the three results is added to the output result of this multiplication to obtain the output of the exp2 piecewise quadratic polynomial approximation module.

7. A system, characterized in that, Comprising the single-precision floating-point Nth root calculation architecture based on piecewise quadratic polynomial approximation according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Geographic information data decryption method

    CN108629190A

  • A calculation system based on a type-2 hyperbolic CORDIC arbitrary exponential function

    CN109739470A

  • Error-unbiased approximate multiplier for normalized floating-point numbers and its implementation method

    JP7016559B1

  • Implementation method and device for calculating sine or cosine function

    WO2022001722A1