Improved implementation of a power-like nonlinear function

CN120892012BActive Publication Date: 2026-09-25CLP KESHENTAI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511314144.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-09-25
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

LUT存储的采样点越多,采样点之间的距离越小,引入的近似误差就越小,但额外的存储开销也越大;反之亦然

Benefits of technology

[0017]本发明的上述技术方案相比现有技术具有以下优点:本发明所述的类幂非线性函数实现方法,通过解耦非线性函数的符号位、指数位以及尾数位的计算,将非线性函数的整个定义域映射到一个较小的子集上,能够更有效地实现LUT;且仅在计算非线性函数尾数位近似值时查表,仅需要输入值尾数位以及额外的一位查表,降低了地址位宽;将输入值尾数位构建为一个新的输入值,这个新的输入值一定位于实现LUT的区间内,避免了输入上溢出或者下溢出的场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892012B_ABST
    Figure CN120892012B_ABST
Patent Text Reader

Abstract

The present application relates to an improved implementation method of a class of power-like nonlinear functions. The method significantly compresses the interval required by the lookup table (LUT) by decoupling the calculation of the sign bit, the exponent bit and the mantissa bit of the function value. Firstly, the exponent bit and the sign bit of the function value are determined directly from the exponent bit and the sign bit of the input value by mathematical formula. Then, the input mantissa bit is constructed as a new input value, which is mapped to a predetermined compressed interval. A LUT is established on the interval, and the approximate value of the mantissa bit is calculated by linear interpolation. Finally, the final result is obtained by combining the three parts. The present application is applicable to functions such as square root, reciprocal and reciprocal square root, and supports IEEE 754 format floating point number processing. The advantages of the present application are that the LUT storage overhead is greatly reduced, the interpolation error is reduced, and the operation overflow is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital circuit technology, and in particular to an improved method for implementing power-like nonlinear functions. Background Technology

[0002] Nonlinear functions play an irreplaceable role in mathematical modeling, signal processing, artificial intelligence, and engineering control. They can describe the dynamic behavior of complex systems; for example, activation functions in neural networks introduce expressive power through nonlinear transformations, enabling models to fit highly nonlinear data relationships. However, the computation of nonlinear functions typically involves complex operations such as exponential, logarithmic, and trigonometric functions. Direct real-time computation consumes significant computational resources, increasing system latency and hardware costs. To balance computational accuracy and efficiency, lookup tables (LUTs) have become a widely used and efficient method.

[0003] The core idea of ​​a LUT is to pre-store a series of sampling points for a nonlinear function. During actual computation, the input value is quickly looked up in the table, and an interpolation method is used to calculate an approximate result. In existing technologies, the input value is typically looked up by two adjacent sampling points on the left and right sides of the x-axis. These two sampling points are used to construct a linear function, and the input value is substituted into this linear function to calculate an approximate value of the nonlinear function.

[0004] This method significantly reduces the resources required for nonlinear function computation, making it particularly suitable for resource-constrained scenarios such as FPGAs or embedded systems. However, it also introduces additional storage overhead and a certain amount of approximation error. The more sampling points a LUT stores and the smaller the distance between the sampling points, the smaller the introduced approximation error, but the greater the additional storage overhead; conversely, the smaller the distance between sampling points and the greater the additional storage overhead.

[0005] Furthermore, for nonlinear functions with very large domains or those that vary drastically within a certain interval, balancing storage overhead and approximation error is extremely difficult, because a small number of sampling points cannot capture the characteristics of the entire nonlinear function.

[0006] Therefore, there is an urgent need for an implementation method to reduce the approximation error of power-like nonlinear function LUT results in hardware with limited storage resources. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides an improved method for implementing power-like nonlinear functions. In the hardware circuit, the power-like nonlinear function is efficiently implemented by decoupling the calculation of the sign bit, exponent bit, and mantissa bit of the function value; the method includes the following steps: Step S1: Receive a floating-point input value conforming to the IEEE 754 standard and separate its sign bit, exponent bit, and mantissa bit; Step S2: Through mathematical formula derivation, directly calculate the exponent of the function value based on the exponent of the input value, and determine the sign of the function value based on the sign of the input value; Step S3: Based on the last few digits of the input value, construct a new input value located within a predetermined compression range by concatenating a fixed prefix or logical shifting; Step S4: In the lookup table (LUT) pre-built on the predetermined compression interval, find two sampling points adjacent to the new input value and their corresponding function values; Step S5: Using linear interpolation, calculate the approximate value of the last digit of the function value based on the new input value and the two adjacent sampling points; Step S6: Combine the approximate values ​​of the sign bit, exponent bit, and mantissa bit of the calculated function value to form the final function approximation result; The predetermined compression interval is much smaller than the original domain of the power-like nonlinear function.

[0008] In one embodiment of the present invention, the power-like nonlinear function includes a reciprocal function. Square root function and the reciprocal square root function , where x is the input value.

[0009] In one embodiment of the present invention, in step S2, the operation of calculating the exponent of the function value is implemented by an adder and a shifter.

[0010] In one embodiment of the present invention, in step S3, the method for constructing the new input value is to append a prefix value of fixed width, determined by the parity of the exponent of the input value, to the mantissa of the input value.

[0011] In one embodiment of the present invention, the lookup table (LUT) is stored in a RAM with a depth equal to the number of sampling points N and a width equal to the mantissa of the function value.

[0012] In one embodiment of the present invention, the process of finding adjacent sampling points in step S4 is implemented by a judgment logic circuit. The logic circuit determines the interval segment in which the new input value is located by comparing the new input value with the stored interval boundary value, and outputs the addresses of the sampling points at both ends of the interval segment in RAM.

[0013] In one embodiment of the present invention, the read logic of the RAM is as follows: based on the address output by the judgment logic circuit, the two adjacent sampling points and their corresponding function values ​​are read out simultaneously and input to the linear interpolation operation module.

[0014] In one embodiment of the present invention, in step S5, the linear interpolation formula is as follows: in For the new input value, , They are respectively Two adjacent sampling points, and ≥ , < .

[0015] In one embodiment of the present invention, the computation module for performing the linear interpolation includes: A divider for calculating the slope: ; A multiplier is used to calculate the product of the slope and the offset: ; At least one adder is used to calculate the difference in function values, the difference in sample points, and the final sum-product: , , , .

[0016] In one embodiment of the present invention, it is applicable to processing scalar numbers, non-scalar numbers, and zero input values, and avoids overflow or underflow during the calculation process by constructing new input values.

[0017] Compared with the prior art, the above-mentioned technical solution of the present invention has the following advantages: The power-like nonlinear function implementation method of the present invention decouples the calculation of the sign bit, exponent bit and mantissa bit of the nonlinear function, maps the entire domain of the nonlinear function to a smaller subset, and can implement LUT more effectively; and only looks up the table when calculating the approximate value of the mantissa bit of the nonlinear function, only the mantissa bit of the input value and an additional bit are needed to look up the table, which reduces the address bit width; and constructs a new input value by the mantissa bit of the input value, and this new input value must be within the interval of implementing LUT, avoiding the scenario of input overflow or underflow. Attached Figure Description

[0018] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0019] Figure 1 This is a flowchart illustrating the improved method for implementing power-like nonlinear functions according to the present invention.

[0020] Figure 2 This is a schematic flowchart illustrating a specific embodiment of the present invention.

[0021] Figure 3 This is a hardware structure diagram of a specific embodiment of the present invention.

[0022] Figure 4 The power function described in this invention A comparison chart of the absolute error distribution.

[0023] Figure 5 The power function described in this invention A comparison chart of relative error distributions.

[0024] Figure 6 The power function described in this invention A comparison chart of the absolute error distribution.

[0025] Figure 7 The power function described in this invention A comparison chart of relative error distributions.

[0026] Figure 8 This is a schematic diagram illustrating the classification of FP32 floating-point number formats described in this invention.

[0027] Figure 9 Regarding the present invention A diagram showing the comparison of error values.

[0028] Figure 10 Regarding the present invention Error comparison diagram. Detailed Implementation

[0029] like Figures 1-3 As shown, this embodiment provides an improved method for implementing a power-like nonlinear function, including the following steps: Now using power functions This nonlinear function is implemented assuming its data type is FP32. FP32 data consists of 1 sign bit S, 8 exponent bits E, and 23 mantissa bits M.

[0030] Specifically, such as Figure 8 The text lists ten different classifications of FP32; among them, the formula for representing non-standard numbers is... The formula for representing the specification number is: .

[0031] This embodiment only addresses scenarios where the input value is 0, a non-standard number, or a standard number. In these three scenarios, the power function... The function value must be greater than or equal to 0, that is, S is 0. Therefore, the sign bit will be ignored in the following parts to simplify the explanation.

[0032] Scenario 1, input value is 0: .

[0033] Scenario 2, input value is the number of specifications: Input value function value At this time, the function value This refers to the number of specifications.

[0034] Assumption , .

[0035] make ,but , .in, , .

[0036] When the exponent is odd .because , , ,so , .therefore , .

[0037] At this point, the function value sign bit =0; function value exponent By input values exponent Add a constant 127 and then shift right by one bit to obtain the result; by inputting the value mantissa A new input value is constructed by extending the input with a nine-bit binary number 0b001111111. Then to The results were obtained using a lookup table-based method and linear interpolation. , The last 23 bits of the binary representation are the function value. mantissa Therefore, it is only necessary to... range Implement the power function LUT.

[0038] When the exponent is even .because , , ,so , .therefore , .

[0039] At this point, the function value sign bit =0; function value exponent By input values exponent Add a constant 126 and then shift right by one bit to obtain the result; by inputting the value mantissa A new input value is constructed by extending the input with a nine-bit binary number 0b010000000. Then to The results were obtained using a lookup table-based method and linear interpolation. , The last 23 bits of the binary representation are the function value. mantissa Therefore, it is only necessary to... range Implement the power function LUT.

[0040] In summary, in scenario two, it is only necessary to be within the range... Implement the power function LUT, not input value range This greatly reduces the lookup range of the LUT.

[0041] Scenario 3, Input value is a non-standard number: Input value function value At this time, the function value This refers to the number of specifications.

[0042] Assumption , .

[0043] at this time, In order to construct the form of the specification number representation formula, it is necessary to ensure that Even number and ,Right now The specific implementation method is as follows: first, Initialize to 0; then, input value mantissa Perform a left shift; each left shift... The value is incremented by 1; when When the most significant 1 is shifted left to the 24th bit, a judgment is made. Is it an even number?

[0044] if If it is even, then and after the shift A new input value can be constructed by extending the lower 23 bits with a nine-bit binary number 0b001111111. ,Right now The new input value is then processed using a lookup table-based method and linear interpolation to obtain an intermediate result, the lower 23 bits of which are the input value. ;if If it is an odd number, then it is necessary to... Transform into This satisfies the format of the specification number. And after the shift A new input value can be constructed by extending the lower 23 bits with a nine-bit binary number 0b010000000. ,Right now The new input value is then processed using a lookup table-based method and linear interpolation to obtain an intermediate result, the lower 23 bits of which are the input value. .

[0045] In summary, in scenario three, it is only necessary to be within the range... Implement the power function LUT, not input value range This coincides with the LUT lookup interval in Scenario 2.

[0046] Based on the analysis and summary of the above three scenarios, this invention will use the power function. LUT lookup range from Shrink to In a LUT table of the same depth, the lookup step size is less than one percent of the original, which greatly improves the fitting accuracy.

[0047] Figure 4 and Figure 5 The invention and traditional lookup tables were compared in their calculation of power functions. The absolute error curve and relative error curve at that time are shown in the figure. Figure 9 As shown. From Figure 9 From this, it can be concluded that the present invention will use the power function The calculation error was reduced by more than three orders of magnitude, significantly improving the calculation accuracy; and as Figure 9 As shown, The error.

[0048] For power functions Implementation instructions: power function The function value must be greater than 0, meaning S is 0. Therefore, the sign bit will be ignored in the following sections to simplify the explanation.

[0049] Scenario 1, input value is the number of specifications: Input value function value At this time, the function value This refers to the number of specifications.

[0050] Assumption , .

[0051] make ,but , .in, , .

[0052] When the exponent is odd and hour, .because , ,so , .therefore , .

[0053] At this point, the function value sign bit =0; function value exponent By subtracting the input value from a constant 381 exponent Then shift right by one position to obtain the function value. mantissa It is 0.

[0054] When the exponent is odd and hour, .because , , ,so , .therefore , .

[0055] At this point, the function value sign bit =0; function value exponent By subtracting the input value from a constant 379 exponent Then shift right by one bit to obtain the result; by inputting the value... mantissa A new input value is constructed by extending the input with a nine-bit binary number 0b001111101. Then to The results were obtained using a lookup table-based method and linear interpolation. , The last 23 bits of the binary representation are the function value. mantissa Therefore, it is only necessary to... range Implement the power function LUT.

[0056] When the exponent is even .because , , ,so , .therefore , .

[0057] At this point, the function value sign bit =0; function value exponent By subtracting the input value from a constant 380 exponent Then shift right by one bit to obtain the result; by inputting the value... mantissa A new input value is constructed by extending the input with a nine-bit binary number 0b001111110. Then to The results were obtained using a lookup table-based method and linear interpolation. , The last 23 bits of the binary representation are the function value. mantissa Therefore, it is only necessary to... range Implement the power function LUT.

[0058] In summary, in scenario one, it is only necessary to be within the range... Implement the power function LUT, not input value range This greatly reduces the lookup range of the LUT.

[0059] Scenario 2, input value is a non-standard number: Input value function value At this time, the function value This refers to the number of specifications.

[0060] Assumption , .

[0061] at this time, In order to construct the form of the specification number representation formula, it is necessary to ensure that Even number and ,Right now The specific implementation method is as follows: first, Initialize to 0; then, input value mantissa Perform a left shift; each left shift... The value is incremented by 1; when When the most significant 1 is shifted left to the 24th bit, a judgment is made. Is it an even number?

[0062] if Even number and after shift If the lower 23 are all 0, then and ,Right now , ;if Even number and after shift If the lower 23 are not all zero, then Therefore, it is necessary to Transform into Thus ,at this time, And after the shift By extending the lower 23 bits with a nine-bit binary number 0b001111101, a new input value can be constructed. The new input value is then processed using a lookup table-based method and linear interpolation to obtain an intermediate result, the lower 23 bits of which are the input value. ;if If it is an odd number, then it is necessary to... Transform into This satisfies the format of the specification number. And after the shift A new input value can be constructed by extending the lower 23 bits with a nine-bit binary number 0b001111110. ,Right now The new input value is then processed using a lookup table-based method and linear interpolation to obtain an intermediate result, the lower 23 bits of which are the input value. .

[0063] In summary, in scenario two, it is only necessary to be within the range... Implement the power function LUT, not input value range This coincides with the LUT lookup range in Scenario 1.

[0064] Based on the analysis and summary of the above three scenarios, this invention will use the power function. LUT lookup range from Shrink to In a LUT table of the same depth, the lookup step size is less than one percent of the original, which greatly improves the fitting accuracy.

[0065] Figure 6 and Figure 7 The invention and traditional lookup tables were compared in their calculation of power functions. The absolute and relative error curves at different times (for ease of comparison, only the range with larger errors is shown), the results are as follows: Figure 10 As shown. From Figure 10 From this, it can be concluded that the present invention will use the power function The calculation error was reduced by more than seven orders of magnitude, significantly improving the calculation accuracy; for example Figure 10 middle The error is visible.

[0066] In summary, the method described in this embodiment is applicable to functions such as square root, reciprocal, and reciprocal square root, and supports IEEE 754 format floating-point number processing. Its advantages include significantly reducing LUT storage overhead, minimizing interpolation errors, and avoiding computational overflow.

[0067] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. An improved method for implementing a power-like nonlinear function, the method being implemented in FPGA or ASIC hardware circuits, characterized in that, In hardware circuits, a power-like nonlinear function is implemented by decoupling the calculation of the sign bit, exponent bit, and mantissa bit of the function value; this includes the following steps: Step S1: Receive a floating-point input value conforming to the IEEE 754 standard and separate its sign bit, exponent bit, and mantissa bit; Step S2: Through mathematical formula derivation, the exponent of the function value is directly calculated based on the exponent of the input value, and the sign of the function value is determined based on the sign of the input value; in step S2, the operation of calculating the exponent of the function value is implemented by an adder and a shifter; Step S3: Based on the last few digits of the input value, construct a new input value located within a predetermined compression range by concatenating a fixed prefix or logical shifting; Step S4: In the lookup table pre-built on the predetermined compression interval, find two sampling points adjacent to the new input value and their corresponding function values; the process of finding adjacent sampling points in step S4 is implemented by a judgment logic circuit, which determines the interval segment in which the new input value is located by comparing the new input value with the stored interval boundary value, and outputs the addresses of the sampling points at both ends of the interval segment in RAM; Step S5: Using linear interpolation, calculate the approximate value of the function value's last digits based on the new input value and two adjacent sampling points; Step S6: Combine the approximate values ​​of the sign bit, exponent bit, and mantissa bit of the calculated function value to form the final function approximation result; The predetermined compression interval is much smaller than the original domain of the power-like nonlinear function.

2. The method for implementing a power-like nonlinear function according to claim 1, characterized in that: The power-like nonlinear function includes the reciprocal function. Square root function and the reciprocal square root function , where x is the input value.

3. The method for implementing a power-like nonlinear function according to claim 1, characterized in that: In step S3, the method for constructing the new input value is to append a prefix value with a fixed width determined by the parity of the exponent of the input value to the last few digits of the input value.

4. The method for implementing a power-like nonlinear function according to claim 1, characterized in that: The lookup table is stored in a RAM with a depth equal to the number of sampling points N and a width equal to the mantissa of the function value.

5. The method for implementing a power-like nonlinear function according to claim 1, characterized in that: The read logic of the RAM is as follows: based on the address output by the judgment logic circuit, the two adjacent sampling points and their corresponding function values ​​are read out simultaneously and input to the linear interpolation operation module.

6. The method for implementing a power-like nonlinear function according to claim 1, characterized in that: The computation module for performing the linear interpolation includes: A divider for calculating the slope: ; A multiplier is used to calculate the product of the slope and the offset: ; At least three adders are used to calculate the difference in function values, the difference in sample points, and the final sum-product: , , , .

7. The method for implementing an exponential nonlinear function according to claim 1, characterized in that: It is suitable for handling standard numbers, non-standard numbers, and zero input values, and avoids overflow or underflow during the calculation process by constructing the new input value.

Citation Information

Patent Citations

  • Nonlinear layer data processing method and device and storage medium

    CN117332196A

  • Systems and methods for accelerating the computation of the reciprocal function and the reciprocal-square-root function

    US20230161554A1