Transcendental function processing method, circuit and device, electronic equipment and chip
By using Chebischev and endpoint interpolation methods in polynomial fitting, the processing of transcendent functions is optimized, and the problems of low calculation accuracy and complex hardware implementation in the prior art are solved, and efficient and low-power transcendent functions are realized.
Patent Information
- Application Number
- CN202510097099.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
When processing transcendent functions, the calculation accuracy is not high, the area and power consumption of hardware are large, and performance is limited in high-precision computing scenarios.
By using Chebischev and endpoint interpolation when fitting polynomials, the table entries are reduced, the subtraction operation is optimized, the calculation period is compressed, and the software iterative solution to the sub-interval division is further compressed.
The accuracy of polynomial fitting to transcendent functions is improved, the bit width of lookup tables, multipliers, and accumulators during hardware implementation is reduced, and it is compatible with various types of operators, which improves the computing performance of transcendent functions, and optimizes the area and performance of the computing circuit.
Smart Images

Figure CN120045015A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of chip design, and particularly to a transcendental function processing method, circuit, device, electronic device, and chip. Background Art
[0002] Transcendental Functions refer to functions whose relationships between variables cannot be expressed by a finite number of addition, subtraction, multiplication, division, exponentiation, and root extraction operations, such as exponential functions, logarithmic functions, trigonometric functions, etc. Transcendental functions are widely used in various practical problems such as signal processing, control system design, and structural optimization. Therefore, optimizing the processing performance of transcendental functions plays an important role in the technological development of various fields. Summary of the Invention
[0003] The present disclosure aims to solve at least one of the technical problems in the related art to some extent.
[0004] A first aspect embodiment of the present disclosure provides a transcendental function processing method, including:
[0005] Determining a transcendental function to be processed and function description information;
[0006] Based on the function description information, fitting the transcendental function to determine a plurality of target sub-intervals corresponding to the transcendental function and polynomial coefficients corresponding to each target sub-interval;
[0007] Generating a lookup table according to the interval range and polynomial coefficients corresponding to each target sub-interval;
[0008] Determining configuration parameters of each circuit component in a calculation circuit for processing the transcendental function according to the polynomial coefficients and the lookup table.
[0009] A second aspect embodiment of the present disclosure provides a transcendental function processing circuit, including:
[0010] An address decoder, a lookup table, and an arithmetic component;
[0011] Wherein, the address decoder is configured to parse the input data to determine the interval to which the input data belongs; according to the interval to which the input data belongs, determine a target address and data to be operated, and obtain target polynomial coefficients from the lookup table based on the target address;
[0012] The arithmetic component is configured to process the data to be operated based on the target polynomial coefficients to obtain an operation result corresponding to the input data.
[0013] A third aspect embodiment of the present disclosure provides a transcendental function processing device, including:
[0014] A first determination module, configured to determine a transcendental function to be processed and function description information;
[0015] A second determination module, configured to fit the transcendental function based on the function description information, determine a plurality of target sub-intervals corresponding to the transcendental function, and polynomial coefficients corresponding to each of the target sub-intervals;
[0016] A generation module, configured to generate a look-up table according to the interval range and polynomial coefficients corresponding to each of the target sub-intervals;
[0017] A third determination module, configured to determine configuration parameters of each circuit component in a calculation circuit for processing the transcendental function according to the polynomial coefficients and the look-up table.
[0018] An embodiment of the fourth aspect of the present disclosure provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the transcendental function processing method proposed in the embodiment of the first aspect of the present disclosure is implemented.
[0019] An embodiment of the fifth aspect of the present disclosure provides a chip, where the chip includes a processing circuit and an interface circuit; wherein, the interface circuit is configured to obtain an instruction and send the instruction to the processing circuit, and the processing circuit is configured to execute the instruction to implement the transcendental function processing method proposed in the embodiment of the first aspect of the present disclosure.
[0020] An embodiment of the sixth aspect of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, the transcendental function processing method proposed in the embodiment of the first aspect of the present disclosure is implemented.
[0021] An embodiment of the seventh aspect of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the transcendental function processing method proposed in the embodiment of the first aspect of the present disclosure is implemented.
[0022] The transcendental function processing method, circuit, device, electronic device, and chip provided by the present disclosure have the following beneficial effects:
[0023] In the embodiments of the present disclosure, based on the function description information corresponding to the transcendental function to be processed, a plurality of target sub-intervals for fitting the transcendental function and the polynomial coefficients corresponding to each target sub-interval are determined. Then, based on the interval range corresponding to the target sub-interval and the mapping relationship between the polynomial coefficients, a look-up table is generated. Furthermore, the look-up table and the polynomial coefficients are used to configure the computing circuit for processing the transcendental function, and the parameters of each circuit component configuration, such as the bit width, are adjusted. Thus, the accuracy of the polynomial fitting the transcendental function can be improved, providing conditions for improving the computing efficiency of the hardware implementation, optimizing the area and performance of the computing circuit, and further enhancing the robustness of the transcendental function processing.
[0024] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and / or additional aspects and advantages of the present disclosure will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0026] Figure 1 is a schematic flowchart of a method for processing a transcendental function provided by an embodiment of the present disclosure;
[0027] Figure 2 is a schematic flowchart of a method for processing a transcendental function provided by another embodiment of the present disclosure;
[0028] Figure 3a is a schematic diagram of the function image before and after translation when the first sub-interval is on the positive semi-axis provided by an embodiment of the present disclosure;
[0029] Figure 3b is a schematic diagram of the function image before and after translation when the first sub-interval is on the negative semi-axis provided by an embodiment of the present disclosure;
[0030] Figure 4 is a schematic flowchart of a method for processing a transcendental function provided by another embodiment of the present disclosure;
[0031] Figure 5 is a schematic diagram of a transcendental function processing circuit provided by an embodiment of the present disclosure;
[0032] Figure 6 is a schematic structural diagram of a transcendental function processing device provided by another embodiment of the present disclosure;
[0033] Figure 7 shows a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure;
[0034] Figure 8It is a schematic structural diagram of a chip proposed by an embodiment of the present disclosure. Detailed implementation manners
[0035] The embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as a limitation to the present disclosure.
[0036] Currently, transcendental functions are mostly used in operators such as activation functions and normalization in artificial intelligence (AI) algorithms, such as the sigmoid operator in a neural network (NN) and the softmax operator in a deep learning model Transformer based on the self-attention mechanism. Usually, the amount of computation is huge and it depends on high computational precision. There is a need to improve the computational performance of transcendental functions in parallel computing units such as a graphics processing unit (GPU) and a neural processing unit (NPU).
[0037] The computational implementation of transcendental functions is mainly divided into two types: software and hardware. Software is based on basic multiply-add instructions and is implemented by polynomial fitting such as Taylor expansion. Software implementation often requires dozens of instructions for calculation and has poor performance. Hardware implementation can use polynomial fitting calculation similar to software implementation and can support more types of operators, but is limited by the parallel processing architecture (PPA, i.e., Parallel Processing Architecture) and has low computational precision. Alternatively, the CORDIC (Coordinate Rotation Digital Computer) algorithm can also be used to calculate by iteratively rotating the angle of the transcendental algorithm. This algorithm does not require a multiplier for operations of trigonometric functions, square roots, and exponential functions. The area and power consumption of hardware implementation are very friendly, but the operators are limited by their own rotation modes and need to be iterated multiple times in high-precision calculation scenarios, and the performance is limited by this.
[0038] Therefore, the present disclosure provides a transcendental function processing method. By adopting Chebyshev and endpoint interpolation during polynomial fitting, the number of table entries obtained by solving is less than that of Taylor expansion, which is beneficial to improving the calculation speed of transcendental functions. And by using the offset method, the subtraction operation is optimized, the calculation period is compressed, and the sub-interval division is solved by using the software iteration method, further compressing the coefficient bit width. It can reduce the bit widths of the look-up table, multiplier, and accumulator in the hardware implementation of the transcendental function calculation unit, be compatible with various types of operators, improve the calculation performance of the transcendental function, and optimize PPA at the same time.
[0039] It should be noted that the transcendental function processing method provided by the present disclosure can be used for the design of the arithmetic logic unit (ALU) in a chip, specifically for designing the calculation unit of the transcendental function, which can also be called a non-linear function or a special function unit (SFU). Or it can also be extended to other algorithmic technical fields outside chip design.
[0040] Next, the transcendental function processing method, circuit, device, electronic device, and chip of the embodiments of the present disclosure will be described with reference to the accompanying drawings.
[0041] Figure 1 It is a schematic flowchart of a transcendental function processing method provided by an embodiment of the present disclosure.
[0042] It should be noted that the transcendental function processing method of the embodiments of the present disclosure can be applied to a transcendental function processing device. In some possible embodiments, the transcendental function processing device can be configured in an electronic device or a chip, so that the electronic device or the chip can perform the function of calculating transcendental functions.
[0043] As Figure 1 shown, the transcendental function processing method may include the following steps:
[0044] Step 101, determine the transcendental function to be processed and function description information.
[0045] In some embodiments, the function description information may include one or more of the fitting range, input data format, reference error, initial sub-interval length, and target data format, etc.
[0046] It should be noted that for different types of transcendental functions, the method of obtaining the fitting range is different. When the transcendental function is a symmetric or periodic function, the fitting range is the period of the transcendental function closest to zero. The curve in this range can cover the complete function image through flipping and copying. For example, when the transcendental function is F(x) = sin(2πx), the fitting range can be selected as [0, 0.25).
[0047] When the transcendental function is a function whose output tends to saturate, first obtain the saturation region according to the calculation error. Since the function output value grows very slowly or remains almost unchanged in the saturation region, the non-saturation region can be used as the fitting range of the transcendental function. For example, when the transcendental function is F(x) = tanh(x), the fitting range can be [0, 4).
[0048] Alternatively, when the transcendental function is a function corresponding to the floating-point representation structure, the fitting range of the transcendental function can be determined according to the floating-point normalized representation in the IEEE binary floating-point arithmetic standard, as shown in the following formula (1).
[0049] (-1) sign ×2 Exp-bias ×1.Mantissa (1)
[0050] Among them, (-1) sign is the sign bit of the floating-point number, indicating the positive or negative of the floating-point number. sign being 0 indicates a positive number, and sign being 1 indicates a negative number. 2 Exp-bias is the exponent part of the floating-point number, representing the power of the floating-point number. Usually, it is an exponent with a base of 2. The exponent needs to subtract a bias value bias from the value of the exponent Exp to obtain the actual exponent value. This bias value is 127 for 32-bit floating-point numbers and 1023 for 64-bit floating-point numbers. 1.Mantissa is the mantissa part of the floating-point number, which is a binary decimal used to represent the significant digit part of the floating-point number. The actually stored mantissa part only contains the binary decimal after the highest significant bit is 1.
[0051] For example, when the transcendental function is the reciprocal calculation function F(x) = rcp(x) = 1 / x, the exponent part of the rcp result can be obtained by taking the negative of the exponent part, and then by calculating the reciprocal of the mantissa part (i.e., 1.Mantissa in formula 1), the fitting range can be determined as [1, 2).
[0052] It should be noted that the input data format refers to the precision of the independent variable substituted into the transcendental function for calculation, which can be determined according to actual needs, such as single-precision floating-point numbers (i.e., 32-bit floating-point numbers, FP32), double-precision floating-point numbers (FP64), etc. The target data format is the precision of the output value of the transcendental function calculation, and the target data format should be more precise than the input data format, that is, the bit width is doubled compared to the input data format. For example, if the input data format is FP32, then the calculation result of FP64 is taken as the target data format after fitting is completed.
[0053] It should be noted that the reference error can be only the relative error, that is, the absolute difference between the actual value and the fitted value is required to be less than a fixed value, or it can be only the absolute error, that is, the absolute ratio of the difference between the actual value and the fitted value to the actual value is required to be less than a fixed value. The fixed value can be determined according to the actual fitting needs. Or the reference error can also include both relative error and absolute error.
[0054] In the present disclosure, the reference error can be described by the Unit of Least Precision (ULP). The description of ULP is not completely equivalent to the relative error. In hardware representation, it represents how many Least Significant Bits (LSBs) the distance between the calculation result and the true value is. And when using ULP to describe, the calculation error value of the input algorithm also needs to be adjusted according to the rounding mode. If it is the Round Ties ToEven or Round Ties To Away mode specified by IEEE, the calculation error also needs to subtract 0.5 ULP; if it is the Round Toward Zero, Round Toward Positive or Round TowardNegative mode, since the rounding decision always biases in one direction, which may lead to a larger error, the calculation error also needs to subtract 1 ULP.
[0055] In the embodiments of the present disclosure, the initial sub-interval length is the largest power of 2 less than or equal to the size of the fitting range. For example, for the rcp(x) operator, the fitting range is [1, 2), then the size of the fitting range can be determined as 2 - 1 = 1, and the 0th power of 2 is 1, which is equal to the size of the fitting range, so the initial sub-interval length can be determined as 2 0 。
[0056] Therefore, in the embodiments of the present disclosure, after determining the transcendental function to be processed currently, the function description information corresponding to the transcendental function can be obtained according to the type of the transcendental function and the actual needs of the transcendental function calculation.
[0057] Step 102: Based on the function description information, fit the transcendental function to determine multiple target sub-intervals corresponding to the transcendental function and the polynomial coefficients corresponding to each target sub-interval.
[0058] It should be noted that generally, the fitting range of the transcendental function is relatively large. Fitting directly in a large interval will cause a large fitting error and low calculation accuracy. Therefore, in this disclosure, the initial fitting range needs to be cut multiple times into multiple small sub-intervals first, and then fitting calculations are performed on each sub-interval to obtain multiple sets of polynomial coefficients. Thus, a look-up table can be established based on each sub-interval and the corresponding polynomial coefficients, enabling the fitting result to be obtained through table lookup and polynomial calculation when implemented in hardware.
[0059] In the embodiments of this disclosure, multiple sub-intervals with an interval length equal to the initial sub-interval length can be sequentially divided starting from the upper boundary value of the fitting range according to the initial sub-interval length in the function description information until the difference between the minimum lower boundary value of the multiple sub-intervals and the lower boundary value of the initial fitting range is a power of 2 value, such as 0.5, that is, 2 to the power of -1, so that the sub-interval division can be aligned with powers of 2, and a comparator is not required for subsequent address decoding of the sub-intervals.
[0060] For example, for the transcendental function F(x) = rcp(x) = 1 / x with a fitting range of [1, 2), when the initial sub-interval length is 2 0 a sub-interval [1, 2) is obtained. After iteratively updating the sub-interval length and the sub-interval length is 2 -3 , that is, 0.125, 4 sub-intervals can be obtained, which are [1.5, 1.625), [1.625, 1.75), [1.75, 1.875), and [1.875, 2). At this time, the difference between the minimum lower boundary value 1.5 of the obtained sub-intervals and the lower boundary value 1 of the initial fitting range is 0.5, which is a power of 2 value.
[0061] Then, for each sub-interval obtained after division, interpolation method can be used for fitting, error calculation, and interval division iteration to determine multiple target sub-intervals corresponding to the transcendental function and the polynomial coefficients corresponding to each target sub-interval.
[0062] For example, the Chebyshev interpolation method can be used to perform second-order fitting calculations in each sub-interval with the same length as the initial sub-interval to obtain 3 interpolation points within the sub-interval, as shown in the following formula (2).
[0063]
[0064] where x 0 , x 1 and x 2For three fitting interpolation points in the sub-interval [a i , a i+1 ), a i is the lower bound of the sub-interval, and a i+1 is the upper bound of the sub-interval.
[0065] Then, these 3 interpolation points can be respectively substituted into the transcendental function to calculate the corresponding true values, which can be respectively expressed as f 0 , f 1 and f 2 . Then, for the true values f 0 , f 1 and f 2 corresponding to the 3 interpolation points respectively, Newton interpolation is used, as shown in the following formula (3), to calculate the polynomial coefficients corresponding to this sub-interval.
[0066]
[0067] Among them, C 0 is the constant term coefficient of polynomial fitting, C 1 is the first-order term coefficient, and C 2 is the second-order term coefficient.
[0068] Then, the polynomial function for fitting the transcendental function can be shown as the following formula (4).
[0069] appr(x) = c 2 (x - a i ) 2 + c 1 (x - a i ) + C 0 (4)
[0070] Among them, appr(x) is the polynomial fitting value corresponding to the independent variable x, and the value range of x is [a i , a i+1 ). It can be seen that when fitting the transcendental function using the above formula (4) in different sub-intervals, the polynomial coefficients in formula (4) are different.
[0071] After that, multiple points can be taken in each sub-interval to calculate the error between the result fitted by formula (4) and the result calculated by the transcendental function for each point, and compare it with the reference error to determine whether it is necessary to determine a smaller sub-interval length to further divide and fit this sub-interval. For example, the lower boundary point can be pasted in the sub-interval, 16 points can be taken at the average interval of the sub-interval length, and then the upper boundary point of the sub-interval is added to calculate the error for these 17 points; or according to experience and the characteristics of the transcendental function, points with larger fitting errors can be selected, etc.
[0072] In an embodiment of the present disclosure, when there is at least one point corresponding to an error greater than the reference error in any sub-interval, the initial sub-interval length can be updated, and with a smaller sub-interval length, multiple sub-intervals are redetermined and returned. For example, when the initial sub-interval length is 2 0 , and in the case where there is a corresponding error greater than the reference error in the divided sub-intervals, 2 0 / 2 = 2 -1 is updated to the new sub-interval length, and the fitting range is redetermined for sub-interval division.
[0073] Alternatively, when the errors corresponding to all sub-intervals are less than or equal to the reference error, according to the initial sub-interval length corresponding to the sub-interval, continue to divide the sub-interval downward to obtain the next boundary aligned with the power of 2, and repeat the above operations of calculating interpolation points, polynomial coefficients, and errors for the newly obtained sub-intervals. For example, when the sub-interval length is 2 -3 , and each of the divided sub-intervals [1.5, 1.625), [1.625, 1.75), [1.75, 1.875), and [1.875, 2] satisfies that the corresponding error is less than or equal to the reference error, 1.5 can be used as the upper boundary value, and 2 -3 is used as the sub-interval length, and continue to divide the sub-interval downward to obtain the next boundary aligned with the power of 2 (i.e., the -2 power of 2, 0.25), then the sub-intervals are [1.25, 1.375), [1.375, 1.5), and then calculate the interpolation points, polynomial coefficients, and errors for the sub-intervals.
[0074] By repeating the above operations, target sub-intervals with all errors less than or equal to the reference error can be obtained, as well as the polynomial coefficients corresponding to each target sub-interval.
[0075] It should be noted that for the convenience of hardware processing, after determining the target sub-intervals, the present disclosure can translate the target sub-intervals. Translate the target sub-intervals located on the positive semi-axis in the negative direction so that the lower bound of the sub-interval aligns with the 0 point, and translate the target sub-intervals located on the negative semi-axis in the positive direction so that the upper bound of the sub-interval aligns with the 0 point. Thus, the bit width of the polynomial coefficients can be reduced. When implementing the fitting calculation of transcendental functions in the circuit, there is no need to actually implement a subtractor, nor to record the scaling factor, reducing the hardware calculation.
[0076] It should be noted that the interval lengths of multiple target sub-intervals may be the same or different, but within the fitting range, after removing the bias of the lower bound (i.e., the lower bound of the sub-interval aligns with the 0 point), there is only one sub-interval length between any two powers of 2.
[0077] Step 103: Generate a lookup table according to the interval range and polynomial coefficients corresponding to each target sub-interval.
[0078] In the embodiments of the present disclosure, the interval range corresponding to each target sub-interval can be mapped to the polynomial coefficients obtained after calculating the target sub-interval, so as to generate a look-up table. Thus, by determining which interval range the independent variable to be currently calculated is in, it is possible to determine the coefficients of the polynomial when calculating the value of the independent variable through polynomial fitting, thereby accelerating the calculation and improving the calculation efficiency.
[0079] It should be noted that after generating the look-up table, the high segment of each target sub-interval can be used as the look-up table address by means of truncation, and the low segment is the result of the above formula (4) (x - a i ) for quadratic polynomial calculation.
[0080] Step 104: Determine the configuration parameters of each circuit component in the calculation circuit for processing transcendental functions according to the polynomial coefficients and the look-up table.
[0081] In the embodiments of the present disclosure, a calculation circuit adapted to the above sub-interval division can be set up to implement hardware processing of transcendental functions. To improve the processing efficiency of the calculation circuit, reduce the coefficient bit width to save the size of the multiplier in hardware implementation, and optimize PPA, a look-up table can be configured in the calculation circuit, and the bit widths of the polynomial coefficients, the truncation bit width of the square term, the truncation bit width of the first-order term multiplication, and the bit width of the accumulation alignment can be iteratively solved, so that the bit widths of each term are as small as possible, and the combination satisfies that the error of each target sub-interval is less than or equal to the reference error, to determine the configuration parameters of each circuit component (such as squarer, mantissa normalization component, etc.) in the calculation circuit.
[0082] In this embodiment, by determining multiple target sub-intervals for fitting the transcendental function and the corresponding polynomial coefficients of each target sub-interval based on the function description information corresponding to the transcendental function to be processed, then generating a look-up table based on the mapping relationship between the interval range corresponding to the target sub-interval and the polynomial coefficients, and then using the look-up table and the polynomial coefficients to configure the calculation circuit for processing the transcendental function and adjusting the parameters of each circuit component configuration, such as the bit width, etc. Thus, the accuracy of polynomial fitting for transcendental functions can be improved, providing conditions for improving the calculation efficiency of hardware implementation, optimizing the area and performance of the calculation circuit, and further improving the robustness of transcendental function processing.
[0083] Figure 2 It is a schematic flowchart of a method for processing transcendental functions provided by an embodiment of the present disclosure. As Figure 2 shown, the method for processing transcendental functions may include the following steps:
[0084] Step 201: Determine the fitting range of the transcendental function according to the type of the transcendental function.
[0085] In the embodiments of the present disclosure, it can be seen that for different types of transcendental functions, the characteristics of their images are different. Therefore, the selection methods of the intervals to be fitted are also different.
[0086] For example, when the transcendental function is a symmetric or periodic function, the fitting range is the period of the transcendental function closest to zero. The curve in this range can cover the complete function image through flipping and copying; when the transcendental function is a function with an output tending to saturation, such as the sigmoid function, etc., the saturation region can be obtained first according to the calculation error. Since the function output value grows very slowly or remains almost unchanged in the saturation region, the non-saturation region can be used as the fitting range of the transcendental function. Or, when the transcendental function is a function corresponding to the floating-point representation structure, the fitting range of the transcendental function can be determined according to the floating-point normalized representation in the IEEE binary floating-point arithmetic standard.
[0087] Step 202: Determine a first value that is less than or equal to the size of the fitting range and is the nth power of a preset number as the initial sub-interval length.
[0088] Wherein, the size of the fitting range is the difference between the upper boundary value and the lower boundary value of the fitting range.
[0089] Wherein, the preset number is the base number for sub-interval division in the form of the nth power determined according to needs. In the embodiments of the present disclosure, the preset number can be determined by the exponent part in the floating-point normalized representation, such as 2, or it can also be other values.
[0090] Wherein, the difference between the nth power of the preset number and the size of the fitting range is less than the difference between other powers of the preset number and the size of the fitting range, and n is an integer.
[0091] In the embodiments of the present disclosure, after determining the fitting range of the transcendental function, the difference between the upper boundary value and the lower boundary value of the fitting range can be calculated to obtain the size of the fitting range. Then, determine the maximum value of n that can satisfy the condition that the value of the nth power of the preset number is less than or equal to the size of the fitting range, so as to determine the first value of the nth power of the preset number corresponding to this n value as the initial sub-interval length.
[0092] For example, if the fitting range is [1, 2), it can be determined that the size of the fitting range is 2 - 1 = 1. When the preset number is 2, it can be known that the 0th power of 2 is 1, which is equal to 1, and the difference from the size of the fitting range is 0. When n is less than 0 or greater than 0, the difference between the nth power of 2 and the size of the fitting range is greater than 0. Then, the initial sub-interval length can be determined to be 2. 0, that is, the first value is 1. Or, if the fitting range is [0.5, 2), the size of the fitting range can be determined as 2 - 0.5 = 1.5. When the preset number is 2, it can be known that the 0th power of 2 is 1 which is less than 1.5, and the difference from the size of the fitting range is 0.5. When n is less than 0 or greater than 0, the difference between the nth power of 2 and the size of the fitting range is greater than 0.5. Then the initial sub-interval length can be determined as 2 0 , that is, the first value is 1.
[0093] Step 203, based on the initial sub-interval length, with the upper boundary of the fitting range as the upper boundary of the first first sub-interval, sequentially obtain multiple first sub-intervals from within the fitting range.
[0094] Among them, the difference between the lower boundary value of the last first sub-interval and the lower boundary value of the fitting range is the power value of the preset number.
[0095] For example, the initial sub-interval length is 2 0 = 1, and the fitting range is [0.5, 2). Then the upper boundary of the fitting range is the upper boundary of the first first sub-interval, and the first first sub-interval obtained is [1, 2). Since the difference between the lower boundary value 1 of [1, 2) and the lower boundary value 0.5 of the fitting range is 0.5, which is the -1th power of 2, the first sub-interval can be determined as [1, 2). In some possible embodiments, the initial sub-interval length can be iteratively updated to 2 -3 , then in the fitting range [0.5, 2), multiple first sub-intervals [1.875, 2), [1.75, 1.875), [1.625, 1.75), and [1.5, 1.625) are sequentially divided. Since 1.5 - 0.5 = 1, and the 0th power of 2 is 1, it can be determined that the division of the first sub-interval is completed.
[0096] Step 204, determine the polynomial coefficients and multiple verification points corresponding to each first sub-interval.
[0097] Among them, the verification point refers to the point used to calculate the error between the fitting value of the polynomial and the true value of the transcendental function within the sub-interval, and the accuracy of the polynomial fitting can be judged through the error corresponding to the verification point.
[0098] In the embodiments of the present disclosure, the Chebyshev interpolation method can be used to calculate the coefficients of the polynomial when performing polynomial fitting on the transcendental function within each first sub-interval. And multiple verification points corresponding to each first sub-interval can be determined according to which points in the transcendental function may have a large error during fitting.
[0099] Optionally, multiple first interpolation points corresponding to each first sub-interval can be determined first.
[0100] In the embodiments of the present disclosure, the first interpolation points are determined by the Chebyshev interpolation method, and the calculation formula of the first interpolation points is as shown in formula (2) in the above embodiments.
[0101] It should be noted that when using Chebyshev interpolation to perform polynomial fitting on transcendental functions, the order of the polynomial is determined by the number of selected Chebyshev interpolation points. When 3 first interpolation points are selected as shown in formula (2), a second-order polynomial fitting is performed. It is also possible to select 2 first interpolation points, as shown in the following formula (5), and at this time, a first-order polynomial fitting is performed on the transcendental function.
[0102]
[0103] Then, in the case where both boundary values of any first sub-interval are non-zero, according to the boundary value closest to the 0 point in any first sub-interval, the translation parameter and translation direction corresponding to any first sub-interval are determined.
[0104] In the embodiments of the present disclosure, in the case where the boundary value closest to the 0 point in any first sub-interval is the lower boundary value, the lower boundary of the any first sub-interval can be aligned with the 0 point. That is to say, at this time, the translation direction corresponding to the first sub-interval is the negative axis direction, and the translation parameter is to translate the same distance as the lower boundary value. In the case where the boundary value closest to the 0 point in any first sub-interval is the upper boundary value, the upper boundary of the any first sub-interval can be aligned with the 0 point. That is to say, at this time, the translation direction corresponding to the first sub-interval is the positive axis direction, and the translation parameter is to translate the same distance as the absolute value of the upper boundary value.
[0105] Next, with reference to Figure 3, it is described how to translate the graph of the transcendental function according to the first sub-interval. Figure 3a is a schematic diagram of the function graph before and after translation when the first sub-interval is on the positive half-axis, Figure 3b is a schematic diagram of the function graph before and after translation when the first sub-interval is on the negative half-axis.
[0106] It should be noted that in Figure 3a and Figure 3b , the graph of the transcendental function func(x) is only for illustrative purposes and should not be regarded as a limitation to the present disclosure. For different transcendental functions, the translation of their sub-intervals can be determined similarly.
[0107] In Figure 3a , the graph of the transcendental function before translation is shown on the left side of the arrow. For the first sub-interval [a i , a i+1 in the transcendental function func(x), a i is the lower boundary value of the first sub-interval, a i+1 is the upper boundary value of the first sub-interval, and ai and a i+1 are both non - zero. Since a i is the closest to the origin (0), the graph of the function can be translated by a distance of a in the negative x - axis direction. After translation, we get the function func(x + a i ) and its graph, as shown on the right side of the arrow. At this time, the first sub - interval [a i , a i will correspond to the translated [0, a i+1 - a i+1 . i )
[0108] It should be noted that when the lower - boundary value a i of the first sub - interval is greater than 0, the implementation of (x - a i ) in hardware is equivalent to truncating the higher - order bits of the input, which is easy to implement.
[0109] In Figure 3b , the graph of the transcendental function before translation is shown on the left side of the arrow. For the first sub - interval [a i , a i+1 in the transcendental function func(x), a i is the lower - boundary value of the first sub - interval, a i+1 is the upper - boundary value of the first sub - interval, a i and a i+1 are both non - zero and less than 0. Since the representation of floating - point negative numbers is only through the sign, and the representation of exponent plus mantissa is always positive, and the absolute value of a i is greater than the absolute value of any x in this first sub - interval, so (x - a i ) actually calculates |a i | - |x|, and the negative operation on |x| will bring a relatively large latency.
[0110] In the embodiments of the present disclosure, to avoid subtraction operations, when the sub - interval [a i , a i+1 falls on the negative axis, the right - hand endpoint of the graph of func(x) in this sub - interval can be translated to the origin (0), that is, the graph of the function is translated by a distance of |a i+1 | in the positive x - axis direction. After translation, the fitted function is func(x + a i+1 ), and its graph is as shown on the right side of the arrow in Figure 3b . At this time, the first sub - interval [a i , a i+1 will correspond to the translated [0, a i - a i+1 ). Thus, before fitting the transcendental function, it is necessary to calculate x′ = x - a i+1, what is actually calculated is -(|x| - |a i+1 |), which can be implemented by truncation on hardware, and the negation operation of the result can be incorporated into the sign of coefficient C 1 .
[0111] After that, based on the translation parameter and the translation direction, the coordinates of multiple first interpolation points can be transformed respectively to obtain multiple second interpolation points.
[0112] In the embodiments of the present disclosure, when the first interpolation points are as shown in the above formula (2), based on the translation parameter and the translation direction, the coordinates of multiple first interpolation points are transformed respectively to obtain multiple second interpolation points.
[0113] When the translation direction is translation towards the negative axis, the expressions of multiple second interpolation points can be as shown in the following formula (6).
[0114]
[0115] Among them, B 0 is the interval length of the first sub-interval.
[0116] When the translation direction is translation towards the positive axis, the expressions of multiple second interpolation points can be as shown in the following formula (7).
[0117]
[0118] Among them, B size also represents the length of the sub-interval, and N r is the parameter of the coordinate transformation.
[0119] Then, based on multiple second interpolation points, the polynomial coefficients corresponding to any first sub-interval are determined.
[0120] In the embodiments of the present disclosure, these multiple second interpolation points can be respectively substituted into the translated transcendental function to calculate the true value corresponding to each second interpolation point. Since the values corresponding to the first interpolation point and the second interpolation point in the transcendental function should be the same, they can also be respectively expressed as f 0 、f 1 and f 2 , and the polynomial coefficients corresponding to this sub-interval are calculated using the formula (3) shown in the above embodiments.
[0121] Finally, according to the length of any first sub-interval, multiple verification points can be obtained from any first sub-interval, where multiple verification points include the maximum error point corresponding to the transcendental function.
[0122] In the embodiments of the present disclosure, in order to ensure that the selected verification points can more accurately and reliably reflect the error between the polynomial and the transcendental function, at least the maximum error point corresponding to the transcendental function may be included in the multiple verification points, and the selection of the maximum error point may be determined by the difference method used during fitting. For example, when the second-order Chebyshev interpolation method is adopted, the maximum error points are the sub-interval endpoints and the 1 / 4 points.
[0123] It should be noted that, in addition to the maximum error points, the verification points may also include some other points to avoid the influence of floating-point truncation errors. Therefore, empirically, the verification points can be taken as the 16-equal division points, which can include the maximum error points and are convenient for iterative calculation.
[0124] Step 205: Based on the polynomial coefficients corresponding to each first sub-interval, determine the error value corresponding to each verification point within the first sub-interval.
[0125] In the embodiments of the present disclosure, the polynomial coefficients corresponding to each first sub-interval and the selected verification points can be respectively substituted into formula (4) in the above embodiments to calculate the polynomial fitting value corresponding to each verification point within each first sub-interval. And the true value of the transcendental function corresponding to the verification point can be calculated, and the absolute error and / or relative error between the fitting value and the true value is determined as the error value corresponding to each verification point within the first sub-interval. Here, when calculating the error, the error calculation method should correspond to the specified reference error.
[0126] Step 206: When the error value corresponding to at least one verification point is greater than the reference error, update the initial sub-interval length to obtain the updated sub-interval length.
[0127] Among them, the updated sub-interval length is less than the length of the initial sub-interval.
[0128] In the embodiments of the present disclosure, when the error value corresponding to at least one verification point is greater than the reference error, it can be determined that the fitting accuracy of the transcendental function in the first sub-interval is relatively low, and then the length of the divided sub-interval can be shortened to improve the polynomial fitting accuracy.
[0129] Optionally, the quotient of the initial sub-interval length and a preset number can be determined as the updated sub-interval length.
[0130] For example, when the initial sub-interval length is 2 0 = 1, and the preset number is 2, then the updated sub-interval length is 1 / 2 = 0.5 = 2 -1 , which can ensure that when the sub-interval division is iterated with the updated sub-interval length, the length of the sub-interval is aligned with the power of the preset number.
[0131] Step 207: Based on the updated sub-interval length, return to perform the operation of sequentially obtaining multiple first sub-intervals within the fitting range until multiple second sub-intervals are obtained for which the error values corresponding to all verification points are less than the reference error.
[0132] In the embodiments of the present disclosure, after determining the updated sub-interval length, multiple first sub-intervals can be sequentially obtained again from within the fitting range with the updated sub-interval length, and the error values corresponding to multiple verification points within each re-divided first sub-interval are judged again. If there is still at least one error value greater than the reference error, the sub-interval length is continuously updated and the sub-interval is re-divided. On the contrary, in the case where the error values of all verification points are less than the reference error, the multiple first sub-intervals divided by the sub-interval length at this time can be determined as the second sub-intervals.
[0133] Step 208: Determine the multiple second sub-intervals as multiple target sub-intervals corresponding to the transcendental function.
[0134] In the embodiments of the present disclosure, since the error values corresponding to the multiple second sub-intervals are all less than the reference error, it can be determined that when the transcendental function is divided by the second sub-interval and then fitted, the accuracy is relatively high. Therefore, the multiple second sub-intervals can be determined as the multiple target sub-intervals corresponding to the transcendental function.
[0135] Optionally, after determining the multiple target sub-intervals, the operation of obtaining multiple first sub-intervals can also be returned based on the updated sub-interval length and the remaining range within the fitting range until all the target sub-intervals corresponding to the transcendental function are determined.
[0136] It can be understood that in the present disclosure, when dividing the first sub-interval from the upper boundary of the fitting range, the lower boundary value of the last first sub-interval divided is not the lower boundary value of the fitting range, but a power value whose difference from the lower boundary value of the fitting range is a preset number. This can ensure that the minimum lower boundary of the multiple target sub-intervals determined each time is aligned with the power value of the preset number, and there is only one sub-interval length between any two power values of the preset number.
[0137] Therefore, after the first sub-interval division starts from the upper boundary of the fitting range to obtain multiple target sub-intervals, the remaining range within the fitting range is the range between the lower boundary value of the fitting range and the minimum lower boundary of the multiple target sub-intervals. Then, within this remaining range, the updated sub-interval length can be used to continue dividing the sub-intervals, and the last sub-interval divided also satisfies that the difference between the lower boundary value and the lower boundary value of the fitting range is a power value of a preset number. Repeat the above operations such as selecting sub-interval verification points, calculating errors, and updating sub-interval lengths until the only remaining interval within the fitting range is equal to the updated sub-interval length, and the error in this remaining interval is less than the reference error, then all target sub-intervals corresponding to the transcendental function can be obtained.
[0138] For example, the updated sub-interval length is 2 -3 = 0.125. Starting from the upper boundary of the fitting range for division, the multiple target sub-intervals obtained are [1.5, 1.625), [1.625, 1.75), [1.75, 1.875), and [1.875, 2). Then, the remaining range within the fitting range can be determined as [1, 1.5). Return to execute the operation of obtaining multiple first sub-intervals to get [1.25, 1.375) and [1.375, 1.5). At this time, 1.25 - 1 = 0.25, which is equal to 2 to the power of -2. Then, calculate the error for each sub-interval. When the errors are all less than the reference error, [1.25, 1.375) and [1.375, 1.5) are also determined as target sub-intervals, and continue to divide the remaining range of [1, 1.25) into multiple sub-intervals with a length of 2 -3 , and so on, until all target sub-intervals corresponding to the transcendental function are obtained.
[0139] Step 209, generate a lookup table according to the interval range and polynomial coefficients corresponding to each target sub-interval.
[0140] Step 210, determine the configuration parameters of each circuit component in the calculation circuit for processing the transcendental function according to the polynomial coefficients and the lookup table.
[0141] For the detailed descriptions of the above Step 209 and Step 210, reference can be made to the above embodiments of the present disclosure, which will not be elaborated here.
[0142] Figure 4 It is a schematic flowchart of a method for processing a transcendental function provided by an embodiment of the present disclosure. As Figure 4 shown, the method for processing the transcendental function may include the following steps:
[0143] Step 401, determine the transcendental function to be processed and function description information.
[0144] Step 402: Based on the function description information, fit the transcendental function to determine multiple target sub-intervals corresponding to the transcendental function and the polynomial coefficients corresponding to each target sub-interval.
[0145] For the detailed descriptions of the above steps 401 and 402, reference can be made to the above embodiments of the present disclosure, which will not be elaborated here.
[0146] Step 403: Determine the address index corresponding to each target sub-interval according to the lower boundary value of the interval range corresponding to each target sub-interval.
[0147] In the embodiments of the present disclosure, since the lower boundary value of the interval range corresponding to each target sub-interval is aligned with the power value of a preset number, in the present disclosure, an address index can be assigned to the lower boundary value, i.e., the high segment, of each target sub-interval, such as 001, 010, etc.
[0148] It should be noted that in the present disclosure, there may be a situation where the lengths of the sub-intervals among all the target sub-intervals are inconsistent, and leading 0s can be added before the address index to distinguish sub-intervals of different sizes.
[0149] Step 404: Generate a lookup table based on the address index and polynomial coefficients corresponding to each target sub-interval.
[0150] In the embodiments of the present disclosure, the address index and polynomial coefficients corresponding to each target sub-interval can be mapped to generate a lookup table. Thus, by configuring the lookup table in the computing circuit, when calculating the transcendental function, the coefficients of polynomial fitting can be quickly determined in the lookup table through the address index.
[0151] Step 405: Determine the configuration parameters of each circuit component in the computing circuit for processing the transcendental function according to the polynomial coefficients and the lookup table.
[0152] It should be noted that the configuration parameters can be the bit widths of each circuit component during calculation. In the embodiments of the present disclosure, by controlling variables, the available range of the bit width of each circuit component during calculation can be determined, and then the optimal configuration parameters can be determined through permutation and combination.
[0153] Optionally, the candidate bit widths corresponding to each parameter to be calculated can be determined first according to the input data format and the polynomial corresponding to each target sub-interval.
[0154] Among them, the parameters to be calculated can include the quadratic term coefficient, the linear term coefficient, the constant term coefficient, the truncation of the square term, the truncation of the linear term multiplication, and the accumulation alignment.
[0155] Among them, the candidate bit width is the value of all the maximum significant bit widths less than or equal to the corresponding calculation parameter.
[0156] For example, the significant bits of fp64 is 52 bits (excluding the implicit 1). After iteration of sub-intervals, it can be considered that the maximum significant bit width of the coefficients of the polynomial is (52, 52, 52). The truncation bit width of the quadratic term, after x - a i and then entering the highest - order exponent of the quadratic - term fitting calculation is known. When it is fp32, the maximum of this bit width is 24 bits, and the result of the square calculation is 48 bits. To save the bit width of subsequent multiplication calculations, the quadratic term is truncated once. The rule is to retain a certain number of bits starting from the highest bit. The initial value is 2 times the mantissa length of fp32 with implicit bits, that is, the candidate bit width is 48 bits. The truncation bit width of the linear - term multiplication is the same as that of the quadratic term, saving the shifter for accumulation alignment. The initial value is set to the target data format, and the mantissa length of fp64 with implicit bits is 53. The bit width of accumulation alignment, after alignment, only retains a certain number of bits counted backward from the highest - order exponent of the constant term, and directly discards the shifted - out bits, which is equivalent to truncation. The initial value is set to the adder length of fp64 fma, which is 111 bits.
[0157] Then, based on the reference error, it is possible to traverse the candidate bit widths corresponding to each parameter in turn to determine the available bit widths corresponding to each parameter.
[0158] In the embodiments of the present disclosure, a parameter to be calculated can be selected, the bit widths of other parameters to be calculated are fixed, starting from the maximum candidate bit width, gradually reducing the bit width of the parameter to be calculated. On all target sub - intervals, verify whether the verification points of each target sub - interval are less than or equal to the parameter error, and find the minimum bit width of the parameter to be calculated that meets the error requirement, so as to determine the available bit widths corresponding to each parameter, that is, the available bit width is greater than or equal to the minimum value and less than or equal to the maximum candidate bit width.
[0159] After that, it is possible to traverse the available bit widths corresponding to each parameter to be calculated to determine the minimum bit width combination corresponding to all parameters.
[0160] In the embodiments of the present disclosure, for the available bit widths corresponding to each parameter to be calculated, through permutation and combination, traverse and find the bit width combinations that meet the error requirements. First, select the combination with the smallest sum of the bit width of the quadratic - term coefficient and the bit width of the quadratic term. Secondly, select the combination with the smallest sum of the bit width of the linear term and the minimum truncation of the linear - term multiplication, so as to obtain the minimum bit width combination corresponding to all parameters. Thus, based on this minimum bit width combination, the configuration parameters of the corresponding circuit components in the calculation circuit for processing transcendental functions can be set.
[0161] In the embodiments of the present disclosure, based on the transcendental - function processing method proposed in the above - mentioned embodiments, a transcendental - function processing circuit is also proposed to implement the fitting calculation of transcendental functions through hardware. Figure 5The following is a schematic structural diagram of a transcendental function processing circuit provided by an embodiment of the present disclosure. As Figure 5 shown, the transcendental function processing circuit may include:
[0162] An address decoder 51, a lookup table 52, and an arithmetic component 53.
[0163] Among them, the address decoder 51 can be used to parse the input data and determine the interval to which the input data belongs. Specifically, the input data can be split and range-compressed to obtain the interval to which the input data belongs.
[0164] The address decoder 51 can also determine the target address and the data to be operated based on the interval to which the input data belongs, and obtain the target polynomial coefficient from the lookup table 52 based on the target address.
[0165] It should be noted that in the present disclosure, after dividing the fitting range of the transcendental function into multiple target sub-intervals, an address index is set based on the upper boundary of the target sub-interval, and the size of the target sub-interval is aligned with the nth power of a preset number. Therefore, the target address in the lookup table can be directly determined according to the interval to which the input data belongs, and the data to be operated (x - a i ) can be determined based on the difference between the input data and the lower boundary value of the interval, and the polynomial coefficient mapped to the target address can be obtained by traversing the lookup table 52. Thus, there is no need to specifically store the boundary, and only the leading 0 recognition of the high segment needs to be performed to decode the address segment and the result of (x - a i ).
[0166] The arithmetic component 53 can be used to process the data to be operated based on the target polynomial coefficient to obtain the operation result corresponding to the input data.
[0167] It should be noted that when calculating quadratic interpolation, the arithmetic component 53 may include a squarer, an accumulator, etc. If it is not calculating quadratic interpolation, other calculation units may also be included in the arithmetic component.
[0168] In Figure 5 the transcendental function processing circuit shown, the arithmetic component 53 may include a mantissa normalization component, a squarer, an accumulator, and a normalization component, etc. The squarer can calculate the square of (x - a i ), and the calculation result will be truncated according to the configured truncation bit width when outputting. Since the highest bit of the result of (x - a i ) is not necessarily 1, the mantissa normalization component can move the leading 1 of the result to the highest bit to ensure that the high bit of the first-term multiplication calculation contains as many valid bits as possible.
[0169] It should be noted that by inserting mantissa normalization into the circuit structure, more significant bits can be retained in the high-order bits at the output end of the multiplier as much as possible. Under the same precision implementation, the corresponding term coefficients and the multiplier bit width can be compressed, and the PPA optimization in the multi-operator compatible sfu is very obvious.
[0170] The result of (x - a i ) output by the mantissa normalization component can be multiplied by the linear term coefficient obtained from the look-up table, and the multiplication result will be truncated once, and more significant digit information can be retained. The result of (x - a i ) 2 output by the squarer can be multiplied by the quadratic term coefficient obtained from the look-up table. Whether the multiplication result of the quadratic term is truncated can be determined according to the area of the actual implementation. And after the multiplication calculation, the linear term and the quadratic term can perform shift alignment on the constant term, the result bit width is fixed, and the shifted bits are directly discarded (i.e., truncated). Then, by adding the truncated linear term, quadratic term, and constant term, the calculation of the quadratic polynomial is completed and normalized to a floating-point format for output, so that the fitting calculation result of the transcendental function by the calculation circuit can be obtained.
[0171] To implement the above embodiments, the present disclosure also proposes a transcendental function processing device.
[0172] Figure 6 It is a schematic structural diagram of the transcendental function processing device provided by the embodiments of the present disclosure.
[0173] As Figure 6 shown, the transcendental function processing device 600 may include:
[0174] A first determination module 601, configured to determine a transcendental function to be processed and function description information;
[0175] A second determination module 602, configured to fit the transcendental function based on the function description information, determine multiple target sub-intervals corresponding to the transcendental function, and polynomial coefficients corresponding to each target sub-interval;
[0176] A generation module 603, configured to generate a look-up table according to the interval range and polynomial coefficients corresponding to each target sub-interval;
[0177] A third determination module 604, configured to determine configuration parameters of each circuit component in the calculation circuit for processing the transcendental function according to the polynomial coefficients and the look-up table.
[0178] In some possible embodiments, the first determination module 601 may specifically be configured to:
[0179] Determine the fitting range of the transcendental function according to the type of the transcendental function;
[0180] Determine a first value that is less than or equal to the size of the fitting range and is the nth power of a preset number as the initial sub-interval length, where the difference between the nth power of the preset number and the size of the fitting range is less than the difference between other powers of the preset number and the size of the fitting range, and n is an integer.
[0181] In some possible embodiments, the second determination module 602 may specifically be configured to:
[0182] Based on the initial sub-interval length, with the upper boundary of the fitting range as the upper boundary of the first first sub-interval, sequentially obtain a plurality of first sub-intervals from within the fitting range, where the difference between the lower boundary value of the last first sub-interval and the lower boundary value of the fitting range is the power value of the preset number;
[0183] Determine the polynomial coefficients and a plurality of verification points corresponding to each first sub-interval;
[0184] Based on the polynomial coefficients corresponding to each first sub-interval, determine the error value corresponding to each verification point within the first sub-interval;
[0185] In the case where the error value corresponding to at least one verification point is greater than the reference error, update the initial sub-interval length to obtain an updated sub-interval length, where the updated sub-interval length is less than the initial sub-interval length;
[0186] Based on the updated sub-interval length, return to perform the operation of sequentially obtaining a plurality of first sub-intervals from within the fitting range until a plurality of second sub-intervals corresponding to all verification points having error values less than the reference error are obtained;
[0187] Determine the plurality of second sub-intervals as the plurality of target sub-intervals corresponding to the transcendental function.
[0188] In some possible embodiments, the second determination module 602 may specifically be configured to:
[0189] Determine the quotient of the initial sub-interval length and the preset number as the updated sub-interval length.
[0190] In some possible embodiments, the second determination module 602 may specifically be configured to:
[0191] Determine a plurality of first interpolation points corresponding to each first sub-interval;
[0192] In the case where both boundary values of any first sub-interval are non-zero, determine the translation parameter and the translation direction corresponding to any first sub-interval according to the boundary value closest to the 0 point in any first sub-interval;
[0193] Based on the translation parameter and the translation direction, transform the coordinates of the plurality of first interpolation points respectively to obtain a plurality of second interpolation points;
[0194] Based on multiple second interpolation points, determine the polynomial coefficients corresponding to any first sub-interval;
[0195] According to the length of any first sub-interval, obtain multiple verification points from within any first sub-interval, where the multiple verification points include the maximum error point corresponding to the transcendental function.
[0196] In some possible embodiments, the second determination module 602 may also be used for:
[0197] Based on the updated sub-interval length and the remaining range within the fitting range, return to perform the operation of obtaining multiple first sub-intervals until all target sub-intervals corresponding to the transcendental function are determined.
[0198] In some possible embodiments, the generation module 603 may specifically be used for:
[0199] According to the lower boundary value of the interval range corresponding to each target sub-interval, determine the address index corresponding to the target sub-interval;
[0200] Based on the address index and polynomial coefficients corresponding to each target sub-interval, generate a lookup table.
[0201] In some possible embodiments, the third determination module 604 may specifically be used for:
[0202] According to the input data format and the polynomial corresponding to each target sub-interval, determine the candidate bit-widths corresponding to each parameter to be calculated;
[0203] Based on the reference error, sequentially traverse the candidate bit-widths corresponding to each parameter to determine the available bit-widths corresponding to each parameter;
[0204] Traverse the available bit-widths corresponding to each parameter to be calculated to determine the minimum bit-width combination corresponding to all parameters.
[0205] For the functions and specific implementation principles of the above-mentioned modules in the embodiments of the present disclosure, reference may be made to the above-mentioned method embodiments, and details are not described herein again.
[0206] The transcendental function processing device according to the embodiments of the present disclosure determines a plurality of target sub-intervals for fitting the transcendental function and the polynomial coefficients corresponding to each target sub-interval based on the function description information corresponding to the transcendental function to be processed, then generates a lookup table based on the interval range corresponding to the target sub-interval and the mapping relationship between the polynomial coefficients, and then uses the lookup table and the polynomial coefficients to configure a calculation circuit for processing the transcendental function and adjust the parameters of each circuit component configuration, such as the bit width. Therefore, the accuracy of the polynomial fitting the transcendental function can be improved, which provides conditions for improving the calculation efficiency of the hardware implementation, optimizing the area and performance of the calculation circuit, and further improving the robustness of the transcendental function processing.
[0207] To implement the above embodiments, the present disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the transcendental function processing method proposed in the foregoing embodiments of the present disclosure.
[0208] Figure 7 The block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure is shown. Figure 7 The illustrated electronic device 12 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0209] As Figure 7 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0210] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnection (PCI) bus.
[0211] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0212] The memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 7 not shown, commonly referred to as a "hard disk drive"). Although Figure 7 not shown in the figure, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a compact disc read only memory (CD-ROM), a digital video disc read only memory (DVD-ROM) or other optical media) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data media interfaces. The memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present disclosure.
[0213] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in the memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present disclosure.
[0214] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0215] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the foregoing embodiments.
[0216] To implement the foregoing embodiments, the present disclosure also proposes a chip, which includes a processing circuit and an interface circuit; wherein, the interface circuit is used to obtain an instruction and send the instruction to the processing circuit, and the processing circuit is used to execute the instruction to implement the transcendental function processing method proposed in the foregoing embodiments of the present disclosure.
[0217] Figure 8 is a schematic structural diagram of the chip proposed in the embodiments of the present disclosure. Reference can be made to Figure 8 the schematic structural diagram of the chip 800 shown, but not limited thereto.
[0218] The chip 800 includes a processing circuit 801, and the processing circuit 801 is configured to execute any of the above methods.
[0219] In some embodiments, the chip 800 further includes one or more interface circuits 802. Optionally, the interface circuit 802 is connected to a memory 803. The interface circuit 802 can be used to receive signals from the memory 803 or other devices, and the interface circuit 802 can be used to send signals to the memory 803 or other devices. For example, the interface circuit 802 can read the instructions stored in the memory 803 and send the instructions to the processing circuit 801.
[0220] In some embodiments, the interface circuit 802 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 801 performs other steps.
[0221] In some embodiments, terms such as interface circuit, interface, transceiver pin, transceiver, etc. may be used interchangeably.
[0222] In some embodiments, the chip 800 further includes one or more memories 803 for storing instructions. Optionally, all or part of the memories 803 may be outside the chip 800.
[0223] To implement the above embodiments, the present disclosure also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the transcendental function processing method as proposed in the foregoing embodiments of the present disclosure.
[0224] To implement the above embodiments, the present disclosure also provides a computer program product including a computer program that, when executed by a processor, implements the transcendental function processing method as proposed in the foregoing embodiments of the present disclosure.
[0225] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0226] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0227] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the involved functions or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present disclosure pertain.
[0228] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0229] It should be understood that the various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0230] Those of ordinary skill in the art can understand that all or part of the steps carried out in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0231] In addition, in each of the various embodiments of the present disclosure, the functional units can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0232] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A transcendental function processing method, characterized in that: include: Determine the transcendental function to be processed and function description information; Based on the function description information, the transcendental function is fitted to determine a plurality of target subintervals corresponding to the transcendental function and a polynomial coefficient corresponding to each of the target subintervals; Generate a lookup table according to the interval range and polynomial coefficients corresponding to each target subinterval; According to the polynomial coefficients and the lookup table, configuration parameters of each circuit component in a calculation circuit for processing the transcendental function are determined.
2. The method according to claim 1, characterized in that The step of determining the transcendental function and function description information to be processed includes: Determining a fitting range of the transcendental function according to the type of the transcendental function; A first numerical value that is less than or equal to the size of the fitting range and is the nth power of a preset number is determined as the initial sub-interval length, wherein the difference between the nth power of the preset number and the size of the fitting range is smaller than the difference between other powers of the preset number and the size of the fitting range, and n is an integer.
3. The method according to claim 2, characterized in that The function description information includes a fitting range, a reference error, and an initial subinterval length. The fitting of the transcendental function based on the function description information to determine a plurality of target subintervals corresponding to the transcendental function includes: Based on the length of the initial sub-interval, taking the upper boundary of the fitting range as the upper boundary of the first first sub-interval, sequentially obtaining multiple first sub-intervals from the fitting range, wherein the difference between the lower boundary value of the last first sub-interval and the lower boundary value of the fitting range is the power value of the preset number; Determine a polynomial coefficient and a plurality of verification points corresponding to each of the first subintervals; Determine, based on the polynomial coefficient corresponding to each of the first subintervals, an error value corresponding to each verification point in the first subinterval; When the error value corresponding to at least one verification point is greater than the reference error, the initial subinterval length is updated to obtain an updated subinterval length, wherein the updated subinterval length is less than the initial subinterval length; Based on the updated subinterval length, returning to execute the operation of sequentially acquiring multiple first subintervals from the fitting range until multiple second subintervals corresponding to all verification points whose error values are smaller than the reference error are obtained; The multiple second sub-intervals are determined as multiple target sub-intervals corresponding to the transcendental function.
4. The method according to claim 3, characterized in that The updating of the initial sub-interval length to obtain an updated sub-interval length includes: The quotient of the initial sub-interval length and the preset number is determined as the updated sub-interval length.
5. The method according to claim 3, characterized in that The determining of the polynomial coefficients and the plurality of verification points corresponding to each of the first subintervals includes: Determine a plurality of first interpolation points corresponding to each first subinterval; In the case where both boundary values of any first subinterval are non-zero, determining the translation parameter and translation direction corresponding to any first subinterval according to a boundary value in any first subinterval that is closest to the zero point; Based on the translation parameter and the translation direction, the coordinates of the plurality of first interpolation points are respectively transformed to obtain a plurality of second interpolation points; Determine, based on the plurality of second interpolation points, a polynomial coefficient corresponding to any one of the first subintervals; According to the length of any one of the first subintervals, a plurality of verification points are obtained from within the any one of the first subintervals, wherein the plurality of verification points include a maximum error point corresponding to the transcendental function.
6. The method according to claim 3, characterized in that After determining the plurality of second subintervals as a plurality of target subintervals corresponding to the transcendental function, the method further includes: Based on the updated subinterval length and the remaining range within the fitting range, return to execute the operation of obtaining the plurality of first subintervals until all target subintervals corresponding to the transcendental function are determined.
7. The method according to any one of claims 1 to 6, characterized in that: The step of generating a lookup table according to the interval range and polynomial coefficients corresponding to each target subinterval includes: Determine the address index corresponding to each target sub-interval according to the lower boundary value of the interval range corresponding to the target sub-interval; The lookup table is generated based on the address index and polynomial coefficient corresponding to each target sub-interval.
8. The method according to claim 7, characterized in that The function description information includes an input data format and a reference error, and the configuration parameters of each circuit component in the calculation circuit for processing the transcendental function are determined to include: Determine a candidate bit width corresponding to each parameter to be calculated according to the input data format and the polynomial corresponding to each target subinterval; Based on the reference error, sequentially traverse the candidate bit widths corresponding to each of the parameters to determine an available bit width corresponding to each parameter; The available bit widths corresponding to each of the parameters to be calculated are traversed to determine the minimum bit width combination corresponding to all the parameters.
9. A transcendental function processing circuit, characterized in that: include: Address decoder, lookup table and operation components; The address decoder is used to parse the input data and determine the interval to which the input data belongs; Determine a target address and data to be calculated according to the interval to which the input data belongs, and obtain target polynomial coefficients from the lookup table based on the target address; The operation component is used to process the data to be operated based on the target polynomial coefficients to obtain the operation result corresponding to the input data.
10. A transcendental function processing device, characterized in that: include: A first determination module, used to determine the transcendental function and function description information to be processed; A second determination module is used to fit the transcendental function based on the function description information, and determine a plurality of target subintervals corresponding to the transcendental function and a polynomial coefficient corresponding to each of the target subintervals; A generating module, used for generating a lookup table according to the interval range and polynomial coefficients corresponding to each of the target subintervals; The third determination module is used to determine the configuration parameters of each circuit component in the calculation circuit for processing the transcendental function according to the polynomial coefficients and the lookup table.
11. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the transcendental function processing method as claimed in any one of claims 1 to 8 is implemented.
12. A chip, characterized in that: The chip includes a processing circuit and an interface circuit; wherein the interface circuit is used to obtain instructions and send the instructions to the processing circuit, and the processing circuit is used to execute the instructions to implement the transcendental function processing method as described in any one of claims 1-8.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the transcendental function processing method according to any one of claims 1 to 8.
14. A computer program product, characterized in that It comprises a computer program, which, when executed by a processor, implements the transcendental function processing method described in any one of claims 1 to 8.