Complex bivariate function hardware implementation method based on segmented plane approximate calculation

The complex bivariate functions are segmented by segmented plane approximation method and calculated with plane function in each small area. The problems of large storage resource consumption and uneven error distribution in the prior art are solved, efficient and accurate approximation calculation is achieved, and hardware resources are saved.

CN120179979APending Publication Date: 2025-06-20NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510188068.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has problems such as large storage resource consumption and uneven error distribution when calculating complex bivariate functions, making it difficult to achieve efficient and accurate approximate calculations.

Method used

The segmented plane approximation method is used to segment complex bivariate functions in their definition domains, and approximate calculations are performed with a plane function in each small area. This approximation calculation is achieved through multiplexers, comparators, Booth encoders, compressors and advancing adders.

Benefits of technology

On the premise of ensuring calculation accuracy, the area, delay and power consumption of hardware implementation are reduced, saving 30.30%, 16.67% and 47.24% respectively, and operating speed is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179979A_ABST
    Figure CN120179979A_ABST
Patent Text Reader

Abstract

The invention provides a complex bivariate function hardware implementation method based on segmentation plane approximation calculation, which comprises the following steps of: 1, segmenting a complex bivariate function z = f (x, y) to be calculated in a definition domain by adopting a segmentation plane approximation method to obtain segmentation information of all segments in an xy plane; step 2, using a multiplexer and a comparator, positioning the segment to which the input variables x and y belong according to the values of the input variables x and y, and obtaining segment information of the segment; and step 3, according to the segment information obtained in the step 2, calculating by using a Booth encoder, a compressor and a carry lookahead adder to obtain a value of a plane function of the segment, taking the value as a value of the complex bivariate function to be calculated in the current segment, and completing approximate calculation of the complex bivariate function. On the premise of ensuring the calculation precision, the area is saved by 30.30%, the time delay is saved by 16.67%, and the power consumption is saved by 47.24%, so that the method has considerable advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for hardware implementation of a bivariate function, in particular to a method for hardware implementation of a complex bivariate function based on piecewise planar approximation calculation. Background Art

[0002] The information provided in this section is only background information related to the present disclosure, and it does not necessarily represent prior art.

[0003] Complex function calculations have a wide range of applications in fields such as 3D images, digital signal processing, and speech recognition. However, existing methods mainly focus on the calculation of complex univariate functions in the form of, such as the iterative Newton - Raphson method, coordinate rotation method (CORDIC), and mathematical recursive algorithms, as well as approximation - based table addition and piecewise polynomial approximation methods. The calculation of complex bivariate functions of z = f(x, y) is more difficult. These functions are often used in digital signal processing, such as the atan2 function in digital receivers and complex functions in digital signals. The calculation methods for complex bivariate functions can be divided into two categories: traditional algorithms (TAs) and universal algorithms (UAs). Traditional algorithms can solve not only univariate functions but also calculate bivariate functions. For example, the atan2 function can be directly calculated through the vector mode of the coordinate rotation method (CORDIC), or the bivariate function can be expressed as a combination of multiple univariate functions and basic addition, subtraction, multiplication, and division operations through factorization. Universal algorithms are designed specifically for bivariate functions and include two - dimensional (2D) interpolation and piecewise planar approximation algorithms. Two - dimensional (2D) interpolation estimates the function value through pre - set sample points on a two - dimensional plane, and the sample points are evenly distributed on the x - y plane. However, the 2D interpolation algorithm has problems such as high storage resource consumption and uneven error distribution in hardware implementation. The core idea of piecewise planar approximation is to divide the entire domain into multiple small regions (blocks), and then approximate the original complex bivariate function with a planar function in each small region, aiming to achieve efficient and accurate approximation calculation of complex bivariate functions while reducing storage consumption and increasing operation speed. Compared with existing methods such as two - dimensional (2D) interpolation and coordinate rotation method (CORDIC), this algorithm has obvious advantages in hardware implementation.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a method for hardware implementation of a complex bivariate function based on piecewise planar approximation calculation in view of the deficiencies of the prior art.

[0006] To solve the above technical problems, the present invention discloses a hardware implementation method for a complex bivariate function based on piecewise planar approximation, including the following steps:

[0007] Step 1, using the piecewise planar approximation method, segment the complex bivariate function z = f(x, y) to be calculated within its domain of definition, and obtain the segmentation information seg of all segments in the xy plane xy ;

[0008] Step 2, using a multiplexer and a comparator, locate the segment to which the input variables x and y belong according to their values, and obtain the segmentation information of this segment accordingly;

[0009] Step 3, according to the segmentation information obtained in Step 2, use a Booth encoder, a compressor, and a carry-lookahead adder to calculate the value of the planar function of this segment, and use it as the value of the complex bivariate function to be calculated in the current segment, thereby completing the approximate calculation of the complex bivariate function.

[0010] Furthermore, the segmentation information in Step 1 includes:

[0011] Let the segmentation information of the i-th segment be set as Then the segmentation information of the i-th segment includes: the starting point, the ending point of variables x and y, and three parameters c x 、c y and c of the planar function gi(x, y) of the i-th segment for approximately calculating the complex bivariate function.

[0012] Furthermore, the segmentation in Step 1 includes:

[0013] Step 1-1, perform preliminary segmentation under a preset maximum absolute error constraint;

[0014] Step 1-2, optimize the segmentation information after preliminary segmentation.

[0015] Furthermore, the performing preliminary segmentation under a preset maximum absolute error constraint in Step 1-1 includes:

[0016] Step 1-1-1, set the domain of definition of the complex bivariate function z = f(x, y) as:

[0017] x ∈ [a, b]

[0018] y ∈ [c, d]

[0019] Step 1-1-2: Segment the independent variables x and y of the bivariate function z = f(x, y) in the xy-plane where they are located. Among them, segment evenly in the y-direction and unevenly in the x-direction. Obtain the optimal segmentation scheme under the constraint of the preset maximum absolute error, and segment according to this segmentation scheme.

[0020] Step 1-1-3: After segmenting the complex bivariate function z = f(x, y), in the i-th segment, use the planar function g i (x, y) for approximate calculation. In this segment, the independent variable x ∈ [a i , b i , and y ∈ [c i , d i . Here, a i and b i respectively represent the maximum and minimum values of the independent variable x of the planar function gi(x, y) in this segment, and c i and d i respectively represent the maximum and minimum values of the independent variable y in this segment.

[0021] Among them, the general expression of the planar function g(x, y) is as follows:

[0022] g(x, y) = c x ×x + c y ×y + c

[0023] Among them, c x , c y and c are the three parameters of the planar function; c x represents the undetermined coefficient of the independent variable x, c y represents the undetermined coefficient of the independent variable y, and c represents the constant term (i.e., the intercept term) of the planar function.

[0024] Furthermore, the optimization of the segmentation information after preliminary segmentation in Step 1-2 is to adjust the intercept term c0 of the planar function in each segment by fusing the Four-Point Traversal method FPT and the Taylor Series method TLS to reduce the maximum absolute error after approximate calculation using the planar function.

[0025] Furthermore, the adjustment of the intercept term c0 of the planar function in each segment in Step 1-2 includes:

[0026] Step 1-2-1: In each segment, use the Four-Point Traversal method FPT to calculate the three parameters of the planar function in this segment. Among them, assume the intercept term is c1, and adjust the intercept term to c1' according to the error.

[0027] Step 1-2-2: In each segment, use the Taylor Series method TLS to calculate the three parameters of the planar function in this segment. Among them, assume the intercept term is c2, and adjust the intercept term to c2' according to the error.

[0028] Step 1-2-3: Using the adjusted intercept terms c1' and c2', calculate the maximum absolute value MAE of the approximation error respectively, and take the intercept term when it is the smallest as the finally adjusted intercept term c'.

[0029] Furthermore, in Steps 1-2-1 and 1-2-2, the method of adjusting the intercept term according to the error includes:

[0030] Calculate the distribution of the error δ between the values of the plane function and the complex bivariate function for all points in the current segment, and determine the maximum positive value max(δ) and the minimum negative value mmin(δ) of the error δ;

[0031] Calculate the adjustment distance D as follows:

[0032]

[0033] Adjust the original intercept term c0 to the new intercept term c0' as follows:

[0034] c0' = c0 + D

[0035] ; The calculation method of the maximum absolute value MAE of the approximation error in Step 1-2-3 is as follows:

[0036] MAE = max(|δ|).

[0037] Furthermore, obtaining the optimal segmentation scheme under the preset maximum absolute error constraint described in Step 1-1-2 includes:

[0038] Step 1-1-2-1: Initialize the ranges of the independent variables x and y, and discretize them with a step size of 2 -i where iw is the number of precision digits; Initialize n xy to 0, indicating the number of segments on the xy plane;

[0039] Step 1-1-2-2: Segment evenly in the y direction. According to the preset number of segments n y in the y direction, calculate the number of points nsg y for each segment in the y direction, and determine the starting point ss y and the ending point se y ;

[0040] Step 1-1-2-3: Set the initial search range lf x = 1 and rg x = length(x); where, lf x represents the left interval of each segment in the x direction, rg x represents its right interval, and length(x) represents the length of the independent variable x;

[0041] Step 1-1-2-4, perform non-uniform segmentation in the x direction, and set the starting point ss of the first segment in the x direction x = 1, the ending point of the first segment where floor means rounding down, that is, rounding the input value down to the nearest integer;

[0042] Step 1-1-2-5, determine whether the segmentation starting point ss x is within the segmentation interval, that is, ss x < length(x). If so, execute Step 1-1-2-6; otherwise, jump to the next y-direction segmentation and execute Step 1-1-2-3 until the non-uniform segmentation in the x direction is completed for all y-direction segments, and then execute Step 1-1-2-11;

[0043] Step 1-1-2-6, determine the relationship between the segmentation starting point ss in the x direction x and the ending point se x . If they are not equal, generate coefficients using the fusion coefficient generation method; otherwise, generate the coefficients c x , c y and c of the plane function in this segment using the method compatible with piecewise plane approximation and piecewise linear approximation;

[0044] Step 1-1-2-7, obtain the plane function g(x, y) of the current x-direction segment according to the coefficients c x , c y and c, and calculate the maximum absolute error MAE between this plane function and the complex bivariate function f(x, y) to be calculated c ;

[0045] Step 1-1-2-8, determine the size relationship between MAE c and the preset threshold MAE d . If MAE c < MAE d , execute Step 1-1-2-9; otherwise, update the right interval rg x = se x , and recalculate the ending point se x of the current x-direction segment, and execute Step 1-1-2-6;

[0046] Step 1-1-2-9, determine the relationship between the ending point se x and the right interval rg x of the segment. If they are equal or the ending point se x is equal to the previous point of the right interval rg x of the segment, update the number of segments n on the xy planexy Increment the number of segments by 1 and store the segment information of the current segment in the x - direction. Then execute step 1 - 1 - 2 - 10. Otherwise, update the left interval lf of the current segment in the x - direction x = se x and recalculate the end point se of the current segment in the x - direction x and execute step 1 - 1 - 2 - 6;

[0047] Step 1 - 1 - 2 - 10: Initialize the parameters of the next segment in the x - direction. Set the starting point ss of the next segment in the x - direction x as the point after the end point se of the previous segment in the x - direction x and the left interval lf of the next segment in the x - direction x as the point two points after the end point of the previous segment in the x - direction. The right interval rg of the next segment in the x - direction x = length(x), and update the end point of the next segment in the x - direction Execute step 1 - 1 - 2 - 6;

[0048] Step 1 - 1 - 2 - 11: Complete the segmentation scheme and output the segmentation information seg xy .

[0049] Furthermore, the positioning of the segment to which it belongs according to the values of the input variables x and y in step 2 includes:

[0050] Let the inputs of the complex bivariate function to be calculated be x and y. According to the input value of y, locate the segment in the y - direction to which it belongs by the first multiplexer;

[0051] Use k comparators to locate the segment to which the input x belongs. The value of the input x is compared with the starting point of each segment in the x - direction respectively. If it is greater than or equal to, the comparator outputs 1 and compares with the starting point of the next segment. Otherwise, the comparator outputs 0, and finally determine the segment in the x - direction;

[0052] After determining the segment in the x - direction and the segment in the y - direction, use the second multiplexer to output the three parameters c x 、c u and c.

[0053] Furthermore, the calculation of the value of the plane function of this segment in step 3 includes:

[0054] Step 3 - 1: Use Booth encoding to generate the partial products of c x ×x and c y ×y, and at the same time use the sign - extension method to align the decimal points;

[0055] Step 3-2: The partial product result is compressed by a compressor, which consists of a half adder and a full adder, and is used to compress the partial product and the constant term c into a sum and a carry.

[0056] Step 3-3: The compressed partial products are added using a carry-lookahead adder CLA, and the result is added to the constant term c to obtain the planar function result.

[0057] Step 3-4: Keep the required significant digits of the planar function result to finally achieve the approximate calculation described above.

[0058] Beneficial effects:

[0059] The hardware implementation method for complex bivariate functions based on piecewise planar approximation provided by the present invention is applicable to high-precision calculations and supports the general hardware implementation of different bivariate functions. Compared with the state-of-the-art piecewise planar approximation method, this method saves 30.30%, 16.67%, and 47.24% in area, delay, and power consumption respectively while ensuring the calculation accuracy. Compared with the existing architectures based on two-dimensional interpolation and recursive segmentation, this method has considerable advantages. Description of the drawings

[0060] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0061] Figure 1 is the intercept adjustment method in the fusion coefficient generation method.

[0062] Figure 2 is the coefficient generation method based on the four-point traversal method.

[0063] Figure 3 is the pseudocode based on the adaptive segmenter.

[0064] Figure 4 is the general hardware implementation architecture based on piecewise planar approximation calculation.

[0065] Figure 5 is the calculation module principle based on the fused multiply-add - multiply-add unit. Detailed implementation manners

[0066] The overall idea of the present invention is as follows:

[0067] First, the complex bivariate function z = f(x, y) is divided into multiple small blocks within its domain (assuming the domain is x ∈ [a, b], y ∈ [c, d]) using the piecewise plane approximation method. The division method of the present invention is based on an adaptive segmentation strategy, and the range of each block segment is automatically determined according to the required precision of each block. Within each small block, a plane function is used to approximate the original complex bivariate function. Assuming that in the i-th small block, a plane function gi(x, y) is used to approximate the original function f(x, y), where x ∈ i , b i , y ∈ i , d i , a i , b i represents the range of x of the plane function g i (x, y) in the i-th small block, and c i , d i represents the range of y in the i-th small block. Each small block uses the above method to approximate the original complex bivariate function. Smaller blocks result in higher precision, but also increase the number of blocks, thereby increasing the storage consumption of the plane function coefficients.

[0068] The expression of the plane function g(x, y) used in the present invention is:

[0069] g(x, y) = c x ×x + c y ×y + c (1)

[0070] where c x , c y , c are three coefficients; c x represents the undetermined coefficient of x, c y represents the undetermined coefficient of y, and c represents the constant term (also called the intercept term) of the plane function; x and y are the two independent variables of the plane function. Regarding the three coefficients c x , c y , c, three points are needed for calculation. Assuming that 3 points are known, with coordinates (x1, y1, z1), (x2, y2, z2), and (x3, y3, z3) respectively, where z1 = f(x1, y1), z2 = f(x2, y2), z3 = f(x3, y3), the calculation methods of the three coefficients in formula (1) are:

[0071]

[0072] c0 = z1 - c x ×x1 - c y ×y1 (4)

[0073] where, in formula (4), c0 represents the value of c before intercept adjustment.

[0074] The principle of the Intercept Adjustment Method (IA) is as follows Figure 1 shown. First, methods such as the Four-Point Traversal Method (FPT) or the Taylor Series Method (TLS) are used to calculate the coefficients of the plane function, including c x and c y and the original intercept term c0, obtaining the plane function expression g0(x, y) = c x ×x + c y ×y + c0 before optimization. Since the error distribution is non-uniform in the positive and negative parts at this time, it is considered to unify the positive and negative error distributions by optimizing the intercept value, thereby reducing the Maximum Absolute Error (MAE). Therefore, by calculating the distribution of the error (δ), the maximum positive value (maxδ) and the minimum negative value (minδ) of the error are determined. Then, an adjustment distance D is calculated, which is the average of the maximum positive error and the minimum negative error, that is Finally, the original intercept term (c0) is adjusted to a new intercept term c1, that is, c1 = c0 + D, obtaining the optimized plane function expression g1(x, y) = c x ×x + c y ×y + c1. Figure 1 In it, min(δ′) and max(δ′) represent the maximum positive value and the minimum negative value of the error of the optimized plane function.

[0075] Based on the principle of the Four-Point Traversal Method (FPT), four combinations of four vertices are traversed to calculate four plane coefficients. Then, the four plane functions are optimized by the Intercept Adjustment (IA) to reduce the MAE. Finally, the plane function with the smallest MAE after optimization is selected to approximate the complex bivariate function. The coefficient generation method based on FPT is as follows Figure 2 shown. In this method, the complex bivariate function is adaptively divided into multiple small blocks, and within each small block, a plane function is used to approximate the original complex bivariate function. The core idea of FPT is to start from the four vertices of each small block, and calculate the coefficients c x , c y and c0 of the plane function that can best approximate the function within the small block through these vertices. In Figure 2 , it is assumed that the four vertices of a small block are known as (x1, y1, z1), (x2, y2, z2), (x3, y3, z3) and (x4, y4, z4), where z1 = f(x1, y1), z2 = f(x2, y2), z3 = f(x3, y3), z4 = f(x4, y4). By randomly and without repetition taking three points from the four vertices, the three coefficients under different vertex combinations are calculated respectively according to formulas (2)(3)(4), and the coefficients of the plane function with the smallest MAE are selected as the final coefficients of FPT.

[0076] The coefficient expression of the Taylor Series Method TLS is as follows:

[0077]

[0078] where represents the partial differential symbol;

[0079] Usually, (x0, y0) is the center of the segment in the x - y plane. Similar to IA FPT, TLS optimizes the intercept value c through the intercept adjustment method IA.

[0080] The Fusion Coefficient Generation Method (FCGM) combines the two methods of IA FPT and IA TLS (Reference: Dong, 2020, TVLSI, PLAC: Piecewise linear approximation computation for all nonlinear unary functions). After calculating the plane coefficients by the two methods, the plane approximation MAE is calculated, and finally the coefficient with the smallest MAE is selected as the result of the Fusion Coefficient Generation Method FCGM. The plane approximation MAE of x and y in the range of [1, 1.125] based on different coefficient generation methods is shown in Table 1.

[0081] Table 1 Plane approximation MAE table based on different coefficient generation methods

[0082]

[0083] Although the coefficients generated by the Fusion Coefficient Generation Method FCGM improve the accuracy of the piecewise plane approximation through function (1), in some high - precision cases, even when the size of the block is already minimized, the block still cannot converge. At this time, the convergence is improved by making the piecewise plane approximation method compatible with the piecewise linear approximation method (PWL). The specific measures are as follows: when the minimum block in the x - direction cannot meet the accuracy requirements, by further shortening the division boundary in the x - direction, that is, moving the end - point of the boundary to the starting point, and then by readjusting the segmentation in the y - direction to further compress the size in the x - direction until the accuracy requirements are met. At this time, the coefficient generation method in the piecewise linear approximation method (PWL) needs to be used to recalculate the coefficients. For the segmentation adjustment in the y - direction, assuming that the value of x is fixed as x0, and the adjusted boundary range in the y - direction is y ∈ [y1, y2], then two boundary points (x0, y1, z1) and (x0, y2, z2) are obtained, where z1 = f(x0, y1) and z2 = f(x0, y2). Referring to the coefficient generation method of PWL, the coefficient expression is:

[0084] c x = 0 (8)

[0085]

[0086] c0 = z1 - cy ×y1 (10)

[0087] Optimize c0 to c1 by the intercept adjustment method IA.

[0088] Next, the adaptive segmenter will be introduced. The pseudocode of the proposed segmenter is as Figure 3 shown. First, relevant parameters need to be initialized. Among them, st represents the starting point of this segment, which is equal to the point after the end point of the previous segment; ed represents the end point of the segment; the input f is used to determine the type of the bivariate function; the predefined MAE d is used to limit the acceptable approximation error for hardware implementation; n y represents the number of segments in the y direction; iw is the number of precision bits for the input and output; qx, qy, and qc are the fractional bits of the coefficients cx, cy, and c respectively. Since x and y have iw fractional bits, the values in the x direction and y direction are not continuous, and the interval between values is 2 -iw , so in lines 1 - 2 of the pseudocode, the ranges of x and y need to be discretized. At the same time, n xy is initialized to 0 to store the number of segments in the x - y plane.

[0089] The segmentation of the x - y plane is divided into two steps: segmentation in the y direction and segmentation in the x direction. It is known that uniform segmentation is adopted in the y direction. Through a for loop, calculate the number of points nsg y in each y segment and the starting point ss y and the end point se y of each segment. After the segmentation in the y direction is completed, non - uniform segmentation in the x direction is performed using a for loop and two while loops. The inner for loop performs the segmentation in the x direction for all y segments. The middle while loop always executes until all x are segmented. The inner while loop is used to find the optimal end point in the x direction for the current y segment. Each time the i - th segment in the y direction is processed, the starting point ss x of the first segment in the x direction is set to 1. Given a binary search window [lf x , rg x = [1, length(x)] to limit the possible range of the end point se x . Line 12 of the pseudocode stipulates that the value of se x is always at the center of the binary search window. The middle while loop starts with the condition ss x ≤ length(x), indicating that the starting point ss x of the new segment is within the range [st, ed]. The control variable F is set to 1 to control the inner while loop, which is used to determine the optimal se xValue. Lines 16 - 23 in the pseudocode calculate the MAE by generating coefficients, simulating the hardware implementation, and calculating the exact value and the approximate value through piecewise planar approximation. c . Then, according to the predefined MAE d and the calculated MAE c , adjust the end point se of the current segment in the x - direction x . If MAE c ≤MAE d , it indicates that the current segment meets the accuracy requirements. At this time, if se x is equal to rg x , it means that the binary window cannot be expanded further, then the current segment is determined, and the widest segment under the constraint of MAE d in the x - direction is obtained. At this time, set F = 0 to complete the segmentation in the current x - direction, and store the starting point, ending point, and coefficients of the current segment; if se x is equal to rg x -1, it means that the x - direction can be further expanded, then move the binary window and update the end point se of the current segment x to the right, and continue the new internal while loop. If MAE c ≥MAE d , it indicates that the segment in the current x - direction is too wide to meet the accuracy requirements. At this time, adjust the binary window and update the end point se of the current segment x to the left, and continue the new internal while loop. After completing the internal while loop, initialize the starting point of the next segment in the x - direction as the last point of the current segment, and update lf x , rg x and the end point se x to start a new segmentation in the x - direction.

[0090] Next, the hardware calculation unit is described. The hardware architecture is as Figure 4 shown. The proposed hardware architecture includes a coefficient selector and a calculation module. The coefficient selector locates the segment to which the input x and y belong and determines the coefficients of the corresponding segment according to the adaptive segmenter. Figure 4 The specific process is as follows: First, the coefficient selector takes all the segment start points (Start Points) generated by the adaptive segmenter and the coefficients (Coefficients) of each segment as hardware inputs. Since the y - direction is uniformly segmented and there are a total of n = 2 in the y - direction msegments. Therefore, a y segment is located by using an n:1 multiplexer, which uses the m most significant bits (MSBs) of y to determine the y-direction specific segment. After the y segment is determined, all the x segment coefficients within the current y segment are stored in a second multiplexer. Then, multiple comparators are used in the coefficient selector to locate the x segment. The comparator determines which x segment the actual value is specifically located in by comparing the actual value of the input x with the start points of each x segment. If the actual value is greater than or equal to the start point of the current segment, the comparator outputs 1 and compares it with the start point of the next segment; otherwise, the comparator outputs 0. After determining the specific x segment and y segment, the second multiplexer is used to output the three coefficients of the plane function.

[0091] The calculation module first uses two Booth encoders to generate the partial products of c x ×x and c y ×y, reducing the number of partial products. At the same time, the sign extension technique is used to align the decimal points. Secondly, the partial product results are compressed by a compressor, which consists of multiple half adders and full adders, and is used to compress the partial products and the constant term c into a sum and a carry. The compressed partial products are added by a carry look-ahead adder (CLA), and the result is added to the constant term c to obtain the result of the complex bivariate function.

[0092] The calculation module introduces the fused multiply-add multiply-add strategy (FMAMA), as Figure 5 shown. First is the generation of partial products. The partial products of c x ×x and c y ×y are generated using the principle of Booth Encoding, and the number of bits of the partial products is reduced by discarding the highest bit. In addition, sign extension of the partial products of c x ×x and c y ×y; according to the fractional bits qx, qy, and qc of the coefficients c x 、c y and c, the positions of the partial products and the constant term c are adjusted so that their decimal points can be correctly aligned. Then, the compressor is used to compress the aligned partial products and c into a sum and a carry. Finally, the compressed sum and carry are input into the carry look-ahead adder (CLA) to quickly calculate the final result, and the result is rounded or truncated according to the required precision requirements to ensure that the fractional bits of the result meet the requirements. The final calculation result only retains the required iw significant digits, which is completed by accurately calculating the required number of integer bits. The number of integer bits is defined as:

[0093]

[0094] This method uses Verilog HDL for modeling and is synthesized by Synopsys Design Compiler (DC) using 28nm and 65nm CMOS technologies. Table 2 shows the comparison results between this method and other piecewise planar approximation methods after synthesis based on 28nm CMOS technology, including the optimization results in terms of storage space, area, delay, and power consumption, etc.;

[0095] Table 2 Comparison Results Table of This Design and Other Piecewise Planar Approximation Methods

[0096]

[0097] Table 3 shows the comparison results between this method and 2D interpolation after synthesis based on 65nm CMOS technology. The results show that this method saves 93.36%, 72.99%, and 60.40% respectively in terms of storage space, area, and delay.

[0098] Table 3 Comparison Results Table of This Design and 2D Interpolation

[0099]

[0100] The comprehensive results of the comparison design between this method and the methods based on Piecewise Linear (PWL) and Coordinate Rotation Digital Computer (CORDIC) are shown in Table 4.

[0101] Table 4 Comprehensive Results Table of the Comparison Design of This Design and Other Methods

[0102]

[0103]

[0104] In specific implementation, this application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the content of the invention of a hardware implementation method for a complex bivariate function based on piecewise planar approximation calculation provided by the present invention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0105] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the parts that contribute to the prior art can be embodied in the form of a computer program, that is, a software product. This computer program software product can be stored in a storage medium, including several instructions to enable a device including a data processing unit (which can be a personal computer, server, single-chip microcomputer, MCU or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0106] The present invention provides an idea and method for hardware implementation of a complex bivariate function based on piecewise planar approximation calculation. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A hardware implementation method for complex two-variable functions based on piecewise plane approximate calculation, characterized in that: The following steps are involved: Step 1: Use the segmented plane approximation method to segment the complex two-variable function z=f(x, y) to be calculated within its domain, and obtain the segmentation information seg of all segments in the xy plane. xy ; Step 2, using a multiplexer and a comparator, locate the segment to which the segment belongs according to the values ​​of the input variables x and y, and obtain the segment information of the segment accordingly; Step 3, based on the segmentation information obtained in step 2, use the Booth encoder, compressor and carry-lookahead adder to calculate the value of the plane function of the segment, as the value of the complex two-variable function to be calculated in the current segment, to complete the approximate calculation of the complex two-variable function.

2. The method for hardware implementation of complex two-variable function based on piecewise plane approximate calculation according to claim 1, characterized in that: The segmentation information described in step 1, including: Let the i-th segment information be Then the i-th segment information In the example, the starting point and the ending point of the variables x and y and the i-th plane function g for approximating the complex two-variable function are included. i The three parameters c of (x, y) x 、c y and c.

3. The method for hardware implementation of complex two-variable function based on piecewise plane approximate calculation according to claim 2, characterized in that: The segmentation described in step 1, including: Step 1-1, perform preliminary segmentation under a preset maximum absolute error constraint; Step 1-2, optimizing the segmentation information after the preliminary segmentation.

4. The method for hardware implementation of complex two-variable function based on piecewise plane approximate calculation according to claim 3, characterized in that: The preliminary segmentation described in step 1-1 is performed under the preset maximum absolute error constraint, including: Step 1-1-1, assume that the domain of the complex two-variable function z=f(x, y) is: x∈[a,b] y∈[c,d] Step 1-1-2, segment the independent variables x and y of the two-variable function z=f(x, y) in the xy plane where they are located, wherein the y direction is segmented uniformly and the x direction is segmented non-uniformly, and the optimal segmentation scheme is obtained under the preset maximum absolute error constraint, and segmentation is performed according to the segmentation scheme; Step 1-1-3, after dividing the complex two-variable function z=f(x, y) into segments, in the i-th segment, a plane function g is used i (x, y) to approximate the calculation, the independent variable x∈[a i , b i ],y∈[c i , d i ], a i and b i Respectively represent the plane function g in this segment i The maximum and minimum values ​​of the independent variable x of (x, y), c i and d i Respectively represent the maximum and minimum values ​​of the independent variable y in the segment; Among them, the general expression of the plane function g(x, y) is as follows: g(x,y)=c x ×x+c y ×y+c Among them, c x 、c y and c are the three parameters of the plane function; c x represents the unknown coefficient of the independent variable x, c y represents the unknown coefficient of the independent variable y, and c represents the constant term of the plane function, that is, the intercept term.

5. The method for hardware implementation of complex two-variable function based on piecewise plane approximate calculation according to claim 4, characterized in that: The segmentation information after the preliminary segmentation described in step 1-2 is optimized, that is, by integrating the four-point traversal method FPT and the Taylor series method TLS, the intercept term c0 of the plane function in each segment is adjusted to reduce the maximum absolute error after approximate calculation using the plane function.

6. The method for hardware implementation of complex two-variable function based on piecewise plane approximate calculation according to claim 5, characterized in that: The adjustment of the intercept term c0 of the plane function in each segment described in step 1-2 includes: Step 1-2-1, in each segment, use the four-point traversal method FPT to calculate the three parameters of the plane function of the segment, where the intercept term is set to c1, and the intercept term adjusted according to the error is c1'; Step 1-2-2, in each segment, use the Taylor series method TLS to calculate the three parameters of the plane function of the segment, where the intercept term is set to c2, and the intercept term adjusted according to the error is set to c2'; Step 1-2-3, use the adjusted intercept terms c1' and c2' to calculate the maximum absolute value of the approximate error MAE respectively, and take the intercept term with the minimum value as the final adjusted intercept term c'.

7. The method for hardware implementation of complex two-variable function based on piecewise plane approximate calculation according to claim 6, characterized in that: In step 1-2-1 and step 1-2-2, the method of adjusting the intercept term according to the error includes: Calculate the distribution of the error δ between the values ​​of the plane function and the values ​​of the complex two-variable function of all points in the current segment, and determine the maximum positive value max(δ) and the minimum negative value min(δ) of the error δ; Calculate the adjustment distance D as follows: Adjust the original intercept term c0 to the new intercept term c0' as follows: c0'=c0+D The maximum absolute value MAE of the approximate error in step 1-2-3 is calculated as follows: MAE=max(|δ|).

8. The method for hardware implementation of complex two-variable function based on piecewise plane approximate calculation according to claim 7, characterized in that: The optimal segmentation scheme described in step 1-1-2 is obtained under the preset maximum absolute error constraint, including: Step 1-1-2-1, initialize the range of the independent variables x and y, and divide them into 2 -i Discretize the step size, where iw is the number of digits of precision; initialize n xy 0, indicating the number of segments on the xy plane; Step 1-1-2-2, evenly segment in the y direction, according to the preset number of segments n in the y direction y , calculate the number of points nsg in each segment in the y direction y , and determine the starting point ss of each segment y and end point y ; Step 1-1-2-3, set the initial search range lf in the x direction x =1 and rg x =length(x); where lf x represents the left interval of each segment in the x direction, rgx represents its right interval, and length(x) represents the length of the independent variable x; Step 1-1-2-4, non-uniform segmentation in the x direction, set the starting point ss of the first segment in the x direction x =1, the end point of the first segment Among them, floor means rounding down, that is, rounding the input value down to the nearest integer; Step 1-1-2-5, determine the segment starting point ss x Whether it is within the segment interval, that is, ss x <length(x), if so, execute step 1-1-2-6, otherwise jump to the next y-direction segment and execute step 1-1-2-3 until all y-direction segments have completed the non-uniform segmentation in the x-direction, and execute step 1-1-2-11; Step 1-1-2-6, determine the segment starting point ss in the x direction x With the end point x If the two are not equal, the fusion coefficient generation method is used to generate the coefficients; otherwise, the coefficients c of the plane function in the segment are generated using a method compatible with the piecewise plane approximation and the piecewise linear approximation. x 、c y and c; Step 1-1-2-7, according to the coefficient c x 、c y and c, obtain the segmented plane function g(x, y) in the current x direction and calculate the maximum absolute error MAE between the plane function and the complex two-variable function f(x, y) to be calculated c ; Step 1-1-2-8, determine MAE c Compared with the preset threshold MAE d If MAE c <MAE d , then execute step 1-1-2-9, otherwise update the right interval rg of the current x-direction segment x =se x , and recalculate the end point se of the current x-direction segment x , execute steps 1-1-2-6; Step 1-1-2-9, determine the end point se x With the segment right interval rg x If the two are equal or the end point is se x Equal to the right interval rg x The previous point of , then update the number of segments n on the xy plane xy , add 1 to the number of segments and store the segment information of the current segment in the x direction, and execute steps 1-1-2-10, otherwise update the left interval lf of the current segment in the x direction x =se x , and recalculate the end point se of the current x-direction segment x , execute steps 1-1-2-6; Step 1-1-2-10, initialize the parameters of the next segment in the x direction and set the starting point ss of the next segment in the x direction x The end point of the previous segment in the x direction, se x The next point after the x-direction segment left interval lf x The last two points of the end point of the previous x-direction segment, the right interval rg of the next x-direction segment x =length(x), update the end point of the next segment in the x direction Follow steps 1-1-2-6; Step 1-1-2-11, complete the segmentation plan and output segmentation information seg xy .

9. The method for hardware implementation of complex two-variable function based on piecewise plane approximate calculation according to claim 8, characterized in that: The step 2 of locating the segment according to the values ​​of the input variables x and y includes: Assuming that the inputs of the complex two-variable function to be calculated are x and y, the first multiplexer locates the segment in the y direction according to the value of the input y; Use k comparators to locate the segment to which the input x belongs. The value of the input x is compared with the starting point of each segment in the x direction. If it is greater than or equal to, the comparator outputs 1 and compares it with the starting point of the next segment. Otherwise, the comparator outputs 0, and finally determines the segment in the x direction. After determining the segments in the x direction and the segments in the y direction, use the second multiplexer to output the three parameters c of the plane function x 、c y and c.

10. The method for hardware implementation of complex dual variable functions based on piecewise plane approximate calculation according to claim 9, characterized in that: The calculation described in step 3 obtains the value of the plane function of the segment, including: Step 3-1, use Booth coding to generate c x ×x and c y ×y partial product, using the sign extension method to align the decimal points; Step 3-2, the partial product result is compressed by a compressor, wherein the compressor is composed of a half adder and a full adder, and is used to compress the partial product and the constant term c into a sum and a carry; Step 3-3, the compressed partial products are added using a carry-lookahead adder CLA, and the result is added to the constant term c to obtain a plane function result; Step 3-4, retaining the required number of significant digits of the plane function result, and finally realizing the approximate calculation.