Resource reuse type transcendental function calculation device and calculation method
Through resource multiplexing type transcendent function computing device and pipeline calculation method, the problem of high hardware resource consumption in the prior art is solved, and efficient and accurate transcendent function computing of various data types is realized.
Patent Information
- Application Number
- CN202510360218.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art requires a large amount of hardware resources when calculating transcendent functions, especially when facing multiple data types, the lookup table is too large and it is difficult to calculate efficiently and accurately.
The resource multiplexing type transcendent function calculation device is adopted, through the combination of selector, processing unit PE and result output unit, the multiplexing of lookup table, fixed-point multiplier, floating-point multiplier and floating-point adder, combined with pipeline calculation methods, the transcendent function calculation of various data types is realized.
It effectively reduces hardware resource consumption, reduces on-chip area, and supports multiple data types to transcend function calculations, improving computing efficiency and accuracy.
Smart Images

Figure CN120255848A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to transcendental function calculation, and specifically to a resource - reused transcendental function calculation device and calculation method. Background Art
[0002] A transcendental function refers to a function whose result cannot be obtained through a finite number of four - arithmetic operations and algebraic operations in calculation. Such functions are currently widely used in high - computing - power fields such as deep - learning models, digital signal processing, and dedicated artificial - intelligence accelerators. Common transcendental functions include trigonometric functions, erf functions, natural exponential functions, logarithmic functions, etc.
[0003] In modern general - purpose processors, transcendental functions generally use an iterative method to calculate approximate values, such as using algorithms like Taylor series expansion, Coordinate Rotation Digital Computer (CORDIC), etc. While this iterative method ensures obtaining high - precision calculation results, it often requires more hardware resources such as multipliers and adders in hardware implementation.
[0004] In artificial - intelligence accelerators, the direct - look - up - table method is also widely used in the calculation of transcendental functions. However, in the case of a relatively high input - data bit - width, when facing multiple data types (such as 16 - bit floating - point numbers and 32 - bit floating - point numbers), this method often results in an overly large look - up table in hardware implementation. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] In view of the above - mentioned drawbacks of the prior art, the present invention provides a resource - reused transcendental function calculation device and calculation method, which can effectively overcome the defect that it is difficult to accurately calculate transcendental functions with fewer hardware resources in the prior art.
[0007] (2) Technical Solutions
[0008] To achieve the above object, the present invention is realized through the following technical solutions:
[0009] A resource - reused transcendental function calculation device includes a selector, multiple processing elements (PEs), and a result output unit;
[0010] The selector receives input data, data types, and control signals, and sends the input data and data types to the corresponding processing element (PE) according to the control signals;
[0011] The processing element (PE) can process all types of transcendental functions and perform transcendental function calculations on the input data according to the data types;
[0012] Result output unit, which concatenates the calculation results of multiple processing units PE in sequence to obtain output data;
[0013] Among them, when all processing units PE perform transcendental function calculations, a lookup table, a fixed-point multiplier, a floating-point multiplier, and a floating-point adder are reused.
[0014] Preferably, the processing unit PE includes a preprocessing unit, a lookup table unit, a fitting calculation unit, and a postprocessing unit;
[0015] The preprocessing unit does not process 8-bit fixed-point numbers and directly sends the data itself as the address index of the lookup table to the lookup table unit; for 16-bit / 32-bit floating-point numbers, it performs processing to obtain the sign bit, exponent bit, and mantissa bit, directly sends the sign bit to the postprocessing unit, quantizes the exponent bit and mantissa bit to obtain the integer bit and decimal bit, extracts the address index of the lookup table from the integer bit and decimal bit according to the corresponding index address extraction method for the type of transcendental function, and sends the address index of the lookup table to the lookup table unit, and sends the remaining bits of the decimal bit as the remaining mantissa to the fitting calculation unit;
[0016] The lookup table unit performs a lookup table operation according to the address index of the lookup table. For 8-bit fixed-point numbers, the corresponding fixed-point result is obtained through the lookup table operation and output; for 16-bit / 32-bit floating-point numbers, the function value and the difference are obtained through the lookup table operation, and the function value and the difference are sent to the fitting calculation unit;
[0017] The fitting calculation unit calculates the fixed-point fitting result according to the remaining mantissa, function value, and difference, and sends the fixed-point fitting result to the postprocessing unit;
[0018] The postprocessing unit normalizes the fixed-point fitting result according to the data type to obtain the corresponding floating-point result, and sends the floating-point result to the result output unit.
[0019] Preferably, the preprocessing unit quantizes the exponent bit and mantissa bit to obtain the integer bit and decimal bit, including:
[0020] According to the standard representation form of floating-point numbers, a hidden integer 1 is added to the fractional part of 16-bit / 32-bit floating-point numbers, and it is converted into the form of 1.N;
[0021] Subtract an offset 2 from the exponent bit n-1-1 to obtain the true exponent. Determine the number of digits for the decimal point shift according to the true exponent. When the true exponent is positive, move the decimal point of 1.N to the left by the corresponding number of digits; when the true exponent is negative, move the decimal point of 1.N to the right by the corresponding number of digits to make the true exponent become 0, so as to align the fractional part of the 16-bit / 32-bit floating-point number to the decimal form and obtain the quantized number.
[0022] Determine the integer part and the fractional part of the 16-bit / 32-bit floating-point number according to the quantized number.
[0023] Where N is the number of mantissa bits of the current binary number.
[0024] n is the exponent bit width of the 16-bit / 32-bit floating-point number. For the 16-bit floating-point number of data type bfp16 and the 32-bit floating-point number of data type fp32, the exponent bits both occupy 8 bits and the offsets are both 127; for the 16-bit floating-point number of data type fp16, the exponent bits occupy 5 bits and its offset is 15.
[0025] Preferably, the look-up table is composed of 256 SRAMs with a width of 8 bits and a depth of 256. For 8-bit fixed-point numbers, the corresponding fixed-point results are directly stored in 256 look-up tables.
[0026] For 16-bit / 32-bit floating-point numbers, 4 SRAMs of 8 bits are spliced to form 64 look-up tables capable of storing 32-bit data. The function values and differences are stored in the look-up tables. The first 16 bits of the 32-bit data are function values and the last 16 bits are differences.
[0027] Among them, the data stored in the look-up table are all stored in the 16-bit fixed-point format. Compared with the floating-point format storage, the fixed-point format storage does not require the sign bit and exponent bit of the floating-point number, and all bits represent the mantissa bits of the floating-point number, which improves the bit utilization rate of 16-bit data and greatly improves the representation accuracy.
[0028] Preferably, for 8-bit fixed-point numbers, the processing unit PE can query 256 8-bit fixed-point numbers in one clock cycle, and the table values queried are the corresponding fixed-point results.
[0029] For 32-bit floating-point numbers, the processing unit PE can perform 64 32-bit floating-point number calculations in one clock cycle.
[0030] For 16-bit floating-point numbers, the processing unit PE can perform calculations on 64 16-bit floating-point numbers in one clock cycle. For 2048-bit input data, when divided into 16-bit data types, it can be divided into 128 16-bit floating-point numbers. If the processing unit PE needs to perform calculations on 128 16-bit floating-point numbers simultaneously in one clock cycle, 128 lookup tables are required, that is, the hardware resources of the lookup tables need to be doubled. Therefore, when performing calculations on 16-bit floating-point numbers, considering factors such as consumption and area, without increasing hardware resources, the input data of 16-bit data type is input in two clock cycles to complete the calculation of 128 16-bit floating-point numbers.
[0031] Preferably, the fitting calculation unit calculates the fixed-point fitting result according to the remaining mantissa, function value, and difference, including:
[0032] According to the remaining mantissa, function value, and difference, the fixed-point fitting result is calculated using a fixed-point multiplier in a linear fitting manner, that is, when the value x to be calculated falls within [x k , x k+1 , assuming the distance between x k and x k+1 is 1, the distance of x from x k is t, and the distance of x from x k+1 is 1 - t, then the function value y corresponding to x is:
[0033] y = t(y k+1 - y k ) + y k ;
[0034] Among them, y k , y k+1 are the function values corresponding to x k , x k+1 respectively, y k+1 - y k is the difference between the function values corresponding to x k , x k+1 , and t is the remaining mantissa.
[0035] Preferably, for different types of transcendental functions, the bit width of the remaining mantissa t retained after extracting the address index of the lookup table is different. In order to facilitate the reuse of the fixed-point multiplier, it is necessary to unify the bit width of the remaining mantissa t. Considering that the maximum bit width of the remaining mantissa t retained after extracting the address index of the 16-bit / 32-bit floating-point number lookup table is 19 bits, and at the same time, to ensure the fitting calculation accuracy and avoid the accumulation of precision loss of intermediate results in subsequent calculations due to rounding processing, the bit width of all remaining mantissas t is uniformly extended to 19 bits;
[0036] The data stored in the lookup table is all stored in 16-bit fixed-point format. Therefore, a 19 * 16-bit fixed-point multiplier is used to calculate t(y k+1 -y k ), and the calculation result is processed into a 32-bit fixed-point number. Then, the function value y k obtained through the lookup table operation is extended to a 32-bit fixed-point number by padding with zeros. Finally, the two 32-bit fixed-point numbers are added to obtain a 32-bit fixed-point fitting result.
[0037] Preferably, the post-processing unit normalizes the fixed-point fitting result according to the data type to obtain the corresponding floating-point result, including:
[0038] Determine the position of the highest 1 in the fractional bits of the fixed-point fitting result, and use this highest 1 as the hidden bit. The numbers after the hidden bit are truncated as the fractional bits of the floating-point result;
[0039] Determine the true exponent according to the position of the hidden bit, and add it to the offset corresponding to the data type of the floating-point number to obtain the exponent bit of the floating-point result;
[0040] Process the sign bit sent by the preprocessing unit according to the type of transcendental function to obtain the corresponding floating-point result;
[0041] Among them, for 16-bit floating-point numbers of the data type bfp16, the high 7 bits after the hidden bit are truncated as the fractional bits of the floating-point result; for 16-bit floating-point numbers of the data type fp16, the high 10 bits after the hidden bit are truncated as the fractional bits of the floating-point result; for 32-bit floating-point numbers of the data type fp32, the high 23 bits after the hidden bit are truncated as the fractional bits of the floating-point result;
[0042] For 16-bit / 32-bit floating-point numbers, if the number of bits after the hidden bit is not enough, zeros are directly padded. If there are remaining bits after truncating the fractional bits, rounding is performed using the round-to-even method: when the highest bit of the remaining bits is 1, if the remaining bits are not all 0, carry; if the remaining bits are all 0, judge according to the lowest bit of the truncated fractional bits. If the lowest bit is not 0, carry; if the lowest bit is 0, discard; when the highest bit of the remaining bits is 0, directly discard.
[0043] Preferably, for some types of transcendental functions, during the preprocessing in the preprocessing unit or the postprocessing in the postprocessing unit, multiplication and addition calculations are performed using a 32-bit floating-point multiplier and a 32-bit floating-point adder;
[0044] Among them, for 16-bit floating-point numbers, after the multiplication-addition calculation is completed during the post-processing in the post-processing unit, it is necessary to convert the 32-bit multiplication-addition calculation result into 16 bits to obtain the corresponding floating-point result.
[0045] A resource multiplexing transcendental function calculation method implements pipelined data pipelining calculation, specifically including the following processes:
[0046] DC0 level: Preprocess the input data according to the data type to obtain the sign bit, the address index of the lookup table, and the remaining mantissa;
[0047] DC1 level: Perform a lookup operation according to the address index of the lookup table;
[0048] EX0 level: For 8-bit fixed-point numbers, obtain the corresponding fixed-point result through a lookup operation and output it; for 16-bit / 32-bit floating-point numbers, obtain the function value and the difference through a lookup operation;
[0049] EX1 level: According to the remaining mantissa and the difference, use a fixed-point multiplier to calculate t(y k+1 -y k );
[0050] EX2 level: Combine the function value to calculate the fixed-point fitting result y = t(y k+1 -y k ) + y k ;
[0051] EX3 level: Normalize the fixed-point fitting result according to the data type to obtain the corresponding floating-point result;
[0052] WB level: Concatenate the calculation results of multiple processing units PE in sequence to obtain the output data;
[0053] Among them, when all processing units PE perform transcendental function calculations, they multiplex the lookup table, fixed-point multiplier, floating-point multiplier, and floating-point adder;
[0054] The DC0-DC1 levels represent the processing process of the input data, the EX0-EX3 levels represent the data calculation execution process, the calculation results are obtained through corresponding calculations, and the WB level represents the collation and output of the calculation results of multiple processing units PE;
[0055] The above process adopts a way of pipeline division with unequal lengths: for transcendental functions of the same type, there are differences between the pipelines corresponding to 16-bit floating-point numbers and 32-bit floating-point numbers. Since the 16-bit floating-point numbers are input in two clock cycles, one more stage is added between the DC0-DC1 levels compared to 32-bit floating-point numbers to process the input 16-bit floating-point numbers; for some types of transcendental functions that need to perform multiply-add calculations during the preprocessing in the preprocessing unit or the postprocessing in the postprocessing unit, corresponding extensions exist between the DC0-DC1 levels and the EX0-EX3 levels; for 8-bit fixed-point numbers, they directly enter the WB level after obtaining the corresponding fixed-point results through table lookup operations.
[0056] (III) Beneficial effects
[0057] Compared with the prior art, a resource-reusable transcendental function calculation device and calculation method provided by the present invention have the following beneficial effects:
[0058] 1) The present invention proposes a resource-configurable implementation scheme for transcendental function calculation, and realizes various transcendental function calculations through instruction operations, while supporting input data of multiple data types such as uint8, int8, fp16, bfp16, and fp32;
[0059] 2) The present invention adopts a hardware implementation method of resource reuse, and effectively reduces the consumption of hardware resources and reduces the on-chip area by reusing lookup tables, fixed-point multipliers, floating-point multipliers, and floating-point adders. Description of the drawings
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0061] Figure 1 It is a system schematic diagram of the present invention;
[0062] Figure 2 For the present invention Figure 1 It is a schematic diagram of the hardware connection of the processing unit PE in the present invention;
[0063] Figure 3 For the present invention Figure 2 It is a schematic diagram of the working process of the preprocessing unit in the present invention;
[0064] Figure 4 It is a schematic diagram of the lookup table setting for 8-bit fixed-point numbers in the present invention;
[0065] Figure 5 Schematic diagram of the look-up table setting for 16-bit / 32-bit floating-point numbers in the present invention;
[0066] Figure 6 Schematic diagram of the process of the present invention. Detailed implementation manners
[0067] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0068] A resource multiplexing transcendental function calculation device, as Figure 1 shown, includes a selector, a plurality of processing units PE, and a result output unit;
[0069] The selector receives input data, data type, and a control signal, and sends the input data and data type to the corresponding processing unit PE according to the control signal;
[0070] The processing unit PE can process all types of transcendental functions and perform transcendental function calculations on the input data according to the data type;
[0071] The result output unit splices the calculation results of the plurality of processing units PE in sequence to obtain output data;
[0072] Among them, when all the processing units PE perform transcendental function calculations, a look-up table, a fixed-point multiplier, a floating-point multiplier, and a floating-point adder are multiplexed.
[0073] In the technical solution of the present application, both the input data and the bit width of the input data are 2048 bits, and the data types include uint8, int8, fp16, bfp16, and fp32. According to the data type, the device can perform calculations on 256 8-bit or 128 16-bit or 64 32-bit data.
[0074] As Figure 1 shown, for the convenience of data processing, in the top layer, the transcendental function calculation unit is divided into 64 processing units PE. Each processing unit PE includes a control signal interface, an input data interface, and an output data interface. The bit widths of the input data interface and the output data interface are both 32 bits, and the 2048-bit input data is divided into 64 32-bit data and sent to the 64 processing units PE respectively.
[0075] As Figure 2As shown, the processing unit PE includes a preprocessing unit, a look-up table unit, a fitting calculation unit, and a postprocessing unit;
[0076] The preprocessing unit does not process 8-bit fixed-point numbers and directly sends the data itself as the address index of the look-up table to the look-up table unit; for example, Figure 3 As shown, for 16-bit / 32-bit floating-point numbers, it is processed to obtain the sign bit, exponent bit, and mantissa bit. The sign bit is directly sent to the postprocessing unit. The exponent bit and mantissa bit are quantized to obtain the integer bit and decimal bit. According to the type of transcendental function, the corresponding index address extraction method is used to extract the address index of the look-up table from the integer bit and decimal bit, and the address index of the look-up table is sent to the look-up table unit. The remaining bits of the decimal bit are sent to the fitting calculation unit as the remaining mantissa;
[0077] The look-up table unit performs a look-up table operation according to the address index of the look-up table. For 8-bit fixed-point numbers, the corresponding fixed-point result is obtained through the look-up table operation and output; for 16-bit / 32-bit floating-point numbers, the function value and difference are obtained through the look-up table operation, and the function value and difference are sent to the fitting calculation unit;
[0078] The fitting calculation unit calculates the fixed-point fitting result according to the remaining mantissa, function value, and difference, and sends the fixed-point fitting result to the postprocessing unit;
[0079] The postprocessing unit normalizes the fixed-point fitting result according to the data type to obtain the corresponding floating-point result, and sends the floating-point result to the result output unit.
[0080] ① The preprocessing unit quantizes the exponent bit and mantissa bit to obtain the integer bit and decimal bit, including:
[0081] According to the standard representation form of floating-point numbers, a hidden integer 1 is added to the fractional part of 16-bit / 32-bit floating-point numbers to convert them into the form of 1.N;
[0082] Subtract an offset of 2 n-1 -1 from the exponent bit to obtain the true exponent. According to the true exponent, determine the number of digits for the decimal point to shift. When the true exponent is positive, move the decimal point of 1.N to the left by the corresponding number of digits; when the true exponent is negative, move the decimal point of 1.N to the right by the corresponding number of digits to make the true exponent become 0, so as to align the fractional part of 16-bit / 32-bit floating-point numbers to the decimal form and obtain the quantized number;
[0083] Determine the integer bit and decimal bit of 16-bit / 32-bit floating-point numbers according to the quantized number;
[0084] Among them, N is the mantissa bit of the current binary number;
[0085] n is the exponent bit width of 16-bit / 32-bit floating-point numbers. For 16-bit floating-point numbers of data type bfp16 and 32-bit floating-point numbers of data type fp32, the exponent bits both occupy 8 bits and the offsets are both 127; for 16-bit floating-point numbers of data type fp16, the exponent bits occupy 5 bits and its offset is 15.
[0086] ② The lookup table consists of 256 SRAMs with a width of 8 bits and a depth of 256. As Figure 4 shown, for 8-bit fixed-point numbers, the corresponding fixed-point results are directly stored in 256 lookup tables;
[0087] As Figure 5 shown, for 16-bit / 32-bit floating-point numbers, 4 SRAMs of 8 bits are concatenated to form 64 lookup tables capable of storing 32-bit data. The function values and differences are stored in the lookup tables. The first 16 bits of the 32-bit data are function values and the last 16 bits are differences;
[0088] Among them, the data stored in the lookup tables are all stored in 16-bit fixed-point format. Compared with the floating-point format storage, the fixed-point format storage does not require the sign bit and exponent bit of the floating-point number, and all bits represent the mantissa bits of the floating-point number, which improves the bit utilization rate of 16-bit data and greatly improves the representation accuracy.
[0089] In the technical solution of this application, for 8-bit fixed-point numbers, the processing unit PE can perform queries on 256 8-bit fixed-point numbers in one clock cycle, and the table values queried are the corresponding fixed-point results;
[0090] For 32-bit floating-point numbers, the processing unit PE can perform calculations on 64 32-bit floating-point numbers in one clock cycle;
[0091] For 16-bit floating-point numbers, the processing unit PE can perform calculations on 64 16-bit floating-point numbers in one clock cycle. For 2048-bit input data, when divided into 16-bit data types, it can be divided into 128 16-bit floating-point numbers. If the processing unit PE needs to perform calculations on 128 16-bit floating-point numbers simultaneously in one clock cycle, 128 lookup tables are required, that is, the hardware resources of the lookup tables need to be increased by 1 time. Therefore, when performing calculations on 16-bit floating-point numbers, considering the factors of consumption and area, without increasing hardware resources, the input data of 16-bit data type is input in two clock cycles to complete the calculation of 128 16-bit floating-point numbers.
[0092] ③ The fitting calculation unit calculates the fixed-point fitting result according to the remaining mantissa, function value and difference, including:
[0093] According to the remaining mantissa, function value, and difference, use a fixed-point multiplier to calculate the fixed-point fitting result by means of linear fitting. That is, when the value x to be calculated falls within [x k , x k+1 , assume that the distance between x k and x k+1 is 1, the distance of x from x k is t, and the distance of x from x k+1 is 1 - t. Then the function value y corresponding to x is:
[0094] y = t(y k+1 - y k ) + y k ;
[0095] Where y k , y k+1 are the function values corresponding to x k , x k+1 respectively, y k+1 - y k is the difference between the function values corresponding to x k , x k+1 , and t is the remaining mantissa.
[0096] For different types of transcendental functions, the bit width of the remaining mantissa t retained after extracting the address index of the lookup table is different. To facilitate the reuse of the fixed-point multiplier, it is necessary to unify the bit width of the remaining mantissa t. Considering that the maximum bit width of the remaining mantissa t retained after extracting the address index of the 16bit / 32bit floating-point number lookup table is 19 bits, and at the same time, to ensure the fitting calculation accuracy and avoid the accumulation of precision loss of intermediate results caused by rounding processing in subsequent calculations, the bit width of all remaining mantissas t is uniformly extended to 19 bits;
[0097] The data stored in the lookup table is all stored in 16bit fixed-point format. Therefore, use a 19 * 16bit fixed-point multiplier to calculate t(y k+1 - y k ), and process the calculation result into a 32bit fixed-point number. Then, use the method of padding with 0 to extend the function value y k obtained through the lookup table operation into a 32bit fixed-point number. Finally, add the two 32bit fixed-point numbers to obtain a 32bit fixed-point fitting result.
[0098] ④ The post-processing unit normalizes the fixed-point fitting result according to the data type to obtain the corresponding floating-point result, including:
[0099] Determine the position of the highest bit 1 in the mantissa bit of the fixed-point fitting result, and use this highest bit 1 as the hidden bit. Truncate the number after the hidden bit as the mantissa bit of the floating-point result;
[0100] Determine the true exponent according to the position of the hidden bit, and add the offset corresponding to the data type of the floating-point number to obtain the exponent bit of the floating-point result;
[0101] Process the sign bit sent by the preprocessing unit according to the type of transcendental function to obtain the corresponding floating-point result;
[0102] Among them, for the 16-bit floating-point number with the data type of bfp16, intercept the high 7 bits after the hidden bit as the mantissa bit of the floating-point result; for the 16-bit floating-point number with the data type of fp16, intercept the high 10 bits after the hidden bit as the mantissa bit of the floating-point result; for the 32-bit floating-point number with the data type of fp32, intercept the high 23 bits after the hidden bit as the mantissa bit of the floating-point result;
[0103] For 16-bit / 32-bit floating-point numbers, if the number of bits after the hidden bit is not enough, directly fill with 0. If there are remaining bits after intercepting the mantissa bit, perform rounding processing in the way of rounding to even: when the highest bit of the remaining bits is 1, if the remaining bits are not all 0, carry; if the remaining bits are all 0, judge according to the lowest bit of the intercepted mantissa bit. If the lowest bit is not 0, carry; if the lowest bit is 0, discard; when the highest bit of the remaining bits is 0, directly discard.
[0104] In the technical solution of this application, for some types of transcendental functions, during the preprocessing in the preprocessing unit or the postprocessing in the postprocessing unit, use a 32-bit floating-point multiplier and a 32-bit floating-point adder for multiply-add calculation (for example, when calculating the logarithmic function lnx, during the postprocessing in the processing unit, it is necessary to add Eln2 to the floating-point result, so a floating-point multiplier and a floating-point adder are required; when calculating the natural exponential function e x x, during the preprocessing in the preprocessing unit, it is necessary to multiply the input data by 1 / ln2, so a floating-point multiplier is required);
[0105] Among them, for 16-bit floating-point numbers, after completing the multiply-add calculation during the postprocessing in the postprocessing unit, it is necessary to convert the 32-bit multiply-add calculation result into 16 bits to obtain the corresponding floating-point result.
[0106] In the technical solution of this application, on the basis of the above-disclosed resource-reuse type transcendental function calculation device, a resource-reuse type transcendental function calculation method is also disclosed, which realizes the pipelined calculation of data in a pipelined manner, as Figure 6 shown, and specifically includes the following processes:
[0107] DC0 level: Preprocess the input data according to the data type to obtain the sign bit, the address index of the lookup table, and the remaining mantissa;
[0108] DC1 level: Perform a table lookup operation according to the address index of the lookup table;
[0109] EX0 level: For 8-bit fixed-point numbers, obtain the corresponding fixed-point result through a table lookup operation for output; for 16-bit / 32-bit floating-point numbers, obtain the function value and the difference through a table lookup operation;
[0110] EX1 level: According to the remaining mantissa and the difference, use a fixed-point multiplier to calculate t(y k+1 -y k );
[0111] EX2 level: Combine the function value to calculate the fixed-point fitting result y = t(y k+1 -y k ) + y k ;
[0112] EX3 level: Normalize the fixed-point fitting result according to the data type to obtain the corresponding floating-point result;
[0113] WB level: Concatenate the calculation results of multiple processing units PE in sequence to obtain the output data;
[0114] Among them, when all processing units PE perform transcendental function calculations, the lookup table, fixed-point multiplier, floating-point multiplier, and floating-point adder are reused;
[0115] DC0 to DC1 levels represent the processing process of the input data, EX0 to EX3 levels represent the data calculation execution process, and the calculation results are obtained through corresponding calculations. The WB level represents the collation and output of the calculation results of multiple processing units PE;
[0116] The above process adopts a way of unequal-length pipeline division: for the same type of transcendental function, there are differences between the pipelines corresponding to 16-bit floating-point numbers and 32-bit floating-point numbers. Since 16-bit floating-point numbers are input in two clock cycles, there is one more level between DC0 and DC1 levels compared to 32-bit floating-point numbers for processing the input 16-bit floating-point numbers; for some types of transcendental functions that require multiplication and addition calculations during the preprocessing in the preprocessing unit or the postprocessing in the postprocessing unit, there are corresponding extensions between DC0 to DC1 levels and EX0 to EX3 levels; for 8-bit fixed-point numbers, after obtaining the corresponding fixed-point result through a table lookup operation, they directly enter the WB level.
[0117] To better illustrate the technical solution of this application, the following will combine 10 types of transcendental functions supported by this application to elaborate in detail on the preprocessing process and the postprocessing process of these 10 types of transcendental functions:
[0118] 1) Sigmoid function
[0119] When the input is a 16-bit / 32-bit floating-point number, the sign bit, exponent bit, and mantissa bit are extracted through preprocessing. The mantissa is quantized so that the decimal point is aligned to the decimal system, resulting in the form of K.M, where K is the integer part and M is the fractional part (the quantization principle of other functions is the same). Due to the characteristics of the Sigmoid function, when the input is greater than 16, the value is approximately 1. Therefore, only the data between [0, 16) needs to be calculated. The index of the lookup table is obtained based on the integer part and the high 4 bits of the fractional part, and an 8-bit index value composed of the low 4 bits of the integer and the high 4 bits of the fractional part is used for the lookup operation. The remaining bits of the fractional part are used as the t value. Then, the fitting calculation is completed according to the above fitting process to obtain the fitting result. Finally, the final result is obtained through post-processing based on the fitting value. The post-processing process needs to consider the processing of the sign bit. For positive numbers, the fitting result is directly processed, and for negative numbers, 1 - the fitting result is processed.
[0120] 2) swish function
[0121] The swish function multiplies the input data of the sigmoid function by the finally calculated function result, and the product is the output of the swish function. Therefore, the core processing process of the swish function preprocessing is basically the same as that of the sigmoid function. The difference lies in an additional floating-point multiplication calculation and a renormalization process in the post-processing stage.
[0122] 3) erf function
[0123] When the input is a 16-bit / 32-bit floating-point number, the sign bit, exponent bit, and mantissa bit are extracted through preprocessing. The mantissa is quantized so that the decimal point is aligned to the decimal system. Due to the characteristics of the erf function, when the input is greater than 4, the value is approximately 1. Therefore, only the data between [0, 4) needs to be calculated. Two bits of the integer part and six bits of the fractional part of the quantization result are used as the lookup index, and the remaining mantissa bits are regarded as t. Then, the fitting calculation and post-processing are completed to obtain the result. Since erf is an odd function, only the positive value needs to be calculated, and the sign bit is added to the result.
[0124] 4) Reciprocal function
[0125] When the input is a 16-bit / 32-bit floating-point number, the sign bit, exponent bit, and mantissa bit are extracted through preprocessing. The first 8 bits of the mantissa are used as the index, and the remaining mantissa is used as t. The fitting calculation is completed to obtain the fitting value. In the post-processing stage, the exponent obtained from the preprocessing is negated and added to the exponent of the fitting value to obtain a new exponent, and the sign is processed so that the sign of the result is the same as the sign of the input data.
[0126] 5) Natural exponential function
[0127] For the convenience of calculation, the exponential function e x can be processed as In this way, calculating the exponential function e x is actually converted into calculating 2 x , so in the preprocessing stage for the input floating-point number, first multiply the input by 1 / ln2 to get the intermediate value, then extract the sign bit, exponent bit and mantissa bit, quantize it into the form of K.M, take the first 7 bits of M and 1 bit of the sign bit to form the index of the lookup table, and use the remaining mantissa as the t value to complete the fitting calculation. In the postprocessing stage, add the exponent of the fitting result to the exponent obtained in the preprocessing. If the sign bit is negative, subtract the exponents;
[0128] 6) Square root function
[0129] In the preprocessing stage, extract the sign bit S, exponent bit E and mantissa bit M of the input floating-point number. Take the first 8 bits of the mantissa as the index value, and the remaining mantissa bits as t, and perform fitting calculation to get the result. In the postprocessing, judge according to the parity of the exponent E. When E is even, add E / 2 to the exponent of the fitting result. When E is odd, multiply the exponent of the fitting result and then add (E - 1) / 2;
[0130] 7) sin function
[0131] For the trigonometric function sin, multiply the input by to make its period become 4, and then perform preprocessing. Extract the sign bit, exponent bit and mantissa bit, and quantize to get K.M. Judge which quadrant the input is in according to the last two digits of the integer part K. When the last two digits are 0, the input is in the first quadrant. When it is 1, the input is in the second quadrant. When it is 2, it is in the third quadrant. When it is 3, it is in the fourth quadrant. Take the first 8 bits of the mantissa as the lookup table index, and use the remaining mantissa as t, and perform fitting calculation. If the input is in the 1 / 3 quadrant, the index value remains unchanged and the t value remains unchanged. If the input is in the 2 / 4 quadrant, perform an inversion operation on the index value and the remaining mantissa t; in the postprocessing stage, if the input is in the first and second quadrants, the result is positive. If the input is in the third and fourth quadrants, the result is negative. Since the sin function is a centrally symmetric function, for the part less than 0, directly take the absolute value and then take the opposite of the function value;
[0132] 8) cos function
[0133] For the trigonometric function cos, multiply the input by Make its period become 4, and then perform preprocessing. Take the sign bit, exponent bit, and mantissa bit, and quantize to obtain K.M. Determine which quadrant the input is in according to the last two digits of the integer part K. When the last two digits are 0, the input is in the first quadrant; when they are 1, the input is in the second quadrant; when they are 2, it is in the third quadrant; when they are 3, it is in the fourth quadrant. Take the first 8 bits of the mantissa as the lookup table index, and regard the remaining mantissa as t, and perform fitting calculation. If the input is in the 1st / 3rd quadrant, the index value remains unchanged and the t value remains unchanged. If the input is in the 2nd / 4th quadrant, the index value and the remaining mantissa t are inverted. In the post-processing stage, if the input is in the first and fourth quadrants, the result is positive. If the input is in the second and third quadrants, the result is negative. Since the sin function is a centrosymmetric function, for the part less than 0, directly take the absolute value and then take the opposite of the function value;
[0134] 9) Logarithmic function
[0135] For the logarithmic function lnx, when the input floating-point number only considers positive numbers, perform preprocessing on the input to obtain the exponent bit E and the mantissa bit M. Use the first 8 bits of the mantissa as the index, and the remaining mantissa as t, and perform fitting calculation to obtain the fitting result. In the post-processing stage, for lnx, it can be transformed into lnx = ln(1.M * 2 E ) = ln(1.M) + Eln2. Therefore, adding Eln2 to the fitting result can obtain the final result;
[0136] 10) Tanh function
[0137] When the input is a 16-bit / 32-bit floating-point number, take out the sign bit, exponent bit, and mantissa bit through preprocessing, perform quantization processing on the mantissa to align the decimal point to the decimal system. According to the characteristics of the tanh function, when the input is greater than 8, the value is approximately 1. Therefore, only need to calculate the data between [0, 8). After quantization, take 3 integer bits and 5 decimal bits as the lookup table index, and regard the remaining mantissa bits as t. Then complete the fitting calculation and post-processing to obtain the result. Since tanh is an odd function, only need to calculate the positive value, and add the sign bit to the result.
[0138] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A resource reuse type transcendental function calculation device, characterized in that: It includes a selector, multiple processing units PE, and a result output unit; The selector receives input data, data types, and control signals, and sends the input data and data types to the corresponding processing units PE according to the control signals; The processing unit PE can process all types of transcendental functions and perform transcendental function calculations on the input data according to the data types; The result output unit splices the calculation results of multiple processing units PE in sequence to obtain output data; Among them, when all processing units PE perform transcendental function calculations, they reuse the look-up table, fixed-point multiplier, floating-point multiplier, and floating-point adder.
2. The resource reuse type transcendental function calculation device according to claim 1, wherein: The processing unit PE includes a preprocessing unit, a look-up table unit, a fitting calculation unit, and a post-processing unit; The preprocessing unit does not process 8-bit fixed-point numbers and directly sends the data itself as the address index of the look-up table to the look-up table unit; for 16-bit / 32-bit floating-point numbers, it processes them to obtain the sign bit, exponent bit, and mantissa bit, directly sends the sign bit to the post-processing unit, quantizes the exponent bit and mantissa bit to obtain the integer bit and decimal bit, extracts the address index of the look-up table from the integer bit and decimal bit according to the corresponding index address extraction method of the transcendental function type, and sends the address index of the look-up table to the look-up table unit, and sends the remaining bits of the decimal bit as the remaining mantissa to the fitting calculation unit; The look-up table unit performs look-up table operations according to the address index of the look-up table. For 8-bit fixed-point numbers, the corresponding fixed-point results are obtained through look-up table operations and output; for 16-bit / 32-bit floating-point numbers, the function value and difference are obtained through look-up table operations and sent to the fitting calculation unit; The fitting calculation unit calculates the fixed-point fitting result according to the remaining mantissa, function value, and difference, and sends the fixed-point fitting result to the post-processing unit; The post-processing unit normalizes the fixed-point fitting result according to the data type to obtain the corresponding floating-point result, and sends the floating-point result to the result output unit.
3. The resource reuse type transcendental function calculation device according to claim 2, wherein: The preprocessing unit quantizes the exponent bit and mantissa bit to obtain the integer bit and decimal bit, including: According to the standard representation form of floating-point numbers, add a hidden integer 1 to the fractional part of 16-bit / 32-bit floating-point numbers and convert them to the form of 1.N; Subtract an offset of 2 from the exponent bits n-1 -1 to obtain the true exponent. Determine the number of decimal point shifts based on the true exponent. When the true exponent is positive, shift the decimal point of 1.N to the left by the corresponding number of digits; when the true exponent is negative, shift the decimal point of 1.N to the right by the corresponding number of digits to make the true exponent 0, thereby aligning the fractional part of the 16-bit / 32-bit floating-point number to the decimal form and obtaining the quantized number; Determine the integer bit and decimal bit of 16-bit / 32-bit floating-point numbers according to the quantized number; Among them, N is the mantissa bit of the current binary number; n is the exponent bit width of 16-bit / 32-bit floating-point numbers. For 16-bit floating-point numbers of data type bfp16 and 32-bit floating-point numbers of data type fp32, the exponent bits both occupy 8 bits and the offsets are both 127; for 16-bit floating-point numbers of data type fp16, the exponent bits occupy 5 bits and its offset is 15.
4. The resource reuse type transcendental function calculation device according to claim 3, characterized in that: The look-up table is composed of 256 SRAMs with a width of 8 bits and a depth of 256. For 8-bit fixed-point numbers, the corresponding fixed-point results are directly stored in 256 look-up tables; For 16-bit / 32-bit floating-point numbers, four 8-bit SRAMs are spliced to form 64 lookup tables capable of storing 32-bit data. The function values and differences are stored in the lookup tables. The first 16 bits of the 32-bit data are function values, and the last 16 bits are differences. Among them, the data stored in the lookup tables are all stored in 16-bit fixed-point format. Compared with the floating-point format storage, the fixed-point format storage does not require the sign bit and exponent bit of the floating-point number, and all bits represent the mantissa bits of the floating-point number, improving the bit utilization rate of 16-bit data and greatly improving the representation accuracy.
5. The resource reuse type transcendental function calculation device according to claim 4, characterized in that: For 8-bit fixed-point numbers, the processing unit PE can query 256 8-bit fixed-point numbers in one clock cycle, and the table values queried are the corresponding fixed-point results. For 32-bit floating-point numbers, the processing unit PE can perform 64 32-bit floating-point number calculations in one clock cycle. For 16-bit floating-point numbers, the processing unit PE can perform 64 16-bit floating-point number calculations in one clock cycle. For 2048-bit input data, when divided by 16-bit data type, it can be divided into 128 16-bit floating-point numbers. If the processing unit PE needs to perform 128 16-bit floating-point number calculations simultaneously in one clock cycle, 128 lookup tables are required, that is, the hardware resources of the lookup tables need to be doubled. Therefore, when performing 16-bit floating-point number calculations, considering the factors of consumption and area, without increasing the hardware resources, the input data of 16-bit data type is input in two clock cycles to complete the calculation of 128 16-bit floating-point numbers.
6. The resource reuse type transcendental function calculation device according to claim 4, wherein: The fitting calculation unit calculates the fixed-point fitting result according to the remaining mantissa, function value and difference, including: According to the remaining mantissa, function value, and difference, use a fixed-point multiplier to calculate the fixed-point fitting result by linear fitting. That is, when the value x to be calculated falls within [x k , x k+1 , assume that the distance between x k and x k+1 is 1, the distance from x to x k is t, and the distance from x to x k+1 is 1 - t. Then the function value y corresponding to x is: y = t(y k+1 -y k ) + y k ; Among them, y k and y k+1 are the function values corresponding to x k and x k+1 respectively. y k+1 −y k is the difference between the function values corresponding to x k and x k+1 , and t is the remaining mantissa.
7. The resource reuse type transcendental function calculation device according to claim 6, characterized in that: For transcendental functions of different types, the bit width of the remaining mantissa t retained after extracting the address index of the lookup table is different. In order to facilitate the reuse of the fixed-point multiplier, it is necessary to unify the bit width of the remaining mantissa t. Considering that the maximum bit width of the remaining mantissa t retained after extracting the address index of the 16-bit / 32-bit floating-point number lookup table is 19 bits, and at the same time to ensure the fitting calculation accuracy and avoid the accumulation of precision loss of the intermediate result in subsequent calculations due to rounding processing, the bit width of all remaining mantissas t is uniformly extended to 19 bits. The data stored in the lookup table is all stored in 16-bit fixed-point format. Therefore, a 19 * 16-bit fixed-point multiplier is used to calculate t(y k+1 -y k ), and the calculation result is processed into a 32-bit fixed-point number. Then, the function value y k obtained through the lookup table operation is extended to a 32-bit fixed-point number by padding with zeros. Finally, the two 32-bit fixed-point numbers are added together to obtain a 32-bit fixed-point fitting result.
8. The resource reuse type transcendental function calculation device according to claim 6, wherein: The post-processing unit performs normalization processing on the fixed-point fitting result according to the data type to obtain the corresponding floating-point result, including: Determine the position of the highest bit 1 of the mantissa bits in the fixed-point fitting result, and use this highest bit 1 as the hidden bit, and intercept the numbers after the hidden bit as the mantissa bits of the floating-point result. Determine the true exponent according to the position of the hidden bit, and add it to the offset corresponding to the data type of the floating-point number to obtain the exponent bit of the floating-point result. Process the sign bit sent by the preprocessing unit according to the type of the transcendental function to obtain the corresponding floating-point result. Among them, for 16-bit floating-point numbers of data type bfp16, the high 7 bits after the hidden bit are intercepted as the mantissa bits of the floating-point result; for 16-bit floating-point numbers of data type fp16, the high 10 bits after the hidden bit are intercepted as the mantissa bits of the floating-point result; for 32-bit floating-point numbers of data type fp32, the high 23 bits after the hidden bit are intercepted as the mantissa bits of the floating-point result. For 16-bit / 32-bit floating-point numbers, if the number of bits after the hidden bit is insufficient, 0 is directly filled. If there are remaining bits after intercepting the mantissa bits, rounding is performed using the round-to-even method: when the highest bit of the remaining bits is 1, if the remaining bits are not all 0, carry is performed; if the remaining bits are all 0, it is judged according to the lowest bit of the intercepted mantissa bits. If the lowest bit is not 0, carry is performed; if the lowest bit is 0, it is discarded; when the highest bit of the remaining bits is 0, it is directly discarded.
9. The resource reuse type transcendental function calculation device according to claim 8, wherein: For some types of transcendental functions, during the preprocessing in the preprocessing unit or the postprocessing in the postprocessing unit, multiplication and addition calculations are performed using a 32-bit floating-point multiplier and a 32-bit floating-point adder. Among them, for 16-bit floating-point numbers, after the multiplication and addition calculations are completed during the postprocessing in the postprocessing unit, the 32-bit multiplication and addition calculation result needs to be converted to 16 bits to obtain the corresponding floating-point result.
10. A resource reuse type transcendental function calculation method, applied to the resource reuse type transcendental function calculation device described in claim 1, characterized in that: The pipelining method is used to implement the pipelined calculation of data, which specifically includes the following processes: DC0 stage: The input data is preprocessed according to the data type to obtain the sign bit, the address index of the lookup table, and the remaining mantissa. DC1 stage: The lookup table operation is performed according to the address index of the lookup table. EX0 stage: For 8-bit fixed-point numbers, the corresponding fixed-point result is obtained through the lookup table operation and output; for 16-bit / 32-bit floating-point numbers, the function value and the difference are obtained through the lookup table operation. Level EX1: Based on the remaining mantissa and the difference, calculate t(y k+1 -y k ) using a fixed-point multiplier; EX2 level: Combine function values to calculate the fixed-point fitting result y = t(y k+1 - y k ) + y k ; EX3 stage: The fixed-point fitting result is normalized according to the data type to obtain the corresponding floating-point result. WB stage: The calculation results of multiple processing elements PE are concatenated in sequence to obtain the output data. Among them, when all processing elements PE perform transcendental function calculations, the lookup table, fixed-point multiplier, floating-point multiplier, and floating-point adder are reused. The DC0 to DC1 stages represent the processing process of the input data, the EX0 to EX3 stages represent the data calculation execution process, the calculation results are obtained through the corresponding calculations, and the WB stage represents the collation and output of the calculation results of multiple processing elements PE. The above process adopts a way of unequal-length pipeline partitioning: for transcendental functions of the same type, there are differences between the pipelines corresponding to 16-bit floating-point numbers and 32-bit floating-point numbers. Since the 16-bit floating-point numbers are input in two clock cycles, there is one more stage between DC0 and DC1 levels compared to 32-bit floating-point numbers to process the input 16-bit floating-point numbers; for some types of transcendental functions that need to perform multiply-accumulate calculations during the preprocessing in the preprocessing unit or the postprocessing in the postprocessing unit, there are corresponding expansions between DC0 and DC1 levels and between EX0 and EX3 levels; for 8-bit fixed-point numbers, they directly enter the WB level after obtaining the corresponding fixed-point results through table lookup operations.
Citation Information
Cited By
Implementation circuit, method and application of S-type activation function based on ASIC (Application Specific Integrated Circuit)
CN121349255A
An implementation circuit, method and application of an S-shaped activation function based on ASIC
CN121349255B