Exponential type low bit width calculation acceleration method and system based on lookup table optimization
By converting the multiplication and division of exponential floating-point data into addition and subtraction operations of exponential and mantissa, and using lookup table optimization and exponential offset adjustment, the limitations of FP8 in calculation error and efficiency are solved, and efficient low-bit width calculation is achieved.
Patent Information
- Application Number
- CN202510451439.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-08
AI Technical Summary
The 8-bit floating point format (FP8) of the existing IEEE-754 standard has many limitations in calculation error, efficiency and accuracy, especially in high-precision sensitive scenarios, facing systematic accuracy defects and efficiency bottlenecks, especially in addition and subtraction operations, which are prone to error accumulation and error amplification.
The exponential low-bit width calculation method based on lookup table optimization is adopted to convert the multiplication and division of exponential floating-point data into addition and subtraction operations of exponential and mantissa. The addition and subtraction process is optimized through the lookup table, combined with the exponential offset adjustment mechanism, to ensure that the calculation results are within the expressible range.
It significantly improves computing efficiency, reduces matrix calculation error and calculation delay, reduces hardware resource usage, and maintains constant relative errors, and is suitable for communication systems and artificial intelligence fields.
Smart Images

Figure CN120276705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer data processing, and specifically to an exponential low-bitwidth calculation acceleration method and system optimized based on a lookup table. Background Art
[0002] Currently, the floating-point storage scheme based on the IEEE-754 standard (IEEE Standard for Floating-point Arithmetic, IEEE-754) always has some limitations. Taking the 8-bit floating-point format (FP8) as an example, its quantization characteristics are restricted by the trade-off between the dynamic range of the exponent field and the mantissa resolution. The quantization step of FP8 shows non-linear growth, resulting in a stepwise jump of the storage error as the value increases, and the compression of the exponent bitwidth limits the exponent coverage range. When the value exceeds the maximum representable range, the error is amplified sharply and even causes the numerical representation to break down; in terms of the analysis of calculation error characteristics and operation efficiency, the accumulation of basic operation errors is manifested as in addition and subtraction operations, the rounding error accumulation effect of FP8 is significant due to insufficient mantissa bits, especially in the accumulation operation, it is easy to cause the "error stagnation phenomenon"; in multiplication, division, and complex-domain matrix multiplication, the error propagation shows a non-linear amplification trend, for example, in complex multiplication, the effective bits are lost due to the truncation of intermediate results; in multiplication, division, and square root operations, the operator complexity is relatively high, further introducing calculation latency; in terms of the limitations of application scenarios, although the low-bitwidth characteristic of FP8 can reduce the bandwidth requirement in scenarios pursuing calculation throughput, its significant calculation error increases the risk of model convergence failure. Generally speaking, due to the quantization mechanism and limited bitwidth constraints, FP8 faces systematic accuracy defects and efficiency bottlenecks in most high-precision sensitive scenarios. Therefore, the present invention proposes an exponential low-bitwidth calculation acceleration method and system optimized based on a lookup table. Summary of the Invention
[0003] The purpose of the present invention is to provide an exponential low-bitwidth calculation acceleration method and system optimized based on a lookup table to solve the problems raised in the above background art.
[0004] To achieve the above purpose, the present invention provides the following technical solution: An exponential low-bitwidth calculation acceleration method optimized based on a lookup table, including the following steps:
[0005] Receive exponential floating-point data, and the basic format of the exponential floating-point data includes a sign bit, an exponent bit, and a mantissa bit;
[0006] Perform different arithmetic processing operations on the exponential floating-point data according to different arithmetic relationships. The arithmetic relationships specifically include negation and reciprocal operations, multiplication and division operations, square root operations, and addition and subtraction operations, and in the multiplication and division operations, directly convert the multiplication and division operations of the exponential floating-point data into addition and subtraction operations of the exponent and the mantissa;
[0007] Design a lookup table and optimize it to convert the addition and subtraction operation process into a data lookup process;
[0008] Set the initial range value of the exponent offset, and determine whether the calculation results obtained from each operation process exceed the upper and lower limits of the initial range value. If they exceed, automatically adjust the exponent offset.
[0009] Furthermore, the basic format of the exponent floating-point data is represented as follows:
[0010]
[0011] In the formula, value_dec represents the decimal value, s represents the sign bit, base represents the base number, and the specific value is 2, e represents the exponent value, bias represents the exponent offset, m represents the mantissa value, and m_bit represents the bit width.
[0012] Furthermore, when performing the negation and reciprocal operations, the specific operation processes are as follows:
[0013] Negation: Invert the sign bit;
[0014] Reciprocal: Keep the sign bit unchanged, invert the mantissa bit and add 1, and perform different addition and subtraction operations on the exponent bit according to different bit widths.
[0015] Furthermore, when performing multiplication and division operations, the specific operation processes are as follows:
[0016] Multiplication: For two exponent floating-point data to be multiplied, first perform an exclusive OR operation on the sign bits, add the exponent bits and mantissa bits of the two exponent floating-point data. If there is a carry in the mantissa bit, add 1 to the exponent bit;
[0017] Division: For two exponent floating-point data to be divided, first perform the reciprocal operation, that is, invert the mantissa bit and add 1, and perform addition and subtraction operations on the exponent bit according to different bit widths, and then calculate according to the steps of multiplication.
[0018] Furthermore, when performing the square root operation, the specific operation process is as follows:
[0019] Perform the operation of dividing by 2 on the exponent part n In the process of binary representation, it is equivalent to shifting right by n + 1 bits.
[0020] Furthermore, when performing addition and subtraction operations, the specific operation processes are as follows:
[0021] Addition:
[0022] If the two addends have the same sign, keep the sign bit unchanged. Depending on the difference in the exponent bits, select different lookup tables and determine the row of the lookup table based on the mantissa bits of the two addends. The column is the result;
[0023] If the two addends have different signs, first determine the sign bit by comparing the magnitudes of the two numbers. Depending on the difference in the exponent bits, select different lookup tables and determine the row of the lookup table based on the mantissa bits of the two addends. The column is the result;
[0024] Subtraction:
[0025] Invert the sign bit of the subtrahend and convert it to an addition operation.
[0026] Furthermore, set an initial range value for the exponent offset. Determine whether the calculation results obtained from each operation process exceed the upper and lower limits of the initial range value. If they exceed, automatically adjust the exponent offset. The specific operation mechanism is as follows:
[0027] (71) Initial value setting of the exponent offset: Set an initial exponent offset value at the beginning of the calculation according to the value range and precision requirements;
[0028] (72) Dynamic adjustment mechanism: When the value during the calculation process exceeds the upper and lower limits of the initial range value, the system automatically adjusts the exponent offset value, readjusts the value to the representable range, and recalculates the exponent bits with the adjusted new offset value so that the calculation result of the mantissa bits can be retained;
[0029] (73) Use an additional exponent offset register or memory for each value for subsequent calls.
[0030] According to the second aspect of the present invention, the present invention provides an exponent-based low-bitwidth calculation acceleration system optimized based on a lookup table for implementing the above-mentioned exponent-based low-bitwidth calculation acceleration method optimized based on a lookup table, including:
[0031] A receiving module for receiving exponent floating-point data. The basic format of the exponent floating-point data includes a sign bit, an exponent bit, and a mantissa bit;
[0032] A data processing module for performing different operation processing operations on the exponent floating-point data according to different operation relationships. The operation relationships specifically include inversion and reciprocal operations, multiplication and division operations, square root operations, and addition and subtraction operations. And in the multiplication and division operations, directly convert the multiplication and division operations of the exponent floating-point data into addition and subtraction operations of the exponent and mantissa;
[0033] A lookup table design module for designing and optimizing the lookup table, and converting the addition and subtraction operation process into a data lookup process;
[0034] An automatic adjustment exponent offset module is used to set the initial range value of the exponent offset, determine whether the calculation results obtained from each operation process exceed the upper and lower limits of the initial range value, and automatically adjust the exponent offset if it exceeds.
[0035] According to the third aspect of the present invention, the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The computer program stored in the memory can run on the processor. When the processor loads and executes the computer program, the above-mentioned exponent-based low-width calculation acceleration method optimized based on the lookup table is adopted.
[0036] According to the fourth aspect of the present invention, the present invention provides a storage medium containing computer-executable instructions. The computer-executable instructions are used to execute the above-mentioned exponent-based low-width calculation acceleration method optimized based on the lookup table when executed by a computer processor.
[0037] The present invention at least has the following beneficial effects:
[0038] 1. The present invention fully considers problems such as storage errors, calculation accuracy, and long time delays in existing FP schemes. In the multiplication, division, and square root operations of the present invention's EFP, by utilizing the self-simplification advantage of the exponential form, complex multiplication and division operations are converted into simple addition and subtraction operations, significantly improving the calculation efficiency. In the addition and subtraction operations, a lookup table is innovatively designed to convert the data calculation process into a data transfer process, and a series of methods for simplifying the lookup table are designed to further improve the operation efficiency of addition and subtraction. Compared with traditional floating-point numbers, EFP significantly reduces the matrix calculation error and reduces the calculation time delay; the algorithm design and performance of EFP can be widely applied in fields such as communication systems, artificial intelligence, and machine learning.
[0039] 2. In the architecture of the present invention, the division operation can share the circuit with the multiplication operation, reducing the occupancy of hardware resources while improving the operation efficiency. At the same time, compared with FP, the relative error of EFP remains constant without obvious fluctuations.
[0040] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. Description of the Drawings
[0041] Figure 1 is a schematic flowchart of the calculation method of the present invention;
[0042] Figure 2 is a schematic diagram of the mantissa distribution of the present invention taking EFP8 as an example;
[0043] Figure 3 is a schematic diagram comparing the average relative error of a single real number addition and subtraction calculation between EFP and FP of the present invention;
[0044] Figure 4 It is a schematic diagram comparing the average relative error of the EFP and FP in the first real number multiplication and division calculations of the present invention;
[0045] Figure 5 It is a schematic diagram of the average relative error of the EFP and FP in the first real number square root calculation of the present invention;
[0046] Figure 6 It is a schematic diagram of the average relative error of the EFP and FP in the vector inner product of the present invention;
[0047] Figure 7 It is a schematic diagram of the average relative error of the EFP and FP in matrix multiplication of the present invention;
[0048] Figure 8 It is a schematic diagram of the average relative error of the EFP and FP in matrix inversion of the present invention;
[0049] Figure 9 It is a schematic diagram comparing the time delays of the EFP and FP in the first real number multiplication, division, and square root calculations of the present invention. Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0051] Embodiment 1:
[0052] In this embodiment, for application scenarios with strict requirements for computing accuracy and energy efficiency ratio, such as edge computing devices and embedded AI acceleration chips, by reconstructing the exponential representation paradigm of floating-point numbers, the common problems of computing accuracy loss and high complexity of arithmetic units existing in traditional floating-point calculations under low-bit width constraints are solved.
[0053] Please refer to Figure 1 , the technical solution provided by the present invention: an exponential low-bit width calculation acceleration method optimized based on a lookup table, including the following steps:
[0054] S1. Receive exponential floating-point data. The basic format of the exponential floating-point data includes a sign bit, an exponent bit, and a mantissa bit. The specific format is as follows:
[0055]
[0056] Wherein, value_dec represents the decimal value, s represents the sign bit, base represents the base number, and its specific value is 2, e represents the exponent value, bias represents the exponent offset, m represents the mantissa value, and m_bit represents the bit width;
[0057] S2. Perform different arithmetic processing operations on the exponential floating-point data according to different arithmetic relationships. The arithmetic relationships specifically include negation and reciprocal operations, multiplication and division operations, square root operations, and addition and subtraction operations. And in the multiplication and division operations, directly convert the multiplication and division operations of the exponential floating-point data into the addition and subtraction operations of the exponent and mantissa;
[0058] S21. Implementation scheme of negation and reciprocal:
[0059] Figure 2 The mantissa distribution of EFP8 is shown in a circle. Considering the sign bit and the mantissa bit, the EFP8 with the same exponent bit in 8 bits is surrounded by a circle. The first bit (red) represents the sign bit, and the 2nd - 5th bits (black) represent the mantissa bit. It can be found that for the negation operation, only need to horizontally exchange the value around the vertical axis and negate the first bit;
[0060] For the reciprocal operation, only need to vertically exchange the value around the horizontal axis, that is, the sign bit remains unchanged, the mantissa bit is negated and incremented by 1, and the exponent bit performs different addition and subtraction operations according to different bit widths. The specific addition and subtraction values depend on the size of the bit width;
[0061] Taking EFP8 E2M5:{0.00.11111} as an example, its representation in decimal is Its reciprocal is Converted to the EFP8 form: {0.01.00001}. In binary, only need to negate the mantissa bit 11111 of the original value to 00000, and then increment by 1 to get the mantissa bit 00001 of the result;
[0062] S22. Implementation scheme of addition and subtraction:
[0063] Different from the traditional floating point (FP), which directly applies the binary addition and subtraction algorithm because its digit part follows the uniform coding of the fixed point number, for the exponential floating-point data (Exponential Floating Point, EFP), since its mantissa part shows an exponential distribution, the direct addition is hindered. At the same time, to simplify the process and improve the calculation efficiency, a lookup table method is adopted and improved. The advantage of this method is that when performing addition and subtraction operations, only one lookup table operation is required to directly obtain the result without additional calculation steps.
[0064] For addition and subtraction, only addition needs to be considered because subtraction can be converted to addition by simply inverting the sign bit of the subtrahend. At the same time, only the cases of the addends having the same sign and different signs need to be distinguished in the addition case.
[0065] Addends with the same sign: 1. The sign bit remains unchanged. 2. Determine the lookup table according to the difference in the exponent bits. 3. Determine the calculation result according to the lookup table;
[0066] Addends with different signs: 1. Determine the sign bit by judging the magnitudes of the two numbers. 2. Determine the lookup table according to the difference in the exponent bits. 3. Determine the calculation result according to the lookup table;
[0067] Taking EFP8 E4M3 as an example, when the two addends have the same sign and the same exponent bits, the sign bit remains unchanged, and the designed lookup table is shown in Table 1 below:
[0068] Table 1 Lookup table when the addends have the same sign and the same exponent bits in EFP8 E4M3
[0069]
[0070] Table 2 Lookup table when the addends have the same sign and the exponent bits differ by 1 in EFP8 E4M3
[0071]
[0072] If the mantissa bits of the two addends are e1 and e2 (e1 ≤ e2) respectively, then Figure 2 the value in the e1-th row and e2-th column of
[0073] is the result. The first digit of the corresponding data represents the increment of the exponent bit, and the last 3 digits are the mantissa bit result;
[0074] When the difference in the exponent bits is 1, take the mantissa bit with the larger exponent as e1 and the other mantissa bit as e2. Then the value in the e1-th row and e2-th row of Table 2 is the result. The first digit represents whether the larger exponent bit is incremented by 1, and the last 3 digits are the mantissa bit result. When the difference in the exponent bits is 2, the operation is the same as when the difference is 1; m_bit *2 m_bit ;
[0075] For the case of different signs, the lookup table also needs to store the reduced exponent bits and the calculated results of the mantissa bits. It should be noted that for the case of different signs, if defined in the same way as for the same sign, the lookup table does not satisfy the variation rule of the lookup table required for compression. For the variation of the mantissa bits, it satisfies adding 1 successively, but if the exponent bits are defined as the reduced exponent bits, it does not satisfy the rule of adding 1 successively. Therefore, a constant equal to 15 (the maximum offset value of the exponent bits for more than two addends) needs to be defined. Subtracting the reduced exponent bits from the constant can satisfy the rule of adding 1 successively. Table 3 shows the lookup table when the exponent bits of the addends with different signs EFP8 E4M3 are the same. The first row of the table is the minuend, and the second row is the subtrahend. When the exponent bits are the same, only the upper triangular matrix needs to be stored. When doing subtraction and finding the corresponding data, the value of the first four bits needs to be subtracted from the constant 15, and the obtained value is the reduced value of the exponent bits.
[0076] Table 3 Lookup Table when the Exponent Bits of the Addends with Different Signs EFP8 E4M3 are the Same
[0077]
[0078] Consider how many kinds of exponent bit differences need to be designed according to the mantissa bit width and the base. Once the mantissa bit width and the base are determined, the number of lookup tables is also determined. Table 4 shows the number of lookup tables required for addition and subtraction when the base is 2 and 10.
[0079] Table 4 Number of Lookup Tables Required for Addition and Subtraction with Different Mantissa Bit Widths when the Base is 2 and 10
[0080]
[0081] Taking the addition of numbers with the same sign as an example (subtraction can be converted into addition by simply taking the inverse of the subtrahend), the specific algorithm is as follows:
[0082]
[0083] S23. Implementation Scheme for Multiplication and Division Operations
[0084] Since the mantissa of EFP is in exponential form, the multiplication and division of EFP can be converted into addition and subtraction of exponents and mantissas. The specific algorithm for multiplication is as follows. For division, only change the addition in the multiplication process to subtraction:
[0085]
[0086] Specifically, for multiplication: for two exponentially floating-point data to be multiplied, first perform an exclusive OR operation on the sign bits, add the exponent bits and the mantissa bits of the two exponentially floating-point data. If there is a carry in the mantissa bits, add 1 to the exponent bits;
[0087] Division: For two exponent floating-point data in division, first perform the reciprocal operation, that is, invert the mantissa bits and then add 1, and perform addition and subtraction operations on the exponent bits according to different bit widths, and convert it into the calculation steps of multiplication;
[0088] S24. Square root operation implementation plan:
[0089] Based on the characteristics of EFP (the mantissa part adopts the exponential form), so when performing the square root operation, it can be simplified to divide by 2 on the exponent part n operation. In the process of binary representation, it is equivalent to shifting n + 1 bits to the right. For example, dividing by 2 is equivalent to shifting 1 bit to the right. Based on this principle, a square root algorithm can be designed, and its complexity will be greatly reduced compared with FP. The algorithm is as follows:
[0090]
[0091] S3. Design a lookup table and optimize it to convert the addition and subtraction operation process into a data lookup process;
[0092] Design of the lookup table:
[0093] Since the mantissa of EFP shows an exponential distribution and cannot be directly added in the addition and subtraction process, so choose to directly store the addition and subtraction results of the mantissa bits in the lookup table, thus converting the calculation process into a data lookup process, and at the same time not introducing additional calculation errors and reducing the complexity of the addition and subtraction calculations.
[0094] Optimization of the lookup table:
[0095] In the EFP system design, as the number of mantissa bits increases (especially when the designed mantissa bits reach 10 bits or higher), the size of the lookup table will show exponential growth, which puts great pressure on the storage space. Therefore, it becomes crucial to simplify and compress the storage of the lookup table. Since the values in the EFP lookup table show a regular pattern of gradually increasing along the diagonal (the values increase by a fixed step of 1 along a specific direction), so the lookup table only needs to store the first row and the first column. When querying the middle elements, they can be calculated from the values in the first row and the first column. Therefore, the size of the lookup table can be reduced from 2 m_bit ×2 m_bit compressed into a matrix of size 2×2 m_bit This kind of compression greatly reduces the storage requirements of the lookup table.
[0096] Specifically, in the design of the EFP system, when the number of mantissa bits is set to 5, the size of the LUT is 32×32, which is still within a manageable range in practical applications. However, as the number of mantissa bits increases, the size of the lookup table will increase exponentially, resulting in excessive storage requirements. Therefore, the compression technology of the lookup table has become a key challenge and a necessary step in the evolution of EFP from low bit-width to high bit-width;
[0097] Through in-depth analysis of the lookup table, it is observed that the values in the EFP lookup table show a regular pattern of gradually increasing along the diagonal. Taking Table 2 as an example, the values increase with a fixed step (such as 1) along a specific direction (as shown by the blue line); based on this regularity, optimization and compression strategies for the lookup table can be implemented to significantly reduce storage requirements and improve computational efficiency. This optimization technology is crucial for the practicality and efficiency of implementing high-bit-width EFP systems. It has been demonstrated that regardless of the difference between the base and exponent bits, the lookup table always increases by 1 along the diagonal. Therefore, only the first row and the first column of the lookup table need to be retained. If the mantissa bits of the two addends are e1 and e2 (e1 ≤ e2), the result corresponding to the e1-th row and e2-th column of the lookup table can be obtained from the values in the first row and the first column and the regularity;
[0098] Based on this rule, the lookup table can be compressed from a m_bit * m_bit matrix of 2 m_bit ×2 to a matrix with a size of only 2×2, such a compression method greatly reduces the storage requirements of the lookup table while ensuring the computational efficiency and accuracy;
[0099] Table 5 details the comparison of the complexities of EFP, FP, and LNS (logarithmic number system) in addition and subtraction operations. Among them, the LNS scheme shows the highest complexity in performing addition and subtraction, involving two table lookup operations and one direct addition and subtraction calculation. In contrast, the EFP scheme only requires one table lookup operation, while the FP scheme involves one addition operation. It should be noted that these complexity evaluations are also affected by parameters such as simulation software and hardware structure.
[0100] Table 5 Comparison of the Complexities of EFP, FP, and LNS in Addition and Subtraction
[0101]
[0102] When deeply analyzing these storage schemes, it is found that FP requires additional operations such as alignment and rounding during the addition operation, which results in the actual bit width used exceeding its representation bit width. However, the lookup table method adopted by EFP ensures that the actual bit width used is equal to the representation bit width. Therefore, taking 8-bit width as an example, EFP can achieve a true 8-bit width without additional bit width expansion, and this feature is particularly important in resource-constrained or high-precision calculation applications.
[0103] S4. Set the initial range value of the exponent offset, and determine whether the calculation results obtained from each operation process exceed the upper and lower limits of the initial range value. If they exceed, automatically adjust the exponent offset.
[0104] S41. Automatic exponent offset, that is, by setting the initial value of the exponent offset, dynamically adjust the numerical range and precision of the numerical representation. Specifically, during the numerical representation process, set an initial exponent offset value. When the value of the calculation result exceeds the set upper and lower limits, automatically adjust the exponent offset. This mechanism effectively compresses the bit width of the exponent bit, and redistributes the saved bit width resources to the mantissa bit, thereby enhancing the representation ability and precision of the mantissa;
[0105] Its operation mechanism is as follows:
[0106] Setting the initial value of the exponent offset: Set an initial exponent offset value at the beginning of the calculation according to the numerical range and precision requirements.
[0107] Dynamic adjustment mechanism: When the value during the calculation process exceeds the current representation range (exceeds the set upper and lower limits), the system automatically adjusts the exponent offset value, readjusts the value to the representable range, and recalculates the exponent bit with the adjusted new offset value, so that the calculation result of the mantissa bit can be retained.
[0108] Additional exponent offset register or other forms of memory: In order to ensure that the exponent offset status of each value is completely tracked and recorded during the calculation process to support the reuse of data, it is necessary to store the numerical offset corresponding to this value in a register or memory, but the exponent offset value does not directly participate in the numerical calculation, but serves as auxiliary information
[0109] The specific algorithm is as follows:
[0110]
[0111] S42. Implementation scheme for vector inner product calculation:
[0112] Without considering the quantization error, taking two vectors of 1×2 and 2×1 as examples, a =
[0113] [x1 y1], Among them, a and b are already the values after 16-bit quantization of high-precision values. Furthermore, the process of one 16-bit vector inner product calculation is as follows:
[0114] ((x1x2) 16bit +(y1y2) 16bit ) 32bit =(x mul +y mul ) 32bit =c 32bit =c 16bit
[0115] For the simulation calculation of the vector inner product of EFP and FP, the extended-bit addition FMS is adopted. If the quantization error is not considered, the relative error calculation formula is as follows:
[0116] a 16bit ·b 16bit =c 16bit
[0117] a 64bit ·b 64bit =a 16bit ·b 16bit =c 64bit
[0118]
[0119] If the quantization error is considered, the relative error calculation formula is:
[0120] a 16bit ·b 16bit =c 16bit
[0121] a 64bit ·b 64bit =c 64bit
[0122]
[0123] Figure 5 , comparing the average relative error of the vector inner product calculation of EFP and FP, it can be found that in the vector inner product calculation, if the relative error is not considered, EFP reduces the relative error by about 83% compared to FP. If the relative error is considered, EFP reduces the relative error by about 33% compared to FP.
[0124] Table 6 compares the complexities of the EPF and FP operators:
[0125] Table 6 Comparison of the Complexities of EFP and FP Operators
[0126]
[0127] Table 7 respectively compares the detailed data of the relative errors of the inner products of EFP and FP with and without considering quantization errors, as well as the relative error reduction rate of EFP relative to FP:
[0128] Table 7 Relative Errors of Inner Products of EFP and FP Vectors with and without Considering Quantization Errors
[0129]
[0130] Figure 3 , which is a simulation comparison of real number calculation errors. Two binary sequences with the same bit width are randomly generated, and EFP and FP operations are performed. The relative error of the calculated result compared with high-precision calculation shows that the relative error of EFP is slightly higher than that of FP.
[0131] Figure 4 , which is a comparison between EFP and FP in multiplication operations. The relative error of multiplication and division calculations of EFP is stable at 10 -17 , while the relative error of multiplication and division calculations of FP represented by the dashed line fluctuates between 10 -1 and 10 -5 .
[0132] Figure 5 , which compares the relative errors of EFP and FP in square root operations. Both follow the same binary shift principle in square root operations, so their relative errors show consistency.
[0133] Figure 7 , which compares the average relative errors of EFP and FP in matrix multiplication. The relative error of EFP16 is reduced by about 20% compared with FP16.
[0134] Table 8 respectively compares the relative errors of the inner products of EFP and FP with and without considering quantization errors, as well as the specific data of the relative error reduction rate of EFP relative to FP.
[0135] Table 8 Relative Errors of EFP and FP Matrix Multiplications with and without Considering Quantization Errors
[0136]
[0137] Figure 8 , which compares the relative errors of matrix inversion of EFP16 and FP16. The relative error of EFP16 is reduced by about 16% compared with FP16.
[0138] Figure 9 , which compares the calculation time delays of EFP and FP in multiplication, division, and square root operations. The performance of EFP in multiplication and division is improved by more than 10 times, and the performance improvement in square root operations is even more than 1000 times.
[0139] Example 2:
[0140] This embodiment provides an exponential low-bitwidth calculation acceleration system optimized based on a lookup table, which is used to implement the above-mentioned exponential low-bitwidth calculation acceleration method optimized based on a lookup table, and includes:
[0141] A receiving module, which is used to receive exponential floating-point data, and the basic format of the exponential floating-point data includes a sign bit, an exponent bit, and a mantissa bit;
[0142] A data processing module, which is used to perform different arithmetic processing operations on the exponential floating-point data according to different arithmetic relationships. The arithmetic relationships specifically include inversion and reciprocal operations, multiplication and division operations, square root operations, and addition and subtraction operations. And in the multiplication and division operations, the multiplication and division operations of the exponential floating-point data are directly converted into addition and subtraction operations of the exponent and the mantissa;
[0143] A lookup table design module, which is used to design and optimize the lookup table, and convert the addition and subtraction operation process into a data lookup process;
[0144] An automatic exponent offset adjustment module, which is used to set the initial range value of the exponent offset, and judge whether the calculation results obtained by each arithmetic processing exceed the upper and lower limits of the initial range value. If so, the exponent offset is automatically adjusted.
[0145] Specifically, the above-mentioned receiving module, data processing module, lookup table design module, and automatic exponent offset adjustment module can be embedded in a computer processing system. The computer, based on the above-provided exponential low-bitwidth calculation acceleration method optimized based on a lookup table, calls the above-mentioned modules to complete the task of solving the common problems of calculation accuracy loss and high complexity of arithmetic components existing in traditional floating-point calculations under low-bitwidth constraints; the above-mentioned receiving module, data processing module, lookup table design module, and automatic exponent offset adjustment module can perform operations according to the specific steps given by the above-mentioned exponential low-bitwidth calculation acceleration method optimized based on a lookup table.
[0146] It should be noted that the division of each module of the above system is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. Moreover, these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the shared remote driving system construction module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above signal processing module can be called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with the ability to process signals. During the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit of the hardware in the processor element or the instruction in the form of software.
[0147] For example, the above-mentioned modules can be one or more integrated circuits configured to implement the above method. For example: one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Singnal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0148] Embodiment 3:
[0149] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The computer program capable of running on the processor is stored in the memory. When the processor loads and executes the computer program, the above-mentioned exponential low-bitwidth calculation acceleration method optimized based on the lookup table is adopted.
[0150] It should be noted that the terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server. Moreover, the terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may further include input / output devices, network access devices, and a bus, etc.
[0151] Furthermore, the processor can be a central processing unit (CPU). Of course, according to the actual usage situation, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be used. The general-purpose processor can be a microprocessor or any conventional processor, etc. This application does not make any restrictions in this regard.
[0152] Embodiment 4:
[0153] The present invention provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the above-mentioned exponential low-bitwidth calculation acceleration method optimized based on a lookup table when executed by a computer processor.
[0154] Among them, the computer program can be stored in a computer-readable medium. The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some middleware form, etc. The computer-readable medium includes any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the computer-readable medium includes, but is not limited to, the above-mentioned components.
[0155] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0156] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. When an element is referred to as being "assembled on", "mounted on", "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only implementation.
[0157] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0158] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
Claims
1. An exponential low-bitwidth calculation acceleration method optimized based on a lookup table, characterized in that, It includes the following steps: Receive exponential floating-point data, the basic format of which includes a sign bit, an exponent bit, and a mantissa bit; Perform different arithmetic processing operations on the exponential floating-point data according to different arithmetic relations. The arithmetic relations specifically include negation and reciprocal operations, multiplication and division operations, square root operation, and addition and subtraction operations. And in the multiplication and division operations, directly convert the multiplication and division operations of the exponential floating-point data into addition and subtraction operations of the exponent and mantissa; Utilize the design and optimization of the lookup table to convert the addition and subtraction operation process into a data lookup process; Set the initial range value of the exponent offset, and judge whether the calculation results obtained from each arithmetic processing exceed the upper and lower limits of the initial range value. If so, automatically adjust the exponent offset.
2. The exponential low-bitwidth calculation acceleration method optimized based on a lookup table according to claim 1, wherein: The basic format of the exponential floating-point data is represented as follows: In the formula, value_dec represents the decimal value, s represents the sign bit, base represents the base number, specifically taking the value of 2, e represents the exponent bit value, bias represents the exponent offset, m represents the mantissa bit value, and m_bit represents the bit width.
3. The exponential low-bitwidth calculation acceleration method optimized based on a lookup table according to claim 2, wherein When performing negation and reciprocal operations, the specific arithmetic processing operations are as follows: Negation: Invert the sign bit; Reciprocal: Keep the sign bit unchanged, invert the mantissa bit and add 1, and perform different addition and subtraction operations on the exponent bit according to different bit widths.
4. The exponential low-bitwidth calculation acceleration method optimized based on a lookup table according to claim 3, characterized in that When performing multiplication and division operations, the specific arithmetic processing operations are as follows: Multiplication: For two exponential floating-point data to be multiplied, first perform an exclusive OR operation on the sign bits, add the exponent bits and mantissa bits of the two exponential floating-point data. If there is a carry in the mantissa bit, add 1 to the exponent bit; Division: For two exponential floating-point data to be divided, first perform the reciprocal operation, that is, invert the mantissa bit and add 1, and perform addition and subtraction operations on the exponent bit according to different bit widths, and then calculate according to the steps of multiplication.
5. The exponential low-bitwidth calculation acceleration method optimized based on a lookup table according to claim 4, characterized in that When performing square root operation, the specific arithmetic processing operations are as follows: Perform the operation of dividing by 2n on the exponent part. In the process of binary representation, it is equivalent to shifting right by n + 1 bits.
6. The exponential low-bitwidth calculation acceleration method optimized based on a lookup table according to claim 5, wherein When performing addition and subtraction operations, the specific arithmetic processing operations are as follows: Addition: If the two addends have the same sign, keep the sign bit unchanged. According to the difference of the exponent bits, select different lookup tables and determine the row of the lookup table according to the mantissa bits of the two addends, and the column is the result; If the two addends have different signs, first judge the magnitudes of the two numbers, determine the sign bit, select different lookup tables according to the difference of the exponent bits, and determine the row of the lookup table according to the mantissa bits of the two addends, and the column is the result; Subtraction: Invert the sign bit of the subtrahend and convert it into an addition operation.
7. The exponential low-bitwidth calculation acceleration method optimized based on a lookup table according to claim 6, wherein, Set the initial range value of the exponent offset, and judge whether the calculation results obtained from each arithmetic processing exceed the upper and lower limits of the initial range value. If so, automatically adjust the exponent offset. The specific operation mechanism is as follows: (71) Initial value setting of exponent offset: Set an initial exponent offset value at the beginning of the calculation according to the numerical range and precision requirements; (72) Dynamic adjustment mechanism: When the value in the calculation process exceeds the upper and lower limits of the initial range value, the system automatically adjusts the exponent offset value, readjusts the value to the representable range, and recalculates the exponent bit with the adjusted new offset value so that the calculation result of the mantissa bit can be retained; (73)An additional exponent offset register or memory for each value is used for re - call.
8. An exponential low-bitwidth calculation acceleration system optimized based on a lookup table, which is used to implement the exponential low-bitwidth calculation acceleration method optimized based on a lookup table according to any one of claims 1 to 7, and is characterized in that Comprising: A receiving module for receiving exponent floating - point data, the basic format of the exponent floating - point data including a sign bit, an exponent bit, and a mantissa bit; A data processing module for performing different arithmetic processing operations on the exponent floating - point data according to different arithmetic relationships, the arithmetic relationships specifically including negation and reciprocal operations, multiplication and division operations, square - root operations, and addition and subtraction operations, and directly converting the multiplication and division operations of the exponent floating - point data into addition and subtraction operations of the exponent and the mantissa in the multiplication and division operations; A look - up table design module for designing and optimizing a look - up table, and converting the addition and subtraction operation process into a data look - up process; An automatic exponent offset adjustment module for setting an initial range value of the exponent offset, determining whether the calculation results obtained from each arithmetic processing exceed the upper and lower limits of the initial range value, and automatically adjusting the exponent offset if it exceeds; 9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, the method for accelerating exponent - type low - width calculation based on look - up table optimization described in any one of claims 1 to 7 is adopted.
10. A storage medium containing computer-executable instructions, characterized in that, The computer - executable instructions are used to execute the method for accelerating exponent - type low - width calculation based on look - up table optimization described in any one of claims 1 to 7 when executed by a computer processor.