Numerical stability-oriented continuous low-bit wide floating point storage method and system
By reconstructing the exponential-mandnumber coupling mechanism of floating point numbers, using exponential floating point format and automatic exponential offset adjustment, the dynamic range limitation and error accumulation problems of low-bit wide floating point storage in IEEE-754 standard are solved, achieving more efficient and stable calculation results.
Patent Information
- Application Number
- CN202510453106.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-08
AI Technical Summary
The existing IEEE-754 standard floating-point number storage scheme has problems such as limited dynamic range, nonlinear accumulation of errors, and discontinuous numerical distribution when the low bit width is low, resulting in significant calculation errors, especially in high-precision sensitive scenarios, facing systemic accuracy defects and efficiency bottlenecks.
By reconstructing the exponential-mandnumber coupling mechanism of floating point numbers, the exponential floating point (EFP) format is adopted, the exponential offset range value is set, the exponential offset range is automatically adjusted to keep the value within a reasonable range, and an exponential floating point encoding is generated to ensure the reasonable allocation of sign bits, exponential bits and mandnumber bits.
It effectively reduces the computational complexity, improves the computing efficiency, reduces hardware resource usage, maintains constant relative errors, and improves the stability and accuracy of calculations, especially in multiplication and division and square calculations, which significantly reduces the calculation delay.
Smart Images

Figure CN120276679A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer storage, and in particular to a continuous low-width floating-point storage method and system for numerical stability. Background Art
[0002] The floating-point storage scheme based on the IEEE-754 standard (IEEE Standard for Floating-point Arithmetic, IEEE-754) always has some limitations.
[0003] Taking the 8-bit floating-point format (FP8) as an example, its quantization characteristics are restricted by the trade-off between the dynamic range of the exponent field and the mantissa resolution. The quantization step of FP8 shows non-linear growth, resulting in a stepped jump in the storage error as the value increases, and the compression of the exponent bit width limits the exponent coverage range. When the value exceeds the maximum representable range, the error will be amplified sharply and may even cause an interruption in the numerical representation; in terms of the analysis of computational error characteristics and operation efficiency, the cumulative basic operation error is manifested as in addition and subtraction operations, the rounding error cumulative effect of FP8 is significant due to insufficient mantissa bits, especially in the accumulation operation, it is easy to cause the "error stagnation phenomenon"; in multiplication, division and complex domain matrix multiplication, the error propagation shows a non-linear amplification trend. For example, in complex multiplication, the effective bits are lost due to the truncation of intermediate results, and the operator complexity of multiplication, division and square root operations is relatively high, which will further introduce computational delay; in terms of the limitations of application scenarios, although the low-width characteristic of FP8 can reduce the bandwidth requirement in scenarios pursuing computational throughput, its significant computational error leads to an increased risk of model convergence failure. Generally speaking, due to the quantization mechanism and limited bit-width constraints, FP8 faces systematic precision defects and efficiency bottlenecks in most high-precision sensitive scenarios. Therefore, the present invention proposes a continuous low-width floating-point storage method and system for numerical stability. Summary of the Invention
[0004] The purpose of the present invention is to provide a continuous low-width floating-point storage method and system for numerical stability, and by reconstructing the exponent-mantissa coupling mechanism of floating-point numbers, to solve the common problems existing in traditional floating-point storage at low bit-widths, such as limited dynamic range, non-linear error accumulation, and discontinuous numerical distribution.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A continuous low-width floating-point storage method for numerical stability, comprising the following steps:
[0006] Receive floating-point data and determine the data type, where the basic format of the floating-point data includes a sign bit, an exponent bit, and a mantissa bit;
[0007] If the data type is a decimal value, perform decimal-to-exponential floating-point processing operations, specifically including calculating the exponent bits, calculating the mantissa bits, and combining the sign bit, exponent bits, and mantissa bits to generate an exponential floating-point binary code;
[0008] If the data type is a floating-point binary value, perform floating-point binary-to-exponential floating-point processing, specifically including keeping the exponent bits unchanged, converting the mantissa bits, and combining the sign bit, exponent bits, and mantissa bits to generate an exponential floating-point code;
[0009] In decimal-to-exponential floating-point processing and floating-point binary-to-exponential floating-point processing, set the exponential offset range value. If it exceeds the upper and lower limits of the initial range value, perform automatic exponential offset adjustment. If it is within the initial range value, keep it unchanged and output the stored result.
[0010] Furthermore, the exponential floating-point retains the basic format of traditional floating-point, namely the sign bit, exponent bits, and mantissa bits. The specific representation format is as shown in Formula 2.1. The floating-point number formula based on the IEEE-754 standard is 2.2:
[0011]
[0012] Where base represents the base, s represents the sign bit, e represents the exponent bit value, bias represents the exponential offset, m represents the mantissa bit value, and m_bit represents the mantissa bit width; in FP, the base is 2. The exponent part of EFP is the same as that of FP, which replaces the base 2 of the FP exponent part with a variable base, and its base can be any positive integer; for the mantissa part, EFP adopts an exponential form, while FP adopts a linear uniform distribution in the mantissa part.
[0013] Furthermore, perform decimal-to-exponential floating-point processing operations, specifically including:
[0014] (31) Calculate the exponent bits: Where N is the input value and bias is the current exponential offset value;
[0015] (32) Calculate the mantissa bits: m = round(log2|mantissa| * 2 m-bit )
[0016] (33) Combine the sign bit s, exponent bits e, and mantissa bits m to generate an EFP binary code.
[0017] Furthermore, perform floating-point binary-to-exponential floating-point processing, specifically as follows:
[0018] (41) Extract the exponent bits e_FP of the floating-point number and directly use it as the exponent bits e_EFP of the exponential floating-point number;
[0019] (42) Convert the mantissa bits:
[0020] Calculation:
[0021] where x is the EFP mantissa bit width and y is the FP mantissa bit width;
[0022] (43) Combine the sign bit s, the exponent bit e_EFP, and the mantissa bit m_EFP to generate the EFP code. Further, the mantissa conversion in step (42) includes an inverse conversion process, that is, exponent floating-point to floating-point:
[0023] (51) Input the exponent floating-point code and extract the mantissa bit m_EFP;
[0024] (52) Calculate the floating-point mantissa bit:
[0025] (53) Combine the floating-point standard format for output.
[0026] Further, in the decimal-to-exponent floating-point processing and the floating-point binary-to-exponent floating-point processing, set the exponent offset range value. If it exceeds the upper and lower limits of the initial range value, perform automatic exponent offset adjustment, as follows:
[0027] (61) Exponent offset initial value setting: Set an initial exponent offset value at the beginning of the calculation according to the numerical range and precision requirements;
[0028] (62) For the calculated exponent bit e, if e > 2^(e_bit) - 1, that is, the exponent bit e exceeds the upper limit:
[0029] then update bias_new = -(e - bias) + 2^(e_bit) - 1;
[0030] (63) If the exponent bit e < 0, that is, the exponent bit e is below the lower limit:
[0031] then update bias_new = -(e - bias);
[0032] (64) If the exponent bit e is within the initial range value, keep bias unchanged;
[0033] (65) Output the calculation result and record the current adjusted exponent offset bias = bias_new.
[0034] According to the second aspect of the present invention, the present invention provides a continuous low-width floating-point storage system for numerical stability, which is used to implement the above-mentioned continuous low-width floating-point storage method for numerical stability, including:
[0035] A receiving module, configured to receive floating-point data and determine the data type, where the basic format of the floating-point data includes a sign bit, an exponent bit, and a mantissa bit;
[0036] A decimal-to-exponential floating-point processing module is used to perform decimal-to-exponential floating-point processing operations if the data type is a decimal value, specifically including calculating the exponent bits, calculating the mantissa bits, and combining the sign bit, exponent bits, and mantissa bits to generate an exponential floating-point binary code.
[0037] A floating-point binary-to-exponential floating-point processing module is used to perform floating-point binary-to-exponential floating-point processing if the data type is a floating-point binary value, specifically including keeping the exponent bits unchanged, converting the mantissa bits, and combining the sign bit, exponent bits, and mantissa bits to generate an exponential floating-point code.
[0038] An automatic exponent offset module is used to set the exponent offset range value in the decimal-to-exponential floating-point processing and floating-point binary-to-exponential floating-point processing. If it exceeds the upper and lower limits of the initial range value, automatic exponent offset adjustment is performed. If it is within the initial range value, it remains unchanged, and the stored result is output.
[0039] According to the third aspect of the present invention, the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, the above-mentioned continuous low-width floating-point storage method for numerical stability is adopted.
[0040] The present invention at least has the following beneficial effects:
[0041] 1. The present invention fully considers the problems of storage error, calculation accuracy, and long delay in the existing FP scheme. This storage scheme adopts the EFP storage method, which can effectively reduce the complexity in subsequent calculations.
[0042] 2. In the subsequent architecture of the present invention, the division operation can share the circuit with the multiplication operation, which improves the operation efficiency while reducing the occupation of hardware resources. At the same time, compared with FP, the relative error of EFP remains constant and there is no obvious fluctuation.
[0043] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the storage method described in the present invention;
[0045] Figure 2 It is a schematic diagram comparing the numerical distributions of EFP and FP of the present invention;
[0046] Figure 3 It is a schematic diagram comparing the local numerical distributions of EFP8 E4M3 and FP8 E4M3 of the present invention;
[0047] Figure 4 It is a schematic diagram of the storage error comparison between EFP8 and FP8 of the present invention;
[0048] Figure 5 It is a schematic diagram of the relative storage error comparison between EFP8 and FP8 of the present invention;
[0049] Figure 6 It is a schematic diagram of the quantization error of EFP16 to FP64 of the present invention;
[0050] Figure 7 It is a schematic diagram of the quantization error of FP16 to FP64 of the present invention. Detailed implementation manners
[0051] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0052] Please refer to Figure 1 , the present invention provides a technical solution: a continuous low-width floating-point storage method for numerical stability, including the following steps:
[0053] S1. Receive floating-point data and judge the data type, where the basic format of the floating-point data includes a sign bit, an exponent bit, and a mantissa bit;
[0054] Exponential Floating Point (EFP) retains the basic format of Floating Point (FP), that is, the sign bit, the exponent bit, and the mantissa bit. The specific representation format of EFP is as follows:
[0055]
[0056] Among them, value_dec represents the decimal value, s represents the sign bit, base represents the base, e represents the exponent bit value, bias represents the exponent offset, m represents the mantissa bit value, and m_bit represents the bit width. The EFP exponent part is basically the same as that of FP. It replaces the base 2 of the FP exponent part with a variable base, and the base can be any positive integer (the variable base uses 2 below, which is the same as FP);
[0057] For the mantissa part, EFP uses an exponential form, while FP uses a linear uniform distribution in the mantissa part. In subsequent calculations, the advantages brought by the exponentiation of the mantissa part will be seen. It can convert complex multiplication and division operations into simple addition operations in exponential representation, achieving higher calculation efficiency;
[0058] S2. If the data type is a decimal value, perform decimal-to-exponential floating-point processing operations, specifically including calculating the exponent bit, calculating the mantissa bit, and combining the sign bit, exponent bit, and mantissa bit to generate an exponential floating-point binary code, as follows:
[0059] (S21) Calculate the exponent bit: where N is the input value and bias is the current exponent bias;
[0060] (S22) Calculate the mantissa bit: m = round(log2|mantissa| * 2 m-bit )
[0061] (S23) Combine the sign bit s, exponent bit e, and mantissa bit m to generate the EFP binary code;
[0062] S3. If the data type is a floating-point binary value, perform floating-point binary-to-exponential floating-point processing, specifically including keeping the exponent bit unchanged, converting the mantissa bit, and combining the sign bit, exponent bit, and mantissa bit to generate an exponential floating-point code, as follows:
[0063] (S31) Extract the exponent bit e_FP of the floating point and directly use it as the exponent bit e_EFP of the exponential floating point;
[0064] (S32) Convert the mantissa bit:
[0065] Calculate:
[0066] where x is the width of the EFP mantissa bit and y is the width of the FP mantissa bit;
[0067] (S33) Combine the sign bit s, exponent bit e_EFP, and mantissa bit m_EFP to generate the EFP code; It should be noted that for the conversion between FP and EFP, there are certain differences in the format definitions of FP and EFP, and the following supplements are required:
[0068] a) Taking the EFP8 E4M3 with an exponent width of 4 bits, a mantissa width of 3 bits, and a base of 2 as an example, analyze the numerical characteristics of EFP; Table 4 shows the comparison of the data formats of EFP8 E4M3 and FP8 E4M3; When representing NaN and Zeros, FP8 does not consider the sign bit, resulting in redundant representation. To make full use of the 8-bit bit space, EFP8 sets NaN to 1.0000.000 and Zeros to 0.0000.000, thus releasing the representation space and expanding the storage range;
[0069] Table 1. Comparison of the data formats of EFP8 E4M3 and FP8 E4M3
[0070]
[0071] b) Table 2 shows the decimal meanings represented by the mantissa widths of the two under different numerical values. It can be seen that the mantissa part of FP is encoded in a uniformly distributed manner, which is beneficial to addition and subtraction operations, but the complexity increases sharply for multiplication, division, and square root operations; The mantissa part of EFP is distributed in exponential form, and the calculation results of multiplication and division can still be accurately represented;
[0072] Table 2. Comparison of the meanings represented by the mantissa bits of EFP8 E4M3 and FP8 E4M3
[0073]
[0074] S4. In the decimal-to-exponential floating-point processing and floating-point binary-to-exponential floating-point processing, set the exponential offset range value. If it exceeds the upper and lower limits of the initial range value, perform automatic exponential offset adjustment. If it is within the initial range value, keep it unchanged and output the stored result;
[0075] Automatic exponential offset means dynamically adjusting the numerical range and precision of numerical representation by setting an initial value of exponential offset. Specifically, during the numerical representation process, set an initial exponential offset value. When the numerical value of the calculation result exceeds the set upper and lower limits, automatically adjust the exponential offset. This mechanism effectively compresses the width of the exponent bits and redistributes the saved bit width resources to the mantissa bits, thereby enhancing the representation ability and precision of the mantissa;
[0076] Its operation mechanism is as follows:
[0077] (41) Initial value setting of exponential offset: Set an initial exponential offset value at the beginning of the calculation according to the numerical range and precision requirements;
[0078] (42) Dynamic adjustment mechanism: When the value during the calculation exceeds the current representation range (exceeds the set upper and lower limits), the system automatically adjusts the exponent offset value, readjusts the value back into the representable range, recalculates the exponent bit based on the adjusted new offset value, so that the calculation result of the mantissa bit can be retained;
[0079] (43) Additional exponent offset register or other forms of memory: To ensure that the exponent offset status of each value is completely tracked and recorded during the calculation and support the re - call of data, the value offset corresponding to this value needs to be stored in a register or memory, but the exponent offset value does not directly participate in the numerical calculation, but serves as auxiliary information;
[0080] The specific algorithm is as follows:
[0081]
[0082] Next, the present invention will be further elaborated in combination with specific embodiments:
[0083] Take 8 - bit EFP (E4M3 format) as an example
[0084] 1. Format definition
[0085] Total bit width: 8 bits
[0086] Sign bit (s): 1 bit
[0087] Exponent bit (e): 4 bits (range: 0 - 15)
[0088] Mantissa bit (m): 3 bits (range: 0 - 7)
[0089] Exponent offset (bias): Default bias = 2^(4 - 1)-1 = 7 (calculated according to e_bit = 4)
[0090] Value range: 8.5E - 3 to 469.5
[0091] Largest representable value: ±2^(15 - 7)×2^(7 / 8)≈±469.5
[0092] Smallest representable value (non - zero): ±2^(-7)×2^(1 / 8)≈±8.5E - 3
[0093] 2. Example of converting decimal to EFP
[0094] Input: N = 12.5
[0095] Step 1: Determine that the sign bit s = 0 (positive number)
[0096] Step 2: Calculate the exponent bit
[0097] a)
[0098] b)
[0099] Step 3: Calculate the mantissa digit
[0100] mantissa = 12.5 / 2^3 = 1.5625
[0101] m = round(log2|mantissa| * 2 m-bit = 5;
[0102] Step 4: Combine the EFP encoding
[0103] - s = 0, e = 1010, m = 101 → EFP encoding: 0 1010 101
[0104] Dynamic offset detection:
[0105] If e_actual > 8 during the conversion process, then adjust bias = -e + 2^(e_bit) - 1;
[0106] 3. Example of FP8 to EFP8 (E4M3)
[0107] Input: FP8 value: 0 1010 101
[0108] Step 1: Extract the FP8 fields
[0109] Sign bit s = 0, exponent e_FP = 1010 (actual exponent 10 - 7 = 3), mantissa m_FP = 101 (implied leading 1, actual value 1.625)
[0110] Step 2: Convert the exponent bit
[0111] e_EFP = e_FP - FP_bias + EFP_bias = 10 (binary 1010)
[0112] Step 3: Convert the mantissa digit (x = 3, y = 3), where x is the EFP mantissa digit width and y is the FP mantissa digit width:
[0113]
[0114] Step 4: Combine the EFP encoding
[0115] s = 0, e = 1010, m = 110 → EFP encoding: 0 1010 110
[0116] 4. Example of EFP (E4M3) to FP8
[0117] Input: EFP encoding: 0 1101 110
[0118] Step 1: Parse the field
[0119] s = 0, e = 13, m = 6
[0120] Step 2: Calculate the actual exponent
[0121] e_actual = 13 - 7 = 6
[0122] Step 3: Calculate the mantissa, x = y = 3
[0123]
[0124] Step 4: Combine FP8
[0125] Sign s = 0, exponent e = 6 + 7 = 13 (1101), mantissa m_FP = 5 (101) → FP8 encoding: 0 1101101
[0126] 5. Example of dynamic exponent offset scenario
[0127] Initial state: bias = 7
[0128] Step 1: Input 600, theoretical exponent
[0129] Step 2: Detect overflow (16 > 15)
[0130] Step 3: Adjust bias = -e_actual + 2^(e_bit) – 1 = -9 + 15 = 6
[0131] Step 4: Recalculate exponent e_new = 9 + 6 = 15 (legal range)
[0132] Result: Output the adjusted bias_c = 6, and use the new offset value for subsequent calculations
[0133] 6. Hardware implementation suggestions
[0134] 6.1. Register configuration:
[0135] EFP_BIAS_REG: Store the current dynamic offset value;
[0136] EFP_OVERFLOW_FLAG: Overflow flag bit.
[0137] This embodiment fully demonstrates the conversion logic and dynamic offset mechanism of 8-bit EFP (E4M3).
[0138] Figure 2Compare the numerical distributions of different exponent bit widths and mantissa bit widths of EFP8 and FP8. The black dashed line in the figure represents the numerical distribution of EFP8 E4M3 with a base of 10, an exponent bit width of 4 bits, and a mantissa bit width of 3 bits. It can be seen that compared with the blue dashed line (EFP8 E4M3 with a base of 2), the numerical range of EFP8 with a base of 10 is larger but the precision is lower. The pink dashed line (EFP8 E5M2), the red dashed line (EFP8 E4M3), and the green dashed line (EFP8 E3M4) change according to the trend of gradually decreasing exponent bit width and gradually increasing mantissa bit width. It can be seen that the numerical range becomes lower and the precision becomes higher.
[0139] Figure 3 Show the local numerical distributions of EFP8 E4M3 and FP8 E4M3, where the x-axis represents the 8-bit binary sequence and the y-axis represents the decimal value of the sequence. The numerical distribution of EFP follows the constant relative error curve, while for FP, as the value gets larger, it deviates more from the initial constant relative error, and at some points, the distribution is suddenly discontinuous and the relative error expands by 10 times.
[0140] Figure 4 Compare the storage errors of EFP8 E4M3 and FP8 E4M3. The blue dashed line part represents EFP8, and the red solid line part represents FP8. The x-axis represents the actual value of the high bit width, and the y-axis represents the error after converting the actual value to the corresponding low bit width representation. The storage error of FP floating-point numbers shows a stepped increase. This oscillating precision causes mutations during the numerical conversion process, affecting the stability and consistency of calculations. In contrast, the storage error of EFP8 maintains a relatively stable growth slope. As the value increases, the storage error of EFP8 shows a linear growth trend. This stability enables EFP8 to avoid sudden jumps during the numerical conversion process, improving the reliability and precision consistency of calculations.
[0141] Figure 5 Compare the relative errors of two storage formats, EFP8 E4M3 and FP8 E4M3. Among them, the relative error of FP8 also shows a kind of oscillating phenomenon, that is, in some numerical ranges, the relative error is very low, while in other ranges, the relative error increases significantly. This volatility may lead to unstable situations during the calculation process, especially when dealing with operations involving large changes in numerical precision. The relative error of EFP8 remains constant and does not fluctuate significantly with the change of the value. Such a constant relative error ensures the stability of the calculation.
[0142] Figure 6 , Figure 7It respectively and intuitively shows the quantization error situations when FP64 is converted to EFP16 and when FP64 is converted to FP16. Through observation, it can be found that the quantization errors of both conversions are maintained at the order of magnitude of 10^-4, and EFP16 shows a smoother characteristic when dealing with the quantization of high-precision numerical values, that is, the fluctuation of its quantization error is relatively small.
[0143] Specifically, in terms of storage, EFP uses an exponential form of mantissa, which has a constant relative precision compared with the linear uniform distribution of FP. By flexibly adjusting the base, exponent bit width, and mantissa bit width, EFP can dynamically adjust the storage range and precision, making it have better applicability in different scenarios. In terms of calculation, the multiplication and division results of EFP can be accurately represented, forming a closed domain, and the relative error is stably at a very low level (10 -17 ) and is independent of the mantissa bit width. Compared with the multiplication and division of FP, its precision stability is significantly improved. In matrix multiplication and inverse calculation, the relative error of EFP is reduced by about 16% and 20% respectively, further verifying its advantage in calculation precision. In terms of calculation latency, the multiplication and division operations of EFP are performed by binary addition and subtraction operations, greatly improving the calculation efficiency. The square root operation is simplified to an operation of dividing by 2 in the exponent part, and the complexity is also significantly reduced. Specifically in terms of latency performance, the performance of the multiplication and division of EFP is improved by more than 10 times, and the square root performance is improved by more than 1000 times. These results highlight its superiority in high-efficiency calculation tasks.
[0144] Generally speaking, EFP is suitable for scenarios that require low-precision and high-efficiency calculations, especially outstanding in tasks involving complex operations such as multiplication, division, and square root. In scientific computing, engineering fields, and applications that perform precise calculations on large-scale data, EFP can improve the calculation efficiency and precision. In addition, its algorithm design and performance also enable it to have broad application prospects in fields such as communication systems, artificial intelligence, and machine learning.
[0145] Embodiment 2:
[0146] This embodiment provides a continuous low-precision floating-point storage system for numerical stability, which is used to implement the above-mentioned continuous low-precision floating-point storage method for numerical stability, including:
[0147] A receiving module, which is used to receive floating-point data and judge the data type, where the basic format of the floating-point data includes a sign bit, an exponent bit, and a mantissa bit;
[0148] A decimal-to-exponential floating-point processing module, which is used to perform decimal-to-exponential floating-point processing operations if the data type is a decimal numerical value, specifically including calculating the exponent bit, calculating the mantissa bit, and combining the sign bit, exponent bit, and mantissa bit to generate an exponential floating-point binary code;
[0149] A floating-point binary to exponential floating-point processing module, which is used to perform floating-point binary to exponential floating-point processing if the data type is a floating-point binary value. Specifically, it includes keeping the exponent bits unchanged, converting the mantissa bits, and combining the sign bit, exponent bits, and mantissa bits to generate an exponential floating-point code;
[0150] An automatic exponent offset module, which is used to set the exponent offset range value in the decimal to exponential floating-point processing and the floating-point binary to exponential floating-point processing. If it exceeds the upper and lower limits of the initial range value, automatic exponent offset adjustment is performed. If it is within the initial range value, it remains unchanged and the stored result is output.
[0151] Specifically, the above receiving module, decimal to exponential floating-point processing module, floating-point binary to exponential floating-point processing module, and automatic exponent offset module can be embedded in a computer processing system. The computer calls the above modules according to the above-provided continuous low-width floating-point storage method for numerical stability to complete the task of storing and representing a new type of exponential floating-point; the above receiving module, decimal to exponential floating-point processing module, floating-point binary to exponential floating-point processing module, and automatic exponent offset module can perform operations according to the specific steps given by the above-mentioned continuous low-width floating-point storage method for numerical stability.
[0152] It should be noted that it should be understood that the division of each module of the above system is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; they can also be partially implemented in the form of software called by a processing element and partially implemented in the form of hardware. For example, the shared remote driving system construction module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above signal processing module is called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit of the hardware in the processor element or the instruction in the form of software.
[0153] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Singnal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain above module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0154] Embodiment 3:
[0155] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, the above-mentioned continuous low-width floating-point storage method for numerical stability is adopted.
[0156] It should be noted that the terminal device may be a computer device such as a desktop computer, a laptop computer, or a cloud server, and the terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may further include input / output devices, network access devices, and a bus, etc.
[0157] Further, the processor may adopt a Central Processing Unit (CPU). Of course, according to actual usage, other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), off-the-shelf Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. may also be adopted. The general-purpose processor may adopt a microprocessor or any conventional processor, etc. The present application does not make any restrictions in this regard.
[0158] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0159] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. When an element is referred to as being "assembled on", "mounted on", "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be intermediate elements at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only embodiments.
[0160] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
[0161] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
Claims
1. A continuous low-width floating-point storage method for numerical stability, characterized in that, It includes the following steps: Receive floating-point data and determine the data type, where the basic format of the floating-point data includes a sign bit, an exponent bit, and a mantissa bit; If the data type is a decimal value, perform decimal-to-exponential floating-point processing operations, specifically including calculating the exponent bit, calculating the mantissa bit, and combining the sign bit, exponent bit, and mantissa bit to generate an exponential floating-point binary code; If the data type is a floating-point binary value, perform floating-point binary-to-exponential floating-point processing, specifically including keeping the exponent bit unchanged, converting the mantissa bit, and combining the sign bit, exponent bit, and mantissa bit to generate an exponential floating-point code; In the decimal-to-exponential floating-point processing and floating-point binary-to-exponential floating-point processing, set the exponent offset range value. If it exceeds the upper and lower limits of the initial range value, perform automatic exponent offset adjustment. If it is within the initial range value, keep it unchanged and output the stored result.
2. The continuous low-width floating-point storage method for numerical stability according to claim 1, characterized in that: The exponential floating-point retains the basic format of the traditional floating-point, that is, the sign bit, exponent bit, and mantissa bit. The specific representation format is as shown in Formula 2.1, and the floating-point number formula based on the IEEE-754 standard is 2.2: EFP: FP: Where base represents the base, s represents the sign bit, e represents the exponent bit value, bias represents the exponent offset, m represents the mantissa bit value, and m_bit represents the mantissa bit width; in FP, the base is 2. The exponent part of EFP is the same as that of FP, which replaces the base 2 of the FP exponent part with a variable base, and its base can be any positive integer; for the mantissa part, EFP uses the exponential form, while FP uses a linear uniform distribution in the mantissa part.
3. The continuous low-width floating-point storage method for numerical stability according to claim 1, characterized in that: Perform decimal-to-exponential floating-point processing operations, specifically including: (31) Calculate the exponent bits: where N is the input value and bias is the current exponent bias; (32) Calculate the number of mantissa digits: m = round(log2|mantissa| * 2 m-bit ); (33) Combine the sign bit s, exponent bit e, and mantissa bit m to generate an EFP binary code.
4. The method for storing continuous low-precision floating-point numbers oriented to numerical stability according to claim 3, characterized in that: Perform floating-point binary-to-exponential floating-point processing, specifically as follows: (41) Extract the exponent bit e_FP of the floating point and directly use it as the exponent bit e_EFP of the exponential floating point; (42) Convert the mantissa bit: Calculation: In the formula, x is the EFP mantissa bit width and y is the FP mantissa bit width; (43) Combine the sign bit s, exponent bit e_EFP, and mantissa bit m_EFP to generate an EFP code.
5. The continuous low-width floating-point storage method for numerical stability according to claim 4, characterized in that The mantissa conversion in step (42) includes a reverse conversion process, that is, exponential floating-point to floating point: (51) Input the exponential floating-point code and extract the mantissa bit m_EFP; (52) Calculate the number of floating-point mantissa bits: (53) Combine the floating-point standard format for output.
6. The continuous low-width floating-point storage method for numerical stability according to claim 5, characterized in that: In the decimal-to-exponential floating-point processing and floating-point binary-to-exponential floating-point processing, set the exponent offset range value. If it exceeds the upper and lower limits of the initial range value, perform automatic exponent offset adjustment, specifically as follows: (61) Exponent offset initial value setting: Set an initial exponent offset value at the beginning of the calculation according to the numerical range and precision requirements; (62) For the calculated exponent bit e, if e > 2^(e_bit) - 1, that is, the exponent bit e exceeds the upper limit: then update bias_new = -(e - bias) + 2^(e_bit) - 1; (63) If the exponent bit e < 0, that is, the exponent bit e is below the lower limit: Then update bias_new = -(e - bias); (64) If the exponent bit e is within the initial range value, keep bias unchanged; (65)Output the calculation result and record the currently adjusted exponent offset bias = bias_new.
7. A continuous low-width floating-point storage system for numerical stability, which is used to implement the continuous low-width floating-point storage method for numerical stability described in any one of claims 1 to 6, characterized in that, It includes: A receiving module, configured to receive floating-point data and determine the data type, wherein the basic format of the floating-point data includes a sign bit, an exponent bit, and a mantissa bit; A decimal-to-exponent floating-point processing module, configured to perform decimal-to-exponent floating-point processing operations if the data type is a decimal value, specifically including calculating the exponent bit, calculating the mantissa bit, and combining the sign bit, exponent bit, and mantissa bit to generate an exponent floating-point binary code; A floating-point binary-to-exponent floating-point processing module, configured to perform floating-point binary-to-exponent floating-point processing if the data type is a floating-point binary value, specifically including keeping the exponent bit unchanged, converting the mantissa bit, and combining the sign bit, exponent bit, and mantissa bit to generate an exponent floating-point code; An automatic exponent offset module, configured to set an exponent offset range value in the decimal-to-exponent floating-point processing and floating-point binary-to-exponent floating-point processing. If the upper and lower limits of the initial range value are exceeded, automatic exponent offset adjustment is performed. If it is within the initial range value, it remains unchanged, and the stored result is output.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, A computer program capable of running on a processor is stored in the memory. When the processor loads and executes the computer program, the continuous low-width floating-point storage method for numerical stability described in any one of claims 1 to 6 is adopted.