Data processing method and device, AI accelerator card, electronic equipment and storage medium
By splitting the target double-precision floating-point number into multiple single-precision floating-point numbers and performing corresponding operations, the problem that AI accelerator cards that do not support double-precision floating-point operations cannot meet the needs of high-precision computing, and achieve higher-precision floating-point operations and more accurate computing results.
Patent Information
- Application Number
- CN202510041613.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
In quantum physics simulation, climate prediction, financial analysis and other scenarios, AI accelerator cards that do not support double-precision floating-point operations cannot meet a small number of higher-precision floating-point operations, resulting in inaccurate calculation results.
By splitting the target double-precision floating-point number into multiple single-precision floating-point numbers, and performing the operation type according to the same mantissa digit segments and then adding them, the operation of double-precision floating-point numbers is replaced by operations between multiple single-precision floating-point numbers.
This enables AI accelerator cards that do not support double-precision floating-point operations to perform higher-precision floating-point operations, meeting the operation needs of specific scenarios and improving the accuracy of the operation results.
Smart Images

Figure CN119960724A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of AI acceleration cards, and more specifically, to a data processing method, device, AI acceleration card, electronic device, and storage medium. Background Art
[0002] At present, most AI accelerator cards tend to support single-precision floating-point operations, but not double-precision floating-point operations for cost considerations. In scenarios such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time-precision operations, sonar analysis, and seismic wave analysis, there are often a small number of higher-precision floating-point operation requirements. For these scenarios, if an AI accelerator card that supports double-precision floating-point operations is used to implement them, the cost will be very unfriendly. If an AI accelerator card that does not support double-precision floating-point operations is used to implement them, there may be situations where the computing requirements are not met, which often leads to inaccurate computing results for the task. Summary of the invention
[0003] The purpose of the embodiments of the present application is to provide a data processing method, device, AI acceleration card, electronic device and storage medium to solve the problem that in scenarios where there is a small amount of higher-precision floating-point computing requirements in quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time-precision computing, sonar analysis, seismic wave analysis, etc., AI acceleration cards that do not support double-precision floating-point operations will not meet the computing requirements, resulting in inaccurate computing results of the task.
[0004] The present application provides a data processing method, which is applied to an AI accelerator card. The method includes: obtaining multiple target double-precision floating-point numbers to be processed, the operation type of this operation task, and the target accuracy of this operation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data; each double-precision floating-point number is processed according to the operation type and the target accuracy. The target double-precision floating-point number is split into multiple single-precision floating-point numbers; wherein: different single-precision floating-point numbers split from the same target double-precision floating-point number correspond to different mantissa segments of the target double-precision floating-point number; the sum of the single-precision floating-point numbers split from the same target double-precision floating-point number is equal to the target double-precision floating-point number; between different target double-precision floating-point numbers, at least part of the single-precision floating-point numbers with the same mantissa segment after splitting can be offset; each of the target double-precision floating-point numbers is replaced by the multiple single-precision floating-point numbers split, and the operation type is respectively performed according to the same mantissa segment and then added.
[0005] In the above implementation process, each target double-precision floating-point number that can be offset by the single-precision floating-point numbers of at least part of the same mantissa bit segment after the split is split into multiple single-precision floating-point numbers, and each target double-precision floating-point number is replaced by the split multiple single-precision floating-point numbers and the operation types are respectively performed according to the same mantissa bit segment and then added. In this way, the operation of double-precision floating-point numbers can be replaced by the operation between multiple single-precision floating-point numbers, so that AI acceleration cards that do not support double-precision floating-point operations can also perform higher-precision floating-point operations, meeting the small amount of higher-precision floating-point operation requirements in scenarios such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time precision operations, sonar analysis, seismic wave analysis, etc., and improving the accuracy of the operation results. In addition, since at least part of the single-precision floating-point numbers with the same mantissa bit segment after splitting can offset each other between different target double-precision floating-point numbers, after replacing each target double-precision floating-point number with multiple split single-precision floating-point numbers and performing operation types respectively according to the same mantissa bit segment, the values of this part of the mantissa bit segment will offset each other, so that the final result can reach the required data range and precision.
[0006] Further, each of the target double-precision floating-point numbers is split into multiple single-precision floating-point numbers according to the operation type and the target precision, including: determining the number of single-precision floating-point numbers to be split according to the operation type and the target precision; determining the number of mantissas of each single-precision floating-point number according to the target precision and the number of single-precision floating-point numbers to be split; for each of the target double-precision floating-point numbers, taking out the mantissa containing the fixed-point number required by the target precision from the target double-precision floating-point number; and sequentially splitting the mantissa taken out from the target double-precision floating-point number into each single-precision floating-point number according to the number of mantissas of each single-precision floating-point number.
[0007] In floating-point operations, the number of bits required for different operations may be different, and a single-precision floating-point number has a maximum of 23 mantissas. Therefore, the number of single-precision floating-point numbers that need to be split is determined according to the operation type and target precision, and then the number of mantissas of each single-precision floating-point number is determined. This can more reasonably achieve the splitting of the target double-precision floating-point number and avoid too many single-precision floating-point numbers being split.
[0008] Furthermore, after splitting the mantissa taken from the target double-precision floating-point number into each single-precision floating-point number in sequence according to the number of mantissas of each single-precision floating-point number, the method further includes: for the i-th single-precision floating-point number, if the highest bit of the mantissa is 0, removing the highest bit until the highest bit of the mantissa becomes 1; wherein i is a value greater than or equal to 2 and less than or equal to the number of single-precision floating-point numbers to be split.
[0009] In the above implementation, the highest-order 0s are continuously removed, which makes the length of the single-precision floating-point number smaller, thereby further reducing the computational complexity and improving the computational efficiency.
[0010] Furthermore, the method also includes: for each single-precision floating-point number: configuring the exponent of the first single-precision floating-point number to be the exponent of the target double-precision floating-point number to be split; configuring the exponent of the i-th single-precision floating-point number to be: the exponent of the target double-precision floating-point number to be split - (i-1)*L - the number of highest bits removed from the i-th single-precision floating-point number; wherein L is the number of mantissas of the single-precision floating-point number.
[0011] In the above implementation, by setting the exponent of the ith single-precision floating-point number to: the exponent of the split target double-precision floating-point number - (i-1)*L - the number of highest bits removed from the ith single-precision floating-point number, the exponent of the ith single-precision floating-point number can be adapted to the ith single-precision floating-point number, and the sum of the single-precision floating-point numbers can be equal to the target double-precision floating-point number, thereby ensuring the reliability of the solution.
[0012] Further, the number of single-precision floating-point numbers to be split is determined according to the operation type and the target precision, including: if the operation type is multiplication, and the target precision is a precision lower than or equal to 46 significant digits, then the number of single-precision floating-point numbers to be split is determined to be 3; if the operation type is multiplication, and the target precision is a precision higher than 46 significant digits, then the number of single-precision floating-point numbers to be split is determined to be 4; if the operation type is not multiplication, and the target precision is a precision lower than or equal to 46 significant digits, then the number of single-precision floating-point numbers to be split is determined to be 2; if the operation type is not multiplication, and the target precision is a precision higher than 46 significant digits, then the number of single-precision floating-point numbers to be split is determined to be 3.
[0013] For floating-point numbers, the number of bits required for multiplication and non-multiplication operations is different, and a single-precision floating-point number has a maximum of 23 digits. Therefore, splitting in the above manner can more reasonably achieve the splitting of the target double-precision floating-point number and avoid too many single-precision floating-point numbers being split out.
[0014] Further, the number of mantissas of each single-precision floating-point number is determined according to the target precision and the number of single-precision floating-point numbers to be split, including: rounding N / M to obtain the number of mantissas L of each single-precision floating-point number; wherein N is the number of mantissas of the fixed-point number required by the target precision; M is the number of single-precision floating-point numbers to be split; wherein, if N / M is a decimal, the number of mantissas of the first M*LN single-precision numbers is determined to be L bits, and the number of mantissas of the following M-(M*LN) single-precision numbers is determined to be L-1 bits.
[0015] The embodiment of the present application also provides a data processing method, which is applied to an AI accelerator card, and the method includes: obtaining multiple target double-precision floating-point numbers to be processed, and the target number of significant digits of this calculation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data; taking out the mantissa of the target number of significant digits from each target double-precision floating-point number and converting it into a long integer; determining the magnification when the mantissa of each target double-precision floating-point number is converted to a long integer; performing calculations on each of the long integers, and reducing the calculation results of each of the long integers according to the magnification of each target double-precision floating-point number to obtain a final calculation result.
[0016] In the above implementation process, the mantissa of the target number of significant digits of each target double-precision floating-point number is taken out and converted into a long integer, and then each long integer is operated, and the operation result of each long integer is reduced according to the magnification factor of each target double-precision floating-point number. In this way, the double-precision floating-point operation is converted into the long integer operation, so that AI accelerator cards that do not support double-precision floating-point operations can also perform higher-precision floating-point operations, meeting the small amount of higher-precision floating-point operation requirements in scenarios such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time precision operations, sonar analysis, seismic wave analysis, etc., and improving the accuracy of the operation results.
[0017] Furthermore, the operation performed on each of the long integers is a periodic function operation; the operation performed on each of the long integers includes: performing the periodic function operation after taking the modulus of each of the long integers.
[0018] In the above implementation, the periodicity of the function is utilized to amplify the target double-precision floating-point number to a long integer and then take the modulus, thereby reducing the numerical range of the data, preventing the long integer numbers converted from each target double-precision floating-point number from exceeding the long integer range after calculation, and improving the calculation reliability.
[0019] Furthermore, the operation performed on each of the long integers is a square root operation; the operation performed on each of the long integers includes: after performing square root operations on each of the long integers, calculating the product K of the square root operation results of each long integer; reducing the operation results of each of the long integers according to the magnification factors of each target double-precision floating-point number to obtain a final operation result, including: calculating K / sqrt(S) to obtain the final operation result; S is a total magnification determined according to the magnification factors of each target double-precision floating-point number; and sqrt represents the square root operation.
[0020] In the above implementation, by converting the square root operation between double-precision floating-point numbers into a square root and then a product operation of long integer numbers, the numerical range of the intermediate results generated in the calculation process can be effectively reduced, thereby making the final calculation result more accurate.
[0021] The embodiment of the present application also provides a data processing device, which is applied to an AI accelerator card, and the device includes: a first acquisition module, which is used to acquire multiple target double-precision floating-point numbers to be processed, the operation type of this operation task, and the target accuracy of this operation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data; a splitting module, which is used to split the target double-precision floating-point number according to the operation type and the target accuracy. Each of the target double-precision floating-point numbers is split into multiple single-precision floating-point numbers; wherein: different single-precision floating-point numbers split from the same target double-precision floating-point number correspond to different mantissa segments of the target double-precision floating-point number; the sum of the single-precision floating-point numbers split from the same target double-precision floating-point number is equal to the target double-precision floating-point number; between different target double-precision floating-point numbers, at least part of the single-precision floating-point numbers with the same mantissa segment after splitting can be offset; a first operation module is used to replace each of the target double-precision floating-point numbers with the multiple single-precision floating-point numbers split, respectively perform the operation type according to the same mantissa segment, and then add them.
[0022] In the above implementation, each target double-precision floating-point number that can be offset by the single-precision floating-point numbers of at least part of the same mantissa segment after the split is split into multiple single-precision floating-point numbers, and each target double-precision floating-point number is replaced by the split multiple single-precision floating-point numbers and the operation types are respectively performed according to the same mantissa segment and then added. In this way, the operation of double-precision floating-point numbers can be replaced by the operation between multiple single-precision floating-point numbers, so that AI acceleration cards that do not support double-precision floating-point operations can also perform higher-precision floating-point operations, meeting the small amount of higher-precision floating-point operation requirements in scenarios such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time precision operations, sonar analysis, seismic wave analysis, etc., and improving the accuracy of the operation results. In addition, since at least part of the single-precision floating-point numbers with the same mantissa bit segment after splitting can offset each other between different target double-precision floating-point numbers, after replacing each target double-precision floating-point number with multiple split single-precision floating-point numbers and performing operation types respectively according to the same mantissa bit segment, the values of this part of the mantissa bit segment will offset each other, so that the final result can reach the required data range and precision.
[0023] The embodiment of the present application also provides a data processing device, which is applied to an AI acceleration card, and the device includes: a second acquisition module, which is used to obtain multiple target double-precision floating-point numbers to be processed, and a target number of significant digits for this calculation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data; a conversion module, which is used to extract the mantissa of the target number of significant digits from each target double-precision floating-point number and convert it into a long integer; a second calculation module, which is used to determine the magnification when the mantissa of each target double-precision floating-point number is converted to a long integer, and to calculate each of the long integers, and to reduce the calculation results of each of the long integers according to the magnification of each target double-precision floating-point number to obtain a final calculation result.
[0024] In the above implementation, the mantissa of the target effective digits of each target double-precision floating-point number is taken out and converted into a long integer, and then each long integer is operated, and the operation result of each long integer is reduced according to the magnification of each target double-precision floating-point number. In this way, the operation of double-precision floating-point numbers is converted into the operation of long integers, so that AI accelerator cards that do not support double-precision floating-point operations can also perform higher-precision floating-point operations, meet the small amount of higher-precision floating-point operation requirements in scenarios such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time precision operations, sonar analysis, and seismic wave analysis, and improve the accuracy of the operation results. In addition, since the single-precision floating-point numbers of at least part of the same mantissa bit segment after splitting can offset each other between different target double-precision floating-point numbers, after each target double-precision floating-point number is replaced with multiple single-precision floating-point numbers split out and the operation type is performed separately according to the same mantissa bit segment, the values of the mantissa bit segment of this part will offset each other, so that the final result can reach the required data range and accuracy.
[0025] An embodiment of the present application also provides an AI acceleration card, including a processor and a memory; the processing unit is used to execute one or more programs stored in the memory to implement any of the above-mentioned data processing methods.
[0026] An embodiment of the present application also provides an electronic device, including the above-mentioned AI acceleration card.
[0027] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores one or more programs. The one or more programs can be executed by one or more processors to implement any of the above-mentioned data processing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0029] Figure 1 A schematic diagram of a flow chart of a first data processing method provided in an embodiment of the present application;
[0030] Figure 2 A schematic diagram of a second data processing method provided in an embodiment of the present application;
[0031] Figure 3 A schematic diagram of the structure of a first data processing device provided in an embodiment of the present application;
[0032] Figure 4 A schematic diagram of the structure of a second data processing device provided in an embodiment of the present application;
[0033] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0035] Embodiment 1:
[0036] In order to solve the problem that AI accelerator cards that do not support double-precision floating-point operations will not meet the computing requirements in scenarios with a small amount of higher-precision floating-point operations such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time-precision computing, sonar analysis, and seismic wave analysis, resulting in inaccurate computing results of tasks, this embodiment provides a data processing method that can be applied to AI accelerator cards that support double-precision floating-point operations. Figure 1 As shown, Figure 1 The following is a flow chart of a data processing method provided in an embodiment of the present application, including:
[0037] S101: Acquire a plurality of target double-precision floating-point numbers to be processed, the operation type of this operation task, and the target precision of this operation task.
[0038] In an embodiment of the present application, the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, and double-precision seismic wave data.
[0039] In the embodiment of the present application, the multiple target double-precision floating-point numbers to be processed may be input by a user, or may be transmitted by an upper-layer application or other device (eg, a data acquisition device).
[0040] In the embodiment of the present application, the operation type of this operation task and the target accuracy of this operation task may be given together with the operation task.
[0041] In the embodiment of the present application, the target precision may be a precision greater than single precision (ie, 23 effective digits) and less than or equal to double precision (ie, 52 effective digits). The target precision is determined by the task requirements.
[0042] S102: Split each target double-precision floating-point number into multiple single-precision floating-point numbers according to the operation type and the target precision.
[0043] In the embodiment of the present application, different single-precision floating-point numbers split from the same target double-precision floating-point number correspond to different mantissa segments of the target double-precision floating-point number. The sum of the single-precision floating-point numbers split from the same target double-precision floating-point number is equal to the target double-precision floating-point number. Between different target double-precision floating-point numbers, at least some of the single-precision floating-point numbers with the same mantissa segment after splitting can be offset.
[0044] In the embodiment of the present application, the single-precision floating-point numbers in the same mantissa bit segment can be offset, which means that the value of the single-precision floating-point numbers in the same mantissa bit segment after operation is within a preset data range. The preset data range may be that the effective digits of the value after operation do not exceed 5 digits.
[0045] For example, assuming that the two single-precision floating-point numbers are 1.55 and 1.49999 respectively, and the operation is a subtraction operation, the calculation result is 0.00001, and the significant digit is 1, so the two can be offset.
[0046] In the embodiment of the present application, the number of single-precision floating-point numbers to be split can be determined according to the operation type and the target precision, and the number of mantissas of each single-precision floating-point number can be determined according to the target precision and the number of single-precision floating-point numbers to be split.
[0047] For each target double-precision floating-point number, the mantissa including the fixed-point number required by the target precision can be taken out from the target double-precision floating-point number, and the mantissa taken out from the target double-precision floating-point number can be split into each single-precision floating-point number in turn according to the number of mantissas of each single-precision floating-point number.
[0048] Exemplarily, if the operation type is multiplication, and the target precision is less than or equal to 46 significant digits, then the number of single-precision floating-point numbers to be split is determined to be three.
[0049] If the operation type is multiplication and the target precision is a precision greater than 46 significant digits, the number of single-precision floating-point numbers to be split is determined to be 4.
[0050] If the operation type is not multiplication and the target precision is less than or equal to 46 significant digits, the number of single-precision floating-point numbers to be split is determined to be 2.
[0051] If the operation type is not multiplication and the target precision is greater than 46 significant digits, the number of single-precision floating-point numbers to be split is determined to be 3.
[0052] For floating-point numbers, there is a difference in the number of bits required for multiplication and non-multiplication (the non-multiplication described in this application, i.e., the operation types that are not multiplication include addition, division and subtraction). Multiplication requires more bits, and a single-precision floating-point number has a mantissa of at most 23 bits. Therefore, by splitting in the above manner, the target double-precision floating-point number can be split more reasonably, avoiding too many single-precision floating-point numbers being split out.
[0053] In the embodiment of the present application, N / M can be rounded up to obtain the number of mantissas of each single-precision floating-point number L, where N is the number of mantissas of the fixed-point number required by the target precision, and M is the number of single-precision floating-point numbers to be split.
[0054] If N / M is a decimal, the number of digits of the mantissa of the first M*LN single-precision numbers is determined to be L, and the number of digits of the mantissa of the following M-(M*LN) single-precision numbers is determined to be L-1.
[0055] In an embodiment of the present application, for the i-th single-precision floating-point number, if the highest bit of the mantissa is 0, the highest bit is removed until the highest bit of the mantissa becomes 1; where i is a value greater than or equal to 2 and less than or equal to the number of single-precision floating-point numbers to be split.
[0056] It can be understood that in the embodiment of the present application, the mantissa taken from the target double-precision floating-point number contains the fixed-point number 1, so the highest bit of the first single-precision floating-point number must be 1, so there is no need to perform the above removal operation.
[0057] After processing based on the above implementation method, the length of the single-precision floating-point number can be made smaller, thereby further reducing the calculation complexity and improving the calculation efficiency.
[0058] Accordingly, in order to ensure that the target double-precision floating-point number can be obtained by adding the single-precision floating-point numbers, for each single-precision floating-point number: the exponent of the first single-precision floating-point number can be configured as the exponent of the split target double-precision floating-point number; the exponent of the i-th single-precision floating-point number can be configured as: the exponent of the split target double-precision floating-point number - (i-1) * L - the number of highest bits removed from the i-th single-precision floating-point number; wherein L is the number of mantissas of the single-precision floating-point number.
[0059] For example, the mantissa of each single-precision number is (assuming N / M is an integer):
[0060] float 1 (the first single-precision floating point number): bits 1 to L of the mantissa of the input data;
[0061] float 2 (the second single-precision floating-point number): Enter the L+1 to 2L bits of the mantissa of the data, and continuously remove the highest bit 0.
[0062] float M (Mth single-precision floating-point number): input the (M-1)*L+1 to M*L bits of the mantissa of the data, and continuously remove the highest bit 0.
[0063] The order code is:
[0064] float 1: the exponent of the target double-precision floating point number;
[0065] float 2: the exponent of the target double-precision floating point number - L - the number of highest bits to be removed;
[0066] float M: exponent of input data - (M-1)*L - number of highest bits to be removed.
[0067] Exemplary:
[0068] Assume the input data is: -198.1290583332232
[0069] The corresponding target double-precision floating-point number is: 11000000 01101000 11000100 001000010011111011110001 00001111 00001010
[0070] Sign bit: 1
[0071] Order code: 1000000 0110 decimal 1030
[0072] Mantissa (including fixed point 1): 1 1000 11000100 00100001 00111110 111100010000111100001010
[0073] Assume that the 32-bit mantissa is taken and divided into two single-precision floating-point numbers, then:
[0074] The mantissa of float1: 1 1000 11000100 001
[0075] The mantissa of float2: 0 0001 00111110 111 The 4 most significant zeros need to be removed, and the result is: 100111110 111
[0076] The exponent of float1: 1030 (binary 1000000 0110),
[0077] The exponent of float2: 1030-16-4=1010
[0078] S103: Replace each target double-precision floating-point number with the split multiple single-precision floating-point numbers, perform operation types on them respectively according to the same mantissa bit segment, and then add them together.
[0079] For example, take the operation AB+C as an example:
[0080] Assuming A and C are double-precision floating-point numbers and B is a single-precision floating-point number, the processing is as follows:
[0081] A=A1+A2
[0082] C=C1+C2
[0083] Then, when performing the operation AB+C, since A1 and C1 belong to the same mantissa segment, and A2 and C2 belong to the same mantissa segment, AB+C will be converted to (A1*B+C1)+(A2*B+C2).
[0084] Assuming that A1*B+C1 has a canceling effect, the exponent of (A1*B+C1) will be reduced to a level close to (A2*B+C2) after the operation, so that it can be added with (A2*B+C2). The interval of valid data during the operation will not exceed the interval of single-precision floating-point numbers, so that effective calculation can be performed, thereby ensuring the accuracy of the calculation result.
[0085] Based on the data processing method provided in this embodiment, each target double-precision floating-point number that can be offset by the single-precision floating-point numbers of at least part of the same mantissa bit segment after the splitting is split into multiple single-precision floating-point numbers, and each target double-precision floating-point number is replaced by the split multiple single-precision floating-point numbers and the operation types are respectively performed according to the same mantissa bit segment and then added. In this way, the operation of double-precision floating-point numbers can be replaced by the operation between multiple single-precision floating-point numbers, so that AI acceleration cards that do not support double-precision floating-point operations can also perform higher-precision floating-point operations, meet the small amount of higher-precision floating-point operation requirements in scenarios such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time precision operations, sonar analysis, seismic wave analysis, etc., and improve the accuracy of the operation results. In addition, since at least part of the single-precision floating-point numbers with the same mantissa bit segment after splitting can offset each other between different target double-precision floating-point numbers, after replacing each target double-precision floating-point number with multiple split single-precision floating-point numbers and performing operation types respectively according to the same mantissa bit segment, the values of this part of the mantissa bit segment will offset each other, so that the final result can reach the required data range and precision.
[0086] Embodiment 2:
[0087] In order to solve the problem that AI accelerator cards that do not support double-precision floating-point operations will not meet the computing requirements in scenarios where there is a small amount of higher-precision floating-point computing requirements such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time-precision computing, sonar analysis, and seismic wave analysis, resulting in inaccurate computing results of tasks, this embodiment provides another data processing method that can be applied to AI accelerator cards that support double-precision floating-point operations. See Figure 2 As shown, Figure 2 The following is a flow chart of a data processing method provided in an embodiment of the present application, including:
[0088] S201: Acquire multiple target double-precision floating-point numbers to be processed and the target number of significant digits for this calculation task.
[0089] In an embodiment of the present application, the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, and double-precision seismic wave data.
[0090] In the embodiment of the present application, the multiple target double-precision floating-point numbers to be processed may be input by a user, or may be transmitted by an upper-layer application or other device (eg, a data acquisition device).
[0091] In the embodiment of the present application, the target number of significant digits for this calculation task may be given together with the calculation task.
[0092] In the embodiment of the present application, the target number of significant digits may be a value greater than single precision (ie, 23 significant digits) and less than or equal to double precision (ie, 52 significant digits). The target number of significant digits is determined by the task requirements.
[0093] S202: Take out the mantissa of the target number of significant digits from each target double-precision floating-point number and convert it into a long integer.
[0094] In the embodiment of the present application, when extracting the mantissa from each target double-precision floating-point number, the fixed-point number 1 must be included.
[0095] In the embodiment of the present application, it can be determined whether J-1023 is greater than or equal to N-1, where J is the exponent of the target double-precision floating point number, and N is the number of mantissa bits to be taken out (including the fixed-point number 1).
[0096] If J-1023 is greater than or equal to N-1, it indicates that the mantissa extracted is the integer part of the double-precision floating point number. In this case, the converted long integer is equivalent to the double-precision floating point number with the decimal point shifted left by ((M-1023)–(N-1)) and then truncated.
[0097] If J-1023 is less than N-1, it indicates that the extracted mantissa contains the decimal part of the double-precision floating-point number. In this case, the converted long integer is equivalent to the double-precision floating-point number whose decimal point is shifted right by ((M-1023)–(N-1)) and then truncated.
[0098] S203: Determine the magnification factor when the mantissa of each target double-precision floating-point number is converted into a long integer.
[0099] It can be understood that if J-1023 is greater than N-1, the magnification is: 1 / ((M-1023)–(N-1)). If J-1023 is less than N-1, the magnification is: ((M-1023)–(N-1)). In particular, if J-1023 is equal to N-1, the magnification is 1.
[0100] S204: Calculate each long integer number, and reduce the calculation result of each long integer number according to the magnification factor of each target double-precision floating-point number to obtain a final calculation result.
[0101] When reducing the calculation results of each long integer number, the total magnification factor should be determined according to the magnification factors of each target double-precision floating-point number to reduce the number, so as to obtain the final calculation result.
[0102] Exemplarily, when the operation performed on each long integer is a periodic function operation, when performing the operation on each long integer, the periodic function operation can be performed after taking the modulus of each long integer.
[0103] For example, taking the cos function operation as an example, assume that cos(A*B*C) is operated, and assume that A, B, and C are all double-precision data, and A, B, and C are all converted into long integer data X, Y, and Z. Assuming that A, B, and C are all converted into long integer data X, Y, and Z, the total magnification is n times. 2π is magnified m times and converted into a long integer number P.
[0104] Then cos(A*B*C)
[0105] =cos((X*Y*Z / n)%(P / m))
[0106] =cos((X*Y*Z%P) / (m / n))
[0107] =cos(((X%P)*(Y%P)*(Z%P))%P) / (m / n))
[0108] That is, cos(A*B*C) can be converted to calculate cos(((X%P)*(Y%P)*(Z%P))%P) / (M / N)). Among them, % is the modulus operation, and m / n is a floating point number.
[0109] In this way, by utilizing the periodicity of the function, the target double-precision floating-point number is enlarged to a long integer and then modulo it, thereby reducing the numerical range of the data, preventing the long integer numbers converted from each target double-precision floating-point number from exceeding the long integer range after calculation, and improving calculation reliability.
[0110] Exemplarily, when the operation performed on each long integer is a square root operation, when performing the operation on each long integer, the square root operation can be performed on each long integer respectively, and then the product K of the square root operation results of each long integer can be calculated. When calculating the final result, K / sqrt(S) can be calculated to obtain the final operation result. Among them, S is the total magnification determined according to the magnification of each target double-precision floating-point number, and sqrt represents the square root operation.
[0111] For example, take sqrt(A*B) as an example, A and B are both double-precision floating point numbers. Convert A and B to long integer numbers X and Y, and determine that the total magnification when A and B are converted to long integer numbers is n, then:
[0112] sqrt(A*B)
[0113] =sqrt((X*Y) / n)
[0114] =sqrt(X*Y) / sqrt(n)
[0115] =sqrt(X)*sqrt(Y) / sqrt(n)
[0116] That is, when calculating sqrt(A*B), it will be converted to calculating sqrt(X)*sqrt(Y) / sqrt(n).
[0117] In this way, by converting the square root operation between double-precision floating-point numbers into a square root and then multiplication operation of long integer numbers, the numerical range of the intermediate results generated in the calculation process can be effectively reduced, thereby making the final calculation result more accurate.
[0118] Based on the solution of this embodiment, the mantissa of the target number of significant digits of each target double-precision floating-point number is taken out and converted into a long integer, and then each long integer is operated, and the operation result of each long integer is reduced according to the magnification factor of each target double-precision floating-point number. In this way, the operation of double-precision floating-point numbers is converted into the operation of long integer numbers, so that AI accelerator cards that do not support double-precision floating-point operations can also perform higher-precision floating-point operations, meet the small amount of higher-precision floating-point operation requirements in scenarios such as quantum physics simulation, climate forecasting, financial analysis, natural language analysis, image analysis, time precision operations, sonar analysis, seismic wave analysis, etc., and improve the accuracy of the operation results.
[0119] In the embodiment of the present application, the solutions of the above embodiment 1 and embodiment 2 can be set in the AI accelerator card at the same time, and the method of embodiment 1 or the method of embodiment 2 can be selected for processing according to specific business needs.
[0120] Embodiment three:
[0121] Based on the same inventive concept, the present application also provides a data processing device 300 and a data processing device 400. Figure 3 and Figure 4 As shown, Figure 3 Shows the use of Figure 1 A data processing apparatus according to the method, Figure 4 Shows the use of Figure 2 The data processing device of the method shown. It should be understood that the specific functions of the device 300 and the device 400 can be referred to the description above, and the detailed description is appropriately omitted here to avoid repetition. The device 300 and the device 400 include at least one software function module that can be stored in the memory in the form of software or firmware or fixed in the operating system of the device 300 and the device 400. Specifically:
[0122] See also Figure 3As shown, the device 300 is applied to an AI acceleration card, and includes: a first acquisition module 301, a splitting module 302 and a first operation module 303. Among them:
[0123] The first acquisition module 301 is used to acquire a plurality of target double-precision floating-point numbers to be processed, the operation type of this operation task, and the target precision of this operation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data;
[0124] The splitting module 302 is used to split each of the target double-precision floating-point numbers into multiple single-precision floating-point numbers according to the operation type and the target precision; wherein: different single-precision floating-point numbers split from the same target double-precision floating-point number correspond to different mantissa segments of the target double-precision floating-point number; the sum of the single-precision floating-point numbers split from the same target double-precision floating-point number is equal to the target double-precision floating-point number; between different target double-precision floating-point numbers, at least some of the split single-precision floating-point numbers in the same mantissa segment can be offset;
[0125] The first operation module 303 is used to replace each of the target double-precision floating-point numbers with the split multiple single-precision floating-point numbers and perform the operation type respectively according to the same mantissa bit segment and then add them.
[0126] In a feasible implementation of the embodiment of the present application, the splitting module 302 is specifically used to: determine the number of single-precision floating-point numbers to be split according to the operation type and the target precision; determine the number of mantissas of each single-precision floating-point number according to the target precision and the number of single-precision floating-point numbers to be split; for each target double-precision floating-point number, take out the mantissa containing the fixed-point number required by the target precision from the target double-precision floating-point number; and split the mantissa taken out from the target double-precision floating-point number into each single-precision floating-point number in sequence according to the number of mantissas of each single-precision floating-point number.
[0127] In this feasible implementation manner, the splitting module 302 is also used to split the mantissa taken from the target double-precision floating-point number into each single-precision floating-point number in sequence according to the number of mantissas of each single-precision floating-point number. For the i-th single-precision floating-point number, if the highest bit of the mantissa is 0, remove the highest bit until the highest bit of the mantissa becomes 1; wherein i is a value greater than or equal to 2 and less than or equal to the number of single-precision floating-point numbers to be split.
[0128] In this feasible implementation manner, the splitting module 302 is further used for: for each single-precision floating-point number: configuring the exponent of the first single-precision floating-point number to be the exponent of the target double-precision floating-point number to be split; configuring the exponent of the i-th single-precision floating-point number to be: the exponent of the target double-precision floating-point number to be split - (i-1)*L - the number of highest bits removed from the i-th single-precision floating-point number; wherein L is the number of mantissas of the single-precision floating-point number.
[0129] In this feasible implementation manner, the splitting module 302 is specifically used for: if the operation type is multiplication, and the target precision is a precision lower than or equal to 46 significant digits, then determining that the number of single-precision floating-point numbers to be split is 3; if the operation type is multiplication, and the target precision is a precision higher than 46 significant digits, then determining that the number of single-precision floating-point numbers to be split is 4; if the operation type is not multiplication, and the target precision is a precision lower than or equal to 46 significant digits, then determining that the number of single-precision floating-point numbers to be split is 2; if the operation type is not multiplication, and the target precision is a precision higher than 46 significant digits, then determining that the number of single-precision floating-point numbers to be split is 3.
[0130] In this feasible implementation manner, the splitting module 302 is specifically used to: round N / M to obtain the number of mantissas L of each single-precision floating-point number; wherein N is the number of mantissas of the fixed-point number required by the target precision; M is the number of single-precision floating-point numbers to be split; wherein, if N / M is a decimal, then it is determined that the number of mantissas of the first M*LN single-precision numbers is L, and the number of mantissas of the following M-(M*LN) single-precision numbers is L-1.
[0131] See also Figure 4 As shown, the device 400 is applied to an AI acceleration card, and includes: a second acquisition module 401, a conversion module 402, and a second operation module 403. Among them:
[0132] The second acquisition module 401 is used to acquire multiple target double-precision floating-point numbers to be processed and the target number of effective digits of this operation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data;
[0133] A conversion module 402 is used to extract the mantissa of the target number of significant digits from each target double-precision floating point number and convert it into a long integer;
[0134] The second operation module 403 is used to determine the magnification factor when the mantissa of each target double-precision floating-point number is converted into a long integer number, and to perform operations on each of the long integer numbers, and to reduce the operation results of each of the long integer numbers according to the magnification factor of each target double-precision floating-point number to obtain a final operation result.
[0135] In a feasible implementation of the embodiment of the present application, the operation performed on each of the long integers is a periodic function operation; the second operation module 403 is specifically used to perform the periodic function operation after taking the modulus of each of the long integers.
[0136] In a feasible implementation of an embodiment of the present application, the operation performed on each of the long integers is a square root operation; the second operation module 403 is specifically used to: after performing square root operations on each of the long integers, calculate the product K of the square root operation results of each long integer; calculate K / sqrt(S) to obtain the final operation result; S is the total magnification determined according to the magnification of each target double-precision floating-point number; the sqrt represents the square root operation.
[0137] It should be understood that, for the sake of brevity, some of the contents described in Embodiment 1 and Embodiment 2 are not repeated in this embodiment.
[0138] Embodiment 4:
[0139] Based on the same inventive concept, this embodiment provides an AI accelerator card, see Figure 5 As shown, it includes a processing unit 501 and a memory 502. Among them:
[0140] The processing unit 501 is used to execute one or more programs stored in the memory 502 to implement the above data processing method.
[0141] The processing unit 501 may include a core processor, which may be a processor implemented based on technologies such as ASIC (Application Specific Integrated Circuit), GPU (Graphics Processing Unit) or FPGA (Field Programmable Gate Array). Alternatively, it may also include TensorCore / CUDACore, where TensorCore is a computing unit specifically used to accelerate AI tasks such as deep learning, while CUDACore is a general computing unit.
[0142] The memory 502 may be a RAM (Random Access Memory), a ROM (Read-Only Memory), a flash memory, a video memory (eg, HBM (High Bandwidth Memory)), etc., but is not limited thereto.
[0143] It is understandable. Figure 5The structure shown is for illustration only. The AI accelerator card may also include Figure 5 More or fewer components as shown, or with Figure 5 For example, it may also have an internal communication bus for realizing communication between the processor 501 and the memory 502; for another example, it may also have an external communication interface, such as a PCIe (Peripheral Component Interconnect Express) interface, a USB (Universal Serial Bus) interface, a CAN (Controller Area Network) bus interface, etc.; for another example, it may also have a power supply component, but this is not a limitation.
[0144] Based on the same inventive concept, this embodiment also provides an electronic device, which includes the aforementioned AI acceleration card.
[0145] In the embodiment of the present application, the electronic device may be, but is not limited to, a server, a host, or other devices.
[0146] Based on the same inventive concept, this embodiment also provides a computer-readable storage medium, such as a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash disk, an SD (Secure Digital Memory Card) card, an MMC (Multimedia Card) card, etc., in which one or more programs for implementing the above steps are stored, and the one or more programs can be executed by one or more processors to implement the above data processing method. No further details will be given here.
[0147] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0148] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0149] Furthermore, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0150] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0151] As used herein, a plurality refers to two or more than two.
[0152] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data processing method, characterized in that: Applied to an AI accelerator card, the method includes: Acquire multiple target double-precision floating-point numbers to be processed, the operation type of this operation task, and the target precision of this operation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data; Each of the target double-precision floating-point numbers is split into a plurality of single-precision floating-point numbers according to the operation type and the target precision; wherein: different single-precision floating-point numbers split from the same target double-precision floating-point number correspond to different mantissa segments of the target double-precision floating-point number; the sum of the single-precision floating-point numbers split from the same target double-precision floating-point number is equal to the target double-precision floating-point number; between different target double-precision floating-point numbers, at least some of the split single-precision floating-point numbers in the same mantissa segment can be offset; Each of the target double-precision floating-point numbers is replaced by the split multiple single-precision floating-point numbers, and the operation type is respectively performed according to the same mantissa bit segment and then added.
2. The data processing method according to claim 1, characterized in that: Splitting each of the target double-precision floating-point numbers into multiple single-precision floating-point numbers according to the operation type and the target precision includes: Determine the number of single-precision floating-point numbers to be split according to the operation type and the target precision; Determine the number of mantissas of each single-precision floating-point number according to the target precision and the number of single-precision floating-point numbers to be split; For each of the target double-precision floating-point numbers, extract the mantissa containing the fixed-point number required by the target precision from the target double-precision floating-point number; The mantissa taken from the target double-precision floating-point number is split into each single-precision floating-point number in sequence according to the number of mantissas of each single-precision floating-point number.
3. The data processing method according to claim 2, characterized in that: After sequentially splitting the mantissa taken from the target double-precision floating-point number into each single-precision floating-point number according to the number of mantissas of each single-precision floating-point number, the method further includes: For the i-th single-precision floating-point number, if the highest bit of the mantissa is 0, remove the highest bit until the highest bit of the mantissa becomes 1; Here, i is a value greater than or equal to 2 and less than or equal to the number of single-precision floating-point numbers to be split.
4. The data processing method according to claim 3, characterized in that: The method further comprises: For each single-precision floating-point number: Configure the exponent of the first single-precision floating-point number to be the exponent of the split target double-precision floating-point number; The exponent of the i-th single-precision floating-point number is configured as: the exponent of the split target double-precision floating-point number - (i-1)*L - the number of highest bits removed from the i-th single-precision floating-point number; Where L is the number of mantissas of a single-precision floating-point number.
5. The data processing method according to claim 2, characterized in that: Determining the number of single-precision floating-point numbers to be split according to the operation type and the target precision includes: If the operation type is multiplication, and the target precision is less than or equal to 46 significant digits, then the number of single-precision floating-point numbers to be split is determined to be 3; If the operation type is multiplication, and the target precision is a precision higher than 46 significant digits, then the number of single-precision floating-point numbers to be split is determined to be 4; If the operation type is not multiplication, and the target precision is less than or equal to 46 significant digits, then the number of single-precision floating-point numbers to be split is determined to be 2; If the operation type is not multiplication, and the target precision is a precision higher than 46 significant digits, then the number of single-precision floating-point numbers that need to be split is determined to be 3.
6. The data processing method according to claim 2, characterized in that: Determining the number of mantissas of each single-precision floating-point number according to the target precision and the number of single-precision floating-point numbers to be split includes: N / M is rounded up to obtain the number of mantissas L of each single-precision floating-point number; wherein N is the number of mantissas of the fixed-point number required by the target precision; and M is the number of single-precision floating-point numbers to be split; If N / M is a decimal, the number of digits of the mantissa of the first M*LN single-precision numbers is determined to be L, and the number of digits of the mantissa of the following M-(M*LN) single-precision numbers is determined to be L-1.
7. A data processing method, characterized in that: Applied to an AI accelerator card, the method includes: Acquire multiple target double-precision floating-point numbers to be processed and the target number of significant digits of this computing task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data; Take the mantissa of the target number of significant digits from each target double-precision floating-point number and convert it into a long integer; Determine the magnification factor when the mantissa of each target double-precision floating-point number is converted into a long integer; The long integers are calculated, and the calculation results of the long integers are reduced according to the magnification factors of the target double-precision floating-point numbers to obtain the final calculation results.
8. The data processing method according to claim 7, characterized in that: The operation performed on each of the long integers is a periodic function operation; The operation on each of the long integers includes: The periodic function operation is performed after taking the modulus of each long integer.
9. The data processing method according to claim 7, characterized in that: The operation performed on each of the long integers is a square root operation; The operation on each of the long integers includes: After performing square root operations on each of the long integers, respectively, a product K of the square root operation results of each long integer is calculated; The operation results of each long integer number are reduced according to the magnification multiple of each target double-precision floating-point number to obtain a final operation result, including: K / sqrt(S) is calculated to obtain a final operation result; S is a total magnification determined according to the magnification of each target double-precision floating-point number; and sqrt represents a square root operation.
10. A data processing device, characterized in that: Applied to an AI accelerator card, the device includes: The first acquisition module is used to acquire a plurality of target double-precision floating-point numbers to be processed, the operation type of this operation task, and the target precision of this operation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data; A splitting module is used to split each of the target double-precision floating-point numbers into multiple single-precision floating-point numbers according to the operation type and the target precision; wherein: different single-precision floating-point numbers split from the same target double-precision floating-point number correspond to different mantissa segments of the target double-precision floating-point number; the sum of the single-precision floating-point numbers split from the same target double-precision floating-point number is equal to the target double-precision floating-point number; between different target double-precision floating-point numbers, at least part of the single-precision floating-point numbers in the same mantissa segment after splitting can be offset; The first operation module is used to replace each of the target double-precision floating-point numbers with the split multiple single-precision floating-point numbers and perform the operation type respectively according to the same mantissa bit segment and then add them.
11. A data processing device, characterized in that: Applied to an AI accelerator card, the device includes: A second acquisition module is used to acquire multiple target double-precision floating-point numbers to be processed and the target number of effective digits of this operation task; wherein the target double-precision floating-point number is one of the following: double-precision quantum physics simulation data, double-precision climate data, double-precision financial data, double-precision text or voice data, double-precision image data, double-precision time data, double-precision sonar data, double-precision seismic wave data; A conversion module, used for taking out the mantissa of the target number of significant digits from each target double-precision floating-point number and converting it into a long integer; The second operation module is used to determine the magnification factor when the mantissa of each target double-precision floating-point number is converted into a long integer number, and to perform operations on each of the long integer numbers, and to reduce the operation results of each of the long integer numbers according to the magnification factor of each target double-precision floating-point number to obtain a final operation result.
12. An AI accelerator card, characterized in that: It comprises a processing unit and a memory; the processing unit is used to execute one or more programs stored in the memory to implement the data processing method as described in any one of claims 1 to 9.
13. An electronic device, characterized in that: Including the AI acceleration card as described in claim 12.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the data processing method according to any one of claims 1 to 9.