Arithmetic logic unit and floating-point number calculation method for processor

By introducing a format conversion module and a hybrid precision multiplier in floating-point number calculation, the exponential range loss during conversion is compensated with exponential auxiliary information, and the accuracy loss problem when high-precision floating-point numbers are converted to low-precision floating-point numbers in mixed-point number calculation is solved, achieving higher calculation accuracy and resource savings.

WO2025175868A1PCT designated stage Publication Date: 2025-08-28HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/135737
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-22
Filing Date
2024-11-29
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

In general matrix multiplication calculation, when using mixed precision calculation, there is a problem of large accuracy loss in the process of converting high-precision floating-point numbers to low-precision floating-point numbers, especially when the exponential range of high-precision floating-point numbers is inconsistent, resulting in large accuracy loss after matrix multiplication.

Method used

By introducing a format conversion module and a mixed precision multiplier in floating point calculation, the exponential auxiliary information of low-precision floating point numbers is retained, and the exponential auxiliary information is used to compensate for the exponential range of the conversion during the multiplication and addition operation, reducing the accuracy loss of the final multiplication and addition result.

Benefits of technology

It effectively reduces the accuracy loss in the calculation process of floating point numbers and ensures the accuracy of multiplication and addition results. Especially when high-precision floating point numbers are converted into low-precision floating point numbers, the accuracy of the final result can be maintained as much as possible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135737_28082025_PF_FP_ABST
    Figure CN2024135737_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of floating-point numbers. Provided are an arithmetic logic unit and a floating-point number calculation method for a processor. The arithmetic logic unit comprises a format conversion module and a mixed-precision multiplier-accumulator, wherein the format conversion module performs a format conversion on a high-precision floating-point number to obtain a low-precision floating-point number, the exponent lost in the conversion process is reserved as exponent auxiliary information, and the mixed-precision multiplier-accumulator uses the exponent auxiliary information for compensating the lost exponent when performing a multiply-accumulate operation on the low-precision floating-point number. Thus, the precision loss of floating-point numbers can be reduced during use of the mixed-precision multiplier-accumulator.
Need to check novelty before this filing date? Find Prior Art

Description

Processor arithmetic logic unit and floating point calculation method

[0001] This application claims priority to Chinese patent application No. 202410201925.6 filed on February 22, 2024, entitled “Arithmetic logic unit of processor and method for floating-point calculation”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of floating-point number technology, and in particular to an arithmetic logic unit of a processor and a floating-point number calculation method. Background Art

[0003] General matrix-to-matrix multiplication (GEMM) is a critical computationally intensive process in scenarios such as artificial intelligence (AI) and high-performance computing (HPC), and mixed-precision computing is a popular computational optimization technique. For example, in AI scenarios, single-precision multiplications and additions can be replaced with half-precision multiplications, while maintaining single-precision additions. This reduces the computational power required for this scenario while maintaining comparable accuracy.

[0004] When using mixed-precision computing for general matrix multiplication, the exponent range of the high-precision floating-point numbers in the matrix is ​​uniformly scaled. These high-precision floating-point numbers are then converted to low-precision floating-point numbers, ensuring that the exponent range of the high-precision floating-point numbers does not overflow. The low-precision floating-point numbers are multiplied together to obtain an intermediate result. This intermediate result is then fed into a standard high-precision adder for subsequent accumulation to obtain the multiplication-addition result. Finally, the exponent range of the multiplication-addition result is inversely scaled to obtain the final multiplication-addition result.

[0005] The exponent ranges of the high-precision floating-point numbers in the matrix may be different. Therefore, if the high-precision floating-point numbers in the matrix are uniformly scaled, some of the high-precision floating-point numbers may overflow the exponent range when converted to low-precision floating-point numbers, resulting in a significant loss of precision after matrix multiplication. Summary of the Invention

[0006] The present application provides an arithmetic logic unit of a processor and a floating-point calculation method, which can reduce precision loss during floating-point calculation.

[0007] In a first aspect, the present application provides an arithmetic logic unit of a processor, which includes a format conversion module and a mixed-precision multiplier and adder. The format conversion module is used to receive multiple first floating-point numbers, convert each first floating-point number into a second floating-point number, obtain multiple second floating-point numbers, and obtain exponential auxiliary information corresponding to each second floating-point number. The precision of the multiple first floating-point numbers is the same, the precision of the multiple second floating-point numbers is the same, the precision of the first floating-point numbers is higher than the precision of the second floating-point numbers, and the exponential auxiliary information corresponding to each second floating-point number is used to compensate for the exponent range lost during the conversion. Multiple combinations of multiple second floating-point numbers and the exponential auxiliary information corresponding to each second floating-point number are input into the mixed-precision multiplier and adder, each combination including two second floating-point numbers for multiplication. The mixed-precision multiplier and adder is used to perform multiplication and addition operations on the multiple combinations based on the exponential auxiliary information corresponding to each second floating-point number to obtain multiplication and addition results of the multiple combinations.

[0008] In the solution shown in the present application, when a high-precision floating-point number is converted into a low-precision floating-point number, the exponent auxiliary information corresponding to the low-precision floating-point number is retained, and the exponent auxiliary information is used in multiplication and addition operations to compensate for part of the exponent range or the entire exponent range lost during the conversion. Therefore, by retaining the exponent auxiliary information, the precision loss of the final multiplication and addition result can be reduced as much as possible.

[0009] In an optional manner, for each second floating-point number, if the first exponent is not within the exponent range corresponding to the second floating-point number, if the first exponent is greater than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is less than or equal to the first exponent and greater than the exponent of the second floating-point number; if the first exponent is less than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is greater than or equal to the first exponent and less than the exponent of the second floating-point number, the first exponent is the exponent of the first floating-point number converted to obtain the second floating-point number, and the compensation exponent is the exponent indicated by the exponent auxiliary information of the second floating-point number. In this way, the exponent indicated by the exponent auxiliary information can compensate for all or part of the precision loss.

[0010] In an optional manner, the format conversion module is used to determine, for each first floating-point number, a first exponent division range to which the exponent of the first floating-point number belongs; in the correspondence between the exponent division range and the exponent auxiliary information, determine the exponent auxiliary information corresponding to the first exponent division range as the exponent auxiliary information corresponding to the second floating-point number obtained by converting the first floating-point number, the exponent division range being obtained based on the exponent range corresponding to the first floating-point number and the exponent range corresponding to the second floating-point number; subtract a first product from the exponent of the first floating-point number to obtain the exponent of the converted second floating-point number, the first product being equal to the product of the exponent auxiliary information corresponding to the converted second floating-point number and a first numerical value, the first numerical value being the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

[0011] In the solution shown in the present application, the exponent range of the high-precision floating-point number is divided into a first exponent division range, and a correspondence between the exponent division range and the exponent auxiliary information is established. When the high-precision floating-point number is converted into a low-precision floating-point number, the exponent auxiliary information can be quickly determined by using this correspondence, and the exponent of the converted low-precision floating-point number can be quickly calculated.

[0012] In an optional manner, the mixed-precision multiplier and adder includes a multiplication unit and an addition unit; the multiplication unit is used to receive multiple combinations, and based on the exponential auxiliary information corresponding to each second floating-point number in each combination, perform a multiplication operation on the two second floating-point numbers in each combination to obtain a first multiplication result corresponding to each combination; the addition unit is used to sum the first multiplication results corresponding to the multiple combinations to obtain a multiplication-addition result. Alternatively, the multiplication unit is used to receive multiple combinations, and perform a multiplication operation on the two second floating-point numbers in each combination to obtain a second multiplication result corresponding to each combination; the addition unit is used to sum the second multiplication results corresponding to the multiple combinations based on the exponential auxiliary information corresponding to each second floating-point number in each combination to obtain a multiplication-addition result.

[0013] In the solution shown in the present application, when exponential auxiliary information is used for compensation, exponential compensation can be performed during multiplication operations or during addition operations, so as to minimize the precision loss of the final multiplication and addition results.

[0014] In an optional manner, the multiplication unit includes an exponent processor, a multiplier and an exclusive-OR operator; the exponent processor is used to, for each combination, determine the first compensation exponent corresponding to each second floating-point number in the combination based on the exponential auxiliary information corresponding to each second floating-point number in the combination; add the exponents of the two second floating-point numbers in the combination to the corresponding two first compensation exponents to obtain the exponent part of the first multiplication result corresponding to the combination; the multiplier is used to, for each combination, multiply the mantissas of the two second floating-point numbers in the combination to obtain the mantissa part of the first multiplication result corresponding to the combination; the exclusive-OR operator is used to, for each combination, perform an exclusive-OR operation on the signs of the two second floating-point numbers in the combination to obtain the sign part of the first multiplication result corresponding to the combination.

[0015] In the solution shown in the present application, an exponent processor is separately provided in the multiplication unit. For each combination, the exponent processor determines the first compensation exponent corresponding to each second floating-point number in the combination based on the exponential auxiliary information corresponding to each second floating-point number in the combination, adds the exponents of the two second floating-point numbers in the combination to the corresponding two first compensation exponents, and obtains the exponential part of the first multiplication result corresponding to the combination. In this way, the compensation exponent is used to compensate the exponential part in the multiplication result, so that the precision loss is reduced.

[0016] In an optional manner, the addition unit includes an exponent processor and an addition sub-unit. The exponent processor is used to, for each combination, determine the first compensation exponent corresponding to each second floating-point number in the combination based on the exponential auxiliary information corresponding to each second floating-point number in the combination; add the first compensation exponent corresponding to each second floating-point number in the combination to the exponent of the corresponding second multiplication result to obtain the updated second multiplication result corresponding to the combination; and the addition sub-unit is used to perform a sum operation on the updated second multiplication results corresponding to the multiple combinations to obtain the multiplication-addition result.

[0017] In the solution shown in the present application, an exponent processor is set in the addition unit. For each combination, the exponent processor determines the first compensation exponent corresponding to each second floating-point number in the combination based on the exponential auxiliary information corresponding to each second floating-point number in the combination, adds the first compensation exponents corresponding to the two second floating-point numbers in the combination to the exponent of the second multiplication result to perform exponential compensation, and obtains an updated second multiplication result. The updated second multiplication result is used to perform the final multiplication and addition operation, so that the precision loss is reduced.

[0018] In an optional manner, the exponent processor is used to, for each second floating-point number in each combination, determine the product of the exponential auxiliary information corresponding to the second floating-point number and the first numerical value as the first compensation index corresponding to the second floating-point number, where the first numerical value is the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

[0019] In an optional manner, the mixed-precision multiplier includes a stream control unit and a multiplication-addition unit. The stream control unit is used to receive the multiple combinations, add the exponents of the two second floating-point numbers in each combination to obtain the exponent sum corresponding to each combination, determine the combination to be calculated in the multiple combinations based on the exponent sum corresponding to each combination, input the combination to be calculated into the multiplication-addition unit, and perform multiplication-addition operations on the combination to be calculated based on the exponential auxiliary information corresponding to each second floating-point number in the combination to be calculated to obtain the multiplication-addition result.

[0020] In the solution shown in the present application, when performing multiplication and addition calculations on multiple combinations, the multiplication and addition operations are performed using exponents and retaining combinations that contribute to the final multiplication and addition results, dynamically saving unnecessary calculations while still obtaining calculation results with sufficient accuracy, which can save computing resources and reduce power consumption.

[0021] In an optional manner, the flow control unit is used to subtract the index sum corresponding to each combination from the maximum index sum corresponding to the multiple combinations to obtain multiple differences, and determine the combination whose index sum among the multiple differences is less than or equal to the target threshold as the combination to be calculated.

[0022] In the scheme shown in the present application, when the products of multiple combinations are accumulated, the product of the combination with a larger exponent sum contributes to the final multiplication and addition result output, and the product of the combination with a smaller exponent sum contributes relatively little to the final multiplication and addition result output. Therefore, the difference between the sum of the exponents of each combination and the maximum sum of the exponents is used to screen the combinations to which the sum of the exponents with a smaller difference belongs as the combinations to be calculated, thereby screening out the combinations that contribute to the final multiplication and addition result output, and thus making the determined multiplication and addition result more accurate while saving power consumption.

[0023] In an optional embodiment, the target threshold is a multiple of the mantissa length of the first floating-point number. In this way, since the mantissa length stored in the memory is generally a multiple of the mantissa length of the first floating-point number, mantissa bits exceeding the target threshold length are not stored and therefore can be ignored during addition, thereby saving computing resources and reducing power consumption.

[0024] In a second aspect, the present application provides a method for floating-point number calculation, the method comprising:

[0025] A plurality of first floating-point numbers are obtained, each of the first floating-point numbers is converted into a second floating-point number to obtain a plurality of second floating-point numbers, and exponent auxiliary information corresponding to each second floating-point number is obtained, wherein the plurality of first floating-point numbers have the same precision, the plurality of second floating-point numbers have the same precision, the precision of the first floating-point number is higher than the precision of the second floating-point number, and the exponent auxiliary information corresponding to each second floating-point number is used to compensate for an exponent range lost during conversion. Based on the exponent auxiliary information corresponding to each second floating-point number, a multiplication-addition operation is performed on a plurality of combinations of the plurality of second floating-point numbers to obtain multiplication-addition results of the plurality of combinations, wherein each combination includes two second floating-point numbers subjected to multiplication operation.

[0026] In an optional manner, for each second floating-point number, if the first exponent is not within the exponent range corresponding to the second floating-point number, if the first exponent is greater than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is less than or equal to the first exponent and greater than the exponent of the second floating-point number; if the first exponent is less than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is greater than or equal to the first exponent and less than the exponent of the second floating-point number; the first exponent is the exponent of the first floating-point number converted to obtain the second floating-point number, and the compensation exponent is the exponent indicated by the exponent auxiliary information of the second floating-point number.

[0027] In an optional manner, the converting of each first floating-point number into a second floating-point number to obtain multiple second floating-point numbers and obtaining exponent auxiliary information corresponding to each second floating-point number includes: for each first floating-point number, determining a first exponent division range to which the exponent of the first floating-point number belongs, and in the correspondence between the exponent division range and the exponent auxiliary information, determining the exponent auxiliary information corresponding to the first exponent division range as the exponent auxiliary information corresponding to the second floating-point number obtained by converting the first floating-point number, the exponent division range being obtained based on the exponent range corresponding to the first floating-point number and the exponent range corresponding to the second floating-point number; subtracting a first product from the exponent of the first floating-point number to obtain the exponent of the converted second floating-point number, the first product being equal to the product of the exponent auxiliary information corresponding to the converted second floating-point number and a first numerical value, the first numerical value being the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

[0028] In an optional manner, the performing of multiplication and addition operations on multiple combinations of the multiple second floating-point numbers based on the exponential auxiliary information corresponding to each second floating-point number to obtain the multiplication and addition results of the multiple combinations includes: performing a multiplication operation on the two second floating-point numbers in each combination based on the exponential auxiliary information corresponding to each second floating-point number to obtain a first multiplication result corresponding to each combination; performing a summation operation on the first multiplication results corresponding to the multiple combinations to obtain the multiplication and addition result; or performing a multiplication operation on the two second floating-point numbers in each combination to obtain a second multiplication result corresponding to each combination; performing a summation operation on the second multiplication results corresponding to the multiple combinations based on the exponential auxiliary information corresponding to each second floating-point number to obtain the multiplication and addition result.

[0029] In an optional manner, the multiplication operation is performed on the two second floating-point numbers in each combination based on the exponent auxiliary information corresponding to each second floating-point number to obtain the first multiplication result corresponding to each combination, including: for each combination, based on the exponent auxiliary information corresponding to each second floating-point number in the combination, determining the first compensation exponent corresponding to each second floating-point number in the combination; adding the exponents of the two second floating-point numbers in the combination to the corresponding two first compensation exponents to obtain the exponent part of the first multiplication result corresponding to the combination; multiplying the mantissas of the two second floating-point numbers in the combination to obtain the mantissa part of the first multiplication result corresponding to the combination; and performing an exclusive-OR operation on the signs of the two second floating-point numbers in the combination to obtain the sign part of the first multiplication result corresponding to the combination.

[0030] In an optional manner, the second multiplication results corresponding to the multiple combinations are summed up based on the exponential auxiliary information corresponding to each second floating-point number to obtain the multiplication-addition result, including: for each combination, based on the exponential auxiliary information corresponding to each second floating-point number in the combination, determining the first compensation exponent corresponding to each second floating-point number in the combination; adding the first compensation exponent corresponding to each second floating-point number in the combination to the exponent of the corresponding second multiplication result to obtain an updated second multiplication result corresponding to the combination; and summing the updated second multiplication results corresponding to the multiple combinations to obtain the multiplication-addition result.

[0031] In an optional manner, for each combination, based on the exponential auxiliary information corresponding to each second floating-point number in the combination, determining the first compensation exponent corresponding to each second floating-point number in the combination includes: for each second floating-point number in each combination, multiplying the exponential auxiliary information corresponding to the second floating-point number by a first numerical value to determine the first compensation exponent corresponding to the second floating-point number, where the first numerical value is the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

[0032] In an optional manner, the performing multiplication and addition operations on multiple combinations of the multiple second floating-point numbers based on the exponential auxiliary information corresponding to each second floating-point number to obtain the multiplication and addition results of the multiple combinations includes: adding the exponents of the two second floating-point numbers in each combination to obtain the exponential sum corresponding to each combination; determining a combination to be calculated in the multiple combinations based on the exponential sum corresponding to each combination; and performing multiplication and addition operations on the combination to be calculated based on the exponential auxiliary information corresponding to each second floating-point number in the combination to be calculated to obtain the multiplication and addition result.

[0033] In an optional manner, the combination to be calculated among the multiple combinations is determined based on the index sum corresponding to each combination, including: subtracting the index sum corresponding to each combination from the maximum index sum corresponding to the multiple combinations to obtain multiple differences; and determining the combination to which the index sum among the multiple differences has a difference less than or equal to a target threshold as the combination to be calculated.

[0034] In an optional manner, the target threshold is a multiple of the mantissa length of the first floating-point number.

[0035] In a third aspect, the present application provides a computing device comprising a processor and a memory, wherein the memory stores at least one computer instruction, and the computer instruction is loaded and executed by the processor to implement the operations performed by the method for floating-point number calculation as described in the second aspect and any optional manner of the second aspect.

[0036] In a fourth aspect, the present application provides a processor comprising an arithmetic logic unit as in the first aspect and any optional method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG1 is a schematic diagram of floating-point numbers of three precisions provided by an exemplary embodiment of the present application;

[0038] FIG2 is a comparative diagram of index ranges provided by an exemplary embodiment of the present application;

[0039] FIG3 is a comparative schematic diagram of dynamic range provided by an exemplary embodiment of the present application;

[0040] FIG4 is a schematic diagram of the structure of an arithmetic logic unit provided by an exemplary embodiment of the present application;

[0041] FIG5 is a schematic diagram of floating point number conversion provided by an exemplary embodiment of the present application;

[0042] FIG6 is a schematic diagram of index auxiliary information provided by an exemplary embodiment of the present application;

[0043] FIG7 is a schematic diagram of a framework of floating-point multiplication and addition operations provided by an exemplary embodiment of the present application;

[0044] FIG8 is a schematic diagram of a framework of a multiplication unit performing a multiplication operation according to an exemplary embodiment of the present application;

[0045] FIG9 is another schematic diagram of a framework of floating-point multiplication and addition operations provided by an exemplary embodiment of the present application;

[0046] FIG10 is a schematic diagram of a framework of an addition unit performing an addition operation according to an exemplary embodiment of the present application;

[0047] FIG11 is a schematic diagram of matrix multiplication provided by an exemplary embodiment of the present application;

[0048] FIG12 is another schematic diagram of the structure of an arithmetic logic unit provided by an exemplary embodiment of the present application;

[0049] FIG13 is a schematic diagram of a power saving method provided by an exemplary embodiment of the present application;

[0050] FIG14 is a schematic diagram of a power saving method provided by another exemplary embodiment of the present application;

[0051] FIG15 is a schematic diagram of the structure of an arithmetic logic unit provided by another exemplary embodiment of the present application;

[0052] FIG16 is a flow chart of a method for floating-point number calculation provided by an exemplary embodiment of the present application;

[0053] FIG17 is a flow chart of a method for floating-point number calculation provided by another exemplary embodiment of the present application;

[0054] FIG18 is a schematic diagram of the structure of a computing device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0056] The following is an explanation of some terminology concepts involved in the embodiments of this application.

[0057] 1. Half-precision floating-point (FP16), consisting of a 1-bit sign, a 5-bit exponent (also called exponent), and a 10-bit mantissa, see Figure 1. In addition, there is an omitted 1-bit integer.

[0058] 2. Single-precision floating-point (FP32), consisting of a 1-bit sign, an 8-bit exponent, and a 23-bit mantissa, see Figure 1. In addition, there is an omitted 1-bit integer.

[0059] 3. Double-precision floating-point (FP64), consisting of a 1-bit sign, an 11-bit exponent, and a 52-bit mantissa, see Figure 1. In addition, there is an omitted 1-bit integer.

[0060] 4. The exponent range of a floating-point number refers to the range of values ​​that the exponent of a floating-point number can take. The exponent range varies for floating-point numbers of different precisions. For example, as shown in Figure 2, the exponent range of a double-precision floating-point number is -1023 to 1024, the exponent range of a single-precision floating-point number is -127 to 128, and the exponent range of a half-precision floating-point number is -15 to 16.

[0061] 5. The dynamic range of floating point numbers refers to the range of floating point values. The dynamic range of floating point numbers of different precisions is different. For example, referring to Figure 3, which shows double-precision floating point numbers and single-precision floating point numbers, the dynamic range of double-precision floating point numbers is -2 1023 to -2 -1022 , 2 -1022 to 2 1023 , the dynamic range of single-precision floating point numbers is -2 127 to -2 -126 , 2 -126 to 2 127 .

[0062] Floating-point calculations are used in many fields, including graphics processing, astronomy, and medicine. Floating-point calculations consume a lot of power, and the higher the precision of the floating-point number, the greater the power consumption. Currently, mixed-precision multipliers are used to reduce the power consumption of floating-point calculations. The multiplication and addition process involves converting the exponent of a high-precision floating-point number to that of a low-precision floating-point number, rounding the mantissa of the high-precision floating-point number to that of a low-precision floating-point number, and then performing the calculation using a low-precision floating-point multiplier. This converts the high-precision floating-point multiplication into a low-precision floating-point multiplication, reducing the power consumption of the floating-point multiplication. A standard high-precision floating-point adder is then used for accumulation. This calculation method results in two aspects of precision loss: the loss of precision caused by floating-point conversion and the loss of precision caused by multiplying low-precision floating-point numbers. The precision loss caused by multiplying low-precision floating-point numbers is relatively small compared to the precision loss caused by conversion. This is because the two input floating-point numbers are N / 2 bits long, and the output after multiplication is N bits long. Multiplying two N / 2-bit floating-point numbers preserves the resulting floating-point number intact, resulting in minimal precision loss. For example, a half-precision floating-point number with (sign, exponent, mantissa) = (1, 5, 10) can be multiplied to a floating-point precision of approximately (1, 6, 21). Alternatively, a brain floating point number (BF16) with (sign, exponent, mantissa) = (1, 8, 7) can be multiplied to a floating-point precision of approximately (1, 9, 15). The output can be stored as a single-precision floating-point number (1, 8, 23) with minimal precision loss. Therefore, when using mixed-precision multipliers for calculations, the loss of precision is primarily due to the conversion. For example, as shown in Figure 3, the exponent ranges of double-precision floating-point numbers and single-precision floating-point numbers are different. When the exponent range of a double-precision floating-point number falls within the exponent range of a single-precision floating-point number, the dynamic range does not overflow when converting the double-precision floating-point number to single-precision floating-point, resulting in minimal precision loss. However, when the exponent range of a double-precision floating-point number falls outside the exponent range of a single-precision floating-point number, the dynamic range overflows when converting the double-precision floating-point number to single-precision floating-point, resulting in significant precision loss.

[0063] To reduce the precision loss caused by floating-point conversion, high-precision floating-point numbers are currently uniformly multiplied by a coefficient. This ensures that the exponent range of the high-precision floating-point numbers does not overflow when the high-precision floating-point numbers are converted to low-precision floating-point numbers. As a result, the exponent ranges of the high-precision floating-point numbers may vary. Therefore, they are uniformly scaled to high-precision floating-point numbers. This means that when some high-precision floating-point numbers are converted to low-precision floating-point numbers, the exponent range will still overflow. This makes it suitable for low-precision multiplication and addition scenarios. For example, training and inference scenarios in the AI ​​field. However, for scenarios with higher precision requirements, the precision loss is still relatively large and cannot meet the accuracy requirements.

[0064] Based on this, an embodiment of the present application provides an arithmetic logic unit of a processor, which, when converting a high-precision floating-point number into a low-precision floating-point number, retains the exponent auxiliary information corresponding to each low-precision floating-point number to compensate for part of the exponent range or the entire exponent range lost during the conversion. The exponent auxiliary information is used during multiplication and addition operations to compensate for part of the exponent range or the entire exponent range lost during the conversion. Therefore, by retaining the exponent auxiliary information, the precision loss of the final multiplication and addition result can be reduced as much as possible, thereby ensuring that the precision of the final multiplication and addition result is lossless as much as possible.

[0065] In the embodiments of the present application, multiplication and addition operations are performed on multiple combinations of multiple first floating-point numbers as an example. Each combination includes two first floating-point numbers to be multiplied. The two first floating-point numbers in each combination are multiplied to obtain multiplication results corresponding to the multiple combinations. The multiplication results corresponding to the multiple combinations are accumulated to obtain a final multiplication and addition result. For example, if matrix A and matrix B are multiplied, the row elements and column elements in matrix A are multiplied and then added. The multiple combinations are combinations of row elements and column elements that are multiplied.

[0066] The embodiments of the present application provide an arithmetic logic unit (ALU) of a processor, which may be a chip such as a central processing unit (CPU), an AI, or a graphics processing unit (GPU). The processor may be a processor in a computing device. The computing device may be any device that needs to perform floating-point multiplication and addition calculations. For example, the computing device may be a mobile terminal such as a mobile phone or tablet computer, a computer device such as a desktop or laptop computer, or a server.

[0067] Referring to Figure 4, the arithmetic logic unit includes a format conversion module and a mixed precision multiplier-adder, and the format conversion module receives multiple first floating-point numbers, and the precision of the multiple first floating-point numbers is the same. The sign of each first floating-point number is determined to be the sign of the second floating-point number, the exponent e2 of each first floating-point number is converted to the exponent e1 required by the second floating-point number, and the mantissa is converted to the mantissa required by the second floating-point number by truncation (truncate) and rounding (rounding) mode, to obtain multiple second floating-point numbers, the precision of the multiple second floating-point numbers is the same, and the precision of the first floating-point number is higher than the precision of the second floating-point number. For example, the first floating-point number is a double-precision floating-point number, the symbol s2 is 1 bit, the exponent e2 is 11 bits, and the mantissa m2 is 52 bits, and the second floating-point number is a single-precision floating-point number, the symbol s1 is 1 bit, the exponent e1 is 8 bits, and the mantissa m1 is 23 bits. For each first floating-point number, the format conversion module splits the exponent e2 of each first floating-point number into two parts: one part is the exponent of the second floating-point number, denoted as e1, and the other part is exponent auxiliary information, denoted as e1sb. The exponent auxiliary information corresponding to each second floating-point number is used to compensate for the exponent range lost during conversion.

[0068] The format conversion module inputs multiple combinations and exponential auxiliary information corresponding to each second floating-point number into a mixed-precision multiplier-adder. The mixed-precision multiplier receives the multiple combinations and the exponential auxiliary information and performs multiplication and addition operations on the multiple combinations based on the exponential auxiliary information to obtain multiplication and addition results of the multiple combinations.

[0069] It should be noted that the first floating-point number and the second floating-point number may be floating-point numbers of adjacent precisions, or may not be floating-point numbers of adjacent precisions. For example, the first floating-point number may be a double-precision floating-point number, and the second floating-point number may be a single-precision floating-point number. For another example, the first floating-point number may be a double-precision floating-point number, and the second floating-point number may be a half-precision floating-point number.

[0070] Optionally, referring to FIG5 , the format conversion module includes an exponent splitter and a truncation and rounding unit. The exponent splitter splits the exponent of the first floating-point number into the exponent of the second floating-point number and exponent auxiliary information, and the truncation and rounding unit truncates and rounds the mantissa of the first floating-point number to the mantissa of the second floating-point number. In this way, a 2N-bit high-precision floating-point number is converted into an N-bit low-precision floating-point number.

[0071] In an optional manner, the process of converting the exponent of the first floating-point number to the exponent of the second floating-point number and determining the exponent auxiliary information is as follows:

[0072] For each first floating-point number, the format conversion module obtains the correspondence between the exponent division range and the exponent auxiliary information. The exponent division range is obtained based on the exponent range of the first floating-point number (referred to as the first exponent range) and the exponent range of the second floating-point number (referred to as the second exponent range). The second exponent range is the first exponent division range, including the lower boundary value, but excluding the upper boundary value. The maximum value of the second exponent range is the starting point of the upward division, and the minimum value is the starting point of the downward division. Every first value is divided to obtain an exponent division range, and the exponent division range includes the lower boundary value but does not include the upper boundary value. The maximum value of the last exponent division range divided upward is the maximum value of the first exponent range, and the minimum value of the last exponent division range divided downward is the minimum value of the first exponent range, and the exponent division range includes the upper boundary value but does not include the lower boundary value. Based on this strategy, the distance between the maximum and minimum values ​​of the last exponent division range is not equal to the first value. The first value is the distance between the maximum and minimum values ​​of the second exponent range. Each exponent division range corresponds to different exponent auxiliary information.

[0073] The format conversion module then determines the exponent partition range to which the exponent of the first precision floating point number belongs, referred to as the first exponent partition range. Based on the correspondence between the exponent auxiliary information and the exponent partition range, the exponent auxiliary information corresponding to the first exponent partition range is determined. The exponent of the second floating point number is then calculated using formula (1). e1 = e2 - a* e1 sb (1)

[0074] In formula (1), e1 represents the exponent of the second floating-point number, e2 represents the exponent of the first floating-point number, a represents the first value, and e1sb represents the exponent auxiliary information corresponding to the second floating-point number.

[0075] For example, referring to Figure 6, the first floating-point number is a single-precision floating-point number, and the first exponent ranges from -127 to 128. The second floating-point number is a half-precision floating-point number, and the second exponent ranges from -15 to 16. The first numerical value corresponding to the second floating-point number is equal to the distance between -15 and 16, that is, the value is 31. If the exponent of the first floating-point number is between -15 and 16, excluding 16, then the exponent auxiliary information corresponding to the second floating-point number is determined to be 0. If the exponent of the first floating-point number is between 16 and 47, including 16 but excluding 47, then the exponent auxiliary information corresponding to the second floating-point number is determined to be 1. If the exponent of the first floating-point number is between 47 and 78, including 47, then the exponent auxiliary information corresponding to the second floating-point number is determined to be 2. If the exponent of the first floating-point number is between 77 and 109, including 77 but excluding 109, then the exponent auxiliary information corresponding to the second floating-point number is determined to be 3. If the exponent of the first floating-point number is between 109 and 128, inclusive, the exponent auxiliary information corresponding to the second floating-point number is determined to be 4. If the exponent of the first floating-point number is between -46 and -15, inclusive, and excluding -46, the exponent auxiliary information corresponding to the second floating-point number is determined to be -1. If the exponent of the first floating-point number is between -77 and -46, inclusive, and excluding -77, the exponent auxiliary information corresponding to the second floating-point number is determined to be -2. If the exponent of the first floating-point number is between -108 and -77, inclusive, and excluding -108, the exponent auxiliary information corresponding to the second floating-point number is determined to be -3. If the exponent of the first floating-point number is between -127 and -108, inclusive, and excluding -127 and -108, the exponent auxiliary information corresponding to the second floating-point number is determined to be -4.

[0076] It should be noted that the above correspondence can be stored in a table or other manner. In addition, when storing exponential auxiliary information, the exponential auxiliary information can be stored, and the correspondence between the exponential auxiliary information and the compensation index can also be stored, and the compensation index is the product of the exponential auxiliary information and the first numerical value. The division method of the above-mentioned exponential division range can be set according to actual needs, and the embodiment of the present application is not limited. For example, the exponential range corresponding to the second floating-point number is divided into two exponential division ranges, the division above zero is the first exponential division range, and the division below zero is the second exponential division range, and then the maximum value of the exponential range is divided upward, and the size of each exponential division range is equal to the size of the first exponential division range, and the minimum value of the exponential range is divided downward, and the size of each exponential division range is equal to the size of the second exponential division range. Each exponential division range corresponds to different exponential auxiliary information.

[0077] In another optional manner, the process of converting the exponent of the first floating-point number to the exponent of the second floating-point number and determining the exponent auxiliary information is as follows:

[0078] When converting the exponent using the saturation truncation method, if the exponent of the first floating-point number exceeds the upper boundary of the exponent range of the second floating-point number, the upper boundary value of the exponent range is determined as the exponent of the second floating-point number, and the exponent auxiliary information is the exponent of the first floating-point number minus the exponent of the second floating-point number. If the exponent of the first floating-point number exceeds the lower boundary of the exponent range of the second floating-point number, the lower boundary value of the exponent range is determined as the exponent of the second floating-point number, and the exponent auxiliary information is the exponent of the first floating-point number minus the exponent of the second floating-point number. If the exponent of the first floating-point number does not exceed the exponent range of the second floating-point number, the exponent of the first floating-point number is simply represented using the bits required by the second floating-point number, the exponent does not change, and the exponent auxiliary information is 0.

[0079] It should be noted that when converting the mantissa by truncation and rounding, either method can be used, and the embodiments of the present application do not limit this.

[0080] In an optional manner, from the perspective of the plurality of second floating-point numbers as a whole, the exponent auxiliary information corresponding to the plurality of second floating-point numbers is used to compensate for the exponent range lost during conversion, thereby compensating for the dynamic range lost during conversion. From the perspective of each second floating-point number in the plurality of second floating-point numbers individually, the exponent auxiliary information corresponding to the second floating-point number indicates a compensation exponent. If the exponent of the first floating-point number converted to obtain the second floating-point number is the first exponent, and when the first exponent does not fall within the exponent range of the second floating-point number, if the first exponent is greater than 0, then the sum of the compensation exponent and the exponent of the second floating-point number is less than or equal to the first exponent and greater than the exponent of the second floating-point number. When the sum of the compensation exponent and the exponent of the second floating-point number is equal to the first exponent, it indicates that the entire exponent range loss when obtaining the second floating-point number has been compensated. When the sum of the compensation exponent and the exponent of the second floating-point number is less than the first exponent, it indicates that a portion of the exponent range loss when obtaining the second floating-point number has been compensated. If the first exponent is less than 0, then the sum of the compensation exponent and the exponent of the second floating-point number is greater than or equal to the first exponent and less than the exponent of the second floating-point number. When the sum of the compensation exponent and the exponent of the second floating-point number is equal to the first exponent, it indicates that all exponent range loss in obtaining the second floating-point number has been compensated. When the sum of the compensation exponent and the exponent of the second floating-point number is greater than the first exponent, it indicates that part of the exponent range loss in obtaining the second floating-point number has been compensated. When the first exponent falls within the exponent range of the second floating-point number, the compensation exponent indicated by the exponent auxiliary information is 0.

[0081] Furthermore, to ensure conversion accuracy without loss, exponent auxiliary information is set such that, when the first exponent is greater than 0, the sum of the compensation exponent and the exponent of the second floating-point number may be greater than the first exponent, but may not exceed the exponent range corresponding to the first floating-point number. When the first exponent is less than 0, the sum of the compensation exponent and the exponent of the second floating-point number may be less than the first exponent, but may not exceed the exponent range corresponding to the first floating-point number.

[0082] In an optional manner, there are multiple ways to use exponential auxiliary information to perform multiplication and addition operations on multiple combinations. Two feasible ways are provided below.

[0083] Method 1: During multiplication, exponential auxiliary information is used for compensation. Accordingly, the mixed-precision multiplier includes a multiplication unit and an addition unit, see Figure 7. The multiplication unit receives multiple combinations and exponential auxiliary information sent by the format conversion module. For each combination, the exponential auxiliary information corresponding to the two second floating-point numbers in the combination is used to multiply the two second floating-point numbers in the combination to obtain the multiplication result corresponding to the combination, which is called the first multiplication result. The multiplication unit inputs the first multiplication results corresponding to the multiple combinations to the addition unit, and the addition unit sums the first multiplication results corresponding to the multiple combinations to obtain the multiplication and addition results corresponding to the multiple combinations. Among them, the addition unit includes an adder and an accumulator. Each time, the first multiplication results corresponding to two combinations are input, and the summation operation is performed to obtain an intermediate result. The intermediate result is then stored in the accumulator, and then the intermediate result is read from the accumulator to the adder, and summed with the first multiplication result of another combination until the final multiplication and addition result is obtained. In Figure 7, the exponent and mantissa of the first floating-point number are represented as (e2, m2), and the exponent and mantissa of the second floating-point number are represented as (e1, m1). The addition unit is also represented as (e2, m2) to illustrate that the addition part uses high-precision addition. For example, if the first floating-point number is a single-precision floating-point number, e2 and m2 are 8 and 23 respectively, and the second floating-point number is a half-precision floating-point number, e1 and m1 are 5 and 10 respectively.

[0084] Optionally, during a multiplication operation, the multiplication unit includes an exponent processor, a multiplier, and an exclusive-OR operator, as shown in FIG8 . For each combination, the exponent processor uses the exponential auxiliary information corresponding to each second floating-point number in the combination to determine a first compensation exponent corresponding to each second floating-point number in the combination, adds the exponent of each second floating-point number in the combination to the corresponding first compensation exponent to obtain an addition result, and then adds the addition results of the two second floating-point numbers in the combination to obtain the exponential portion of the first multiplication result corresponding to the combination.

[0085] For each combination, the multiplier performs a multiplication operation on the mantissas of the two second floating-point numbers in the combination to obtain a mantissa portion of the first multiplication result corresponding to the combination. For each combination, the exclusive-OR operator performs an exclusive-OR operation on the signs of the two second floating-point numbers in the combination to obtain a sign portion of the first multiplication result corresponding to the combination.

[0086] Optionally, during multiplication, the multiplication unit further includes a normalizer (see FIG8 ). For each combination, the multiplier inputs the mantissa corresponding to the combination to the normalizer, and the exponent processor inputs the exponent corresponding to the combination to the normalizer. The normalizer normalizes the mantissa and the exponent corresponding to the combination, ensuring that the first multiplication result corresponding to the combination is within a specified range and has the same precision as the first floating-point number. During accumulation, the first multiplication results corresponding to multiple combinations are all normalized first multiplication results.

[0087] It should be noted that for each combination, if the exponent part of the first multiplication result corresponding to the combination exceeds the second value, the exponent part of the first multiplication result corresponding to the combination is updated to the second value, and the second value is the maximum value of the exponent range corresponding to the first floating-point number minus one. For example, the first floating-point number is a single-precision floating-point number with an exponent range of -127 to 128. The exponent part of the first multiplication result is 150, and the exponent part of the first multiplication result is updated to 127. This is because the exponent of the first floating-point number is 8 bits and can represent a maximum of 128, but 128 is represented as infinity, so the exponent part is the maximum value of the exponent range corresponding to the first floating-point number minus one. If the exponent part of the first multiplication result corresponding to the combination is less than the third value, the exponent part of the first multiplication result corresponding to the combination is updated to the third value, and the third value is the minimum value of the exponent range corresponding to the first floating-point number plus one. For example, if the first floating-point number is a single-precision floating-point number with an exponent range of -127 to 128, and the exponent of the first multiplication result is -140, the exponent of the first multiplication result is updated to -126. This is because the exponent of the first floating-point number is 8 bits and can represent a maximum of -127, but -127 is represented as 0, so the minimum value of the exponent range corresponding to the first floating-point number is added by one.

[0088] Alternatively, for each combination, if the exponent of the first multiplication result corresponding to the combination exceeds the second value, the first multiplication result corresponding to the combination is recorded as infinity. If the exponent of the first multiplication result corresponding to the combination is less than the third value, the first multiplication result corresponding to the combination is recorded as 0.

[0089] Method 2: During addition operations, exponential auxiliary information is used for compensation. Accordingly, the mixed-precision multiplier includes a multiplication unit and an addition unit, see Figure 9. The multiplication unit receives multiple combinations sent by the format conversion module. For each combination, the exponents of the two second floating-point numbers in the combination are added, the signs are XORed, and the mantissas are multiplied to obtain the multiplication result corresponding to the combination, which is called the second multiplication result. The multiplication unit inputs the second multiplication results corresponding to the multiple combinations to the addition unit. The addition unit receives the second multiplication results corresponding to the multiple combinations and receives the exponential auxiliary information sent by the format conversion module. The addition unit uses the exponential auxiliary information corresponding to each second floating-point number to perform a sum operation on the second multiplication results corresponding to the multiple combinations to obtain the multiplication and addition results corresponding to the multiple combinations. The addition unit includes an exponent processor, an addition subunit, and an accumulator. Each time, the second multiplication results corresponding to two combinations are input. The addition subunit and the exponent processor cooperate to perform a summation operation to obtain an intermediate result, which is then stored in the accumulator. The intermediate result is then read from the accumulator to the adder and summed with another combination until the final multiplication and addition result is obtained. In Figure 9, the exponent and mantissa of the first floating-point number are represented as (e2, m2), and the exponent and mantissa of the second floating-point number are represented as (e1, m1). The addition unit position is also represented by (e2, m2) to illustrate that the addition part uses a high-precision addition operation. For example, if the first floating-point number is a double-precision floating-point number, e2 and m2 are 11 and 52 respectively, and the second floating-point number is a single-precision floating-point number, e1 and m1 are 8 and 23 respectively.

[0090] Optionally, the addition subunit includes a subtractor, a shifter, a sign operator, an adder, and a normalizer, as shown in Figure 10. For each combination, the exponent processor uses the exponential auxiliary information corresponding to each second floating-point number in the combination to determine the first compensation exponent corresponding to each second floating-point number in the combination. The exponent processor adds the first compensation exponents corresponding to the two second floating-point numbers in the combination to obtain an addition result, and then adds the addition result to the exponent of the second multiplication result corresponding to the combination to obtain an updated exponent of the second multiplication result, i.e., an updated second multiplication result for the combination. Assuming that the multiple combinations include a first combination and a second combination, the second multiplication result corresponding to the first combination and the second multiplication result corresponding to the second combination are added as an example for explanation. Assuming that the sum of the second multiplication result corresponding to the first combination and the second multiplication result corresponding to the second combination represents a third result, the updated exponent of the second multiplication result corresponding to the first combination is e3, the mantissa is m3, and the sign is s3, and the updated exponent of the second multiplication result corresponding to the second combination is e4, the mantissa is m4, and the sign is s4. The subtractor calculates the difference between e3 and e4 and inputs this difference to the shifter. The shifter uses this difference to shift m3 so that the decimal points are aligned, obtaining a mantissa, represented as m3-1. This mantissa is input to the adder, which adds m3-1 to m4 to obtain the mantissa of the third result, which is then input to the normalizer. The sign operator calculates s3 and s4 to obtain the sign of the third result, which is then input to the normalizer. The normalizer normalizes the mantissa and exponent of the third result to obtain the final third result.

[0091] It should be noted that for each second multiplication result, if the updated exponent exceeds the second value, the updated exponent is updated to the second value, where the second value is the maximum value of the exponent range corresponding to the first floating-point number minus one. If the updated exponent is less than the third value, the updated exponent is updated to the third value, where the third value is the minimum value of the exponent range corresponding to the first floating-point number plus one. Alternatively, for each second multiplication result, if the updated exponent portion exceeds the second value, the second multiplication result is recorded as infinity. If the updated exponent portion is less than the third value, the second multiplication result is recorded as 0.

[0092] Optionally, in the two aforementioned compensation methods using the exponential auxiliary information, the process of the exponent processor determining the first compensation exponent is as follows:

[0093] The exponent processor obtains a first value, where the first value is the distance between a maximum value and a minimum value of an exponent range corresponding to the second floating-point number. For example, if the second floating-point number is FP16 data, the exponent range is -15 to 16, and the first value is 31, or if the second floating-point number is FP32 data, the exponent range is -127 to 128, and the first value is 255.

[0094] For each second floating point number in each combination, the exponent processor determines the first compensation exponent corresponding to the second floating point number using formula (2). b = a* e1 sb (2)

[0095] In formula (2), b represents the first compensation exponent, a represents the first value, and e1 sb represents the exponential auxiliary information corresponding to the second floating-point number.

[0096] For example, the second floating-point number is FP16 data, the first floating-point number is FP32 data, the exponent auxiliary information is represented by e1 sb, the first value is 31, and the first compensation exponent is equal to 31*e1 sb.

[0097] Alternatively, in the correspondence between the exponent auxiliary information and the compensation exponent, a first compensation exponent corresponding to the exponent auxiliary information of the second floating-point number is obtained.

[0098] In an optional approach, when multiple combinations of products are accumulated, only the first few relatively large products contribute to the final multiplication-addition output, while the remaining products under the effective digit have a relatively low contribution to the final multiplication-addition output. Therefore, products that contribute little to the multiplication-addition output can be deleted. For example, in large-scale dense matrix multiplication calculations, the inner products of multiple long vectors can be used to complete the calculation. Referring to Figure 11, the dimensions of matrix A are [m, k], the dimensions of matrix B are [k, n], and AxB is represented as matrix C. The main calculation of AxB is the inner product of vectors with a vector length of k. When k is relatively large, there are some redundant calculations for the vector inner product. For example, when dozens of double-precision floating-point numbers are multiplied and accumulated to obtain a double-precision result, only the first few relatively large products contribute to the final vector inner product output, while the remaining products under the effective digit have a relatively low contribution to the final vector inner product output.

[0099] Referring to Figure 12, the mixed precision multiplier and adder includes a stream control unit and a multiplication and addition unit. The stream control unit receives multiple combinations, adds the exponents of the two second floating-point numbers in each combination, and obtains the exponent sum corresponding to each combination. Based on the exponent sum corresponding to each combination, the combination to be calculated among the multiple combinations is determined, and the combination to be calculated is input to the multiplication and addition unit. The combination to be calculated is a combination that contributes relatively more to the final multiplication and addition result. The multiplication and addition unit receives the combination to be calculated, and based on the exponential auxiliary information corresponding to each second floating-point number in the combination to be calculated, performs multiplication and addition operations on the combination to be calculated to obtain the final multiplication and addition result. The calculation process is described in Figures 7 to 10 above and will not be repeated here. For example, referring to Figure 13, when performing a multiplication operation on four combinations, the exponent and mantissa of the first combination are (e0, m0) and (i0, n0), the exponent and mantissa of the second combination are (e1, m1) and (i1, i1), the exponent and mantissa of the third combination are (e3, m3) and (i3, n3), and the exponent and mantissa of the fourth combination are (e4, m4) and (i4, i4). The bottom rectangle represents the comparison of the maximum exponent sum with the exponent sums corresponding to each combination. From the exponent sum comparison, it can be seen that the exponent sums corresponding to the first two combinations differ slightly. When performing the accumulation operation, the first two combinations are the combinations to be calculated, and the other two combinations contribute very little to the final multiplication and addition result and are not calculated.

[0100] In this way, when performing multiplication and addition calculations on multiple combinations, the multiplication and addition operations are performed using exponents and retaining the combinations that contribute to the final multiplication and addition results, dynamically saving unnecessary calculations while still obtaining calculation results with sufficient accuracy, which can save computing resources and reduce power consumption.

[0101] Optionally, in FIG12 , the format conversion module inputs the exponential auxiliary information to the multiplication-addition unit, and the multiplication-addition unit uses the exponential auxiliary information to perform multiplication-addition operations on the combination to be calculated. Alternatively, the format conversion module inputs the exponential auxiliary information to the flow control unit, and the flow control unit inputs the exponential auxiliary information when inputting the combination to be calculated to the multiplication-addition unit.

[0102] Optionally, the flow control unit uses the sum of the exponentials corresponding to the multiple combinations to determine the unit to be calculated as follows:

[0103] The flow control unit obtains a target threshold, which is used to determine the combinations to be calculated. The target threshold is set based on actual needs. The flow control unit subtracts the index sum corresponding to each combination from the maximum index sum corresponding to the multiple combinations to obtain multiple differences. The flow control unit determines the differences among these multiple differences that are less than the target threshold and determines the combination containing the index sum of the difference less than the target threshold as the combination to be calculated. Here, the difference can be understood as the distance between the maximum index and each of the other index sums.

[0104] Optionally, the target threshold is an integer multiple of the mantissa length of the first floating point number. For example, if the first floating point number is a double-precision floating point number, the mantissa length of the double-precision floating point number is used as the target threshold, or the length of two double-precision mantissas is used as the target threshold.

[0105] It should be noted that for a mixed-precision multiplier, the length of the mantissa stored in the memory is an integer multiple of the length of the mantissa of the floating-point number. When accumulating, the exponent needs to be aligned. When aligning the exponent, the mantissa of the floating-point number with a smaller exponent needs to be right-shifted, and the number of bits right-shifted is equal to the distance between the maximum exponent and the mantissa. The principle of right-shifting the mantissa is to add the distance number of 0s to the highest bit of the current mantissa, and the corresponding mantissa bits at the end of the original mantissa are directly discarded. The number of bits shifted is discarded (the reason for direct discarding is that the space of the memory of the storage location is fixed). In this way, when the multiplication results of multiple combinations are accumulated, if the distance between the sum of the exponents and the sum of the maximum exponents is greater, the number of bits of the mantissa right-shifted is relatively large. The more bits of right-shifted, the smaller the mantissa, which has little effect on the final accumulated result. Therefore, the combination with a smaller exponent sum has little effect on the accumulated result. When the number of bits right-shifted reaches the storage space of the memory, the mantissa is all 0, which has no effect on the final accumulated result. Therefore, the target threshold takes the mantissa length of the double-precision floating-point number as the target threshold, or takes the length of two double-precision mantissas as the target threshold, etc.

[0106] In one implementation, a mixed-precision multiplier / adder includes a stream control unit, a multiplication unit, and an addition unit. The stream control unit receives multiple combinations, adds the exponents of the two second floating-point numbers in each combination, and obtains the sum of the exponents corresponding to the multiple combinations. Based on the sum of the exponents corresponding to each combination, a combination to be calculated among the multiple combinations is determined. The stream control unit inputs the multiple combinations to the multiplication unit, and the multiplication unit uses the above-mentioned method 1 to multiply the two second floating-point numbers in each combination to obtain a first multiplication result corresponding to each combination. The multiplication unit inputs the first multiplication result corresponding to the combination to be calculated to the addition unit, and the addition unit accumulates the first multiplication results corresponding to the combination to be calculated to obtain a final multiplication-addition result. Alternatively, the multiplication unit uses the above-mentioned method 2 to multiply the two second floating-point numbers in each combination to obtain a second multiplication result corresponding to each combination, and inputs the second multiplication result corresponding to the combination to be calculated to the addition unit. The addition unit uses the above-mentioned method 2 to accumulate the second multiplication results corresponding to the combination to be calculated to obtain a final multiplication-addition result. For example, referring to FIG14, when performing an accumulation operation on four combinations and then performing the accumulation, two of the combinations are combinations to be calculated, and the other two combinations contribute very little to the final multiplication and addition result and are not calculated. FIG14 is similar to FIG13 and will not be described again here.

[0107] In this way, unnecessary accumulation calculations can be dynamically saved during the accumulation processing, while still obtaining calculation results with sufficient accuracy, thus saving computing resources.

[0108] It should be noted that, in FIG13 and FIG14 , the process of the multiplication-addition unit performing multiplication-addition operations on the combination to be calculated is described in the foregoing text and will not be repeated here.

[0109] In the embodiment of the present application, a relatively large amount of precision can be retained to meet the precision required by the scenario. In addition, the computing power of the computing center is extremely large, having reached the E level and continuing to develop towards the Z level. The total power consumption has become a bottleneck. In the embodiment of the present application, the amount of calculation is reduced while meeting the precision requirements, which can significantly reduce power consumption. For example, the power consumption of a double-precision floating-point multiplier is approximately 10:1 compared to the power consumption of a half-precision floating-point multiplier. Using a half-precision floating-point multiplier to replace a double-precision floating-point multiplier can significantly reduce the total power consumption. For example, if 20% of the double-precision multiplications are performed using half-precision multiplications, the power consumption can be effectively reduced by about 18%. Therefore, the precision can be adjusted according to the power consumption budget, such as reducing the total power consumption to 50%, but allowing a larger precision error. In this way, the above-mentioned arithmetic logic unit can be applied to any scenario with relatively high computing power requirements. In these scenarios, the power consumption of resources can be reduced. For example, scenarios with large computing power in the cloud and scenarios with large computing power at the edge can both reduce power consumption to extend the battery's standby time.

[0110] In addition, using low-precision multipliers instead of high-precision multipliers can reduce chip costs.

[0111] An embodiment of the present application also provides an arithmetic logic unit of a processor that can reduce power consumption without reducing the accuracy of calculations. For example, in computationally intensive scientific calculations (such as large-scale dense matrix multiplication), the accuracy requirements are relatively high, and double-precision floating-point numbers are usually required to meet the needs of the computing scenario. If the accuracy is directly reduced and single precision or half precision is used for calculation, it may not be possible to meet the accuracy requirements. However, in these large-scale dense matrix multiplication calculations, the inner product of multiple long vectors can be used to complete it. Only the first few relatively large products contribute to the final vector inner product output, and the remaining products under the effective number of digits contribute relatively little to the final vector inner product output. Therefore, the products that contribute little to the vector inner product output can be deleted. The corresponding processing is:

[0112] Referring to Figure 15, the arithmetic logic unit includes a stream control unit and a multiplication and addition unit. The stream control unit receives a plurality of first floating-point numbers, and the plurality of first floating-point numbers form a plurality of combinations, each of the plurality of combinations including two multiplied first floating-point numbers. For each combination, the stream control unit calculates the exponential sum of the two first floating-point numbers in the combination, determines the combination to be calculated in the plurality of combinations based on the exponential sum corresponding to each combination, and inputs the combination to be calculated into the multiplication and addition unit. The multiplication and addition unit receives the combination to be calculated, performs multiplication and addition operations on the combination to be calculated, and obtains a final multiplication and addition result. In this way, when performing multiplication and addition calculations on a plurality of combinations, the exponential sum is used to retain the combination that contributes to the final multiplication and addition result, and the multiplication and addition operation is performed, thereby dynamically saving unnecessary calculations while still obtaining a calculation result of sufficient accuracy, thereby saving computing resources and reducing power consumption.

[0113] Optionally, the process of the flow control unit determining the combination to be calculated is described in the above text and will not be repeated here.

[0114] Based on the same technical concept, embodiments of the present application provide a method for floating-point number calculation. This method can be implemented by a computing device, which can be any device that needs to perform floating-point multiplication and addition calculations. For example, the computing device can be a mobile terminal such as a mobile phone or tablet computer, a computer device such as a desktop computer or laptop computer, or a server.

[0115] FIG16 provides a method flow for floating point number calculation, see steps 1601 to 1602 .

[0116] Step 1601: Acquire multiple first floating-point numbers, convert each first floating-point number into a second floating-point number, obtain multiple second floating-point numbers, and obtain exponential auxiliary information corresponding to each second floating-point number, wherein the multiple first floating-point numbers have the same precision, the multiple second floating-point numbers have the same precision, the precision of the first floating-point number is higher than the precision of the second floating-point number, and the exponential auxiliary information corresponding to the multiple second floating-point numbers is used to compensate for the exponent range lost during the conversion.

[0117] Step 1602 : Based on the exponential auxiliary information corresponding to each second floating-point number, perform multiplication and addition operations on multiple combinations of the multiple second floating-point numbers to obtain multiplication and addition results of the multiple combinations, where each combination includes two second floating-point numbers that are subjected to multiplication operations.

[0118] In an optional manner, for each second floating-point number, if the first exponent is not within the exponent range corresponding to the second floating-point number, if the first exponent is greater than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is less than or equal to the first exponent and greater than the exponent of the second floating-point number; if the first exponent is less than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is greater than or equal to the first exponent and less than the exponent of the second floating-point number; the first exponent is the exponent of the first floating-point number converted to obtain the second floating-point number, and the compensation exponent is the exponent indicated by the exponent auxiliary information of the second floating-point number.

[0119] In an optional manner, converting each first floating-point number into a second floating-point number to obtain a plurality of second floating-point numbers, and obtaining exponential auxiliary information corresponding to each second floating-point number includes:

[0120] For each first floating-point number, determining a first exponent partition range to which the exponent of the first floating-point number belongs, and in a correspondence between exponent partition ranges and exponent auxiliary information, determining the exponent auxiliary information corresponding to the first exponent partition range as exponent auxiliary information corresponding to a second floating-point number obtained by converting the first floating-point number; the exponent partition range being obtained by partitioning the exponent range corresponding to the first floating-point number and the exponent range corresponding to the second floating-point number;

[0121] Subtract a first product from the exponent of the first floating-point number to obtain the exponent of the converted second floating-point number, where the first product is equal to the product of the exponential auxiliary information corresponding to the converted second floating-point number and a first numerical value, where the first numerical value is the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

[0122] In an optional manner, performing multiplication and addition operations on multiple combinations of the multiple second floating-point numbers based on the exponential auxiliary information corresponding to each second floating-point number to obtain multiplication and addition results of the multiple combinations includes:

[0123] performing a multiplication operation on the two second floating-point numbers in each combination based on the exponential auxiliary information corresponding to each second floating-point number, to obtain a first multiplication result corresponding to each combination;

[0124] performing a sum operation on the first multiplication results corresponding to the plurality of combinations to obtain the multiplication-addition result; or,

[0125] Performing a multiplication operation on the two second floating-point numbers in each combination to obtain a second multiplication result corresponding to each combination;

[0126] Based on the exponential auxiliary information corresponding to each second floating-point number, a summation operation is performed on the second multiplication results corresponding to the multiple combinations to obtain the multiplication-addition result.

[0127] In an optional manner, performing a multiplication operation on the two second floating-point numbers in each combination based on the exponential auxiliary information corresponding to each second floating-point number to obtain a first multiplication result corresponding to each combination includes:

[0128] For each combination, determining a first compensation exponent corresponding to each second floating-point number in the combination based on the exponent auxiliary information corresponding to each second floating-point number in the combination; adding the exponents of two second floating-point numbers in the combination to the corresponding two first compensation exponents to obtain an exponential portion of a first multiplication result corresponding to the combination;

[0129] performing a multiplication operation on the mantissas of the two second floating-point numbers in the combination to obtain a mantissa portion of a first multiplication result corresponding to the combination;

[0130] An exclusive OR operation is performed on the signs of the two second floating-point numbers in the combination to obtain a sign portion of the first multiplication result corresponding to the combination.

[0131] In an optional manner, performing a summation operation on the second multiplication results corresponding to the multiple combinations based on the exponent auxiliary information corresponding to each second floating-point number to obtain the multiplication-addition result includes:

[0132] For each combination, determining a first compensation exponent corresponding to each second floating-point number in the combination based on the exponent auxiliary information corresponding to each second floating-point number in the combination; adding the first compensation exponent corresponding to each second floating-point number in the combination to the exponent of the corresponding second multiplication result to obtain an updated second multiplication result corresponding to the combination;

[0133] A sum operation is performed on the updated second multiplication results corresponding to the multiple combinations to obtain the multiplication-addition result.

[0134] In an optional manner, for each combination, determining, based on the exponent auxiliary information corresponding to each second floating-point number in the combination, the first compensation exponent corresponding to each second floating-point number in the combination includes:

[0135] For each second floating-point number in each combination, the product of the exponential auxiliary information corresponding to the second floating-point number and a first numerical value is determined as a first compensation exponent corresponding to the second floating-point number, where the first numerical value is the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

[0136] In an optional manner, performing multiplication and addition operations on multiple combinations of the multiple second floating-point numbers based on the exponential auxiliary information corresponding to each second floating-point number to obtain multiplication and addition results of the multiple combinations includes:

[0137] Add the exponents of the two second floating-point numbers in each combination to obtain the exponential sum corresponding to each combination;

[0138] Determining a combination to be calculated among the multiple combinations based on the sum of the indices corresponding to each combination;

[0139] Based on the exponential auxiliary information corresponding to each second floating-point number in the combination to be calculated, a multiplication and addition operation is performed on the combination to be calculated to obtain the multiplication and addition result.

[0140] In an optional manner, determining the combination to be calculated among the multiple combinations based on the sum of the exponents corresponding to each combination includes:

[0141] subtracting the sum of exponents corresponding to each combination from the maximum sum of exponents corresponding to the multiple combinations to obtain a plurality of differences;

[0142] The indices of the plurality of differences whose difference values ​​are less than or equal to the target threshold and the combination to which they belong are determined as the combination to be calculated.

[0143] In an optional manner, the target threshold is a multiple of the mantissa length of the first floating-point number.

[0144] It should be noted that the method description of floating-point calculation is the same as the floating-point calculation process in the previous article, except that it is executed by a computing device. The specific processing steps can be found in the description in the previous article and will not be repeated here.

[0145] Based on the same technical concept, the present invention also provides another method for floating-point calculation. This method can be implemented by a computing device, which can be any device that needs to perform floating-point multiplication and addition calculations. Figure 17 provides a flow chart of the floating-point calculation method, see steps 1701 to 1704.

[0146] Step 1701: Acquire multiple combinations of multiple first floating-point numbers, each combination including two first floating-point numbers to be multiplied.

[0147] Step 1702: Add the exponents of the two first floating-point numbers in each combination to obtain a sum of the exponents corresponding to the multiple combinations.

[0148] Step 1703: Determine a combination to be calculated among the multiple combinations based on the sum of the exponents corresponding to each combination.

[0149] Step 1704: perform multiplication and addition operations on the combinations to be calculated to obtain multiplication and addition results corresponding to the multiple combinations.

[0150] In an optional manner, determining the combination to be calculated among the multiple combinations based on the sum of the exponents corresponding to each combination includes:

[0151] Subtract the sum of the exponents corresponding to each combination from the maximum sum of the exponents corresponding to the multiple combinations to obtain multiple differences;

[0152] The indexes of the multiple differences whose difference values ​​are less than or equal to the target threshold and the combination to which they belong are determined as the combination to be calculated.

[0153] In an optional manner, the target threshold is a multiple of the mantissa length of the first floating-point number.

[0154] This application also provides a computing device 100. Referring to FIG. 18 , computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. Processor 104, memory 106, and communication interface 108 communicate with each other via bus 102. Computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors or memories in computing device 100.

[0155] Bus 102 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG18 shows only one line, but this does not imply that there is only one bus or only one type of bus. Bus 104 may include a path for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, and communication interface 108).

[0156] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0157] The memory 106 may include volatile memory, such as random access memory (RAM). The memory 106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0158] The memory 106 stores executable program codes, and the processor 104 executes the executable program codes to implement the floating-point number calculation method. In other words, the memory 106 stores instructions for executing the floating-point number calculation method.

[0159] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0160] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform a floating-point calculation method, or instructs a computing device to perform a floating-point calculation method.

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. An arithmetic logic unit of a processor, characterized in that: Includes format conversion module and mixed precision multiplier-accumulator; The format conversion module is configured to receive a plurality of first floating-point numbers, convert each of the first floating-point numbers into a second floating-point number to obtain a plurality of second floating-point numbers, and obtain exponent auxiliary information corresponding to each second floating-point number, wherein the plurality of first floating-point numbers have the same precision, the plurality of second floating-point numbers have the same precision, the precision of the first floating-point numbers is higher than the precision of the second floating-point numbers, and the exponent auxiliary information corresponding to each second floating-point number is used to compensate for an exponent range lost during conversion; Inputting a plurality of combinations of the plurality of second floating-point numbers and exponential auxiliary information corresponding to each second floating-point number into the mixed-precision multiplier-adder, wherein each combination includes two second floating-point numbers to be multiplied; The mixed-precision multiplier-adder is used to perform multiplication-addition operations on the multiple combinations based on the exponential auxiliary information corresponding to each second floating-point number to obtain multiplication-addition results of the multiple combinations.

2. The arithmetic logic unit according to claim 1, wherein: For each of the second floating-point numbers, if the first exponent is not within the exponent range corresponding to the second floating-point number, if the first exponent is greater than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is less than or equal to the first exponent and greater than the exponent of the second floating-point number; if the first exponent is less than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is greater than or equal to the first exponent and less than the exponent of the second floating-point number, the first exponent is the exponent of the first floating-point number converted to obtain the second floating-point number, and the compensation exponent is the exponent indicated by the exponent auxiliary information of the second floating-point number.

3. The arithmetic logic unit according to claim 1 or 2, characterized in that: The format conversion module is configured to determine, for each first floating-point number, a first exponent division range to which the exponent of the first floating-point number belongs, and, in a correspondence between the exponent division range and the exponent auxiliary information, determine the exponent auxiliary information corresponding to the first exponent division range as the exponent auxiliary information corresponding to the second floating-point number obtained by converting the first floating-point number, the exponent division range being obtained by dividing the exponent range corresponding to the first floating-point number and the exponent range corresponding to the second floating-point number; Subtract a first product from the exponent of the first floating-point number to obtain the exponent of the converted second floating-point number, where the first product is equal to the product of the exponential auxiliary information corresponding to the converted second floating-point number and a first numerical value, where the first numerical value is the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

4. The arithmetic logic unit according to any one of claims 1 to 3, characterized in that: The mixed-precision multiplier-adder includes a multiplication unit and an addition unit; The multiplication unit is configured to receive the multiple combinations, and perform a multiplication operation on the two second floating-point numbers in each combination based on the exponential auxiliary information corresponding to each second floating-point number in each combination to obtain a first multiplication result corresponding to each combination; The adding unit is configured to perform a sum operation on the first multiplication results corresponding to the plurality of combinations to obtain the multiplication-addition result; or The multiplication unit is configured to receive the multiple combinations, perform a multiplication operation on the two second floating-point numbers in each combination, and obtain a second multiplication result corresponding to each combination; The adding unit is configured to perform a sum operation on the second multiplication results corresponding to the plurality of combinations based on the exponential auxiliary information corresponding to each second floating-point number in each combination to obtain the multiplication-addition result.

5. The arithmetic logic unit according to claim 4, characterized in that The multiplication unit includes an exponential processor, a multiplier and an XOR operator; The exponent processor is configured to, for each combination, determine a first compensation exponent corresponding to each second floating-point number in the combination based on exponential auxiliary information corresponding to each second floating-point number in the combination; and add the exponents of two second floating-point numbers in the combination to the corresponding two first compensation exponents to obtain an exponential portion of a first multiplication result corresponding to the combination; The multiplier is configured to, for each combination, perform a multiplication operation on the mantissas of the two second floating-point numbers in the combination to obtain a mantissa portion of the first multiplication result corresponding to the combination; The XOR operator is configured to, for each combination, perform an XOR operation on the signs of the two second floating-point numbers in the combination to obtain a sign portion of the first multiplication result corresponding to the combination.

6. The arithmetic logic unit according to claim 4, characterized in that: The adding unit includes an exponential processor and an adding subunit; The exponent processor is configured to, for each combination, determine, based on the exponential auxiliary information corresponding to each second floating-point number in the combination, a first compensation exponent corresponding to each second floating-point number in the combination; and add the first compensation exponent corresponding to each second floating-point number in the combination to the exponent of the corresponding second multiplication result to obtain an updated second multiplication result corresponding to the combination; The adding subunit is used to perform a sum operation on the updated second multiplication results corresponding to the multiple combinations to obtain the multiplication-addition result.

7. The arithmetic logic unit according to claim 5 or 6, characterized in that: The exponent processor is used to, for each second floating-point number in each combination, determine a first compensation exponent corresponding to the second floating-point number by multiplying the exponential auxiliary information corresponding to the second floating-point number by a first numerical value, where the first numerical value is the distance between a maximum value and a minimum value of an exponent range corresponding to the second floating-point number.

8. The arithmetic logic unit according to any one of claims 1 to 7, characterized in that: The mixed precision multiplier-adder includes a stream control unit and a multiply-add unit; The flow control unit is configured to receive the multiple combinations, add the exponents of the two second floating-point numbers in each combination, and obtain a sum of the exponents corresponding to each combination; determining a combination to be calculated among the plurality of combinations based on the sum of the exponents corresponding to each combination, and inputting the combination to be calculated into the multiplication and addition unit; The multiplication-addition unit is configured to perform a multiplication-addition operation on the combination to be calculated based on the exponential auxiliary information corresponding to each second floating-point number in the combination to be calculated, to obtain the multiplication-addition result.

9. The arithmetic logic unit according to claim 8, characterized in that: The flow control unit is configured to subtract the sum of exponents corresponding to each combination from the maximum sum of exponents corresponding to the multiple combinations to obtain a plurality of difference values; The indices of the plurality of differences whose difference values ​​are less than or equal to the target threshold and the combination to which they belong are determined as the combination to be calculated.

10. The arithmetic logic unit according to claim 9, characterized in that: The target threshold is a multiple of the mantissa length of the first floating-point number.

11. A method for floating point number calculation, characterized in that: The method comprises: Acquire multiple first floating-point numbers, convert each of the first floating-point numbers into a second floating-point number to obtain multiple second floating-point numbers, and obtain exponent auxiliary information corresponding to each second floating-point number, wherein the multiple first floating-point numbers have the same precision, the multiple second floating-point numbers have the same precision, the precision of the first floating-point numbers is higher than the precision of the second floating-point numbers, and the exponent auxiliary information corresponding to each second floating-point number is used to compensate for an exponent range lost during the conversion; Based on the exponential auxiliary information corresponding to each second floating-point number, multiplication and addition operations are performed on multiple combinations of the multiple second floating-point numbers to obtain multiplication and addition results of the multiple combinations, each combination including two second floating-point numbers subjected to multiplication operations.

12. The method according to claim 11, characterized in that For each of the second floating-point numbers, if the first exponent is not within the exponent range corresponding to the second floating-point number, if the first exponent is greater than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is less than or equal to the first exponent and greater than the exponent of the second floating-point number; if the first exponent is less than 0, then the sum of the exponent of the second floating-point number and the compensation exponent is greater than or equal to the first exponent and less than the exponent of the second floating-point number, the first exponent is the exponent of the first floating-point number converted to obtain the second floating-point number, and the compensation exponent is the exponent indicated by the exponent auxiliary information of the second floating-point number.

13. The method according to claim 11 or 12, characterized in that The converting each first floating-point number into a second floating-point number to obtain a plurality of second floating-point numbers, and obtaining exponential auxiliary information corresponding to each second floating-point number, includes: For each first floating-point number, determining a first exponent partition range to which the exponent of the first floating-point number belongs, and in a correspondence between exponent partition ranges and exponent auxiliary information, determining the exponent auxiliary information corresponding to the first exponent partition range as exponent auxiliary information corresponding to a second floating-point number obtained by converting the first floating-point number; the exponent partition range being obtained by partitioning the exponent range corresponding to the first floating-point number and the exponent range corresponding to the second floating-point number; Subtract a first product from the exponent of the first floating-point number to obtain the exponent of the converted second floating-point number, where the first product is equal to the product of the exponential auxiliary information corresponding to the converted second floating-point number and a first numerical value, where the first numerical value is the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

14. The method according to any one of claims 11 to 13, characterized in that The performing multiplication and addition operations on a plurality of combinations of the plurality of second floating-point numbers based on the exponential auxiliary information corresponding to each second floating-point number to obtain multiplication and addition results of the plurality of combinations includes: performing a multiplication operation on the two second floating-point numbers in each combination based on the exponential auxiliary information corresponding to each second floating-point number, to obtain a first multiplication result corresponding to each combination; performing a sum operation on the first multiplication results corresponding to the plurality of combinations to obtain the multiplication-addition result; or, Performing a multiplication operation on the two second floating-point numbers in each combination to obtain a second multiplication result corresponding to each combination; Based on the exponential auxiliary information corresponding to each second floating-point number, a summation operation is performed on the second multiplication results corresponding to the multiple combinations to obtain the multiplication-addition result.

15. The method according to claim 14, characterized in that The multiplication operation is performed on the two second floating-point numbers in each combination based on the exponential auxiliary information corresponding to each second floating-point number to obtain a first multiplication result corresponding to each combination, including: For each combination, determining a first compensation exponent corresponding to each second floating-point number in the combination based on the exponent auxiliary information corresponding to each second floating-point number in the combination; adding the exponents of two second floating-point numbers in the combination to the corresponding two first compensation exponents to obtain an exponential portion of a first multiplication result corresponding to the combination; performing a multiplication operation on the mantissas of the two second floating-point numbers in the combination to obtain a mantissa portion of a first multiplication result corresponding to the combination; An exclusive OR operation is performed on the signs of the two second floating-point numbers in the combination to obtain a sign portion of the first multiplication result corresponding to the combination.

16. The method according to claim 14, characterized in that The step of performing a summation operation on the second multiplication results corresponding to the plurality of combinations based on the exponential auxiliary information corresponding to each second floating-point number to obtain the multiplication-addition result includes: For each combination, determining a first compensation exponent corresponding to each second floating-point number in the combination based on the exponent auxiliary information corresponding to each second floating-point number in the combination; adding the first compensation exponent corresponding to each second floating-point number in the combination to the exponent of the corresponding second multiplication result to obtain an updated second multiplication result corresponding to the combination; A sum operation is performed on the updated second multiplication results corresponding to the multiple combinations to obtain the multiplication-addition result.

17. The method according to claim 15 or 16, characterized in that The step of determining, for each combination, a first compensation exponent corresponding to each second floating-point number in the combination based on exponent auxiliary information corresponding to each second floating-point number in the combination includes: For each second floating-point number in each combination, the product of the exponential auxiliary information corresponding to the second floating-point number and a first numerical value is determined as a first compensation exponent corresponding to the second floating-point number, where the first numerical value is the distance between the maximum value and the minimum value of the exponent range corresponding to the second floating-point number.

18. The method according to any one of claims 11 to 17, characterized in that The performing multiplication and addition operations on a plurality of combinations of the plurality of second floating-point numbers based on the exponential auxiliary information corresponding to each second floating-point number to obtain multiplication and addition results of the plurality of combinations includes: Add the exponents of the two second floating-point numbers in each combination to obtain the exponential sum corresponding to each combination; Determining a combination to be calculated among the multiple combinations based on the sum of the indices corresponding to each combination; Based on the exponential auxiliary information corresponding to each second floating-point number in the combination to be calculated, a multiplication and addition operation is performed on the combination to be calculated to obtain the multiplication and addition result.

19. The method according to claim 18, characterized in that The determining, based on the sum of the indices corresponding to each combination, a combination to be calculated among the multiple combinations includes: subtracting the sum of exponents corresponding to each combination from the maximum sum of exponents corresponding to the multiple combinations to obtain a plurality of differences; The indices of the plurality of differences whose difference values ​​are less than or equal to the target threshold and the combination to which they belong are determined as the combination to be calculated.

20. The method according to claim 19, characterized in that The target threshold is a multiple of the mantissa length of the first floating-point number.

21. A computing device, characterized in that The computing device includes a processor and a memory, wherein the memory stores at least one computer instruction, and the computer instruction is loaded and executed by the processor to implement the operation performed by the floating-point number calculation method according to any one of claims 11 to 20.

22. A processor, characterized in that: The processor comprises an arithmetic logic unit as claimed in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Arithmetic logic unit and floating-point number multiplication calculation method and device

    CN113138750A

  • Arithmetic logic unit, floating-point number processing method, GPU chip and electronic equipment

    CN114461176A

  • Floating point multiply-add unit with fusion precision conversion function and application method of floating point multiply-add unit

    CN115390790A