Floating-point number processing method and device, equipment, storage medium and program product

Through loop instructions and format conversion, FP32 floating-point numbers are converted to TF32 floating-point numbers, solving the problem of low FP32 computing efficiency, realizing efficient floating-point multiplication operations, improving performance and reducing hardware and storage resource requirements.

CN120653220AActive Publication Date: 2025-09-16MOORE THREADS TECHNOLOGY (SHANGHAI) CO LTD

Patent Information

Application Number
CN202511114623.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-16
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

When existing technologies implement FP32 floating-point operations at the hardware level, they face complex calculation logic, low efficiency, and major changes to existing matrix calculation units, making it difficult to perform floating-point operations efficiently on hardware.

Method used

The loop indicator instructs the tensor computing engine on the processing mode and number of loops for performing format conversion, converts FP32 floating-point numbers into simplified TF32 floating-point numbers, and uses high-bit truncation and low-bit rounding to perform format conversion to implement floating-point multiplication operations.

Benefits of technology

Without changing the structure of the tensor computing engine, a 2.66-fold performance improvement was achieved, which improved accuracy and reduced software development complexity and storage resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653220A_ABST
    Figure CN120653220A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a floating-point number processing method and device, equipment, a storage medium and a program product. The floating-point number processing method comprises the steps that a first floating-point number and a second floating-point number which need to be processed currently are obtained; the data formats of the first floating-point number and the second floating-point number are first data formats; performing format conversion on the first floating-point number and the second floating-point number based on the loop indication, and determining a first operation result after the first floating-point number and the second floating-point number are multiplied based on the converted first floating-point number and the converted second floating-point number; the loop indication is used for indicating the tensor calculation engine to execute a processing mode of format conversion and the number of loops; the converted first floating-point number and the converted second floating-point number are in a second data format, and the second data format is the simplification of the first data format. Therefore, the structure of the tensor calculation engine does not need to be changed, and 2.66 times of performance improvement can be obtained only by fewer hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of computer technology, and in particular to a floating-point number processing method, apparatus, device, storage medium, and program product. Background Art

[0002] Currently, most artificial intelligence (AI) models are trained using single-precision floating point (FP32) data format. However, FP32 has a large number of digits, which makes its floating point operations challenging to implement at the hardware level.

[0003] Some chips designed specifically for artificial intelligence applications have integrated efficient matrix calculation units. In these chips, if FP32 floating-point operations are to be implemented through hardware, more computing logic needs to be introduced, which requires significant changes to the existing matrix calculation units. In addition, the actual operation of the chip is less efficient. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure at least provide a floating-point number processing method, apparatus, device, storage medium, and program product.

[0005] The technical solution of the embodiment of the present disclosure is implemented as follows: In one aspect, an embodiment of the present disclosure provides a floating-point number processing method, which is applied to a tensor computing engine. The floating-point number processing method includes: Obtaining a first floating-point number and a second floating-point number to be processed currently; the data formats of the first floating-point number and the second floating-point number are the first data format; Based on the loop indication, the first floating-point number and the second floating-point number are format converted respectively, and based on the converted first floating-point number and the converted second floating-point number, a first operation result after multiplying the first floating-point number and the second floating-point number is determined; the loop indication is used to indicate the processing method and the number of loops for the tensor computing engine to perform the format conversion; the converted first floating-point number and the converted second floating-point number are in a second data format, and the second data format is a simplification of the first data format.

[0006] On the other hand, an embodiment of the present disclosure provides a tensor computing engine, the tensor computing engine comprising: a format conversion unit, configured to perform format conversion on the first floating-point number and the second floating-point number respectively based on a loop indication; the loop indication is used to indicate a processing method and a number of loops for the tensor computing engine to perform the format conversion; the data format of the first floating-point number and the second floating-point number is the first data format; The dot product calculation unit is used to determine a first operation result after multiplying the first floating-point number and the second floating-point number based on the converted first floating-point number and the converted second floating-point number; the converted first floating-point number and the converted second floating-point number are in a second data format, and the second data format is a simplification of the first data format.

[0007] In the disclosed embodiment, since the loop indication is used to indicate the processing method and number of loops for the tensor computing engine to perform format conversion, based on the loop indication, it is possible to quickly determine how many loops are required for the floating-point multiplication operation, and which processing method is used in each loop to perform format conversion on the first floating-point number and the second floating-point number, respectively. Since the data format of the first floating-point number and the second floating-point number is the first data format, the converted first floating-point number and the converted second floating-point number are the second data format, and the second data format is a simplification of the first data format; therefore, based on the loop indication, the first floating-point number and the second floating-point number are format converted, respectively, and the simplified first floating-point number (the converted first floating-point number) and the simplified second floating-point number (the converted second floating-point number) corresponding to different loops can be obtained. Thus, based on the converted first floating-point number and the converted second floating-point number, the first operation result after the multiplication of the first floating-point number and the second floating-point number is determined. In this way, it is only necessary to add a format conversion processing method in the tensor computing engine, and different processing methods can be called in different loops through loop instructions to perform format conversion on the first floating-point number and the second floating-point number respectively (simplification of the data format), and based on the converted first floating-point number and the converted second floating-point number, the multiplication result of the first floating-point number and the second floating-point number is determined to achieve higher-precision floating-point multiplication operations. There is no need to change the structure of the tensor computing engine, and only fewer hardware resources are needed to achieve a 2.66-fold performance improvement. Compared with the solution of directly truncating the FP32 mantissa, it has higher precision, reduces the complexity of software development compared to the software implementation solution, and reduces the demand for storage resources.

[0008] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0010] Figure 1 A schematic diagram of a floating-point data format provided in an embodiment of the present disclosure; Figure 2 A schematic diagram of the implementation process of a floating point number processing method provided in the embodiment of the present disclosure Figure 1 ; Figure 3 A schematic diagram of a floating-point number storage implementation in a floating-point number processing method provided in an embodiment of the present disclosure; Figure 4 A schematic diagram of the implementation process of a floating point number processing method provided in the embodiment of the present disclosure Figure 2 ; Figure 5 A schematic diagram of the implementation process of a floating point number processing method provided in the embodiment of the present disclosure Figure 3 ; Figure 6 A schematic diagram of the composition structure of a tensor storage engine in a floating-point number processing method provided in an embodiment of the present disclosure; Figure 7 A schematic diagram of the structure of a tensor calculation engine in a floating-point number processing method provided in an embodiment of the present disclosure; Figure 8 A schematic diagram of the calculation flow of DOT in a floating-point number processing method provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0011] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the technical solutions of the present disclosure are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting the present disclosure. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0012] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0013] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure pertains. The terms used herein are for the purpose of describing the present disclosure only and are not intended to limit the present disclosure.

[0015] Before further describing the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are first described. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations.

[0016] like Figure 1 As shown in Figure 1, half-precision floating point (FP16) is a 16-bit floating point format that complies with the IEEE 754 standard. FP16 is designed to provide lower precision than single-precision floating point (FP32) while reducing memory usage and increasing computation speed. The FP16 data format consists of 1 sign bit, 5 exponent bits, and 10 mantissa bits.

[0017] Brain Floating Point 16 (BF16) is a 16-bit floating-point format that has the same numerical range as FP32 but slightly lower precision. BF16 is particularly well-suited for deep learning, as deep learning models are typically more sensitive to numerical range and can have relatively lower precision requirements. The BF16 data format includes 1 sign bit, 8 exponent bits, and 7 mantissa bits.

[0018] Tensor Float 32 (TF32) is a computing format for Tensor Cores introduced on Ampere architecture GPUs. In certain tasks, TF32 can provide up to 10 times the performance of FP32. TF32 is a truncated Float32 data format that truncates the 23 mantissa bits in FP32 to 10 bits, while the exponent bits remain at 8 bits, for a total length of 19 bits (including a sign bit). This way, TF32 maintains the same precision as FP16 (10-bit mantissa bits) while also preserving the dynamic range of FP32 (8-bit exponent bits).

[0019] FP32 is a floating-point number representation standard widely used in computer science and follows the IEEE 754 standard. The FP32 data format includes: 1 sign bit, 8 exponent bits, and 23 mantissa bits.

[0020] In order to better understand the floating-point number processing method provided by the embodiments of the present disclosure, the solutions in the related art are first described below.

[0021] Under the constraints of chip area and power consumption, when using the neural network processing unit (NPU) and graphics processing unit (GPU) integrated on the chip to process artificial intelligence models, in order to support higher computing power, most NPUs and GPUs usually use matrix calculation units with lower precision than FP32, such as FP16, BF16, INT8 and FP8 data formats.

[0022] If the AI ​​model is run in FP16 format, since FP16 and FP32 have certain similarities in data format, in most cases, the data format conversion can be performed directly; alternatively, if the lower-precision INT8 format is selected, the model needs to be quantized. Quantization mainly includes two methods: post-training quantization (PTQ) and quantization-aware training (QAT).

[0023] Whether directly converting data formats or using quantization, these operations add additional complexity and workload.

[0024] TF32 uses the same 10-bit mantissa as FP16 and has the same 8-bit exponent as FP32, making it a viable alternative to FP32 in many scenarios. One approach is to store TF32 in FP32 format and then truncate the trailing 13 bits during calculations. This truncation method is not suitable for certain scenarios requiring higher precision. Another approach is to convert FP32 to TF32 using a software solution. However, converting FP32 at the software level requires not only complex instruction scheduling but also additional space to store the converted data. Furthermore, format conversion at runtime also imposes an additional burden, making it difficult for the final performance improvement to reach the theoretical value. If the hardware directly implements the FP32 matrix calculation unit, due to the larger mantissa of FP32, more calculation logic needs to be introduced, and the existing matrix calculation unit will require significant changes.

[0025] To this end, the present disclosure provides a floating point number processing method, which is applied to a tensor computing engine. Figure 2 As shown, the method includes the following steps 201 to 202: Step 201: Obtain a first floating-point number and a second floating-point number to be processed currently; the data format of the first floating-point number and the second floating-point number is a first data format.

[0026] The first data format may be an FP32 data format. The first floating-point number and the second floating-point number are at least two FP32 floating-point numbers to be processed. The Tensor Compute Engine (TCE) is used to perform floating-point multiplication operations.

[0027] In some implementations, step 201 may be specifically implemented by reading the first floating-point number and the second floating-point number from a local memory based on metadata information of the tensor and information about the data block.

[0028] Step 202: format-convert the first floating-point number and the second floating-point number based on a loop indication, and determine a first operation result after multiplying the first floating-point number and the second floating-point number based on the converted first floating-point number and the converted second floating-point number; the loop indication is used to indicate the processing method and the number of loops for the tensor computing engine to perform format conversion; the converted first floating-point number and the converted second floating-point number are in a second data format, which is a simplification of the first data format.

[0029] The second data format is a simplification of the first data format, which may mean that the second data format has fewer mantissa bits than the first data format. The second data format may be a TF32 data format. The loop indicator is used to reflect how many loops are required to process the current floating-point number, and whether the first floating-point number and the second floating-point number are converted to the first type of TF32 floating-point number or the second type of TF32 floating-point number in each loop. The first type may include high-order truncated data, obtained by truncating the high order of data in the first data format, such as big TF32. The second type may include low-order rounded data, which may be obtained by rounding the remaining low-order portion of the first type of data using a rounding algorithm (i.e., rounding to the nearest even number) after obtaining the first type of data, such as small TF32.

[0030] For example, to convert an FP32 floating-point number into two TF32 floating-point numbers, the following multiplication formula can be used: A * B = (big_TF32_A + small_TF32_A) * (big_TF32_B + small_TF32_B) = big_TF32_A * big_TF32_B + big_TF32_A * small_TF32_B + small_TF32_A * big_TF32_B + small_TF32_A * small_TF32_B. A represents the first floating-point number, and B represents the second floating-point number. big_TF32_A represents the first sub-floating-point number (i.e., the first floating-point number of the first type after conversion), small_TF32_A represents the fourth sub-floating-point number (i.e., the first floating-point number of the second type after conversion); big_TF32_B represents the second sub-floating-point number (i.e., the second floating-point number of the first type after conversion), and small_TF32_B represents the third sub-floating-point number (i.e., the second floating-point number of the second type after conversion).

[0031] The format conversion processing method includes: converting an FP32 floating-point number into high-order truncated data (for example, a big_TF32 floating-point number) and converting an FP32 floating-point number into low-order rounded data (for example, a small_TF32_A floating-point number).

[0032] For example, the conversion of an FP32 floating-point number to a big_TF32 floating-point number can be expressed as: big_TF32 = FP32&0xffffe000; the conversion of an FP32 floating-point number to a small_TF32_A floating-point number can be expressed as: small_TF32 = RNE((FP32&0x1fff) + 0x1000). 0xffffe000 is a hexadecimal mask used to preserve specific bits during the conversion process. 0x1fff represents a specific exponent value used to adjust the exponent during the conversion to ensure that the exponent maintains the correct range and offset. 0x1000 represents the offset of the mantissa in the TF32 format, used to adjust the position of the mantissa during the conversion process to accommodate the TF32 format requirements.

[0033] In some embodiments, step 202 may be specifically implemented by performing format conversion on the first floating-point number and the second floating-point number in different loops based on a loop indication, and performing a multiplication operation on the converted first floating-point number and the converted second floating-point number in each loop. After the loop ends, the multiplication operation results of the multiple loops are added to obtain a first operation result.

[0034] In some embodiments, step 202 may be specifically implemented as follows: first, format conversion is performed on the first floating-point number and the second floating-point number, respectively; then, in different loops, different converted first floating-point numbers and different converted second floating-point numbers are directly called; and in each loop, multiplication is performed on the converted first floating-point number and the converted second floating-point number; after the loop ends, the multiplication results of the multiple loops are added to obtain a first operation result. The combination of the converted first floating-point number and the converted second floating-point number corresponding to each loop is different.

[0035] In some embodiments, the converted first floating-point number includes a first sub-floating-point number (e.g., big_TF32_A) and a fourth sub-floating-point number (e.g., small_TF32_A). The converted second floating-point number includes a second sub-floating-point number (e.g., big_TF32_B) and a third sub-floating-point number (e.g., small_TF32_B). The first operation result refers to the result of multiplying the first floating-point number by the second floating-point number. At this time, the specific implementation method of step 202 can be: based on the loop indication, converting the first floating-point number A into big_TF32_A and small_TF32_A, converting the second floating-point number B into big_TF32_B and small_TF32_B, and sequentially determining the result of multiplying big_TF32_A and big_TF32_B, the result of multiplying big_TF32_A and small_TF32_B, the result of multiplying small_TF32_A and big_TF32_B, and the result of multiplying small_TF32_A and small_TF32_B based on the loop indication; and then adding the results of the four multiplications to obtain a first operation result after the first floating-point number and the second floating-point number are multiplied.

[0036] In some embodiments, the specific implementation of step 202 can also be: first convert the first floating point number A into big_TF32_A and small_TF32_A, convert the second floating point number B into big_TF32_B and small_TF32_B, and then call big_TF32_A and big_TF32_B, big_TF32_A and small_TF32_B, small_TF32_A and big_TF32_B, small_TF32_A and big_TF32_B, big_TF32_A and big_TF32_B, big_TF32_A and big_TF32_B, and big_TF32_A and big_TF32_B in sequence based on the loop instruction. B, and sequentially determine the result of multiplying big_TF32_A and big_TF32_B, the result of multiplying big_TF32_A and small_TF32_B, the result of multiplying small_TF32_A and big_TF32_B, and the result of multiplying small_TF32_A and small_TF32_B. In this process, after obtaining the result of this multiplication, the result of this multiplication is added to the result of the previous multiplication to obtain a first operation result after the multiplication of the first floating-point number and the second floating-point number.

[0037] In the embodiment of the present disclosure, since the loop indication is used to indicate the processing method and the number of loops for the tensor computing engine to perform format conversion, based on the loop indication, it is possible to quickly determine how many loops are required for the floating-point multiplication operation, and which processing method is used in each loop to perform format conversion on the first floating-point number and the second floating-point number respectively. Since the data format of the first floating-point number and the second floating-point number is the first data format, the converted first floating-point number and the converted second floating-point number are the second data format, and the second data format is a simplification of the first data format; therefore, based on the loop indication, the first floating-point number and the second floating-point number are format converted respectively, and different data format simplifications can be performed on the first floating-point number and the second floating-point number in different loops, thereby determining the first operation result after the multiplication of the first floating-point number and the second floating-point number based on the converted first floating-point number and the converted second floating-point number. In this way, it is only necessary to add a format conversion processing method in the tensor computing engine, and different processing methods can be called in different loops through loop instructions to perform format conversion on the first floating-point number and the second floating-point number respectively (simplification of the data format), and based on the converted first floating-point number and the converted second floating-point number, the multiplication result of the first floating-point number and the second floating-point number is determined to achieve higher-precision floating-point multiplication operations. There is no need to change the structure of the tensor computing engine, and only fewer hardware resources are needed to achieve a 2.66-fold performance improvement. Compared with the solution of directly truncating the FP32 mantissa, it has higher precision, reduces the complexity of software development compared to the software implementation solution, and reduces the demand for storage resources.

[0038] In some embodiments, step 201 may be implemented by steps 2011 to 2022 as follows: Step 2011: Determine a first number of floating-point multiplication operations that the tensor computing engine can process at one time.

[0039] The first number refers to the number of floating-point multiplication operations that the tensor computing engine can process at one time.

[0040] In some implementations, step 2011 may be specifically implemented by determining the first quantity based on the structure of a dot product calculation unit in a tensor calculation engine.

[0041] In some implementations, the dot product calculation unit may include 8*8 dot product units (DP), and in this case, the first number may be 8*8.

[0042] Step 2012: Read the first number of floating-point numbers from a first address of the local memory as the first floating-point number, and read the first number of floating-point numbers from a second address of the local memory as the second floating-point number.

[0043] The data format of the first floating-point number and the second floating-point number is a first data format.

[0044] The first address refers to the address of the first floating-point number in the local memory, and the second address refers to the address of the second floating-point number in the local memory.

[0045] In some embodiments, the first number may be 8*8. In this case, 8*8 floating-point numbers may be read from the first address of the local memory as the first floating-point number, and 8*8 floating-point numbers may be read from the second address of the local memory as the second floating-point number.

[0046] In the embodiment of the present disclosure, the first floating-point number and the second floating-point number are read according to the first number of floating-point multiplication operations that the tensor computing engine can process at one time, and the multiplication operations of the first number of floating-point numbers can be processed at the same time, thereby improving the utilization rate of the tensor computing engine.

[0047] In some embodiments, the above step 202 may be implemented by the following steps 2021 to 2023: Step 2021: Determine the number of loops based on the loop indication; the number of loops is used to indicate the corresponding processing method for the format conversion.

[0048] In some implementations, the loop indication may be controlled by an instruction of the tensor computing engine, and the loop indication may be 3 loops or 4 loops.

[0049] Step 2022: Based on the number of loops, in different loops, respectively, adopt different processing methods for the format conversion to extract different parts of the first floating-point number and the second floating-point number to obtain at least two different simplified first floating-point numbers and at least two different simplified second floating-point numbers; the converted first floating-point number includes the simplified first floating-point number, and the converted second floating-point number includes the simplified second floating-point number.

[0050] Different processing methods for the format conversion include: format conversion based on a first function and format conversion based on a second function. The first function may include a format conversion function for high-order truncation, which implements format conversion by performing high-order truncation on the data in the first data format. Exemplarily, the first function may be a format conversion function for big_TF32, i.e., big_TF32=FP32&0xffffe000. The second function may include a format conversion function for low-order rounding, which implements format conversion by performing a rounding operation on the low-order part of the data in the first data format. Exemplarily, the second function may be a format conversion function for small_TF32, i.e., small_TF32=RNE((FP32&0x1fff) + 0x1000).

[0051] In some embodiments, if the number of loops is 4, then in the first loop, the first function is used to perform format conversion on the first floating-point number and the second floating-point number respectively; in the second loop, the first function is used to perform format conversion on the first floating-point number, and the second function is used to perform format conversion on the second floating-point number; in the third loop, the second function is used to perform format conversion on the first floating-point number, and the first function is used to perform format conversion on the second floating-point number; in the fourth loop, the second function is used to perform format conversion on the first floating-point number and the second floating-point number respectively; thereby obtaining a simplified first floating-point number and a simplified second floating-point number corresponding to each loop.

[0052] In some embodiments, if the number of loops is 3, in the first loop, the first function is used to perform format conversion on the first floating-point number and the second floating-point number respectively; in the second loop, the first function is used to perform format conversion on the first floating-point number, and the second function is used to perform format conversion on the second floating-point number; in the third loop, the second function is used to perform format conversion on the first floating-point number, and the first function is used to perform format conversion on the second floating-point number; thereby obtaining a simplified first floating-point number and a simplified second floating-point number corresponding to each loop.

[0053] Step 2023: Determine the first operation result based on the simplified first floating-point number and the simplified second floating-point number.

[0054] In some implementations, step 2023 may be specifically implemented by performing a calculation on the simplified first floating-point number and the simplified second floating-point number to obtain a first calculation result.

[0055] Based on the above technical solution, the number of floating-point number processing cycles can be determined based on the loop indication, and the format conversion method corresponding to each cycle can be determined based on the number of loops. Then, in different loops, different format conversion processing methods are used to extract different parts of the first floating-point number and the second floating-point number, respectively, to obtain the simplified first floating-point number and the simplified second floating-point number corresponding to each loop, and then the first operation result is obtained by performing operation processing on the simplified first floating-point number and the simplified second floating-point number. In this way, different processing methods can be called in different loops to perform format conversion on the first floating-point number and the second floating-point number respectively, simply by using the loop indication. This method of converting while looping does not require additional space to store the converted floating-point numbers, reducing the demand for storage resources.

[0056] In some embodiments, the above step 2022 may be implemented by the following steps 2022a to 2022d: Step 2022a: In the first loop, the first function is used to perform format conversion on the first floating-point number and the second floating-point number to obtain a first sub-floating-point number and a second sub-floating-point number.

[0057] The first sub-floating-point number is the floating-point number resulting from format conversion of the first floating-point number using the first function, i.e., the simplified first floating-point number obtained using the first function. The second sub-floating-point number is the floating-point number resulting from format conversion of the second floating-point number using the first function, i.e., the simplified second floating-point number obtained using the first function. The number of loops can be represented by loop id 0, so the first loop can be represented by loop id 0.

[0058] In some implementations, the first function may refer to a format conversion function of big_TF32, that is, big_TF32=FP32&0xffffe000. In this case, the first sub-floating-point number may be represented as big_TF32_A, and the second sub-floating-point number may be represented as big_TF32_B.

[0059] Step 2022b: In a second loop, the first function is used to perform format conversion on the first floating-point number to obtain a first sub-floating-point number, and the second function is used to perform format conversion on the second floating-point number to obtain a third sub-floating-point number.

[0060] The third sub-floating-point number is the floating-point number obtained by converting the second floating-point number using the second function. This is the simplified second floating-point number obtained using the second function. The number of loops can be represented by loop, so the second loop can be represented as loop id 1.

[0061] In some embodiments, the second function may refer to a format conversion function of small_TF32, that is, small_TF32=RNE((FP32&0x1fff) + 0x1000), and the third sub-floating-point number may be expressed as small_TF32_B.

[0062] Step 2022c, a third loop, performing format conversion on the first floating-point number using the second function to obtain a fourth sub-floating-point number, and performing format conversion on the second floating-point number using the first function to obtain a second sub-floating-point number.

[0063] The number of loops can be represented by loop, so the third loop can be represented by loop id 2. The fourth sub-floating-point number refers to the floating-point number obtained by format conversion of the first floating-point number using the second function, that is, the simplified first floating-point number obtained based on the second function.

[0064] In some implementations, the first function may refer to a format conversion function of big_TF32, that is, big_TF32=FP32&0xffffe000, and the fourth sub-floating-point number may be represented as small_TF32_A.

[0065] Step 2022d, fourth loop, using the second function to perform format conversion on the first floating-point number and the second floating-point number respectively to obtain a fourth sub-floating-point number and a third sub-floating-point number; the simplified first floating-point number includes the first sub-floating-point number and the fourth sub-floating-point number, and the simplified second floating-point number includes the second sub-floating-point number and the third sub-floating-point number.

[0066] The number of loops can be represented by loop, so the fourth loop can be represented as loop id 3.

[0067] Based on the above technical solution, in the first loop, the first function is used to perform format conversion on the first floating-point number and the second floating-point number respectively to obtain the first sub-floating-point number and the second sub-floating-point number; in the second loop, the first function is used to perform format conversion on the first floating-point number to obtain the first sub-floating-point number, and the second function is used to perform format conversion on the second floating-point number to obtain the third sub-floating-point number; in the third loop, the second function is used to perform format conversion on the first floating-point number to obtain the fourth sub-floating-point number, and the first function is used to perform format conversion on the second floating-point number to obtain the second sub-floating-point number; in the fourth loop, the second function is used to perform format conversion on the first floating-point number and the second floating-point number respectively to obtain the fourth sub-floating-point number and the third sub-floating-point number; in this way, different format conversion processing methods can be used to extract different parts of the first floating-point number and the second floating-point number according to the number of loops, so as to obtain the simplified first floating-point number and the simplified second floating-point number corresponding to each loop. The entire process does not require any changes to the structure of the tensor computing engine, and the floating-point numbers can be converted while looping, thereby obtaining an exponential performance improvement.

[0068] In some embodiments, the above step 2023 may be implemented by the following steps 2023a to 2023b: Step 2023a: Perform a multiplication operation on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results.

[0069] Since the number of floating-point multiplication operations that the tensor computing engine can process at one time is the first number, that is, the number of first floating-point numbers and the number of second floating-point numbers obtained at one time are the first number, and each loop requires the first number of simplified first floating-point numbers and the first number of simplified second floating-point numbers, then the number of second operation results after the multiplication operation is 1 / 2 of the first number, that is, the second number is 1 / 2 of the first number. For example, if the first number is 8, the second number is 4.

[0070] Step 2023b: perform an addition operation on the second number of second operation results to obtain the first operation result.

[0071] Based on the above technical solution, by multiplying the simplified first floating-point number and the simplified second floating-point number in each loop and adding the second number of second operation results after the multiplication, the multiplication result of the first floating-point number and the second floating-point number can be obtained.

[0072] In some embodiments, the above step 202 may also be implemented through the following steps 2024 to 2027: Step 2024: Using different processing methods for the format conversion, extract different parts of the first floating-point number and the second floating-point number to obtain at least two different simplified first floating-point numbers and at least two different simplified second floating-point numbers; the converted first floating-point number includes the simplified first floating-point number, and the converted second floating-point number includes the simplified second floating-point number.

[0073] Step 2025: Based on the number of loops, in different loops, respectively retrieve different simplified first floating-point numbers and different simplified second floating-point numbers; the number of loops is determined based on the loop indication.

[0074] In some embodiments, step 2025 may be specifically implemented as follows: in a first loop, calling the first sub-floating-point number and the second sub-floating-point number; in a second loop, calling the first sub-floating-point number and the third sub-floating-point number; in a third loop, calling the fourth sub-floating-point number and the second sub-floating-point number; in a fourth loop, calling the fourth sub-floating-point number and the third sub-floating-point number; the simplified first floating-point number includes the first sub-floating-point number and the fourth sub-floating-point number, and the simplified second floating-point number includes the second sub-floating-point number and the third sub-floating-point number.

[0075] Step 2026: Perform a multiplication operation on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results.

[0076] Step 2027: perform an addition operation on the second number of second operation results to obtain the first operation result.

[0077] Based on the above technical solution, before the loop, different format conversion processing methods are first used to extract different parts of the first floating-point number and the second floating-point number, and then different simplified first floating-point numbers and different simplified second floating-point numbers are respectively retrieved in different loops based on the number of loops. Then, the simplified first floating-point number and the simplified second floating-point number in each loop are multiplied, and the second number of second operation results after the multiplication operation are added, and the multiplication result of the first floating-point number and the second floating-point number can also be obtained. In this way, the method of converting first and then retrieving can reduce the computational complexity of each loop and improve the processing efficiency of the loop.

[0078] Based on the above embodiment, the floating point number processing method provided by the embodiment of the present disclosure further includes the following steps 203 to 204: Step 203: In response to the configuration instruction, configure metadata information of the tensor and block information of the tensor; each block of the tensor includes multiple floating-point numbers in the first data format.

[0079] Tensors are stored in memory as contiguous blocks of memory. A single element in a tensor can be a single floating-point number. The size of a tensor is directly related to the memory required, and the number of bits in the floating-point number (such as 32 or 16 bits) affects the storage efficiency of the tensor. Tensor metadata may include, but is not limited to, the tensor size (dimensions), data type, storage offset, stride, and content attributes.

[0080] In some implementations, the specific implementation of “configuring metadata information of the tensor” in step 203 may be: configuring the metadata information (TensorInfo) of the tensor using an application programming interface (API) function.

[0081] In some embodiments, configuring the tile information of the tensor in step 203 may be implemented by defining a bit field in the instructions of the tensor storage engine to represent the tile information. The tile information may include, but is not limited to, the coordinates and size of the tile.

[0082] Step 204: Based on the metadata information and the block information, transfer the data in any block of the tensor from the global memory to the local memory.

[0083] Global memory refers to a memory area that can be shared by all threads. Local memory can be a faster, smaller memory area located near the processor in a computer device. Local memory is designed to reduce latency when the processor accesses main memory, thereby increasing data access speed. Local memory can be part of a cache or a dedicated high-speed storage area, such as a register or some type of cache memory.

[0084] In some embodiments, to avoid bank conflicts when the tensor computation engine reads floating-point numbers from local memory, global memory and local memory store floating-point numbers in a target storage manner; the target storage manner can evenly distribute multiple floating-point numbers for each computation across different banks.

[0085] In a feasible implementation, the target storage mode may be a swizzle mode. Swizzle is a technology for optimizing memory access by changing the storage mode of data in memory, which involves address remapping before data is written to memory or after data is read from memory.

[0086] like Figure 3 As shown, the original data cache line (origin data line) includes 16 banks. The number in each bank represents the data stored in each bank; each bank is 16 bytes. Since the same column of different cache lines cannot be read and written simultaneously, only different columns can be read and written; therefore, during calculations, the data of a column cannot be read and written at once. Based on this, the embodiment of the present disclosure adopts a swizzle method, placing 0s in different rows in different columns, 1s in different rows in different columns, ..., 15s in different rows in different columns. In this way, there will be no bank conflict, and the data can be read and written one clock cycle at a time. A clock cycle refers to the shortest time unit required for the GPU to perform a single operation.

[0087] Based on the above technical solution, by configuring the metadata information and block information of the tensor, the data in any block of the tensor can be transferred from the global memory to the local memory based on the metadata information and block information, so as to facilitate subsequent operations on the floating-point numbers in any block.

[0088] Based on the above embodiment, the floating point number processing method provided by the embodiment of the present disclosure may further include the following steps 205 to 206: Step 205: Determine a first storage area from the free memory of the tensor computing engine based on the bit width of the first floating-point number, the data read delay, and the preset first storage bit.

[0089] The first floating-point number has a bit width of 32 bits. Data read latency refers to the time required to read the first floating-point number. The first storage bit refers to the bit width required during floating-point operations; for example, it is necessary to store not only the floating-point number itself but also the sign bit, exponent bit, mantissa bit, extension bit, and sticky bit. The specific setting can be set based on business needs.

[0090] Step 206: Determine a second storage area from the free memory of the tensor computing engine based on the bit width of the second floating-point number, the data read delay, and the preset second storage bit.

[0091] The second floating-point number has a bit width of 32 bits. Data read latency refers to the time required to read the second floating-point number. The second storage bit refers to the bit width required during floating-point calculations. The first storage bit and the second storage bit can be the same or different.

[0092] In some embodiments, after reading the first floating-point number and the second floating-point number, the first floating-point number can be stored in the first storage area, and the second floating-point number can be stored in the second storage area. It should be noted that when existing solutions calculate FP32 floating-point numbers at the software level, because the specific required space is not clear, a larger storage space is allocated when allocating memory; while the disclosed embodiment stores the original FP32 floating-point numbers in the first storage area and the second storage area. Since only the FP32 floating-point numbers currently to be processed read from the local memory need to be stored, only storage space matching the FP32 floating-point numbers needs to be opened up. Compared with existing solutions, less storage space is opened up, thereby saving storage space to the greatest extent.

[0093] In some embodiments, when performing floating-point operations by converting first and then retrieving, the converted first floating-point number (simplified first floating-point number) can be stored in the first storage area, and the converted second floating-point number (simplified second floating-point number) can be stored in the second storage area so that they can be retrieved at any time.

[0094] Based on the above scheme, based on the bit width of the first floating-point number, the data reading delay, and the preset first storage bit, the first storage area is determined from the free memory of the tensor computing engine; based on the bit width of the second floating-point number, the data reading delay, and the preset second storage bit, the second storage area is determined from the free memory of the tensor computing engine. This enables the first floating-point number and the second floating-point number to be retrieved at any time, or the simplified first floating-point number and the simplified second floating-point number to be retrieved at any time, so as to further improve the processing efficiency of floating-point operations.

[0095] It should be noted that each DP in the dot product calculation unit (DOT8 calculation unit) can process 8 floating-point multiplication operations at a time. In order to clearly understand each floating-point multiplication operation, this disclosure takes the multiplication operation of a single first floating-point number and a single second floating-point number as an example for explanation.

[0096] The present disclosure provides a floating point number processing method. Figure 4 As shown, the method includes the following steps 401 to 406: Step 401: Determine a first number of floating-point multiplication operations that a tensor computing engine can process at one time.

[0097] Step 402: Read a first number of floating-point numbers from a first address in the local memory as first floating-point numbers, and read a first number of floating-point numbers from a second address in the local memory as second floating-point numbers.

[0098] Step 403: Determine the number of loops based on the loop indication; the number of loops is used to indicate a corresponding format conversion processing method.

[0099] Step 404: Based on the number of loops, in different loops, different format conversion processing methods are used to extract different parts of the first floating-point number and the second floating-point number to obtain at least two different simplified first floating-point numbers and at least two different simplified second floating-point numbers.

[0100] The converted first floating-point number includes a simplified first floating-point number, and the converted second floating-point number includes a simplified second floating-point number.

[0101] In some embodiments, step 404 may be specifically implemented as follows: in a first loop, the first function is used to perform format conversion on the first floating-point number and the second floating-point number, respectively, to obtain a first sub-floating-point number and a second sub-floating-point number; in a second loop, the first function is used to perform format conversion on the first floating-point number to obtain a first sub-floating-point number, and the second function is used to perform format conversion on the second floating-point number to obtain a third sub-floating-point number; in a third loop, the second function is used to perform format conversion on the first floating-point number to obtain a fourth sub-floating-point number, and the first function is used to perform format conversion on the second floating-point number to obtain a second sub-floating-point number; in a fourth loop, the second function is used to perform format conversion on the first floating-point number and the second floating-point number, respectively, to obtain a fourth sub-floating-point number and a third sub-floating-point number; the simplified first floating-point number includes the first sub-floating-point number and the fourth sub-floating-point number, and the simplified second floating-point number includes the second sub-floating-point number and the third sub-floating-point number.

[0102] Step 405: Perform a multiplication operation on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results.

[0103] Step 406: Perform an addition operation on the second number of second operation results to obtain a first operation result.

[0104] The present disclosure provides a floating point number processing method. Figure 5 As shown, the method includes the following steps 501 to 506: Step 501: Determine a first number of floating-point multiplication operations that a tensor computing engine can process at one time.

[0105] Step 502: Read a first number of floating-point numbers from a first address in the local memory as first floating-point numbers, and read a first number of floating-point numbers from a second address in the local memory as second floating-point numbers.

[0106] Step 503 : adopt different format conversion processing modes to extract different parts of the first floating-point number and the second floating-point number to obtain at least two different simplified first floating-point numbers and at least two different simplified second floating-point numbers.

[0107] Step 504: Based on the number of loops, different simplified first floating-point numbers and different simplified second floating-point numbers are respectively retrieved in different loops.

[0108] In some embodiments, step 504 may be specifically implemented as follows: in a first loop, the first sub-floating-point number and the second sub-floating-point number are retrieved; in a second loop, the first sub-floating-point number and the third sub-floating-point number are retrieved; in a third loop, the fourth sub-floating-point number and the second sub-floating-point number are retrieved; in a fourth loop, the fourth sub-floating-point number and the third sub-floating-point number are retrieved; the simplified first floating-point number includes the first sub-floating-point number and the fourth sub-floating-point number, and the simplified second floating-point number includes the second sub-floating-point number and the third sub-floating-point number.

[0109] Step 505: Perform a multiplication operation on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results.

[0110] Step 506: perform an addition operation on the second number of second operation results to obtain a first operation result.

[0111] An embodiment of the present disclosure provides a tensor computing engine, which includes: a format conversion unit, configured to perform format conversion on a first floating-point number and a second floating-point number respectively based on a loop indication; the loop indication is configured to indicate a processing method and a number of loops for the tensor computing engine to perform the format conversion; the data format of the first floating-point number and the second floating-point number is a first data format; a dot product calculation unit, configured to determine a first operation result after multiplying the first floating-point number and the second floating-point number based on the converted first floating-point number and the converted second floating-point number; the converted first floating-point number and the converted second floating-point number are in a second data format, which is a simplification of the first data format.

[0112] The TensorMemoryEngine (TME) is used to move tensor data from global memory to local memory.

[0113] In some embodiments, the dot product calculation unit may be a DOT8 calculation unit. The DOT8 calculation unit may include eight dot product calculation units for performing dot product operations in vector or matrix operations. For example, the eight dot products performed by the DOT8 calculation unit may be represented as: A0*B0 + A1*B1 + ... + A7*B7.

[0114] In some embodiments, a format conversion unit for format conversion of floating-point numbers can be set in at least the following three positions: 1. The format conversion unit is set on the path after reading the first floating-point number and the second floating-point number and before storing the converted first floating-point number and the second floating-point number; 2. The format conversion unit is set on the path between the first storage area and the second storage area and the DOT8 calculation unit; 3. The format conversion unit is integrated into the DOT8 calculation unit.

[0115] In some embodiments, the format conversion unit is specifically configured to perform format conversion on the first floating-point number using a first function to obtain a first sub-floating-point number; perform format conversion on the second floating-point number using the first function to obtain a second sub-floating-point number; perform format conversion on the first floating-point number using a second function to obtain a fourth sub-floating-point number; and perform format conversion on the second floating-point number using a second function to obtain a third sub-floating-point number.

[0116] In some embodiments, the dot product calculation unit is specifically used to determine the number of loops based on the loop indication; the number of loops is used to indicate the corresponding format conversion processing method; based on the number of loops, in different loops, different format conversion processing methods are used to extract different parts of the first floating-point number and the second floating-point number to obtain at least two different simplified first floating-point numbers and at least two different simplified second floating-point numbers; the converted first floating-point number includes the simplified first floating-point number, and the converted second floating-point number includes the simplified second floating-point number; and a first operation result is determined based on the simplified first floating-point number and the simplified second floating-point number.

[0117] In some embodiments, the dot product calculation unit is specifically used to call different simplified first floating-point numbers and different simplified second floating-point numbers in different loops based on the number of loops; perform multiplication operations on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results; and perform addition operations on the second number of second operation results to obtain a first operation result.

[0118] In some embodiments, the tensor computing engine further includes: a first storage area and a second storage area; the first storage area is used to store the first floating-point number, or to store the converted first floating-point number; the second storage area is used to store the second floating-point number, or to store the converted second floating-point number.

[0119] In some embodiments, the dot product calculation unit also includes: a multiplier, used to perform multiplication operations on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results; an exponent alignment unit, used to perform exponent alignment on the second number of second operation results; an adder, used to perform addition operations on the second number of second operation results after exponent alignment to obtain the first operation result.

[0120] The following describes the application of the floating-point number processing method provided by the embodiments of the present disclosure in actual scenarios.

[0121] This disclosed embodiment uses a hardware implementation and simple logic control, adding format conversion logic and loop control (loop instructions) to the Tensor Compute Engine (TCE) to implement FP32 computations using TF32 computation logic. This reduces software development complexity and the memory requirements of the software implementation. This solution requires only minimal hardware resources, yet achieves a 2.66x performance improvement and offers higher accuracy than direct TF32 conversion.

[0122] like Figure 6 As shown in the figure, the Tensor Memory Engine (TME) is responsible for moving tensor data from global memory to local memory. The Async Barrier synchronization mechanism is used to control the execution order of multiple asynchronous tasks, ensuring that they can be executed synchronously at a certain point. A register file is a storage device consisting of a set of registers that can be used to quickly access and store temporary data; for example, it can be used to store local variables and intermediate calculation results. Constant memory is a read-only memory shared by all threads and can be cached to improve performance. It is suitable for storing data that does not change, such as lookup tables or configuration parameters.

[0123] like Figure 7As shown, one implementation is to triple the size of the original A and B buffers. This way, the tensor compute engine stores the converted FP32 floating-point number A (big_TF32_A and small_TF32_A) in the A buffer, and stores the converted FP32 floating-point number B (big_TF32_B and small_TF32_B) in the B buffer. The tensor compute engine then loops through four iterations. The first loop executes big_TF32_A * big_TF32_B, storing the result in the C buffer. The next loop executes big_TF32_A * small_TF32_B and adds it to the previous result, calculating the remaining small_TF32_A * big_TF32_B and small_TF32_A * small_TF32_B. A control bit (loop indicator) is added to the tensor compute engine instructions to control whether small_TF32_A * small_TF32_B is required. Another implementation does not increase the size of the A / B buffer. Instead, the original data is read from the A buffer and B buffer in each loop and then the format conversion is performed.

[0124] The format conversion function is as follows: big_TF32 = FP32&0xffffe000; small_TF32 = RNE((FP32&0x1fff) + 0x1000).

[0125] The specific implementation scheme of the embodiment of the present disclosure is as follows: By configuring the metadata information (TensorInfo) of the tensor, such as the dimensions, strides, and data type of the tensor. At the same time, the data block information is configured in the tensor storage engine instructions, such as the coordinates and size of the data block. The tensor storage engine moves the tensor data from the global memory to the local memory. In order to avoid bank conflicts when the tensor computing engine reads data from the local memory, the tensor storage engine and the tensor computing engine use the same swizzle method, such as Figure 3 shown.

[0126] like Figure 7As shown, the DOT8 compute unit in the tensor compute engine has an 8x8 structure. Each time, the tensor compute engine reads at least 8x8 Matrix A elements and 8x8 Matrix B elements from local memory or registers. The sizes of the A and B buffers need to be configured based on data read latency and local memory / register bandwidth. To minimize data read latency, the A and B buffers are typically set large, but the degree of largeness is manageable.

[0127] In the first method, buffer A and buffer B store the converted TP32 data. In this case, the data format conversion unit can be set on the path after reading the original FP32 data and before storing the converted TP32 data.

[0128] In the second method, A buffer and B buffer still store the original FP32 data, and a data format conversion unit is added between A buffer and B buffer and the DOT8 calculation logic.

[0129] The third method is to integrate the format conversion unit with the DOT8 calculation unit.

[0130] The present disclosure embodiment takes the third method as an example for explanation. Figure 8 As shown in the DOT8 computation pipeline, the 0th level conv unit is used to convert raw data formats such as FP16, BF16, and TF32 into an internal computation format that uses more mantissa bits than FP32. The disclosed embodiment integrates the FP32 to big TF32 and small TF32 format conversion logic with the 0th level. The tensor computation engine can use peripheral control logic to determine whether the 0th level output is big TF32 or small TF32.

[0131] For example, the peripheral control logic adds 3 or 4 loops, which are controlled by the instructions of the tensor computing engine, and passes the loop id (0,1,2,3) to the conv unit at level 0, which outputs according to the loop id. Specifically, When loop id is 0, the conv unit converts the formats of FP32 floating-point numbers A and B respectively, and outputs big_TF32_A and big_TF32_B at the same time; In loop id 1, the conv unit converts the format of FP32 floating point number A and outputs big_TF32_A, and converts the format of FP32 floating point number B and outputs small TF32_B; In loop id 2, the conv unit converts the format of FP32 floating point number A and outputs small TF32_A, and outputs big TF32_B for FP32 floating point number B. In loop id 3, the conv unit converts the format of FP32 floating-point numbers A and B, and outputs small TF32_A and small TF32_B.

[0132] The calculation result of each loop is used as the input c of the next calculation. After the loop is completed, a final 8x8 calculation result matrix can be obtained.

[0133] The mul unit multiplies two converted TF floating-point numbers. The max.exp unit aligns the exponents of the first-level operation results. The align unit adds the second-level operation results.

[0134] The solution of the disclosed embodiment achieves higher-precision FP32 matrix calculation without changing the structure of the existing tensor computing engine (TCE). Compared with the solution of directly truncating the FP32 mantissa, it has higher precision and reduces the complexity of software development compared with the software implementation solution, while also reducing the demand for storage resources.

[0135] It should be understood that “one embodiment” or “an embodiment” mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, “in one embodiment” or “in an embodiment” appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present disclosure, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure are for description only and do not represent the advantages and disadvantages of the embodiments.

[0136] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0137] In the several embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0138] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0139] In addition, all functional units in the embodiments of the present disclosure may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0140] Those skilled in the art will understand that all or part of the steps of the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0141] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0142] The above is only an embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the present disclosure, and they should all be covered by the protection scope of the present disclosure.

Claims

1. A floating point number processing method, characterized in that: Applied to the tensor computing engine, the floating-point number processing method includes: Obtaining a first floating-point number and a second floating-point number to be currently processed; the data formats of the first floating-point number and the second floating-point number are a first data format; Based on the loop indication, the first floating-point number and the second floating-point number are format converted respectively, and based on the converted first floating-point number and the converted second floating-point number, a first operation result after multiplying the first floating-point number and the second floating-point number is determined; the loop indication is used to indicate the processing method and the number of loops for the tensor computing engine to perform format conversion; the converted first floating-point number and the converted second floating-point number are in a second data format, and the second data format is a simplification of the first data format.

2. The floating-point number processing method according to claim 1, wherein The obtaining of the first floating-point number and the second floating-point number to be processed currently includes: determining a first number of floating-point multiplication operations that the tensor computation engine can process at one time; The first number of floating-point numbers is read from a first address of a local memory as the first floating-point number, and the first number of floating-point numbers is read from a second address of the local memory as the second floating-point number.

3. The floating point number processing method according to claim 1, wherein The format conversion of the first floating-point number and the second floating-point number is performed based on the loop indication, and a first operation result after multiplication of the first floating-point number and the second floating-point number is determined based on the converted first floating-point number and the converted second floating-point number, including: Determining a number of cycles based on the cycle indication; the number of cycles is used to indicate a corresponding processing method for the format conversion; Based on the number of loops, in different loops, different processing modes of the format conversion are respectively adopted to extract different parts of the first floating-point number and the second floating-point number to obtain at least two different simplified first floating-point numbers and at least two different simplified second floating-point numbers; the converted first floating-point number includes the simplified first floating-point number, and the converted second floating-point number includes the simplified second floating-point number; The first operation result is determined based on the simplified first floating-point number and the simplified second floating-point number.

4. The floating point number processing method according to claim 3, wherein The method further comprises: extracting different parts of the first floating-point number and the second floating-point number by using different processing modes of the format conversion in different cycles based on the number of cycles to obtain at least two different simplified first floating-point numbers and at least two different simplified second floating-point numbers. In a first loop, the first function is used to perform format conversion on the first floating-point number and the second floating-point number respectively to obtain a first sub-floating-point number and a second sub-floating-point number; In a second loop, the first floating-point number is format-converted using the first function to obtain a first sub-floating-point number, and the second floating-point number is format-converted using the second function to obtain a third sub-floating-point number. In a third loop, the second function is used to perform format conversion on the first floating-point number to obtain a fourth sub-floating-point number, and the first function is used to perform format conversion on the second floating-point number to obtain a second sub-floating-point number; In a fourth loop, the second function is used to perform format conversion on the first floating-point number and the second floating-point number, respectively, to obtain a fourth sub-floating-point number and a third sub-floating-point number; the simplified first floating-point number includes the first sub-floating-point number and the fourth sub-floating-point number, and the simplified second floating-point number includes the second sub-floating-point number and the third sub-floating-point number.

5. The floating point number processing method according to claim 4, wherein Determining the first operation result based on the simplified first floating-point number and the simplified second floating-point number includes: Performing a multiplication operation on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results; An addition operation is performed on the second number of second operation results to obtain the first operation result.

6. The floating point number processing method according to claim 1, wherein The format conversion of the first floating-point number and the second floating-point number is performed based on the loop indication, and a first operation result after multiplication of the first floating-point number and the second floating-point number is determined based on the converted first floating-point number and the converted second floating-point number, including: Using different processing methods for the format conversion, extracting different parts of the first floating-point number and the second floating-point number to obtain at least two different simplified first floating-point numbers and at least two different simplified second floating-point numbers; the converted first floating-point number includes the simplified first floating-point number, and the converted second floating-point number includes the simplified second floating-point number; Based on a number of loops, in different loops, respectively retrieving different simplified first floating-point numbers and different simplified second floating-point numbers; the number of loops is determined based on the loop indication; Performing a multiplication operation on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results; An addition operation is performed on the second number of second operation results to obtain the first operation result.

7. The floating point number processing method according to claim 6, wherein: The method of respectively retrieving different simplified first floating-point numbers and different simplified second floating-point numbers in different loops based on the number of loops includes: In the first loop, the first and second sub-floating-point numbers are retrieved; In the second loop, the first and third sub-floating-point numbers are retrieved. In the third loop, the fourth sub-floating point number and the second sub-floating point number are retrieved; In the fourth loop, the fourth sub-floating-point number and the third sub-floating-point number are retrieved; the simplified first floating-point number includes the first sub-floating-point number and the fourth sub-floating-point number, and the simplified second floating-point number includes the second sub-floating-point number and the third sub-floating-point number.

8. The floating point number processing method according to any one of claims 1 to 7, characterized in that: The floating point number processing method further includes: Determining a first storage area from a free memory of the tensor computing engine based on a bit width of the first floating-point number, a data read latency, and a preset first storage bit; Based on the bit width of the second floating-point number, the data read delay, and a preset second storage bit, a second storage area is determined from the free memory of the tensor calculation engine.

9. The floating point number processing method according to any one of claims 1 to 7, characterized in that: The floating point number processing method further includes: In response to the configuration instruction, configure metadata information of the tensor and block information of the tensor; each block of the tensor includes a plurality of floating-point numbers in a first data format; Based on the metadata information and the block information, data in any block of the tensor is transferred from the global memory to the local memory.

10. The floating point number processing method according to claim 9, wherein: The global memory and the local memory store floating-point numbers in a target storage manner; the target storage manner can evenly distribute multiple floating-point numbers calculated each time in different storage areas.

11. A tensor computing engine, characterized in that: The tensor computing engine includes: a format conversion unit, configured to perform format conversion on a first floating-point number and a second floating-point number respectively based on a loop indication; the loop indication is used to indicate a processing method and a number of loops for the tensor computing engine to perform the format conversion; the data format of the first floating-point number and the second floating-point number is a first data format; A dot product calculation unit is used to determine a first operation result after multiplying the first floating-point number and the second floating-point number based on the converted first floating-point number and the converted second floating-point number; the converted first floating-point number and the converted second floating-point number are in a second data format, and the second data format is a simplification of the first data format.

12. The tensor computing engine according to claim 11, wherein: The tensor calculation engine further includes: a first storage area and a second storage area; a first storage area, configured to store the first floating-point number, or to store the converted first floating-point number; The second storage area is used to store the second floating-point number, or to store the converted second floating-point number.

13. The tensor computing engine according to claim 11 or 12, characterized in that: The dot product calculation unit further includes: a multiplier, configured to perform a multiplication operation on the simplified first floating-point number and the simplified second floating-point number in each loop to obtain a second number of second operation results; the converted first floating-point number includes the simplified first floating-point number, and the converted second floating-point number includes the simplified second floating-point number; an exponent alignment unit, configured to perform exponent alignment on the second number of second operation results; An adder is used to perform an addition operation on the second number of second operation results after exponent alignment to obtain the first operation result.

Citation Information

Patent Citations

  • Floating-point operand computing method and device using the method

    CN106557299A

  • Method and apparatus for realizing floating point arithmetic operation

    CN106997284A

  • Operation instruction execution method, device and circuit, processor and equipment

    CN116795432A

  • Chip including fused multiply-accumulator, device and control method of data operation

    CN117420982A

  • Operation unit, floating-point number operation method and device

    CN118915995A

Cited By

  • Data processing method and device applied to database system

    CN121233072A