Operator execution method and device, computer equipment and storage medium

By using preset data layout and in-thread data exchange in the fusion operator, the problem of inability to operate directly after data accuracy conversion is solved, and performance improvement and hardware overhead reduction is achieved.

CN120276769AActive Publication Date: 2025-07-08SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510733136.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-08
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

In the fusion operator, due to the inconsistent data accuracy between adjacent operations, the data arrangement method is changed and subsequent operations cannot be performed directly. The prior art reorders through instructions for data exchange between threads, resulting in large hardware overhead and affecting performance.

Method used

The preset data layout method is adopted to store the output results in the thread-local register array, and data exchange is reduced between threads by exchanging data between TLRs corresponding to the same thread, and directly meets the data layout requirements of the next operation.

Benefits of technology

It reduces the hardware overhead of data reordering, improves the performance of converged operators, simplifies the programming process, and improves the efficiency of instruction execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276769A_ABST
    Figure CN120276769A_ABST
Patent Text Reader

Abstract

The invention discloses an operator execution method and device, computer equipment and a storage medium, and belongs to the technical field of artificial intelligence, in the method, after an output result of first precision of a first operation in a fusion operator is obtained, the output result is stored in a TLR array according to a preset data arrangement mode, and the output result of the first precision of the first operation in the fusion operator is obtained; and converting the data in the TLR array from the first precision to a second precision based on a preset data arrangement mode, the second precision being an input precision corresponding to a second operation in the fusion operator, and after determining that the converted data meets a data arrangement requirement corresponding to the second operation, processing the data through each thread to obtain a calculation result of the fusion operator, the processing is determined according to the category of the second operation. Thus, the data arrangement mode of the output result of the previous operation in the fusion operator during output is changed, the data arrangement requirement corresponding to the next operation can be met after precision conversion, data exchange between threads is not needed, hardware overhead is small, and therefore the performance of the fusion operator can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to an operator execution method, apparatus, computer device, and storage medium. Background Art

[0002] Generally, a fused operator includes multiple tensor operations and scalar operations. When the data precisions between adjacent operations in the fused operator are different, precision conversion needs to be implemented in the fused operator. Since the data layout of the output result of the previous operator changes after precision conversion, and the changed layout does not meet the data layout requirements corresponding to the next operator, subsequent operations cannot be directly performed. To avoid affecting subsequent operations, in related technologies, instructions for data exchange between multiple threads are used to rearrange the data. However, the hardware overhead of data exchange between threads is relatively large, which will seriously affect the performance of the fused operator. Summary of the Invention

[0003] Embodiments of this application provide an operator execution method, apparatus, computer device, and storage medium, so as to reduce the hardware overhead of data rearrangement in a fused operator and improve the performance of the fused operator.

[0004] In a first aspect, an embodiment of this application provides an operator execution method. The fused operator includes a first operation and a second operation. The output of the first operation corresponds to a first precision, and the input of the second operation corresponds to a second precision. The method includes: After obtaining the output result of the first operation, store the output result in a thread local register (TLR) array according to a preset data layout when converting from the first precision to the second precision. Each row of TLR in the TLR array corresponds to a thread; Based on the preset data layout, convert the data in the TLR array from the first precision to the second precision; When it is determined that the data in the TLR array after precision conversion meets the data layout requirements corresponding to the second operation, process the data in the TLR array through each thread to obtain the calculation result of the fused operator, where the processing is determined according to the type of the second operation.

[0005] In some embodiments, the output result includes Q rows of data, the TLR array includes M rows of TLR, and the i-th row of TLR in the TLR array corresponds to the i-th thread, where Q and M are positive integers, and 0 ≤ i < Q - 1; The preset data layout satisfies that the data in the i-th row and the (i + Q / 2)-th row of the output result are arranged in the i-th k to the (i + 1)-th In the TLRs corresponding to k threads, the data in the i-th row and the (i + Q / 2)-th row are arranged in different TLR columns, and the data from the 0-th row to the (Q / 2 - 1)-th row are arranged in the same way, and the data from the Q / 2-th row to the (Q - 1)-th row are arranged in the same way, where k is the number of threads corresponding to each row of data in the output result; When the first precision is lower than the second precision, the preset data arrangement method further satisfies that each row of data in the output result is continuously arranged in TLR columns, and is sequentially arranged in k TLRs corresponding to each TLR column until half of the storage space of each TLR is filled, and then each TLR is sequentially filled.

[0006] In some embodiments, when the first precision is lower than the second precision, before converting the data in the TLR array from the first precision to the second precision based on the preset data arrangement method, it further includes: Expanding the TLR columns in the TLR array according to the space required to store the output result of the second precision; For each TLR column with data in the TLR array, according to the rule of moving data between TLRs corresponding to the same thread, move half of the data in the TLR column to an idle TLR column, where the half of the data is the data occupying half of the storage space at the head of each TLR, or the half of the data is the data occupying half of the storage space at the tail of each TLR.

[0007] In some embodiments, when the first precision is higher than the second precision, the preset data arrangement method further satisfies that each row of data in the output result is continuously arranged in TLR groups, and is sequentially arranged row by row in the order of increasing row numbers in the k-row 2-column TLRs corresponding to each TLR group, and after filling one TLR in the 2 TLRs of each row, then arrange the other TLR.

[0008] In some embodiments, when the first precision is higher than the second precision, it further includes: In the case where it is determined that the data in the TLR array after precision conversion does not meet the data arrangement requirements, according to the data arrangement requirements, exchange the data in the TLR array according to the rule of exchanging data in TLRs corresponding to the same thread; and After it is determined that the data in the TLR array after exchange meets the data arrangement requirements, process the data in the TLR array through each thread to obtain the calculation result of the fusion operator.

[0009] In some embodiments, after converting from the first precision to the second precision, each TLR has free storage space. The data arrangement requirement is that each row of data in the output result is continuously arranged in the TLR columns and there is no free storage space between the TLR rows; According to the data arrangement requirement, exchange the data in the TLR array according to the rule of exchanging the data in the TLRs corresponding to the same thread, including: Exchange the data in the TLRs corresponding to each thread according to the rule that the free storage space of each TLR is filled and the continuous data in the same row in the output result is arranged in each TLR; Exchange the data in the TLRs corresponding to each thread according to the rule that there is no free storage space between the TLR rows, so that the data in the TLR array meets the data arrangement requirement.

[0010] In some embodiments, process the data in the TLR array through each thread to obtain the calculation result of the fusion operator, including: When the category of the second operation is a scalar operation, calculate the data in the TLR array through each thread to obtain the calculation result of the fusion operator; When the category of the second operation is a tensor operation, write the data in the TLR array into the tensor core through each thread, and the tensor core is used to perform tensor operations on the written data to obtain the calculation result of the fusion operator.

[0011] In some embodiments, the first operation is a tensor operation, the second operation is a scalar operation, or the first operation is a tensor operation, the second operation is a tensor operation, or the first operation is a scalar operation, the second operation is a tensor operation, or the first operation is a scalar operation, the second operation is a scalar operation.

[0012] In a second aspect, an operator execution device is provided in an embodiment of the present application. The fusion operator includes a first operation and a second operation. The output of the first operation corresponds to a first precision, and the input of the second operation corresponds to a second precision. The device includes: An output module, configured to store the output result into a thread local register TLR array according to a preset data arrangement manner when converting from the first precision to the second precision after obtaining the output result of the first operation. Each row of TLRs in the TLR array corresponds to a thread; A conversion module, configured to convert the data in the TLR array from the first precision to the second precision based on the preset data arrangement manner; A processing module, configured to process the data in the TLR array through each thread to obtain the calculation result of the fusion operator when it is determined that the data in the TLR array after precision conversion meets the data arrangement requirements corresponding to the second operation, where the processing is determined according to the type of the second operation.

[0013] In a third aspect, an embodiment of the present application provides a computer device, including a memory, an artificial intelligence chip, and a computer program stored in the memory and executable on the artificial intelligence chip. When the artificial intelligence chip executes the computer program, the above-mentioned any operator execution method is implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a storage medium. When the computer program in the storage medium is executed by a processor of a computer device, the computer device can execute the above-mentioned any operator execution method.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned any operator execution method is implemented.

[0016] In the embodiment of the present application, the fusion operator includes a first operation and a second operation. The output of the first operation corresponds to a first precision, and the input of the second operation corresponds to a second precision. After obtaining the output result of the first operation, the output result is stored in the TLR array according to a preset data arrangement method when converting from the first precision to the second precision. Each row of TLR in the TLR array corresponds to a thread. Based on the preset data arrangement method, the data in the TLR array is converted from the first precision to the second precision. When it is determined that the data in the TLR array after precision conversion meets the data arrangement requirements corresponding to the second operation, the data in the TLR array is processed through each thread to obtain the calculation result of the fusion operator, where the processing is determined according to the type of the second operation. In this way, by changing the data arrangement method of the output result of the previous operation in the fusion operator during output, the data arrangement requirements corresponding to the next operation in the fusion operator can be met after precision conversion based on the preset data arrangement method, and subsequent operations can be directly performed without exchanging data between threads. Therefore, the hardware overhead of data rearrangement can be saved, thereby improving the performance of the fusion operator. Description of the Drawings

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings: Figure 1 It is a schematic structural diagram of an artificial intelligence chip provided by an embodiment of the present application; Figure 2Schematic diagram of the output result of an operator provided by an embodiment of the present application; Figure 3 Schematic diagram of the arrangement of an output result in a TLR array in the related art; Figure 4 Schematic diagram of data in a TLR array in the related art; Figure 5 Another schematic diagram of data in a TLR array in the related art; Figure 6 Another schematic diagram of data in a TLR array in the related art; Figure 7 Another schematic diagram of data in a TLR array in the related art; Figure 8 Schematic diagram of the execution process of an operator execution method provided by an embodiment of the present application; Figure 9 Schematic diagram of the arrangement of an output result in a TLR array provided by an embodiment of the present application; Figure 10 Another schematic diagram of the arrangement of an output result in a TLR array provided by an embodiment of the present application; Figure 11 For an embodiment of the present application Figure 9 Schematic diagram after precision conversion of the data in; Figure 12 For an embodiment of the present application Figure 10 Schematic diagram after precision conversion of the data in; Figure 13 Another schematic diagram of the execution process of an operator execution method provided by an embodiment of the present application; Figure 14a And Figure 14b Two schematic diagrams of the arrangement of output results in a TLR array provided by an embodiment of the present application; Figure 15a 、 Figure 15b And Figure 15c Three schematic diagrams of the arrangement of output results in a TLR array provided by an embodiment of the present application; Figure 16 For an embodiment of the present application Figure 14a Schematic diagram after data exchange of the data in; Figure 17 Schematic diagram after data exchange of the data in 16 provided by an embodiment of the present application; Figure 18 For an embodiment of the present application Figure 15a Schematic diagram after data exchange of the data in; Figure 19A schematic diagram after data exchange for 18 types of data provided by an embodiment of the present application; Figure 20 A schematic structural diagram of an operator execution device provided by an embodiment of the present application; Figure 21 A schematic hardware structure diagram of a computer device for implementing an operator execution method provided by an embodiment of the present application. Detailed implementation manners

[0018] In order to reduce the hardware overhead of data rearrangement in a fused operator and improve the performance of the fused operator, an embodiment of the present application provides an operator execution method, device, computer device, and storage medium.

[0019] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only for illustrating and explaining the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0020] For the convenience of understanding the present application, in the technical terms involved in the present application: 1. Fused operator.

[0021] A fused operator usually includes multiple tensor operations and scalar operations. Among them, tensor operations such as matrix multiplication operations, convolution operations, etc., and scalar operations such as addition operations, subtraction operations, etc. Compared with tensor operations, scalar operations have a relatively low computational complexity. In order to improve the overall computational performance of the fused operator, two types of hardware with different computing capabilities, tensor cores and scalar cores, are usually used to process tensor operations and scalar operations respectively. Specifically, the tensor core processes the tensor operation and outputs the operation result to the TLR of the scalar core with an appropriate precision, such as 32-bit floating point (32-bit Floating Point, FP32). Then, the scalar core performs scalar operations on the data in the TLR.

[0022] Taking a fused operator including a matrix multiplication and a subsequent activation function (ReLU) as an example, the tensor core is responsible for quickly completing the matrix multiplication operation and outputting the operation result to the TLR of the scalar core in FP32. Then, the scalar core applies ReLU to each data in the TLR to obtain the final calculation result of the fused operator.

[0023] In this way, both the calculation speed of the fused operator can be guaranteed and the necessary numerical accuracy can be ensured, thereby improving the overall computational performance of the fused operator.

[0024] 2. Data layout method.

[0025] The data layout method is used to describe how the output result of an operation, such as a matrix multiplication operation, is arranged in the TLR array. Among them, each row of TLR in the TLR array corresponds to a thread, and each TLR is usually used to store a FP32 data. Subsequently, the example of each TLR storing a FP32 data is also used for introduction. However, it should be noted that each TLR can also store a data of other precisions, such as FP8 and FP16. Moreover, with the development of technology, in the future, each TLR may also store a FP64 or FP128 data.

[0026] 3. Precision conversion.

[0027] In an artificial intelligence chip, the precision of the processed data is mostly floating-point type, such as FP8, FP16, FP32, FP64, FP128, etc. Moreover, precision conversion usually occurs between adjacent precisions, such as precision conversion between FP8 and FP16, precision conversion between FP16 and FP32, precision conversion between FP32 and FP64, or precision conversion between FP64 and FP128.

[0028] 4. Data layout requirements.

[0029] The data layout requirements refer to the layout requirements when data is output.

[0030] In the fused operator, there are generally two cases for data output: data output from the TLR array to the tensor core and data output from the TLR array to the shared memory. The former corresponds to the case where the next operator is a tensor operation, such as a matrix multiplication operation, and the latter corresponds to the case where the next operation is a scalar operation. Therefore, the data layout requirements can also be considered as the data layout requirements corresponding to the next operator. To simplify development, the technical personnel set the data layout requirements in these two cases to be the same.

[0031] As an example, the data layout requirements are that when data is output, each row of data is continuously stored in the k rows of TLR corresponding to itself in the TLR column, and there is no idle storage space between the TLR rows.

[0032] When implementing large model calculations using an artificial intelligence chip, in order to save memory bandwidth and reduce the latency data exchange between the main memory and each level of cache, developers will fuse multiple small operators together to form a fused operator.

[0033] Generally, a fusion operator includes multiple tensor operations and scalar operations. When the data precisions between adjacent operations in the fusion operator are different, precision conversion needs to be implemented in the fusion operator. Since the data layout of the output result of the previous operator will change after precision conversion, the changed layout does not meet the data layout requirements of the next operator, and subsequent operations such as output to shared memory and output to tensor cores cannot be directly performed.

[0034] In order not to affect subsequent operations, in the related art, instructions for data exchange between multiple threads are used to rearrange the data. Taking the fusion operator including matrix multiplication operation and addition operation in sequence, and the output precision of the matrix multiplication operation is FP32 and the input precision of the addition operation is FP16 as an example, the related art will be introduced below.

[0035] See Figure 1 , Figure 1 which is a schematic structural diagram of an artificial intelligence chip 100 provided by an embodiment of the present application, including a tensor core, a scalar core, and a load store cache (LSC). Among them, the tensor core and the scalar core can communicate through an interface, and the scalar core and the load store cache can also communicate through an interface.

[0036] In the related art, the tensor core performs a matrix multiplication operation to obtain an output result of FP32. Assuming the output result is data of 16 rows and 16 columns as shown in Figure 2 , Figure 2 where R represents a row, C represents a column, and RiCi represents the data in the i-th row and j-th column of the output result. The numbers of i and j both range from 0 to 15.

[0037] After that, the tensor core can store the output result in the TLR array of the scalar core. The i-th row TLR in the TLR array corresponds to the i-th thread, and each TLR can store one FP32 data. Assuming the TLR array contains 32 rows and 8 columns of TLRs, that is, there are 32 threads, each thread corresponds to one row of TLRs, and the numbers of each row of TLRs range from 0 to 7. TLRi represents the TLR in the i-th column. Then, Figure 2 the data layout form of the data shown in Figure 3 is as shown in

[0038] where part of the data is shown in gray shading to facilitate distinguishing the data in the same row of the output result. Figure 3 After that, the scalar core can convert the data shown in

[0039] from FP32 to FP16. At this time, each TLR frees up half of its storage space. Figure 3Exchange data between adjacent threads. For example, exchange the data in the TLR corresponding to thread 1 to the TLR corresponding to thread 0, exchange the data in the TLR corresponding to thread 3 to the TLR corresponding to thread 2... exchange the data in the TLR corresponding to thread 31 to the TLR corresponding to thread 30. After the exchange, the data in the TLR array is as Figure 4 shown.

[0040] Next, the scalar core can continue to exchange data between adjacent threads in Figure 4 through the shuffle instruction. For example, exchange the data in the TLR corresponding to thread 2 to the TLR corresponding to thread 1, exchange the data in the TLR corresponding to thread 6 to the TLR corresponding to thread 5... exchange the data in the TLR corresponding to thread 30 to the TLR corresponding to thread 29. After the exchange, the data in the TLR array is as Figure 5 shown.

[0041] Then, the scalar core can continue to exchange data between threads in Figure 5 through the shuffle instruction. For example, exchange the data in TLRi corresponding to thread 0 to TLRi corresponding to thread 2, exchange the data in TLRi corresponding to thread 1 to TLRi corresponding to thread 3... exchange the data in TLRi corresponding to thread 28 to TLRi corresponding to thread 30, and exchange the data in TLRi corresponding to thread 29 to TLRi corresponding to thread 31, where i is an odd number (i.e., 1, 3, 5, 7). After the exchange, the data in the TLR array is as Figure 6 shown.

[0042] Next, the scalar core can exchange the data in the TLRs corresponding to the same thread. After the exchange, the data in the TLR array is as Figure 7 shown, Figure 7 and the data shown satisfies the data arrangement requirements for the addition operation.

[0043] Finally, the arithmetic logic unit (ALU) in the scalar core uses the data shown by each thread to Figure 7 perform an addition operation, thereby obtaining the calculation result of the fusion operator.

[0044] From the above process, it can be seen that after the precision conversion, three data exchanges between threads are performed through three shuffle instructions. See Figure 1, in practical applications, each time data is exchanged between threads through the shuffle instruction, the data needs to be stored from the TLR array into the shared memory of the load-store cache. Then, the data is rearranged in the buffer of the load-store cache, and after rearrangement, it is stored back into the TLR array through the shared memory. This process of data exchange between threads needs to be repeated three times, which is not only time-consuming but also occupies TLR resources for a long time, prolonging the total running time of the fusion operator, and thus seriously affecting the performance of the fusion operator.

[0045] To solve the above problems, the inventor thought of designing a special data arrangement method (i.e., the preset data arrangement method). When the output result of the previous operation in the fusion operator is output in the preset data arrangement method, after precision conversion based on the preset data arrangement method, it can directly meet the data arrangement requirements corresponding to the next operation in the fusion operator, that is, subsequent operations can be directly performed without exchanging data between threads. Therefore, the hardware overhead of data rearrangement can be saved, thereby improving the performance of the fusion operator.

[0046] The execution entity of the operator execution method provided in this application can be an artificial intelligence chip 100, such as a graphics processing unit (GPU), a general-purpose computing on graphics processing units (GPGPU), a domain specific architecture (DSA), etc. The structure of the artificial intelligence chip 100 can be seen Figure 1 , which will not be elaborated here. And, the artificial intelligence chip 100 includes tensor cores and scalar cores, so some steps in this method can be executed by tensor cores and some by scalar cores. The following will introduce it in the way that some steps in this method are executed by tensor cores and some by scalar cores.

[0047] See Figure 8 , Figure 8 is a schematic diagram of the execution process of an operator execution method provided in an embodiment of this application, including the following steps.

[0048] In step 801, the tensor core executes the first operation in the fusion operator and obtains the output result with the first precision.

[0049] Among them, the first operation can be a tensor operation such as matrix multiplication operation, or a scalar operation such as addition operation and subtraction operation. And, the output result of the first operation is usually in matrix form, that is, it contains multiple rows and multiple columns of data.

[0050] In step 802, the tensor core stores the output result in the TLR array according to the preset data arrangement mode when converting from the first precision to the second precision. Each TLR in each row of the TLR array corresponds to a thread.

[0051] Among them, the second precision is the input precision of the second operation in the fusion operator. The second operation is the next operation of the first operation, and the second operation can be a tensor operation or a scalar operation.

[0052] Suppose the output result of the first operation includes Q rows and P columns of data, the TLR array includes M rows and N columns of TLRs, and the TLR in the i-th row of the TLR array corresponds to the i-th thread, where Q, P, M, and N are integers greater than zero, and 0 ≤ i < Q - 1.

[0053] In practical applications, the value of Q can be 8, 16, 32, etc., and the value of M can be twice that of Q. Subsequently, an example with Q being 16 will be introduced. When Q is 16, M is 32, and P can be 16, 32, 64, etc. The value of N is jointly determined by M, the storage space of a single TLR, and the storage space required for the output result of the first operation.

[0054] Then, the preset data arrangement mode satisfies that the data in the i-th row and the (i + Q / 2)-th row of the output result are arranged in the TLRs corresponding to the i-th k to the (i + 1)-th k threads. The data in the i-th row and the (i + Q / 2)-th row are arranged in different TLR columns. Moreover, the arrangement mode of the data from the 0-th row to the Q / 2 - 1-th row is the same, and the arrangement mode of the data from the Q / 2-th row to the Q - 1-th row is the same. k is the number of threads corresponding to each row of data in the output result, such as k = 4.

[0055] And when the first precision is lower than the second precision, the preset data arrangement mode can also satisfy that each row of data in the output result is continuously arranged in the TLR columns, and is sequentially arranged in each TLR until half of the storage space of each TLR is full, and then sequentially fills each TLR. When arranging continuously in the TLR columns, the arrangement order between the TLR columns can be any order, that is, when each row of data is arranged in multiple TLR columns, the arrangement order of the multiple TLR columns can be any order.

[0056] Next, several examples will be used to introduce the preset data arrangement mode when the first precision is lower than the second precision.

[0057] Example 1: The fusion operator sequentially includes matrix multiplication operation 1 and matrix multiplication operation 2. The output precision of matrix multiplication operation 1 is FP16, and the input precision of matrix multiplication operation 2 is FP32.

[0058] Suppose the output result of matrix multiplication operation 1 is as Figure 2As shown, then, Figure 2 The preset data arrangement pattern of the data shown in the TLR array can be as Figure 9 shown. Where TLRi represents the i-th column TLR, and the value of i ranges from 0 to 3. Taking the data in the 0th row of the output result as an example, the data in the 0th row is arranged in TLR0 and TLR2 corresponding to threads 0 to 3. The arrangement order between TLR columns is: TLR0 is arranged first and then TLR2 (the data from the 0th row to the 7th row are all in this arrangement order). Then, the data in the 0th row is sequentially arranged in TLR0 until half of the storage space of each TLR is full, and then each TLR is filled up sequentially. Then, the data in the 0th row is sequentially arranged in TLR2 until half of the storage space of each TLR is full, and then each TLR is filled up sequentially. Taking the data in the 8th row of the output result as an example, the data in the 8th row is arranged in TLR1 and TLR3 corresponding to threads 0 to 3. The arrangement order between TLR columns is: TLR1 is arranged first and then TLR3 (the data from the 8th row to the 15th row are all in this arrangement order). Then, the data in the 8th row is sequentially arranged in TLR1 until half of the storage space of each TLR is full, and then each TLR is filled up sequentially. Then, the data in the 8th row is sequentially arranged in TLR3 until half of the storage space of each TLR is full, and then each TLR is filled up sequentially.

[0059] It should be noted that Figure 9 This is only an example, and there can be other data arrangement patterns. For example, Figure 9 perform at least one exchange between TLR columns on the data in Figure 9 Taking the exchange of the data in TLR0 and TLR2 in Figure 9 as an example, it is equivalent to changing the arrangement order between TLR columns of the data from the 0th row to the 7th row. Taking the exchange of the data in TLR1 and TLR3 in Figure 9 as an example, it is equivalent to changing the arrangement order between TLR columns of the data from the 8th row to the 15th row. Taking the exchange of the data in TLR0 and TLR1 in

[0060] Example 2: The fusion operator sequentially includes a subtraction operation and a matrix multiplication operation. The output precision of the subtraction operation is FP8, and the input precision of the matrix multiplication operation is FP16.

[0061] Assume that the output result of the scalar operation is as Figure 2 shown, then, Figure 2 The preset data arrangement pattern of the data shown in the TLR array can be as Figure 10 shown. Where TLRi represents the i-th column TLR, and the value of i ranges from 0 to 1. In order to clearly show Figure 10For the data in , the data in the TLRs corresponding to threads 16 to 31 are drawn in another column. In fact, the TLRs corresponding to threads 16 to 31 are connected below the TLRs corresponding to threads 0 to 15. Taking the data in the 0th row of the output result as an example, the data in the 0th row is arranged in TLR0 corresponding to threads 0 to 3. In TLR0, it is arranged sequentially until half of the storage space of each TLR is empty, and then each TLR is filled sequentially; taking the data in the 8th row of the output result as an example, the data in the 8th row is arranged in TLR1 corresponding to threads 0 to 3. In TLR1, it is arranged sequentially until half of the storage space of each TLR is empty, and then each TLR is filled sequentially.

[0062] It should be noted that Figure 10 The illustration is only for example, and there can be other data arrangement methods. For example, Figure 10 interchange the data among TLR columns in .

[0063] In step 803, the scalar core converts the data in the TLR array from the first precision to the second precision based on a preset data arrangement method.

[0064] Since more storage space is required after the precision is increased, therefore, the TLR columns in the TLR array can be expanded first according to the space required to store the output result of the second precision. Then, for each TLR column with data in the TLR array, according to the rule of moving data among the TLRs corresponding to the same thread, move half of the data in this TLR column to an idle TLR column. Among them, half of the data can be the data occupying half of the storage space at the head of each TLR in this TLR column, or the data occupying half of the storage space at the tail of each TLR in this TLR column. That is, for each TLR column with data, the first half of the data in the i-th TLR in the TLR column can be moved to the i-th TLR in an idle TLR column, or the second half of the data in the i-th TLR in the TLR column can be moved to the i-th TLR in an idle TLR column. Among them, moving half of the data can be completed in the TLR array without the help of a load-store cache, and the execution complexity of the instruction is relatively low.

[0065] The precision conversion will be introduced below in combination with the foregoing examples.

[0066] For Example 1: In order to convert the Figure 9 data shown from FP16 to FP32, 4 columns of TLRs need to be expanded. Then, the second half of the data in each TLR in the original 4 columns of TLRs can be moved to a column of idle TLRs to obtain the Figure 11 data shown, and then the precision conversion is performed.

[0067] For Example 2: In order to convert the Figure 10The data shown is converted from FP8 to FP16, and 2 columns of TLR need to be extended. Then, the second half of the data of each TLR in the original 2 columns of TLR can be moved to an idle TLR column. Figure 12 It is a schematic diagram of the arranged data after the movement. Then, the precision conversion of the data in the TLR array can be performed.

[0068] In step 804, when the scalar core determines that the data in the TLR array after precision conversion meets the data arrangement requirements corresponding to the second operation, the data in the TLR array is processed by each thread to obtain the calculation result of the fusion operator, and the processing is determined according to the category of the second operation.

[0069] Among them, the data arrangement requirements are that each row of data in the output result is continuously stored in the TLR column and there is no idle storage space between the TLR rows. And when the category of the second operation is a scalar operation, the data in the TLR array can be calculated by each thread to obtain the calculation result of the fusion operator (corresponding to Figure 8 804a therein); when the category of the second operation is a tensor operation, the data in the TLR array can be written into the tensor core by each thread, and the tensor core performs tensor operations on the written data to obtain the calculation result of the fusion operator (corresponding to Figure 8 804b therein).

[0070] For Example 1, Figure 11 it already meets the data arrangement requirements. Therefore, the data in the TLR array can be directly processed by each thread to obtain the calculation result of the fusion operator. Specifically, the scalar core can store the data in the TLR array into the tensor core, and the tensor core performs matrix multiplication operation 2, thereby obtaining the calculation result of the fusion operator.

[0071] Similarly, for Example 2, Figure 12 it already meets the data arrangement requirements. Therefore, the data in the TLR array can be directly processed by each thread to obtain the calculation result of the fusion operator. Specifically, the scalar core can store the data in the TLR array into the tensor core, and the tensor core performs matrix multiplication operation, thereby obtaining the calculation result of the fusion operator.

[0072] In the embodiments of the present application, when the precisions between different operations in the fusion operator are different, the output data of the previous operator is output in a preset data arrangement manner, and combined with an instruction for data exchange in the TLR corresponding to the same thread (an instruction for moving half of the data), the conversion from low-precision data to high-precision data and data rearrangement can be achieved quickly and efficiently. Since the hardware overhead for data exchange between TLRs corresponding to the same thread is much smaller than that for data exchange between TLRs of different threads, the hardware overhead for conversion and data rearrangement can be saved, thereby improving the performance of the fusion operator. In addition, due to the reduction in the number of instructions and the complexity of instruction execution, the programming process can be simplified and the instruction execution efficiency can be improved.

[0073] See Figure 13 , Figure 13 which is a schematic diagram of the execution process of another operator execution method provided by the embodiments of the present application, including the following steps.

[0074] In step 1301, the tensor core executes the first operation in the fusion operator and obtains the output result with the first precision.

[0075] Among them, the first operation can be a tensor operation such as matrix multiplication operation, or a scalar operation such as addition operation and subtraction operation. And the output result of the first operation is usually in matrix form, that is, it contains multiple rows and multiple columns of data.

[0076] In step 1302, the tensor core stores the output result in the TLR array according to the preset data arrangement manner when converting from the first precision to the second precision. Each row of TLRs in the TLR array corresponds to a thread.

[0077] Among them, the second precision is the input precision of the second operation in the fusion operator. The second operation is the next operation of the first operation, and the second operation can be a tensor operation or a scalar operation.

[0078] Suppose the output result of the first operation includes Q rows and P columns of data, the TLR array includes M rows and N columns of TLRs, and the i-th row of TLRs in the TLR array corresponds to the i-th thread, where Q, P, M, and N are positive integers, and 0 ≤ i < Q - 1.

[0079] In practical applications, the value of Q can be 8, 16, 32, etc., and the value of M can be twice that of Q. Subsequently, an example with Q being 16 will be introduced. When Q is 16, M is 32, and P can be values such as 16, 32, 64, etc. The value of N is jointly determined by M, the storage space of a single TLR, and the storage space required for the output result of the first operation.

[0080] Then, the preset data arrangement manner satisfies that the data in the i-th row and the (i + Q / 2)-th row of the output result are arranged in the i k to the (i + 1) In the TLRs corresponding to k threads, the data in the i-th row and the (i + Q / 2)-th row are arranged in different TLR columns. Moreover, the data arrangement patterns of the rows from the 0-th row to the (Q / 2 - 1)-th row are the same, and the data arrangement patterns of the rows from the Q / 2-th row to the (Q - 1)-th row are the same. k is the number of threads corresponding to each row of data in the output result, such as k = 4.

[0081] Furthermore, when the first precision is higher than the second precision, the preset data arrangement pattern can also satisfy that each row of data in the output result is continuously arranged by TLR groups, and in the 2-column TLR corresponding to k rows in each TLR group, they are arranged row by row in ascending order of row numbers, and after filling one TLR in the 2 TLRs of each row, the other TLR is arranged. Here, each TLR group includes 2-column TLRs, and the 2-column TLRs can be adjacent or non-adjacent in the TLR array. When arranged continuously by TLR groups, the arrangement order between groups can be any arrangement order, that is, when there are multiple TLR groups, the arrangement order between multiple TLR groups can be any order.

[0082] The following introduces the preset data arrangement pattern when the first precision is higher than the second precision through several examples.

[0083] Example 3: The fusion operator sequentially includes matrix multiplication operation and subtraction operation. The output precision of the matrix multiplication operation is FP32, and the input precision of the subtraction operation is FP16.

[0084] Assume that the output result of the matrix multiplication operation is as Figure 2 shown. Then, Figure 2 the preset data arrangement pattern of the data shown in the TLR array can be as Figure 14a shown. To clearly show the data belonging to the same row in the output result, Figure 14a part of the data is marked in gray in

[0085] Taking the data in the 0-th row of the output result as an example, the data in the 0-th row is arranged in TLR0, TLR1, TLR4, and TLR5 corresponding to threads 0 to 3. Moreover, TLR0 and TLR1 are TLR group 0, and TLR4 and TLR5 are TLR group 1. The arrangement order between groups is: arrange TLR group 0 first and then TLR group 1 (the arrangement order between groups for the data from the 0-th row to the 7-th row is the same). Then, the data in the 0-th row is first arranged row by row in ascending order of row numbers in the 4-row 2-column TLR corresponding to TLR group 0, and after filling one TLR in the 2 TLRs of each row, the other TLR is arranged. Then, in the 4-row 2-column TLR corresponding to TLR group 1, it is arranged row by row in ascending order of row numbers, and after filling one TLR in the 2 TLRs of each row, the other TLR is arranged.

[0086] Taking the data in the 8th row of the output result as an example, the data in the 8th row is arranged in TLR2, TLR3, TLR6, and TLR7 corresponding to threads 0 to 3. Moreover, TLR2 and TLR3 are TLR group 2, and TLR6 and TLR7 are TLR group 3. The arrangement order between groups is: arrange TLR group 2 first and then TLR group 3 (the data from the 8th row to the 15th row are all in this arrangement order between groups). Then, the data in the 8th row is first arranged row by row in the 4-row 2-column TLR corresponding to TLR group 2 in ascending order of row numbers, and after filling one TLR in the two TLRs in each row, the other TLR is arranged. Then, in the 4-row 2-column TLR corresponding to TLR group 3, it is arranged row by row in ascending order of row numbers, and after filling one TLR in the two TLRs in each row, the other TLR is arranged.

[0087] It should be noted that Figure 14a the above is only an example. In fact, there can be other data arrangement methods, such as Figure 14a performing at least one TLR column exchange on the data in Figure 14a Taking the exchange of the data in TLR1 and TLR5 in Figure 14b as an example, the data obtained after the exchange is as shown in

[0088] Example 4: The fusion operator sequentially includes scalar operation 1 and scalar operation 2, and the output precision of scalar operation 1 is FP16, and the input precision of scalar operation 2 is FP8.

[0089] Assuming that the output result of the matrix multiplication operation is as shown in Figure 2 then, Figure 2 the preset data arrangement method of the data shown in Figure 15a can be as shown. Taking the data in the 0th row of the output result as an example, the data in the 0th row is arranged in TLR0 and TLR1 corresponding to threads 0 to 3. Moreover, TLR0 and TLR1 are a TLR group. Then, the data in the 0th row is arranged row by row in the k-row 2-column TLR corresponding to this TLR group in ascending order of row numbers, and after filling one TLR in the two TLRs in each row, the other TLR is arranged. Taking the data in the 8th row of the output result as an example, the data in the 8th row is arranged in TLR2 and TLR3 corresponding to threads 0 to 3. Moreover, TLR2 and TLR3 are a TLR group. Then, the data in the 8th row is arranged row by row in this TLR group in ascending order of row numbers, and after filling one TLR in the two TLRs in each row, the other TLR is arranged.

[0090] It should be noted that Figure 15a the above is only an example. In fact, there can be other data arrangement methods, such as Figure 15a performing at least one TLR column exchange on the data inFigure 15a Taking the exchange of data in TLR1 and TLR3 as an example, the data as shown in Figure 15b is obtained. Taking Figure 15a the exchange of data in TLR1 and TLR3 as an example, the data as shown in Figure 15b is obtained. Taking Figure 15a the exchange of data in TLR1, TLR2 and TLR3 as an example, the data as shown in Figure 15c is obtained.

[0091] In step 1303, based on a preset data arrangement method, the scalar core converts the data in the TLR array from the first precision to the second precision.

[0092] Since less storage space is required after the precision reduction, there will be free storage space in each TLR in the TLR array, generally half of the free storage space.

[0093] In step 1304, when it is determined that the data in the TLR array after precision conversion does not meet the data arrangement requirements corresponding to the second operation, according to the data arrangement requirements, the data in the TLR array is exchanged according to the rule of exchanging the data in the TLRs corresponding to the same thread.

[0094] Among them, the data arrangement requirements are that each row of data is continuously stored by column in the TLRs corresponding to k threads and there is no free storage space between the TLR rows. Then, according to the rule that the free storage space of each TLR is filled and the same-row continuous data in the output result is arranged in each TLR, the data in the TLRs corresponding to each thread is exchanged. Furthermore, according to the rule that there is no free storage space between the TLRs of the same thread, the data in the TLRs corresponding to each thread is exchanged, so that the data in the TLR array meets the data arrangement requirements.

[0095] It should be noted that the exchange of the data in the TLRs corresponding to the same thread can be completed in the TLR array without relying on the load / store cache. That is to say, the above two exchanges of the data in the TLRs corresponding to each thread can be directly performed in the TLR array, and the execution complexity of the instruction is relatively low.

[0096] The data exchange process is introduced below in combination with the above examples.

[0097] For Example 3, taking Figure 14a as an example, Figure 14aIf the data arrangement requirements are not met, the data in the TLR corresponding to each thread can be exchanged according to the rule that the idle storage space of each TLR is filled and the consecutive data in the same row of the output result is arranged in each TLR. For example, for the i-th thread, the data in each odd-column TLR in the i-th row TLR corresponding to it is exchanged to the adjacent previous even-column TLR. The data in the TLR array after the exchange is as Figure 16 shown.

[0098] Then, according to the rule that there is no idle storage space between TLR rows, the data in the TLRs corresponding to the same thread in Figure 16 is exchanged. For example, the data in TLR2 corresponding to each thread is exchanged to TLR1, the data in TLR4 is exchanged to TLR2, and the data in TLR6 is exchanged to TLR3. The data in the TLR array after the exchange is as Figure 17 shown.

[0099] It should be noted that Figure 14b also does not meet the data arrangement requirements, and the steps of data exchange for the data shown in Figure 14b are similar and will not be elaborated here.

[0100] For Example 4, taking Figure 15a as an example, Figure 15a if the data arrangement requirements are not met, the data in the TLR corresponding to each thread can be exchanged according to the rule that the idle storage space of each TLR is filled and the consecutive data in the same row of the output result is arranged in each TLR. For example, the data in TLR1 corresponding to each thread is exchanged to TLR0, and the data in TLR3 is exchanged to TLR2. The data in the TLR array after the exchange is as Figure 18 shown (in order to clearly show the data in Figure 18 , the data in the TLRs corresponding to threads 16 to 31 are drawn in another column. In fact, the TLRs corresponding to threads 16 to 31 are connected below the TLRs corresponding to threads 0 to 15).

[0101] Then, according to the rule that there is no idle storage space between TLR rows, the data in the TLRs corresponding to the same thread in Figure 18 is exchanged. For example, the data in TLR2 corresponding to each thread is exchanged to TLR1. The data in the TLR array after the exchange is as Figure 19 shown (in order to clearly show the data in Figure 19 , the data in the TLRs corresponding to threads 16 to 31 are drawn in another column. In fact, the TLRs corresponding to threads 16 to 31 are connected below the TLRs corresponding to threads 0 to 15).

[0102] It should be noted that Figure 15b ,Figure 15c Nor does it meet the data arrangement requirements. For the Figure 15b , Figure 15c data shown, the steps of data exchange are similar and will not be elaborated here.

[0103] In step 1305, the data in the TLR array is processed by each thread to obtain the calculation result of the fusion operator, and the processing is determined according to the category of the second operation.

[0104] Among them, when the category of the second operation is a scalar operation, the data in the TLR array can be calculated by each thread to obtain the calculation result of the fusion operator (corresponding to Figure 13 1305a in ); when the category of the second operation is a tensor operation, the data in the TLR array can be written into the tensor core by each thread, and the tensor core performs tensor operations on the written data to obtain the calculation result of the fusion operator (corresponding to Figure 13 1305b in ).

[0105] For Example 3, Figure 17 it already meets the data arrangement requirements. Therefore, the data in the TLR array can be directly processed by each thread to obtain the calculation result of the fusion operator. Specifically, the scalar core can use each thread to perform a subtraction operation on the data in the TLR array, thereby obtaining the calculation result of the fusion operator.

[0106] Similarly, for Example 4, Figure 19 it already meets the data arrangement requirements. Therefore, the data in the TLR array can be directly processed by each thread to obtain the calculation result of the fusion operator. Specifically, the scalar core can use each thread to perform scalar operation 2 on the data in the TLR array, thereby obtaining the calculation result of the fusion operator.

[0107] In the embodiments of the present application, when the precisions between different operations in the fusion operator are different, the output data of the previous operator is output in a preset data arrangement manner, and by combining two instructions for data exchange in the TLRs corresponding to the same thread (referring to the instructions for exchanging the data in the TLRs corresponding to each thread twice), the conversion from high-precision data to low-precision data and data rearrangement can be quickly and efficiently realized. Since the hardware overhead for exchanging data between the TLRs of the same thread is much smaller than that for exchanging data between the TLRs of different threads, the hardware overhead for data rearrangement can be saved, thereby improving the performance of the fusion operator. In addition, due to the reduction in the number of instructions and the reduction in the complexity of instruction execution, the programming process can be simplified and the instruction execution efficiency can be improved.

[0108] Based on the same technical concept, the embodiments of the present application also provide an operator execution device. The principle of the operator execution device for solving problems is similar to that of the above operator execution method. Therefore, the implementation of the operator execution device can refer to the implementation of the operator execution method, and the repeated parts will not be elaborated.

[0109] See Figure 20 , Figure 20 which is a schematic structural diagram of an operator execution device provided by an embodiment of the present application, including: An output module 201, configured to, after obtaining the output result of the first operation, store the output result into a thread local register TLR array according to a preset data arrangement manner when converting from the first precision to the second precision, where each row of TLR in the TLR array corresponds to a thread; A conversion module 202, configured to convert the data in the TLR array from the first precision to the second precision based on the preset data arrangement manner; A processing module 203, configured to, when determining that the data in the TLR array after precision conversion meets the data arrangement requirements corresponding to the second operation, process the data in the TLR array through each thread to obtain the calculation result of the fusion operator, where the processing is determined according to the type of the second operation.

[0110] In some embodiments, the output result includes Q rows of data, the TLR array includes M rows of TLR, and the i-th row of TLR in the TLR array corresponds to the i-th thread, where Q and M are integers greater than zero, and 0 ≤ i < Q - 1; The preset data arrangement manner satisfies that the data in the i-th row and the (i + Q / 2)-th row of the output result are arranged in the TLRs corresponding to the i-th k to the (i + 1)-th k threads, the data in the i-th row and the (i + Q / 2)-th row are arranged in different TLR columns, and the arrangement manners of the data from the 0-th row to the Q / 2 - 1-th row are the same, the arrangement manners of the data from the Q / 2-th row to the Q - 1-th row are the same, and k is the number of threads corresponding to each row of data in the output result; When the first precision is lower than the second precision, the preset data arrangement manner further satisfies that each row of data in the output result is continuously arranged in TLR columns, and after being sequentially arranged to half of the storage space of each TLR corresponding to k TLRs in each TLR column, it is sequentially arranged to fill each TLR.

[0111] In some embodiments, when the first precision is lower than the second precision, the conversion module 202 is further configured to: Before converting the data in the TLR array from the first precision to the second precision based on the preset data arrangement manner, expand the TLR columns in the TLR array according to the space required to store the output result of the second precision. For each TLR column with data in the TLR array, according to the rule of moving data between TLRs corresponding to the same thread, move half of the data in the TLR column to an idle TLR column, where the half of the data is the data occupying half of the storage space at the head of each TLR, or the half of the data is the data occupying half of the storage space at the tail of each TLR.

[0112] In some embodiments, when the first precision is higher than the second precision, the preset data arrangement method further satisfies that each row of data in the output result is continuously arranged by TLR group, and is arranged row by row in ascending order of row numbers in the k-row 2-column TLR corresponding to each TLR group. After filling one TLR in the 2 TLRs of each row, the other TLR is arranged.

[0113] In some embodiments, when the first precision is higher than the second precision, it further includes: An exchange module 204, configured to, when it is determined that the data in the TLR array does not meet the data arrangement requirements after precision conversion, exchange the data in the TLR array according to the data arrangement requirements and the rule of exchanging data in TLRs corresponding to the same thread; A processing module 203 is further configured to, after it is determined that the data in the TLR array meets the data arrangement requirements after exchange, process the data in the TLR array through each thread to obtain the calculation result of the fusion operator.

[0114] In some embodiments, after converting from the first precision to the second precision, each TLR has idle storage space, and the data arrangement requirement is that each row of data in the output result is continuously arranged by TLR column and there is no idle storage space between TLR rows; the exchange module 204 is specifically configured to: Exchange the data in the TLRs corresponding to each thread according to the rule that the idle storage space of each TLR is filled and the continuous data in the same row in the output result is arranged in each TLR; Exchange the data in the TLRs corresponding to each thread according to the rule that there is no idle storage space between TLR rows, so that the data in the TLR array meets the data arrangement requirements.

[0115] In some embodiments, the processing module 203 is specifically configured to: When the category of the second operation is a scalar operation, calculate the data in the TLR array through each thread to obtain the calculation result of the fusion operator; When the category of the second operation is a tensor operation, write the data in the TLR array into a tensor core through each thread, and the tensor core is used to perform tensor operations on the written data to obtain the calculation result of the fusion operator.

[0116] In some embodiments, the first operation is a tensor operation and the second operation is a scalar operation; alternatively, the first operation is a tensor operation and the second operation is a tensor operation; alternatively, the first operation is a scalar operation and the second operation is a tensor operation; alternatively, the first operation is a scalar operation and the second operation is a scalar operation.

[0117] The division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, each functional module in the embodiments of the present application may be integrated in a processor, may exist separately physically, or two or more modules may be integrated in one module. The coupling between each module can be realized through some interfaces, and these interfaces are usually electrical communication interfaces, but mechanical interfaces or other forms of interfaces are not excluded. Therefore, the modules described as separate components may or may not be physically separated, and may be located in one place or distributed at different positions of the same or different devices. The above integrated modules can be implemented in the form of hardware or in the form of software function modules.

[0118] After introducing the operator execution method and apparatus according to the exemplary embodiments of the present application, next, a computer device according to another exemplary embodiment of the present application will be introduced.

[0119] Based on the same technical concept, the embodiments of the present application provide a computer device, as Figure 21 shown, including at least one artificial intelligence chip 100 and a memory 200 connected to the at least one artificial intelligence chip 100. In the embodiments of the present application, the specific connection medium between the artificial intelligence chip 100 and the memory 200 is not limited. Figure 21 Taking the connection between the artificial intelligence chip 100 and the memory 200 through a bus as an example. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0120] In the embodiments of the present application, the memory 200 stores instructions executable by the at least one artificial intelligence chip 100. The at least one artificial intelligence chip 100 can execute the steps of the above operator execution method by executing the instructions stored in the memory 200.

[0121] Among them, the artificial intelligence chip 100 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and circuits. By running or executing the instructions stored in the memory 200 and calling the data stored in the memory 200, the operator execution can be realized. Optionally, the artificial intelligence chip 100 may include one or more processing units. The artificial intelligence chip 100 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the artificial intelligence chip 100. In some embodiments, the artificial intelligence chip 100 and the memory 200 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.

[0122] The artificial intelligence chip 100 can be a general-purpose processor, such as a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0123] The memory 200, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 200 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disc, and so on. The memory 200 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 200 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0124] In an exemplary embodiment, a storage medium is further provided. When the computer program in the storage medium is executed by a processor of a computer device, the computer device can execute any of the above operator execution methods. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0125] In an exemplary embodiment, a computer program product is further provided. When the computer program product is executed by a computer device, the computer device can implement any of the exemplary methods provided by the present application.

[0126] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0127] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0128] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also includes these changes and modifications.

Claims

1. An operator execution method, characterized in that The fusion operator includes a first operation and a second operation. The output of the first operation corresponds to a first precision, and the input of the second operation corresponds to a second precision. The method includes: After obtaining the output result of the first operation, store the output result in the thread local register (TLR) array according to a preset data arrangement method when converting from the first precision to the second precision. Each row of TLR in the TLR array corresponds to a thread. Based on the preset data arrangement method, convert the data in the TLR array from the first precision to the second precision. When it is determined that the data in the TLR array after precision conversion meets the data arrangement requirements corresponding to the second operation, process the data in the TLR array through each thread to obtain the calculation result of the fusion operator, where the processing is determined according to the type of the second operation.

2. The method according to claim 1, wherein The output result includes Q rows of data, the TLR array includes M rows of TLR, and the i-th row of TLR in the TLR array corresponds to the i-th thread, where Q and M are integers greater than zero, and 0 ≤ i < Q - 1. The preset data arrangement method satisfies that the data in the i-th row and the (i + Q / 2)-th row of the output result are arranged in the i-th k to the (i + 1)-th k TLRs corresponding to threads. The data in the i-th row and the (i + Q / 2)-th row are arranged in different TLR columns, and the data arrangement methods of the 0-th row to the (Q / 2 - 1)-th row are the same, and the data arrangement methods of the Q / 2-th row to the (Q - 1)-th row are the same. k is the number of threads corresponding to each row of data in the output result; When the first precision is lower than the second precision, the preset data arrangement method further satisfies that each row of data in the output result is continuously arranged in TLR columns, and in k TLRs corresponding to each TLR column, it is arranged in order until half of the storage space of each TLR is full, and then each TLR is filled in order.

3. The method according to claim 1 or 2, characterized in that, When the first precision is lower than the second precision, before converting the data in the TLR array from the first precision to the second precision based on the preset data arrangement method, it further includes: Expand the TLR columns in the TLR array according to the space required to store the output result of the second precision. For each TLR column with data in the TLR array, move half of the data in the TLR column to an idle TLR column according to the rule of moving data between TLRs corresponding to the same thread. The half of the data is the data occupying half of the storage space at the head of each TLR, or the half of the data is the data occupying half of the storage space at the tail of each TLR.

4. The method according to claim 2, wherein When the first precision is higher than the second precision, the preset data arrangement method further satisfies that each row of data in the output result is continuously arranged in TLR groups, and in k rows and 2 columns of TLRs corresponding to each TLR group, it is arranged row by row in ascending order of row numbers. After filling one TLR in each row of the 2 TLRs, the other TLR is arranged.

5. The method according to claim 1 or 4, characterized in that When the first precision is higher than the second precision, it further includes: When it is determined that the data in the TLR array after precision conversion does not meet the data arrangement requirements, exchange the data in the TLR array according to the data arrangement requirements and the rule of exchanging data in TLRs corresponding to the same thread; and When it is determined that the data in the TLR array after exchange meets the data arrangement requirements, process the data in the TLR array through each thread to obtain the calculation result of the fusion operator.

6. The method according to claim 5, wherein After converting from the first precision to the second precision, each TLR has free storage space, and the data arrangement requirement is that the data in each row of the output result is continuously arranged in TLR columns and there is no free storage space between TLR rows; According to the data arrangement requirement, exchange the data in the TLR array according to the rule of exchanging the data in the TLRs corresponding to the same thread, including: Exchange the data in the TLRs corresponding to each thread according to the rule that the free storage space of each TLR is filled and the data in the same row of the output result is arranged continuously in each TLR; Exchange the data in the TLRs corresponding to each thread according to the rule that there is no free storage space between TLR rows, so that the data in the TLR array meets the data arrangement requirement.

7. The method according to claim 1, characterized in that Process the data in the TLR array through each thread to obtain the calculation result of the fusion operator, including: When the category of the second operation is a scalar operation, calculate the data in the TLR array through each thread to obtain the calculation result of the fusion operator; When the category of the second operation is a tensor operation, write the data in the TLR array into the tensor core through each thread, and the tensor core is used to perform tensor operations on the written data to obtain the calculation result of the fusion operator.

8. The method according to claim 1, wherein the first operation is a tensor operation and the second operation is a scalar operation, or the first operation is a tensor operation and the second operation is a tensor operation, or the first operation is a scalar operation and the second operation is a tensor operation, or the first operation is a scalar operation and the second operation is a scalar operation.

9. An operator execution device, characterized in that, The fusion operator includes a first operation and a second operation. The output of the first operation corresponds to the first precision, and the input of the second operation corresponds to the second precision. The device includes: An output module, configured to, after obtaining the output result of the first operation, store the output result in a thread local register TLR array according to a preset data arrangement method when converting from the first precision to the second precision. Each row of TLRs in the TLR array corresponds to a thread; A conversion module, configured to convert the data in the TLR array from the first precision to the second precision based on the preset data arrangement method; A processing module, configured to, when it is determined that the data in the TLR array after precision conversion meets the data arrangement requirement corresponding to the second operation, process the data in the TLR array through each thread to obtain the calculation result of the fusion operator, wherein the processing is determined according to the category of the second operation.

10. A computer device, comprising a memory, an artificial intelligence chip, and a computer program stored in the memory and executable on the artificial intelligence chip, characterized in that, The artificial intelligence chip implements the steps of the method according to any one of claims 1-8 when executing the computer program.

11. A storage medium, characterized in that, When the computer program in the storage medium is executed by the processor of the computer device, the computer device can execute the method according to any one of claims 1-8.

12. A computer program product, characterized in that, Including a computer program, which implements the method according to any one of claims 1-8 when executed by a processor.

Citation Information

Patent Citations

  • Floating point multiply-add unit with fusion precision conversion function and application method of floating point multiply-add unit

    CN115390790A

  • Method for accelerating random precision sparse matrix multiplication and addition operation based on tensor core

    CN119646369A

  • Generation of vector codes for tensor convolutions

    US11422781B1

  • Propagating reduced-precision on computation graphs

    US20200249924A1

  • Multi-precision tensor multiplication in neural network

    WO2025091335A1

Cited By

  • Processor, chip product, computer equipment and tensor processing method

    CN121785664A