A method, device, equipment and storage medium for fusing fork operators

Through coroutine programming mode and multi-package collaboration function, the forked operators are fused into fusion operators and executed concurrently, which solves the problem of low performance of forked operators in the calculation diagram and improves the utilization and operation efficiency of hardware computing power.

CN118152980BActive Publication Date: 2025-07-18SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410339440.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-07-18
Estimated Expiration
2044-03-22

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively integrate the fork operators in the calculation diagram, affecting its performance.

Method used

Using coroutine programming mode and multi-package collaboration function, multiple forked operators are fused into fusion operators and executed concurrently. The input data is loaded from the video memory to the cache through a loading operation to avoid repeated loading.

Benefits of technology

It greatly improves performance bandwidth, makes full use of hardware computing power, shortens running time, and improves the operating efficiency and performance of computing graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118152980B_ABST
    Figure CN118152980B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, device, equipment, and storage medium for fusing fork operators, which relates to the field of artificial intelligence technology. The method includes: obtaining a plurality of fork operators with the same input data; performing operator fusion on the plurality of fork operators to obtain a fused operator, and the execution process of the fused operator includes: loading the input data into the cache; and concurrently executing the plurality of fork operators based on the input data in the cache. In this way, the input data is loaded from the video memory to the cache through one loading operation, and it is not necessary to load the input data from the video memory when executing each fork operator, avoiding repeated loading of the input data, which greatly improves the performance bandwidth and fully utilizes the hardware computing power. Since the plurality of fork operators are executed concurrently, the calculation time of the plurality of fork operators is shortened to the calculation time of the fork operator with the longest time-consuming, thereby greatly shortening the running time and improving the running efficiency and computational graph performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method, device, equipment and storage medium for fusing fork operators. Background Art

[0002] An artificial intelligence model usually refers to a neural network model trained to perform inference and prediction, such as an image inference model, a speech inference model, etc. The operations of an artificial intelligence model can be implemented by operators in a computation graph. A computation graph is a multi-graph structure used to represent the computational tasks and data flow process of an artificial intelligence model. An operator refers to various operations performed on tensors of each layer in an artificial intelligence model. For example, the convolution operation performed by the convolution layer of an artificial intelligence model on the input data of the artificial intelligence model is a convolution operator. A tensor can be understood as a multi-dimensional array, which can have any number of dimensions, and different tensors can have different data types and shapes. The computation graph of an artificial intelligence model can include operators that perform a variety of operations on tensors.

[0003] In practical applications, the computation graph contains upstream and downstream operators with strong dependencies. To reduce data movement during the calculation of upstream and downstream operators, related technologies adopt traditional fusion strategies to fuse upstream and downstream operators. In this way, when the upstream operator is executed, the intermediate result obtained by calculating the upstream operator is saved in the cache, instead of writing the intermediate result back to the global memory. The downstream operator directly obtains the intermediate result from the cache and performs subsequent calculations, thereby improving the performance of the operator.

[0004] However, in addition to the upstream and downstream operators with strong dependencies, the computation graph also contains fork operators without dependency relationships. And traditional fusion strategies are difficult to fuse fork operators well, thus affecting the performance of fork operators. Summary of the Invention

[0005] The embodiments of the present application provide a method, device, equipment and storage medium for fusing fork operators, which are used to fuse fork operators and improve the performance of fork operators.

[0006] On the one hand, the embodiments of the present application provide a method for fusing fork operators, and the method includes:

[0007] Obtain a plurality of fork operators, and the input data of the plurality of fork operators is the same;

[0008] Fuse the plurality of fork operators to obtain a fused operator, and the execution process of the fused operator includes: loading the input data into the cache; and concurrently executing the plurality of fork operators based on the input data in the cache.

[0009] On the one hand, an embodiment of the present application provides a forking operator fusion device, which includes:

[0010] An acquisition module, configured to acquire a plurality of forking operators, where the input data of the plurality of forking operators is the same;

[0011] A fusion module, configured to perform operator fusion on the plurality of forking operators to obtain a fusion operator, and the execution process of the fusion operator includes: loading the input data into a cache; and concurrently executing the plurality of forking operators based on the input data in the cache.

[0012] Optionally, the plurality of forking operators include: a first forking operator executed on a tensor calculation unit and a second forking operator executed on a vector calculation unit;

[0013] Specifically, the fusion module is configured to:

[0014] Adopt a coroutine programming mode to perform operator fusion on the first forking operator and the second forking operator to obtain a fusion operator.

[0015] Optionally, the cache is: a shared cache of the tensor calculation unit and the vector calculation unit;

[0016] Specifically, the fusion module is configured to:

[0017] Obtain the input data from the shared cache through the tensor calculation unit, and execute the first forking operator based on the input data;

[0018] Obtain the input data from the shared cache through the vector calculation unit, and execute the second forking operator based on the input data.

[0019] Optionally, the plurality of forking operators include: a first forking operator executed on a tensor calculation unit and a plurality of second forking operators executed on a vector calculation unit;

[0020] Specifically, the fusion module is configured to:

[0021] Based on the multi-pack cooperation function of the vector calculation unit, perform operator fusion on the plurality of second forking operators to obtain a preliminary fusion result;

[0022] Adopt a coroutine programming mode to perform operator fusion on the first forking operator and the preliminary fusion result to obtain a fusion operator.

[0023] Optionally, the cache is: a shared cache of the tensor calculation unit and the vector calculation unit;

[0024] Specifically, the fusion module is configured to:

[0025] The input data is fetched from the shared cache by the tensor calculation unit, and the first forking operator is executed based on the input data;

[0026] The input data is fetched from the shared cache by the vector calculation unit, and the multiple second forking operators are concurrently executed based on the input data.

[0027] Optionally, the tensor calculation unit and the vector calculation unit are heterogeneous computing cores.

[0028] Optionally, the multiple forking operators are operators executed on the vector calculation unit;

[0029] The fusion module is specifically configured to:

[0030] Based on the multi-pack cooperation function of the vector calculation unit, the multiple forking operators are fused to obtain a fused operator.

[0031] Optionally, the cache is: the shared cache of the tensor calculation unit and the vector calculation unit; or, the exclusive cache of the vector calculation unit;

[0032] The fusion module is specifically configured to:

[0033] The input data is fetched from the shared cache or the exclusive cache by the vector calculation unit, and the multiple forking operators are concurrently executed based on the input data.

[0034] On the one hand, an embodiment of the present application provides a computer device, including a memory, a processor chip, and a computer program stored on the memory and executable on the processor chip. When the processor chip executes the program, the steps of the above forking operator fusion method are implemented.

[0035] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is enabled to execute the steps of the above forking operator fusion method.

[0036] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is enabled to execute the steps of the above forking operator fusion method.

[0037] In the embodiments of the present application, multiple forking operators are fused into a fused operator, so that when the fused operator is executed, multiple forking operators are executed concurrently. Moreover, for the same input data corresponding to multiple forking operators, the input data is loaded from the video memory to the cache through a single loading operation, instead of loading the input data from the video memory when each forking operator is executed, thus avoiding repeated loading of the input data, greatly improving the performance bandwidth, and making full use of the hardware computing power. Secondly, since multiple forking operators are executed concurrently during the execution of the fused operator, the calculation time of multiple forking operators is shortened to the calculation time of the forking operator with the longest time consumption, thereby greatly shortening the running time and improving the running efficiency and the performance of the computational graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0039] Figure 1 Schematic structural diagram of a processor chip provided by an embodiment of the present application;

[0040] Figure 2 Schematic diagram of a computational graph in the related art;

[0041] Figure 3 Schematic diagram of a computational graph in the related art;

[0042] Figure 4 Schematic flow chart of a method for fusing forking operators provided by an embodiment of the present application;

[0043] Figure 5 Schematic diagram of a computational graph provided by an embodiment of the present application;

[0044] Figure 6 Schematic flow chart of an execution method of a forking operator in the related art;

[0045] Figure 7 Schematic flow chart of a method for fusing forking operators provided by an embodiment of the present application;

[0046] Figure 8 Schematic flow chart of a method for fusing forking operators provided by an embodiment of the present application;

[0047] Figure 9 Schematic flow chart of a method for fusing forking operators provided by an embodiment of the present application;

[0048] Figure 10Schematic diagram of a structure of a fork operator fusion device provided by an embodiment of the present application;

[0049] Figure 11 Schematic diagram of a structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0051] The design concept of the embodiments of the present application will be introduced below.

[0052] Refer to Figure 1 , which is a structural diagram of a processor chip applicable to the embodiments of the present application. The processor chip 100 at least includes: a video memory 101 and a plurality of execution units 102. Among them, each execution unit 102 includes: a cache 103, a tensor calculation unit 104, and a vector calculation unit 105. The tensor calculation unit can also be called a tensor calculation core, and the vector calculation unit can also be called a vector calculation core. The tensor calculation unit 104 and the vector calculation unit 105 are heterogeneous calculation cores.

[0053] The video memory 101 can be a High Bandwidth Memory (HBM for short), or other types of memories. The cache 102 is a temporary memory, and its capacity is smaller than that of the video memory 101, but the data exchange speed is faster than that of the video memory 101.

[0054] The cache 102 includes: a shared cache of the tensor calculation unit 104 and the vector calculation unit 105 (for example, a Gemm Main Buffer (GMB for short)), and also includes: an exclusive cache of the tensor calculation unit 104 or the vector calculation unit 105.

[0055] In addition, in addition to using the cache 103 to store data in the execution unit 102, data can also be stored through registers. Similarly, the types of registers include: registers shared by the tensor calculation unit 104 and the vector calculation unit 105, and exclusive registers of the tensor calculation unit 104 or the vector calculation unit 105. Compared with the cache 103, the capacity of the register is smaller than that of the cache 103, but the data exchange speed is faster than that of the cache 103.

[0056] In addition to the above structures, the processor chip 100 in the present application may also include other structures, and the present application does not make specific limitations thereto.

[0057] The processor chip 100 can be: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a General-purpose computing on graphics processing units (GPGPU), a Domain Specific Architecture (DSA), etc.

[0058] In practical applications, the computational graph contains upstream and downstream operators with strong dependencies. To reduce data movement during the calculation of upstream and downstream operators, related technologies use traditional fusion strategies to fuse upstream and downstream operators. In this way, when the upstream operator is executed, the intermediate result obtained by the upstream operator calculation is saved in the cache, instead of writing the intermediate result back to the global memory. The downstream operator directly obtains the intermediate result from the cache to perform subsequent calculations, thereby improving the performance of the operator.

[0059] For example, refer to Figure 2 , operators A, B, and C are upstream and downstream operators with strong dependencies. Fuse operators A, B, and C to obtain the fused operator (ABC). When executing the fused operator (ABC), first execute operator A to obtain the first intermediate result, and save the first intermediate result in the cache. Read the first intermediate result from the cache, and based on the first intermediate result, execute operator B to obtain the second intermediate result. Save the second intermediate result in the cache. Read the second intermediate result from the cache, and based on the second intermediate result, execute operator C to obtain the calculation result of the fused operator (ABC).

[0060] However, in addition to the above-mentioned upstream and downstream operators with strong dependencies, the computational graph also contains fork operators with no dependencies. For example, refer to Figure 3 , the computational graph includes operators A, B, C, and D, where operators B and D are fork operators with no dependencies.

[0061] In this case, traditional fusion strategies are difficult to fuse fork operators well, thus affecting the performance of fork operators.

[0062] In view of this, based on the Figure 1 system architecture diagram shown, this application provides a process for a fork operator fusion method, as shown in Figure 4 . The process of this method is executed by the processor chip and includes the following steps:

[0063] Step 401: Obtain multiple fork operators with the same input data.

[0064] Specifically, the multiple fork operators have no dependency relationship and can be executed concurrently.

[0065] In some embodiments, the calculation results of the multiple fork operators serve as the input of a subsequent operator. For example, Figure 3 in the shown computation graph, the calculation results of operator B and operator D serve as the input of operator C.

[0066] In some embodiments, the calculation result of each fork operator serves as the input of a subsequent operator. For example, Figure 5 in the shown computation graph, the calculation result of operator B serves as the input of operator E, and the calculation result of operator D serves as the input of operator F.

[0067] It should be noted that the form of the operators associated with the subsequent fork operators is not limited to the above several types and can be adjusted according to the actual situation. In this regard, the present application does not make specific limitations.

[0068] Step 402: Perform operator fusion on the multiple fork operators to obtain a fused operator. The execution process of the fused operator includes: loading the input data into the cache; based on the input data in the cache, concurrently executing the multiple fork operators.

[0069] Specifically, by defining the execution processes of the multiple fork operators in a single kernel function, the fusion of the multiple fork operators is achieved. The execution processes of the multiple fork operators defined in this kernel function are also the execution processes of the fused operator. That is, when executing the fused operator, the input data is loaded from the video memory into the cache allocated for the fused operator. When concurrently executing the multiple fork operators in the fused operator, the input data can be read from the cache for calculation.

[0070] After obtaining the calculation results of the multiple fork operators, the obtained multiple calculation results can be merged to obtain a merged result, and then the merged result can be saved to the video memory so that subsequent operators can directly obtain the merged result from the video memory to perform subsequent calculations. Alternatively, the obtained multiple calculation results can be saved to the video memory, and subsequent operators can obtain the multiple calculation results from the video memory for merged calculation to obtain a merged result. In this regard, the present application does not make specific limitations.

[0071] In the embodiments of the present application, multiple forking operators are fused into a fused operator, so that when the fused operator is executed, multiple forking operators are executed concurrently. Moreover, for the same input data corresponding to multiple forking operators, the input data is loaded from the video memory to the cache through one loading operation, instead of loading the input data from the video memory when each forking operator is executed, thus avoiding repeated loading of the input data, greatly improving the performance bandwidth, and making full use of the hardware computing power. Secondly, since multiple forking operators are executed concurrently during the execution of the fused operator, the calculation time of multiple forking operators is shortened to the calculation time of the forking operator with the longest time consumption, thereby greatly shortening the running time and improving the running efficiency and computational graph performance.

[0072] For the first forking operator executed in the tensor computing unit and the second forking operator executed in the vector computing unit. In the related art, that is, before operator fusion, refer to Figure 6 , the first forking operator corresponds to a first kernel function, and the second forking operator corresponds to a second kernel function.

[0073] When the first forking operator is executed, the input data is loaded from the video memory to the cache allocated for the first kernel function. The tensor computing unit reads the input data from the cache and executes the first forking operator based on the input data. At this time, the vector computing unit is in an idle state.

[0074] When the second forking operator is executed, the input data is loaded from the video memory to the cache allocated for the second kernel function. The vector computing unit reads the input data from the cache and executes the second forking operator based on the input data. At this time, the tensor computing unit is in an idle state.

[0075] It can be seen that before the forking operator fusion, it is necessary to repeatedly load the input data from the video memory to the cache, resulting in a waste of hardware resources. In addition, when calculating the first forking operator, the vector computing unit is in an idle state, that is, the computing power of the vector computing unit is wasted; similarly, when calculating the second forking operator, the tensor computing unit is in an idle state, that is, the computing power of the tensor computing unit is wasted.

[0076] To solve the above problems, the present application proposes to adopt the coroutine programming mode to fuse the first forking operator and the second forking operator to obtain a fused operator. The fused operator corresponds to a kernel function. When the fused operator is executed, the input data is loaded into the cache allocated for the fused operator, and this cache is a shared cache for the tensor computing unit and the vector computing unit.

[0077] The input data is fetched from the shared cache by the tensor computing unit, and the first forking operator is executed based on the input data; the input data is fetched from the shared cache by the vector computing unit, and the second forking operator is executed based on the input data; the first forking operator and the second forking operator are executed concurrently. The computing time consumed by the first forking operator and the second forking operator is the computing time consumed by the forking operator with the longest computing time among the first forking operator and the second forking operator.

[0078] For example, refer to Figure 7 , the computing graph includes operator A, operator B, operator C, and operator D, where operator B and operator D are forking operators. Operator B is an operator executed by the tensor computing unit, and operator D is an operator executed by the vector computing unit.

[0079] Before the fusion of the forking operators, each of operator B and operator D corresponds to a kernel function.

[0080] When executing operator B, the computing result of operator A is loaded from the video memory into the cache allocated for operator B. The tensor computing unit reads the computing result of operator A from the cache and executes operator B based on the computing result of operator A. At this time, the vector computing unit is idle.

[0081] When executing operator D, the computing result of operator A is loaded from the video memory into the cache allocated for operator D. The vector computing unit reads the computing result of operator A from the cache and executes operator D based on the input data. At this time, the tensor computing unit is idle.

[0082] In this application, the coroutine programming mode is adopted to fuse operator B and operator D into a fused operator (BD), and the fused operator (BD) corresponds to a fused kernel function. When executing the fused operator (BD), the computing result of operator A is loaded from the video memory into the cache allocated for the fused operator (BD), and this cache is the shared cache of the tensor computing unit and the vector computing unit.

[0083] The tensor computing unit reads the computing result of operator A from the shared cache and executes operator B based on the computing result of operator A to obtain the computing result of operator B. The vector computing unit reads the computing result of operator A from the shared cache and executes operator D based on the computing result of operator A to obtain the computing result of operator D. The computing results of operator B and operator D are saved in the video memory.

[0084] Subsequently, operator C is executed based on the computing results of operator B and operator D. Specifically, when executing operator C, the computing results of operator B and operator D are loaded from the video memory into the cache. Then the computing results of operator B and operator D are read from the cache to execute operator C to obtain the computing result of operator C.

[0085] In the embodiments of the present application, the coroutine programming mode is adopted to fuse the first forking operator executed by the tensor computing unit and the second forking operator executed by the vector computing unit into a fused operator, so as to implement the concurrent execution of the first forking operator and the second forking operator. Moreover, when executing the fused operator, there is no need to repeatedly load the input data, which greatly improves the performance bandwidth and makes full use of the hardware computing power. Secondly, the tensor computing unit and the vector computing unit concurrently execute the first forking operator and the second forking operator, so that the computing power of the tensor computing unit and the vector computing unit is fully utilized, thereby greatly shortening the running time and improving the running efficiency.

[0086] In some embodiments, the multiple forking operators include: a first forking operator executed by the tensor computing unit and multiple second forking operators executed by the vector computing unit.

[0087] In this scenario, when executing the first forking operator or each second forking operator, it is necessary to load the input data from the video memory into the cache, that is, it is necessary to repeatedly load the input data from the video memory into the cache, resulting in a waste of hardware resources. In addition, when calculating the first forking operator, the vector computing unit is in an idle state, that is, the computing power of the vector computing unit is wasted; similarly, when calculating multiple second forking operators, the tensor computing unit is in an idle state, that is, the computing power of the tensor computing unit is wasted.

[0088] To solve the above problems, based on the multi-pack cooperation (Cwarp) function of the vector computing unit, the present application fuses multiple second forking operators to obtain a preliminary fusion result and realizes the concurrency of multiple second forking operators. The coroutine programming mode is adopted to fuse the first forking operator and the preliminary fusion result to obtain a fused operator, realizing the concurrency of the first forking operator and multiple second forking operators.

[0089] Specifically, the fused operator corresponds to a kernel function. When executing this fused operator, the input data is loaded into the cache allocated for the fused operator, and this cache is a shared cache of the tensor computing unit and the vector computing unit.

[0090] The tensor computing unit obtains the input data from the shared cache and executes the first forking operator based on the input data; and, the vector computing unit obtains the input data from the shared cache and concurrently executes multiple second forking operators based on the input data. The first forking operator and multiple second forking operators are executed concurrently. The calculation time of the first forking operator and multiple second forking operators is the calculation time of the forking operator with the longest time-consuming among the first forking operator and multiple second forking operators.

[0091] For example, see Figure 8, the computational graph includes operator A, operator B, operator C, operator D, operator E, and operator F. Among them, operator B, operator D, operator E, and operator F are forking operators. Operator B is an operator executed in the tensor computing unit, and operator D, operator E, and operator F are operators executed in the vector computing unit.

[0092] Use the multi-pack cooperation function of the vector computing unit to fuse operator D, operator E, and operator F to obtain a preliminary fusion result. Adopt the coroutine programming mode to fuse operator B and the preliminary fusion result to obtain a fused operator (BDEF), and the fused operator (BDEF) corresponds to a fused kernel function.

[0093] When executing the fused operator (BDEF), load the calculation result of operator A from the video memory into the cache allocated for the fused operator (BDEF), and this cache is a shared cache for the tensor computing unit and the vector computing unit.

[0094] The tensor computing unit reads the calculation result of operator A from the shared cache and executes operator B based on the calculation result of operator A to obtain the calculation result of operator B.

[0095] The vector computing unit reads the calculation result of operator A from the shared cache and concurrently executes operator D, operator E, and operator F based on the calculation result of operator A to obtain the respective calculation results of operator D, operator E, and operator F. Save the respective calculation results of operator B, operator D, operator E, and operator F in the video memory.

[0096] Subsequently, continue to execute operator C based on the respective calculation results of operator B, operator D, operator E, and operator F. Specifically, when executing operator C, load the respective calculation results of operator B, operator D, operator E, and operator F into the cache. Then read the respective calculation results of operator B, operator D, operator E, and operator F from the cache to execute operator C to obtain the calculation result of operator C.

[0097] In the embodiments of this application, use the multi-pack cooperation function of the vector computing unit to fuse multiple second forking operators to achieve the concurrency of multiple second forking operators. Adopt the coroutine programming mode to fuse the forking operator executed by the tensor computing unit and the forking operator executed by the vector computing unit to obtain a fused operator, achieving the concurrency of the first forking operator and multiple second forking operators. In this way, the calculation time of multiple forking operators is shortened to the calculation time of the forking operator with the longest time-consuming, thus greatly shortening the running time and improving the running efficiency. Secondly, the tensor computing unit and the vector computing unit concurrently execute the first forking operator and multiple second forking operators, which makes the computing power of the tensor computing unit and the vector computing unit fully utilized and also improves the performance of the computational graph.

[0098] In some embodiments, multiple forking operators are operators executed by a vector computing unit. In this scenario, when executing multiple forking operators, it is necessary to load input data from the video memory into the cache, that is, it is necessary to repeatedly load input data from the video memory into the cache, resulting in a waste of hardware resources.

[0099] To solve the above problems, based on the multi-pack cooperation function of the vector computing unit, this application fuses multiple forking operators to obtain a fused operator, realizing the concurrency of multiple forking operators.

[0100] Specifically, the fused operator corresponds to a kernel function. When executing this fused operator, load the input data into the cache allocated for the fused operator, and this cache is a shared cache of the tensor computing unit and the vector computing unit; or, an exclusive cache of the vector computing unit.

[0101] The vector computing unit obtains the input data from the shared cache or the exclusive cache, and concurrently executes multiple forking operators based on the input data, realizing the concurrent execution of multiple forking operators. The computing time consumption of multiple forking operators is the computing time consumption of the forking operator with the longest time consumption among multiple forking operators.

[0102] For example, see Figure 9 , the computation graph includes operator A, operator C, operator D, operator E, and operator F, where operator D, operator E, and operator F are forking operators executed by the vector computing unit.

[0103] Based on the multi-pack cooperation function of the vector computing unit, fuse operator D, operator E, and operator F into a fused operator (DEF), and the fused operator (DEF) corresponds to a fused kernel function.

[0104] When executing the fused operator (DEF), load the calculation result of operator A from the video memory into the cache allocated for the fused operator (DEF), and this cache is a shared cache of the tensor computing unit and the vector computing unit.

[0105] The vector computing unit reads the calculation result of operator A from the shared cache, and concurrently executes operator D, operator E, and operator F based on the calculation result of operator A, obtaining the respective calculation results of operator D, operator E, and operator F. Save the respective calculation results of operator D, operator E, and operator F in the video memory.

[0106] Subsequently, continue to execute operator C based on the respective calculation results of operator D, operator E, and operator F. Specifically, when executing operator C, load the respective calculation results of operator D, operator E, and operator F into the cache. Then read the respective calculation results of operator D, operator E, and operator F from the cache to execute operator C, obtaining the calculation result of operator C.

[0107] In the embodiments of the present application, multiple forked operators executed by a vector calculation unit are fused, enabling the concurrency of multiple forked operators. In this way, the calculation time of multiple forked operators is shortened to the calculation time of the forked operator with the longest time consumption, thus greatly shortening the running time and improving the running efficiency. Secondly, the input data is loaded from the video memory to the cache through a single load operation, eliminating the need to load the input data from the video memory when executing each forked operator, avoiding repeated loading of the input data, greatly enhancing the performance bandwidth, and making full use of the hardware computing power.

[0108] Based on the same technical concept, the embodiments of the present application provide a structural schematic diagram of a forked operator fusion device, as Figure 10 shown. The forked operator fusion device 1000 includes:

[0109] An acquisition module 1001, configured to acquire multiple forked operators with the same input data;

[0110] A fusion module 1002, configured to fuse the multiple forked operators to obtain a fused operator. The execution process of the fused operator includes: loading the input data into the cache; and concurrently executing the multiple forked operators based on the input data in the cache.

[0111] Optionally, the multiple forked operators include: a first forked operator executed by a tensor calculation unit and a second forked operator executed by a vector calculation unit;

[0112] The fusion module 1002 is specifically configured to:

[0113] Adopt a coroutine programming mode to fuse the first forked operator and the second forked operator to obtain a fused operator.

[0114] Optionally, the cache is: a shared cache of the tensor calculation unit and the vector calculation unit;

[0115] The fusion module 1002 is specifically configured to:

[0116] Obtain the input data from the shared cache through the tensor calculation unit and execute the first forked operator based on the input data;

[0117] Obtain the input data from the shared cache through the vector calculation unit and execute the second forked operator based on the input data.

[0118] Optionally, the multiple forked operators include: a first forked operator executed by a tensor calculation unit and multiple second forked operators executed by a vector calculation unit;

[0119] The fusion module 1002 is specifically configured to:

[0120] Based on the multi-pack cooperation function of the vector calculation unit, fuse the multiple second bifurcation operators to obtain a preliminary fusion result;

[0121] Adopt a coroutine programming mode to fuse the first bifurcation operator and the preliminary fusion result to obtain a fusion operator.

[0122] Optionally, the cache is: a shared cache of the tensor calculation unit and the vector calculation unit;

[0123] The fusion module 1002 is specifically configured to:

[0124] Obtain the input data from the shared cache through the tensor calculation unit, and execute the first bifurcation operator based on the input data;

[0125] Obtain the input data from the shared cache through the vector calculation unit, and concurrently execute the multiple second bifurcation operators based on the input data.

[0126] Optionally, the tensor calculation unit and the vector calculation unit are heterogeneous computing cores.

[0127] Optionally, the multiple bifurcation operators are operators executed on the vector calculation unit;

[0128] The fusion module 1002 is specifically configured to:

[0129] Based on the multi-pack cooperation function of the vector calculation unit, fuse the multiple bifurcation operators to obtain a fusion operator.

[0130] Optionally, the cache is: a shared cache of the tensor calculation unit and the vector calculation unit; or, an exclusive cache of the vector calculation unit;

[0131] The fusion module 1002 is specifically configured to:

[0132] Obtain the input data from the shared cache or the exclusive cache through the vector calculation unit, and concurrently execute the multiple bifurcation operators based on the input data.

[0133] In the embodiments of the present application, multiple fork operators are fused into a fused operator, such that when the fused operator is executed, multiple fork operators are executed concurrently. Moreover, for the same input data corresponding to the multiple fork operators, the input data is loaded from the video memory to the cache through one loading operation, instead of loading the input data from the video memory when each fork operator is executed, thus avoiding repeated loading of the input data, greatly improving the performance bandwidth, and making full use of the hardware computing power. Secondly, since multiple fork operators are executed concurrently during the execution of the fused operator, the calculation time of the multiple fork operators is shortened to the calculation time of the fork operator with the longest time consumption, thereby greatly shortening the running time and improving the running efficiency and the computational graph performance.

[0134] Based on the same technical concept, the embodiments of the present application provide a computer device, as Figure 11 shown, including at least one processor chip 100 and a memory 1101 connected to the at least one processor chip. In the embodiments of the present application, the specific connection medium between the processor chip 100 and the memory 1101 is not limited. Figure 11 Taking the example that the processor chip 100 and the memory 1101 are connected through a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0135] In the embodiments of the present application, the memory 1101 stores instructions executable by the at least one processor chip 100. By executing the instructions stored in the memory 1101, the at least one processor chip 100 can execute the steps of the above-mentioned fork operator fusion method.

[0136] Among them, the processor chip 100 is the control center of the computer device, and can connect various parts of the computer device by using various interfaces and lines. By running or executing the instructions stored in the memory 1101 and calling the data stored in the memory 1101, the fork operator fusion can be realized. Optionally, the processor chip 100 may include one or more processing units. The processor chip 100 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor chip 100. In some embodiments, the processor chip 100 and the memory 1101 may be implemented on the same chip, and in some embodiments, they may also be separately implemented on independent chips.

[0137] The processor chip 100 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0138] As a non-volatile computer-readable storage medium, the memory 1101 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1101 can include at least one type of storage medium, for example, it can include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disc, and so on. The memory 1101 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 1101 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0139] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is caused to execute the steps of the above-mentioned fork operator fusion method.

[0140] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above-mentioned fork operator fusion method.

[0141] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0142] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer device or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0144] These computer program instructions can also be loaded onto a computer device or other programmable data processing devices, such that a series of operation steps are executed on the computer device or other programmable devices to generate a process implemented by the computer device. Thus, the instructions executed on the computer device or other programmable devices provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0145] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0146] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.

Claims

1. A method for fusing bifurcation operators, characterized in that Including: Obtain a plurality of forking operators, where the plurality of forking operators have no dependency relationship and the input data is the same; The plurality of forking operators include: a first forking operator executed on a tensor computing unit and a second forking operator executed on a vector computing unit; Adopt a coroutine programming mode to fuse the plurality of forking operators to obtain a fused operator. The execution process of the fused operator includes: loading the input data from the video memory to the cache allocated for the fused operator through one loading operation; the cache is: the shared cache of the tensor computing unit and the vector computing unit; Based on the input data in the cache, concurrently execute the first forking operator and the second forking operator through the tensor computing unit and the vector computing unit.

2. The method according to claim 1, wherein There are multiple second forking operators; The adopting a coroutine programming mode to fuse the plurality of forking operators to obtain a fused operator includes: Based on the multi-pack cooperation function of the vector computing unit, fuse multiple second forking operators to obtain a preliminary fusion result; Adopt a coroutine programming mode to fuse the first forking operator and the preliminary fusion result to obtain a fused operator.

3. The method according to claim 2, wherein The based on the input data in the cache, concurrently execute the first forking operator and the second forking operator through the tensor computing unit and the vector computing unit includes: Obtain the input data from the shared cache through the tensor computing unit and execute the first forking operator based on the input data; Obtain the input data from the shared cache through the vector computing unit and concurrently execute the multiple second forking operators based on the input data.

4. The method according to any one of claims 1-3, characterized in that, The tensor computing unit and the vector computing unit are heterogeneous computing cores.

5. A forked operator fusion device, characterized in that Including: An obtaining module, configured to obtain a plurality of forking operators, where the plurality of forking operators have no dependency relationship and the input data is the same; The plurality of forking operators include: a first forking operator executed on a tensor computing unit and a second forking operator executed on a vector computing unit; A fusing module, configured to adopt a coroutine programming mode to fuse the plurality of forking operators to obtain a fused operator. The execution process of the fused operator includes: loading the input data from the video memory to the cache allocated for the fused operator through one loading operation; the cache is: the shared cache of the tensor computing unit and the vector computing unit; Based on the input data in the cache, concurrently execute the first forking operator and the second forking operator through the tensor computing unit and the vector computing unit.

6. A computer device, comprising a memory, a processor chip, and a computer program stored in the memory and executable on the processor chip, characterized in that, When the processor chip executes the program, it implements the steps of the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 4.

8. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions that, when executed by a computer device, cause the computer device to perform the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Operator fusion method and device

    CN116089895A

  • Operator fusion method and device, electronic equipment and storage medium

    CN116523023A