Method, device and equipment for determining operator execution provider and storage medium

CN122819503APending Publication Date: 2026-09-25MOORE THREADS TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611265175.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-19
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]相关技术中,深度学习推理框架采用预设执行提供者优先策略,即使算子节点可以同时被多个执行提供者支持,框架也优先选择预设执行提供者来执行该算子节点,例如,采用图形处理器(Graphics Processing Unit,GPU)优先策略,即使算子节点可以同时被图形处理器和中央处理器(Central Processing Unit,CPU)支持,框架也优先选择GPU来执行该算子节点,决策维度单一,预设执行提供者可能被分配一些它并不擅长的算子节点,而其余执行提供者的算力潜能却被闲置;并且,同一算子在不同硬件平台上的最优执行提供者可能不同,同一算子在某一硬件平台上的最优执行提供者为GPU,而同一算子在另一硬件平台上的最优执行提供者为CPU,但相关技术采用固定的GPU优先策略,无法适应硬件平台的实际性能特性

Benefits of technology

[0016]本公开实施例中,通过校准集确定算子节点在各执行提供者中推理校准集的目标耗时,并基于每个算子节点在不同执行提供者中推理校准集的目标耗时,确定每个算子节点的目标执行提供者。由此,可以根据算子节点在各执行提供者中的运行性能,灵活确定算子节点的目标执行提供者,可以充分利用执行提供者的算力潜能,提高算子节点的执行效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819503A_ABST
    Figure CN122819503A_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and device for determining an operator execution provider, an equipment and a storage medium, and relates to the technical field of deep learning. The present disclosure obtains a calibration set, wherein the calibration set comprises dimension information of a plurality of tensors; for each operator node in a computation graph of a model, each dimension information is inferred in different execution providers based on the operator node to obtain a target time consumption of the operator node in the different execution providers for inferring the calibration set; and a target execution provider of each operator node is determined based on the target time consumption of the operator node in the different execution providers for inferring the calibration set. Thus, the target execution provider of the operator node can be flexibly determined according to the running performance of the operator node in each execution provider, the computing power potential of the execution provider can be fully utilized, and the execution effect of the operator node is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of deep learning technology, and in particular to a method, apparatus, device, and storage medium for determining an operator execution provider. Background Technology

[0002] Deep learning inference frameworks are responsible for efficiently deploying trained models to run on target hardware. Logically, a model is represented as a computational graph, which includes operator nodes and edges. Operator nodes represent operators (such as convolution and matrix multiplication), and edges represent the flow of data between operator nodes. Mainstream deep learning inference frameworks employ an execution provider (EP) mechanism, scheduling operator nodes in the model's computational graph to different execution providers (such as CPUs and GPUs) for execution.

[0003] In related technologies, deep learning inference frameworks employ a pre-defined execution provider priority strategy. Even if an operator node can be supported by multiple execution providers simultaneously, the framework prioritizes the pre-defined execution provider to execute that operator node. For example, a Graphics Processing Unit (GPU) priority strategy might be used, where even if an operator node can be supported by both a GPU and a Central Processing Unit (CPU), the framework prioritizes the GPU. This decision-making dimension is singular, and the pre-defined execution provider may be assigned some operator nodes it is not good at, while the computing power potential of other execution providers is idle. Furthermore, the optimal execution provider for the same operator may differ on different hardware platforms. The optimal execution provider for the same operator on one hardware platform might be the GPU, while on another hardware platform, the optimal execution provider might be the CPU. However, related technologies use a fixed GPU priority strategy, which cannot adapt to the actual performance characteristics of hardware platforms. Therefore, the strategy for determining the execution provider of operator nodes in related technologies is rigid and fixed, making it difficult to fully utilize the computing power potential of execution providers and reducing the execution performance of operator nodes. Summary of the Invention

[0004] To address the problems existing in the aforementioned related technologies, this disclosure provides a method, apparatus, device, and storage medium for determining an operator execution provider.

[0005] The first aspect of this disclosure provides a method for determining an operator execution provider, comprising:

[0006] Obtain the calibration set, which includes dimensional information of multiple tensors;

[0007] For each operator node in the computation graph of the model, inference is performed on the information of each dimension based on the operator node in different execution providers to obtain the target time of the operator node inference calibration set in different execution providers;

[0008] Based on the target execution time of each operator node inference calibration set across different execution providers, the target execution provider for each operator node is determined.

[0009] A second aspect of this disclosure provides an apparatus for determining an operator execution provider, comprising:

[0010] The acquisition module is used to acquire the calibration set, which includes the dimensional information of multiple tensors;

[0011] The first inference module is used to infer information of each dimension based on the operator node in the computation graph of the model in different execution providers, so as to obtain the target time of the operator node inference calibration set in different execution providers.

[0012] The determination module is used to determine the target execution provider for each operator node based on the target execution time of the inference calibration set in different execution providers.

[0013] A third aspect of this disclosure provides a computer device, including a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements the method for determining an operator execution provider as described in the first aspect.

[0014] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for determining an operator execution provider described in the first aspect.

[0015] The technical solution provided in this disclosure has the following advantages compared with related technologies:

[0016] In this embodiment of the disclosure, the target execution time of an operator node for inference calibration sets across various execution providers is determined through calibration sets. Based on the target execution time of each operator node for inference calibration sets across different execution providers, the target execution provider for each operator node is determined. Therefore, the target execution provider for an operator node can be flexibly determined according to its performance across different execution providers, fully utilizing the computing power potential of the execution providers and improving the execution effect of the operator node.

[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0019] To more clearly illustrate the technical solutions in the embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of a method for determining an operator execution provider provided in an embodiment of this disclosure;

[0021] Figure 2 This is a flowchart of another method for determining an operator execution provider provided in an embodiment of this disclosure;

[0022] Figure 3 This is a flowchart of another method for determining an operator execution provider provided in this disclosure embodiment;

[0023] Figure 4 This is a flowchart of another method for determining an operator execution provider provided in this disclosure embodiment;

[0024] Figure 5 This is a flowchart of another method for determining an operator execution provider provided in this disclosure embodiment;

[0025] Figure 6 This is a flowchart of another method for determining an operator execution provider provided in this disclosure embodiment;

[0026] Figure 7 This is a flowchart of another method for determining an operator execution provider provided in this disclosure embodiment;

[0027] Figure 8 This is a schematic diagram of the structure of an operator execution provider determination device provided in an embodiment of this disclosure;

[0028] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Detailed Implementation

[0029] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0030] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0034] To better understand the inventive concept of the embodiments of this disclosure, the technical solutions of the embodiments of this disclosure will be described below in conjunction with exemplary embodiments.

[0035] The method for determining the operator execution provider provided in this disclosure can be executed by a computer device. This device can be understood as any device with processing and computing capabilities, including but not limited to electronic devices such as smartphones, laptops, tablets, in-vehicle terminals, wearable devices, digital TVs, desktop computers, smart home devices, etc.

[0036] Figure 1 This is a flowchart illustrating a method for determining an operator execution provider according to an embodiment of this disclosure. This method can be executed by a computer device, such as... Figure 1 As shown, the method for determining the operator execution provider provided in this embodiment includes the following steps:

[0037] Step 110: Obtain the calibration set, which includes dimensional information of multiple tensors.

[0038] In this embodiment of the disclosure, the computer device can acquire a calibration set.

[0039] For example, users can configure a calibration set as needed and then upload the calibration set to a computer device through the interface provided by the deep learning inference framework. The computer device can then receive the calibration set uploaded by the user.

[0040] For example, the data structure of the calibration set can be:

[0041] CalibrationSet = [(Shape_1), (Shape_2), ..., (Shape_n)];

[0042] Where CalibrationSet represents the calibration set; Shape_1 represents the dimension information of the first tensor; Shape_2 represents the dimension information of the second tensor; and Shape_n represents the dimension information of the nth tensor.

[0043] A calibration set can include dimensional information (shape) of multiple tensors, which can also be referred to as tensor shape data. The dimensional information of a tensor can be used to characterize its structural features.

[0044] Step 120: For each operator node in the computation graph of the model, infer the information of each dimension based on the operator node in different execution providers to obtain the target time of the operator node inferring the calibration set in different execution providers.

[0045] In this embodiment of the disclosure, after obtaining the calibration set, the computer device can perform inference on the information of each dimension in the calibration set based on the operator node in different execution providers for each operator node in the computation graph of the model, so as to obtain the target time for the operator node to infer the calibration set in different execution providers.

[0046] An execution provider can be understood as a hardware device or software module capable of data processing. For example, an execution provider may include a central processing unit (CPU), a graphics processing unit (GPU), etc.

[0047] For example, if the computation graph of the model contains operator node A and operator node B, the calibration set includes three dimensions of information: Shape1, Shape2, and Shape3, and the execution providers include CPU and GPU;

[0048] For operator node A, inference is performed on Shape1, Shape2, and Shape3 respectively based on operator node A in the CPU to obtain the target time cost_CPU1 for inference of the calibration set by operator node A in the CPU; inference is performed on Shape1, Shape2, and Shape3 respectively based on operator node A in the GPU to obtain the target time cost_GPU1 for inference of the calibration set by operator node A in the GPU.

[0049] For operator node B, inference is performed on Shape1, Shape2, and Shape3 respectively based on operator node B in the CPU to obtain the target time cost_CPU2 for inference of the calibration set by operator node B in the CPU; inference is performed on Shape1, Shape2, and Shape3 respectively based on operator node B in the GPU to obtain the target time cost_CPU2 for inference of the calibration set by operator node B in the GPU.

[0050] The time taken by an operator node to infer the target calibration set in the execution provider can characterize the efficiency and performance of the execution provider in executing the operator node.

[0051] Step 130: Based on the target execution time of each operator node in the inference calibration set in different execution providers, determine the target execution provider for each operator node.

[0052] In this embodiment of the disclosure, after obtaining the target time of each operator node in the inference calibration set in different execution providers, the computer device can determine the target execution provider of each operator node based on the target time of each operator node in the inference calibration set in different execution providers.

[0053] In this embodiment, the target execution time of an operator node for inference calibration sets across different execution providers is determined through a calibration set. Based on the target execution time of each operator node for inference calibration sets across different execution providers, the target execution provider for each operator node is determined. Therefore, the target execution provider for an operator node can be flexibly determined according to its operating performance across different execution providers, fully utilizing the computing power potential of the execution providers and improving the execution performance of the operator node.

[0054] Figure 2 This is a flowchart illustrating a method for determining an operator execution provider according to an embodiment of this disclosure. This method can be executed by a computer device, such as... Figure 2 As shown, the method for determining the operator execution provider provided in this embodiment includes the following steps:

[0055] Step 210: Obtain the calibration set, which includes dimensional information of multiple tensors.

[0056] Step 220: For each operator node and each dimension information in the computation graph of the model, perform a preset number of inferences on the dimension information based on the operator node in different execution providers, and obtain the total time spent by the operator node in performing a preset number of inferences on the dimension information in different execution providers.

[0057] The number of preset rounds can be set as needed, such as 3 rounds, but there is no limit here.

[0058] For example, if the computation graph of the model contains operator node A and operator node B, the calibration set includes three dimensions of information: Shape1, Shape2, and Shape3, and the execution providers include CPU and GPU;

[0059] For operator node A and Shape1, the total time CPUt11 is obtained by performing a preset number of inference rounds on Shape1 based on operator node A in the CPU; and the total time GPUt11 is obtained by performing a preset number of inference rounds on Shape1 based on operator node A in the GPU.

[0060] For operator node A and Shape2, the total time CPUt12 is obtained by performing a preset number of inference rounds on Shape2 based on operator node A in the CPU; and the total time GPUt12 is obtained by performing a preset number of inference rounds on Shape2 based on operator node A in the GPU.

[0061] For operator node A and Shape3, the total time CPUt13 is obtained by performing a preset number of inference rounds on Shape3 based on operator node A in the CPU; and the total time GPUt13 is obtained by performing a preset number of inference rounds on Shape3 based on operator node A in the GPU.

[0062] For operator node B and Shape1, in the CPU, a preset number of inference rounds are performed on Shape1 based on operator node B, and the total time CPUt21 for operator node B to perform preset number of inference rounds on Shape1 in the CPU is obtained; in the GPU, a preset number of inference rounds are performed on Shape1 based on operator node B, and the total time GPUt21 for operator node B to perform preset number of inference rounds on Shape1 in the GPU is obtained.

[0063] For operator node B and Shape2, the total time CPUt22 is obtained by performing a preset number of inference rounds on Shape2 based on operator node B in the CPU; and the total time GPUt22 is obtained by performing a preset number of inference rounds on Shape2 based on operator node B in the GPU.

[0064] For operator node B and Shape3, the total time CPUt23 is obtained by performing a preset number of inference rounds on Shape3 based on operator node B in the CPU; and the total time GPUt23 is obtained by performing a preset number of inference rounds on Shape3 based on operator node B in the GPU.

[0065] Step 230: For each operator node, each execution provider, and each dimension information, determine the average time taken by the operator node to infer the dimension information in the execution provider based on the ratio of the total time taken by the operator node to infer the dimension information in the preset number of preset rounds.

[0066] In some embodiments, the computer device may determine, for each operator node, each execution provider, and each dimension information, the ratio of the total time spent by the operator node in reasoning about the dimension information in the execution provider for a preset number of rounds to the preset number of rounds, and determine the ratio as the average time spent by the operator node in reasoning about the dimension information in the execution provider.

[0067] Continuing with the example above, if the preset number of rounds is 3:

[0068] For operator node A and Shape1, the total time spent by operator node A in performing 3 rounds of inference on Shape1 in the CPU is CPUt11. The ratio of CPUt11 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node A in inferring Shape1 in the CPU, Cost_CPU11. The total time spent by operator node A in performing 3 rounds of inference on Shape1 in the GPU is GPUt11. The ratio of GPUt11 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node A in inferring Shape1 in the GPU, Cost_GPU11.

[0069] For operator node A and Shape2, the total time spent by operator node A in performing 3 rounds of inference on Shape2 in the CPU is CPUt12. Then, the ratio of CPUt12 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node A in inferring Shape2 in the CPU, Cost_CPU12. The total time spent by operator node A in performing 3 rounds of inference on Shape2 in the GPU is GPUt12. Then, the ratio of GPUt12 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node A in inferring Shape2 in the GPU, Cost_GPU12.

[0070] For operator node A and Shape3, the total time spent by operator node A in performing 3 rounds of inference on Shape3 in the CPU is CPUt13. Then, the ratio of CPUt13 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node A in inferring Shape3 in the CPU, Cost_CPU13. The total time spent by operator node A in performing 3 rounds of inference on Shape3 in the GPU is GPUt13. Then, the ratio of GPUt13 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node A in inferring Shape3 in the GPU, Cost_GPU13.

[0071] For operator node B and Shape1, the total time spent by operator node B in performing 3 rounds of inference on Shape1 in the CPU is CPUt21. Then, the ratio of CPUt21 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node B in inferring Shape1 in the CPU, Cost_CPU21. The total time spent by operator node B in performing 3 rounds of inference on Shape1 in the GPU is GPUt21. Then, the ratio of GPUt21 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node B in inferring Shape1 in the GPU, Cost_GPU21.

[0072] For operator node B and Shape2, the total time spent by operator node B in performing 3 rounds of inference on Shape2 in the CPU is CPUt22. Then, the ratio of CPUt22 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node B in inferring Shape2 in the CPU, Cost_CPU22. The total time spent by operator node B in performing 3 rounds of inference on Shape2 in the GPU is GPUt22. Then, the ratio of GPUt22 to the preset number of rounds 3 is calculated to obtain the average time spent by operator node B in inferring Shape2 in the GPU, Cost_GPU22.

[0073] For operator node B and Shape3, the total time spent by operator node B in performing 3 rounds of inference on Shape3 in the CPU is CPUt23. Then, the ratio of CPUt23 to the preset round 3 is calculated to obtain the average time spent by operator node B in inferring Shape3 in the CPU, Cost_CPU23. The total time spent by operator node B in performing 3 rounds of inference on Shape3 in the GPU is GPUt23. Then, the ratio of GPUt23 to the preset round 3 is calculated to obtain the average time spent by operator node B in inferring Shape3 in the GPU, Cost_GPU23.

[0074] Step 240: For each operator node and each execution provider, determine the target time for the operator node to infer the calibration set in the execution provider based on the average time taken by the operator node to infer each dimension of information in the execution provider.

[0075] In this embodiment of the disclosure, for each operator node and each execution provider, the computer device can determine the target time for the operator node to infer the calibration set in the execution provider based on the average time taken by the operator node to infer information of each dimension in the execution provider.

[0076] Continuing with the example above, for operator node A and the CPU, we can add the average time Cost_CPU11 for operator node A to infer Shape1 in the CPU, the average time Cost_CPU12 for operator node A to infer Shape2 in the CPU, and the average time Cost_CPU13 for operator node A to infer Shape3 in the CPU to obtain the target time Cost_CPU1 for operator node A to infer the calibration set in the CPU:

[0077] Cost_CPU1=(Cost_CPU11+Cost_CPU12+Cost_CPU13);

[0078] For operator node A and the GPU, the average time cost (Cost_GPU11) for operator node A to infer Shape1, the average time cost (Cost_GPU12) for operator node A to infer Shape2, and the average time cost (Cost_GPU13) for operator node A to infer Shape3 in the GPU can be added together to obtain the target time cost (Cost_GPU1) for operator node A to infer the calibration set in the GPU:

[0079] Cost_GPU1=Cost_GPU11+Cost_GPU12+Cost_GPU13;

[0080] For operator node B and the CPU, the average time cost (Cost_CPU21) for operator node B to infer Shape1 in the CPU, the average time cost (Cost_CPU22) for operator node B to infer Shape2 in the CPU, and the average time cost (Cost_CPU23) for operator node B to infer Shape3 in the CPU can be added together to obtain the target time cost (Cost_CPU2) for operator node B to infer the calibration set in the CPU:

[0081] Cost_CPU2=Cost_CPU21+Cost_CPU22+Cost_CPU23;

[0082] For operator node B and the GPU, the average time (Cost_CPU21) for operator node B to infer Shape1 on the GPU, the average time (Cost_CPU22) for operator node B to infer Shape2 on the CPU, and the average time (Cost_CPU23) for operator node B to infer Shape3 on the CPU can be added together to obtain the target time (Cost_CPU2) for operator node B to infer the calibration set on the CPU:

[0083] Cost_CPU2=Cost_CPU21+Cost_CPU22+Cost_CPU23.

[0084] Step 250: Based on the target execution time of each operator node in the inference calibration set in different execution providers, determine the target execution provider for each operator node.

[0085] Therefore, by performing a preset number of rounds of inference on the information of each dimension of the calibration set based on the operator node in different execution providers, the average time consumed by the operator node inferring the information of each dimension of the calibration set in different execution providers can be obtained. Based on the average time consumed, the target time consumed by the operator node inferring the calibration set in different execution providers can be determined. This can improve the accuracy of determining the target time consumed by the operator node inferring the calibration set in different execution providers and improve the accuracy of determining the running performance of the operator node in different execution providers. Furthermore, based on the performance of each operator node running in different execution providers, the target execution provider of each operator node can be flexibly determined, which can make full use of the computing power potential of the execution providers and improve the execution effect of the operator node.

[0086] Figure 3 This is a flowchart illustrating a method for determining an operator execution provider according to an embodiment of this disclosure. This method can be executed by a computer device, such as... Figure 3 As shown, the method for determining the operator execution provider provided in this embodiment includes the following steps:

[0087] Step 310: Obtain the calibration set, which includes the dimensional information of multiple tensors and the weights corresponding to each dimension.

[0088] In this embodiment of the disclosure, the weight corresponding to each dimension can be used to characterize the importance of that dimension. The greater the weight of the dimension, the greater its importance.

[0089] The weight of each dimension is positively correlated with its frequency of occurrence in the actual reasoning scenario. The higher the frequency of occurrence of a dimension in the actual reasoning scenario, the greater its weight.

[0090] Alternatively, the weight of each dimension is positively correlated with its priority. The higher the priority of a dimension, the greater its weight.

[0091] Therefore, the weight of dimensional information can be determined based on the frequency of its appearance in actual reasoning scenarios or the priority of dimensional information, which can improve the accuracy of dimensional information weight.

[0092] Step 320: For each operator node and each dimension information in the computation graph of the model, perform a preset number of inferences on the dimension information based on the operator node in different execution providers, and obtain the total time spent by the operator node in performing a preset number of inferences on the dimension information in different execution providers.

[0093] Step 330: For each operator node and each execution provider, calculate the ratio of the total time spent by the operator node in reasoning the dimension information in the execution provider to the preset number of rounds, and obtain the average time spent by the operator node in reasoning the dimension information in the execution provider.

[0094] Step 340: For each operator node and each execution provider, based on the weight corresponding to each dimension information, the average time spent by the operator node in reasoning for each dimension information in the execution provider is weighted and summed to obtain the target time spent by the operator node in reasoning for the calibration set in the execution provider.

[0095] In this embodiment of the disclosure, for each operator node and each execution provider, the computer device can perform a weighted summation of the average time taken by the operator node to infer each dimension information in the execution provider based on the weight corresponding to each dimension information, so as to obtain the target time taken by the operator node to infer the calibration set in the execution provider.

[0096] For example, if the computation graph of the model contains operator node A and operator node B; the calibration set includes dimension information Shape1, Shape2, Shape3 and the weights W1 corresponding to Shape1, W2 corresponding to Shape2 and W3 corresponding to Shape3; and the execution providers include CPU and GPU.

[0097] For operator node A and the CPU, the average time spent by operator node A inferring Shape1 in the CPU is Cost_CPU11; the average time spent by operator node A inferring Shape2 in the CPU is Cost_CPU12; and the average time spent by operator node A inferring Shape3 in the CPU is Cost_CPU13. Based on the weights corresponding to each dimension of information, the average time spent by operator node A inferring each dimension of information in the CPU is weighted and summed to obtain the target time Cost_CPU1 for operator node A inferring the calibration set in the CPU:

[0098] Cost_CPU1=(Cost_CPU11×W1+Cost_CPU12×W2+Cost_CPU13×W3);

[0099] For operator node A and the GPU, the average time for operator node A to infer Shape1 on the GPU is Cost_GPU11; the average time for operator node A to infer Shape2 on the GPU is Cost_GPU12; and the average time for operator node A to infer Shape3 on the GPU is Cost_GPU13. Based on the weights corresponding to each dimension of information, the average time for operator node A to infer each dimension of information on the GPU is weighted and summed to obtain the target time Cost_GPU1 for operator node A to infer the calibration set on the GPU:

[0100] Cost_GPU1=(Cost_GPU11×W1+Cost_GPU12×W2+Cost_GPU13×W3);

[0101] For operator node B and the CPU, the average time spent by operator node B inferring Shape1 in the CPU is Cost_CPU21; the average time spent by operator node B inferring Shape2 in the CPU is Cost_CPU22; and the average time spent by operator node B inferring Shape3 in the CPU is Cost_CPU23. Based on the weights corresponding to each dimension of information, the average time spent by operator node B inferring each dimension of information in the CPU is weighted and summed to obtain the target time Cost_CPU2 for operator node B inferring the calibration set in the CPU:

[0102] Cost_CPU2=(Cost_CPU21×W1+Cost_CPU22×W2+Cost_CPU23×W3);

[0103] For operator node B and the GPU, the average time for operator node B to infer Shape1 in the GPU is Cost_GPU21; the average time for operator node B to infer Shape2 in the GPU is Cost_GPU22; and the average time for operator node B to infer Shape3 in the GPU is Cost_GPU23. Based on the weights corresponding to each dimension of information, the average time for operator node B to infer each dimension of information in the GPU is weighted and summed to obtain the target time Cost_GPU2 for operator node B to infer the calibration set in the GPU:

[0104] Cost_GPU2=(Cost_GPU21×W1+Cost_GPU22×W2+Cost_GPU23×W3).

[0105] Step 350: Based on the target execution time of each operator node in the inference calibration set in different execution providers, determine the target execution provider for each operator node.

[0106] Therefore, based on the weights corresponding to the information in each dimension of the calibration set, the average time spent by the operator node in reasoning about each dimension of the information in the execution provider can be weighted and summed to obtain the target time spent by the operator node in reasoning about the calibration set in the execution provider. This can further improve the accuracy of determining the target time spent by the operator node in reasoning about the calibration set in each execution provider, and further improve the accuracy of determining the running performance of the operator node in each execution provider. In addition, based on the performance of each execution provider running each operator node, the target execution provider of each operator node can be flexibly determined, which can make full use of the computing power potential of the execution provider and improve the execution effect of the operator node.

[0107] In some embodiments of this disclosure, the above-mentioned weighted summation of the average time spent by the operator node in reasoning about each dimension of information in the execution provider, based on the weights corresponding to each dimension of information, to obtain the target time spent by the operator node in reasoning about the calibration set in the execution provider, may include S11-S12:

[0108] S11. Normalize the weights corresponding to each dimension to obtain the target weights corresponding to each dimension.

[0109] In some embodiments, the sum of the target weights corresponding to each dimension of information is equal to 1.

[0110] In some embodiments, the weights corresponding to each dimension can be summed to obtain a weight sum value; for each weight corresponding to each dimension, the ratio of that weight to the weight sum value is calculated to obtain the target weight corresponding to that dimension.

[0111] For example, the weight_i corresponding to the i-th dimension information.

[0112] For example, the calibration set includes dimensional information Shape1, Shape2, Shape3, and the weights W1 corresponding to Shape1, W2 corresponding to Shape2, and W3 corresponding to Shape3;

[0113] The target weight corresponding to weight W1 is Weight_1 = W1 / (W1 + W2 + W3);

[0114] The target weight corresponding to weight W2 is Weight_2 = W2 / (W1 + W2 + W3);

[0115] The target weight corresponding to weight W3 is Weight_3 = W3 / (W1 + W2 + W3).

[0116] S12. Based on the target weight corresponding to each dimension of information, the average time spent by the operator node in reasoning for each dimension of information in the execution provider is weighted and summed to obtain the target time spent by the operator node in reasoning for the calibration set in the execution provider.

[0117] Therefore, the weights corresponding to each dimension of information can be normalized, and the weights corresponding to multiple dimensions can be adjusted to the same scale, thereby improving the comparability of each weight and thus improving the accuracy of the operator node in determining the target time consumption of the inference calibration set in each execution provider.

[0118] In some embodiments of this disclosure, before performing a preset number of inference rounds based on the operator node for each operator node and each dimension information in each execution provider, the computer device may perform at least one number of inference rounds based on the operator node for each operator node and each dimension information in different execution providers. The inference time consumed in this process is not included in the statistics; that is, the time consumed during the warm-up phase is not included in the statistics.

[0119] Therefore, before performing a preset number of inference rounds on each dimension information based on each operator node in different execution providers, at least one number of inference rounds on the dimension information based on the operator node can be performed on each operator node and each dimension information in different execution providers. This can warm up the execution of operator nodes in the execution providers and eliminate the performance jitter caused by the cold start of the execution providers.

[0120] In some embodiments of the present disclosure, the above determining the target execution provider for each operator node based on the target time consumed by each operator node to infer the calibration set in different execution providers can be performed by a computer device Figure 4 is a flow chart of a method for determining an operator execution provider provided by an embodiment of the present disclosure, as Figure 4 shown, the method for determining an operator execution provider provided by this embodiment includes the following steps:

[0121] Step 410, for each operator node, determine the minimum time consumption of the operator node from the target time consumed by the operator node inferring the calibration set in different execution providers.

[0122] Step 420, determine the target execution provider for each operator node based on the execution provider corresponding to the minimum time consumption of each operator node.

[0123] That is, for each operator node, the target execution provider of the operator node can be determined based on the execution provider corresponding to the minimum time consumption of the operator node.

[0124] Thereby, the target execution provider of the operator node can be determined based on the execution provider corresponding to the minimum time consumption of the operator node, the target execution provider of the operator node can be flexibly determined according to the operation performance of the operator node in each execution provider, the computing power potential of the execution provider can be fully utilized, and the execution effect of the operator node can be improved.

[0125] In some embodiments, for each operator node, the execution provider corresponding to the minimum time consumption of the operator node can be determined as the target execution provider of the operator node.

[0126] Thereby, the execution provider corresponding to the minimum time consumption of the operator node can be determined as the target execution provider of the operator node, the target execution provider of the operator node can be flexibly determined according to the operation performance of the operator node in each execution provider, the computing power potential of the execution provider can be fully utilized, and the execution effect of the operator node can be improved.

[0127] Continuing with the above example, the target time consumed by operator node A inferring the calibration set on CPU is Cost_CPU1, and the target time consumed by operator node A inferring the calibration set on GPU is Cost_GPU1. If Cost_CPU1<Cost_GPU1, then Cost_CPU1 is the minimum time consumption of operator node A, and the CPU corresponding to the minimum time consumption Cost_CPU1 can be determined as the target execution provider of operator node A;

[0128] The target execution time of operator node B inferring the calibration set in the CPU is Cost_CPU2, and the target execution time of operator node B inferring the calibration set in the GPU is Cost_GPU2. If Cost_CPU2 > Cost_GPU2, then Cost_GPU2 is the minimum execution time of operator node B, and the GPU corresponding to the minimum execution time Cost_GPU2 can be determined as the target execution provider of operator node B.

[0129] In some embodiments of this disclosure, the target execution provider for each operator node is determined based on the target time consumption of the inference calibration set for each operator node across different execution providers, and the computer device can execute... Figure 5 This is a flowchart of a method for determining an operator execution provider provided in an embodiment of this disclosure, such as... Figure 5 As shown, the method for determining the operator execution provider provided in this embodiment includes the following steps:

[0130] Step 510: For each operator node, determine the minimum execution time of the operator node from the target execution time of the inference calibration set in different execution providers.

[0131] Step 520: For each operator node, determine the execution provider corresponding to the minimum execution time of the operator node as the current execution provider of the operator node.

[0132] Step 530: According to the node topology order of the computation graph, each operator node in the computation graph is taken as the current operator node in turn, and the current execution provider of the first operator node in the computation graph is determined as the target execution provider of the first operator node.

[0133] The topological order of nodes in a computation graph refers to the execution order of operator nodes in the computation graph.

[0134] After determining the current execution provider of each operator node, the computer device can sequentially take each operator node in the computation graph as the current operator node according to the node topology order of the computation graph, and determine the current execution provider of the first operator node in the computation graph as the target execution provider of the first operator node.

[0135] Step 540: If the current execution provider of the successor operator of the current operator is different from the current execution provider of the current operator, determine the executor switching benefit value of the successor operator based on the absolute value of the difference between the minimum execution time of the successor operator and the first target execution time. If the executor switching benefit value of the successor operator is less than the preset switching penalty threshold, update the current execution provider of the successor operator to the current execution provider of the current operator, and determine the current execution provider of the current operator as the target execution provider of the successor operator. The first target execution time is the target execution time of the successor operator in the inference calibration set of the current execution provider of the current operator.

[0136] In this embodiment of the disclosure, the successor operator of the current operator node can be understood as the execution of the successor operator node depending on the output data of the current operator node, and the input data of the successor operator node being the output data of the current operator node.

[0137] The first target time is the target time for the successor operator node of the current operator node to infer the calibration set in the current execution provider of the current operator node.

[0138] In some embodiments, if the current execution provider of the successor operator of the current operator is different from the current execution provider of the current operator, the absolute value of the difference between the minimum execution time of the successor operator and the first target execution time can be determined, and then the absolute value can be determined as the executor switching benefit value of the successor operator.

[0139] In this embodiment of the disclosure, if the executor switching benefit value of the successor operator of the current operator node is less than a preset switching penalty threshold, it indicates that the successor operator node has low sensitivity to the execution provider. The computer device can update the current execution provider of the successor operator node to the current execution provider of the current operator node, and determine the current execution provider of the current operator node as the target execution provider of the successor operator node. At this time, the target execution provider of the successor operator node is the same as the target execution provider of the current operator node. That is, the target execution provider of the operator node with low sensitivity to the execution provider can be merged with the target execution provider of its predecessor operator node.

[0140] In this embodiment, the preset switching penalty threshold can be set as needed, for example, 0.5ms, and is not limited here. A larger preset switching penalty threshold tends to reduce the number of switching operations by the execution provider.

[0141] Step 550: If the executor switching benefit value of the successor operator of the current operator node is greater than or equal to the preset switching penalty threshold, the current execution provider of the successor operator node is determined as the target execution provider of the successor operator node.

[0142] In this embodiment of the disclosure, if the executor switching benefit value of the successor operator of the current operator node is greater than or equal to the preset switching penalty threshold, it indicates that the successor operator node is highly sensitive to the execution provider, and the current execution provider of the successor operator node can be kept unchanged, and the current execution provider of the successor operator node is determined as the target execution provider of the successor operator node.

[0143] For example, the computation graph includes five operator nodes: Node1, Node2, Node3, Node4, and Node5. The topological order of the nodes in the computation graph is Node1→Node2→Node3→Node4→Node5. Node2 is the successor operator node of Node1, Node3 is the successor operator node of Node2, Node4 is the successor operator node of Node3, and Node5 is the successor operator node of Node4.

[0144] The current execution provider for Node1 is GPU, for Node2 it is CPU, for Node3 it is CPU, for Node4 it is GPU, and for Node5 it is GPU.

[0145] According to the node topology order of the computation graph, Node1, Node2, Node3, Node4, and Node5 are selected as the current operator nodes in sequence.

[0146] First, Node1 is designated as the current operator node. Since Node1 is the first operator node in the computation graph, its current execution provider GPU is determined as the target execution provider for Node1. Because the current execution provider CPU of Node2 is different from the current execution provider GPU of Node1, the executor switching benefit value Δ(Node2) of Node2 is determined based on the absolute value of the difference between Node2's minimum execution time and Node2's first target execution time. Node2's first target execution time is the target execution time of Node2 inference calibration set in Node1's current execution provider GPU. If Δ(Node2) is greater than or equal to a preset switching penalty threshold, then Node2's current execution provider CPU is determined as Node2's target execution provider; that is, Node2's target execution provider is the CPU.

[0147] Then, Node2 is set as the current operator node. Since the current execution provider CPU of Node3 is the same as the current execution provider CPU of Node2, the current execution provider CPU of Node3 is determined as the target execution provider of Node3. That is, the target execution provider of Node3 is the CPU.

[0148] Next, Node3 is set as the current operator node. Node4's current execution provider is a GPU, which is different from Node3's current execution provider CPU. At this point, the executor switching benefit value Δ(Node4) of Node4 is determined based on the absolute value of the difference between Node4's minimum execution time and Node4's first target execution time. Node4's first target execution time is the target execution time of Node4 in the inference calibration set in Node3's current execution provider CPU. If Δ(Node4) is less than a preset switching penalty threshold, then Node4's current execution provider GPU is updated to Node3's current execution provider CPU, which means Node4's current execution provider is updated to CPU. Node3's current execution provider CPU is then determined as Node4's target execution provider, i.e., Node4's target execution provider is CPU.

[0149] Then, Node4 is set as the current operator node. The current execution provider of Node5 is a GPU, which is different from the current execution provider of Node4, CPU. Since the current execution provider of Node5, GPU, is different from the current execution provider of Node4, the executor switching benefit value Δ(Node5) of Node5 is determined based on the absolute value of the difference between the minimum execution time of Node5 and the first target execution time of Node5. The first target execution time of Node5 is the target execution time of Node5 in the inference calibration set in the current execution provider CPU of Node4. If Δ(Node5) is greater than or equal to the preset switching penalty threshold, the current execution provider GPU of Node5 is determined as the target execution provider of Node5, that is, the target execution provider of Node5 is GPU.

[0150] Ultimately, the target execution provider for Node 1 is GPU, the target execution provider for Node 2 is CPU, the target execution provider for Node 3 is CPU, the target execution provider for Node 4 is CPU, and the target execution provider for Node 5 is GPU.

[0151] Step 560: If the current execution provider of the successor operator of the current operator node is the same as the current execution provider of the current operator node, determine the current execution provider of the successor operator node as the target execution provider of the successor operator node.

[0152] In this embodiment of the disclosure, if the current execution provider of the successor operator of the current operator is the same as the current execution provider of the current operator, the current execution provider of the successor operator can be determined as the target execution provider of the successor operator, that is, the execution provider of the successor operator remains unchanged.

[0153] In other words, if the current execution provider of the successor operator is the same as the current execution provider of the current operator, or if the executor switching benefit value of the successor operator of the current operator is greater than or equal to the preset switching penalty threshold, the current execution provider of the successor operator is determined as the target execution provider of the successor operator.

[0154] Therefore, based on the executor switching benefit value of the operator node, the execution provider of the operator node with low sensitivity to the execution provider can be merged with the execution provider of the predecessor operator node of that operator node, while the target execution provider of the operator node with high sensitivity to the execution provider remains unchanged. Alternatively, if the execution provider of the successor operator node is the same as that of the operator node, the execution provider of the successor operator node can be kept unchanged. Under the premise of ensuring the execution effect of the operator node, fragmented execution provider switching between operators can be avoided, which can reduce the number of execution provider switching between operators, reduce the execution provider switching overhead, and improve the execution efficiency of the operator node.

[0155] In some embodiments of this disclosure, after determining the target execution provider for each of the operator nodes, the computer device can call the target execution provider of the target operator node to execute the target operator node for the target operator node in the computation graph of the model that satisfies the first execution condition. The first execution condition may include that the predecessor operator node of the target operator node has been executed.

[0156] The predecessor operator of the target operator node can be understood as the target operator node's execution depending on the output data of the predecessor operator node, and the output data of the predecessor operator node is the input data of the target operator node.

[0157] Therefore, for a target operator node that meets the first execution condition, the target execution provider of the target operator node can be invoked to execute the target operator node, which can make full use of the computing power potential of the target execution provider and improve the execution effect of the operator node.

[0158] In some embodiments of this disclosure, after the target execution provider for each operator node is determined as described above, the computer device can execute... Figure 6 A flowchart of a method for determining an operator execution provider is provided, such as... Figure 6As shown, the method for determining the operator execution provider provided in this embodiment includes the following steps:

[0159] Step 610: Based on the target execution provider of each operator node and the data dependency relationship between different operator nodes, construct at least one operator node with the same target execution provider and continuous data dependency relationship as a subgraph.

[0160] For example, the computation graph of the model includes five operator nodes: Node1, Node2, Node3, Node4, and Node5. Node1's target execution provider is the GPU; Node2's target execution provider is the CPU; Node3's target execution provider is the CPU; Node4's target execution provider is the CPU; and Node5's target execution provider is the GPU. Node1 does not depend on the other operator nodes; Node2 depends on Node1; Node3 depends on Node2; Node4 depends on Node3; and Node5 depends on Node4. Therefore:

[0161] Node1 can be constructed as a subgraph Subgraph_1:GPU[Node1];

[0162] Since Node2, Node3, and Node4 share the same target execution provider and have continuous data dependencies, they can be constructed into a subgraph Subgraph_2: CPU[Node2, Node3, Node4];

[0163] Node5 can be constructed as a subgraph Subgraph_3:GPU[Node5].

[0164] Step 620: For target subgraphs that meet the second execution condition in different subgraphs, call the target execution provider of the target subgraph, and execute the operator nodes in the target subgraph according to the data dependency relationship between different operator nodes in the target subgraph. The second execution condition includes that the predecessor subgraph of the target subgraph has been executed.

[0165] The completion of the predecessor subgraph of the target subgraph can be understood as the completion of the execution of each operator node in the predecessor subgraph of the target subgraph.

[0166] The predecessor subgraph of the target subgraph can be understood as the target subgraph whose execution depends on the output data of the predecessor subgraph, and the output tensor of the predecessor subgraph is the input data of the target subgraph.

[0167] Continuing with the example above, Subgraph_1 is the predecessor subgraph of Subgraph_2; Subgraph_2 is the predecessor subgraph of Subgraph_3.

[0168] If Subgraph_1 has been executed, then Subgraph_2 is the target subgraph that satisfies the second execution condition.

[0169] Therefore, at least one operator node with the same target execution provider and continuous data dependency can be constructed as a subgraph. By executing the operator node according to the subgraph, the computing power potential of the target execution provider can be fully utilized, and the execution effect of the operator node can be improved.

[0170] In some embodiments of this disclosure, after the target execution provider for each operator node is determined as described above, the computer device can execute... Figure 7 A flowchart of a method for determining an operator execution provider is provided, such as... Figure 7 As shown, the method for determining the operator execution provider provided in this embodiment includes the following steps:

[0171] Step 710: Based on the target execution provider of each operator node and the data dependency relationship between different operator nodes, construct at least one operator node with the same target execution provider and continuous data dependency relationship as a subgraph.

[0172] Step 720: For each subgraph, identify the switching boundary between the subgraph and its predecessor subgraph, and insert a virtualized transport operator node in the switching boundary. The switching boundary is used to characterize that the target execution provider of the subgraph is different from the target execution provider of the predecessor subgraph and there is a data transmission direction from the predecessor subgraph to the subgraph.

[0173] Continuing with the example above, if Subgraph_1 has been executed, then Subgraph_2 is the target subgraph that satisfies the second execution condition. The switching boundary 1 between the target subgraph_2 and its predecessor subgraph_1 can be represented as:

[0174] Boundary 1: Subgraph_1 (GPU) → Subgraph_2 (CPU);

[0175] The arrows indicate the data transmission direction from Subgraph_1 to Subgraph_2;

[0176] If Subgraph_2 has been executed, then Subgraph_3 is the target subgraph that satisfies the second execution condition. The switching boundary 2 between the target subgraph_3 and the predecessor subgraph_2 can be represented as:

[0177] Boundary 2: Subgraph_2 (CPU) → Subgraph_3 (GPU);

[0178] The arrows indicate the data transmission direction from Subgraph_2 to Subgraph_3.

[0179] In this embodiment of the disclosure, after determining the switching boundary between the target subgraph and its predecessor subgraph, the computer device can insert a virtualized transmission operator node in the switching boundary.

[0180] The virtualization transport operator node is used to ensure that no physical copy of data occurs during the process of converting the output tensor of the predecessor subgraph on which the target subgraph depends into the input tensor of the target subgraph.

[0181] The virtual transfer operator is also known as the virtual memory copy (Memcpy) operator.

[0182] Step 730: For target subgraphs that meet the second execution condition in different subgraphs, based on the dummy transmission operator node in the switching boundary corresponding to the target subgraph, configure the output tensor of the predecessor subgraph of the target subgraph as the input tensor of the target subgraph, wherein the output tensor and the input tensor share the same target memory space, and the data corresponding to the input tensor and the data corresponding to the output tensor are the same target data.

[0183] In some embodiments, configuring the output tensor of the predecessor subgraph of the target subgraph as the input tensor of the target subgraph based on the dummy transport operator node in the switching boundary corresponding to the target subgraph may include S21-S22:

[0184] S21. Based on the dummy transmission operator node in the switching boundary corresponding to the target subgraph, set the data pointer of the output tensor of the predecessor subgraph of the target subgraph to the data pointer of the input tensor of the target subgraph, so that the data pointer of the input tensor of the target subgraph points to the target memory space corresponding to the output tensor of the predecessor subgraph, wherein the target memory space stores the target data corresponding to the output tensor of the predecessor subgraph.

[0185] The data pointer points to the starting address of the target memory space.

[0186] The data pointer of the input tensor of the target subgraph is the same as the data pointer of the output tensor of the predecessor subgraph of the target subgraph.

[0187] For example, dst.data_ptr = src.data_ptr.

[0188] Where dst is an abbreviation for destination, representing the input tensor of the target subgraph, and dst.data_ptr represents the data pointer of the input tensor of the target subgraph;

[0189] src is short for source, which represents the output tensor of the predecessor subgraph of the target subgraph, and src.data_ptr represents the data pointer of the output tensor of the predecessor subgraph of the target subgraph.

[0190] In some embodiments, after setting the data pointer of the output tensor of the predecessor subgraph to the data pointer of the input tensor of the target subgraph, the computer device may increment the reference count of the target memory space.

[0191] The reference count of the target memory space can be understood as the number of times the target memory space is referenced.

[0192] S22. Send an asynchronous memory prefetch instruction to the target memory space for the first target execution provider. The asynchronous memory prefetch instruction is used to make the target data in the target memory space accessible to the first target execution provider before the target subgraph is executed. The first target execution provider is the target execution provider of the target subgraph.

[0193] For example, if the target execution provider of the predecessor subgraph of the target subgraph is the CPU and the target execution provider of the target subgraph is the GPU, then a first asynchronous memory prefetch (MemPrefetchAsync) instruction for the GPU is sent to the target memory space. This first asynchronous memory prefetch instruction can make the target data in the target memory space in a state that can be accessed by the GPU before the target subgraph is executed, so as to warm up the GPU cache.

[0194] If the target execution provider of the predecessor subgraph of the target subgraph is the GPU and the target execution provider of the target subgraph is the CPU, then a second asynchronous memory prefetch instruction for the CPU is sent to the target memory space. This second asynchronous memory prefetch instruction can make the target data in the target memory space in a state that can be accessed by the CPU before the target subgraph is executed, so as to flush it back to the GPU cache.

[0195] Step 740: Based on the input tensor and the dependencies between the operator nodes in the target subgraph, call the target execution provider of the target subgraph to execute the operator nodes in the target subgraph.

[0196] Therefore, by inserting virtualized transport operator nodes between subgraphs corresponding to different target execution providers, the output tensor of the predecessor subgraph and the input tensor of the target subgraph are associated with the same target memory space. Before the target subgraph is executed, an asynchronous prefetch operation is performed on the target data in the target memory space, targeting the target execution provider. This eliminates the need for physical copying of the output data of the predecessor subgraph, thereby reducing data transfer overhead and memory usage during execution provider switching. It enables fine-grained hybrid scheduling at the execution provider operator level, achieving more precise computational power allocation and improving the execution efficiency of operator nodes.

[0197] Figure 8 This is a schematic diagram of a device for determining an operator execution provider according to an embodiment of this disclosure. This device can be understood as the aforementioned computer equipment or a functional module within the aforementioned computer equipment. Figure 8 As shown, the operator execution provider determination device 800 includes:

[0198] The acquisition module 810 is used to acquire a calibration set, wherein the calibration set includes dimensional information of multiple tensors;

[0199] The first inference module 820 is used to infer information of each dimension based on the operator node in different execution providers for each operator node in the computation graph of the model, so as to obtain the target time of the operator node inferring the calibration set in different execution providers.

[0200] The determination module 830 is used to determine the target execution provider for each operator node based on the target time of the inference calibration set in different execution providers for each operator node.

[0201] Optionally, the first inference module 820 mentioned above includes:

[0202] The inference submodule is used to perform a preset number of inferences on the dimension information based on the operator node in different execution providers for each dimension information, and to obtain the total time spent by the operator node to perform a preset number of inferences on the dimension information in different execution providers.

[0203] The average time determination submodule is used to determine the average time for each execution provider and each dimension information by comparing the total time spent by the operator node in reasoning about the dimension information in the execution provider for a preset number of rounds with the preset number of rounds; and,

[0204] The target time determination submodule is used to determine the target time for the operator node to infer the calibration set in the execution provider based on the average time taken by the operator node to infer information for each dimension in the execution provider.

[0205] Optionally, the calibration set mentioned above may also include weights corresponding to each dimension.

[0206] The above-mentioned target time determination submodule includes:

[0207] The summation unit is used to perform a weighted summation of the average time taken by the operator node to infer each dimension of information in the execution provider, based on the weight corresponding to each dimension of information, to obtain the target time taken by the operator node to infer the calibration set in the execution provider.

[0208] Optionally, the weight corresponding to each of the above-mentioned dimensions is positively correlated with the frequency of the dimension information in the actual reasoning scenario; or, the weight corresponding to each dimension is positively correlated with the priority of the dimension information.

[0209] Optionally, the above summation unit includes:

[0210] The normalization subunit is used to normalize the weights corresponding to each dimension information to obtain the target weights corresponding to each dimension information.

[0211] The summation subunit is used to perform a weighted summation of the average time taken by the operator node to infer each dimension of information in the execution provider, based on the target weight corresponding to each dimension of information, to obtain the target time taken by the operator node to infer the calibration set in the execution provider.

[0212] Optionally, the means for determining the operator execution provider mentioned above further includes:

[0213] The second inference module is used to perform at least one round of inference on each dimension information based on operator nodes in different execution providers before performing a preset number of rounds of inference on the dimension information based on operator nodes in different execution providers.

[0214] Optionally, the determining module 830 mentioned above includes:

[0215] The minimum execution time determination submodule is used to determine the minimum execution time of each operator node from the target execution time of the operator node in the inference calibration set across different execution providers.

[0216] The target execution provider determination submodule is used to determine the target execution provider for each operator node based on the execution provider corresponding to the minimum execution time of each operator node.

[0217] Optionally, the target execution provider determination submodule mentioned above includes:

[0218] The first determining unit is used to determine the execution provider corresponding to the minimum execution time of each operator node as the target execution provider of the operator node.

[0219] Optionally, the target execution provider determination submodule mentioned above includes:

[0220] The second determining unit is used to determine the execution provider corresponding to the minimum execution time of each operator node as the current execution provider of the operator node.

[0221] The third determining unit is used to sequentially take each operator node in the computation graph as the current operator node according to the node topology order of the computation graph, and determine the current execution provider of the first operator node in the computation graph as the target execution provider of the first operator node.

[0222] The fourth determining unit is used to determine the executor switching benefit value of the successor operator node based on the absolute value of the difference between the minimum execution time of the successor operator node and the first target execution time when the current execution provider of the successor operator node is different from the current execution provider of the current operator node. If the executor switching benefit value is less than the preset switching penalty threshold, the current execution provider of the successor operator node is updated to the current execution provider of the current operator node, and the current execution provider of the current operator node is determined as the target execution provider of the successor operator node. The first target execution time is the target execution time of the successor operator node in the inference calibration set of the current execution provider of the current operator node.

[0223] The fifth determining unit is used to determine the current execution provider of the successor operator as the target execution provider of the successor operator when the current execution provider of the successor operator is the same as the current execution provider of the current operator, or when the executor switching benefit value of the successor operator is greater than or equal to a preset switching penalty threshold.

[0224] Optionally, the means for determining the operator execution provider mentioned above further includes:

[0225] The calling module is used to determine the target execution provider for each operator node, and then, for the target operator node in the computation graph that meets the first execution condition, call the target execution provider of the target operator node to execute the target operator node. The first execution condition includes that the predecessor operator node of the target operator node has been executed.

[0226] Optionally, the means for determining the operator execution provider mentioned above further includes:

[0227] The building module is used to construct a subgraph based on the target execution provider of each operator node and the data dependency relationship between different operator nodes, after determining the target execution provider of each operator node. At least one operator node with the same target execution provider and continuous data dependency relationship is constructed as a subgraph.

[0228] The execution module is used to call the target execution provider of the target subgraph for different subgraphs that meet the second execution condition, and execute the operator nodes in the target subgraph according to the data dependency relationship between different operator nodes in the target subgraph. The second execution condition includes that the predecessor subgraph of the target subgraph has been executed.

[0229] Optionally, the means for determining the operator execution provider mentioned above further includes:

[0230] The identification module is used to construct at least one operator node with the same target execution provider and continuous data dependency into a subgraph, and for each subgraph, identify the switching boundary between the subgraph and its predecessor subgraph, and insert a virtualized transmission operator node in the switching boundary. The switching boundary is used to characterize that the target execution provider of the subgraph is different from the target execution provider of the predecessor subgraph and there is a data transmission direction from the predecessor subgraph to the subgraph.

[0231] The above execution module includes:

[0232] The configuration submodule is used to configure the output tensor of the predecessor subgraph of the target subgraph as the input tensor of the target subgraph based on the dummy transport operator node in the switching boundary corresponding to the target subgraph. The output tensor and the input tensor share the same target memory space, and the data corresponding to the input tensor and the data corresponding to the output tensor are the same target data.

[0233] The execution submodule is used to call the target execution provider of the target subgraph based on the input tensor and the dependencies between the operator nodes in the target subgraph, and execute the operator nodes in the target subgraph.

[0234] Optionally, the above configuration submodules include:

[0235] The setting unit is used to set the data pointer of the output tensor of the predecessor subgraph of the target subgraph to the data pointer of the input tensor of the target subgraph based on the dummy transmission operator node in the switching boundary corresponding to the target subgraph, so that the data pointer of the input tensor points to the target memory space corresponding to the output tensor, wherein the target memory space stores the target data corresponding to the output tensor;

[0236] The sending unit is used to send an asynchronous memory prefetch instruction to the target memory space for the first target execution provider. The asynchronous memory prefetch instruction is used to make the target data in the target memory space accessible to the first target execution provider before the target subgraph is executed. The first target execution provider is the target execution provider of the target subgraph.

[0237] The operator execution provider determination device provided in this disclosure can implement the method of any of the above embodiments, and its execution mode and beneficial effects are similar, so they will not be described again here.

[0238] This disclosure also provides a computer device, which includes a processor and a memory, wherein the memory stores a computer program. When the computer program is executed by the processor, it can implement the methods of any of the above embodiments. The execution method and beneficial effects are similar, and will not be described again here.

[0239] The computer device in this disclosure can be understood as any device with processing and computing capabilities, including but not limited to electronic devices such as smartphones, laptops, tablets, in-vehicle terminals, wearable devices, digital TVs, desktop computers, smart home devices, etc.

[0240] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure, such as... Figure 9 As shown, the computer device 900 may include a processor 910 and a memory 920. The memory 920 stores a computer program 921. When the computer program 921 is executed by the processor 910, it can implement the method provided in any of the above embodiments. The execution method and beneficial effects are similar and will not be described again here.

[0241] Of course, for the sake of simplicity, Figure 9 Only some of the components of the computer device 900 relevant to the present invention are shown in this illustration; components such as buses, input / output interfaces, input devices, and output devices are omitted. In addition, the computer device 900 may include any other suitable components depending on the specific application.

[0242] This disclosure provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the methods of any of the above embodiments. The execution method and beneficial effects are similar, and will not be described again here.

[0243] The aforementioned computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0244] The computer program described above can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer device, partially on the user's device, as a standalone software package, partially on the user's computer device and partially on a remote computer device, or entirely on a remote computer device or server.

[0245] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0246] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0247] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for determining an operator execution provider, characterized in that, include: Obtain a calibration set, wherein the calibration set includes dimensional information of multiple tensors; For each operator node in the computation graph of the model, reasoning is performed on the dimensional information based on the operator node in different execution providers to obtain the target time for the operator node to reason the calibration set in different execution providers; The target execution provider for each operator node is determined based on the target time consumption of the calibration set inferenced across different execution providers.

2. The method according to claim 1, characterized in that, The step of reasoning about each dimension information based on the operator node in different execution providers to obtain the target time for the operator node to reason about the calibration set in different execution providers includes: For each of the aforementioned dimension information, the operator node is used to perform a preset number of inference rounds on the dimension information in different execution providers to obtain the total time consumed by the operator node to perform the preset number of inference rounds on the dimension information in different execution providers; For each execution provider and each dimension information, the average time taken by the operator node to infer the dimension information in the execution provider is determined based on the ratio of the total time taken by the operator node to perform the preset number of inference rounds in the execution provider to the preset number of rounds; and, Based on the average time taken by the operator node to infer each dimension information in the execution provider, the target time taken by the operator node to infer the calibration set in the execution provider is determined.

3. The method according to claim 2, characterized in that, The calibration set also includes the weights corresponding to each of the dimensional information; Determining the target time for the operator node to infer the calibration set in the execution provider based on the average time spent by the operator node inferring each dimension information in the execution provider includes: Based on the weight corresponding to each dimension information, the average time taken by the operator node to infer each dimension information in the execution provider is weighted and summed to obtain the target time taken by the operator node to infer the calibration set in the execution provider.

4. The method according to claim 3, characterized in that, The weight corresponding to each dimension is positively correlated with the frequency of occurrence of the dimension in the actual reasoning scenario; Alternatively, the weight corresponding to each of the said dimension information is positively correlated with the priority of the said dimension information.

5. The method according to claim 3, characterized in that, The step of weighted summing of the average time taken by the operator node to infer each dimension information in the execution provider based on the weight corresponding to each dimension information, to obtain the target time taken by the operator node to infer the calibration set in the execution provider, includes: The weights corresponding to each dimension are normalized to obtain the target weights corresponding to each dimension. Based on the target weight corresponding to each dimension information, the average time taken by the operator node to infer each dimension information in the execution provider is weighted and summed to obtain the target time taken by the operator node to infer the calibration set in the execution provider.

6. The method according to claim 2, characterized in that, Before performing a preset number of inference rounds on the dimension information based on the operator node in different execution providers for each dimension information, the method further includes: For each of the aforementioned dimension information, inference is performed on the dimension information at least once per round based on the operator node in different execution providers.

7. The method according to claim 1, characterized in that, The step of determining the target execution provider for each operator node based on the target time consumption of inferring the calibration set across different execution providers includes: For each operator node, the minimum execution time of the operator node is determined from the target execution time of the calibration set inferred from the different execution providers; Based on the execution provider corresponding to the minimum execution time of each operator node, the target execution provider for each operator node is determined.

8. The method according to claim 7, characterized in that, The step of determining the target execution provider for each operator node based on the execution provider corresponding to the minimum execution time of each operator node includes: For each operator node, the execution provider corresponding to the minimum execution time of the operator node is determined as the target execution provider of the operator node.

9. The method according to claim 7, characterized in that, The step of determining the target execution provider for each operator node based on the execution provider corresponding to the minimum execution time of each operator node includes: For each operator node, the execution provider corresponding to the minimum execution time of the operator node is determined as the current execution provider of the operator node; According to the node topology order of the computation graph, each operator node in the computation graph is taken as the current operator node in turn, and the current execution provider of the first operator node in the computation graph is determined as the target execution provider of the first operator node. If the current execution provider of the successor operator of the current operator is different from the current execution provider of the current operator, the executor switching benefit value of the successor operator is determined based on the absolute value of the difference between the minimum execution time of the successor operator and the first target execution time. If the executor switching benefit value is less than a preset switching penalty threshold, the current execution provider of the successor operator is updated to the current execution provider of the current operator, and the current execution provider of the current operator is determined as the target execution provider of the successor operator. The first target execution time is the target execution time of the successor operator inferring the calibration set from the current execution provider of the current operator. If the current execution provider of the successor operator is the same as the current execution provider of the current operator, or if the executor switching benefit value of the successor operator is greater than or equal to the preset switching penalty threshold, the current execution provider of the successor operator shall be determined as the target execution provider of the successor operator.

10. The method according to any one of claims 1-9, characterized in that, After determining the target execution provider for each operator node, the method further includes: For a target operator node in the computation graph that satisfies the first execution condition, the target execution provider of the target operator node is invoked to execute the target operator node, wherein the first execution condition includes that the predecessor operator node of the target operator node has been executed.

11. The method according to any one of claims 1-9, characterized in that, After determining the target execution provider for each operator node, the method further includes: Based on the target execution provider of each operator node and the data dependency relationship between different operator nodes, at least one operator node with the same target execution provider and continuous data dependency relationship is constructed as a subgraph; For target subgraphs that satisfy the second execution condition in different subgraphs, the target execution provider of the target subgraph is invoked, and the operator nodes in the target subgraph are executed according to the data dependency relationship between different operator nodes in the target subgraph. The second execution condition includes that the predecessor subgraph of the target subgraph has been executed.

12. The method according to claim 11, characterized in that, After constructing at least one operator node with the same target execution provider and continuous data dependency into a subgraph, the method further includes: For each subgraph, a switching boundary is identified between the subgraph and its predecessor subgraph, and a virtualized transport operator node is inserted in the switching boundary. The switching boundary is used to characterize that the target execution provider of the subgraph is different from the target execution provider of the predecessor subgraph and there is a data transmission direction from the predecessor subgraph to the subgraph. The target execution provider that invokes the target subgraph executes the operator nodes in the target subgraph according to the data dependencies between different operator nodes in the target subgraph, including: Based on the dummy transport operator node in the switching boundary corresponding to the target subgraph, the output tensor of the predecessor subgraph of the target subgraph is configured as the input tensor of the target subgraph, wherein the output tensor and the input tensor share the same target memory space, and the data corresponding to the input tensor and the data corresponding to the output tensor are the same target data; Based on the input tensor and the dependencies between the operator nodes in the target subgraph, the target execution provider of the target subgraph is invoked to execute the operator nodes in the target subgraph.

13. The method according to claim 12, characterized in that, The step of configuring the output tensor of the predecessor subgraph of the target subgraph as the input tensor of the target subgraph based on the dummy transport operator node in the switching boundary corresponding to the target subgraph includes: Based on the dummy transport operator node in the switching boundary corresponding to the target subgraph, the data pointer of the output tensor of the predecessor subgraph of the target subgraph is set as the data pointer of the input tensor of the target subgraph, so that the data pointer of the input tensor points to the target memory space corresponding to the output tensor, wherein the target memory space stores the target data corresponding to the output tensor; An asynchronous memory prefetch instruction is sent to the target memory space for a first target execution provider, wherein the asynchronous memory prefetch instruction is used to make the target data in the target memory space accessible to the first target execution provider before the target subgraph is executed, and the first target execution provider is the target execution provider of the target subgraph.

14. A device for determining an operator execution provider, characterized in that, include: An acquisition module is used to acquire a calibration set, wherein the calibration set includes dimensional information of multiple tensors; The first inference module is used to infer the dimensional information of each operator node in the computation graph of the model in different execution providers, and obtain the target time of the operator node inferring the calibration set in different execution providers. The determination module is used to determine the target execution provider for each operator node based on the target time consumption of the calibration set inferenced among different execution providers for each operator node.

15. A computer device, characterized in that, include: A memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements the method for determining an operator execution provider as described in any one of claims 1-13.

16. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method for determining the operator execution provider as described in any one of claims 1-13.