Task deployment method and task deployment device

By deploying tasks based on the computing power of the computing unit, the problem of unbalanced load of the computing unit in the computing system is solved, and more efficient computing efficiency and lower data migration overhead are achieved.

CN120045306APending Publication Date: 2025-05-27HUAWEI TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311609712.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When the prior art deploys computing tasks in a computing system, it is easy to cause load imbalance between computing units, extend business computing time, and affect computing efficiency.

Method used

By deploying tasks based on the computing power of the computing unit, the computing tasks are assigned to each computing unit, so that their computing power load matches the computing power, and load balancing is achieved.

Benefits of technology

Effectively balance the load between computing units, improve overall computing efficiency, reduce unnecessary data migration, and reduce data handling overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045306A_ABST
    Figure CN120045306A_ABST
Patent Text Reader

Abstract

The invention provides a task deployment method and a task deployment device. The method comprises the following steps: determining a computational graph to be deployed; according to the computing power of a plurality of computing units corresponding to the computing graph, a plurality of computing tasks in the computing graph are distributed to the plurality of computing units, and the computing power of at least two computing units in the plurality of computing units is different. Moreover, the difference between the ratio of the computing power loads of the plurality of computing units and the ratio of the computing power of the plurality of computing units is equal to 0 or smaller than a first threshold value, and the computing power loads of the computing units are determined according to the sum of the computing power required by the computing tasks corresponding to the computing units. According to the scheme, the calculation tasks are distributed according to the calculation power of the calculation units, the sum of the calculation power needed by the calculation tasks executed by the calculation units is matched with the calculation power of the calculation units, loads among the calculation units can be balanced, and therefore the overall calculation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a task deployment method and a task deployment device. Background Art

[0002] The demand for computing power of computing systems is increasing, and a large number of existing and emerging businesses need to deploy computing tasks on computing systems.

[0003] In the related art, computing tasks are usually deployed equally to each computing unit according to the number of computing units in the computing system. That is, in one deployment, the number of computing tasks processed by each computing unit is equal or approximately equal. Taking the computing unit as a processor as an example, when the server includes three processors, the computing graph including multiple computing tasks is specifically divided into three subsets, and the computing tasks in each subset are equal or approximately equal, and then the three processors respectively execute the computing tasks in the three subsets. Since the processing capabilities of multiple computing units may be different, this deployment scheme of dividing the computing graph equally according to the number of computing units may cause unbalanced loads on each computing unit, prolong the overall computing time of the business, and affect the overall computing efficiency. Summary of the invention

[0004] The present application provides a task deployment method and a task deployment device, which deploy computing tasks based on the computing power of computing units, can balance the load between computing units and improve overall computing efficiency.

[0005] In the first aspect, the present application provides a task deployment method. The method includes: determining a computational graph to be deployed, wherein the computational graph includes multiple computational tasks; allocating the multiple computational tasks to the multiple computational units according to the computational power of the multiple computational units corresponding to the computational graph, wherein the computational power of at least two of the multiple computational units is different, the ratio of the computational power loads of the multiple computational units and the difference between the ratio of the computational power of the multiple computational units are equal to 0 or less than a first threshold, and the computational power load of each computational unit is determined according to the sum of the computational power required for the computational tasks corresponding to the respective computational units.

[0006] In the above scheme, computing tasks are allocated according to the computing power of the computing units, so that the sum of the computing power required for the computing tasks executed by the computing units matches the computing power of the computing units, thereby balancing the load between multiple computing units and improving the overall computing efficiency.

[0007] In a possible implementation of the first aspect, allocating the multiple computing tasks to the multiple computing units according to the computing power of the multiple computing units includes: determining the number N of divisions of the computing graph according to the computing power of the multiple computing units; dividing the computing graph into N subsets according to N, each subset in the N subsets including one or more of the multiple computing tasks, the difference in computing power required by any two subsets in the N subsets is equal to 0 or less than a second threshold, and the computing power required by each subset is the sum of the computing power required by the computing tasks in the each subset; allocating the N subsets to the multiple computing units according to the ratio of the computing power of the multiple computing units.

[0008] In the above scheme, the concept of meta-computing power is introduced. According to the total meta-computing power N of multiple processors, the computational graph is divided into N subsets according to the computing power. Then, each subset is allocated according to the ratio of the computing power of multiple computing units. This can evenly distribute computing tasks to the computing units to achieve the purpose of load balancing.

[0009] In a possible implementation of the first aspect, determining the number N of divisions of the computational graph according to the computing power of the multiple computing units includes: determining a meta-computing power according to a common divisor of the computing power of the multiple computing units, the meta-computing power representing the unit computing power of the multiple computing units; calculating the quotient of the computing power of the multiple computing units and the meta-computing power to obtain the number of meta-computing powers of the multiple computing units; and calculating the sum of the number of meta-computing powers of the multiple computing units to obtain N.

[0010] In a possible implementation of the first aspect, before determining the element computing power according to the common divisor of the computing power of the multiple computing units, the method further includes: updating the computing power of the multiple computing units according to the time when the multiple computing units perform the same computing task.

[0011] In a possible implementation of the first aspect, dividing the computation graph into N subsets according to N includes: dividing the computation graph with the goal of minimizing the data transmission amount of the subsets to obtain the N subsets, wherein the data transmission amount of the subsets is the sum of the data transmission amounts of the computing tasks in the subsets.

[0012] In a possible implementation of the first aspect, before allocating the multiple computing tasks to the multiple computing units based on the computing power of the multiple computing units, the method also includes: removing target computing tasks from the multiple computing tasks, wherein the target computing tasks include predetermined tasks to be performed by each of the computing units.

[0013] In a possible implementation of the first aspect, the computing unit includes a server, a processor unit in the server, a processor in the server, or a core in the processor.

[0014] In a second aspect, the present application also provides a task deployment device, which includes: an analysis module and a deployment module.

[0015] Among them, the analysis module is used to determine the computing graph to be deployed and the computing power of multiple computing units corresponding to the computing graph, and the computing graph includes multiple computing tasks.

[0016] Among them, the deployment module is used to allocate the multiple computing tasks to the multiple computing units according to the computing power of the multiple computing units corresponding to the computing graph, wherein the computing power of at least two of the multiple computing units is different, the ratio of the computing power loads of the multiple computing units and the difference between the ratio of the computing power of the multiple computing units are equal to 0 or less than a first threshold, and the computing power load of each computing unit is determined according to the sum of the computing power required for the computing tasks corresponding to each computing unit.

[0017] In a possible implementation of the second aspect, the deployment module is specifically used to: determine the number N of divisions of the computational graph according to the computing power of the multiple computing units; divide the computational graph into N subsets according to the N, each subset in the N subsets includes one or more of the multiple computing tasks, the difference in computing power required for any two subsets in the N subsets is equal to 0 or less than a second threshold, and the computing power required for each subset is the sum of the computing power required for the computing tasks in the each subset; and allocate the N subsets to the multiple computing units.

[0018] In a possible implementation of the second aspect, the deployment module is specifically used to: determine the meta-computing power according to the common divisor of the computing power of the multiple computing units, the meta-computing power representing the unit computing power of the multiple computing units; calculate the quotient of the computing power of the multiple computing units and the meta-computing power to obtain the number of meta-computing powers of the multiple computing units; calculate the sum of the number of meta-computing powers of the multiple computing units to obtain N.

[0019] In a possible implementation of the second aspect, the deployment module is specifically used to: before determining the meta-computing power according to the common divisor of the computing powers of the multiple computing units, update the computing powers of the multiple computing units according to the time when the multiple computing units perform the same computing task.

[0020] In a possible implementation of the second aspect, the deployment module is specifically used to: divide the computational graph with the goal of minimizing the data transmission volume of the subsets to obtain the N subsets, and the data transmission volume of the subsets is the sum of the data transmission volumes of the computing tasks in the subsets.

[0021] In a possible implementation of the second aspect, the deployment module is specifically used to: before allocating the multiple computing tasks to the multiple computing units based on the computing power of the multiple computing units, remove the target computing tasks from the multiple computing tasks, and the target computing tasks include predetermined tasks to be performed by each of the computing units.

[0022] In a possible implementation of the second aspect, the computing unit includes a server, a processor unit in the server, a processor in the server, or a core in the processor.

[0023] In a third aspect, the present application further provides a computing device. The computing device includes: a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the task deployment method provided in the first aspect or any possible implementation method of the first aspect.

[0024] In a fourth aspect, the present application also provides a computer-readable storage medium, which stores instructions, and when the computer-readable storage medium is executed on a computer, enables the computer to implement the task deployment method provided by the first aspect or any possible implementation method of the first aspect.

[0025] In a fifth aspect, the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to implement the task deployment method provided in the first aspect or any possible implementation manner of the first aspect.

[0026] Any of the above-mentioned devices, computer storage media or computer program products are used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding schemes in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flowchart of a task deployment method provided in an embodiment of the present application;

[0028] Figure 2 This is a method provided by the embodiment of the present application. Figure 1 A detailed schematic diagram of the task deployment method shown;

[0029] Figure 3a This is a schematic diagram of a task deployment for a central processing unit (CPU) and a graphics processing unit (GPU) provided in an embodiment of the present application;

[0030] Figure 3b is a schematic diagram of a calculation graph provided in an embodiment of the present application;

[0031] Figure 3cIt is a schematic diagram of the order in which each computing unit performs computing tasks obtained by pre-scheduling provided in an embodiment of the present application;

[0032] Figure 4 This embodiment of the present application provides a method based on Figure 1 A flowchart of a hierarchical deployment method of the illustrated method;

[0033] Figure 5 This is a schematic diagram of a task deployment for a central processing unit (CPU) and a graphics processing unit (NPU) provided in an embodiment of the present application.

[0034] Figure 6 It is a structural diagram of a task deployment device provided in an embodiment of the present application;

[0035] Figure 7 It is a structural diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0037] Before introducing the embodiments of the present application, the nouns appearing in the embodiments of the present application are introduced below.

[0038] A computing unit refers to a device in a computing system that performs specific computing tasks. According to the hierarchy of computing units in the computing system, computing can specifically include computing devices, processor units, processors, or processor cores. Among them, the processor unit can be determined based on the memory area accessed by the processor. The processors accessing the same memory area are called a processor unit, and a processor unit may include one or more processors. It can be understood that a computing system usually includes multiple computing units. The computing system can specifically be a distributed computing cluster including multiple computing devices, or a computing device including multiple processors. Computing devices may include servers and terminal devices, and terminal devices may include laptops, smart phones, etc.

[0039] The processor of the computing device may include multiple processors of the same structure or multiple processors of heterogeneity, each processor including at least one core. The processor may specifically include a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU). It can be understood that isomorphic and heterogeneous refer to the same or different structures of the processors. For example, any two of the CPU, GPU, and NPU are heterogeneous processors, multiple CPUs with different structures, or multiple GPUs with different structures, or multiple NPUs with different structures also belong to heterogeneous processors, and multiple CPUs with the same structure, or multiple GPUs with the same structure, or multiple NPUs with the same structure belong to isomorphic processors.

[0040] Computing power refers to the computing power of a computing unit. Computing power can be expressed in floating-point operations per second (FLOPS or FP), and can be measured in 6 million floating-point operations per second (MFP), 1 billion floating-point operations per second (GFP), or 1 trillion floating-point operations per second (TFP).

[0041] Meta-computing power can represent the unit computing power of a computing unit. The number of meta-computing power can measure the computing power of a computing unit.

[0042] Task deployment refers to the process of assigning computing tasks, that is, the process of determining the computing unit that will perform the computing task. After the task is deployed, the computing task is dispatched to the corresponding computing unit for calculation according to the deployment result.

[0043] A computational graph is a directed acyclic graph used to describe mathematical calculations. It includes multiple nodes and edges connecting the nodes. A node represents a computational task, and an edge connecting two nodes represents the transfer of the computational result of one computational task to another. A computational graph represents the computation and data transfer process through a series of nodes and edges.

[0044] In the related art, task deployment is to divide the computational graph into equal parts by the number of computing units to obtain subsets equal to the number of computing units, each subset includes equal or approximately equal computing tasks, and one computing unit executes a computing task in one subset.

[0045] For example, a series of objects (computing tasks) are created in the deployment software. These objects can synchronize information by sending and receiving messages to transmit the calculation results. After compiling a series of objects to obtain a calculation graph, the calculation graph is submitted to the runtime environment of the deployment software. The runtime environment will divide the calculation graph into equal parts according to the number of processors to implement task deployment. After deployment, the runtime environment will schedule the calculation tasks that can be executed concurrently to execute on multiple processors, and migrate the calculation results of the calculation tasks between the processors. However, this deployment scheme will cause a more serious load imbalance problem between processors. Due to the load imbalance between processors, the overall computing efficiency is affected. In addition, migrating calculation results between processors will bring greater data handling overhead.

[0046] To this end, an embodiment of the present application provides a task deployment method that can solve the above-mentioned problems.

[0047] In the task deployment method provided in the embodiment of the present application, the computing tasks in the computing graph are allocated to each computing unit according to the computing power of each computing unit, and the difference between the ratio of the computing power load of each computing unit and the ratio of the computing power of each computing unit is equal to 0 or less than the first threshold value, that is, the ratio of the computing power load of each computing unit is equal to or approximately equal to the ratio of the computing power of each computing unit, so as to ensure that the computing power load of each computing unit matches its computing power. This method performs task deployment according to the computing power of the computing unit and the computing power required for the computing task, so that the computing power load of each computing unit can be balanced, so that the computing unit can fully exert its computing power when executing the corresponding computing task, thereby improving the overall computing efficiency. Due to the balanced computing power load of the computing unit, unnecessary data migration between computing units can be reduced, thereby reducing data handling overhead.

[0048] The following takes the computing unit as a processor as an example. Figure 1 and Figure 2 The task deployment method provided in the embodiment of the present application is introduced in detail.

[0049] Figure 1 is a flowchart of a task deployment method provided by an embodiment of the present application. The method can be executed by a task deployment device, and the task deployment device can be any computing device that needs to be deployed, such as a server or a smart phone that needs to be deployed. Figure 1 As shown, the method includes S101-S102.

[0050] S101, the task deployment device determines the computation graph to be deployed.

[0051] In this embodiment, Figure 2As shown, the task deployment device can use the computational graph generation software to process the pre-written computer high-level language, and compile it through the function flow run time (FFRT) to obtain a computational graph suitable for hardware storage and execution of the computing system. The computational graph includes multiple nodes and edges corresponding to multiple nodes. The node represents a computing task (or operator) in the high-level language, and the edge corresponding to the node represents the data transmission volume of the computing task.

[0052] In the process of generating the computational graph, Figure 2 As shown in the figure, determine the weight of each node and the weight of each edge in the computational graph. The weight of the node is used to reflect the computing power required for the computing task, that is, the computing amount of the computing task; the larger the weight of the node, the greater the computing power required for the computing task. The weight of the edge is used to reflect the data transmission volume of the computing task; the larger the weight of the edge, the greater the data transmission volume.

[0053] Profiling is performed on the computing tasks represented by each node to obtain an operator library, which includes the measured values ​​of the computing power required for each computing task and the measured values ​​of the data transmission volume of each computing task.

[0054] The weight of the node can be obtained by the following formula (1) based on the measured value and theoretical value of the computing power of the computing task represented by each node in the operator library.

[0055] w_node (i)=Scale[Profiling(i)·Analysis(i)] (1)

[0056] In the above formula (1), w_node(i) represents the weight of the i-th node in the computational graph, Profiling(i) represents the measured value of the computing power of the computing task represented by the i-th node, and Analysis(i) represents the theoretical value of the computing power of the computing task represented by the i-th node. The Scale function is a function that standardizes data. The Scale function standardizes Profiling(i)·Analysis(i) to obtain the weight of node i.

[0057] The weight of the edge corresponding to the node can also be obtained by the following formula (2) based on the measured value of the data transmission volume of the computing task represented by each node in the operator library and the theoretical value of the data transmission volume.

[0058] w_link (i→j)= Profiling (i→j) ·Analysis (i→j) / EdgeSlack (i→j)(2)

[0059] In the above formula (2), w_link(i→j) represents the weight of the edge between the i-th node and the j-th node in the computation graph, Profiling(i→j) represents the measured value of the data transmission volume between the computation task represented by the i-th node and the computation task represented by the j-th node, Analysis(i→j) represents the theoretical value of the data transmission volume between the computation task represented by the i-th node and the computation task represented by the j-th node, EdgeSlack(i→j) represents the time slack between the i-th node and the j-th node, EdgeSlack(i→j)=LST(j)-EST(i)+1-p(i), where LST(j) represents the latest start execution time of the computation task represented by the j-th node, EST(i) represents the earliest start execution time of the computation task represented by the i-th node, p(i) represents the duration of the computation task represented by the i-th node when it is executed, and the value range of EdgeSlack(i→j) is between 1 and lv, where lv is the number of time-dependent layers of the computation graph.

[0060] S102, the task deployment device allocates multiple computing tasks to multiple computing units according to the computing power of the multiple computing units, wherein the difference between the ratio of the computing power loads of the multiple computing units and the ratio of the computing power of the multiple computing units is equal to 0 or less than a first threshold.

[0061] In this embodiment, Figure 2 As shown, the task deployment device can analyze the computing resources to obtain the computing power of the computing resources, where the computing resources indicate the computing units that perform the calculations, and then allocate multiple computing tasks to each computing unit from the perspectives of the computing power of the processor and the computing power required for the computing tasks, so that the difference between the ratio of the computing power load of multiple computing units and the ratio of the computing power of multiple computing units is equal to 0 or less than a first threshold, that is, the ratio of the computing power load of multiple computing units and the ratio of the computing power of multiple computing units are equal or approximately equal, thereby achieving the purpose of balancing the computing power load of the computing units.

[0062] Specifically, Figure 2 As shown in , the computation graph is first divided into N subsets using the balanced minimum cut method; the balanced minimum cut method can make the sum of the weights of the nodes in each subset equal or approximately equal, and the sum of the weights of the edges between the nodes is minimal. Then, as Figure 2 As shown, the subsets are aggregated to obtain subsets corresponding to each computation graph. The unbalanced minimum cut is realized by the balanced minimum cut + aggregation method, thereby achieving the above purpose. In this embodiment, S102 can be specifically realized by the following three steps S1021-S1023.

[0063] S1021, determine the number N of segments of the computation graph according to the computing power of multiple processors.

[0064] In this step, the task deployment device analyzes the computing power of multiple processors, calculates the greatest common divisor of the computing power of each processor based on the computing power of each processor, and obtains the meta-computing power of each processor. Then, the computing power of each processor is divided by the meta-computing power to obtain the number of meta-computing powers of each processor. Finally, the sum of the number of meta-computing powers of each processor is calculated to obtain the number of splits N.

[0065] For example, if the CPU and GPU are multiple processors, and the computing power of the CPU is 100T FP and the computing power of the GPU is 40T FP, the greatest common divisor of 100 and 40 is 20, and the corresponding meta-computing power of the CPU and GPU is 20T FP. Figure 3a As shown, it can be determined that the number of meta-computing powers of the CPU is 5, the number of meta-computing powers of the GPU is 2, and the number of splits N is 7.

[0066] In addition, since there may be errors between the nominal computing power and the actual computing power of the processor. Therefore, in this step, the computing power of the processor can also be adjusted according to the time it takes for each processor to run the same task. Taking the above-mentioned CPU and GPU as an example, when the difference between the inverse ratio of the time it takes for the CPU and GPU to perform the same task and the ratio of the computing power of the CPU and GPU is greater than the third threshold, the computing power of the processor can be adjusted so that the ratio of the computing power of the processor is close to the inverse ratio of the aforementioned time, that is, the difference between the inverse ratio of the time it takes for the CPU and GPU to perform the same task and the ratio of the computing power of the CPU and GPU is equal to or close to the third threshold.

[0067] S1022, dividing the computation graph into N subsets according to N, wherein each subset in the N subsets includes one or more of the multiple computing tasks, and the difference in computing power required by any two subsets in the N subsets is equal to 0 or less than a second threshold.

[0068] In this step, with the goal of minimizing the amount of data transmitted by the subset, the computational graph is balanced and minimum cut is performed on N pairs of the computational graph, and the computational graph is divided into N subsets, so that the computing power required for any two subsets in the N subsets is equal or approximately equal, that is, the difference in the computing power required for any two subsets is equal to 0 or less than the second threshold. In other words, or, this embodiment divides the computational graph into equal parts according to the computing power of the computational graph, and the number of equal parts is determined according to the computing power of the processor. Among them, the amount of data transmitted by the subset is the sum of the amount of data transmitted by the computing tasks in the subset, and the computing power required for the subset is the sum of the computing power required for the computing tasks in the subset. Among them, the division with the goal of minimizing the amount of data transmitted by the subset can reduce the amount of data transmitted between computing units.

[0069] S1023, allocate N subsets to multiple processors.

[0070] In this step, if Figure 2As shown, the subsets are aggregated according to the ratio of the computing power of multiple processors, and N subsets are allocated to multiple processors, so that the difference between the ratio of the computing power load of multiple processors and the ratio of the computing power of multiple processors is equal to 0 or less than the first threshold. Since the computing power required for each subset is equal or approximately equal, the division according to the ratio of the computing power of the processors can ensure that the ratio of the computing power load of each processor is equal or approximately equal to the ratio of the operators, so that the computing load of each processor matches its computing power.

[0071] Taking the above CPU and GPU as an example, the ratio of CPU and GPU computing power is 5:2 (simplified to 100:40), and N is 7. Figure 3a As shown in the figure, the computation graph is divided into 7 subsets, so 5 of the 7 subsets can be divided into CPUs and the remaining 2 subsets can be divided into GPUs according to a ratio of 5:2. Since the computing power required for each of the 7 subsets is equal or approximately equal, the ratio of the CPU and GPU computing load to the computing power is also equal or approximately equal, and the CPU and GPU computing loads match their computing power.

[0072] In this embodiment, after the deployment result is obtained, the dependency relationship of the computing tasks executed by each computing unit may be pre-scheduled to obtain the scheduling order of the computing tasks executed by each computing unit.

[0073] by Figure 3b As an example, the calculation diagram shown in the figure includes calculation tasks 1 to 10. Figure 1 In the method shown, computing tasks 1, 3, and 5 are assigned to computing unit P1, computing tasks 2, 4, and 7 are assigned to computing unit P3, computing tasks 8 and 9 are assigned to computing unit P2, and computing tasks 10 and 6 are assigned to computing unit P4. Figure 3b As shown in the dependency relationship between the computing tasks, the execution order of P1 is task 1, task 3, task 5, the execution order of P2 is task 8, task 9, the execution order of P3 is task 2, task 4, task 7, and the execution order of P4 is task 10, task 6.

[0074] The scheduling order of computing units P1, P2, P3 and P4 can be determined by pre-scheduling. Figure 3c As shown, that is: P1 is scheduled to execute task 1 and P4 to execute task 10 in parallel; after P1 finishes executing task 1, P1 is scheduled to execute task 3 and P2 is scheduled to execute task 2; after P1 finishes executing task 3 and P4 finishes executing task 10, P4 is scheduled to execute task 6; after P4 finishes executing task 6, P2 is scheduled to execute task 8; after P1 finishes executing task 3, P1 is scheduled to execute task 5; after P2 finishes executing task 8, P2 is scheduled to execute task 9.

[0075] In some application scenarios, the computation graph may contain computation tasks that have strong affinity with each processor. If a computation task is only suitable for execution on a specific processor, the computation task is a computation task with strong affinity for the processor. Therefore, before splitting the computation graph, these pre-determined target computation tasks with strong affinity can be removed, that is, these target computation tasks do not participate in graph splitting and task deployment.

[0076] As mentioned above, depending on the level of the computing system, the computing unit can be a computing device, a processor unit, a processor, or a processor core. Figure 1 The method embodiment shown, the embodiment of the present application also provides a hierarchical deployment method. Taking the computing device as a server as an example, multiple computing tasks of the application scenario are first deployed to each server, and then multiple computing tasks corresponding to the server are deployed to each processor in the server, and finally multiple computing tasks corresponding to the processor are deployed to each core of the processor.

[0077] Figure 4 is a flowchart of the hierarchical deployment method provided by the embodiment of the present application. The method can be executed by a task deployment device, and the task deployment device can be any computing device that needs to be deployed, such as a server or a smart phone that needs to be deployed. Figure 4 As shown, the method may include S401-S405.

[0078] S401, the task deployment device determines the computation graph to be deployed.

[0079] In this embodiment, the specific process of S401 can refer to the above Figure 1 The introduction of S101 in the illustrated method embodiment will not be repeated here.

[0080] S402, allocating multiple computing tasks to multiple servers according to the computing power of the multiple servers, wherein the difference between the ratio of the computing power loads of the multiple servers and the ratio of the computing power of the multiple servers is equal to 0 or less than a first threshold.

[0081] S403, allocating multiple computing tasks of the server to multiple processor units according to the computing power of the multiple processor units in the server, wherein the difference between the ratio of the computing power loads of the multiple processor units and the ratio of the computing power of the multiple processor units is equal to 0 or less than a first threshold.

[0082] S404, allocating multiple computing tasks of the processor unit to multiple processors according to the computing power of multiple processors in the processor unit, wherein the difference between the ratio of the computing power loads of the multiple processors and the ratio of the computing power of the multiple processors is equal to 0 or less than a first threshold.

[0083] S405, allocating multiple computing tasks of the processor to the multiple cores according to the computing power of the multiple cores in the processor, wherein the difference between the ratio of the computing power load of the multiple cores and the ratio of the computing power of the multiple cores is equal to 0 or less than a first threshold.

[0084] In this embodiment, the specific processes of S402 to S405 can refer to the above Figure 1 The introduction of S102 in the illustrated method embodiment will not be repeated here.

[0085] It should be noted that if only one computing device needs to be deployed, such as a server or a mobile phone, S402 is not executed and execution starts from S403; if the computing device includes only one processor, S403 and S404 are not executed and S405 is directly executed; if a processor unit of the computing device includes only one processor, the deployment of the processor unit does not execute S404 and directly executes S405. If a processor of the computing device includes only one core, the deployment of the processor does not execute S405.

[0086] In one embodiment, Figure 1 and Figure 4 The method shown can be applied to a graph engine (GE), which runs in the above-mentioned task deployment device. Figure 1 or Figure 4 By using this method, GE can significantly improve the efficiency of the hybrid computing system when executing artificial intelligence (AI) model training scenarios. The computing system includes 8 identical CPU chips and 8 identical NPU chips, such as Figure 5 As shown, the computing power of each CPU is about 62.6T FP64, and the computing power of each NPU is about 600TFP16. Converting the CPU computing power to FP16 is equivalent to 250TFP16 computing power. Therefore, the "meta computing power" can be obtained as 50T FP16. Therefore, a single CPU has 5 meta computing powers and a single NPU has 12 meta computing powers. After splitting the computational graph of a business in an artificial intelligence (AI) training scenario based on the total number of meta computing powers and aggregating the subsets, after testing, the efficiency of the computing system in executing the computational graph of the business can be significantly improved.

[0087] based on Figure 1 and Figure 4 The method embodiment shown in the embodiment of the present application also provides a task deployment device. The task deployment device is used to perform the above Figure 2 or Figure 4 The various steps in the method embodiment are shown.

[0088] Figure 6 600 is a schematic diagram of a task deployment device 600 provided in an embodiment of the present application. Figure 6 As shown, the method may include an analysis module 601 and a deployment module 602. The analysis module 601 and the deployment module 602 may be software programs or hardware devices. When the analysis module 601 and the deployment module 602 are software, the analysis module 601 and the deployment module 602 may be set in different hardware devices, or may be set in different hardware devices.

[0089] The analysis module 601 is used to determine the computation graph to be deployed. In the process of determining the computation graph, the weights of each node and each edge in the computation graph are determined.

[0090] The deployment module 602 is used to allocate multiple computing tasks to multiple computing units according to the computing power of multiple computing units corresponding to the computing graph. Figure 4 In the illustrated method embodiment, the deployment module 602 is executed multiple times to implement hierarchical deployment, thereby balancing the computing load between computing units of different computing powers in a finer granularity, so that the computing load matches its computing power.

[0091] It should be noted that Figure 6 The task deployment device 600 provided in the embodiment shown in the figure only uses the division of the above-mentioned functional modules as an example when executing the task deployment method. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Figure 1 or Figure 4 The task deployment method embodiments shown belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0092] Figure 7 It is a schematic diagram of the hardware structure of a computing device 700 provided in an embodiment of the present application.

[0093] The computing device 700 may include the task deployment device described above. Figure 7 The computing device 700 includes a processor 701, a memory 702, a communication interface 703, and a bus 704. The processor 701, the memory 702, and the communication interface 703 are connected to each other via the bus 704. The processor 701, the memory 702, and the communication interface 703 may also be connected in other connection modes besides the bus 704.

[0094] The processor 701 may be a general-purpose processor, which may be a processor that performs specific steps and / or operations by reading and executing the contents stored in a memory (e.g., the memory 702). For example, the general-purpose processor may be a central processing unit (CPU). The processor 701 may include at least one circuit to perform Figure 1 or Figure 4 The illustrated embodiment provides all or part of the steps of the task deployment method. The processor 701 may include one or more cores.

[0095] The memory 702 may be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical storage, hard disk, etc. Figure 7 As shown, the memory 702 can be specifically used for analyzing the program codes corresponding to the module and the deployment module. When the processor 701 runs the program code, the above-mentioned task deployment method is implemented.

[0096] The communication interface 703 includes an input / output (I / O) interface, a physical interface, and a logical interface, etc., which are used to interconnect devices within the computing device 700, and an interface for interconnecting the computing device 700 with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.

[0097] The bus 704 may be any type of communication bus for interconnecting the processor 701 , the memory 702 , and the communication interface 703 , such as a system bus.

[0098] The above devices may be arranged on independent chips, or at least partially or completely on the same chip. Whether to arrange each device independently on different chips or to integrate them on one or more chips often depends on the needs of product design. The embodiments of the present application do not limit the specific implementation form of the above devices.

[0099] Figure 7 The computing device 700 shown is merely exemplary. During implementation, the computing device 700 may further include other components, which are not listed one by one herein.

[0100] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.

[0101] It is understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0102] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the present application should be included in the scope of protection of the present application.

Claims

1. A task deployment method, characterized in that, the method includes: Determine a computation graph to be deployed, where the computation graph includes multiple computation tasks; According to the computing power of multiple computing units corresponding to the computation graph, allocate the multiple computation tasks to the multiple computing units, where at least two of the multiple computing units have different computing powers, and the difference between the computing power loads ratio and the computing power ratio of the multiple computing units is equal to 0 or less than a first threshold, and the computing power load of each computing unit is determined according to the sum of the computing powers required by the computation tasks corresponding to each computing unit.

2. The method according to claim 1, characterized in that, the step of allocating the multiple computation tasks to the multiple computing units according to the computing power of the multiple computing units includes: Determine the number of partitions N of the computation graph according to the computing power of the multiple computing units; Divide the computation graph into N subsets according to the N, where each of the N subsets includes one or more of the multiple computation tasks, and the difference in the computing power required by any two of the N subsets is equal to 0 or less than a second threshold, and the computing power required by each subset is the sum of the computing powers required by the computation tasks in each subset; Allocate the N subsets to the multiple computing units.

3. The method according to claim 2, characterized in that, the step of determining the number of partitions N of the computation graph according to the computing power of the multiple computing units includes: Determine a meta-computing power according to the greatest common divisor of the computing powers of the multiple computing units, and the meta-computing power represents the unit computing power of the multiple computing units; Calculate the quotient of the computing powers of the multiple computing units and the meta-computing power to obtain the number of meta-computing power of the multiple computing units; Calculate the sum of the number of meta-computing power of the multiple computing units to obtain the N.

4. The method according to claim 3, characterized in that, before determining the meta-computing power according to the greatest common divisor of the computing powers of the multiple computing units, the method further includes: Update the computing power of the multiple computing units according to the time for the multiple computing units to execute the same computation task.

5. The method according to any one of claims 2-4, characterized in that, the step of dividing the computation graph into N subsets according to the N includes: Partition the computation graph with the goal of minimizing the data transfer amount of the subsets to obtain the N subsets, and the data transfer amount of the subset is the sum of the data transfer amounts of the computation tasks in the subset.

6. The method according to any one of claims 1-5, characterized in that, before allocating the multiple computation tasks to the multiple computing units according to the computing power of the multiple computing units, the method further includes: Remove a target computation task from the multiple computation tasks, and the target computation task includes the tasks executed by each of the computing units determined in advance.

7. The method according to any one of claims 1-6, characterized in that, the computing unit includes a server, a processor in the server, or a core in the processor.

8. A task deployment device, characterized in that, the device includes: An analysis module for determining a computation graph to be deployed, where the computation graph includes multiple computing tasks; A deployment module for allocating the multiple computing tasks to the multiple computing units according to the computing power of the multiple computing units corresponding to the computation graph, where at least two of the multiple computing units have different computing powers, and the difference between the ratio of the computing power loads of the multiple computing units and the ratio of the computing powers of the multiple computing units is equal to 0 or less than a first threshold, and the computing power load of each computing unit is determined according to the sum of the computing powers required by the computing tasks corresponding to each computing unit.

9. The apparatus according to claim 8, wherein, the deployment module is specifically configured to: determine the number of partitions N of the computation graph according to the computing power of the multiple computing units; divide the computation graph into N subsets according to the N, each of the N subsets includes one or more of the multiple computing tasks, and the difference between the computing powers required by any two of the N subsets is equal to 0 or less than a second threshold, and the computing power required by each subset is the sum of the computing powers required by the computing tasks in each subset; allocate the N subsets to the multiple computing units.

10. The apparatus according to claim 9, wherein, the deployment module is specifically configured to: determine a meta-computing power according to the greatest common divisor of the computing powers of the multiple computing units, and the meta-computing power represents the unit computing power of the multiple computing units; calculate the quotient of the computing powers of the multiple computing units and the meta-computing power to obtain the number of meta-computing power of the multiple computing units; calculate the sum of the number of meta-computing power of the multiple computing units to obtain the N.

11. The apparatus according to claim 10, wherein, the deployment module is specifically configured to: before determining the meta-computing power according to the greatest common divisor of the computing powers of the multiple computing units, update the computing powers of the multiple computing units according to the time for the multiple computing units to execute the same computing task.

12. The apparatus according to any one of claims 9-11, wherein, the deployment module is specifically configured to: partition the computation graph with the goal of minimizing the data transfer amount of the subsets to obtain the N subsets, and the data transfer amount of the subset is the sum of the data transfer amounts of the computing tasks in the subset.

13. The apparatus according to any one of claims 8-12, wherein, the deployment module is specifically configured to: before allocating the multiple computing tasks to the multiple computing units according to the computing power of the multiple computing units, remove a target computing task from the multiple computing tasks, and the target computing task includes the tasks executed by the respective computing units determined in advance.

14. The apparatus according to any one of claims 8-13, wherein, the computing unit includes a server, a processor unit in the server, a processor in the server, or a core in the processor.

15. A computing device, wherein, the computing device includes: a processor and a memory, and the processor is configured to execute a computer program stored in the memory to implement the method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that, it includes instructions which, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Dynamic calculation graph optimization and deployment method and system based on heterogeneous calculation platform

    CN121681123A

  • Task deployment method and task deployment apparatus

    EP4804026A1

  • Task deployment method and task deployment apparatus

    WO2025113318A1