Task deployment method and task deployment apparatus

By allocating the computing power of the calculation unit based on the computing power distribution calculation task, the required computing power is matched with the computing power of the calculation unit, the problem of unbalanced load in the computing system is solved and the overall computing efficiency is improved.

WO2025113318A1PCT designated stage expired Publication Date: 2025-06-05HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133597
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-21
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When the prior art deploys computing tasks in a computing system, it is easy to cause load imbalance between computing units, extend business computing time, and affect the overall computing efficiency.

Method used

By allocating the calculation tasks in the calculation diagram based on the computing power of the calculation unit, the sum of the calculation forces required for the calculation tasks performed by the calculation unit is matched with its computing power, thereby equalizing the load between the calculation units.

Benefits of technology

The load balance between computing units is realized, the overall computing efficiency is improved, unnecessary data migration between computing units is reduced, and data handling overhead is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133597_05062025_PF_FP_ABST
    Figure CN2024133597_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a task deployment method and a task deployment apparatus. The method comprises: determining a computing graph to be deployed; and on the basis of computing power of a plurality of computing units corresponding to the computing graph, allocating a plurality of computing tasks in the computing graph to the plurality of computing units, wherein the computing power of at least two of the plurality of computing units is different, the difference between the ratio of computing power loads of the plurality of computing units and the ratio of the computing power of the plurality of computing units is equal to 0 or less than a first threshold value, and the computing power load of each computing unit is determined on the basis of the sum of the computing power required by the computing tasks corresponding to the computing unit. According to the solution, computing tasks are allocated on the basis of computing power of computing units, so that the sum of the computing power required by the computing tasks executed by each computing unit matches the computing power of the computing unit, and loads among the plurality of computing units can be balanced, thereby improving the overall computing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

A task deployment method and task deployment device

[0001] This application claims priority to Chinese patent application number 202311609712.9, filed with the State Intellectual Property Office of China on November 27, 2023, entitled "A Task Deployment Method and Task Deployment Device," the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a task deployment method and a task deployment device. Background Art

[0003] The demand for computing power of computing systems is increasing, and a large number of existing and emerging businesses need to deploy computing tasks on computing systems.

[0004] In the related art, computing tasks are usually deployed equally to each computing unit according to the number of computing units in the computing system. That is, in one deployment, the number of computing tasks processed by each computing unit is equal or approximately equal. Taking the computing unit as a processor as an example, when the server includes three processors, the computing graph including multiple computing tasks is divided into three subsets, and the computing tasks in each subset are equal or approximately equal, and then the three processors respectively execute the computing tasks in the three subsets. Since the processing capabilities of multiple computing units may be different, this deployment scheme of dividing the computing graph by the number of computing units may cause unbalanced loads on each computing unit, prolong the overall computing time of the business, and affect the overall computing efficiency. Summary of the Invention

[0005] The present application provides a task deployment method and a task deployment device, which deploy computing tasks based on the computing power of computing units, can balance the load between computing units and improve overall computing efficiency.

[0006] In a first aspect, the present application provides a task deployment method. The method includes: determining a computation graph to be deployed, the computation graph including multiple computation tasks; allocating the multiple computation tasks to the multiple computation units according to the computation power of the multiple computation units corresponding to the computation graph, wherein at least two of the multiple computation units have different computation powers, the ratio of the computation power loads of the multiple computation units and the difference between the ratio of the computation power of the multiple computation units are equal to 0 or less than a first threshold, and the computation power load of each computation unit is determined according to the sum of the computation powers required for the computation tasks corresponding to each computation unit.

[0007] In the above scheme, computing tasks are allocated according to the computing power of the computing units, so that the sum of the computing power required for the computing tasks executed by the computing units matches the computing power of the computing units, thereby balancing the load among multiple computing units and improving the overall computing efficiency.

[0008] In a possible implementation of the first aspect, allocating the multiple computing tasks to the multiple computing units based on the computing power of the multiple computing units includes: determining the number N of splits of the computing graph based on the computing power of the multiple computing units; dividing the computing graph into N subsets according to N, each subset in the N subsets including one or more of the multiple computing tasks, the difference in computing power required by any two subsets in the N subsets is equal to 0 or less than a second threshold, and the computing power required by each subset is the sum of the computing power required by the computing tasks in each subset; allocating the N subsets to the multiple computing units according to the ratio of the computing power of the multiple computing units.

[0009] In the above scheme, the concept of meta-computing power is introduced. According to the total meta-computing power N of multiple processors, the computation graph is divided into N subsets according to the computing power. Then, each subset is allocated according to the ratio of the computing power of multiple computing units. This can evenly distribute computing tasks to the computing units to achieve the purpose of load balancing.

[0010] In a possible implementation of the first aspect, determining the number N of splits of the computational graph based on the computing power of the multiple computing units includes: determining a meta-computing power based on a common divisor of the computing power of the multiple computing units, the meta-computing power representing the unit computing power of the multiple computing units; calculating the quotient of the computing power of the multiple computing units and the meta-computing power to obtain the number of meta-computing powers of the multiple computing units; and calculating the sum of the number of meta-computing powers of the multiple computing units to obtain N.

[0011] In a possible implementation of the first aspect, before determining the meta-computing power based on the common divisor of the computing power of the multiple computing units, the method further includes: updating the computing power of the multiple computing units based on the time it takes for the multiple computing units to perform the same computing task.

[0012] In a possible implementation of the first aspect, dividing the computation graph into N subsets according to N includes: dividing the computation graph with the goal of minimizing the data transmission volume of the subsets to obtain the N subsets, and the data transmission volume of the subsets is the sum of the data transmission volumes of the computing tasks in the subsets.

[0013] In a possible implementation of the first aspect, before allocating the multiple computing tasks to the multiple computing units based on the computing power of the multiple computing units, the method further includes: removing target computing tasks from the multiple computing tasks, and the target computing tasks include predetermined tasks to be performed by each of the computing units.

[0014] In a possible implementation of the first aspect, the computing unit includes a server, a processor unit in the server, a processor in the server, or a core in the processor.

[0015] In a second aspect, the present application further provides a task deployment device, which includes an analysis module and a deployment module.

[0016] Among them, the analysis module is used to determine the computing graph to be deployed and the computing power of multiple computing units corresponding to the computing graph, and the computing graph includes multiple computing tasks.

[0017] Among them, the deployment module is used to allocate the multiple computing tasks to the multiple computing units according to the computing power of the multiple computing units corresponding to the computing graph, wherein the computing power of at least two computing units among the multiple computing units is different, the ratio of the computing power load of the multiple computing units and the difference between the ratio of the computing power of the multiple computing units are equal to 0 or less than the first threshold, and the computing power load of each computing unit is determined according to the sum of the computing power required for the computing tasks corresponding to each computing unit.

[0018] In a possible implementation of the second aspect, the deployment module is specifically used to: determine the number N of splits of the computational graph according to the computing power of the multiple computing units; divide the computational graph into N subsets according to N, each subset in the N subsets includes one or more of the multiple computing tasks, the difference in computing power required by any two subsets in the N subsets is equal to 0 or less than a second threshold, and the computing power required for each subset is the sum of the computing power required for the computing tasks in each subset; and allocate the N subsets to the multiple computing units.

[0019] In a possible implementation of the second aspect, the deployment module is specifically used to: determine the meta-computing power based on the common divisor of the computing power of the multiple computing units, where the meta-computing power represents the unit computing power of the multiple computing units; calculate the quotient of the computing power of the multiple computing units and the meta-computing power to obtain the number of meta-computing powers of the multiple computing units; calculate the sum of the number of meta-computing powers of the multiple computing units to obtain N.

[0020] In a possible implementation of the second aspect, the deployment module is specifically used to: before determining the meta-computing power based on the common divisor of the computing power of the multiple computing units, update the computing power of the multiple computing units based on the time it takes for the multiple computing units to perform the same computing task.

[0021] In a possible implementation of the second aspect, the deployment module is specifically used to: divide the computational graph with the goal of minimizing the data transmission volume of the subsets to obtain the N subsets, and the data transmission volume of the subsets is the sum of the data transmission volumes of the computational tasks in the subsets.

[0022] In a possible implementation of the second aspect, the deployment module is specifically used to: before allocating the multiple computing tasks to the multiple computing units based on the computing power of the multiple computing units, remove the target computing tasks from the multiple computing tasks, and the target computing tasks include predetermined tasks to be performed by each of the computing units.

[0023] In a possible implementation of the second aspect, the computing unit includes a server, a processor unit in the server, a processor in the server, or a core in the processor.

[0024] In a third aspect, the present application further provides a computing device. The computing device includes a processor and a memory. The processor is configured to execute a computer program stored in the memory to implement the task deployment method provided in the first aspect or any possible implementation of the first aspect.

[0025] In a fourth aspect, the present application also provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is executed on a computer, the computer implements the task deployment method provided by the first aspect or any possible implementation method of the first aspect.

[0026] In a fifth aspect, the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to implement the task deployment method provided in the first aspect or any possible implementation method of the first aspect.

[0027] Any of the devices, computer storage media, or computer program products provided above are used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding schemes in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] FIG1 is a flowchart of a task deployment method provided in an embodiment of the present application;

[0029] FIG2 is a detailed schematic diagram of a task deployment method shown in FIG1 provided in an embodiment of the present application;

[0030] FIG3a is a schematic diagram of a task deployment for a central processing unit (CPU) and a graphics processing unit (GPU) provided by an embodiment of the present application;

[0031] FIG3 b is a schematic diagram of a calculation graph provided in an embodiment of the present application;

[0032] FIG3c is a schematic diagram of a sequence of computing tasks performed by various computing units obtained through pre-scheduling according to an embodiment of the present application;

[0033] FIG4 is a flowchart of a hierarchical deployment method based on the method shown in FIG1 provided in an embodiment of the present application;

[0034] FIG5 is a schematic diagram of a task deployment for a central processing unit (CPU) and a graphics processing unit (NPU) provided by an embodiment of the present application;

[0035] FIG6 is a schematic diagram of the structure of a task deployment device provided in an embodiment of the present application;

[0036] FIG7 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0038] Before introducing the embodiments of the present application, the nouns appearing in the embodiments of the present application are first introduced below.

[0039] A computing unit refers to a device in a computing system that performs specific computing tasks. Depending on the level of computing units in the computing system, computing can specifically include computing devices, processor units, processors, or processor cores. A processor unit can be determined based on the memory area accessed by the processor. Processors accessing the same memory area are called a processor unit, and a processor unit can include one or more processors. It is understood that a computing system typically includes multiple computing units. A computing system can specifically be a distributed computing cluster including multiple computing devices, or a computing device including multiple processors. Computing devices can include servers and terminal devices, and terminal devices can include laptops, smartphones, and the like.

[0040] The processor of the computing device may include multiple homogeneous processors or multiple heterogeneous processors, each processor including at least one core. The processor may specifically include a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU). It can be understood that homogeneous and heterogeneous refer to the same or different structures of the processors. For example, any two of the CPU, GPU, and NPU are heterogeneous processors, and multiple CPUs with different structures, or multiple GPUs with different structures, or multiple NPUs with different structures are also heterogeneous processors. Multiple CPUs with the same structure, or multiple GPUs with the same structure, or multiple NPUs with the same structure are homogeneous processors.

[0041] Computing power refers to the computing capacity of a computing unit. It can be expressed in floating-point operations per second (FLOPS or FP), and can be measured in units of 6 million floating-point operations per second (MFP), 1 billion floating-point operations per second (GFP), or 1 trillion floating-point operations per second (TFP).

[0042] Meta-computing power can represent the unit computing power of a computing unit. The number of meta-computing power can measure the computing power of a computing unit.

[0043] Task deployment is the process of assigning computing tasks, that is, determining the computing units that will perform the computing tasks. After the task is deployed, the computing tasks are dispatched to the corresponding computing units for calculation according to the deployment results.

[0044] A computation graph is a directed acyclic graph used to describe mathematical computations. It consists of multiple nodes and edges connecting them. A node represents a computation task, and an edge connecting two nodes represents the transfer of the results of one computation task to another. A computation graph represents the computation and data transfer process through a series of nodes and edges.

[0045] In related technologies, task deployment is to divide the computational graph into equal parts by the number of computing units to obtain subsets equal to the number of computing units. Each subset includes equal or approximately equal computing tasks, and one computing unit executes a computing task in one subset.

[0046] For example, a series of objects (computing tasks) are created in the deployment software. These objects can synchronize information by sending and receiving messages to transmit computation results. After compiling these objects to form a computation graph, the graph is submitted to the deployment software's runtime environment. This runtime environment divides the computation graph equally according to the number of processors to implement task deployment. After deployment, the runtime environment schedules concurrently executable computation tasks to be executed on multiple processors and migrates the computation results of the computation tasks between processors. However, this deployment solution can lead to significant load imbalance between processors. This imbalance in load affects overall computational efficiency. Furthermore, migrating computation results between processors incurs greater data transfer overhead.

[0047] To this end, an embodiment of the present application provides a task deployment method that can solve the above-mentioned problems.

[0048] In the task deployment method provided in the embodiment of the present application, the computing tasks in the computing graph are allocated to each computing unit according to the computing power of each computing unit, and the difference between the ratio of the computing power load of each computing unit and the ratio of the computing power of each computing unit is equal to 0 or less than the first threshold value, that is, the ratio of the computing power load of each computing unit and the ratio of the computing power of each computing unit are equal or approximately equal, so as to ensure that the computing power load of each computing unit matches its computing power. This method performs task deployment according to the computing power of the computing unit and the computing power required for the computing task, which can balance the computing power load of each computing unit and enable the computing unit to fully exert its computing power when executing the corresponding computing task, thereby improving the overall computing efficiency. Due to the balanced computing power load of the computing unit, unnecessary data migration between computing units can be reduced, thereby reducing data handling overhead.

[0049] The following takes the computing unit as an example, and combines Figures 1 and 2 to specifically introduce the task deployment method provided in the embodiment of the present application.

[0050] Figure 1 is a flowchart of a task deployment method provided by an embodiment of the present application. This method can be executed by a task deployment device, which can be any computing device to be deployed, such as a server or smartphone. As shown in Figure 1, the method includes S101-S102.

[0051] S101: The task deployment device determines a computation graph to be deployed.

[0052] In this embodiment, as shown in Figure 2, the task deployment device can use computational graph generation software to process pre-written high-level computer languages ​​and compile them using the Function Flow Runtime (FFRT) to generate a computational graph suitable for hardware storage and execution in the computing system. The computational graph includes multiple nodes and edges corresponding to the nodes. A node represents a computational task (or operator) in the high-level language, and the edges corresponding to the nodes represent the data transfer volume of the computational task.

[0053] During computational graph generation, as shown in Figure 2, the weights of each node and edge in the graph are determined. Node weights reflect the computing power required for the computational task, that is, the computational effort of the task; a larger node weight indicates a greater computational power requirement. Edge weights reflect the data transfer volume of the computational task; a larger edge weight indicates a greater data transfer volume.

[0054] Profiling is performed on the computing tasks represented by each node to generate an operator library. The operator library includes the measured values ​​of the computing power required for each computing task and the measured values ​​of the data transmission volume for each computing task.

[0055] The weight of a node can be obtained by the following formula (1) based on the measured value and theoretical value of the computing power of the computing task represented by each node in the operator library.

[0056] w_node (i)=Scale[Profiling(i)·Analysis(i)] (1)

[0057] In the above formula (1), w_node(i) represents the weight of the i-th node in the computation graph, Profiling(i) represents the measured value of the computing power of the computing task represented by the i-th node, and Analysis(i) represents the theoretical value of the computing power of the computing task represented by the i-th node. The Scale function is a function that standardizes data. The Scale function standardizes Profiling(i)·Analysis(i) to obtain the weight of node i.

[0058] The weight of the edge corresponding to the node can also be obtained by the following formula (2) based on the measured value and theoretical value of the data transmission volume of the computing task represented by each node in the operator library.

[0059] w_link (i→j)= Profiling (i→j) ·Analysis (i→j) / EdgeSlack (i→j) (2)

[0060] In the above formula (2), w_link(i→j) represents the weight of the edge between the i-th node and the j-th node in the computation graph, Profiling(i→j) represents the measured value of the data transmission volume between the computation task represented by the i-th node and the computation task represented by the j-th node, Analysis(i→j) represents the theoretical value of the data transmission volume between the computation task represented by the i-th node and the computation task represented by the j-th node, EdgeSlack(i→j) represents the temporal slack between the i-th node and the j-th node, EdgeSlack(i→j)=LST(j)-EST(i)+1-p(i), where LST(j) represents the latest start execution time of the computation task represented by the j-th node, EST(i) represents the earliest start execution time of the computation task represented by the i-th node, p(i) represents the duration of the computation task represented by the i-th node when it is executed, and the value range of EdgeSlack(i→j) is between 1 and lv, where lv is the number of temporal layers of the computation graph.

[0061] S102: The task deployment device allocates the plurality of computing tasks to the plurality of computing units according to the computing power of the plurality of computing units, wherein the difference between the ratio of the computing power loads of the plurality of computing units and the ratio of the computing power of the plurality of computing units is equal to 0 or less than a first threshold.

[0062] In this embodiment, as shown in Figure 2, the task deployment device can analyze the computing resources to obtain the computing power of the computing resources, where the computing resources indicate the computing units that perform the calculations. Then, based on the computing power of the processor and the computing power required for the computing tasks, multiple computing tasks are allocated to each computing unit, so that the ratio of the computing power load of multiple computing units and the difference between the ratio of the computing power of multiple computing units are equal to 0 or less than the first threshold, that is, the ratio of the computing power load of multiple computing units and the ratio of the computing power of multiple computing units are equal or approximately equal, thereby achieving the purpose of balancing the computing power load of the computing units.

[0063] Specifically, as shown in Figure 2, the computation graph is first divided using the balanced minimum cut method to obtain N subsets; the balanced minimum cut can ensure that the sum of the weights of the nodes in each subset is equal or approximately equal, and the sum of the weights of the edges between the nodes is minimized. Then, as shown in Figure 2, the subsets are aggregated to obtain subsets corresponding to each computation graph. The balanced minimum cut + aggregation method is used to achieve an unbalanced minimum cut, thereby achieving the above purpose. In this embodiment, S102 can be specifically implemented through the following three steps S1021-S1023.

[0064] S1021, determine the number N of segments of the computation graph according to the computing power of multiple processors.

[0065] In this step, the task deployment device analyzes the computing power of multiple processors and calculates the greatest common divisor of their computing power based on their computing power, obtaining the meta-computing power of each processor. The computing power of each processor is then divided by this meta-computing power to obtain the number of meta-computing powers for each processor. Finally, the sum of the meta-computing powers of each processor is calculated to obtain the number of partitions, N.

[0066] Taking multiple processors as CPUs and GPUs as an example, when the computing power of the CPU is 100T FP and the computing power of the GPU is 40T FP, the greatest common divisor of 100 and 40 is 20, and the corresponding meta-computing power of the CPU and GPU is 20T FP. As shown in Figure 3a, it can be determined that the number of meta-computing powers of the CPU is 5 and the number of meta-computing powers of the GPU is 2, and the number of splits N can be determined to be 7.

[0067] In addition, since there may be errors between the nominal computing power and the actual computing power of the processor, in this step, the computing power of the processor can also be adjusted according to the time it takes for each processor to run the same task. Taking the above-mentioned CPU and GPU as an example, when the difference between the inverse ratio of the time it takes for the CPU and GPU to perform the same task and the ratio of the computing power of the CPU and GPU is greater than the third threshold, the computing power of the processor can be adjusted so that the ratio of the computing power of the processor is close to the inverse ratio of the aforementioned time, that is, the difference between the inverse ratio of the time it takes for the CPU and GPU to perform the same task and the ratio of the computing power of the CPU and GPU is equal to or close to the third threshold.

[0068] S1022, dividing the computation graph into N subsets according to N, wherein each subset in the N subsets includes one or more of the multiple computing tasks, and the difference in computing power required by any two subsets in the N subsets is equal to 0 or less than a second threshold.

[0069] In this step, with the goal of minimizing the amount of data transmitted by the subset, a balanced minimum cut is performed on the computation graph according to N pairs, and the computation graph is divided into N subsets, so that the computing power required by any two subsets in the N subsets is equal or approximately equal, that is, the difference in the computing power required by any two subsets is equal to 0 or less than the second threshold. In other words, this embodiment divides the computation graph into equal parts according to the computing power of the computation graph, and the number of equal parts is determined according to the computing power of the processor. Among them, the amount of data transmitted by the subset is the sum of the amount of data transmitted by the computing tasks in the subset, and the computing power required by the subset is the sum of the computing power required by the computing tasks in the subset. Among them, dividing the subsets with the goal of minimizing the amount of data transmitted by the subset can reduce the amount of data transmitted between computing units.

[0070] S1023: Allocate the N subsets to multiple processors.

[0071] In this step, as shown in Figure 2, subsets are aggregated according to the ratio of the computing power of multiple processors, and N subsets are assigned to multiple processors so that the difference between the ratio of the computing power load of the multiple processors and the ratio of the computing power of the multiple processors is equal to 0 or less than a first threshold. Because the computing power required by each subset is equal or approximately equal, dividing the subset according to the computing power ratio of the processors can ensure that the computing power load ratio of each processor is equal or approximately equal to the ratio of the operators, thereby matching the computing power load of each processor with its computing power.

[0072] Taking the CPU and GPU mentioned above as an example, the CPU-GPU computing power ratio is 5:2 (simplified to 100:40), and N is 7. As shown in Figure 3a, the computation graph is divided into 7 subsets. Therefore, 5 of the 7 subsets can be allocated to the CPU according to the 5:2 ratio, and the remaining 2 subsets can be allocated to the GPU. Because the computing power required by each of the 7 subsets is equal or approximately equal, the ratio of the CPU and GPU computing power load to the computing power ratio is also equal or approximately equal, and the CPU and GPU computing power loads match their computing power.

[0073] In this embodiment, after the deployment result is obtained, the dependency relationship of the computing tasks executed by each computing unit can be pre-scheduled to obtain the scheduling order of the computing tasks executed by each computing unit.

[0074] Taking the computation graph shown in Figure 3b as an example, which includes computation tasks 1 through 10, after the method shown in Figure 1, computation tasks 1, 3, and 5 are assigned to computation unit P1; computation tasks 2, 4, and 7 are assigned to computation unit P3; computation tasks 8 and 9 are assigned to computation unit P2; and computation tasks 10 and 6 are assigned to computation unit P4. Based on the dependencies between the computation tasks shown in Figure 3b, the execution order for P1 is task 1, task 3, and task 5; the execution order for P2 is task 8, and task 9; the execution order for P3 is task 2, task 4, and task 7; and the execution order for P4 is task 10, and task 6.

[0075] Through pre-scheduling, the scheduling order of computing units P1, P2, P3 and P4 can be determined as shown in Figure 3c, that is: P1 is scheduled to execute task 1 and P4 to execute task 10 in parallel; after P1 finishes executing task 1, P1 is scheduled to execute task 3 and P2 is scheduled to execute task 2; after P1 finishes executing task 3 and P4 finishes executing task 10, P4 is scheduled to execute task 6; after P4 finishes executing task 6, P2 is scheduled to execute task 8; after P1 finishes executing task 3, P1 is scheduled to execute task 5; after P2 finishes executing task 8, P2 is scheduled to execute task 9.

[0076] In some application scenarios, a computation graph may contain tasks with strong affinity for various processors. If a task is only suitable for execution on a specific processor, it is considered to have strong affinity for that processor. Therefore, before splitting the computation graph, these pre-determined target tasks with strong affinity can be removed. In other words, these target tasks are not involved in graph splitting and task deployment.

[0077] As mentioned above, according to the different levels in the computing system, the computing unit can be a computing device, a processor unit, a processor, or a core of a processor. Therefore, based on the method embodiment shown in Figure 1, the embodiment of the present application also provides a layered deployment method. Taking the computing device as a server as an example, multiple computing tasks of the application scenario are first deployed to each server, and then the multiple computing tasks corresponding to the server are deployed to each processor in the server, and finally the multiple computing tasks corresponding to the processor are deployed to each core of the processor.

[0078] Figure 4 is a flow chart of a hierarchical deployment method provided by an embodiment of the present application. This method can be executed by a task deployment device, which can be any computing device to be deployed, such as a server or smartphone. As shown in Figure 4, the method may include steps S401-S405.

[0079] S401: The task deployment device determines a computation graph to be deployed.

[0080] In this embodiment, the specific process of S401 can refer to the introduction of S101 in the method embodiment shown in Figure 1 above, and will not be repeated here.

[0081] S402: Allocate the plurality of computing tasks to the plurality of servers according to the computing power of the plurality of servers, wherein the difference between the ratio of the computing power loads of the plurality of servers and the ratio of the computing power of the plurality of servers is equal to 0 or less than a first threshold.

[0082] S403: Allocate multiple computing tasks of the server to the multiple processor units based on the computing power of the multiple processor units in the server, wherein a difference between a ratio of the computing power loads of the multiple processor units and a ratio of the computing power of the multiple processor units is equal to 0 or less than a first threshold.

[0083] S404: Allocate the plurality of computing tasks of the processor unit to the plurality of processors according to the computing power of the plurality of processors in the processor unit, wherein the difference between the ratio of the computing power loads of the plurality of processors and the ratio of the computing power of the plurality of processors is equal to 0 or less than a first threshold.

[0084] S405: Allocate the plurality of computing tasks of the processor to the plurality of cores according to the computing power of the plurality of cores in the processor, wherein the difference between the ratio of the computing power loads of the plurality of cores and the ratio of the computing power of the plurality of cores is equal to 0 or less than a first threshold.

[0085] In this embodiment, the specific processes of S402 to S405 can all refer to the introduction of S102 in the method embodiment shown in FIG1 , and will not be repeated here.

[0086] It should be noted that if only one computing device, such as a server or a mobile phone, is to be deployed, S402 is skipped and the process starts from S403. If the computing device only has one processor, S403 and S404 are skipped and S405 is executed directly. If a processor unit of the computing device only has one processor, the deployment of that processor unit does not skip S404 and the process goes directly to S405. If a processor of the computing device only has one core, the deployment of that processor does not skip S405.

[0087] In one embodiment, the methods shown in Figures 1 and 4 can be applied to a graph engine (GE), which runs in the task deployment device described above. By applying the methods of Figures 1 or 4, the GE can significantly improve the efficiency of a hybrid computing system when executing artificial intelligence (AI) model training scenarios. The computing system includes 8 identical CPU chips and 8 identical NPU chips. As shown in Figure 5, each CPU has a computing power of approximately 62.6T FP64 and each NPU has a computing power of approximately 600T FP16. Converting the CPU computing power to FP16 is equivalent to 250T FP16 computing power. This results in a "meta-computing power" of 50T FP16. Therefore, a single CPU has 5 meta-computing powers and a single NPU has 12 meta-computing powers. After segmenting the computational graph of a business in an artificial intelligence (AI) training scenario based on the total meta-computing power and aggregating the subsets, testing has shown that the efficiency of the computing system when executing the computational graph of the business can be significantly improved.

[0088] Based on the method embodiments shown in Figures 1 and 4 , the present application also provides a task deployment device. The task deployment device is used to execute each step in the method embodiment shown in Figures 2 or 4 .

[0089] FIG6 is a schematic diagram of the structure of a task deployment apparatus 600 provided in an embodiment of the present application. As shown in FIG6 , the method may include an analysis module 601 and a deployment module 602. The analysis module 601 and the deployment module 602 may be software programs or hardware devices. When the analysis module 601 and the deployment module 602 are software, the analysis module 601 and the deployment module 602 may be provided in different hardware devices, or may be provided in different hardware devices.

[0090] The analysis module 601 is used to determine the computation graph to be deployed. In the process of determining the computation graph, the weights of each node and each edge in the computation graph are determined.

[0091] The deployment module 602 is used to allocate multiple computing tasks to multiple computing units based on the computing power of the multiple computing units corresponding to the computation graph. In the embodiment of the method shown in Figure 4, the deployment module 602 is executed multiple times to achieve layered deployment, thereby balancing the computing power load between computing units of different computing power at a finer granularity, so that the computing power load matches their computing power.

[0092] It should be noted that the task deployment device 600 provided in the embodiment shown in FIG6 only uses the division of the above-mentioned functional modules as an example when executing the task deployment method. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the computing device provided in the above embodiment and the task deployment method embodiment shown in FIG1 or FIG4 belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0093] FIG7 is a schematic diagram of the hardware structure of a computing device 700 provided in an embodiment of the present application.

[0094] The computing device 700 may include the aforementioned task deployment device. Referring to FIG7 , the computing device 700 includes a processor 701, a memory 702, a communication interface 703, and a bus 704. The processor 701, the memory 702, and the communication interface 703 are interconnected via the bus 704. The processor 701, the memory 702, and the communication interface 703 may also be connected using other connection methods besides the bus 704.

[0095] Processor 701 may be a general-purpose processor, which may be a processor that performs specific steps and / or operations by reading and executing content stored in a memory (e.g., memory 702). For example, a general-purpose processor may be a central processing unit (CPU). Processor 701 may include at least one circuit to execute all or part of the steps of the task deployment method provided in the embodiment shown in FIG. 1 or FIG. 4 . Processor 701 may include one or more cores.

[0096] Memory 702 can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical storage, hard disk, etc. As shown in Figure 7, memory 702 can be used to analyze the program code corresponding to the module and the deployment module. When the processor 701 executes the program code, the above-mentioned task deployment method is implemented.

[0097] Communication interface 703 includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, which are used to interconnect components within computing device 700, as well as interfaces for interconnecting computing device 700 with other devices (e.g., other computing devices or user equipment). Physical interfaces can include Ethernet interfaces, fiber optic interfaces, ATM interfaces, and the like.

[0098] The bus 704 may be any type of communication bus for interconnecting the processor 701 , the memory 702 , and the communication interface 703 , such as a system bus.

[0099] The above-mentioned devices can be provided on separate chips, or at least partially or entirely on the same chip. Whether to provide each device independently on different chips or to integrate them on one or more chips often depends on the product design requirements. The embodiments of this application do not limit the specific implementation of the above-mentioned devices.

[0100] The computing device 700 shown in FIG. 7 is merely exemplary. During implementation, the computing device 700 may further include other components, which are not listed here one by one.

[0101] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0102] It is understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not intended to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the order of the sequence numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0103] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.

Claims

1. A task deployment method, characterized in that: The method comprises: Determine a computation graph to be deployed, wherein the computation graph includes multiple computation tasks; According to the computing power of multiple computing units corresponding to the computing graph, the multiple computing tasks are allocated to the multiple computing units, wherein the computing power of at least two of the multiple computing units is different, the ratio of the computing power loads of the multiple computing units and the difference between the ratio of the computing power of the multiple computing units are equal to 0 or less than a first threshold, and the computing power load of each computing unit is determined according to the sum of the computing power required for the computing tasks corresponding to the each computing unit.

2. The method according to claim 1, characterized in that The allocating the plurality of computing tasks to the plurality of computing units according to the computing power of the plurality of computing units comprises: Determine the number N of segments of the computation graph according to the computing power of the plurality of computing units; Divide the computation graph into N subsets according to N, each subset in the N subsets includes one or more of the multiple computing tasks, the difference in computing power required by any two subsets in the N subsets is equal to 0 or less than a second threshold, and the computing power required by each subset is the sum of the computing power required by the computing tasks in each subset; The N subsets are assigned to the plurality of computing units.

3. The method according to claim 2, characterized in that The determining, according to the computing power of the plurality of computing units, the number N of segments of the computing graph includes: Determine a meta-computing power according to a common divisor of the computing powers of the multiple computing units, where the meta-computing power represents a unit computing power of the multiple computing units; Calculating the quotient of the computing power of the multiple computing units and the meta-computing power to obtain the number of meta-computing powers of the multiple computing units; The sum of the number of element computing powers of the multiple computing units is calculated to obtain the N.

4. The method according to claim 3, characterized in that Before determining the element computing power according to the common divisor of the computing power of the plurality of computing units, the method further comprises: The computing power of the multiple computing units is updated according to the time when the multiple computing units execute the same computing task.

5. The method according to any one of claims 2 to 4, characterized in that: The dividing the computation graph into N subsets according to N includes: The computation graph is divided with the goal of minimizing the data transmission amount of the subsets to obtain the N subsets, where the data transmission amount of the subsets is the sum of the data transmission amounts of the computing tasks in the subsets.

6. The method according to any one of claims 1 to 5, characterized in that: Before allocating the plurality of computing tasks to the plurality of computing units according to the computing power of the plurality of computing units, the method further includes: A target computing task is removed from the multiple computing tasks, wherein the target computing task includes predetermined tasks to be performed by each of the computing units.

7. The method according to any one of claims 1 to 6, characterized in that: The computing unit includes a server, a processor in the server, or a core in the processor.

8. A task deployment device, characterized in that: The device comprises: An analysis module, used to determine a computational graph to be deployed, wherein the computational graph includes multiple computational tasks; A deployment module is used to allocate the multiple computing tasks to the multiple computing units according to the computing power of the multiple computing units corresponding to the computing graph, wherein the computing power of at least two of the multiple computing units is different, the difference between the ratio of the computing power loads of the multiple computing units and the ratio of the computing power of the multiple computing units is equal to 0 or less than a first threshold, and the computing power load of each computing unit is determined according to the sum of the computing power required for the computing tasks corresponding to each computing unit.

9. The device according to claim 8, characterized in that The deployment module is specifically used for: Determine the number N of segments of the computation graph according to the computing power of the plurality of computing units; Divide the computation graph into N subsets according to N, each subset in the N subsets includes one or more of the multiple computing tasks, the difference in computing power required by any two subsets in the N subsets is equal to 0 or less than a second threshold, and the computing power required by each subset is the sum of the computing power required by the computing tasks in each subset; The N subsets are assigned to the plurality of computing units.

10. The device according to claim 9, characterized in that The deployment module is specifically used for: Determine a meta-computing power according to a common divisor of the computing powers of the multiple computing units, where the meta-computing power represents a unit computing power of the multiple computing units; Calculating the quotient of the computing power of the multiple computing units and the meta-computing power to obtain the number of meta-computing powers of the multiple computing units; The sum of the number of element computing powers of the multiple computing units is calculated to obtain the N.

11. The device according to claim 10, characterized in that The deployment module is specifically used to: before determining the meta-computing power according to the common divisor of the computing power of the multiple computing units, update the computing power of the multiple computing units according to the time when the multiple computing units perform the same computing task.

12. The device according to any one of claims 9 to 11, characterized in that: The deployment module is specifically used for: The computation graph is divided with the goal of minimizing the data transmission amount of the subsets to obtain the N subsets, where the data transmission amount of the subsets is the sum of the data transmission amounts of the computing tasks in the subsets.

13. The device according to any one of claims 8 to 12, characterized in that: The deployment module is specifically used to: before allocating the multiple computing tasks to the multiple computing units according to the computing power of the multiple computing units, remove the target computing tasks from the multiple computing tasks, and the target computing tasks include the predetermined tasks to be performed by the respective computing units.

14. The device according to any one of claims 8 to 13, characterized in that: The computing unit includes a server, a processor unit in the server, a processor in the server, or a core in the processor.

15. A computing device, characterized in that: The computing device comprises: a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Task deployment method and task deployment device

    CN120045306A

  • Neural network model processing method and device

    CN116187391A

  • Neural network computational graph compiling method and device

    CN116931941A

  • Method and system for distributed computation

    US20110041136A1

  • Graph computing method and apparatus

    US20220043675A1