Resource scheduling method and device for task flow, electronic equipment and storage medium

By iteratively updating the directed acyclic graph of the computing task, recombining heterogeneous computing resources and adjusting the relationship between the task nodes, the problem of unreasonable scheduling of heterogeneous computing resources is solved, the task execution efficiency is improved and the system delay is reduced.

CN120144293APending Publication Date: 2025-06-13CISDI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226093.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The inability to reasonably schedule various heterogeneous computing resources leads to high system delays and low task execution efficiency.

Method used

By obtaining the initial directed acyclic graph of the computing task, monitoring the resource occupancy rate and system delay of the task nodes, iteratively update the graph structure to recombinate heterogeneous computing resources and adjust the serial and parallel relationship of the task nodes, and optimizing resource allocation to reduce system delay.

Benefits of technology

It improves the rationality of heterogeneous computing resources and task execution efficiency, and reduces the system delay of computing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144293A_ABST
    Figure CN120144293A_ABST
Patent Text Reader

Abstract

The invention provides a task flow resource scheduling method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining an initial directed acyclic graph of a calculation task, executing an initial task flow according to the initial directed acyclic graph, and monitoring the resource occupancy rate of each task node in the initial task flow, if it is detected that at least one resource occupancy rate is larger than or equal to a corresponding preset occupancy rate threshold value, iterative updating is carried out on the initial directed acyclic graph to obtain a current directed acyclic graph, and iterative updating comprises recombining heterogeneous computing resources distributed to task nodes and updating at least one of the serial-parallel relation of the task nodes; if it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, monitoring the system time delay of the initial directed acyclic graph, so as to reduce the system time delay through the iterative update of the initial directed acyclic graph; according to the method, the rationality of heterogeneous computing resource allocation and the task execution efficiency are improved, and the system time delay of the computing task is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of task scheduling, and in particular, to a resource scheduling method, device, electronic device, and storage medium for a task flow. Background Art

[0002] In recent years, deep neural networks (DNNs) have become the core force driving the development of artificial intelligence technology and are widely used in many fields such as autonomous driving, intelligent monitoring, speech recognition, and image classification. Nowadays, models trained by various deep learning frameworks, such as TensorFlow Lite, PyTorch, and MindSpore, can be deployed on servers with graphics processing unit (GPU) computing power and also begin to support deployment on edge devices with artificial intelligence (AI) computing power through adaptation. Edge computing enables intelligent devices to perform complex inference tasks locally, avoiding high-latency cloud processing and driving the development of the AI edge computing field. With the availability of AI computing power in edge computing, how to improve the deployment efficiency and inference efficiency of end-side models on edge devices has become a key issue in model deployment research.

[0003] Regarding the problem of efficient and fast deployment, a solution to establish an abstract deployment process of a syntax tree has been proposed based on the differences between the training environment and the running environment. For example, a method of abstracting the additional processing process and the inference process of the model through a syntax tree and merging them into a target code for direct deployment on edge devices. This method solves the problem of low model deployment efficiency caused by the differences in the running environment of the model on end devices.

[0004] To address the problem of low latency in inference, the general idea is to divide the inference task into multiple directed acyclic task flows to form a directed acyclic graph (DAG), mainly to improve the inference efficiency in the division strategy of inference tasks and the resource allocation strategy of task flows. The mainstream task division methods are mostly to divide the model itself into task flows or perform task compression, such as model compression, model segmentation, etc. Model compression aims to reduce the computational workload and storage requirements of the neural network while maintaining the accuracy of the model as much as possible; the resource allocation strategy of the task flow is mainly cross-device and cross-system allocation, that is, to decompose the large neural network model into multiple smaller sub-models or modules, and distribute these sub-models to multiple edge devices or work together between edge devices and the cloud to achieve model inference acceleration. However, cross-device collaboration places high demands on the scheduling strategy of the task flow. Edge devices usually have lower computing resources and bandwidth limitations. Early edge devices had significant bottlenecks in the central processing unit (CPU) and memory. However, with the development of chip technology and AI technology, CPU and memory are no longer the main resource bottlenecks in edge computing devices. In order to achieve low-latency processing, edge computing integrates a variety of heterogeneous computing resources. At present, it is not possible to reasonably schedule various heterogeneous computing resources to reduce inference latency and improve model inference efficiency. Summary of the invention

[0005] The present invention provides a method, device, electronic device and storage medium for resource scheduling of task flow to solve the above-mentioned technical problem that various heterogeneous computing resources cannot be reasonably scheduled to reduce system latency and improve task execution efficiency.

[0006] In one embodiment of the present application, the present application provides a method for resource scheduling of a task flow, comprising: obtaining an initial directed acyclic graph of a computing task, the initial directed acyclic graph being used to characterize a mapping relationship between an initial task flow and allocated heterogeneous computing resources; executing the initial task flow according to the initial directed acyclic graph, and monitoring the resource occupancy rate of each task node in the initial task flow; if it is detected that at least one resource occupancy rate is greater than or equal to a corresponding preset occupancy rate threshold, iteratively updating the initial directed acyclic graph to obtain a current directed acyclic graph, the iterative update comprising at least one of recombining heterogeneous computing resources allocated to the task node and updating a serial-parallel relationship of the task node; if it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, monitoring the system delay of the initial directed acyclic graph to reduce the system delay through iterative updating of the initial directed acyclic graph, the system delay being used to characterize the processing delay of the longest dependent path from the start node to the end node of the initial task flow.

[0007] In an embodiment of the present application, reducing the system latency through iterative update of the initial directed acyclic graph includes: if the reduction of the system latency meets a preset convergence condition, or the total number of iterative updates is greater than or equal to a preset number threshold, determining the initial directed acyclic graph as the target directed acyclic graph; if the reduction of the system latency does not meet the preset convergence condition, or the total number of iterative updates is less than the preset number threshold, performing iterative update on the initial directed acyclic graph to obtain the current directed acyclic graph.

[0008] In an embodiment of the present application, after obtaining the current directed acyclic graph, it further includes: using the current directed acyclic graph as the initial directed acyclic graph; repeatedly executing the initial task flow according to the initial directed acyclic graph and monitoring the resource occupancy rate of each task node in the initial task flow. If it is detected that at least one resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, performing iterative update on the initial directed acyclic graph to obtain the current directed acyclic graph. If it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, monitoring the system latency of the initial directed acyclic graph to reduce the system latency through iterative update of the initial directed acyclic graph, and using the current directed acyclic graph as the initial directed acyclic graph.

[0009] In an embodiment of the present application, recombining the heterogeneous computing resources allocated to the task nodes includes: polling the idle heterogeneous computing resources and establishing a new mapping relationship between the idle heterogeneous computing resources and the first optimization node, where the first optimization node is used to represent a task node whose resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold; or, adjusting the allocation ratio and allocation type of the heterogeneous computing resources of the second optimization node according to the longest dependency path of the initial task flow and the preset allocation adjustment step, where the second optimization node is used to represent a task node on the longest dependency path.

[0010] In an embodiment of the present application, before at least one of recombining the heterogeneous computing resources allocated to the task nodes and updating the serial-parallel relationship of the task nodes, it further includes: adjusting at least one of the heterogeneous computing resources allocated to each sub-operation and the serial-parallel relationship of the sub-operation based on the scheduling order of multiple sub-operations in the third optimization node to obtain the target optimization node; updating the target optimization node to the initial directed acyclic graph to continue monitoring the resource occupancy rate of each task node in the updated initial directed acyclic graph; where the third optimization node is used to represent a task node with sub-operations in the first optimization node and the second optimization node.

[0011] In an embodiment of the present application, the determination of the system latency includes: determining the node latency of a task node according to the workload and resource allocation amount of the task node, where the workload and the resource allocation amount are inversely proportional; and determining the sum of the corresponding node latencies as the system latency according to the longest dependency path of the initial task flow.

[0012] In an embodiment of the present application, the determination of the initial directed acyclic graph includes: splitting the computing task into multiple task steps according to the business scenario of the computing task; obtaining an initial task flow according to the dependency relationships of the task steps, where the initial task flow includes multiple task nodes, and the edges between the task nodes are used to represent the data flow and dependency relationships; integrating the resource values of heterogeneous computing resources; and establishing a mapping relationship between the heterogeneous computing resources and the initial task flow based on the integrated resource values to obtain an initial directed acyclic graph, so as to allocate the heterogeneous computing resources to each task node.

[0013] In an embodiment of the present application, the present application provides a resource scheduling device for a task flow, including: a task acquisition module, configured to acquire an initial directed acyclic graph of a computing task, where the initial directed acyclic graph is used to represent the mapping relationship between an initial task flow and allocated heterogeneous computing resources; an execution monitoring module, configured to execute the initial task flow according to the initial directed acyclic graph and monitor the resource occupancy rate of each task node in the initial task flow; an update module, configured to, if it is detected that at least one resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, perform iterative update on the initial directed acyclic graph to obtain a current directed acyclic graph, and recombine at least one of the heterogeneous computing resources allocated to the task nodes and update the serial-parallel relationship of the task nodes; and a latency monitoring module, configured to, if it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, monitor the system latency of the initial directed acyclic graph, so as to reduce the system latency through the iterative update of the initial directed acyclic graph, where the system latency is used to represent the processing latency of the longest dependency path of the initial task flow from the start node to the end node.

[0014] In an embodiment of the present application, the present application provides an electronic device, where the electronic device includes: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the electronic device to implement the resource scheduling method for a task flow according to any one of the above embodiments.

[0015] In an embodiment of the present application, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of a computer, enables the computer to execute the resource scheduling method for a task flow according to any one of the above embodiments.

[0016] Advantages of the embodiments of the present invention: The present application provides a method, apparatus, electronic device, and storage medium for resource scheduling of a task flow. In the embodiments of the present invention, when the resource occupancy rate is greater than or equal to a preset occupancy rate threshold, the initial directed acyclic graph can be iteratively updated first. On the basis that the resource occupancy rate is less than the preset occupancy rate threshold, with the goal of reducing the system latency, through the iterative update of the initial directed acyclic graph, the scheduling of heterogeneous computing resources is completed, improving the rationality of heterogeneous computing resource allocation and the task execution efficiency, and reducing the system latency of computing tasks.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are incorporated herein and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0019] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown;

[0020] Figure 2 A flowchart showing the resource scheduling method of a task flow according to an embodiment of the present application is shown;

[0021] Figure 3 A schematic diagram of an initial directed acyclic graph according to an embodiment of the present application is shown;

[0022] Figure 4 An implementation flowchart showing the resource scheduling method of a task flow according to an embodiment of the present application is shown;

[0023] Figure 5 A schematic diagram of an initial directed acyclic graph according to another embodiment of the present application is shown;

[0024] Figure 6 A schematic diagram of a current directed acyclic graph according to an embodiment of the present application is shown;

[0025] Figure 7 A schematic diagram of a current directed acyclic graph according to another embodiment of the present application is shown;

[0026] Figure 8 A block diagram showing the resource scheduling apparatus of a task flow according to an embodiment of the present application is shown;

[0027] Figure 9 The figure shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0028] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0029] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The form, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout form may also be more complex.

[0030] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.

[0031] Please refer to Figure 1 , Figure 1 The figure shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied. As Figure 1 shown, the system architecture 100 may include a resource acquisition module 101, a resource monitoring module 102, a DAG generation module 103, a dynamic balancer 104, a task scheduling module 105, and a latency monitoring module 106. Among them, the resource acquisition module 101 is used to acquire the types and available numbers of heterogeneous computing resources, the resource monitoring module 102 is used to monitor the resource occupancy rate of each task node, the DAG generation module 103 is used to generate the current directed acyclic graph, the dynamic balancer 104 is used for iterative updating of the initial directed acyclic graph, the task scheduling module 105 is used to execute the initial task flow according to the initial directed acyclic graph to achieve the scheduling of the initial task flow, and the latency monitoring module 106 is used to monitor the processing latency of each task node to obtain the system latency.

[0032] Exemplarily, an initial directed acyclic graph of a computing task is obtained. The initial directed acyclic graph is used to represent the mapping relationship between an initial task flow and allocated heterogeneous computing resources; the initial task flow is executed according to the initial directed acyclic graph, and the resource occupancy rate of each task node in the initial task flow is monitored; if it is detected that the resource occupancy rate of at least one is greater than or equal to the corresponding preset occupancy rate threshold, the initial directed acyclic graph is iteratively updated to obtain the current directed acyclic graph. The iterative update includes at least one of recombining the heterogeneous computing resources allocated to the task nodes and updating the serial-parallel relationship of the task nodes; if it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, the system latency of the initial directed acyclic graph is monitored to reduce the system latency through the iterative update of the initial directed acyclic graph. The system latency is used to represent the processing latency of the longest dependency path of the initial task flow from the start node to the end node.

[0033] In the related art, various heterogeneous computing resources cannot be reasonably scheduled to reduce the system latency and improve the task execution efficiency.

[0034] To solve the above technical problems, the present application provides a resource scheduling method, device, electronic device and storage medium for a task flow. The implementation details of the technical solutions of the embodiments of the present application are elaborated in detail below.

[0035] Please refer to Figure 2 , Figure 2 which shows a schematic flowchart of a resource scheduling method for a task flow according to an embodiment of the present application. As Figure 2 shown, in an exemplary embodiment, the resource scheduling method for a task flow at least includes steps S210 to S240, which are introduced in detail as follows:

[0036] Step S210, obtain an initial directed acyclic graph of a computing task.

[0037] Among them, the initial directed acyclic graph is used to represent the mapping relationship between the initial task flow and the allocated heterogeneous computing resources.

[0038] In an embodiment of the present application, the computing task includes an inference task, a scientific computing task, etc. The computing task may also include non-visual tasks such as a vision task or a natural language processing task.

[0039] In an embodiment of the present application, the determination of the initial directed acyclic graph includes: splitting the computing task into multiple task steps according to the business scenario of the computing task; obtaining an initial task flow according to the dependency relationship of each task step. The initial task flow includes multiple task nodes, and the edges between the task nodes are used to represent the data flow and the dependency relationship; integrating the resource values of the heterogeneous computing resources; establishing a mapping relationship between the heterogeneous computing resources and the initial task flow based on the integrated resource values to obtain the initial directed acyclic graph, so as to allocate the heterogeneous computing resources to each task node.

[0040] In one embodiment of the present application, the task steps include at least one of preprocessing, model inference, tensor decoding, and postprocessing.

[0041] In one embodiment of the present application, the resource value for integrating heterogeneous computing resources includes normalizing the resource values of various types of heterogeneous computing resources.

[0042] In one embodiment of the present application, in an AI vision inference system, heterogeneous computing resources may include at least one of memory, Graphics Processing Unit (GPU), Tensor Processing Unit (TPU), Central Processing Unit (CPU), Vision Processing Unit (VPU), and Look-Up Table (LUT) in a Field-Programmable Gate Array (FPGA).

[0043] In one embodiment of the present application, the CPU is a general-purpose computing unit suitable for multitasking, and is good at processing complex logic, control flow, and various types of operations. The main frequency and the number of cores are resource values; the TPU is dedicated to accelerating tensor operations in deep learning models, such as matrix multiplication and convolution operations, and is generally measured in TOPS, which represents the number of operations that can be processed per second, in trillions; the VPU can efficiently process image and video data, such as image scaling, denoising, video encoding and decoding, etc., and is generally related to computing memory; the FPGA is good at parallel computing of customized algorithms, and generally uses the number of LUTs to represent the resource value; the memory is divided into system memory, TPU memory, and VPU memory. The memory is a storage unit in the system, used to store data, instructions, and cache intermediate data. Normalize and convert the resource values of various types of heterogeneous computing resources for convenient modeling.

[0044] In one embodiment of the present application, the dependencies include the serial order, parallel order, and data transfer direction of each task node.

[0045] In one embodiment of the present application, the present application is applicable to any directed acyclic task scenario with a serial or parallel combination.

[0046] In one embodiment of the present application, analyze the characteristics of each task node and the characteristics of heterogeneous computing resources, and allocate different heterogeneous computing resources to each task node to form an initial Directed Acyclic Graph (DAG). Please refer to Figure 3 , Figure 3A schematic diagram of an initial directed acyclic graph according to an embodiment of the present application is shown. As Figure 3 shown, T1-T4 are task nodes, and the same task nodes represent a parallel relationship, indicating that the task can be parallelized in multiple threads; R1-R4 represent heterogeneous computing resources available, and x is used to characterize the allocation ratio; the dynamic balancer is used to iteratively update the initial directed acyclic graph.

[0047] In an embodiment of the present application, after obtaining the initial directed acyclic graph of the computing task, it further includes: decomposing a task node into multiple sub-operations; determining the scheduling order of each sub-operation according to the dependency relationship of each sub-operation; and reallocating the heterogeneous computing resources of the task node to each sub-operation based on the scheduling order.

[0048] In an embodiment of the present application, for example, the sub-operations corresponding to the pre-processing task include at least one of resizing the picture, normalizing, and data augmentation.

[0049] Step S220, execute the initial task flow according to the initial directed acyclic graph, and monitor the resource occupancy rate of each task node in the initial task flow.

[0050] In an embodiment of the present application, the resource monitoring module monitors the resource occupancy rate of each task node and transmits the monitoring result to the dynamic balancer to iteratively update the initial directed acyclic graph.

[0051] Step S230, if it is detected that the resource occupancy rate of at least one is greater than or equal to the corresponding preset occupancy rate threshold, then iteratively update the initial directed acyclic graph to obtain the current directed acyclic graph, and re-combine and allocate at least one of the heterogeneous computing resources allocated to the task nodes and update the serial-parallel relationship of the task nodes.

[0052] In an embodiment of the present application, if the resource occupancy rate of a task node is greater than or equal to the corresponding preset occupancy rate threshold, it is determined that the heterogeneous computing resources allocated to the task node are nearly fully loaded, and then the initial directed acyclic graph is iteratively updated through the dynamic balancer.

[0053] Step S240, if it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, then monitor the system latency of the initial directed acyclic graph to reduce the system latency through the iterative update of the initial directed acyclic graph.

[0054] Wherein, the system latency is used to characterize the processing latency of the longest dependency path of the initial task flow from the start node to the end node.

[0055] In one embodiment of the present application, the system latency is reduced by iterative updating of the initial directed acyclic graph, including: if the reduction of the system latency meets the preset convergence condition, or the total number of iterative updates is greater than or equal to the preset number threshold, then the initial directed acyclic graph is determined as the target directed acyclic graph; if the reduction of the system latency does not meet the preset convergence condition, or the total number of iterative updates is less than the preset number threshold, then the initial directed acyclic graph is iteratively updated to obtain the current directed acyclic graph.

[0056] In one embodiment of the present application, the determination of the system latency includes: determining the node latency of a task node according to the workload and resource allocation amount of the task node, where the workload and the resource allocation amount are inversely proportional; and determining the sum of the corresponding node latencies as the system latency according to the longest dependency path of the initial task flow.

[0057] In one embodiment of the present application, the preset convergence condition is used to determine that the system latency has gradually decreased and tends to be stable, and subsequent iterative updates will not significantly reduce the system latency. For example, the reduction change of the system latency for consecutive multiple times is less than the preset change threshold, or the system latency for consecutive multiple times is less than the preset latency threshold, etc.

[0058] In one embodiment of the present application, the system latency is determined as follows:

[0059]

[0060] where, L total is the system latency, Path k is the k-th dependency path of the initial task flow, w i is the workload of task node i, x ij is the resource percentage allocated to task node i on heterogeneous computing resource j, R j is the resource value of heterogeneous computing resource j, max k is the longest dependency path of the initial task flow.

[0061] In one embodiment of the present application, the workload is characterized by the amount of computation or the amount of data.

[0062] In one embodiment of the present application, 0 ≤ x ij ≤ 100%.

[0063] In one embodiment of the present application, if the heterogeneous computing resource is a CPU, then 0 ≤ x ij ≤ 100% × the number of cores, and the resource value is characterized by the available core rate of the CPU.

[0064] In one embodiment of the present application, allocating more resource allocation amounts to a task node can reduce the node delay of the task node. The present application needs to minimize the processing delay on the longest dependency path, so as to reduce the system delay of the computing task.

[0065] In one embodiment of the present application, recombining the heterogeneous computing resources allocated to the task node includes: polling for idle heterogeneous computing resources and establishing a new mapping relationship between the idle heterogeneous computing resources and a first optimized node, where the first optimized node is used to represent a task node with a resource occupancy rate greater than or equal to the corresponding preset occupancy rate threshold; or, adjusting the allocation ratio and allocation type of the heterogeneous computing resources of a second optimized node according to the longest dependency path of the initial task flow and a preset allocation adjustment step, where the second optimized node is used to represent a task node on the longest dependency path.

[0066] In one embodiment of the present application, when the heterogeneous computing resources allocated to a task node are approaching full load, idle heterogeneous computing resources can be allocated to the task node.

[0067] In one embodiment of the present application, after adjusting the allocation ratio and allocation type of the heterogeneous computing resources of the second optimized node, a second optimized node can also be split, or a second optimized node can be merged with a task node identical to the second optimized node, such as splitting a decoding task into two parallel decoding tasks.

[0068] In one embodiment of the present application, before recombining at least one of the heterogeneous computing resources allocated to the task node and updating the serial-parallel relationship of the task node, it further includes: adjusting at least one of the heterogeneous computing resources allocated to each sub-operation and the serial-parallel relationship of the sub-operation based on the scheduling order of multiple sub-operations in a third optimized node to obtain a target optimized node; updating the target optimized node to the initial directed acyclic graph to continue monitoring the resource occupancy rate of each task node in the updated initial directed acyclic graph; where the third optimized node is used to represent a task node with sub-operations among the first optimized node and the second optimized node.

[0069] In one embodiment of the present application, before adjusting the heterogeneous computing resources and the serial-parallel relationship between task nodes, the heterogeneous computing resources and the serial-parallel relationship between sub-operations within the task node can be adjusted first.

[0070] In one embodiment of the present application, the allocation of heterogeneous computing resources is adjusted through a dynamic balancer to minimize the system delay of the longest dependency path. To control the amount of calculation, the preset allocation adjustment step is set to 10%.

[0071] In one embodiment of the present application, the preset number threshold can be set to 100.

[0072] After obtaining the current directed acyclic graph in an embodiment of the present application, the method further includes: using the current directed acyclic graph as the initial directed acyclic graph; repeatedly executing an initial task flow according to the initial directed acyclic graph and monitoring the resource occupancy rate of each task node in the initial task flow. If it is detected that at least one resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, the initial directed acyclic graph is iteratively updated to obtain the current directed acyclic graph. If it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, the system latency of the initial directed acyclic graph is monitored, so as to reduce the system latency through the iterative update of the initial directed acyclic graph, and use the current directed acyclic graph as the initial directed acyclic graph in the subsequent step.

[0073] In an embodiment of the present application, please refer to Figure 4 , Figure 4 which shows a schematic implementation flowchart of a resource scheduling method for a task flow according to an embodiment of the present application. As Figure 4 shown, an initial task flow is established: the computing task is split into multiple task steps, and according to the dependency relationships of the task steps, the initial task flow is obtained. For example, it is split into a pre-processing service, an inference service, a decoding service, and a post-processing service. Each task is forward-dependent on the previous task, and there is only one path in this initial task flow, that is, k = 1. The minimum processing latency of this path is the minimum system latency of this initial task flow; heterogeneous computing resources are integrated: the resource values of various heterogeneous computing resources are normalized to complete the resource initialization work. The heterogeneous computing resources include at least one of CPU, TPU, VPU, and memory; a mapping relationship is established: a mapping relationship is established between the initial task flow and the allocated heterogeneous computing resources to obtain an initial directed acyclic graph, so as to allocate the normalized heterogeneous computing resources to each task node in the initial task flow; task scheduling: the initial directed acyclic graph is loaded and each task node is started; resource monitoring: the resource occupancy rate of each task node can be monitored through a resource monitoring module; if it is detected that at least one resource occupancy rate is greater than or equal to the preset occupancy rate threshold, the step of dynamically optimizing the mapping relationship is entered. If it is not detected that the resource occupancy rate is greater than or equal to the preset occupancy rate threshold, the latency monitoring step is entered; dynamically optimizing the mapping relationship: the heterogeneous computing resources can be recombined according to a preset allocation adjustment step by a dynamic balancer to establish a new mapping relationship with the corresponding task node; DAG generation: a current directed acyclic graph is generated according to the new mapping relationship and the dependency relationships between the task nodes in the initial directed acyclic graph, so as to use the current directed acyclic graph as the initial directed acyclic graph and enter the task scheduling step; latency monitoring: if the system latency converges, or the total number of iterative updates is greater than or equal to the preset number threshold, the output step is entered. If the system latency does not converge, or the total number of iterative updates is less than the preset number threshold, the step of dynamically optimizing the mapping relationship is entered; output: the initial directed acyclic graph is determined as the target directed acyclic graph and output.

[0074] In one embodiment of the present application, please refer to Figure 5 , Figure 5 which shows a schematic diagram of an initial directed acyclic graph according to another embodiment of the present application. As Figure 5 shown, the processing process of the initial task flow is single-threaded serial processing, which only uses one CPU core and TPU to execute serial tasks. The serial tasks cannot utilize the computing power of other available cores, resulting in slow processing speed, unable to fully utilize the advantages of hardware resources, and causing resource waste.

[0075] In one embodiment of the present application, please refer to Figure 6 , Figure 6 which shows a schematic diagram of the current directed acyclic graph according to one embodiment of the present application. As Figure 6 shown, in the intermediate stage of allocating heterogeneous computing resources to task nodes through a dynamic balancer, this method uses CPU multi-core processing and uses a VPU unit according to the requirements of task nodes in the pre-processing stage; the allocation of more heterogeneous computing resources speeds up the overall inference speed and reduces the system latency.

[0076] In one embodiment of the present application, please refer to Figure 7 , Figure 7 which shows a schematic diagram of the current directed acyclic graph according to another embodiment of the present application. As Figure 7 shown, when the system latency is close to the convergence state and the decoding task is assigned to the TPU resource, the system latency further decreases in a gradient manner. Compared with the Figure 6 stage, in this stage, the dynamic balancer analyzes and obtains the optimal allocation strategy of task nodes and heterogeneous computing resources, and according to the characteristics of task nodes, allocates them to the optimal heterogeneous computing resources for processing, so that the system latency is the lowest, the inference speed is the fastest, and the heterogeneous computing resources are the largest.

[0077] Please refer to Figure 8 , Figure 8 which shows a block diagram of a resource scheduling device for a task flow according to one embodiment of the present application. This device can be applied to the Figure 1 shown implementation environment. This device can also be applicable to other exemplary implementation environments and is specifically configured in other devices. This embodiment does not limit the implementation environment applicable to this device.

[0078] As Figure 8 shown, a resource scheduling device 800 for a task flow according to one embodiment of the present application includes: a task acquisition module 801, an execution monitoring module 802, an update module 803, and a latency monitoring module 804.

[0079] Among them, the task acquisition module 801 is used to acquire an initial directed acyclic graph of a computing task, and the initial directed acyclic graph is used to represent the mapping relationship between the initial task flow and the allocated heterogeneous computing resources;

[0080] An execution monitoring module 802 is configured to execute an initial task flow according to an initial directed acyclic graph and monitor the resource occupancy rate of each task node in the initial task flow;

[0081] An update module 803 is configured to, if it is detected that at least one resource occupancy rate is greater than or equal to a corresponding preset occupancy rate threshold, iteratively update the initial directed acyclic graph to obtain a current directed acyclic graph, and recombine and allocate at least one of the heterogeneous computing resources allocated to the task nodes and update the serial-parallel relationship of the task nodes;

[0082] A latency monitoring module 804 is configured to, if it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, monitor the system latency of the initial directed acyclic graph, so as to reduce the system latency through the iterative update of the initial directed acyclic graph, and the system latency is used to characterize the processing latency of the longest dependency path of the initial task flow from the start node to the end node.

[0083] It should be noted that the task flow resource scheduling device provided in the above embodiments and the task flow resource scheduling method provided in the above embodiments belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiments, and will not be elaborated here. In practical applications, the task flow resource scheduling device provided in the above embodiments can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here either.

[0084] An embodiment of the present application further provides an electronic device, including: one or more processors; a storage device configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the task flow resource scheduling method provided in each of the above embodiments.

[0085] Please refer to Figure 9 , Figure 9 which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. It should be noted that Figure 9 the computer system 900 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0086] Such as Figure 9As shown, computer system 900 includes a Central Processing Unit (CPU) 901, which can perform various appropriate actions and processes according to programs stored in a Read-Only Memory (ROM) 902 or programs loaded from a storage section 908 into a Random Access Memory (RAM) 903, such as executing the methods described in the above embodiments. In the RAM 903, various programs and data required for system operations are also stored. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An Input / Output (I / O) interface 905 is also connected to the bus 904.

[0087] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read therefrom can be installed into the storage section 908 as needed.

[0088] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by a Central Processing Unit (CPU) 901, various functions defined in the system of the present application are executed.

[0089] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0091] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the units themselves. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of this application.

[0092] Another aspect of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of the computer, the computer is enabled to execute the resource scheduling method of the task flow provided in each of the above embodiments. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device.

[0093] In the above embodiments, unless otherwise specified, when using serial numbers such as "first" and "second" to describe a common object, it only indicates different instances of the same object, rather than indicating that the object to be described must be in a given order, whether in terms of time, space, sorting, or any other way.

[0094] The above embodiments are only used to exemplarily illustrate the principles and effects of this application, rather than to limit this application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in this application should still be covered by the claims of this application.

Claims

1. A method for resource scheduling of a task flow, characterized in that: The method comprises: Acquire an initial directed acyclic graph of the computing task, where the initial directed acyclic graph is used to characterize a mapping relationship between an initial task flow and allocated heterogeneous computing resources; Executing the initial task flow according to the initial directed acyclic graph, and monitoring the resource occupancy rate of each task node in the initial task flow; If it is detected that at least one resource occupancy rate is greater than or equal to a corresponding preset occupancy rate threshold, the initial directed acyclic graph is iteratively updated to obtain a current directed acyclic graph, wherein the iterative update includes at least one of recombining the heterogeneous computing resources allocated to the task node and updating the serial-parallel relationship of the task node; If it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, the system delay of the initial directed acyclic graph is monitored to reduce the system delay through iterative update of the initial directed acyclic graph. The system delay is used to characterize the processing delay of the longest dependent path of the initial task flow from the start node to the end node.

2. The method for resource scheduling of task flow according to claim 1, characterized in that: Reducing the system latency by iteratively updating the initial directed acyclic graph includes: If the reduction of the system delay satisfies a preset convergence condition, or the total number of iterative updates is greater than or equal to a preset number threshold, the initial directed acyclic graph is determined as a target directed acyclic graph; If the reduction of the system delay does not meet the preset convergence condition, or the total number of iterative updates is less than a preset threshold, the initial directed acyclic graph is iteratively updated to obtain a current directed acyclic graph.

3. The method for resource scheduling of task flow according to claim 2, characterized in that: After getting the current directed acyclic graph, it also includes: Using the current directed acyclic graph as an initial directed acyclic graph; Repeat the process of executing the initial task flow according to the initial directed acyclic graph, and monitor the resource occupancy rate of each task node in the initial task flow; if it is detected that at least one resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, iteratively update the initial directed acyclic graph to obtain a current directed acyclic graph; if it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, monitor the system delay of the initial directed acyclic graph to reduce the system delay through iterative update of the initial directed acyclic graph, and use the current directed acyclic graph as the initial directed acyclic graph.

4. The method for resource scheduling of task flow according to any one of claims 1 to 3, characterized in that: Recombining the heterogeneous computing resources allocated to the task node, including: Polling idle heterogeneous computing resources, and establishing a new mapping relationship between the idle heterogeneous computing resources and a first optimization node, where the first optimization node is used to characterize a task node whose resource occupancy rate is greater than or equal to a corresponding preset occupancy rate threshold; or, According to the longest dependency path of the initial task flow and a preset allocation adjustment step, the allocation ratio and allocation type of the heterogeneous computing resources of the second optimization node are adjusted, and the second optimization node is used to characterize the task node on the longest dependency path.

5. The method for resource scheduling of task flow according to claim 4, characterized in that: Before at least one of recombining the heterogeneous computing resources allocated to the task node and updating the serial-parallel relationship of the task node, the method further includes: Based on the scheduling order of multiple sub-operations in the third optimization node, adjust at least one of the heterogeneous computing resources allocated to each of the sub-operations and the serial-parallel relationship of the sub-operations to obtain a target optimization node; Updating the target optimization node to the initial directed acyclic graph to continue monitoring the resource occupancy rate of each task node in the updated initial directed acyclic graph; The third optimization node is used to represent the task node where sub-operations exist in the first optimization node and the second optimization node.

6. The method for resource scheduling of task flow according to any one of claims 1 to 3, characterized in that: The determination of the system delay includes: Determine a node delay of a task node according to a workload and a resource allocation amount of the task node, wherein the workload is inversely proportional to the resource allocation amount; According to the longest dependent path of the initial task flow, the sum of the corresponding node delays is determined as the system delay.

7. The method for resource scheduling of task flow according to any one of claims 1 to 3, characterized in that: The determination of the initial directed acyclic graph includes: Splitting the computing task into multiple task steps according to the business scenario of the computing task; According to the dependency relationship of each task step, an initial task flow is obtained, wherein the initial task flow includes a plurality of task nodes, and the edges between the task nodes are used to represent the data flow direction and the dependency relationship; Integrate the resource value of heterogeneous computing resources; A mapping relationship between the heterogeneous computing resources and the initial task flow is established based on the integrated resource values ​​to obtain an initial directed acyclic graph, so as to allocate the heterogeneous computing resources to each of the task nodes.

8. A resource scheduling device for task flow, characterized in that: The device comprises: A task acquisition module, used to acquire an initial directed acyclic graph of computing tasks, wherein the initial directed acyclic graph is used to characterize a mapping relationship between an initial task flow and allocated heterogeneous computing resources; An execution monitoring module, used for executing the initial task flow according to the initial directed acyclic graph, and monitoring the resource occupancy rate of each task node in the initial task flow; An update module, configured to iteratively update the initial directed acyclic graph to obtain a current directed acyclic graph, recombining the heterogeneous computing resources allocated to the task nodes and updating at least one of the serial-parallel relationship of the task nodes if it is detected that at least one resource occupancy rate is greater than or equal to a corresponding preset occupancy rate threshold; A delay monitoring module is used to monitor the system delay of the initial directed acyclic graph if it is not detected that the resource occupancy rate is greater than or equal to the corresponding preset occupancy rate threshold, so as to reduce the system delay by iterative updating of the initial directed acyclic graph. The system delay is used to characterize the processing delay of the longest dependent path of the initial task flow from the start node to the end node.

9. An electronic device, characterized in that: The electronic device comprises: One or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the electronic device implements the resource scheduling method for the task flow as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the resource scheduling method for task flows according to any one of claims 1 to 7.