Scheduling method and device
By calculating resource utilization index values and anti-fragmentation parameters, selecting appropriate running nodes to perform machine learning tasks, solving the problem of inefficient task scheduling caused by node resource fragmentation, and achieving efficient resource utilization and stable task execution.
Patent Information
- Application Number
- CN202510096845.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
AI Technical Summary
In the machine learning platform, the fragmentation of node resources causes the scheduler to be unable to effectively schedule new tasks to nodes with sufficient resources, resulting in waste of resources and reduced task execution efficiency.
By responding to the inference task scheduling request of the target model, obtain the amount of resources applied by the task, the total amount of resources and available amount of the running node, calculate the resource utilization index value and anti-fragmentation parameters, and select the appropriate running node to execute the task.
It effectively reduces resource fragmentation, improves the efficiency of task scheduling, avoids resource waste, and ensures the balance of internal resource utilization of operating nodes.
Smart Images

Figure CN119938336A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a scheduling method and device. Background Art
[0002] In a machine learning platform, the scheduler can schedule machine learning tasks to appropriate nodes for execution based on the cluster's resource status and task requirements. However, with the frequent scheduling of tasks and the continuous allocation and use of resources, node resource fragmentation is prone to occur (that is, even if the total resources of some nodes in the cluster appear to be sufficient, these resources are occupied by multiple tasks in different sizes, resulting in an inability to find a continuous, large enough memory space to meet the resource requirements of new tasks). In this case, even if the scheduler can recognize the existence of these nodes, it cannot effectively schedule new tasks to them, resulting in a waste of resources and a decrease in the efficiency of machine learning task execution. Summary of the invention
[0003] The technical solutions provided by this application are as follows:
[0004] The first aspect of the present application provides a scheduling method, including:
[0005] In response to a scheduling request of at least one target model-based reasoning task, obtaining a resource application amount of each reasoning task in the at least one target model-based reasoning task;
[0006] Obtaining the total amount and available amount of resources of each running node in at least one running node;
[0007] Based on the resource application amount of the inference task and the total amount and available amount of resources of each of the running nodes, determine the resource utilization index value and anti-fragmentation parameter of each of the running nodes; the resource utilization index value is used to reflect the resource utilization of the running node; the anti-fragmentation parameter is used to reflect the impact of scheduling the inference task to the running node on reducing the degree of resource fragmentation of the running node;
[0008] Based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes, at least one first target running node is selected from each of the running nodes to execute the reasoning task.
[0009] The determining of the resource utilization index value and the anti-fragmentation parameter of each operating node based on the resource application amount of the reasoning task and the total amount and available amount of resources of each operating node includes:
[0010] Determine the utilization rate of each type of resource of each running node based on the application amount of each type of resource for the reasoning task and the total amount and available amount of each type of resource of each running node; the utilization rate represents the degree of utilization of the type of resource when the reasoning task is scheduled to the running node;
[0011] Determining a resource utilization index value of each of the operating nodes based on the utilization rate of each of the types of resources of each of the operating nodes and a first weight factor; the first weight factor is used to reflect the importance of the type of resources to the reasoning task;
[0012] Based on the resource application amount of the inference task and the available amount of resources of each of the running nodes, the expected remaining amount of resources of each of the running nodes is determined, and the expected remaining amount is used as an anti-fragmentation parameter.
[0013] The selecting at least one first target running node from each of the running nodes to perform the reasoning task based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes includes:
[0014] If the anti-fragmentation parameter of the running node is used to reflect that scheduling the inference task to the running node has a positive impact on reducing the resource fragmentation degree of the running node, the resource utilization index value of the running node is improved to obtain an adjusted resource utilization index value of each running node;
[0015] Based on the adjusted resource utilization index value of each of the running nodes, at least one first target running node is selected from each of the running nodes to execute the reasoning task.
[0016] The scheduling method further includes:
[0017] In response to a scheduling request of a task of the target model obtained through training, obtaining a fault parameter of each operating node in at least one operating node;
[0018] Determining a fault tolerance performance index value of the running node based on the fault parameter of the running node; the fault tolerance performance index value is used to reflect the fault tolerance performance of the running node;
[0019] Based on the fault tolerance performance index value of each of the running nodes, at least one second target running node is selected from each of the running nodes to execute the task of obtaining the target model through training.
[0020] Obtaining a fault parameter of each running node in at least one running node includes:
[0021] The number of failures of each running node in at least one running node, the number of faulty nodes on the same route as the running node, and a second weight factor and a third weight factor are obtained; the second weight factor is used to reflect the importance of the impact of the running node on the running stability; the third weight factor is used to reflect the importance of the impact of the faulty nodes on the same route as the running node on the running stability.
[0022] The scheduling method further includes:
[0023] Obtaining the resource usage and non-preemptible resource usage of the tenant corresponding to each of the inference tasks;
[0024] Based on the resource usage and non-preemptible resource amount of the tenants corresponding to each of the inference tasks, at least one first tenant whose resource usage does not exceed the non-preemptible resource amount and at least one second tenant whose resource usage exceeds the non-preemptible resource amount are screened out from the tenants corresponding to each of the inference tasks;
[0025] Prioritizing the scheduling order of each first reasoning task corresponding to the at least one first tenant among the reasoning tasks;
[0026] After the scheduling order of each of the first reasoning tasks is determined, the scheduling order of each of the second reasoning tasks corresponding to the at least one second tenant in the reasoning tasks is determined.
[0027] The determining a scheduling order of each second reasoning task corresponding to the at least one second tenant in each reasoning task includes:
[0028] Obtaining a weight factor and a total amount of resources of each second tenant in the at least one second tenant and a resource application amount of a second reasoning task corresponding to the second tenant;
[0029] Determine the resource occupancy index value of each second tenant based on the weight factor, total resource amount and resource usage amount of each second tenant and the resource application amount of the second reasoning task corresponding to the second tenant; the resource occupancy index value is used to reflect the resource occupancy of the second tenant;
[0030] The scheduling order of each of the second inference tasks is based on the resource occupancy index value of each of the second tenants.
[0031] Before obtaining the fault parameter of each running node in at least one running node, the method further includes:
[0032] Acquire the parallel configuration parameters and resource application amount of the task of obtaining the target model through training; the parallel configuration parameters are used to reflect the demand of the task of obtaining the target model through training for parallel computing resources;
[0033] The parallel configuration parameters and resource application amount of the task of obtaining the target model through training are determined, and if the training requirements of the task of obtaining the target model through training are not met, the parallel configuration parameters and the resource application amount are adjusted.
[0034] The scheduling method further includes:
[0035] Determine the parallel configuration parameters and resource application amount of the task of obtaining the target model through training to meet the training requirements of the task of obtaining the target model through training, and if there is no at least one running node that meets the resource application amount, adjust the parallel configuration parameters and the resource application amount.
[0036] On the other hand, the present application provides a scheduling device, including:
[0037] A first obtaining module, configured to obtain, in response to a scheduling request of at least one reasoning task based on a target model, an application amount of resources for each reasoning task in the at least one reasoning task based on the target model;
[0038] A second obtaining module is used to obtain the total amount and available amount of resources of each running node in at least one running node;
[0039] A first determination module is used to determine the resource utilization index value and anti-fragmentation parameter of each running node based on the resource application amount of the reasoning task and the total amount and available amount of resources of each running node; the resource utilization index value is used to reflect the resource utilization of the running node; the anti-fragmentation parameter is used to reflect the impact of scheduling the reasoning task to the running node on reducing the degree of resource fragmentation of the running node;
[0040] The first selection module is used to select at least one first target running node from each of the running nodes to execute the reasoning task based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0042] Figure 1 A flowchart of a scheduling method provided in Example 1 of the present application;
[0043] Figure 2 A flowchart of a scheduling method provided in Example 2 of the present application;
[0044] Figure 3 A flowchart of a scheduling method provided in Example 3 of the present application;
[0045] Figure 4 A schematic diagram of an implementation scenario of a scheduling method provided in this application;
[0046] Figure 5 A schematic diagram of another implementation scenario of a scheduling method provided in this application;
[0047] Figure 6 A flowchart of a scheduling method provided in Example 4 of the present application;
[0048] Figure 7 A schematic diagram of a selection of a scheduling strategy provided for this application;
[0049] Figure 8 A schematic diagram of another option of a scheduling strategy provided for this application;
[0050] Fig. 9 A schematic diagram of the structure of a scheduling device provided in this application. DETAILED DESCRIPTION
[0051] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0052] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0053] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0054] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0055] Reference Figure 1 , is a flow chart of a scheduling method provided in Example 1 of the present application, such as Figure 1 As shown, the method may include but is not limited to the following steps:
[0056] Step S101 : in response to a scheduling request of at least one reasoning task based on a target model, obtaining a resource application amount of each reasoning task in the at least one reasoning task based on the target model.
[0057] In this embodiment, the target model may be constructed based on machine learning, deep learning or other artificial intelligence algorithms. For example, the target model may include a language model with natural language processing capabilities and deep semantic understanding capabilities.
[0058] The target model-based reasoning task can be used to reason about input data (such as images, texts, instructions, videos, etc.) based on the target model to obtain reasoning results.
[0059] The resource application amount of the inference task may include: the number of computing resources (such as at least one of CPU, memory and GPU) required to execute the inference task. For example, if the number of GPU cards required to execute inference task A is 2, then the GPU card application amount of inference task A is 2; if the number of GPU cards required to execute inference task B is 1, then the GPU card application amount of inference task B is 1.
[0060] Step S102: Obtain the total amount and available amount of resources of each running node in at least one running node.
[0061] In this embodiment, at least one running node whose idle resources satisfy the resource application amount of the reasoning task may be screened from the cluster.
[0062] The total amount of resources of a running node can be understood as the sum of the number of resources available on the running node. For example, if 6 GPU cards can be provided on the running node, the total amount of GPU cards of the running node is 6.
[0063] In the total amount of resources, some resources may have been occupied by other tasks, and the remaining part is the available amount. For example, if the total amount of resources includes: the total number of GPU cards is 6, of which 2 GPU cards have been occupied by other tasks, and the remaining 4 GPU cards can be used, then the available amount is 4.
[0064] The at least one running node may include a node running a Pod (i.e., a container or a group of containers) in a Kubernetes cluster, and the running node may be a physical server or a virtual machine.
[0065] Step S103: Determine the resource utilization index value and anti-fragmentation parameters of each of the running nodes based on the resource application amount of the inference task and the total amount and available amount of resources of each of the running nodes.
[0066] The resource utilization index value may be used to reflect the resource utilization status of the running node.
[0067] The anti-fragmentation parameter may be used to reflect the effect of scheduling the inference task to the execution node on reducing the resource fragmentation degree of the execution node.
[0068] In this embodiment, during the resource allocation process, since multiple tasks occupy resources in different sizes, the resources are divided into more and more small blocks that are difficult to effectively utilize, and the degree of resource fragmentation is higher.
[0069] Step S104: Based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes, select at least one first target running node from each of the running nodes to execute the reasoning task.
[0070] Based on the resource utilization index value and anti-fragmentation parameters of the running node, the balance of resource utilization and the degree of resource fragmentation of the running node can be considered simultaneously in the selection process, so as to select the running node that can not only ensure the balance of resource utilization but also help reduce resource fragmentation in long-term operation as the first target running node.
[0071] In this embodiment, by responding to the scheduling request of at least one reasoning task based on the target model, obtaining the resource application amount of each reasoning task in the at least one reasoning task based on the target model, obtaining the total amount and available amount of resources of each running node in at least one running node, and determining the resource utilization index value and anti-fragmentation parameter of each running node based on the resource application amount of the reasoning task and the total amount and available amount of resources of each running node, it can reflect the resource utilization of each running node and the impact of the scheduling of the reasoning task to the running node on reducing the resource fragmentation degree of the running node. On this basis, based on the resource utilization index value and anti-fragmentation parameter of each running node, at least one first target running node is selected from each running node to execute the reasoning task, which can reduce resource fragmentation caused by unreasonable scheduling, reduce the risk of reasoning task scheduling failure, ensure the execution efficiency of the reasoning task, and ensure the balance of internal resource utilization of the running node.
[0072] As another optional embodiment of the present application, refer to Figure 2 , is a flow chart of a scheduling method provided in Example 2 of the present application. This example is mainly an implementation method of step S103 in Example 1. Figure 2As shown, step S103 may include but is not limited to the following steps:
[0073] Step S1031: Determine the utilization rate of each type of resource of each running node based on the application amount of each type of resource for the inference task and the total amount and available amount of each type of resource of each running node.
[0074] In this embodiment, the used amount of each type of resource of the running node may be determined based on the total amount and the usable amount of each type of resource of the running node.
[0075] In this embodiment, the expected usage of the resource of the type of the running node may be determined based on the application amount of the inference task for the resource of the type and the used amount of the resource of the type of the running node.
[0076] After the expected usage is obtained, the proportion of the expected usage of the type of resource of the running node in the total amount of the type of resource of the running node can be used as the utilization rate of the type of resource of the running node.
[0077] The utilization rate may represent the degree of utilization of the type of resources when the inference task is scheduled to the execution node.
[0078] Step S1032: determine the resource utilization index value of each of the running nodes based on the utilization rate of each of the types of resources of each of the running nodes and a first weight factor; the first weight factor is used to reflect the importance of the type of resources to the reasoning task.
[0079] In this embodiment, the resource utilization index value of the running node can be determined by the following relationship:
[0080]
[0081] S represents the resource utilization index value; C i Indicates the total amount of resources of type i; U i W represents the expected usage of resource type i; i Indicates the first weight factor of the i-th category resource.
[0082] The more important the resource of this type is to the reasoning task, the higher the first weight factor corresponding to the resource of this type is.
[0083] Step S1033: based on the resource application amount of the inference task and the available amount of resources of each of the running nodes, determine the expected remaining amount of resources of each of the running nodes, and use the expected remaining amount as a de-fragmentation parameter.
[0084] In this embodiment, the difference between the available amount of resources of the running node and the amount of resources requested by the inference task can be used as the expected remaining amount of resources of the running node. The expected remaining amount can reflect the expected remaining situation of the resources of the running node when the inference task is scheduled to the running node.
[0085] In this embodiment, the reasoning task may require multiple types of resources, such as CPU, memory, GPU, etc. Each type of resource may have different effects on the reasoning task. Therefore, based on the application amount of each type of resource for the reasoning task and the total amount and available amount of each type of resource of each running node, the utilization rate of each type of resource of each running node is determined. Based on the utilization rate of each type of resource of each running node and the first weight factor, the resource utilization index value of each running node is determined. The resource utilization index value can more comprehensively and accurately reflect the resource utilization of the running node when the reasoning task is scheduled to the running node, so that the selected target running node is more reasonable.
[0086] Furthermore, using the expected remaining amount as an anti-fragmentation parameter can more intuitively reflect the degree of fragmentation of node resources when the reasoning task is scheduled to the running node, which helps to more quickly select at least one target running node from at least one running node, thereby improving the execution efficiency of the reasoning task.
[0087] As another optional embodiment of the present application, refer to Figure 3 , is a flowchart of a scheduling method provided in Example 3 of the present application. This embodiment is mainly an implementation method of the above step S104. Figure 3 As shown, step S104 may include but is not limited to the following steps:
[0088] Step S1041: If the anti-fragmentation parameter of the running node is used to reflect that scheduling the inference task to the running node has a positive impact on reducing the degree of resource fragmentation of the running node, the resource utilization index value of the running node is improved to obtain the adjusted resource utilization index value of each running node.
[0089] For example, the amount of GPU resources requested by an inference task is generally an exponential power of 2 (1, 2, 4, 8, etc.). If the expected remaining amount of the GPU is 0 when the inference task is scheduled to the running node, it means that all GPU resources of the running node have been fully utilized by the current task, leaving no unallocated resources that may cause fragmentation.
[0090] If the expected remaining amount of GPU when the inference task is scheduled to the running node is a certain exponential power of 2, it indicates that the remaining GPU resource block size still meets the standard of allocation in units of exponential powers of 2, which enables the remaining GPU resources to be seamlessly connected to subsequent inference tasks that also apply for resources in units of exponential powers of 2, thereby avoiding fragmentation problems caused by mismatch of resource block sizes.
[0091] Therefore, if the expected remaining amount of the GPU when the inference task is scheduled to the running node is 0 or an exponential power of 2, it can be explained that scheduling the inference task to the running node has a positive impact on reducing the degree of resource fragmentation of the running node, and therefore the resource utilization index value of the running node can be improved.
[0092] In the present application, the specific implementation method of improving the resource utilization index value of the running node is not limited. For example, the resource utilization index value of the running node can be multiplied by a coefficient greater than 1, or an incremental value can be added to the resource utilization index value of the running node.
[0093] Step S1042: Based on the adjusted resource utilization index value of each of the running nodes, select at least one first target running node from each of the running nodes to execute the reasoning task.
[0094] In this embodiment, by comparing the adjusted resource utilization index values of the running nodes, one or more running nodes with relatively higher adjusted resource utilization index values can be selected from the running nodes as the first target running nodes to execute the reasoning task.
[0095] For example, the adjusted resource utilization index values of the running nodes may be sorted in descending order to obtain a sorting result.
[0096] If the inference task is not split into multiple subtasks, the running node ranked first may be selected as the first target running node to execute the entire inference task.
[0097] If the reasoning task is split into multiple subtasks, and the multiple subtasks can be executed in parallel to improve the overall reasoning speed, multiple running nodes with high rankings can be selected as multiple first target running nodes, and the multiple first target running nodes can execute multiple subtasks in parallel.
[0098] In this embodiment, if the anti-fragmentation parameter of the running node is used to reflect that scheduling the inference task to the running node has a positive impact on reducing the degree of resource fragmentation of the running node, the resource utilization index value of the running node is improved, which helps to guide the inference task to be scheduled to the running node that can reduce fragmentation. On this basis, scheduling decisions are made based on the adjusted resource utilization index value, which can ensure that in the scheduling process, priority is given to scheduling the inference task to the running node that can reduce fragmentation, thereby reducing resource fragmentation caused by unreasonable scheduling, reducing the risk of reasoning task scheduling failure, ensuring the execution efficiency of the reasoning task, and ensuring the balance of internal resource utilization of the running node.
[0099] For example, the cluster includes two running nodes, namely node1 and node2. Both node1 and node2 have 6 GPU cards. Node1 has used 3 GPU cards to execute tasks, and node2 has used 4 GPU cards to execute tasks.
[0100] like Figure 4 As shown, if the application amount of GPU resources for the inference task to be executed is 1, the expected remaining amount corresponding to node1 (an implementation method of the anti-fragmentation parameter) is 2, and the expected remaining amount corresponding to node2 is 1, the expected remaining amount of node1 can reflect that the scheduling of the inference task to node1 has a positive impact on reducing the resource fragmentation degree of node1, and the resource utilization index value of node1 is improved. The expected remaining amount of node2 can reflect that the scheduling of the inference task to node2 has no positive impact on reducing the resource fragmentation degree of node2, and the resource utilization index value of node2 can be kept unchanged. After the resource utilization index value of node1 is improved, the resource utilization index value of node1 after the improvement is higher than the resource utilization index value of node2, and node1 can be selected to execute the inference task. Accordingly, node1 has 2 GPU cards left to use, and node2 still has 2 GPU cards left to use.
[0101] If there is a new inference task to be executed with an application amount of 2 GPU resources, and the expected remaining amount of node1 and node2 is 0, both can have a positive impact, and the resource utilization index values of node1 and node2 can be improved. By comparison, the resource utilization index value of node1 after the improvement is higher than the resource utilization index value of node2 after the improvement, and the inference task can be scheduled to node1 for execution. Accordingly, node1 has 0 GPU cards left to use, and node2 still has 2 GPU cards left to use.
[0102] If there is a new inference task to be executed and the application amount for GPU resources is 2, the inference task can be scheduled to be executed on node2.
[0103] Compared with the scheduling scheme that does not consider the expected remaining amount to determine the resource utilization index value, it can reduce fragmentation and reduce the risk of scheduling failure. Figure 5 In the scheduling scheme shown, if the GPU resource application amount of the inference task to be executed is 1, the inference task is scheduled to node2. Accordingly, node1 has 3 GPU cards left and node2 has 1 GPU card left. If there is a new inference task to be executed with an application amount of 2 GPU resources, the inference task is scheduled to node1. Accordingly, node1 has 1 GPU card left and node2 has 1 GPU card left. If there is a new inference task to be executed with an application amount of 2 GPU resources, node1 and node2 have only 1 GPU card left. Both node1 and node2 cannot meet the demand for GPU resources of the inference task, resulting in scheduling failure. Figure 5 The scheduling scheme shown does not consider the expected remaining amount to determine the resource utilization index value. When there are continuous inference tasks applying for GPU resources, some inference tasks fail to be scheduled due to resource fragmentation.
[0104] As another optional embodiment of the present application, refer to Figure 6 , is a flow chart of a scheduling method provided in Example 4 of the present application, such as Figure 6 As shown, the method may include but is not limited to the following steps:
[0105] Step S201: In response to a scheduling request of a task of obtaining a target model through training, a fault parameter of each operating node in at least one operating node is obtained.
[0106] In this embodiment, the node detection service can be called through the health check component to obtain the status of each running node in the cluster, mark the node with a faulty status as faulty, and record the fault parameters of the node (such as fault type, occurrence time, number of faults, etc.) in the parameter cache for subsequent use.
[0107] In the subsequent node status check process, if the node previously marked as faulty has returned to normal, its status can be updated to normal.
[0108] In response to the scheduling request of the task of obtaining the target model through training, at least one running node marked as normal and with idle resources that meet the resource application amount of the task of obtaining the target model through training can be screened out from the cluster, and the fault parameters of each running node in the at least one running node are obtained from the parameter cache. If the running node has not failed before, its fault parameters are empty; if the running node has failed before, its fault parameters contain information related to the fault, such as the number of failures, the type of failure, etc.
[0109] Step S202: determine a fault tolerance performance index value of the running node based on the fault parameters of the running node; the fault tolerance performance index value is used to reflect the fault tolerance performance of the running node.
[0110] In this embodiment, the fault tolerance performance of the operating node can be understood as: the stability and recovery capability of the operating node when facing a fault. The higher the fault tolerance performance index value, the better the fault tolerance performance of the corresponding operating node.
[0111] Step S203: Based on the fault tolerance performance index value of each of the running nodes, select at least one second target running node from each of the running nodes to execute the task of obtaining the target model through training.
[0112] In this embodiment, by comparing the fault tolerance performance index values of each of the running nodes, one or more running nodes with relatively high fault tolerance performance index values can be selected from the running nodes as the second target running nodes to perform the task of obtaining the target model through training.
[0113] For example, the fault tolerance performance index values of the running nodes may be sorted in descending order to obtain a sorting result.
[0114] If the task of obtaining the target model through training is not split into multiple subtasks, the first-ranked running node can be selected as the second target running node to execute the entire task of obtaining the target model through training.
[0115] If the task of obtaining the target model through training is divided into multiple subtasks, and the multiple subtasks can be executed in parallel to improve the overall training speed, multiple running nodes with high rankings can be selected as multiple second target running nodes, and the multiple second target running nodes can execute multiple subtasks in parallel.
[0116] In this embodiment, by responding to the scheduling request of the task of obtaining the target model through training, obtaining the fault parameters of each running node in at least one running node, and determining the fault tolerance performance index value of the running node based on the fault parameters of the running node, the fault tolerance performance of each running node can be evaluated more objectively. On this basis, selecting the second target running node based on the fault tolerance performance index value can ensure that the training task is scheduled to be executed on the running node with higher fault tolerance performance, which can improve the execution success rate of the training task and better meet the high demand for resources and the stability requirements of the operation of the training task.
[0117] Step S204: In response to a scheduling request of at least one target model-based reasoning task, obtain a resource application amount of each reasoning task in the at least one target model-based reasoning task.
[0118] Step S205: Obtain the total amount and available amount of resources of each running node in at least one running node.
[0119] Step S206: Determine the resource utilization index value and anti-fragmentation parameters of each running node based on the resource application amount of the inference task and the total amount and available amount of resources of each running node; the resource utilization index value is used to reflect the resource utilization of the running node; the anti-fragmentation parameter is used to reflect the impact of scheduling the inference task to the running node on reducing the degree of resource fragmentation of the running node.
[0120] Step S207: Based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes, select at least one first target running node from each of the running nodes to execute the reasoning task.
[0121] The detailed process of steps S204-S207 can refer to the relevant introduction of steps S101-S104 in Example 1, which will not be repeated here.
[0122] In this embodiment, in order to efficiently and reasonably allocate running nodes to new tasks, two scheduling strategies can be set: an anti-fragmentation scheduling strategy and a fault-tolerant scheduling strategy. In view of the significant differences in resource requirements of different types of tasks (e.g., training tasks usually require higher resource occupancy (such as multi-GPU parallelism); while reasoning tasks have smaller resource requirements), the amount of resource application for new tasks can be quantified for easy classification. Specifically, the amount of resource application can be normalized to obtain a normalized value. If the normalized value is greater than a set threshold, it can be classified as a training task; if the normalized value is less than a set threshold, it can be classified as an reasoning task.
[0123] For example, Figure 7As shown, when it is determined that the new task is a reasoning task based on the target model, an anti-fragmentation scheduling strategy can be selected from the anti-fragmentation scheduling strategy and the fault-tolerant scheduling strategy, and at least one first target running node is selected from at least one running node based on the anti-fragmentation scheduling strategy to perform the reasoning task (i.e., the process of steps S204-S207).
[0124] like Figure 8 As shown, when it is determined that the new task is a task that obtains the target model through training, a fault-tolerant scheduling strategy can be selected from the anti-fragmentation scheduling strategy and the fault-tolerant scheduling strategy, and at least one second target running node is selected from at least one running node based on the fault-tolerant scheduling strategy to execute the task that obtains the target model through training (i.e., the process of steps S201-S203).
[0125] As another optional embodiment of the present application, a scheduling method is provided in Embodiment 5 of the present application. This embodiment is mainly an implementation method of obtaining the fault parameters of each running node in the at least one running node. Obtaining the fault parameters of each running node in the at least one running node may include but is not limited to the following steps:
[0126] Step S2011: Obtain the number of failures of each running node in at least one running node, the number of failed nodes on the same route as the running node, and the second weight factor and the third weight factor.
[0127] The second weight factor may be used to reflect the importance of the impact of the operating node on the operating stability.
[0128] The third weight factor may be used to reflect the importance of the impact of the faulty node on the same route as the running node on the running stability.
[0129] Corresponding to step S2011, the above step S202 may include but is not limited to:
[0130] Based on the following relationship, the fault tolerance performance index value of the running node is determined:
[0131] S n =S int -W f *N f -W r *N r
[0132] S n represents the fault tolerance performance index value of the running node, S int Represents the fault tolerance performance initialization index value, W f represents the second weight factor, W r represents the third weight factor, N f Indicates the number of failures of the running node, Nr Indicates the number of failed nodes on the same route as the running node.
[0133] In this embodiment, when a running node frequently fails, this directly points to its lower stability and relatively weak fault tolerance. Therefore, it is both reasonable and necessary to consider the fault tolerance performance of the running node using the number of failures as the core fault parameter. Furthermore, the running node does not exist in isolation, and its operating environment, especially the mutual connection with other nodes under the same route, may also have an important impact on its fault tolerance performance. For example, problems such as network congestion or routing failures caused by other nodes under the same route may indirectly affect the operating stability and fault tolerance of the running node. Therefore, taking the number of faulty nodes under the same route as another important parameter helps to understand the stability and reliability of the node more deeply from the routing level, thereby providing a more complete description of the node's fault tolerance performance.
[0134] In order to more accurately quantify the impact of the number of failures of the running node and the number of faulty nodes on the same route as the running node, the second weight factor and the third weight factor can ensure the flexibility and pertinence of determining the fault-tolerant performance index value. It can be seen that by comprehensively considering the number of failures of the running node, the number of faulty nodes on the same route, and the two weight factors, the fault-tolerant performance index value obtained can more comprehensively and accurately reflect the fault-tolerant performance of the running node.
[0135] As another optional embodiment of the present application, a scheduling method is provided in Embodiment 6 of the present application, and the method may include but is not limited to the following steps:
[0136] Step S301: In response to a scheduling request of at least one target model-based reasoning task, obtain the resource usage and non-preemptible resource usage of the tenant corresponding to each reasoning task.
[0137] In a cluster environment, tenants can be used as core units for resource usage. Each tenant has a dedicated resource pool, which can be divided by the cluster according to the specific needs and configurations of the tenant. All users under the tenant can share the resource pool.
[0138] In terms of resource management, the cluster can monitor resources based on the entire tenant level, rather than just focusing on a single user. Therefore, both the amount of resources used and the amount of resources that cannot be preempted can be counted and managed for the entire tenant.
[0139] The resource usage can reflect the total amount of resources currently used by the tenant, which can be obtained by summarizing the resource usage of all users under it.
[0140] The non-preemptible amount of resources can be a resource protection threshold set by tenants to protect their key businesses. This threshold can ensure that tenants can maintain their basic operating requirements even when resources are tight.
[0141] Each user can submit reasoning tasks to the cluster according to actual needs and request the allocation of corresponding running nodes.
[0142] Step S302: Based on the resource usage and non-preemptible resource amount of the tenants corresponding to each of the reasoning tasks, select from the tenants corresponding to each of the reasoning tasks at least one first tenant whose resource usage does not exceed the non-preemptible resource amount, and at least one second tenant whose resource usage exceeds the non-preemptible resource amount.
[0143] In this embodiment, the resource usage and non-preemptible resource amount of the tenant corresponding to each of the reasoning tasks can be compared. If the tenant's resource usage does not exceed the non-preemptible resource amount, the tenant can be classified as a first tenant; if the tenant's resource usage exceeds the non-preemptible resource amount, the tenant can be classified as a second tenant.
[0144] Step S303: preferentially determine a scheduling order of each first reasoning task corresponding to the at least one first tenant among the reasoning tasks.
[0145] When determining the scheduling order of each first reasoning task corresponding to the at least one first tenant in the reasoning tasks,
[0146] The first tenants may be sorted according to their priorities, and the scheduling order of the first reasoning tasks corresponding to the first tenants may be determined in turn according to the sorting results of the first tenants.
[0147] For each first reasoning task in the same first tenant, each first reasoning task may be sorted according to the chronological order of creation time of each first reasoning task to obtain a scheduling order of each first reasoning task.
[0148] After the scheduling order of each first reasoning task is preferentially determined, at least one first target running node may be selected for each first reasoning task in sequence from at least one running node according to the scheduling order of each first reasoning task.
[0149] Step S304: After the scheduling order of each of the first reasoning tasks is determined, determine the scheduling order of each of the second reasoning tasks corresponding to the at least one second tenant in the reasoning tasks.
[0150] After determining the scheduling order of each second reasoning task corresponding to the at least one second tenant in each reasoning task, at least one first target running node can be selected for each second reasoning task from at least one running node in turn according to the scheduling order of each second reasoning task.
[0151] Step S305 , preferentially obtaining the resource application amount of each first reasoning task in sequence according to the scheduling order of each first reasoning task, and obtaining the resource application amount of each second reasoning task in sequence according to the scheduling order of each second reasoning task.
[0152] Step S305 is an implementation of the above step S101.
[0153] Step S306, obtaining the total amount and available amount of resources of each running node in at least one running node.
[0154] The detailed process of step S306 can refer to the relevant introduction of the above step S102, which will not be repeated here.
[0155] Step S307: Based on the amount of resources requested by the first reasoning task or the second reasoning task and the total amount and available amount of resources of each of the running nodes, determine the resource utilization index value and anti-fragmentation parameters of each of the running nodes; the resource utilization index value is used to reflect the resource utilization of the running node; the anti-fragmentation parameter is used to reflect the impact of scheduling the reasoning task to the running node on reducing the degree of resource fragmentation of the running node.
[0156] Step S307 is an implementation of the above step S103.
[0157] Step S308: Based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes, select at least one first target running node from each of the running nodes to execute the first reasoning task or the second reasoning task.
[0158] Step S308 is an implementation of the above step S104.
[0159] In this embodiment, by obtaining the resource usage and non-preemptible resource amount of the tenant corresponding to each of the inference tasks, based on the resource usage and non-preemptible resource amount of the tenant corresponding to each of the inference tasks, at least one first tenant whose resource usage does not exceed the non-preemptible resource amount and at least one second tenant whose resource usage exceeds the non-preemptible resource amount are screened out from the tenants corresponding to each of the inference tasks, and the scheduling order of each first inference task corresponding to the at least one first tenant in each of the inference tasks is preferentially determined, and the first inference task can be preferentially scheduled according to the scheduling order of each first inference task, thereby reducing the risk of resource preemption and ensuring the stability of the business operation of the tenants in the cluster. At the same time, for the second tenant, although its resource usage has exceeded the non-preemptible amount, resources will still be allocated to it according to the scheduling order if resources allow, thereby ensuring the stability of the business operation of all tenants to a certain extent.
[0160] Moreover, based on the resource utilization index value and anti-fragmentation parameters of each of the running nodes, at least one first target running node is selected from each of the running nodes to execute the first reasoning task or the second reasoning task, which can reduce resource fragmentation caused by unreasonable scheduling, reduce the risk of scheduling failure of the first reasoning task or the second reasoning task, ensure the execution efficiency of the first reasoning task or the second reasoning task, and ensure the balance of internal resource utilization of the running nodes.
[0161] As another optional embodiment of the present application, a scheduling method is provided for Embodiment 7 of the present application. This embodiment is mainly an implementation method for determining the scheduling order of each second reasoning task corresponding to the at least one second tenant in each of the reasoning tasks. Determining the scheduling order of each second reasoning task corresponding to the at least one second tenant in each of the reasoning tasks may include but is not limited to the following steps:
[0162] Step S3041: Obtain a weight factor and a total amount of resources of each second tenant in the at least one second tenant and a resource application amount of a second reasoning task corresponding to the second tenant.
[0163] The weight factor of the second tenant can represent the relative importance of the second tenant in cluster resource allocation. The weight factor can be adjusted based on business requirements, resource utilization, and other factors.
[0164] The total amount of resources that the second tenant deserves can be obtained according to its weight factor. Specifically, the total amount of resources of the second tenant can be obtained according to the following relationship:
[0165] The total resources of the second tenant = (weight / total weight)*total resource
[0166] Among them, weight represents the weight factor of the second tenant, total weight represents the sum of the weight factors of all tenants, and total resource represents the total amount of resources in the cluster.
[0167] The total amount of resources of the second tenant can be understood as: the total amount of a certain type of resources of the second tenant, for example, the total amount of GPU resources of the second tenant.
[0168] The resource application amount of the second reasoning task corresponding to the second tenant can be understood as: the application amount of a certain type of resource of the second reasoning task corresponding to the second tenant, for example, the application amount of GPU resources.
[0169] The total amount of resources in a cluster can be understood as the total amount of a certain type of resource in the cluster, for example, the total amount of GPU resources in the cluster.
[0170] Step S3042: Determine the resource occupancy index value of each second tenant based on the weight factor, total resource amount, used resource amount, and resource application amount of the second inference task corresponding to the second tenant.
[0171] The resource occupancy index value may be used to reflect the resource occupancy status of the second tenant.
[0172] In this embodiment, the resource occupancy index value of the second tenant may be determined according to the following relationship:
[0173]
[0174] P i represents the resource occupancy index value of the i-th second tenant, w i Indicates the weight factor of the i-th second tenant, used g Indicates the resource usage of the i-th second tenant, total g Indicates the total amount of resources of the second tenant of the ith tenant, req g Indicates the resource application amount of the second reasoning task corresponding to the i-th second tenant.
[0175] Step S3043: Scheduling order of each second inference task based on the resource occupancy index value of each second tenant.
[0176] In this embodiment, the resource occupancy index values of each second tenant can be sorted from low to high, and the sorting result can be used as the scheduling order of each second reasoning task. By determining the scheduling order of each second reasoning task in this way, the second reasoning task corresponding to the second tenant with less resource occupancy can obtain resources first, so as to ensure more balanced and fair resource utilization among different tenants.
[0177] In this embodiment, the execution of steps S3041-S3043 can be based on a precondition that the resource usage of the second tenant does not exceed the resource upper limit value set by the second tenant. The resource upper limit value can represent the upper limit of the sum of the resource usage within the second tenant. This is a hard constraint used to limit the resource consumption of the tenant and ensure that the tenant does not exceed the range of its available resources.
[0178] In this embodiment, by obtaining the weight factor and the total amount of resources of each second tenant in the at least one second tenant and the resource application amount of the second reasoning task corresponding to the second tenant, the resource occupancy index value of each second tenant is determined based on the weight factor, the total amount of resources and the used amount of resources of each second tenant and the resource application amount of the second reasoning task corresponding to the second tenant. Based on the resource occupancy index value of each second tenant, the scheduling order of each second reasoning task can ensure that the second reasoning task corresponding to the second tenant with less resource occupancy can obtain resources first, avoid excessive concentration and waste of resources, and achieve balanced distribution of resources among different tenants.
[0179] As another optional embodiment of the present application, a scheduling method is provided in Embodiment 8 of the present application, and the method may include but is not limited to the following steps:
[0180] Step S401, in response to a scheduling request of a task that obtains the target model through training, obtain parallel configuration parameters and resource request amounts of the task that obtains the target model through training; the parallel configuration parameters are used to reflect the demand of the task that obtains the target model through training for parallel computing resources.
[0181] In this embodiment, the parallel configuration parameters may include but are not limited to: at least one of data parallelism, tensor parallelism, and pipeline parallelism.
[0182] The resource application amount of the task of obtaining the target model through training can be expressed as: nnodes×proc_per_nodes. nnodes can represent the number of running nodes applied for, and proc_per_nodes can represent the number of resources of a specified specification required on each running node (e.g., the number of GPU cards with a memory size specification of capacityA).
[0183] Step S402: determine the parallel configuration parameters and resource application amount of the task of obtaining the target model through training, and if the training requirements of the task of obtaining the target model through training are not met, adjust the parallel configuration parameters and the resource application amount.
[0184] In this embodiment, the resource capacity required for the task of obtaining the target model through training can be predicted.
[0185] In this embodiment, if the task of obtaining the target model through training does not satisfy the following relationship, it can be determined that the parallel configuration parameters and resource application amount of the task of obtaining the target model through training do not meet the training requirements of the task of obtaining the target model through training:
[0186] memory prdict <tp degree ×pp degree ×capacity A ×(1+θ)
[0187] memory prdict It can represent the resource capacity (e.g., the memory size of the GPU card) required for the task of training the target model. degree Can represent tensor parallelism, pp degree represents the parallelism of the pipeline, capacityA represents the specified specification (e.g., the size of the video memory), and θ represents the threshold value, which is used to ensure that the actual resources can meet the requirements.
[0188] Adjusting the parallel configuration parameters and the resource application amount may include but is not limited to:
[0189] Step S4021, determine whether there are enough idle resources in the cluster to be n times the resource application amount;
[0190] If yes, execute step S4022; if no, execute step S4023.
[0191] Step S4022: multiply the resource application amount by n to obtain the adjusted resource application amount.
[0192] Step S4023, adjust the size of pipeline parallelism and tensor parallelism, the adjusted size of pipeline parallelism is equal to the number of applied running nodes (which can be expressed as nnodes), and the adjusted size of tensor parallelism is equal to the number of resources of specified specifications required on the running node (which can be expressed as proc_per_nodes).
[0193] Step S403: Obtain fault parameters of each running node in at least one running node.
[0194] In this embodiment, the at least one running node may include a running node whose idle resources in the cluster satisfy the resource application amount of the task.
[0195] Step S404: determine a fault tolerance performance index value of the running node based on the fault parameters of the running node; the fault tolerance performance index value is used to reflect the fault tolerance performance of the running node.
[0196] Step S405: Based on the fault tolerance performance index value of each of the running nodes, select at least one second target running node from each of the running nodes to execute the task of obtaining the target model through training.
[0197] Step S406: In response to a scheduling request of at least one target model-based reasoning task, obtain a resource application amount of each reasoning task in the at least one target model-based reasoning task.
[0198] Step S407: Obtain the total amount and available amount of resources of each running node in at least one running node.
[0199] Step S408: Based on the amount of resources requested by the inference task and the total amount and available amount of resources of each of the running nodes, determine the resource utilization index value and anti-fragmentation parameters of each of the running nodes; the resource utilization index value is used to reflect the resource utilization of the running node; the anti-fragmentation parameter is used to reflect the impact of scheduling the inference task to the running node on reducing the degree of resource fragmentation of the running node.
[0200] Step S409: Based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes, select at least one first target running node from each of the running nodes to execute the reasoning task.
[0201] The detailed process of steps S403-S409 can refer to the relevant introduction of steps S201-S207 in Example 4, which will not be repeated here.
[0202] In this embodiment, by obtaining the parallel configuration parameters and resource application amounts of the task of obtaining the target model through training, it is possible to accurately understand the task's demand for parallel computing resources, and when the parallel configuration parameters and resource application amounts of the task of obtaining the target model through training are determined, if the training requirements of the task of obtaining the target model through training are not met, the parallel configuration parameters and the resource application amounts are adjusted to ensure that resources can be efficiently and reasonably allocated according to task requirements, so that the task of obtaining the target model through training can be efficiently and reliably executed.
[0203] In this embodiment, the scheduling method is further described by taking the application of GPU resources as an example. For example, the scheduling method may include the following steps:
[0204] Step S1. The user can submit a task related to the target model in the cluster. The task can apply for GPU resources (the job application for GPU resources is req). The tenant to which the user belongs is tenant1. The tenant configurations are reserved (i.e., the amount of resources that cannot be preempted), limit (the upper limit of resources), and weight (i.e., the weight factor of the tenant), where reserved≤limit; the tenant's resource usage can be expressed as used.
[0205] Step S2: determine whether there is historical data of the task (such as resource application amount, parallel configuration parameters, etc.). If so, execute step S5; if not, execute step S3.
[0206] Step S3: Obtain user application information and current idle resources of the cluster.
[0207] User application information may include: GPU type: typeA; number of GPU cards: N; memory size specification of typeA is capacityA;
[0208] The idle resources of typeA resources included in each running node in the cluster can be seen in Table 1. As shown in Table 1, the number of idle GPU cards included in node1 is x1, the number of idle GPU cards included in node2 is x2, ...
[0209] Table 1
[0210] node1 node2 node3 node4 …… nodeM x1 x2 x3 x4 …… xm
[0211] Step S4: Obtain the parallel configuration parameters of the task (e.g., data parallelism: dp degree , tensor parallelism: tp degree , pipeline parallelism: pp degree ), and configure the parallel configuration parameters and resource application amount of the task (which can be expressed as apply device ==nnodes*proc_per_nodes, nnodes means N running nodes, proc_per_nodes means each running node needs to have M GPU cards, apply device Usually a power of 2) for verification and adjustment:
[0212] The memory memory_predict that the task needs to occupy can be estimated based on the number of parameters of the target model;
[0213] If the task does not satisfy the following relationship, the parallel configuration parameters and resource application amount of the task can be determined to not meet the requirements of the task:
[0214] memory prdict<tp degree *pp degree *capacityA*(1+θ)
[0215] Determine whether there are enough idle GPU resources in the cluster, which is twice the current resource request amount:
[0216] If yes, multiply the resource application amount by 2 to get the adjusted resource application amount;
[0217] If not, adjust the existing parallel strategy: pp degree The maximum number is adjusted to nnodes, tp degree The maximum number is adjusted to proc_per_nodes;
[0218] Step S5, (a) Determine the scheduling order of each task according to the following strategy:
[0219] Tasks whose tenant resource usage does not exceed the non-preemptible resource amount (Reserved) are scheduled first;
[0220] Each task in the first tenant whose resource usage does not exceed the non-preemptible resource amount (Reserved) is scheduled according to the creation time;
[0221] For each task corresponding to the second tenant whose resource usage exceeds the non-preemptible resource amount (Reserved), the tasks corresponding to the second tenant with a lower resource usage ratio are scheduled first according to the resource usage index value of each second tenant.
[0222] (b) Scheduling strategy selection:
[0223] Normalize the user's resource request amount to obtain Resource-normalization, which is between 0 and 1. The higher the value, the more GPU resources are required and the higher the probability of the training task. Set the threshold, for example, 0.5.
[0224] Resource-normalization is greater than 0.5, which is considered to be a task to obtain the target model through training. This training task prefers a stable training environment, and a fault-tolerant scheduling strategy is selected for the training task.
[0225] Resource-normalization is less than or equal to 0.5, which means it is considered to be a reasoning task based on the target model. This reasoning task occupies relatively few resources, so the GPU anti-fragmentation scheduling strategy is selected for the reasoning task.
[0226] Step S6:
[0227] (a) Select at least one running node from the cluster whose idle resources satisfy the resource application amount of the reasoning task.
[0228] (b) Execute the anti-fragmentation scheduling strategy, according to the relation Determine a resource utilization index value of the running node, and select at least one first target running node for the reasoning task from at least one running node according to the resource utilization index value and the anti-fragmentation parameter of each running node.
[0229] Or, implement a fault-tolerant scheduling strategy, according to the relation S n =S int -W f *N f -W r *N r , determine the fault tolerance performance index value of the running node, and based on the fault tolerance performance index value of each running node, select at least one second target running node from at least one running node for the task of obtaining the target model through training.
[0230] As another optional embodiment of the present application, a scheduling method is provided in Embodiment 9 of the present application, and the method may include but is not limited to the following steps:
[0231] Step S501, in response to a scheduling request of a task that obtains the target model through training, obtain parallel configuration parameters and resource request amounts of the task that obtains the target model through training; the parallel configuration parameters are used to reflect the demand of the task that obtains the target model through training for parallel computing resources.
[0232] Step S502: determine the parallel configuration parameters and resource application amount of the task of obtaining the target model through training, and if the training requirements of the task of obtaining the target model through training are not met, adjust the parallel configuration parameters and the resource application amount.
[0233] Step S503: Determine whether there is at least one running node in the cluster with idle resources that meet the resource application amount.
[0234] If it does not exist, execute step S504; if it does exist, execute step S505.
[0235] Step S504: Adjust the parallel configuration parameters and the resource application amount.
[0236] In this embodiment, after adjusting the parallel configuration parameters and the resource application amount, step S503 can be continued to be executed, aiming to filter out at least one running node whose idle resources meet the resource application amount from the cluster after adjusting the parallel configuration parameters and the resource application amount.
[0237] For example, try adjusting the resource request amount to That is, the number of nodes requested by the user is doubled, and the number of resources requested on each node is halved. degree For pp degree ×2,tp degree tp degree / 2.
[0238] Step S505: Obtain fault parameters of each running node in at least one running node.
[0239] Step S506: Determine a fault tolerance performance index value of the running node based on the fault parameters of the running node; the fault tolerance performance index value is used to reflect the fault tolerance performance of the running node.
[0240] Step S507: Based on the fault tolerance performance index value of each of the running nodes, select at least one second target running node from each of the running nodes to execute the task of obtaining the target model through training.
[0241] Step S508: In response to a scheduling request of at least one target model-based reasoning task, obtain a resource application amount of each reasoning task in the at least one target model-based reasoning task.
[0242] Step S509: Obtain the total amount and available amount of resources of each running node in at least one running node.
[0243] Step S510: Based on the amount of resources requested by the inference task and the total amount and available amount of resources of each of the running nodes, determine the resource utilization index value and anti-fragmentation parameters of each of the running nodes; the resource utilization index value is used to reflect the resource utilization of the running node; the anti-fragmentation parameter is used to reflect the impact of scheduling the inference task to the running node on reducing the degree of resource fragmentation of the running node.
[0244] Step S511: Based on the resource utilization index value and anti-fragmentation parameters of each of the running nodes, select at least one first target running node from each of the running nodes to execute the reasoning task.
[0245] The detailed process of steps S403-S409 can refer to the relevant introduction of steps S201-S207 in Example 4, which will not be repeated here.
[0246] In this embodiment, when there are no idle resources in the cluster to meet the resource application amount, the parallel configuration parameters and resource application amount can be adjusted to adapt to the resource status of the current running nodes in the cluster, ensuring that at least one running node with idle resources that meet the resource application amount can be screened out from the cluster, and then the target running node is selected from at least one running node to execute the corresponding task, thereby ensuring stable execution of the task.
[0247] Next, the scheduling device provided in the present application is introduced. The scheduling device introduced below and the scheduling method introduced above can refer to each other.
[0248] Reference Fig. 9 The scheduling device includes: a first obtaining module 100, a second obtaining module 200, a first determining module 300 and a first selecting module 400.
[0249] A first obtaining module, configured to obtain, in response to a scheduling request of at least one reasoning task based on a target model, an application amount of resources for each reasoning task in the at least one reasoning task based on the target model;
[0250] A second obtaining module is used to obtain the total amount and available amount of resources of each running node in at least one running node;
[0251] A first determination module is used to determine the resource utilization index value and anti-fragmentation parameter of each running node based on the resource application amount of the reasoning task and the total amount and available amount of resources of each running node; the resource utilization index value is used to reflect the resource utilization of the running node; the anti-fragmentation parameter is used to reflect the impact of scheduling the reasoning task to the running node on reducing the degree of resource fragmentation of the running node;
[0252] The first selection module is used to select at least one first target running node from each of the running nodes to execute the reasoning task based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes.
[0253] The first determination module 300 may be specifically used for:
[0254] Determine the utilization rate of each type of resource of each running node based on the application amount of each type of resource for the reasoning task and the total amount and available amount of each type of resource of each running node; the utilization rate represents the degree of utilization of the type of resource when the reasoning task is scheduled to the running node;
[0255] Determining a resource utilization index value of each of the operating nodes based on the utilization rate of each of the types of resources of each of the operating nodes and a first weight factor; the first weight factor is used to reflect the importance of the type of resources to the reasoning task;
[0256] Based on the resource application amount of the inference task and the available amount of resources of each of the running nodes, the expected remaining amount of resources of each of the running nodes is determined, and the expected remaining amount is used as an anti-fragmentation parameter.
[0257] The first selection module 400 may be specifically used for:
[0258] If the anti-fragmentation parameter of the running node is used to reflect that scheduling the inference task to the running node has a positive impact on reducing the resource fragmentation degree of the running node, the resource utilization index value of the running node is improved to obtain an adjusted resource utilization index value of each running node;
[0259] Based on the adjusted resource utilization index value of each of the running nodes, at least one first target running node is selected from each of the running nodes to execute the reasoning task.
[0260] In this embodiment, the scheduling device may further include:
[0261] The third obtaining module is used to obtain the fault parameters of each running node in at least one running node in response to the scheduling request of the task of obtaining the target model through training.
[0262] The second determination module is used to determine the fault tolerance performance index value of the running node based on the fault parameters of the running node; the fault tolerance performance index value is used to reflect the fault tolerance performance of the running node.
[0263] The second selection module is used to select at least one second target running node from each of the running nodes based on the fault tolerance performance indicator value of each of the running nodes to perform the task of obtaining the target model through training.
[0264] The third acquisition module can be used to:
[0265] The number of failures of each running node in at least one running node, the number of faulty nodes on the same route as the running node, and a second weight factor and a third weight factor are obtained; the second weight factor is used to reflect the importance of the impact of the running node on the running stability; the third weight factor is used to reflect the importance of the impact of the faulty nodes on the same route as the running node on the running stability.
[0266] The scheduling device may also include:
[0267] The fourth obtaining module is used to obtain the resource usage and non-preemptible resource usage of the tenant corresponding to each of the inference tasks.
[0268] A screening module is used to screen out at least one first tenant whose resource usage does not exceed the non-preemptible amount of resources and at least one second tenant whose resource usage exceeds the non-preemptible amount of resources from the tenants corresponding to each reasoning task based on the resource usage and non-preemptible amount of resources of the tenants corresponding to each reasoning task.
[0269] The third determining module is used to preferentially determine the scheduling order of each first reasoning task corresponding to the at least one first tenant among the reasoning tasks.
[0270] The fourth determining module is used to determine the scheduling order of each second reasoning task corresponding to the at least one second tenant in each of the reasoning tasks after the scheduling order of each of the first reasoning tasks is determined.
[0271] The fourth determining module determines the scheduling order of each second reasoning task corresponding to the at least one second tenant in each reasoning task, which may include:
[0272] Obtaining a weight factor and a total amount of resources of each second tenant in the at least one second tenant and a resource application amount of a second reasoning task corresponding to the second tenant;
[0273] Determine the resource occupancy index value of each second tenant based on the weight factor, total resource amount and resource usage amount of each second tenant and the resource application amount of the second reasoning task corresponding to the second tenant; the resource occupancy index value is used to reflect the resource occupancy of the second tenant;
[0274] The scheduling order of each of the second inference tasks is based on the resource occupancy index value of each of the second tenants.
[0275] The scheduling device may also include:
[0276] The fifth acquisition module is used to obtain the parallel configuration parameters and resource application amount of the task of obtaining the target model through training; the parallel configuration parameters are used to reflect the demand of the task of obtaining the target model through training for parallel computing resources.
[0277] The first adjustment module is used to determine the parallel configuration parameters and resource application amount of the task of obtaining the target model through training. If the training requirements of the task of obtaining the target model through training are not met, the parallel configuration parameters and the resource application amount are adjusted.
[0278] The scheduling device may also include:
[0279] The second adjustment module is used to determine the parallel configuration parameters and resource application amount of the task of obtaining the target model through training, meet the training requirements of the task of obtaining the target model through training, and there is no at least one running node that meets the resource application amount, and adjust the parallel configuration parameters and the resource application amount.
[0280] In another embodiment of the present application, a container cluster management system is provided, which may include: a scheduler and at least one running node.
[0281] Scheduler, which can be used to:
[0282] In response to a scheduling request of at least one target model-based reasoning task, obtaining a resource application amount of each reasoning task in the at least one target model-based reasoning task;
[0283] Obtaining the total amount and available amount of resources of each running node in at least one running node;
[0284] Based on the resource application amount of the inference task and the total amount and available amount of resources of each of the running nodes, determine the resource utilization index value and anti-fragmentation parameter of each of the running nodes; the resource utilization index value is used to reflect the resource utilization of the running node; the anti-fragmentation parameter is used to reflect the impact of scheduling the inference task to the running node on reducing the degree of resource fragmentation of the running node;
[0285] Based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes, at least one first target running node is selected from each of the running nodes to execute the reasoning task.
[0286] At least one first target execution node among the at least one execution node may be used to execute the reasoning task.
[0287] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0288] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0289] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0290] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
Claims
1. A scheduling method, comprising: In response to a scheduling request of at least one target model-based reasoning task, obtaining a resource application amount of each reasoning task in the at least one target model-based reasoning task; Obtaining the total amount and available amount of resources of each running node in at least one running node; Determine the resource utilization index value and anti-fragmentation parameter of each running node based on the resource application amount of the reasoning task and the total amount and available amount of resources of each running node; the resource utilization index value is used to reflect the resource utilization of the running node; The anti-fragmentation parameter is used to reflect the effect of scheduling the inference task to the running node on reducing the resource fragmentation degree of the running node; Based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes, at least one first target running node is selected from each of the running nodes to execute the reasoning task.
2. The scheduling method according to claim 1, wherein the resource utilization index value and the anti-fragmentation parameter of each running node are determined based on the resource application amount of the inference task and the total amount and available amount of resources of each running node, including: Determine the utilization rate of each type of resource of each operating node based on the application amount of each type of resource for the inference task and the total amount and available amount of each type of resource of each operating node; The utilization rate represents the utilization degree of the resource of the type when the reasoning task is scheduled to the running node; Determining a resource utilization index value of each of the operating nodes based on the utilization rate of each type of resource of each of the operating nodes and a first weight factor; The first weight factor is used to reflect the importance of the type of resource to the reasoning task; Based on the resource application amount of the inference task and the available amount of resources of each of the running nodes, the expected remaining amount of resources of each of the running nodes is determined, and the expected remaining amount is used as an anti-fragmentation parameter.
3. The scheduling method according to claim 1, wherein the selecting at least one first target running node from each of the running nodes to execute the reasoning task based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes comprises: If the anti-fragmentation parameter of the running node is used to reflect that scheduling the inference task to the running node has a positive impact on reducing the resource fragmentation degree of the running node, the resource utilization index value of the running node is improved to obtain an adjusted resource utilization index value of each running node; Based on the adjusted resource utilization index value of each of the running nodes, at least one first target running node is selected from each of the running nodes to execute the reasoning task.
4. The scheduling method according to claim 1, further comprising: In response to a scheduling request of a task of the target model obtained through training, obtaining a fault parameter of each operating node in at least one operating node; Determining a fault tolerance performance index value of the running node based on the fault parameters of the running node; The fault tolerance performance index value is used to reflect the fault tolerance performance of the running node; Based on the fault tolerance performance index value of each of the running nodes, at least one second target running node is selected from each of the running nodes to execute the task of obtaining the target model through training.
5. The scheduling method according to claim 4, obtaining the fault parameter of each running node in at least one running node comprises: Obtaining the number of failures of each running node in at least one running node, the number of failed nodes on the same route as the running node, a second weight factor, and a third weight factor; The second weight factor is used to reflect the importance of the impact of the running node on the running stability; The third weight factor is used to reflect the importance of the impact of the faulty node on the same route as the running node on the running stability.
6. The scheduling method according to claim 1, further comprising: Obtaining the resource usage and non-preemptible resource usage of the tenant corresponding to each of the inference tasks; Based on the resource usage and non-preemptible resource amount of the tenants corresponding to each of the inference tasks, at least one first tenant whose resource usage does not exceed the non-preemptible resource amount and at least one second tenant whose resource usage exceeds the non-preemptible resource amount are screened out from the tenants corresponding to each of the inference tasks; Prioritizing the scheduling order of each first reasoning task corresponding to the at least one first tenant among the reasoning tasks; After the scheduling order of each of the first reasoning tasks is determined, the scheduling order of each of the second reasoning tasks corresponding to the at least one second tenant in the reasoning tasks is determined.
7. The scheduling method according to claim 6, wherein determining the scheduling order of each second reasoning task corresponding to the at least one second tenant in each reasoning task comprises: Obtaining a weight factor and a total amount of resources of each second tenant in the at least one second tenant and a resource application amount of a second reasoning task corresponding to the second tenant; Determine the resource occupancy index value of each second tenant based on the weight factor, total resource amount and resource usage amount of each second tenant and the resource application amount of the second reasoning task corresponding to the second tenant; the resource occupancy index value is used to reflect the resource occupancy of the second tenant; The scheduling order of each of the second inference tasks is based on the resource occupancy index value of each of the second tenants.
8. The scheduling method according to claim 4, before obtaining the fault parameter of each running node in at least one running node, further comprising: Acquire the parallel configuration parameters and resource application amount of the task of obtaining the target model through training; the parallel configuration parameters are used to reflect the demand of the task of obtaining the target model through training for parallel computing resources; The parallel configuration parameters and resource application amount of the task of obtaining the target model through training are determined, and if the training requirements of the task of obtaining the target model through training are not met, the parallel configuration parameters and the resource application amount are adjusted.
9. The scheduling method according to claim 8, further comprising: Determine the parallel configuration parameters and resource application amount of the task of obtaining the target model through training to meet the training requirements of the task of obtaining the target model through training, and if there is no at least one running node that meets the resource application amount, adjust the parallel configuration parameters and the resource application amount.
10. A scheduling device, comprising: A first obtaining module, configured to obtain, in response to a scheduling request of at least one reasoning task based on a target model, an application amount of resources for each reasoning task in the at least one reasoning task based on the target model; A second obtaining module is used to obtain the total amount and available amount of resources of each running node in at least one running node; A first determination module is used to determine the resource utilization index value and anti-fragmentation parameters of each of the operating nodes based on the resource application amount of the reasoning task and the total amount and available amount of resources of each of the operating nodes; the resource utilization index value is used to reflect the resource utilization of the operating node; The anti-fragmentation parameter is used to reflect the effect of scheduling the inference task to the running node on reducing the resource fragmentation degree of the running node; The first selection module is used to select at least one first target running node from each of the running nodes to execute the reasoning task based on the resource utilization index value and the anti-fragmentation parameter of each of the running nodes.
Citation Information
Cited By
Power grid multi-source dense state reasoning task scheduling method and device, storage medium and computer equipment
CN122507482A