Task scheduling method and device, electronic equipment and storage medium
By prioritizing tasks based on their priority and arrival time in the AI model training task scheduling system, and rationally allocating computing cluster resources, the problem of low efficiency in training tasks under limited resources is solved, enabling timely task training and efficient resource utilization.
Patent Information
- Application Number
- CN202510928960.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-17
AI Technical Summary
In existing technologies, AI model training task scheduling systems struggle to efficiently and reasonably meet the needs of multiple training tasks under limited computing resources, leading to resource waste and low training efficiency.
By acquiring the resource usage of the computing cluster nodes, the tasks to be trained are sorted according to task priority and arrival time. The highest priority and earliest arriving tasks are scheduled to available nodes for training first, and occupied nodes are released when resources are insufficient to meet training requirements.
This enables high-priority and earliest-arriving tasks to be trained in a timely manner under limited computing resources, improving resource utilization efficiency and the timeliness of training tasks.
Smart Images

Figure CN120803654A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of resource allocation, and in particular to a task scheduling method, device, electronic device, storage medium and computer program product. Background Art
[0002] With the rapid development of artificial intelligence technology, the intelligence level of AI models continues to improve, and the scale of AI model parameters is becoming increasingly large, which has led to a rapid increase in the demand for computing resources, especially high-performance computing resources (such as GPU accelerator cards) during model training.
[0003] In related technologies, task scheduling systems for AI model training use limited computing cluster resources to perform distributed training on AI model training tasks. These systems coordinate the concurrent operation of accelerator cards on multiple nodes in the computing cluster during the same time period to meet the needs of new feature updates or urgent tasks. Therefore, a task scheduling method is urgently needed to meet the training requirements of training tasks while rationally utilizing limited computing resources. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a task scheduling method, apparatus, electronic device, storage medium, and computer program product to meet the training requirements of multiple training tasks while reasonably utilizing limited computing resources.
[0005] According to a first aspect, an embodiment of the present invention provides a task scheduling method, which is applied to a task scheduling system, the method comprising: obtaining a task to be trained and task description information corresponding to the task to be trained, the task description information comprising: a task priority level and information on resources required for training; adding the task to be trained to a queue, wherein the sorting order of the tasks to be trained in the queue is determined according to the task priority level and task arrival time of the task to be trained, the higher the task priority level in the queue, the higher the sorting order of the task to be trained, and the earlier the task arrival time under the same task priority level, the higher the sorting order of the task to be trained; obtaining node resource usage information of a computing cluster used for training and determining available node resources based on the node resource usage information; taking the task to be trained with the highest sorting order in the queue as the target task to be trained, the task to be trained with the highest sorting order representing the task with the highest priority level and the earliest time of joining the queue under the highest priority level; if the available node resources meet the requirements of the resources required for training corresponding to the target task to be trained, scheduling the target task to be trained to the available node resources to perform a training operation on the target task to be trained.
[0006] According to a second aspect, embodiments of the present application provide a task scheduling apparatus applied to a task scheduling system, the apparatus comprising: a first obtaining module configured to obtain a to-be-trained task and task description information corresponding to the to-be-trained task, the task description information comprising: a task priority level and training required resource information; a joining module configured to add the to-be-trained task to a queuing queue, wherein an ordering sequence of the to-be-trained tasks in the queuing queue is determined according to the task priority levels and task arrival times of the to-be-trained tasks, and the to-be-trained task with a higher task priority level in the queuing queue has a more forward ordering sequence, and the to-be-trained task with an earlier task arrival time under the same task priority level has a more forward ordering sequence; a second obtaining module configured to obtain node resource usage information of a computing cluster used for training and determine available node resources based on the node resource usage information; a selecting module configured to select a to-be-trained task with a most forward ordering sequence in the queuing queue as a target to-be-trained task, the to-be-trained task with the most forward ordering sequence representing a task with the highest priority level and the earliest joining queue time under the highest priority level; and a scheduling module configured to, if the available node resources meet the requirements of the training required resources corresponding to the target to-be-trained task, schedule the target to-be-trained task to the available node resources to perform a training operation on the target to-be-trained task.
[0007] According to a third aspect, embodiments of the present application provide an electronic device, comprising: a memory and a processor, which are in communication connection with each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the task scheduling method according to the first aspect or any optional implementation manner of the first aspect.
[0008] According to a fourth aspect, embodiments of the present application provide a computer readable storage medium, which stores computer instructions, and the computer instructions are used to make a computer execute the task scheduling method according to the first aspect or any optional implementation manner of the first aspect.
[0009] According to a fifth aspect, embodiments of the present application provide a computer program product, which comprises computer instructions, and the computer instructions are used to make a computer execute the task scheduling method according to the first aspect or any optional implementation manner of the first aspect.
[0010] The task scheduling method provided by the embodiment of the application can determine the available node resources in the computing cluster by acquiring the node resource usage in the computing cluster, select a task with the highest priority level and the earliest queue joining time under the highest level priority from the queue as a target to-be-trained task, and schedule the target to-be-trained task to the available node resources to perform training operation on the target to-be-trained task if the available node resources meet the requirements of the training resources required by the target to-be-trained task, so that the task with the highest priority level and the earliest queue joining time under the highest level priority can be scheduled and trained in time. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the specific embodiments or the prior art, the drawings needed to be used in the specific embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0012] Figure 1 is an application scenario diagram of the task scheduling method of the embodiment of the application;
[0013] Figure 2 is a flowchart of the task scheduling method of the embodiment of the application;
[0014] Figure 3 is a schematic diagram of the task scheduling method of the embodiment of the application;
[0015] Figure 4 is a schematic diagram of the task scheduling method of the embodiment of the application;
[0016] Figure 5 is a structural block diagram of the task scheduling device of the embodiment of the application;
[0017] Figure 6 is a hardware structure schematic diagram of the electronic device provided by the embodiment of the application. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of the embodiments of the application clearer, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.
[0019] With the rapid development of artificial intelligence technology, the intelligence level of AI models continues to improve, and the parameter scale of AI models is increasingly large, which makes the demand for computing resources, especially high-performance computing resources (such as GPU acceleration cards), rapidly grow during model training. In related technologies, a task scheduling system for AI model training performs distributed training on multiple AI model training tasks to be trained based on limited computing cluster resources, and coordinates the concurrent running of acceleration cards on multiple nodes in the computing cluster in the same time period to meet the use demand of new function updates or urgent tasks. Embodiments of the present application propose a task scheduling method to meet the training requirements of multiple training tasks while reasonably using limited computing resources.
[0020] The task scheduling system in the embodiments of the present application can be in communication connection with a training task submission end. The training task submission end can submit a to-be-trained task through a command line tool, an API or a platform interface, and schedule the selected to-be-trained task to an AI model training service platform in communication connection therewith and integrated with computing cluster resources, and train the corresponding to-be-trained task in combination with training data on the AI model training service platform. Specifically, as shown in Figure 1 The task scheduling system 20 and the AI model training service platform 30 are in communication connection.
[0021] According to the embodiments of the present application, a task scheduling method is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0022] In the present embodiment, a task scheduling method is provided, which can be applied to a task scheduling system. The task scheduling system can be pre-integrated with multiple training task selection strategies. In the embodiments of the present application, the training task selection strategies can include a first-in-first-out strategy (i.e., training the to-be-trained tasks at the front of the queue first) and a polling strategy (polling the to-be-trained tasks in the queue according to the available nodes in the computing cluster to match the tasks that can meet the resource requirements of the to-be-trained tasks for training and perform scheduling training). The number and type of integrated training task selection strategies are not limited in the embodiments of the present application, and can be reasonably selected according to the number of to-be-trained tasks in the task scheduling system, the urgency of the tasks or the training requirements, etc. For example, Figure 2 The flow includes the following steps:
[0023] Step 101, obtaining a to-be-trained task and task description information corresponding to the to-be-trained task, the task description information including: a task priority level and training required resource information;
[0024] Exemplarily, the to-be-trained task and the task description information corresponding to the to-be-trained task can be uploaded to the task scheduling system by a terminal where the task is located. The task description information corresponding to the to-be-trained task in the embodiment of the application is information reflecting a task attribute or information reflecting a task required configuration. For example, the information of the task attribute can include a task priority level, and the information of the task required configuration can include training required resource information, wherein the training required resource information can include information required in the task training process, such as the number of nodes required for training, the number of accelerator cards, and the memory. In the embodiment of the application, the to-be-trained task is encapsulated in the form of a Pod group. By encapsulating the to-be-trained task in the form of a Pod group, the to-be-trained task of distributed training can be queued and scheduled as a scheduling unit of the Pod group, so that a group of associated distributed to-be-trained tasks or all are executed at the same time, or none are executed. For the distributed training task, the problem that the multiple node resources required for distributed training cannot be guaranteed to be synchronized and ready at the same time period, thereby causing the to-be-trained task of distributed training to be long-term stranded in the queuing state, and affecting the timeliness of task training, is avoided. In the embodiment of the application, the Pod group structure is used to store the related information of the to-be-trained task. For example, the Pod group structure can record the information such as the number of nodes required by the to-be-trained task (NodeNum), the queue name (QueueName), the expected running time Duration (unit: second), and the priority (Priority) sequence number. The smaller the priority sequence number is, the higher the priority of the to-be-trained task is. The resource description component (Resource) of the Pod group structure can record the CPU core number, the memory size (unit: byte), and the GPU number required by the to-be-trained task, which represent the training required resource information.
[0025] In step 102, the to-be-trained task is added to a queuing queue. The sorting order of the to-be-trained task in the queuing queue is determined according to the task priority level and the task arrival time of the to-be-trained task. The sorting order of the to-be-trained task with a higher task priority level in the queuing queue is earlier, and the sorting order of the to-be-trained task with an earlier task arrival time under the same task priority level is earlier.
[0026] Exemplarily, the to-be-trained task can set a priority level label for the to-be-trained task on the terminal where the to-be-trained task is located when uploading, and the task scheduling system sets a time stamp for the task arrival time of the received to-be-trained task when receiving the to-be-trained task. In the embodiment of the application, the queuing queue is sorted according to the priority level order of the to-be-trained task and the task arrival time order of the same priority level, and the to-be-trained task with a higher priority level in the queuing queue is sorted in a more front order, and the to-be-trained task with an earlier task arrival time in the same priority level is sorted in a more front order. In the embodiment of the application, the sorting order represents the order of the sorting position of the to-be-trained task in the queuing queue, and the to-be-trained task with a more front sorting position represents that the to-be-trained task can be trained by the task scheduling system earlier. Specifically, the priority level of the task in the embodiment of the application can be divided into three levels, which are high-level priority, medium-level priority and low-level priority. The specific priority setting mode can be set according to the importance, urgency or task association of the to-be-trained task.
[0027] For example, the priority levels of the newly received to-be-trained task A and the to-be-trained task B are both high-level, and the to-be-trained task A and the to-be-trained task B are arranged in the queue position of the high-level task in the queuing queue, and further sorted according to the task arrival time of the to-be-trained task A and the to-be-trained task B in all high-level task queues according to the task arrival time. Since the to-be-trained task A and the to-be-trained task B are newly arrived tasks, the to-be-trained task A and the to-be-trained task B are arranged in the position between the existing high-level priority task and the existing medium-level priority task in the queue. If the to-be-trained task A arrives earlier than the to-be-trained task B, the to-be-trained task A is arranged before the to-be-trained task B, so that the to-be-trained task A is close to the position of the existing high-level priority task, and the to-be-trained task B is close to the position of the existing medium-level priority task. Similarly, the queuing mode of the to-be-trained task of other priority levels is the same.
[0028] In step 103, node resource usage information of a computing cluster used for training is acquired, and available node resources are determined based on the node resource usage information;
[0029] Exemplarily, the node resource usage information can reflect the current usage state information of each node contained in the computing cluster, such as being in a state occupied by a training task or in an idle available state, etc. The node resource usage in the computing cluster can be monitored in a polling manner, all nodes in the computing cluster are acquired first, and the node usage is acquired on each node. Specifically, the nodes in the computing cluster can be divided into occupied nodes (corresponding to NodeInfo as UsingResource) and available nodes (corresponding to NodeInfo as AllocatableResource), and the corresponding state information can monitor the IsSchedulable field in the node information (NodeInfo). When the IsSchedulable field is "true", it indicates that the node is occupied, and when the IsSchedulable field is "false", it indicates that the node is available. The IsSchedulable field of each node can be polled to determine the node resource usage information. For any node, the remaining available resource capacity information of each node can also be determined based on the node resource capacity and the resource capacity usage information (such as accelerator usage or memory usage).
[0030] Step 104, the task with the highest priority level and the earliest time of joining the queue under the highest priority level in the priority processing queue is processed first, and the task is taken as the target training task.
[0031] Step 105, if the available node resource meets the requirement of the training resource required by the target training task, the target training task is scheduled to the available node resource to perform training operation on the target training task.
[0032] Exemplarily, for the determined target training task, the available node resource and the training resource information of the target training task can be combined to determine whether the available node resource can train the target training task. If the available node resource meets the requirement of the training resource required by the target training task, the target training task is scheduled to the available node resource to perform training operation on the target training task.
[0033] The task scheduling method provided by the embodiment of the present application can determine the available node resources in the computing cluster by obtaining the node resource usage in the computing cluster, select the task with the highest priority level and the earliest queue joining time under the highest level priority from the queue as the target training task, and schedule the target training task to the available node resources to perform training operation on the target training task if the available node resources meet the requirements of the training resources corresponding to the target training task, so as to ensure that the task with the highest priority level and the earliest queue joining time under the highest level priority can be scheduled and trained in time.
[0034] As an optional implementation of the embodiment of the present application, the method further comprises: if the available node resources do not meet the requirements of the training resources corresponding to the target training task, responding to the release waiting operation of the occupied node in the computing cluster for training, the occupied node represents the node currently performing the training task.
[0035] For example, if the available node resources do not meet the requirements of the training resources corresponding to the target training task, the responding release waiting operation of the occupied node in the computing cluster for training, the occupied node in the computing cluster represents the node currently training other tasks. When the occupied node finishes the current training task, the resources of the occupied node are released, and the state of the node is changed from occupied to available. The number of nodes in the available state can be counted in real time. If the number of nodes in the available state meets the requirements of the number of nodes required by the target training task, the scheduling operation is responded to the target training task to perform training operation on the target training task, so as to realize the early locking of the available node resources corresponding to the target training task, and ensure that the task with the highest priority level and the earliest queue joining time under the highest level priority can be trained in time, avoiding long waiting time.
[0036] As an optional implementation of the embodiment of the present application, after step 103, the method further comprises: determining whether the available node resources meet the requirements of the training resources corresponding to the training task with the highest priority level and the earliest queue joining time under the highest level priority in the queue; for example, the training task with the highest priority level and the earliest queue joining time under the highest level priority represents the task with the highest priority level and the earliest queue joining time under the highest level priority, and the determination of whether the available node resources meet the requirements of the training resources can be determined according to the comparison result of the available node number in the monitored node resource usage information and the node number contained in the training resource information of the training task with the highest priority level and the earliest queue joining time under the highest level priority.
[0037] If yes, the training task with the highest priority level and the earliest queue joining time under the highest level priority is selected as the target training task.
[0038] For example, if the number of available nodes in the monitored node resource usage information is greater than or equal to the number of nodes required for training the topmost to-be-trained task, it is determined that the required resources for training are met, and the to-be-trained task with the highest ranking order is taken as the target to-be-trained task and is scheduled and trained.
[0039] If not, the new to-be-trained task is obtained in the order from front to back in the queuing queue. For example, if not, the new to-be-trained task is obtained in the order from front to back according to the to-be-trained tasks contained in the queuing queue, that is, the next new to-be-trained task is the task adjacent to the position of the topmost to-be-trained task.
[0040] The available node resources are matched with the training resource requirements corresponding to the new to-be-trained task, and the new to-be-trained task that is successfully matched is taken as the target to-be-trained task.
[0041] For example, the number of nodes required for training corresponding to each new to-be-trained task obtained is compared with the number of available nodes in the computing cluster, and when the number of available nodes in the computing cluster is greater than or equal to the number of nodes required for training corresponding to any new to-be-trained task, it is determined that the matching is successful. The new to-be-trained task that is successfully matched is taken as the target to-be-trained task for scheduling and training operation, which avoids the long waiting state of the node resources in the computing cluster, so that the limited computing resources can be reasonably used.
[0042] As an optional implementation manner of the embodiment of the present application, the method further comprises: in response to receiving an adjustment instruction for the priority of the to-be-trained task in the queuing queue, adjusting the ranking order of the corresponding to-be-trained task in the queuing queue according to the adjustment instruction.
[0043] For example, for the queuing queue obtained by ranking according to the level order of the task priority and the time order of the arrival of the tasks with the same level priority, the ranking order of the corresponding to-be-trained task in the queuing queue can be adjusted based on the received priority adjustment instruction for the to-be-trained task. Specifically, the priority adjustment request for the to-be-trained task whose priority needs to be adjusted can be reported, and when the approval is passed, an adjustment instruction for adjusting the priority label of the corresponding to-be-trained task in the queuing queue is generated. The adjustment instruction can include the identification information of the to-be-trained task whose priority needs to be adjusted and the corresponding priority adjustment information (such as adjusting the priority of the to-be-trained task A from the current "medium level priority" to "high level priority"). By receiving the external adjustment instruction, the priority of the corresponding to-be-trained task in the queuing state is adjusted, so that the ranking order of the corresponding to-be-trained task in the queuing queue is moved forward, which ensures that the to-be-trained task with sudden training demand is trained in time and improves the flexibility of the training task scheduling and training process.
[0044] As an optional implementation of the embodiment of the application, the to-be-trained task includes a single-machine task and a whole-machine task, and the computing cluster includes a first type of node for training the single-machine task and a second type of node for training the whole-machine task, the single-machine task representing a to-be-trained task requiring one node to complete training, and the whole-machine task representing a to-be-trained task requiring multiple nodes to simultaneously perform training.
[0045] Exemplarily, in the embodiment of the application, the multiple nodes in the computing cluster can be divided into the first type of node for training the single-machine task and the second type of node for training the whole-machine task; or the multiple nodes in the computing cluster can be divided into the first type of node for training the single-machine task, the second type of node for training the whole-machine task, and a flexible node, which can be flexibly used based on the training requirements of the single-machine task or the whole-machine task. In the embodiment of the application, the single-machine task represents a to-be-trained task requiring one node to complete training, for example, for a node including eight accelerator cards, the single-machine task only needs to use less than or equal to eight accelerator cards in one node to complete the training task; the whole-machine task represents a to-be-trained task requiring multiple nodes to simultaneously perform training, for example, all accelerator cards in multiple nodes are simultaneously used for distributed training. By dividing the first type of node and the second type of node in the computing cluster, it can be ensured that the single-machine task with a short training time can be timely trained, and the training timeliness of the single-machine task is avoided from being affected by the long-time occupation of the node by the whole-machine task.
[0046] The method provided in the embodiment of the application further includes: when the to-be-trained task is a single-machine task, queuing the to-be-trained task as the single-machine task according to the hierarchical order of the task priority and the time order of the arrival of the tasks with the same priority, to obtain a first queuing sub-queue; and when the to-be-trained task is a whole-machine task, queuing the to-be-trained task as the whole-machine task according to the hierarchical order of the task priority and the time order of the arrival of the tasks with the same priority, to obtain a second queuing sub-queue.
[0047] Exemplarily, in the embodiment of the application, the to-be-trained task received can be distinguished as a single-machine task or a whole-machine task based on the number of nodes required by the to-be-trained task received, if it is a single-machine task, it is queued in the first queuing sub-queue, if it is a whole-machine task, it is queued in the whole-machine task; or the waiting time of the single-machine task in the current task scheduling system can be counted, when the waiting time of the single-machine task in the queuing queue is counted to be longer than a preset time length, the nodes in the computing cluster can be divided to obtain the first type of node and the second type of node, and the single-machine task and the whole-machine task included in the current queuing queue can be identified, the identified single-machine task is queued in the first queuing sub-queue, and the identified whole-machine task is queued in the second queuing sub-queue.
[0048] The single-machine tasks in the first queuing sub-queue are preferentially scheduled on the first type node; and the whole-machine tasks in the second queuing sub-queue are preferentially scheduled on the second type node.
[0049] Exemplarily, by dividing the first queuing sub-queue and the second queuing sub-queue, the single-machine tasks in the first queuing sub-queue can be preferentially scheduled on the first type node to perform training operation on the to-be-trained tasks by using the first type node. If the current first type node resource does not meet the training requirement of the single-machine tasks in the queue, the single-machine tasks can be scheduled to the idle node performing whole-machine task training or to the pre-configured idle mobile node. Similarly, the whole-machine tasks in the second queuing sub-queue are preferentially scheduled on the second type node to perform training operation on the to-be-trained tasks by using the second type node. If the second type node resource does not meet the whole-machine task training requirement, the idle node performing single-machine training task or the idle mobile node can be used, which is not limited in the embodiment of the application. For the computing cluster, 1-2 nodes can be allocated as the first type node, and the others as the second type node. The skilled in the art can flexibly set according to the actual number of single-machine tasks and whole-machine tasks to ensure that the single-machine tasks and the whole-machine tasks can be timely trained.
[0050] As an optional implementation manner of the embodiment of the application, the method further comprises: if the first type node includes multiple nodes, the single-machine tasks in the first queuing sub-queue are sequentially scheduled on the same first type node for training; if the number of remaining accelerator cards in the current first type node does not meet the training requirement of any single-machine task in the sequentially scheduled single-machine tasks, a new first type node is enabled to perform training operation on the target to-be-trained task.
[0051] Exemplarily, taking an example that each node includes 8 GPU accelerator cards, for example, 2 first type nodes A and node B are configured, node A is preferentially scheduled to be used, and the received single-machine tasks are scheduled to node A for training. When the remaining GPU accelerator cards in node A do not meet the single-machine task training requirement, node B is enabled for single-machine task training. For example, node A is used by one single-machine task with 2 GPU accelerator cards, and 8 accelerator cards in node B are not used. When a new single-machine task needs 1 GPU accelerator card, the new single-machine task is scheduled to node A which has been used with 2 GPU accelerator cards. When the new single-machine task needs 7 accelerator cards, the new single-machine task can be scheduled to node B. For the first type node, the number of used GPU accelerator cards can be sorted from large to small. For the new single-machine task, the number of GPU accelerator cards of each node is polled in the order from large to small. When the number of accelerator cards in the node meets the training resource requirement of the single-machine task, the polling is stopped.
[0052] As an optional implementation of the embodiment of the application, the method further comprises: sorting the high-priority training tasks in the second queued sub-queue according to training time consumption to obtain a first sorting result corresponding to the high-priority training tasks; sorting the training tasks of other priority levels in the second queued sub-queue according to training time consumption to obtain a second sorting result corresponding to the training tasks of other priority levels; and reconstructing the second queued sub-queue according to the first sorting result and the second sorting result to obtain a third queued sub-queue.
[0053] Exemplarily, the training time consumption of the training task in the embodiment of the application can be predicted according to the trained task training runtime prediction model, the high-priority training tasks are sorted according to the predicted training time consumption of the training tasks to obtain a first sorting result corresponding to the high-priority training tasks; similarly, the training tasks of other priority levels in the second queued sub-queue are sorted according to the training time consumption to obtain a second sorting result corresponding to the training tasks of other priority levels, and the second queued sub-queue is reconstructed according to the first sorting result and the second sorting result to obtain a third queued sub-queue. By reconstructing the second queued sub-queue according to the training time consumption, training samples of different training time consumption can be flexibly selected for training based on the training time length requirements of the tasks in the third queued sub-queue, so that the node resources are not occupied for a long time by the training tasks with long training time consumption.
[0054] In the embodiment of the application, the task training runtime prediction model can obtain a plurality of completed training tasks, and use the historical training runtime of the corresponding type of task as training data according to the influence parameters of the task type (such as image recognition type or natural language processing type) and the batch size (Batch Size), the number of rounds (Epochs), the training data size, etc. of the corresponding type of training task on the task training runtime, and the XGBoost model in machine learning is trained to obtain a task training runtime prediction model for predicting the task training runtime.
[0055] In a specific training process, the collected training data is randomly shuffled, and a part of the training data (e.g., 20% of the training data) can be reserved. 75% of the remaining 80% of the training data is used for model training, and the remaining 25% of the training data is used for model testing. The XGBoost model is configured with parameters such as ridge regression parameters, maximum depth of tree, learning rate, number of base learners, subsampling ratio, and column sampling ratio. According to the configured parameters, the training data is converted into a format that can be used by the XGBoost model to start model training. According to the test results obtained by combining the test data, the XGBoost model parameters are adjusted until the effect reaches the expected standard. During training, different parameters of the XGBoost model can be set to train multiple models. Each model is used to predict the reserved 20% of the training data, and the model with the highest prediction accuracy is selected as the actual task training running time prediction model. The model training method is not limited in the embodiments of the application, and can be selected as needed by those skilled in the art.
[0056] As an optional implementation of the embodiments of the application, the method further comprises:
[0057] Step a1, the task with the longest training time in the highest priority level and the highest priority level in the third queuing sub-queue is selected as the target training task;
[0058] Step a2, when the number of nodes in the available state in the second type of node does not meet the requirement of the number of nodes required for training corresponding to the target training task, a calculation operation is performed on the remaining occupied time length of the occupied node in the second type of node;
[0059] For example, the calculation operation of the remaining occupied time length of the occupied node in the second type of node can be based on the predicted running time length of the whole machine task running on the current occupied node predicted by the task training running time prediction model. The running time length of the current whole machine task can be obtained by combining the start training time of the current running whole machine task. The remaining occupied time length of the occupied node can be obtained based on the predicted running time length and the running time length. For example, when the training time length is predicted to be 2 days based on the whole machine task related training running parameters, and the current running time length is 1 day, the remaining occupied time length of the occupied node is 1 day.
[0060] Step a3, the remaining occupied time lengths of all occupied nodes are sorted;
[0061] For example, different whole machine tasks running on different occupied nodes can be different, resulting in different remaining occupied time lengths of the occupied nodes. The remaining occupied time length of each occupied node running the whole machine task is calculated, and the remaining occupied time lengths of all occupied nodes are sorted.
[0062] Step a4, select the target number of occupied nodes corresponding to the remaining occupied time length in ascending order as the target available computing nodes, the target number is determined according to the difference between the training required node number corresponding to the target training task and the current number of nodes in the available state;
[0063] For example, in order to train the current target training task in time, the target number of occupied nodes corresponding to the remaining occupied time length is selected in ascending order, that is, a group of occupied nodes that are idle out as soon as possible are selected from the computing cluster as target available computing nodes, and the specific target number is determined according to the difference between the training required node number corresponding to the target training task and the current number of nodes in the available state. For example, the current target training task requires 5 nodes for training, and the computing cluster contains 6 computing nodes, the current available node number is 1 (node A), and the other 5 nodes are in the occupied state. The remaining occupied time length of the 5 occupied nodes (nodes B, C, D, E and F) is 1 hour, 2 hours, 3 hours, 4 hours and 5 hours respectively. Since the target training task needs 5 nodes for training, node A is in the available state, and only 4 occupied nodes need to be selected. According to the remaining occupied time length of the current 5 occupied nodes, it can be seen that the remaining occupied time length of nodes B, C, D and E is the shortest, so nodes B, C, D and E are selected as target available computing nodes.
[0064] Step a5, if the remaining occupied time length of the target available computing node is reduced to zero, the target available computing node is converted from the occupied state to the available state. For example, the remaining occupied time length of nodes B, C, D and E is monitored. When the remaining occupied time length of nodes B, C, D and E is reduced to zero, it indicates that the training task of occupied nodes B, C, D and E is completed, and the target available computing nodes B, C, D and E are converted from the occupied state to the available state.
[0065] Step a6, the target training task is scheduled to the corresponding node resource in the available state to perform training operation on the target training task. For example, when nodes B, C, D and E are converted from the occupied state to the available state, the nodes in the available state in the current computing cluster include nodes A, B, C, D and E, which meet the training requirements of the target training task. The target training task is scheduled to the corresponding node resource in the available state to perform training operation on the target training task.
[0066] As an optional implementation of the embodiment of the present application, the method further comprises: when the remaining occupied time length of all the target available computing nodes decreases to zero, selecting a trainable task from the first sorting result or the second sorting result and training the trainable task using the node in the available state, wherein the trainable task represents that the node in the available state can be used for training, and the latest training completion time of the trainable task is less than the time corresponding to the time when the remaining occupied time length of all the target available computing nodes decreases to zero.
[0067] For example, since the selected target available computing node is the earliest idle occupied node, and the remaining occupied time length of different target available computing nodes can be different, it is necessary to perform resource release waiting operation on each target available computing node, for example, node B needs to wait for one hour, node C needs to wait for two hours, node D needs to wait for three hours, and node E needs to wait for four hours. During the waiting process, a trainable task can be selected from the first sorting result or the second sorting result, and the selected trainable task can be trained using the node in the available state. Specifically, the trainable task can be selected from the side with the minimum time consumption in the first sorting result or the second sorting result, and the trainable task that can be completed before the remaining occupied time length of all the target available computing nodes decreases to zero is screened out for training.
[0068] For example, since the selected target available computing node is the earliest idle occupied node, and the remaining occupied time length of different target available computing nodes can be different, it is necessary to perform resource release waiting operation on each target available computing node, for example, node B needs to wait for one hour, node C needs to wait for two hours, node D needs to wait for three hours, and node E needs to wait for four hours. During the waiting process, a trainable task can be selected from the first sorting result or the second sorting result, and the selected trainable task can be trained using the node in the available state. Specifically, the trainable task can be selected from the side with the minimum time consumption in the first sorting result or the second sorting result, and the trainable task that can be completed before the remaining occupied time length of all the target available computing nodes decreases to zero is screened out for training. Figure 3As shown, the selected target to-be-trained task with the highest priority level and the longest training time under the highest level priority needs to wait for 4 occupied nodes, i.e., node A, node B, node C and node D, and the heights of the 4 rectangles (rectangle A, rectangle B, rectangle C and rectangle D) represent the corresponding remaining occupied time lengths of the occupied nodes. During the waiting for the 4 occupied nodes to be released, the trainable task A that can be trained after node A and node B are converted into available states and the trainable task A completes the training operation before node C is converted into an available state can be found from the first sorting result or the second sorting result. Similarly, the trainable task B that can be trained using node A, node B and node C (or using node B and node C) and completes the training before node D is converted into an available state can be found after node C is converted into an available state. Therefore, during the waiting for the 4 nodes to be released, the waiting time can be fully utilized to train other to-be-trained tasks in the queuing state and meeting the training requirements, and after the 4 nodes are all converted into available states, the 4 nodes are used to train the target to-be-trained task, so that the target to-be-trained task can be trained in time and other trainable tasks can be trained during the training waiting process, and the utilization of limited computing resources is improved. In the embodiment of the application, the first sorting result is preferentially polled, and when no trainable task meeting the training condition is polled, the second sorting result is polled, so as to ensure that the training task with a high level priority is trained in time.
[0069] As an optional implementation manner of the embodiment of the application, the method further includes: counting the queuing time lengths of the to-be-trained tasks other than the target to-be-trained task in the third queuing sub-queue; if there is a to-be-trained task with a queuing time length greater than the target time length, the queuing position of the to-be-trained task in the third queuing sub-queue is moved to a position adjacent to the position of the target to-be-trained task, so that when the target to-be-trained task completes the training, the to-be-trained task with the queuing time length greater than the target time length is taken as a new target to-be-trained task.
[0070] Exemplarily, when the third queue sub-queue sorted according to the training time consumption length, the training of the task with long time consumption is prioritized, so that there may be a task with a queue time longer than the target time and still not trained in the third queue sub-queue. The application embodiment moves the queue position of the task with a queue time longer than the target time in the third queue sub-queue to the adjacent position of the target training task, so that when the target training task is completed, the task with a queue time longer than the target time is used as a new target training task, and the task with a queue time longer than the target time is trained in time. The target time in the application embodiment can be 2 days (48 hours), and those skilled in the art can determine it according to actual needs.
[0071] As an optional implementation of the embodiment of the application, the method further comprises: when the task with a queue time longer than the target time includes multiple tasks, sorting the multiple tasks according to the priority level order and the time order of the tasks with the same priority level; and moving the sorted multiple tasks to the adjacent position of the target training task, wherein the task with the highest priority level and the earliest arrival among the sorted multiple tasks is adjacent to the current target training task.
[0072] Exemplarily, for example, the task with a queue time longer than the target time includes task A, task B and task C, wherein task A and task B are both high-level priority, task A arrives earlier than task B, and task C is medium-level priority, the processing order of the three tasks is to process task A first, then process task B, and finally process task C. Task A is adjacent to the current target training task, so that the current target training task is trained to completion and then task A is processed, followed by task B and finally task C.
[0073] As an optional implementation of the embodiment of the application, the method further comprises: obtaining a task waiting time corresponding to a plurality of completed training tasks, the task waiting time representing the time interval between the task arrival time and the task training start time; calculating a plurality of preset time length evaluation indexes based on the task waiting time; and responding to the task scheduling evaluation operation based on the calculation result.
[0074] Exemplarily, the task waiting time of the plurality of scheduled and completed training tasks in the task scheduling system within a preset interval (such as a month or a week) can be obtained periodically, the average waiting time, the median of the task waiting time, and the 95% quantile of the plurality of task waiting times are calculated, and the scheduling strategy of the task scheduling system is evaluated according to the calculation results. The type of the evaluation index is not limited in the embodiments of the present application, and can be determined by the person skilled in the art according to the actual needs.
[0075] As a specific embodiment of the present application, as shown in Figure 4 The received training task to be trained and the corresponding task description information are constructed in the form of a Pod group structure, so that the overall scheduling is performed in the form of a Pod group structure when training scheduling is performed. The task training running time prediction module is used for predicting the task training running time of the training task to be trained, and the prediction result is written into the Pod group structure. The computing cluster resource monitoring module is used for obtaining the current node resource usage information of the computing cluster to determine the available node resources and the occupied node resource information. The queuing queue management model adds the received training task to be trained to the single-machine task queue or the whole-machine task queue in the queuing queue according to the task description information of the training task to be trained contained in the received Pod group structure, and synchronizes the current node resource usage information of the computing cluster for training collected by the computing cluster resource monitoring module to the queuing queue management module, so that the target training model is scheduled for training based on the current node resource usage information and the task description information of the training task to be trained when the training task to be trained in the queuing queue is scheduled.
[0076] In the embodiments, a task scheduling device is also provided, which is used to implement the above embodiments and preferred embodiments, and will not be described herein. The term "module" used below can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware or a combination of software and hardware is also possible and is contemplated.
[0077] The embodiments provide a task scheduling device, as shown in Figure 5 The device comprises:
[0078] The first obtaining module 501 is configured to obtain a training task to be trained and task description information corresponding to the training task to be trained, wherein the task description information comprises a task priority level and training required resource information.
[0079] Adding module 502, configured to add the task to be trained to a queue, wherein the sorting order of the tasks to be trained in the queue is determined according to the task priority level and the task arrival time of the tasks to be trained, wherein the higher the task priority level in the queue, the higher the sorting order of the tasks to be trained; and the earlier the task arrival time of the tasks to be trained under the same task priority level, the higher the sorting order of the tasks to be trained;
[0080] A second acquisition module 503 is configured to acquire node resource usage information of a computing cluster used for training and determine available node resources based on the node resource usage information;
[0081] A selection module 504 is configured to select the task to be trained that is ranked highest in the queue as the target task to be trained, wherein the task to be trained that is ranked highest in the queue represents the task with the highest priority and is the task that was added to the queue the earliest.
[0082] Scheduling module 505 is configured to schedule the target task to be trained to the available node resources if the available node resources meet the training resource requirements corresponding to the target task to be trained, so as to perform training operations on the target task to be trained. For detailed descriptions, please refer to the corresponding descriptions of the above method embodiments, which will not be repeated here.
[0083] The task scheduling device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0084] The task scheduling device provided in this embodiment can clarify the available node resources in the computing cluster by obtaining the node resource usage in the computing cluster, and select the task with the highest priority level and the earliest time of joining the queue under the highest priority level from the queue as the target task to be trained. If the available node resources meet the requirements of the training resources corresponding to the target task to be trained, the target task to be trained will be scheduled to the available node resources to perform training operations on the target task to be trained, ensuring that the task with the highest priority level and the earliest time of joining the queue under the highest priority level can be scheduled for training in a timely manner.
[0085] As an optional implementation of an embodiment of the present invention, the device also includes: a response module, which is used to respond to the release waiting operation of the occupied nodes in the computing cluster used for training if the available node resources do not meet the requirements of the training resources corresponding to the target task to be trained, and the occupied nodes represent the nodes that are currently executing the training task.
[0086] As an optional implementation of the embodiment of the present application, the device further comprises: a first determining module configured to determine whether the available node resources meet the requirements of the training required resources corresponding to the task with the highest ranking order in the queuing queue; a second determining module configured to, if the requirements are met, take the task with the highest ranking order as the target task to be trained; a third obtaining module configured to, if the requirements are not met, sequentially obtain new tasks to be trained from the queuing queue in the order from front to back; and a third determining module configured to match the available node resources with the training required resources corresponding to the new task to be trained, and take the new task to be trained that is successfully matched as the target task to be trained.
[0087] As an optional implementation of the embodiment of the present application, the device further comprises: an adjusting module configured to, in response to receiving an adjusting instruction for the priority of the task to be trained in the queuing queue, adjust the ranking order of the corresponding task to be trained in the queuing queue according to the adjusting instruction.
[0088] As an optional implementation of the embodiment of the present application, the task to be trained comprises a single-machine task and a whole-machine task, the computing cluster comprises a first type of node for training the single-machine task and a second type of node for training the whole-machine task, the single-machine task represents a task to be trained that needs to be completed by one node, and the whole-machine task represents a task to be trained that needs to be trained by multiple nodes simultaneously; the device further comprises: a first queuing module configured to, when the task to be trained is a single-machine task, queue the task to be trained as a single-machine task in the order of the level of the task priority and the time of arrival of the task with the same level of priority, to obtain a first queuing subqueue, wherein the time of arrival of the task represents the time when the task scheduling system receives the task to be trained; a second queuing module configured to, when the task to be trained is a whole-machine task, queue the task to be trained as a whole-machine task in the order of the level of the task priority and the time of arrival of the task with the same level of priority, to obtain a second queuing subqueue; a second scheduling module configured to preferentially schedule the single-machine task in the first queuing subqueue on the first type of node; and a third scheduling module configured to preferentially schedule the whole-machine task in the second queuing subqueue on the second type of node.
[0089] As an optional implementation of the embodiment of the present application, the device further comprises: a fourth scheduling module configured to, if the first type of node comprises multiple nodes, sequentially schedule the single-machine task in the first queuing subqueue on the same first type of node for training; and an enabling module configured to, if the number of remaining acceleration cards in the current first type of node does not meet the training requirements of any of the sequentially scheduled single-machine tasks, enable a new first type of node to perform the training operation on the target task to be trained.
[0090] As an optional implementation of an embodiment of the present invention, the device also includes: a first sorting module, used to sort the high-priority tasks to be trained in the second queuing sub-queue according to training time, and obtain a first sorting result corresponding to the high-priority tasks to be trained; a second sorting module, used to sort the other-priority tasks to be trained in the second queuing sub-queue according to training time, and obtain a second sorting result corresponding to the other-priority tasks to be trained; and a reconstruction module, used to reconstruct the second queuing sub-queue based on the first sorting result and the second sorting result to obtain a third queuing sub-queue.
[0091] As an optional implementation of the embodiment of the present invention, the apparatus further includes: a first selection module, configured to select the task to be trained with the highest priority level and the longest training time at the highest priority level in the third queuing sub-queue as the target task to be trained; a first calculation module, configured to respond to a calculation operation on the remaining occupied time of the occupied nodes in the second type of nodes when the number of available nodes in the second type of nodes does not meet the requirement for the number of training nodes required for the target task to be trained; a third sorting module, configured to sort the remaining occupied time of all occupied nodes; a first selection module, configured to select, in ascending order, occupied nodes corresponding to the target number of remaining occupied time as target available computing nodes, where the target number is determined based on the difference between the number of training nodes required for the target task to be trained and the number of nodes currently in the available state; a conversion module, configured to convert the target available computing node from an occupied state to an available state if the remaining occupied time of the target available computing node drops to zero; and a fourth scheduling module, configured to schedule the target task to be trained to the corresponding node resource in the available state to perform a training operation on the target task to be trained.
[0092] As an optional implementation manner of an embodiment of the present invention, the device also includes: a second selection module, which is used to select a trainable task from the first sorting result or the second sorting result and use the nodes in an available state for training before the remaining occupied time of all the target available computing nodes drops to zero, wherein the trainable task indicates that training can be performed using the nodes currently in an available state, and the latest completion time of the training of the trainable task is less than the time corresponding to when the target remaining occupied time drops to zero, and the target remaining occupied time indicates the longest remaining occupied time among all the target available computing nodes.
[0093] As an optional implementation of the embodiment of the present application, the device further comprises: a statistical module, configured to count the queuing time length of the training tasks other than the target training task in the third queuing sub-queue; and a first moving module, configured to move the queuing position of the training task in the third queuing sub-queue to the position adjacent to the position of the target training task, if there is a training task with a queuing time length greater than the target time length, so that the training task with the queuing time length greater than the target time length is taken as a new target training task when the target training task is completed.
[0094] As an optional implementation of the embodiment of the present application, the device further comprises: a fourth sorting module, configured to sort the plurality of training tasks according to the priority level order and the time of arrival order of the training tasks with the same priority level, if the training task with the queuing time length greater than the target time length comprises a plurality of training tasks; and a second moving module, configured to move the plurality of sorted training tasks to the positions adjacent to the position of the target training task, wherein the training task with the highest priority level and the earliest arrival among the plurality of sorted training tasks is adjacent to the current target training task.
[0095] As an optional implementation of the embodiment of the present application, the device further comprises: a third obtaining module, configured to obtain a plurality of task waiting time lengths corresponding to the training tasks that have been completed, the task waiting time length representing the time interval length between the time of arrival of the task and the time of starting training of the task; a second calculating module, configured to calculate a plurality of preset time length evaluation indexes based on the task waiting time length; and an evaluation module, configured to respond to the task scheduling evaluation operation based on the calculation result.
[0096] Further function descriptions of the above modules are the same as those of the corresponding embodiments, and will not be described here.
[0097] The embodiment of the present application further provides an electronic device, which has the task scheduling device shown in the above Figure 5 .
[0098] Please refer to Figure 6 , Figure 6 is a structural schematic diagram of a device provided by an optional embodiment of the present application, as shown in Figure 6As shown, the device can include at least one processor 601, such as a central processing unit (CPU), at least one communication interface 603, a memory 604, and at least one communication bus 602. The communication bus 602 is used to realize the connection and communication between the components. The communication interface 603 can include a display, a keyboard, and can also include a standard wired interface and a wireless interface. The memory 604 can be a high-speed volatile random access memory (RAM), and can also be a non-volatile memory, such as at least one disk memory. The memory 604 can also be at least one storage device located away from the aforementioned processor 601. The processor 601 can be combined with Figure 5 The described device, the memory 604 stores an application program, and the processor 601 calls the program code stored in the memory 604 to execute any of the above method steps.
[0099] The communication bus 602 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 602 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0100] The memory 604 can include volatile memory, such as random access memory (RAM); the memory can also include non-volatile memory, such as flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); the memory 604 can also include a combination of the above types of memory.
[0101] The processor 601 can be a central processing unit (CPU), a network processor (NP), or a combination of CPU and NP.
[0102] The processor 601 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0103] Optionally, the memory 604 is also used to store program instructions. The processor 601 can call the program instructions to implement the application Figure 2 The task scheduling method shown in the embodiment.
[0104] An embodiment of the present invention further provides a non-transitory computer storage medium, wherein the computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the processing method of the task scheduling method in any of the above method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk (HDD) or a solid-state drive (SSD); the storage medium can also include a combination of the above types of memory.
[0105] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0106] Although embodiments of the present application have been described in conjunction with the drawings, various modifications and variations can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and variations are also within the scope of the present application as defined by the appended claims.
Claims
1. A task scheduling method, characterized in that: Applied to a task scheduling system, the method includes: Obtaining a task to be trained and task description information corresponding to the task to be trained, wherein the task description information includes: a task priority level and information on resources required for training; Adding the tasks to be trained to a queue, wherein the sorting order of the tasks to be trained in the queue is determined according to the task priority level and the task arrival time of the tasks to be trained, wherein the tasks to be trained with higher task priority levels in the queue are sorted higher, and the tasks to be trained with earlier task arrival times under the same task priority levels are sorted higher; Obtaining node resource usage information of a computing cluster used for training and determining available node resources based on the node resource usage information; The task to be trained that is ranked first in the queue is used as the target task to be trained, wherein the task to be trained that is ranked first represents the task with the highest priority and is the task that was added to the queue the earliest under the highest priority level; If the available node resources meet the requirements of the training resources corresponding to the target task to be trained, the target task to be trained is scheduled to the available node resources to perform a training operation on the target task to be trained.
2. The method according to claim 1, characterized in that The method further comprises: If the available node resources do not meet the requirements of the training resources corresponding to the target task to be trained, a release waiting operation is performed on the occupied nodes in the computing cluster used for training, where the occupied nodes represent nodes that are currently executing the training task.
3. The method according to claim 1, characterized in that After obtaining the node resource usage information of the computing cluster used for training and determining the available node resources based on the node resource usage information, the method further includes: Determining whether the available node resources meet the training resource requirements corresponding to the to-be-trained task that is ranked first in the queue; If the condition is satisfied, the task to be trained with the highest ranking order is used as the target task to be trained; If not, new tasks to be trained are acquired from the queue in order from front to back; The available node resources are matched with the training resource requirements corresponding to the new task to be trained, and the new task to be trained that is successfully matched is used as the target task to be trained.
4. The method according to claim 1, wherein The method further comprises: In response to receiving an instruction to adjust the priority of the tasks to be trained in the queuing queue, the sorting order of the corresponding tasks to be trained in the queuing queue is adjusted according to the adjustment instruction.
5. The method according to claim 1, wherein The tasks to be trained include stand-alone tasks and whole-machine tasks. The computing cluster includes a first type of node for training stand-alone tasks and a second type of node for training whole-machine tasks. Stand-alone tasks represent tasks to be trained that require one node to complete training, and whole-machine tasks represent tasks to be trained that require multiple nodes to train simultaneously. After obtaining the task to be trained, the method further includes: When the task to be trained is a stand-alone task, the task to be trained, which is a stand-alone task, is queued according to the order of task priority and the order of arrival time of tasks of the same priority level to obtain a first queuing sub-queue, wherein the task arrival time represents the time when the task scheduling system receives the task to be trained; When the tasks to be trained are whole-machine tasks, the tasks to be trained as whole-machine tasks are queued according to the order of task priorities and the order of arrival time of tasks of the same priority level to obtain a second queuing sub-queue; Prioritizing scheduling of stand-alone tasks in the first queuing sub-queue on nodes of the first type; The whole machine tasks in the second queuing sub-queue are preferentially scheduled on the second type of nodes.
6. The method according to claim 5, characterized in that The method further comprises: If there are multiple nodes of the first type, the single-machine tasks in the first queuing sub-queue are sequentially scheduled to the same node of the first type for training; If the number of remaining accelerator cards in the current first type node does not meet the training requirement of any single-machine task in the sequentially scheduled single-machine tasks, a new first type node is enabled to perform a training operation on the target task to be trained.
7. The method according to claim 5, characterized in that The method further comprises: Sorting the high-priority tasks to be trained in the second queuing sub-queue according to training time consumption to obtain a first sorting result corresponding to the high-priority tasks to be trained; Sorting the tasks to be trained of other priority levels in the second queuing sub-queue according to training time consumption to obtain a second sorting result corresponding to the tasks to be trained of other priority levels; The second queuing sub-queue is rebuilt according to the first sorting result and the second sorting result to obtain a third queuing sub-queue.
8. The method according to claim 7, characterized in that The method further comprises: The task to be trained with the highest priority level and the longest training time in the third queue sub-queue is used as the target task to be trained; When the number of available nodes in the second type of nodes does not meet the number of training nodes required for the target task to be trained, responding with a calculation operation on the remaining occupied time of the occupied nodes in the second type of nodes; Sort the remaining occupied time of all occupied nodes; Selecting the target number of occupied nodes corresponding to the remaining occupied time as target available computing nodes in ascending order, where the target number is determined by the difference between the number of nodes required for training corresponding to the target task to be trained and the number of nodes currently in available state; If the remaining occupied time of the target available computing node drops to zero, the target available computing node is converted from an occupied state to an available state; The target task to be trained is scheduled to a corresponding node resource in an available state, so as to perform a training operation on the target task to be trained.
9. The method according to claim 8, characterized in that The method further comprises: Before the remaining occupied time of all the target available computing nodes drops to zero, a trainable task is selected from the first sorting result or the second sorting result and trained using the nodes in an available state, wherein the trainable task indicates that training can be performed using the nodes currently in an available state, and the latest completion time of the training of the trainable task is less than the time corresponding to when the target remaining occupied time drops to zero, and the target remaining occupied time indicates the longest remaining occupied time among all the target available computing nodes.
10. The method according to claim 8, characterized in that The method further comprises: Collecting statistics on the queuing time of other tasks to be trained in the third queuing sub-queue except the target task to be trained; If there is a task to be trained whose queuing time is longer than the target time, the queue position of the task to be trained in the third queuing sub-queue is moved forward to a position adjacent to the position of the target task to be trained, so that when the target task to be trained completes training, the task to be trained with a queuing time longer than the target time is used as the new target task to be trained.
11. The method according to claim 10, characterized in that The method further comprises: When the queuing time is longer than the target time, there are multiple tasks to be trained, sorting the multiple tasks to be trained according to the order of priority level and the order of arrival time of tasks of the same priority level; The sorted multiple tasks to be trained are moved to a position adjacent to the position of the target task to be trained, wherein the task to be trained with the highest priority and arriving earliest among the sorted multiple tasks to be trained is adjacent to the current target task to be trained.
12. The method according to any one of claims 1 to 11, characterized in that The method further comprises: Obtaining task waiting times corresponding to multiple tasks that have completed training, where the task waiting time represents the time interval between the task arrival time and the task training start time; Based on the task waiting time, multiple preset time evaluation indicators are calculated; The evaluation operation is scheduled in response to the task based on the calculation results.
13. A task scheduling device, characterized in that: Applied to a task scheduling system, the device comprises: A first acquisition module is used to acquire a task to be trained and task description information corresponding to the task to be trained, wherein the task description information includes: a task priority level and information on resources required for training; An adding module is used to add the task to be trained to a queue, wherein the sorting order of the tasks to be trained in the queue is determined according to the task priority level and the task arrival time of the tasks to be trained. The higher the task priority level in the queue, the higher the sorting order of the tasks to be trained; and the earlier the task arrival time under the same task priority level, the higher the sorting order of the tasks to be trained; A second acquisition module is configured to acquire node resource usage information of a computing cluster used for training and determine available node resources based on the node resource usage information; A selection module is used to select the task to be trained that is ranked highest in the queue as the target task to be trained, wherein the task to be trained that is ranked highest in the queue represents the task with the highest priority level and is the task that was added to the queue the earliest; The scheduling module is used to schedule the target task to be trained to the available node resources if the available node resources meet the requirements of the training resources corresponding to the target task to be trained, so as to perform a training operation on the target task to be trained.
14. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the task scheduling method according to any one of claims 1 to 12 by executing the computer instructions.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the task scheduling method according to any one of claims 1 to 12.
16. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to enable a computer to execute the task scheduling method according to any one of claims 1 to 12.
Citation Information
Cited By
Model training resource scheduling method applied to multiple data centers and related device
CN121029421A
Preemptive distribution method and system for deep learning model training tasks
CN121579156A