Historical task data driven deep learning task scheduling simulation method and system
Patent Information
- Application Number
- CN202510109352.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The existing deep learning task scheduling algorithm has low task scheduling efficiency, the scheduling algorithm development verification environment is complex, and the implementation cost is high.
Through the deep learning task scheduling simulation method driven by historical task data, the scheduling tasks are managed using task access policies, task scheduling policies and task distribution policies, and the data generated by the simulation scheduling process is recorded through the task state database.
The efficiency of research on deep learning task scheduling methods has been improved, the demand for computing power cluster resources has been reduced, and the verification and comparison efficiency of existing scheduling algorithms has been improved.
Smart Images

Figure CN120029734A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of distributed deep learning task scheduling algorithm design, and specifically relates to a deep learning task scheduling simulation method and system driven by historical task data. Background Art
[0002] Scheduling frameworks such as Apache Mesos or Hadoop YARN handle jobs that consist of a few short-running tasks or long-running big data jobs, which run at a high priority and are therefore usually not preempted. Deep learning training tasks usually run for a long time, and their calculations are repeated for a large number of iterations. Therefore, unlike big data scheduling algorithms, deep learning task scheduling algorithms must frequently preempt running jobs to better utilize cluster computing resources.
[0003] At the same time, deep learning task schedulers usually need to access application-level metrics such as function loss values, gradient values, throughput, etc. to support specific properties of deep learning tasks, such as completion time fairness or gradient-based elasticity, which is not easy to implement in existing scheduling frameworks. Therefore, although previous deep learning task scheduling algorithms have been implemented as plug-ins on Kubernetes or YARN, these systems usually need to design additional deep learning task functions to support iteration-level preemption or application-level metric collection, resulting in low task scheduling efficiency. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the historical task data-driven deep learning task scheduling simulation method and system provided by the present invention solve the problems of low task scheduling efficiency of existing deep learning task scheduling algorithms, complex configuration of the scheduling algorithm development and verification environment, and high implementation cost.
[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a deep learning task scheduling simulation method driven by historical task data, comprising the following steps:
[0006] S1. Modify the cluster resource management table according to the newly added or deleted task execution modules through the control module, and update the deep learning task scheduling historical data;
[0007] S2. Generate a set of tasks to be scheduled based on the historical data of deep learning task scheduling through the simulation scheduling module, generate a set of newly schedulable tasks in this round through the task admission strategy, and add them to the set of tasks to be scheduled in this round;
[0008] S3. Sort the task set to be scheduled in this round by the task scheduling strategy to obtain the cluster scheduling task set;
[0009] S4, generating a runnable task sequence and a task set to be suspended according to the cluster scheduling task set through the task distribution strategy;
[0010] S5. Call the task execution module according to the executable task sequence and the task set that needs to be suspended, and update the cluster operation information table and the task operation status information table;
[0011] S6. Refresh the task status database according to the records in the updated cluster operation information table and the task operation status information table.
[0012] Further: In S1, the method for modifying the cluster resource management table is specifically:
[0013] The IP address, process ID, and number of accelerator cards of the newly added task execution module are added to the cluster resource management table, and the records of the deleted or faulty task execution modules are deleted from the cluster resource management table.
[0014] Further: In S2, the method of task admission strategy is specifically:
[0015] According to the generated set of tasks to be scheduled, the tasks in the set of tasks to be scheduled are obtained in a first-in-first-out order, and the set of newly schedulable tasks that can be added in this round is generated under the constraints. The specific expression of the constraints is:
[0016]
[0017] In the formula, f(N jk ) is the N newly added scheduling tasks in this round jk The number of accelerator cards that need to be occupied, n is the number of newly scheduled tasks that can be added in this round, f(R jk ) is the running task R jk The number of accelerator cards that need to be occupied, m is the number of tasks to be scheduled, TotalAccNum is the number of accelerator cards available in the entire cluster, LoadThreshold is the maximum load factor of the cluster, LoadThreshold∈(1.0,2.0].
[0018] Further: In S3, the method of task scheduling strategy is specifically:
[0019] Calculate the first index and the second index of the scheduling task, sort the set of tasks to be scheduled in this round according to the values of the first index and the second index, and use the sorted set of scheduling tasks as the cluster scheduling task set;
[0020] Among them, the first indicator jpr(S ji ) is specifically for task S ji The running priority of
[0021] The expression of the second indicator is specifically:
[0022] jei(S ji )*jit(S ji )*jad(S ji )
[0023] In the formula, jei(S ji ) is task S ji The number of executed iterations, jit(S ji ) is task S ji The single-round iteration running time, jad(S ji ) is task S ji The number of acceleration cards required.
[0024] Further: in said S4, the method of task distribution strategy includes compact task distribution strategy and decentralized task distribution strategy;
[0025] The compact task allocation strategy specifically allocates accelerator card computing resources to the scheduled tasks to be run according to the natural order of the idle accelerator card numbers;
[0026] The decentralized task allocation strategy specifically allocates accelerator card computing resources to the scheduled tasks to be run according to the load balancing principle.
[0027] Further: in said S5, the cluster operation information table includes the execution node identification, IP address, number of accelerator cards and accelerator card information;
[0028] Accelerator card information includes accelerator card ID, local accelerator card ID, execution node ID, memory capacity, usage, and running task list;
[0029] The task running status information table includes the timestamp of this round of scheduling, the list of tasks to be run, the statistics of running task time, the list of completed tasks and the list of suspended tasks.
[0030] The deep learning task scheduling simulation system driven by historical task data includes:
[0031] The simulation scheduling module is used to generate tasks to be scheduled based on the historical data of deep learning task scheduling, trigger task admission strategy, task scheduling strategy and task distribution strategy, and obtain cluster operation information and task operation status information;
[0032] The control module is used to manage the cluster's resource information, including CPU, memory, disk, accelerator card, and network. It is also used to manage the running status information of the scheduled tasks, including the start, stop, preemption, and waiting states, and to distribute the tasks to be scheduled to the responsible task execution module.
[0033] The task execution module is used to trigger the running of tasks, provide management of the accelerator card usage status, and record the task ID, duration, accelerator card utilization, and memory consumption of the accelerator card.
[0034] The task status database is used to store the status information of suspended or started tasks when the system is running.
[0035] The beneficial effects of the present invention are:
[0036] (1) The present invention provides a deep learning task scheduling simulation method and system driven by historical task data. The present invention uses historical task scheduling data to drive the operation of the entire system, uses task admission strategy, task scheduling strategy and task distribution strategy to manage scheduling tasks, and uses a task status database to record the data generated by the new simulation scheduling process, thereby constructing a complete and universal deep learning task scheduling method, so that the research on deep learning task scheduling methods no longer needs to occupy a large amount of computing power cluster resources, thereby improving the efficiency of verification and comparison of existing scheduling algorithms.
[0037] (2) Based on a large amount of deep learning task scheduling historical data, the present invention drives the simulation operation of the deep learning task scheduling algorithm, and provides a low-cost and high-efficiency task scheduling algorithm verification and comparison system, which helps to create new deep learning task scheduling algorithms and improve the development and testing efficiency of existing scheduling algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a flowchart of the historical task data-driven deep learning task scheduling simulation method of the present invention.
[0039] Figure 2 This is a schematic diagram of the structure of the historical task data-driven deep learning task scheduling simulation system of the present invention. DETAILED DESCRIPTION
[0040] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0041] like Figure 1 As shown, in one embodiment of the present invention, a deep learning task scheduling simulation method driven by historical task data includes the following steps:
[0042] S1. Modify the cluster resource management table according to the newly added or deleted task execution modules through the control module, and update the deep learning task scheduling historical data;
[0043] S2. Generate a set of tasks to be scheduled based on the historical data of deep learning task scheduling through the simulation scheduling module, generate a set of newly schedulable tasks in this round through the task admission strategy, and add them to the set of tasks to be scheduled in this round;
[0044] S3. Sort the task set to be scheduled in this round by the task scheduling strategy to obtain the cluster scheduling task set;
[0045] S4, generating a runnable task sequence and a task set to be suspended according to the cluster scheduling task set through the task distribution strategy;
[0046] S5. Call the task execution module according to the executable task sequence and the task set that needs to be suspended, and update the cluster operation information table and the task operation status information table;
[0047] S6. Refresh the task status database according to the records in the updated cluster operation information table and the task operation status information table.
[0048] In S1, the method for modifying the cluster resource management table is specifically as follows:
[0049] The IP address, process ID, and number of accelerator cards of the newly added task execution module are added to the cluster resource management table, and the records of the deleted or faulty task execution modules are deleted from the cluster resource management table.
[0050] In S2, the method of task admission strategy is specifically as follows:
[0051] According to the generated set of tasks to be scheduled, the tasks in the set of tasks to be scheduled are obtained in a first-in-first-out order, and the set of newly schedulable tasks that can be added in this round is generated under the constraints. The specific expression of the constraints is:
[0052]
[0053] In the formula, f(N jk ) is the N newly added scheduling tasks in this round jk The number of accelerator cards that need to be occupied, n is the number of newly scheduled tasks that can be added in this round, f(R jk ) is the running task R jk The number of accelerator cards that need to be occupied, m is the number of tasks to be scheduled, TotalAccNum is the number of accelerator cards available in the entire cluster, LoadThreshold is the maximum load factor of the cluster, LoadThreshold∈(1.0,2.0].
[0054] In this embodiment, the generated set of tasks to be scheduled is defined as {S j1 ,S j2 ...S jm}, the set of all running tasks is defined as {R j1 ,R j2 ...R jn}, the task admission strategy obtains tasks from the set of tasks to be scheduled in a first-in-first-out order, generating a set of newly schedulable tasks for this round, defined as {N j1 ,N j2 ...N jn}.
[0055] In S3, the method of task scheduling strategy is specifically as follows:
[0056] Calculate the first index and the second index of the scheduling task, sort the set of tasks to be scheduled in this round according to the values of the first index and the second index, and use the sorted set of scheduling tasks as the cluster scheduling task set;
[0057] Among them, the first indicator jpr(S ji ) is specifically for task S ji The running priority of
[0058] The expression of the second indicator is specifically:
[0059] jei(S ji )*jit(S ji )*jad(S ji )
[0060] In the formula, jei(S ji ) is task S ji The number of executed iterations, jit(S ji ) is task S ji The single-round iteration running time, jad(S ji ) is task S ji The number of acceleration cards required.
[0061] In this embodiment, all tasks to be scheduled in this round are defined as the task set to be scheduled in this round {S j1 ,S j2 ...S jn}, sort the tasks in the task set to be scheduled in this round as the actual order of cluster scheduling tasks, and obtain the cluster scheduling task set, which is defined as {S jl ,S jn ...S jm}.
[0062] In said S4, the method of task distribution strategy includes compact task distribution strategy and decentralized task distribution strategy;
[0063] The compact task allocation strategy specifically allocates accelerator card computing resources to the scheduled tasks to be run according to the natural order of the idle accelerator card numbers;
[0064] The decentralized task allocation strategy specifically allocates accelerator card computing resources to the scheduled tasks to be run according to the load balancing principle.
[0065] In this embodiment, for all scheduled task sequences that are not in the running state, it is defined as {J 1 ,J 2 ...J n}, the set of all task execution modules including idle accelerator cards is defined as {Executor 0 ,Executor 1 ...Executor m}, the accelerator card collection of a single task execution module is defined as {Accelerator 0 ,Accelerator 1 ...Accelerator 7}, according to the actual task allocation strategy in the simulated scheduling algorithm, these tasks are allocated to the accelerator card of the task execution module. The task distribution strategy adopted by the present invention includes a compact task allocation strategy and a dispersed task allocation strategy. The compact task allocation strategy is to allocate the accelerator card computing resources to the tasks to be run according to the natural order of the idle accelerator card numbers. The dispersed task allocation strategy is to allocate the accelerator card computing resources to the tasks to be run according to the load balancing principle. Since some task resources may not be allocated to the accelerator card resources, after the allocation strategy is completed, an executable task sequence {Active j1 ,Active j2 ...Active jn} and the set of tasks that need to be suspended {Suspend j1 ,Suspend j2 ...Suspend jn}.
[0066] In S5, the cluster operation information table includes the execution node identifier, IP address, number of accelerator cards and accelerator card information;
[0067] Accelerator card information includes accelerator card ID, local accelerator card ID, execution node ID, memory capacity, usage, and running task list;
[0068] The task running status information table includes the timestamp of this round of scheduling, the list of tasks to be run, the statistics of running task time, the list of completed tasks and the list of suspended tasks.
[0069] In this embodiment, the control module calls the task execution submodule of the task execution module to obtain the executable task sequence {Active j1 ,Active j2 ...Active jn} and the set of tasks that need to be suspended {Suspend j1 ,Suspend j2 ...Suspend jn}, call the task execution module to update the cluster operation information table and the task operation status information table. The information that needs to be updated is shown in the following table.
[0070] Table 1 Cluster operation information table
[0071] Node Id Execution node identification IP Addr IP address Accelerator Num Number of accelerator cards Accelerator Infos Accelerator card information
[0072] Table 2 Accelerator card information table
[0073] Accelerator Id Accelerator card identification Local Accelerator Id Local accelerator card ID Node Id Execution node identification Memory Memory capacity In Use Is it in use? Job Ids List of tasks to run
[0074] Table 3 Task running status information table
[0075] Schedule Timestamp Timestamp of this round of scheduling Active Job Ids List of tasks to be run Job Runtime Stats Running task time statistics Finished Job Ids Completed Task List Suspend Job Ids Suspended Tasks List
[0076] like Figure 2 As shown in the figure, the deep learning task scheduling simulation system driven by historical task data includes:
[0077] The simulation scheduling module is used to generate tasks to be scheduled based on the historical data of deep learning task scheduling, trigger task admission strategy, task scheduling strategy and task distribution strategy, and obtain cluster operation information and task operation status information;
[0078] The control module is used to manage the cluster's resource information, including CPU, memory, disk, accelerator card, and network. It is also used to manage the running status information of the scheduled tasks, including the start, stop, preemption, and waiting states, and to distribute the tasks to be scheduled to the responsible task execution module.
[0079] The task execution module is used to trigger the running of tasks, provide management of the accelerator card usage status, and record the task ID, duration, accelerator card utilization, and memory consumption of the accelerator card.
[0080] The task status database is used to store the status information of suspended or started tasks when the system is running.
[0081] In this embodiment, the deep learning task scheduling history data includes the task history record, the CPU, memory and acceleration card usage of the cluster, wherein the task history record includes recording the task ID, the task scheduling time, start time, end time, the acceleration card used and the user name to which the task belongs.
[0082] The beneficial effects of the present invention are as follows: the present invention provides a deep learning task scheduling simulation method and system driven by historical task data. The present invention uses historical task scheduling data to drive the operation of the entire system, manages scheduling tasks using task access strategies, task scheduling strategies and task distribution strategies, and uses a task status database to record the data generated by the new simulation scheduling process, thereby constructing a complete and universal deep learning task scheduling method, so that the research on deep learning task scheduling methods no longer needs to occupy a large amount of computing power cluster resources, thereby improving the efficiency of verification and comparison of existing scheduling algorithms.
[0083] Based on a large amount of deep learning task scheduling historical data, the present invention drives the simulated operation of the deep learning task scheduling algorithm, and provides a low-cost and efficient task scheduling algorithm verification and comparison system, which helps to create new deep learning task scheduling algorithms and improve the development and testing efficiency of existing scheduling algorithms.
[0084] In the description of the present invention, it is necessary to understand that the orientation or positional relationship indicated by the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", and "third" are used only for descriptive purposes, and cannot be understood as indicating or implying the relative importance or the number of implicitly specified technical features. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of the features.
Claims
1. A deep learning task scheduling simulation method driven by historical task data, characterized in that: The following steps are involved: S1. Modify the cluster resource management table according to the newly added or deleted task execution modules through the control module, and update the deep learning task scheduling historical data; S2. Generate a set of tasks to be scheduled based on the historical data of deep learning task scheduling through the simulation scheduling module, generate a set of newly schedulable tasks in this round through the task admission strategy, and add them to the set of tasks to be scheduled in this round; S3. Sort the task set to be scheduled in this round by the task scheduling strategy to obtain the cluster scheduling task set; S4, generating a runnable task sequence and a task set to be suspended according to the cluster scheduling task set through the task distribution strategy; S5. Call the task execution module according to the executable task sequence and the task set that needs to be suspended, and update the cluster operation information table and the task operation status information table; S6. Refresh the task status database according to the records in the updated cluster operation information table and the task operation status information table.
2. The historical task data-driven deep learning task scheduling simulation method according to claim 1 is characterized in that: In S1, the method for modifying the cluster resource management table is specifically as follows: The IP address, process ID, and number of accelerator cards of the newly added task execution module are added to the cluster resource management table, and the records of the deleted or faulty task execution modules are deleted from the cluster resource management table.
3. The historical task data-driven deep learning task scheduling simulation method according to claim 1 is characterized in that: In S2, the method of task admission strategy is specifically as follows: According to the generated set of tasks to be scheduled, the tasks in the set of tasks to be scheduled are obtained in a first-in-first-out order, and the set of newly schedulable tasks that can be added in this round is generated under the constraints. The specific expression of the constraints is: In the formula, f(N jk ) is the N newly added scheduling tasks in this round jk The number of accelerator cards that need to be occupied, n is the number of newly scheduled tasks that can be added in this round, f(R jk ) is the running task R jk The number of accelerator cards that need to be occupied, m is the number of tasks to be scheduled, TotalAccNum is the number of accelerator cards available in the entire cluster, LoadThreshold is the maximum load factor of the cluster, LoadThreshold∈(1.0,2.0].
4. The historical task data-driven deep learning task scheduling simulation method according to claim 3 is characterized in that: In S3, the method of task scheduling strategy is specifically as follows: Calculate the first index and the second index of the scheduling task, sort the set of tasks to be scheduled in this round according to the values of the first index and the second index, and use the sorted set of scheduling tasks as the cluster scheduling task set; Among them, the first indicator jpr(S ji ) Specifically, task S ji The running priority of The expression of the second indicator is specifically: jei(S ji )*jit(S ji )*jad(S ji ) In the formula, jei(S ji ) is task S ji The number of executed iterations, jit(S ji ) is task S ji The single-round iteration running time, jad(S ji ) is task S ji The number of acceleration cards required.
5. The historical task data-driven deep learning task scheduling simulation method according to claim 4 is characterized in that: In said S4, the method of task distribution strategy includes compact task distribution strategy and decentralized task distribution strategy; The compact task allocation strategy specifically allocates accelerator card computing resources to the scheduled tasks to be run according to the natural order of the idle accelerator card numbers; The decentralized task allocation strategy specifically allocates accelerator card computing resources to the scheduled tasks to be run according to the load balancing principle.
6. The historical task data-driven deep learning task scheduling simulation method according to claim 5 is characterized in that: In S5, the cluster operation information table includes the execution node identifier, IP address, number of accelerator cards and accelerator card information; Accelerator card information includes accelerator card ID, local accelerator card ID, execution node ID, memory capacity, usage, and running task list; The task running status information table includes the timestamp of this round of scheduling, the list of tasks to be run, the statistics of running task time, the list of completed tasks and the list of suspended tasks.
7. A deep learning task scheduling simulation system driven by historical task data, applied to a deep learning task scheduling simulation method driven by historical task data as claimed in any one of claims 1 to 6, characterized in that: The system includes: The simulation scheduling module is used to generate tasks to be scheduled based on the historical data of deep learning task scheduling, trigger task admission strategy, task scheduling strategy and task distribution strategy, and obtain cluster operation information and task operation status information; The control module is used to manage the cluster's resource information, including CPU, memory, disk, accelerator card, and network. It is also used to manage the running status information of the scheduled tasks, including the start, stop, preemption, and waiting states. It is also used to distribute the tasks to be scheduled to the responsible task execution module. The task execution module is used to trigger the running of tasks, provide management of the accelerator card usage status, and record the task ID, duration, accelerator card utilization, and memory consumption of the accelerator card. The task status database is used to store the status information of suspended or started tasks when the system is running.
Citation Information
Patent Citations
Resource allocation method and device
CN112181643A
Container cluster resource scheduling method and system based on deep reinforcement learning
CN114443249A
Multi-target dynamic task scheduling method and system based on improved ant colony algorithm
CN114968510A
Deep learning task scheduling method for multi-user GPU cluster
CN118093208A
AI accelerator card resource scheduling method based on double optimization models
CN118426971A