Tail delay optimization job scheduling method and system based on heterogeneous GPU cluster
Through the tail delay optimization job scheduling method based on heterogeneous GPU cluster, we dynamically adapt to hybrid loads, combined with local autonomy and global load balancing mechanisms, the tail delay and throughput loss of data centers in a hybrid load environment is solved, and the resource utilization and service quality are improved.
Patent Information
- Application Number
- CN202510779738.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing technology is difficult to adapt dynamically under hybrid workloads, taking into account tail latency and throughput optimization, and has a scheduling mechanism that provides real-time resource dynamic adjustment capabilities, resulting in data centers frequently breaking service-level targets in hybrid load environments, and serious throughput losses.
The tail delay optimization job scheduling method based on heterogeneous GPU cluster is adopted, the initial scheduling sequence is obtained through the node manager, the task set is constructed and sorted for the heavy-tail load, the best job scheduling sequence is generated, and the load rebalancing is achieved through dynamic task migration across the scheduler, combining local autonomy and global load balancing mechanisms to optimize the job completion time.
Effectively reduce tail latency under mixed loads, improve system resource utilization and throughput, realize cluster-level load balancing, ensure service quality, and maintain efficient resource dynamic adjustment capabilities.
Smart Images

Figure CN120295739A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a tail latency optimization job scheduling method and system based on a heterogeneous GPU cluster. Background Art
[0002] With the rapid development of cloud computing and big data technologies, data centers have become the core infrastructure to support high-concurrency real-time applications such as web search, e-commerce, and social networks. Such applications have stringent microsecond-level tail latency requirements for Quality of Service (QoS), for example, the latency upper limit of 99.9% of requests in the Service Level Objective (SLO) needs to be met. However, the resource scheduling in data centers faces multiple technical bottlenecks: 1. Dynamic imbalance between resource supply and demand: The burstiness and dynamicity of user requests lead to significant fluctuations in resource demand, with short-term resource overload and idle phenomena alternating; 2. Tail effect of distributed architecture: Modern applications are based on a modular microservices architecture, and end-to-end requests need to be processed collaboratively across multiple distributed servers. Their overall response time is limited by the tail latency of the node with the worst performance, forming the "barrel effect"; 3. Challenges of heterogeneous load mixing: Light-tailed tasks (such as real-time inference) and heavy-tailed tasks (such as distributed training) coexist in real scenarios, and a single scheduling strategy is difficult to adapt to the characteristics of mixed loads.
[0003] Current mainstream scheduling algorithms are designed based on classical queuing theory, such as the Shortest Job First (SJF), First-Come-First-Served (FCFS), and Processor Sharing (PS) strategies. Previous studies have shown that SJF can minimize the average job completion time (JCT) in light-tailed load scenarios, while PS achieves fairness in heavy-tailed loads through a time-slicing mechanism. However, in actual production environments, mixed loads are prevalent, and a single strategy cannot balance the low-latency requirements of light-tailed tasks and the resource isolation requirements of heavy-tailed tasks, resulting in frequent breakthroughs of the SLO threshold for tail latency (such as P99 latency). More seriously, traditional optimization methods (such as single-queue priority scheduling or preemptive scheduling) can locally improve tail latency, but they will sacrifice up to 30% of the system throughput, and as the task granularity is refined (such as containerized microservices), the throughput loss increases exponentially.
[0004] In addition, the resource dynamics in the data center further exacerbate the scheduling complexity: 1. Insufficient resource elasticity: Medium-sized data centers are limited by the scale of physical resources, and sudden requests are likely to cause cluster-level overload. Prediction models based on historical data (such as ARIMA and LSTM) fail to capture non-periodic demand changes (such as festival traffic peaks and black swan events), resulting in the failure of resource reservation strategies; 2. Amplification of tail fluctuations by local load hotspots: Factors such as network partitioning, firmware compatibility errors, or uneven task allocation may trigger high-load windows for local resources (such as GPU video memory contention and NVLink bandwidth saturation), causing the amplitude of tail latency fluctuations to increase by 2-5 times; 3. The trade-off dilemma between quality of service and throughput: Existing systems usually adopt service degradation strategies (such as restricting resource quotas for low-priority tasks and discarding timeout requests) to control tail latency, but this directly leads to a decline in user experience and revenue loss (such as a 1% increase in request latency in an e-commerce scenario may cause a 5% loss of orders).
[0005] Therefore, there is a lack of a scheduling mechanism in the existing technology that can dynamically adapt to the characteristics of mixed workloads, balance the optimization of tail latency and throughput, and have the ability to dynamically adjust resources in real time. Summary of the Invention
[0006] Based on the technical problems existing in the background technology, the present invention proposes a tail latency optimization job scheduling method and system based on a heterogeneous GPU cluster, which dynamically selects the optimal scheduling strategy under mixed workloads, thereby improving the system resource utilization and generality while ensuring strict tail latency service level objectives.
[0007] The tail latency optimization job scheduling method based on a heterogeneous GPU cluster proposed by the present invention includes: The node manager obtains the initial scheduling sequences generated by each scheduler, and the initial scheduling sequences are sequences obtained by sorting tasks using the short job first strategy; For the target scheduler with heavy-tailed load, the corresponding initial scheduling sequence is sorted in ascending order to obtain an ordered sequence, and a heavy-tailed load task set is constructed based on this; Generate all possible job permutations according to the heavy-tailed load task set, and take the minimum job completion time of the scheduler with the maximum job completion time as the goal to obtain the optimal job scheduling sequence; Allocate jobs to each scheduler according to the optimal job scheduling sequence for distributed job scheduling.
[0008] Furthermore, the construction formula of the heavy-tailed load task set is as follows: ; where, For the assignments, is the target scheduler, For the top The completion time of the job, To obtain an ordered sequence by arranging all the jobs to be scheduled on the target scheduler in ascending order of completion time, is the total number of all jobs to be scheduled on the target scheduler, is the weight parameter, is the upper limit of the number of short jobs that are filtered out. Indicates a round-down operation. Indicates to extract the first The tasks corresponding to the completion time.
[0009] Furthermore, the generation process of the optimal job scheduling sequence is as follows: Generate all possible job permutations using a heavy tail load task set ; For each arrangement , based on the job completion time of each scheduler, the maximum job completion time is obtained; Find the order that minimizes the job completion time of the scheduler that maximizes the job completion time, assign the order as the optimal job scheduling sequence to each scheduler, and return the scheduling results to each scheduler.
[0010] Furthermore, the optimal job scheduling sequence The generation formula is as follows: ; in, is the job completion time of the scheduler with the largest job completion time, is the sorting method corresponding to the optimal job scheduling sequence, To obtain the maximum value of the job completion time of each scheduler, is the total number of schedulers in the heterogeneous GPU cluster, The index of the scheduler.
[0011] Furthermore, each computing node deploys an independent node manager, which is responsible for monitoring the local resource status and performing initial task scheduling; Multiple schedulers form a scheduler cluster. The scheduler cluster adopts a decentralized design. Each scheduler runs independently and manages the job sequence based on the preset short job priority strategy.
[0012] Tail latency optimization job scheduling system based on heterogeneous GPU clusters, including scheduler, node manager and heterogeneous GPU clusters; Multiple schedulers form a distributed scheduler cluster. Each scheduler runs independently and sorts tasks based on the preset short job first strategy to obtain an initial scheduling sequence; The node manager obtains the initial scheduling sequences generated by each scheduler. For the target scheduler with heavy-tailed load, it sorts the corresponding initial scheduling sequence in ascending order to obtain an ordered sequence, and constructs a heavy-tailed load task set based on this; it generates all possible job permutations according to the heavy-tailed load task set, and aims to minimize the job completion time of the scheduler with the maximum job completion time, so as to obtain the optimal job scheduling sequence; Allocate jobs to each scheduler according to the optimal job scheduling sequence for distributed job scheduling.
[0013] Furthermore, the construction formula of the heavy-tailed load task set is as follows: ; where, is the th job, is the target scheduler, is the job completion time of the previous , is the ordered sequence obtained by sorting all the jobs to be scheduled on the target scheduler in ascending order of completion time, is the total number of all jobs to be scheduled on the target scheduler, is the weight parameter, is the upper limit of the number of short jobs filtered out, represents intercepting the jobs corresponding to the first completion times from the ordered sequence.
[0014] Furthermore, the process of the node manager generating the optimal job scheduling sequence is as follows: Generate all possible job permutations using the heavy-tailed load task set ; For each permutation , obtain the maximum job completion time based on the job completion times of each scheduler; Find the sorting corresponding to the minimum job completion time of the scheduler with the maximum job completion time, and use this sorting as the optimal job scheduling sequence to be allocated to each scheduler, and return the scheduling result to each scheduler.
[0015] Furthermore, the generation formula of the optimal job scheduling sequence is as follows: ; where, is the job completion time of the scheduler with the maximum job completion time, is the sorting method corresponding to the optimal job scheduling sequence, is to take the maximum value of the job completion times of each scheduler, is the total number of schedulers in the heterogeneous GPU cluster, is the index of the scheduler.
[0016] Each computing node deploys an independent node manager, which is responsible for monitoring the local resource status and performing initial task scheduling.
[0017] The advantages of the tail latency optimization job scheduling method and system based on a heterogeneous GPU cluster provided by the present invention are as follows: It can dynamically adapt to the characteristics of mixed workloads, balance tail latency and throughput optimization, and has a scheduling mechanism with real-time resource dynamic adjustment capabilities; for the local scheduler overload problem caused by heavy-tailed loads, the node manager monitors the status of each scheduler queue in real time and realizes load rebalancing through cross-scheduler dynamic task migration. While ensuring local scheduling autonomy, through lightweight global intervention, it effectively eliminates the tail latency bottleneck caused by heavy-tailed loads and finally achieves cluster-level load balancing; by introducing a two-level scheduling mechanism: (b1) local autonomous scheduling and global load rebalancing (step two); (b2) the system only performs cross-node migration on low-resource-occupancy tasks (i.e., small tasks) in the affected schedulers when detecting JCT anomalies caused by heavy-tailed loads (steps three and four). Description of the Drawings
[0018] Figure 1 is the flowchart of the present invention; Figure 2 is the schematic diagram of the tail latency optimization job scheduling scheme. a) is the schematic diagram of the centralized node manager, and b) is the schematic diagram of the distributed scheduler cluster; Figure 3 is the schematic diagram of three job scheduling optimizations. a) is the traditional short job optimization subgraph, b) is the traditional deadline optimization subgraph, and c) is the tail latency optimization subgraph of this embodiment; Figure 4 is the schematic diagram of the scheduling logic of the job scheduling sequence. Detailed Embodiments
[0019] Next, the technical solutions of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0020] Such as Figures 1 to 4As shown in the figure, the tail latency optimization job scheduling method based on a heterogeneous GPU cluster proposed by the present invention includes the following steps: Step 1: The node manager obtains the initial scheduling sequences generated by each scheduler, and the initial scheduling sequences are sequences obtained by sorting tasks using the short job first strategy; Step 2: For the target scheduler with heavy-tailed load, the corresponding initial scheduling sequence is sorted in ascending order to obtain an ordered sequence, and based on this, a heavy-tailed load task set is constructed; Step 3: Generate all possible job permutation methods according to the heavy-tailed load task set, and aim to minimize the job completion time of the scheduler with the maximum job completion time to obtain the optimal job scheduling sequence; Step 4: Allocate according to the optimal job scheduling sequence to each scheduler for distributed job scheduling.
[0021] In this embodiment, by optimizing both the job completion time and the job deadline simultaneously, the optimal scheduling strategy is dynamically selected under mixed workloads, aiming to reduce the tail job completion time and the average job completion time in a high-utilization data center by reducing the variance of the total waiting time between job tasks.
[0022] The purpose of the tail latency task local optimization strategy designed in this embodiment is to (1) perform centralized job management on the tail latency jobs in the initial job scheduling sequence obtained by the short job first strategy, which has been proven to have the best effect under light-tailed workloads, to obtain the optimal job scheduling sequence after optimizing the tail latency task scheduling, , is the scheduling sequence of the th job on the th available scheduler, , , and are the total numbers of jobs and available schedulers respectively. (2) Then allocate to each scheduler for distributed job scheduling to achieve the purpose of optimizing the job completion time (JCT) while maintaining the requirement of meeting the job deadline.
[0023] In one embodiment, step 1 is specifically: The tail-latency task scheduling of this embodiment consists of multi-level collaborative components, specifically including a distributed scheduler, a node manager, and a heterogeneous GPU cluster. An independent node manager is deployed on each computing node, responsible for coordinating local resources and executing task scheduling. When a job is submitted to the system, the main controller allocates it to a specific scheduler according to the job type and resource requirements. Each scheduler uses the Shortest Job First (SJF) policy locally for task orchestration to obtain the initial job scheduling sequence in the scheduler. , as Figure 4 shown in (a) of
[0024] to maximize the scheduling efficiency in the light-tail load scenario. In one embodiment, the node manager further performs global optimization on the distributed scheduling result (i.e., the initial job scheduling sequence ), that is, step two is specifically as follows: For the problem of local scheduler overload caused by heavy-tail load (manifested as a significant increase in job completion time or job deadline timeout), the node manager monitors the queue status of each scheduler in real time, identifies small tasks that occupy less resources , and realizes load rebalancing through cross-scheduler dynamic task migration, specifically as follows: Suppose there are jobs (such as deep learning jobs) that form a set , and available scheduler resources in the heterogeneous GPU cluster form a set . is the th job, is the th available scheduler. For the initial scheduling sequence generated in the scheduler, for the target scheduler with heavy-tail load, a heavy-tail load task set
[0025] needs to be screened out. It should be noted that in this embodiment, the scheduler is the GPU. The scheduler is a functional name, and the GPU is a physical device name. Therefore, the GPU resources in the heterogeneous GPU cluster refer to the scheduler resources.
[0026] Among them, suppose there are pending jobs in the target scheduler , and the job completion time set is . is the job completion time of the th pending job. After sorting the jobs in ascending order of completion time, an ordered sequence is obtained, where , , and are the job completion times at the position, position, position respectively.
[0027] Define the heavy-tailed load task set as: ; (1) where: represents the th job on the target scheduler , is the completion time of job , is the ordered sequence obtained by arranging all the jobs to be scheduled on the target scheduler in ascending order of completion time, is the total number of all jobs to be scheduled on the target scheduler, is the weight parameter, and its default value is set to 50%, that is , indicating to screen the first proportion of jobs with the shortest completion times as small tasks that occupy less resources, is the upper limit of the number of short jobs screened out, represents intercepting the jobs corresponding to the first completion times from the ordered sequence, represents the floor operation to ensure that the number of selected jobs is an integer.
[0028] For the problem of local scheduler overload caused by heavy-tailed load, the node manager monitors the status of each scheduler queue in real time and realizes load rebalancing through dynamic task migration across schedulers. While ensuring local scheduling autonomy, through lightweight global intervention, it effectively eliminates the tail delay bottleneck caused by heavy-tailed load and finally achieves cluster-level load balancing.
[0029] In one embodiment, the total job completion time of the job tasks in the same batch depends on the scheduler with the largest job completion time. Therefore, the optimization goal is to make the job completion time of the scheduler with the largest job completion time as small as possible. That is, step three is specifically: (a1) Generate all possible job permutations using the heavy-tailed load task set obtained by formula (1); (a2) For each permutation , add up the required completion times of the jobs within the th scheduler to obtain the job completion time of scheduler , find the maximum job completion time: ; (a3) Find the sorting corresponding to the minimum job completion time of the scheduler with the maximum job completion time , where is the sorting method corresponding to the optimal job scheduling sequence; (a4) Assign this sorting as the optimal job scheduling sequence to each scheduler, and return the scheduling result to each scheduler.
[0030] Based on (a1) to (a4), the job scheduling sequence obtained by the tail latency task optimization scheduling algorithm is shown in formula (2): ; where is the job completion time of the scheduler with the maximum job completion time, is to take the maximum value of the job completion times of each scheduler, is the total number of schedulers in the heterogeneous GPU cluster, is the index of the scheduler.
[0031] It should be noted that when calculating the optimal job scheduling sequence, the tasks will be re-planned to the schedulers with less task density. Therefore, after obtaining the optimal job scheduling sequence, the jobs can be directly assigned to each scheduler for distributed job scheduling, and the assigned schedulers are idle or schedulers with less task density.
[0032] The optimization goal of this embodiment focuses on the tail latency tasks generated in the distributed scheduling process, that is, the small tasks accumulated in the local scheduler due to the heavy-tailed workload characteristics. By introducing a two-level scheduling mechanism: (b1) local autonomous scheduling and global load rebalancing (step two); (b2) The system only performs cross-node migration on the low-resource occupancy tasks (i.e., small tasks) in the affected schedulers when detecting JCT anomalies caused by heavy-tailed loads (steps three and four).
[0033] In this embodiment, through a decentralized scheduling method, the work of coordinating multiple schedulers is carried out, and the overall completion time is less. However, the amount of tasks remains unchanged, so the waiting time of tail tasks is reduced, thereby improving the overall efficiency of the job. In addition, this embodiment will evaluate the effectiveness of the evaluation queue reordering technique (step two) in optimizing the completion time of tail tasks. This method can ensure that even small tasks can obtain timely resource allocation and reduce waiting time. At the same time, the research will explore how to optimize scheduling decisions by combining the characteristics of jobs and tasks to achieve more accurate and efficient task scheduling. The research will also consider the scalability and adaptability of scheduling strategies to ensure efficient scheduling under different workloads and resource configurations, which is particularly important for processing large-scale and dynamically changing jobs (such as deep learning jobs).
[0034] As an embodiment; Table 1 is an example of a heterogeneous GPU cluster consisting of 4 jobs and 2 different types of GPUs. The execution time and deadline of jobs (job1, job2, job3, job4) on schedulers (i.e., GPUs: 1080Ti, V100); Table 1
[0035] It is assumed that the deadline of the job is specified when the job arrives (as shown in Table 1). For details, see Figure 3 In a), b), and c), the vertical dotted lines are the deadlines of each job.
[0036] As Figure 3 a) in shows the scheduling result under the shortest job first scheduling algorithm. The shortest job first scheduling algorithm is a strategy that preferentially selects jobs with short processing times for scheduling to reduce the overall waiting time. Since the scheduling strategy of the shortest job first scheduling algorithm only relates to the parameter job completion time and completely ignores the deadline of the job, job1 still times out even though it has notified the scheduler of the deadline. Therefore, it can be seen from Figure a) that the job completion time of the shortest job first algorithm will have the situation of job timeout.
[0037] Figure 3 b) in shows the scheduling result under the deadline first scheduling algorithm. The deadline first scheduling algorithm is a method of sorting according to the job deadline and preferentially processing jobs with early deadlines to ensure the timeliness of tasks. It can be seen that, due to only considering the single parameter deadline first and not taking JCT as the optimization goal, the job completion time of the deadline first scheduling algorithm is 15, which is 1.5 times that of the tail latency optimization scheduling algorithm proposed in this application. It causes problems of energy consumption (locally) or operating costs (in the cloud) in actual production. It can be seen from Figure b) that there is a problem of low job completion time efficiency caused by heavy-tailed load in the deadline first algorithm.
[0038] Figure 3 In part (c), it is the local optimization of the tail delay task designed in this embodiment. Since the heavy-tailed load task is the most important reason for deadline timeout, this embodiment has two optimization goals: deadline and job completion time. If it is detected that there will be a timeout according to the initial scheduling sequence of the scheduler, it will be determined as a heavy-tailed task and the scheduling policy will be updated in the node manager, and its optimization effect has been significantly improved.
[0039] As another embodiment; The scheduling policy of this embodiment is implemented based on a multi-level collaborative architecture, combining local autonomous scheduling and global dynamic load balancing mechanisms to optimize the tail delay of jobs (such as deep learning jobs). The specific implementation is as follows: (c1) System architecture and component deployment; The system consists of a distributed scheduler cluster, node managers, and heterogeneous computing resources (such as GPU clusters). Each computing node deploys an independent node manager, which is responsible for monitoring the local resource status and performing initial task scheduling. The scheduler cluster adopts a decentralized design, and each scheduler instance runs independently, managing the job queue based on the short job first policy. The main controller allocates newly submitted jobs to specific schedulers using the short job first policy according to the job type and resource requirements to ensure the initial load diversion.
[0040] (c2) Local autonomous scheduling policy; At the local scheduling layer, each scheduler arranges job tasks using the short job first policy. For light-tailed load tasks (such as real-time inference), tasks with shorter execution times are preferentially scheduled to minimize the average job completion time (JCT). The scheduler continuously maintains the resource status information of the local node (such as GPU video memory, bandwidth utilization), and dynamically updates the job sequence. When allocating tasks, the scheduler selects the optimal node for deployment based on the real-time resource availability and estimated waiting time.
[0041] (c3) Global load balancing mechanism; The node manager realizes global load optimization through cross-scheduler collaborative communication. The specific process is as follows: Status monitoring: The node manager continuously collects the queue status, resource utilization, and task execution progress of each scheduler, and identifies local resource overloads caused by heavy-tailed loads (such as a significant increase in job completion time or queue backlog).
[0042] Tail task identification: For the target scheduler with heavy-tailed load, small tasks with low resource occupancy (such as short-time inference requests) are screened and marked as tail delay optimization targets.
[0043] Dynamic task migration: Migrate the marked tasks to the scheduler nodes with idle resources, and ensure the complete migration of the task context through a lightweight data synchronization mechanism. The migration process follows the principle of minimum interference to avoid affecting the heavy-tailed tasks (such as distributed training jobs) that are being executed. The principle of minimum interference is as follows: During the migration process, ensure that the job tasks being executed are not affected. For example, the node manager considers job2 as a job to be scheduled, but job2 is already in execution. To reduce interference, it will not be scheduled and will wait until it is completed.
[0044] (c4) Integration with the existing scheduling framework; To achieve compatibility with mainstream container orchestration systems, this embodiment extends the existing scheduling framework: Multi-scheduler instantiation: Deploy multiple scheduler instances in the distributed scheduler cluster. Each instance independently manages part of the node resources. Newly submitted tasks are bound to a specific scheduler according to preset rules (such as job type or resource label).
[0045] Dynamic update of resource waiting time: The scheduler maintains a global resource view and records the estimated waiting time of each node in real time. When a task is completed or migrated, update the resource status cache of all schedulers through a collaborative communication protocol to ensure the timeliness of scheduling decisions.
[0046] Elastic queue management: The node local queue adopts a hybrid sorting strategy of priority and timestamp (first use shortest job first (priority), then detect heavy-tailed loads (when the job deadline times out), and finally use the node manager for scheduling. According to the waiting time and completion time of the job, ensure the job deadline (ddl)). When resources are released, give priority to scheduling the task with the longest waiting time, and dynamically adjust the queue length to prevent resource overload.
[0047] Through the above implementation methods, the system can ensure low latency for light-tailed tasks while effectively alleviating the tail latency bottleneck caused by heavy-tailed tasks, achieving a double improvement in cluster-level resource utilization and service quality.
[0048] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A tail latency optimization job scheduling method based on a heterogeneous GPU cluster, characterized in that Including: The node manager obtains the initial scheduling sequences generated by each scheduler, where the initial scheduling sequences are sequences obtained by sorting tasks using the shortest job first policy; For the target scheduler with heavy-tailed load, the corresponding initial scheduling sequence is sorted in ascending order to obtain an ordered sequence, based on which a heavy-tailed load task set is constructed; All possible job permutation ways are generated according to the heavy-tailed load task set, with the goal of minimizing the job completion time of the scheduler with the maximum job completion time, to obtain the optimal job scheduling sequence; Jobs are assigned to each scheduler according to the optimal job scheduling sequence for distributed job scheduling.
2. The tail latency optimization job scheduling method based on a heterogeneous GPU cluster according to claim 1, wherein, Heavy-tailed load task set The construction formula is as follows: ; Among them, is the th job, is the target scheduler, is the job completion time of the above, is the ordered sequence obtained by arranging all the jobs to be scheduled on the target scheduler in ascending order of completion time, is the total number of all the jobs to be scheduled on the target scheduler, is the weight parameter, is the upper limit of the number of short jobs filtered out, represents the floor operation, represents intercepting the jobs corresponding to the first completion times from the ordered sequence.
3. The tail latency optimization job scheduling method based on a heterogeneous GPU cluster according to claim 1, wherein, The generation process of the optimal job scheduling sequence is as follows: Generate all possible job permutations using a heavy-tailed workload task set ; For each permutation , the maximum job completion time is obtained based on the job completion times of the respective schedulers; Find the sorting corresponding to the minimum job completion time of the scheduler with the maximum job completion time, use this sorting as the optimal job scheduling sequence and assign it to each scheduler, and return the scheduling result to each scheduler.
4. The tail latency optimization job scheduling method based on a heterogeneous GPU cluster according to claim 3, wherein Optimal job scheduling sequence The generation formula is as follows: ; wherein, is the job completion time of the scheduler with the maximum job completion time, is the sorting method corresponding to the optimal job scheduling sequence, is to take the maximum value of the job completion times of each scheduler, is the total number of schedulers in the heterogeneous GPU cluster, is the index of the scheduler.
5. The tail latency optimization job scheduling method based on a heterogeneous GPU cluster according to claim 1, wherein Each computing node deploys an independent node manager, which is responsible for monitoring the local resource status and performing initial task scheduling; Multiple schedulers form a distributed scheduler cluster. The distributed scheduler cluster adopts a decentralized design, and each scheduler runs independently and manages the job sequence based on the preset shortest job first policy.
6. A tail latency optimization job scheduling system based on a heterogeneous GPU cluster, characterized in that, Including a scheduler, a node manager, and a heterogeneous GPU cluster; Multiple schedulers constitute a distributed scheduler cluster. Each scheduler runs independently and sorts tasks based on the preset shortest job first policy to obtain an initial scheduling sequence; The node manager obtains the initial scheduling sequences generated by each scheduler. For the target scheduler with heavy-tailed load, the corresponding initial scheduling sequence is sorted in ascending order to obtain an ordered sequence, based on which a heavy-tailed load task set is constructed; all possible job permutation ways are generated according to the heavy-tailed load task set, with the goal of minimizing the job completion time of the scheduler with the maximum job completion time, to obtain the optimal job scheduling sequence; Jobs are assigned to each scheduler according to the optimal job scheduling sequence for distributed job scheduling.
7. The tail latency optimization job scheduling system based on a heterogeneous GPU cluster according to claim 6, wherein Heavy-tailed load task set The construction formula is as follows: ; Among them, is the th job, is the target scheduler, is the job completion time of the above, is the ordered sequence obtained by arranging all the jobs to be scheduled on the target scheduler in ascending order of completion time, is the total number of all the jobs to be scheduled on the target scheduler, is the weight parameter, is the upper limit of the number of short jobs filtered out, means intercepting the jobs corresponding to the first completion times from the ordered sequence.
8. The tail latency optimization job scheduling system based on a heterogeneous GPU cluster according to claim 6, wherein, The process for the node manager to generate the optimal job scheduling sequence is as follows: Generate all possible job permutations using a heavy-tailed load task set ; For each permutation , obtain the maximum job completion time based on the job completion times of each scheduler; Find the sorting corresponding to the minimum job completion time of the scheduler with the maximum job completion time, use this sorting as the optimal job scheduling sequence and assign it to each scheduler, and return the scheduling result to each scheduler.
9. The tail latency optimization job scheduling system based on a heterogeneous GPU cluster according to claim 6, wherein Optimal job scheduling sequence The generation formula is as follows: ; Among them, is the job completion time of the scheduler with the maximum job completion time, is the sorting method corresponding to the optimal job scheduling sequence, is to take the maximum value of the job completion times of each scheduler, is the total number of schedulers in the heterogeneous GPU cluster, is the index of the scheduler.
10. The tail latency optimization job scheduling system based on a heterogeneous GPU cluster according to claim 6, wherein, Each computing node deploys an independent node manager, which is responsible for monitoring the local resource status and performing initial task scheduling.
Citation Information
Patent Citations
Fine-grained task scheduling method under cloud environment
CN106569887A
Method for migrating workload and rack system
CN108268321A
Dynamic resource scheduling method for GPU (Graphics Processing Unit) cluster
CN114647515A
Hybrid load priority distributed scheduling method based on global time wall
CN114968524A
Multitask scheduling method and device based on heterogeneous distributed cluster
CN118227291A