A fast job scheduling method for large-scale parallel machines
By adopting a fast job scheduling method based on the job scheduling architecture in large-scale parallel machines, the efficient scheduling problem of fixed resource scale demand operations is solved, and the job waiting time is reduced and the system resource utilization is improved.
Patent Information
- Application Number
- CN202110325147.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-03-26
AI Technical Summary
In large-scale parallel machines, operations with fixed resource scale requirements are difficult to efficiently schedule, resulting in too long wait time for jobs and low system resource utilization.
A fast job scheduling method based on the job scheduling architecture is adopted, including a job pool, a scheduling policy module and a job start module. The job scheduling process is optimized by setting job wait time thresholds, calculating job priorities, querying available resources in parallel, and resource reservation mechanisms.
It effectively reduces job waiting time, improves system resource utilization, improves user experience, and supports the rapid scheduling of massive fixed resource-scale demand operations in super-large systems.
Smart Images

Figure CN114217912B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a fast job scheduling method for large-scale parallel machines, belonging to the technical field of high-performance computing. Background Art
[0002] The job scheduler is an important part of the job management system. Its function is to select the most suitable job from the jobs waiting to be run as fairly and efficiently as possible to schedule and run, so as to meet the real-time operation requirements of user projects and improve the utilization of system resources.
[0003] Currently, the commonly used task scheduling strategies mainly include traditional improved algorithms, heuristic algorithms and intelligent algorithms. Traditional improved algorithms can meet the basic needs of large-scale parallel machine users, but there are problems such as some large jobs being "starved to death" due to long-term unscheduled operations and low system resource utilization; heuristic algorithms are simple and intuitive, but there are problems such as slow initial solution speed, long search time, and premature convergence; intelligent algorithms can autonomously learn how to effectively schedule independent batch jobs in the system for a given optimization goal, but the algorithms are complex and are currently basically not used in ultra-large-scale systems. Summary of the invention
[0004] The purpose of the present invention is to provide a fast job scheduling method for large-scale parallel machines to solve the problem of efficient scheduling of jobs with fixed resource scale requirements in large-scale parallel machines.
[0005] To achieve the above object, the technical solution adopted by the present invention is: to provide a fast job scheduling method for a large-scale parallel machine, based on a job scheduling framework, the job scheduling framework includes:
[0006] Job pool, used to store basic information of all jobs to be scheduled;
[0007] The scheduling strategy module is used to select appropriate jobs from the job pool according to certain rules for scheduling and running;
[0008] The job start module is used to start the user job that has been scheduled and is about to run;
[0009] The job scheduling method comprises the following steps:
[0010] S1. Set the system job waiting time threshold according to the actual system user job requirements;
[0011] S2. Obtain the basic information of all jobs to be scheduled from the job pool, including the job submission queue, submission time, and resource requirements;
[0012] S3, based on the basic information of the jobs to be scheduled obtained in S2, sort all the jobs to be scheduled from large to small according to the calculated priorities;
[0013] S4: Each queue queries in turn whether the available resources in the queue meet the resource requirements of the job to be scheduled based on the job ranking obtained in S3. Queues can query in parallel. S5. If the number of available resources in the queue meets the amount of resources required by the job, the job start module is called to start the job, and the start result is recorded in the database, and the job scheduling is completed;
[0014] S6. If the number of available resources in the queue does not meet the amount of resources required by the job, determine whether the job waiting time exceeds the threshold set in S1. If so, reserve the resources required for the job. After the resource reservation is successful, call the job start module to start the job, and the job scheduling is completed. If not, the job continues to stay in the job pool and wait for the next scheduling.
[0015] S7. All the jobs to be scheduled are processed and this scheduling ends.
[0016] The further improved scheme in the above technical scheme is as follows:
[0017] 1. In the above scheme, the method for calculating the job priority in S3 is: , where K 1 , K 2 , K 3 are weight factors, P vip is the privilege priority value, initially set to 0, P u is the user priority value, P q is the queue priority value, P m is the preset job priority threshold, T wait The actual waiting time of the system job.
[0018] 2. In the above scheme, when calculating the job priority, first determine whether the job has been set with privileges. If so, its priority value is the privileged priority value, and the job is scheduled to run first; otherwise, take the smaller of the calculated priority value and the priority threshold. When the priority of the job reaches the set threshold, the job needs to make a resource reservation to avoid it being unscheduled for a long time.
[0019] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:
[0020] The job scheduling method of the present invention can concurrently and efficiently select appropriate job scheduling operations in a large-scale parallel machine, reduce the total job waiting time, improve system resource utilization, and improve user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Attached Figure 1 It is the job scheduling framework;
[0022] Attached Figure 2 The flowchart for job scheduling is shown in Figure 2. DETAILED DESCRIPTION
[0023] Embodiment: The present invention provides a fast job scheduling method for a large-scale parallel machine, based on a job scheduling framework, the job scheduling framework comprising:
[0024] Job pool, used to store basic information of all jobs to be scheduled;
[0025] The scheduling strategy module is used to select appropriate jobs from the job pool according to certain rules for scheduling and running;
[0026] The job start module is used to start the user job that has been scheduled and is about to run;
[0027] The job scheduling method comprises the following steps:
[0028] S1. Set the system job waiting time threshold according to the actual system user job requirements;
[0029] S2. Obtain the basic information of all jobs to be scheduled from the job pool, including the job submission queue, submission time, and resource requirements;
[0030] S3, based on the basic information of the jobs to be scheduled obtained in S2, sort all the jobs to be scheduled from large to small according to the calculated priorities;
[0031] S4. Each queue queries in turn whether the number of available resources in the queue meets the resource requirements of the job to be scheduled according to the job ranking obtained in S3. Queues can query in parallel. S5. If the number of available resources in the queue meets the resource requirements of the job, the job startup module is called to start the job, and the startup result is recorded in the database. The job scheduling is completed.
[0032] S6. If the number of available resources in the queue does not meet the amount of resources required by the job, determine whether the job waiting time exceeds the threshold set in S1. If so, reserve the resources required for the job. After the resource reservation is successful, call the job start module to start the job, and the job scheduling is completed. If not, the job continues to stay in the job pool and wait for the next scheduling.
[0033] S7. All the jobs to be scheduled are processed and this scheduling ends.
[0034] The method for calculating job priority in S3 is: , where K 1 , K 2 , K 3 are weight factors, P vipis the privilege priority value, initially set to 0, P u is the user priority value, P q is the queue priority value, P m is the preset job priority threshold. The above parameters are preset by the system administrator when the environment is built. wait The actual waiting time of the system job.
[0035] When calculating the job priority, first determine whether the job has been privileged. If so, its priority value is the privileged priority value, and the job is scheduled to run first. Otherwise, take the smaller of the calculated priority value and the priority threshold. When the job priority reaches the set threshold, the job needs to make a resource reservation to avoid it being unscheduled for a long time.
[0036] The above embodiment is further explained as follows:
[0037] The present invention mainly aims at the rapid scheduling demand of massive fixed-resource-scale demand jobs in ultra-large-scale systems, and proposes a modular job scheduling architecture. To ensure the independence of resource management and use, the architecture logically divides system resources in the form of queues in design. The computing resource configurations between queues are independent of each other and have no intersection, which provides a basis for parallel scheduling. A rapid job scheduling method for large-scale parallel machines that supports reservations is designed to shorten the average waiting time of jobs in a system with no cross-resources between queues and improve system resource utilization.
[0038] During the process of job sorting, resource allocation and scheduling, methods such as querying available resources of each queue based on the condition of non-intersection of resources between queues and the concurrent control mechanism between queues, reserving resources for specific large jobs, suspending low-priority jobs based on backfill technology during the reservation process, and concurrent scheduling of job operations are used. This parallel control method can effectively reduce job waiting time, improve system resource utilization, and improve system job throughput.
[0039] like Figure 1 The job scheduling architecture is mainly composed of three parts: job pool, scheduling strategy and job startup module. Among them, the job pool is responsible for storing the basic information of all jobs to be scheduled. The scheduling strategy module is mainly responsible for selecting appropriate jobs from the job pool according to certain rules. This part can customize the scheduling strategy according to needs to control the job more accurately. The job startup module is mainly responsible for starting and running user jobs that have completed scheduling and are about to run.
[0040] Based on this architecture, a fast job scheduling method that supports reservations is proposed. In order to avoid long waiting time for some large-scale jobs in the system and affect user experience, the maximum waiting time for the job is set. If the job waiting time exceeds the threshold, the job needs to reserve resources and be scheduled to run first.
[0041] The specific operation processing flow is as follows: Figure 2 :
[0042] Step 1: Get all job information from the job pool;
[0043] Step 2: Sort the jobs to be scheduled according to the calculated priorities;
[0044] Step 3: Parallel query whether the available resources of each queue meet the needs of the job to be scheduled; Step 4: If the resources meet the needs, call the job start module to start the job, and record the start result in the database, and the job scheduling is completed;
[0045] Step 5: If the resources do not meet the demand, determine whether the job waiting time exceeds the preset threshold. If so, make a reservation for the resources required for the job. After the resource reservation is successful, call the job start module to start the job, and the job scheduling is completed. If not, the job will continue to stay in the job pool and wait for the next scheduling.
[0046] Step 6: All pending jobs are processed and the scheduling ends.
[0047] The job priority is calculated as:
[0048]
[0049] Among them, K 1 , K 2 , K 3 are weight factors, P vip is the privilege priority value, initially set to 0, P u is the user priority value, P q is the queue priority value, T wait is the job waiting time, P m It is the preset job priority threshold.
[0050] When calculating the job priority, first determine whether the job has been set with privileges. If so, its priority value is the privileged priority value, and the job is scheduled to run first. Otherwise, the smaller of the calculated priority value and the priority threshold is taken. When the job priority reaches the set threshold, the job needs to make a resource reservation to avoid it being unscheduled for a long time.
[0051] When the above-mentioned fast job scheduling method for large-scale parallel machines is adopted, it can concurrently and efficiently select appropriate job scheduling operations in large-scale parallel machines, reduce the total job waiting time, improve system resource utilization, and improve user experience.
[0052] In order to facilitate a better understanding of the present invention, the terms used in this article are briefly explained below:
[0053] Job management system: refers to a management and control system that runs on a large-scale parallel machine and integrates job scheduling, startup, control and recovery functions.
[0054] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with the technology to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be included in the protection scope of the present invention.
Claims
1. A fast job scheduling method for large-scale parallel machines, Features: Based on a job scheduling framework, the job scheduling framework includes: The job pool is used to store the basic information of all jobs to be scheduled; The scheduling strategy module is used to select appropriate jobs from the job pool according to certain rules for scheduling and running; The job start module is used to start the user job that has been scheduled and is about to run; The job scheduling method comprises the following steps: S1. Set the system job waiting time threshold according to the actual system user job requirements; S2. Obtain the basic information of all jobs to be scheduled from the job pool, including the job submission queue, submission time, and resource requirements; S3, based on the basic information of the jobs to be scheduled obtained in S2, sort all the jobs to be scheduled from large to small according to the calculated priorities; S4: Each queue queries in turn whether the available resources in the queue meet the resource requirements of the job to be scheduled based on the job ranking obtained in S3. Queues can query in parallel. S5. If the number of available resources in the queue meets the amount of resources required by the job, the job start module is called to start the job, and the start result is recorded in the database, and the job scheduling is completed; S6. If the number of available resources in the queue does not meet the amount of resources required by the job, determine whether the job waiting time exceeds the threshold set in S1. If so, reserve the resources required for the job. After the resource reservation is successful, call the job start module to start the job, and the job scheduling is completed. If not, the job continues to stay in the job pool and wait for the next scheduling. S7: All pending jobs are processed and the scheduling ends. The method for calculating job priority in S3 is: , where K 1 , K 2 , K 3 are weight factors, P vip is the privilege priority value, initially set to 0, P u is the user priority value, P q is the queue priority value, P m is the preset job priority threshold, T wait The actual waiting time of the system job.
2. A fast job scheduling method for large-scale parallel machines according to claim 1, Features: When calculating the job priority, first determine whether the job has been privileged. If so, its priority value is the privileged priority value, and the job is scheduled to run first. Otherwise, take the smaller of the calculated priority value and the priority threshold. When the job priority reaches the set threshold, the job needs to make a resource reservation to avoid it being unscheduled for a long time.
Citation Information
Patent Citations
Cluster assignment dispatching method and device
CN104123183A
Multi-queue back-filling job scheduling method oriented to living organism gene sequencing calculation task
CN105718312A