Job scheduling method, electronic equipment and storage medium

By predicting the resource requirements of tasks to be started, identifying resource bottlenecks, and seizing resources from ordinary tasks, the delay problem of high-security tasks during resource shortages was solved, ensuring their normal operation and improving the timeliness and reliability of tasks.

CN121008929APending Publication Date: 2025-11-25CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511253104.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing job scheduling methods fail to effectively predict future resource competition, causing high-security jobs to be easily blocked when resources are scarce, resulting in start-up and operation delays and affecting the timeliness and reliability of jobs.

Method used

By obtaining the start time and resource usage duration of jobs to be started, future resource needs can be estimated, resource bottlenecks can be identified, and resources can be seized from ordinary jobs when resources are scarce, freeing up necessary resources for high-security jobs and ensuring their normal operation.

Benefits of technology

It effectively reduces competition and conflict between high-security operations and ordinary operations when resources are scarce, improves the operational efficiency and reliability of high-security operations, and avoids delays and interruptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008929A_ABST
    Figure CN121008929A_ABST
Patent Text Reader

Abstract

The invention discloses a job scheduling method, electronic equipment and a storage medium, and relates to the technical field of job scheduling, and the method comprises the steps: obtaining a to-be-started job, and estimating the starting time and resource occupation duration of the to-be-started job, the to-be-started job comprising a high-guarantee job and a common job; according to the starting moments and the resource occupation durations of all the to-be-started jobs, estimating a first job throughput capacity of a next time period; according to the starting time and the resource occupation duration of the high-guarantee operation, estimating the second operation throughput capacity of the next time period; and under the condition that both the first job throughput capacity and the second job throughput capacity are smaller than a preset resource threshold value, determining a target job from the common jobs running in the next time period, and preempting resources of the target job so as to perform job processing of the high-guarantee job according to the occupied resources. According to the method, the resources are preempted for the high-guarantee operation in advance through time sequence estimation and resource occupation, so that the operation timeliness of the high-guarantee operation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of job scheduling technology, and in particular to job scheduling methods, electronic devices and storage media. Background Technology

[0002] With the rapid development of cloud computing, edge computing, and distributed computing technologies, the scale and processing power of computing infrastructure have been improved. However, at the same time, the requirements of business side for job execution efficiency have become increasingly stringent. If high-guarantee jobs with strict requirements on job success rate, startup time, and runtime are not submitted in a timely manner, the job completion time will be delayed, unpredictable, or even impossible to guarantee. If high-guarantee jobs cannot be guaranteed, the service level of computing infrastructure will be greatly limited.

[0003] Current job scheduling methods primarily make decisions based on the current resource supply and demand status. For example, by monitoring the real-time load of nodes, jobs are allocated to nodes with lower loads. However, if high-guarantee jobs arrive when resources are already occupied by low-priority or long-cycle jobs, the high-guarantee jobs will be forced to wait, causing significant delays in startup time or interruptions in operation, severely limiting the service level of computing infrastructure.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this application is to provide a job scheduling method, electronic device, and storage medium, which aims to solve the technical problem of how to improve the runtime efficiency of high-security jobs.

[0006] To achieve the above objectives, this application proposes a job scheduling method, which includes:

[0007] Obtain jobs to be started and estimate the start time and resource usage duration of the jobs to be started, wherein the jobs to be started include high-security jobs and normal jobs;

[0008] Based on the start time and resource usage duration of all the jobs to be started, estimate the throughput capacity of the first job in the next time period;

[0009] Based on the start time and resource usage duration of the high-security operation, the throughput capacity of the second operation in the next time period is estimated;

[0010] If both the throughput capacity of the first job and the throughput capacity of the second job are less than the preset resource threshold, a target job is determined from the ordinary jobs running in the next time period, and the resources of the target job are preempted to ensure high-security job processing based on the preempted resources.

[0011] In one embodiment, the step of estimating the throughput capacity of the first job in the next time period based on the start time and resource usage duration of all the jobs to be started includes:

[0012] The amount of resources available for allocation in the next time period is determined based on the current resource balance and the resource release amount at the start of the next time period.

[0013] Determine the first time period between the current time and the start time of the next time period, and determine the jobs whose start time is within the first time period and whose resource usage duration is greater than the duration corresponding to the first time period as started jobs;

[0014] Based on the resource requirements of the already started jobs, the resource requirements of jobs to be started in the preset execution queue, and the resource requirements of jobs to be started outside the execution queue whose start time belongs to the next time period, the first resource requirement of the next time period is determined.

[0015] The first operation throughput capacity for the next time period is determined based on the allocable resources and the first resource demand.

[0016] In one embodiment, after the step of estimating the throughput capacity of the first job in the next time period based on the start time and resource usage duration of all the jobs to be started, the method further includes:

[0017] Compare the throughput of the first job with the resource threshold;

[0018] If the throughput capacity of the first job is greater than or equal to the resource threshold, submit the jobs to be started that belong to the next time period to the execution queue.

[0019] If the throughput capacity of the first job is less than the resource threshold, the step of estimating the throughput capacity of the second job in the next time period based on the start time and resource occupation duration of the high-security job is executed.

[0020] In one embodiment, the step of estimating the throughput capacity of the second job in the next time period based on the start time and resource occupancy duration of the high-security job includes:

[0021] The second resource requirement for the next time period is determined based on the resource requirements of the started jobs, the resource requirements of the high-security jobs in the execution queue, and the resource requirements of the high-security jobs outside the execution queue whose start time belongs to the next time period.

[0022] The second operation throughput capacity for the next time period is determined based on the allocable resources and the second resource demand.

[0023] In one embodiment, after the step of estimating the throughput capacity of the second job in the next time period based on the start time and resource usage duration of the high-security job, the method further includes:

[0024] If the throughput of the first job is less than the resource threshold and the throughput of the second job is greater than or equal to the resource threshold, add high-security jobs whose start time belongs to the next time period to the preset execution queue.

[0025] According to the preset scheduling strategy, ordinary jobs whose start time belongs to the next time period are added to the execution queue.

[0026] In one embodiment, the step of estimating the start time and resource usage duration of the job to be started includes:

[0027] The quantile of the historical start time of the job to be started is determined as the start time;

[0028] The resource usage duration will be determined based on the quantile of the historical runtime of the job to be started.

[0029] In one embodiment, the step of estimating the start time of the job to be started further includes:

[0030] Obtain the upstream job of the job to be started;

[0031] The start time of the job to be started is determined based on the runtime of the upstream job.

[0032] In one embodiment, the method further includes:

[0033] If the current time is the start time of the next time period, the target job is interrupted and the target job is redefined as a job to be started;

[0034] Run high-security jobs from the pre-defined execution queue.

[0035] Furthermore, to achieve the above objectives, this application also proposes a job scheduling device, which includes:

[0036] The timing prediction module is used to obtain jobs to be started and predict the start time and resource usage duration of the jobs to be started, wherein the jobs to be started include high-security jobs and normal jobs;

[0037] The first throughput capacity estimation module is used to estimate the throughput capacity of the first job in the next time period based on the start time and resource occupation duration of all the jobs to be started.

[0038] The second throughput capacity estimation module is used to estimate the throughput capacity of the second operation in the next time period based on the start time and resource occupation duration of the high-security operation.

[0039] The resource preemption module is used to determine a target job from the ordinary jobs running in the next time period when both the throughput capacity of the first job and the throughput capacity of the second job are less than a preset resource threshold, and to preempt the resources of the target job so as to perform high-security job processing based on the preempted resources.

[0040] In addition, to achieve the above objectives, this application also proposes an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the job scheduling method described above.

[0041] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the job scheduling method described above.

[0042] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the job scheduling method described above.

[0043] The one or more technical solutions proposed in this application have at least the following technical effects: First, by acquiring the jobs to be started and estimating their start time and resource occupation duration, the future resource demand situation can be predicted in advance, providing a data basis for subsequent scheduling decisions; then, based on the start time and resource occupation duration of all jobs to be started, the throughput capacity of the first job in the next time period is estimated, and the overall resource saturation of the next time period is judged in advance, so as to identify the resource bottleneck of the overall job scheduling; furthermore, based on the start time and resource occupation duration of high-security jobs among the jobs to be started, the throughput capacity of the second job in the next time period is estimated, realizing the independent assessment of the resources required for high-security jobs, thereby ensuring that the resource requirements of high-security jobs can be clearly defined; finally, when both the throughput capacity of the first job and the throughput capacity of the second job are less than the preset resource threshold, the target job is determined from the ordinary jobs running in the next time period and its resources are preempted, freeing up the necessary resources to ensure the operation of high-security jobs, so as to process the high-security jobs according to the preempted resources. This application uses timing prediction to intelligently and selectively postpone or limit the resource consumption of some low-priority jobs (ordinary jobs), effectively reducing the competition and conflict between high-security jobs and ordinary jobs when global resources are scarce, ensuring that they can run normally without being stuck in queues, thereby improving the runtime efficiency and reliability of high-security jobs. Attached Figure Description

[0044] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a flowchart illustrating an embodiment of the job scheduling method of this application.

[0047] Figure 2 A flowchart illustrating the job scheduling process provided in Embodiment 1 of this application;

[0048] Figure 3 This is a schematic diagram of the phased invocation of a job as provided in Embodiment 2 of this application;

[0049] Figure 4 This is a schematic diagram of the module structure of the job scheduling device according to an embodiment of this application;

[0050] Figure 5This is a schematic diagram of the device structure of the hardware operating environment involved in the job scheduling method in the embodiments of this application.

[0051] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0052] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0053] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0054] Current job scheduling methods rely solely on current resource snapshots, lacking the ability to predict future resource contention and failing to proactively reserve resources for high-guarantee jobs. Long-running jobs may be mistakenly accepted due to short-term resource abundance, only to be subsequently blocked due to resource shortages. Therefore, in high-load or resource-constrained scenarios, high-guarantee jobs are highly susceptible to resource contention with numerous regular jobs, potentially leading to unexpected delays or even interruptions in their startup and operation, making their completion times unpredictable and reducing runtime efficiency.

[0055] This application provides a solution. First, by acquiring jobs to be started and estimating their start times and resource usage durations, the future resource demand trend is predicted in advance, providing a data foundation for subsequent scheduling decisions. Then, based on the start times and resource usage durations of the jobs to be started, the throughput capacity of the first job in the next time period is estimated, allowing for an early assessment of the overall resource saturation in the next time period, thus identifying resource bottlenecks in overall job scheduling. Next, based on the start times and resource usage durations of high-security jobs among the jobs to be started, the throughput capacity of the second job in the next time period is estimated, enabling an independent assessment of the resources required by high-security jobs, thereby ensuring that the resource requirements of high-security jobs can be clearly defined. Finally, if both the throughput capacity of the first job and the throughput capacity of the second job are less than a preset resource threshold, a target job is identified from the ordinary jobs running in the next time period, and its resources are preempted to free up necessary resources for the operation of high-security jobs, ensuring that high-security jobs can run normally after being submitted to a preset execution queue. This application uses timing prediction to intelligently and selectively postpone or limit the resource consumption of some low-priority jobs (ordinary jobs), effectively reducing the competition and conflict between high-security jobs and ordinary jobs when global resources are scarce, ensuring that they can run normally without being stuck in queues, thereby improving the runtime efficiency and reliability of high-security jobs.

[0056] It should be noted that the executing entity in this embodiment can be an electronic device with data processing, network communication and program execution functions, such as a tablet computer, personal computer, mobile phone, scheduling system, etc.

[0057] Based on this, embodiments of this application provide a job scheduling method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the job scheduling method of this application.

[0058] In this embodiment, the job scheduling method includes steps S10 to S40:

[0059] Step S10: Obtain the jobs to be started and estimate the start time and resource usage duration of the jobs to be started;

[0060] A pending job refers to a job instance that has been submitted to the scheduling system but has not yet been allocated cluster resources or entered the execution queue. It typically exists in the form of a job description file or database record, including metadata such as the job identifier and resource request specifications. The execution queue is a data structure or logical container used in the scheduling system to temporarily store submitted but not yet started jobs.

[0061] The jobs to be started can include high-guarantee jobs and normal jobs. High-guarantee jobs refer to jobs with high requirements for job success rate, job start time, runtime, and job completion time, and are usually identified by preset tags or metadata fields. Normal jobs, on the other hand, represent the default job type without special markings.

[0062] The start time refers to the predicted future timestamp when a job will begin running, which typically depends on factors such as the current execution queue length and resource release rate.

[0063] Resource occupation duration refers to the predicted duration of a job from its start time until it releases the occupied resources. It is usually the same as the job's runtime and can be estimated based on the job's historical runtime.

[0064] Optionally, each natural day can be divided into multiple time periods. During the current time period, the scheduling system can periodically check the job scheduling requirements and cluster throughput capacity for the next time period, so as to adjust the job scheduling in a timely manner according to changes in requirements and throughput capacity, thereby improving the real-time performance and reliability of job scheduling.

[0065] For example, all jobs to be started in the current time period can be obtained by subscribing to job submission events (such as Kafka messages) or by actively querying the job queue; then, the start time and resource consumption duration of each job to be started can be estimated based on the historical running data of each job to be started.

[0066] Understandably, by estimating the start time and resource usage duration of jobs to be started, proactive collection and prediction of future resource information is achieved, enabling more forward-looking job scheduling.

[0067] Step S20: Based on the start time and resource usage duration of all jobs to be started, estimate the throughput capacity of the first job in the next time period;

[0068] Job throughput capacity refers to the amount of resources that electronic devices can provide for use, while first job throughput capacity represents the amount of resources that the cluster can still consume after carrying all jobs to be started in the next time period, and is used to measure the idle level of resources in the next time period.

[0069] In one feasible implementation, step S20 includes:

[0070] Step S21: Determine the amount of resources available for allocation in the next time period based on the current resource balance and the resource release amount at the start of the next time period.

[0071] The current time refers to the current point in time of the system, represented by a timestamp recorded by the system clock. It is updated in real time and is used to compare with time information such as the start time of the job.

[0072] Resource reserves refer to the total amount of resources in the cluster that have not been allocated or occupied; while resource release for the next time period refers to the amount of resources released by completed jobs from the current time to the start of the next time period.

[0073] Allocable resources refer to the amount of resources that can be allocated to jobs in the execution queue at the start of the next time period.

[0074] Optionally, the sum of the current resource surplus and the resource release amount in the next time period can be determined as the allocatable resource amount in the next time period.

[0075] Step S22: Determine the first time period between the current time and the start time of the next time period, and determine the jobs whose start time is within the first time period and whose resource occupation duration is greater than the duration corresponding to the first time period as started jobs;

[0076] Started jobs represent jobs that were started after the current time and before the start of the next time period, and that will still be running in the next time period.

[0077] Step S23: Determine the first resource requirement for the next time period based on the resource requirements of the started jobs, the resource requirements of the jobs to be started in the execution queue, and the resource requirements of the jobs to be started outside the execution queue whose start time belongs to the next time period.

[0078] The resource requirements for each job can be pre-configured before job scheduling, or dynamically estimated based on the historical operation data of each job.

[0079] The jobs waiting to be started in the execution queue represent jobs that have been submitted, are ready, and are waiting to be started. Jobs waiting to be started outside the execution queue, whose start times belong to the next time slot, are jobs that are not yet ready to run but whose start times belong to the next time slot.

[0080] Resource demand refers to the total amount of resources required to run operations in the next time period. The first resource demand represents the resource demand of all operations to be started in the next time period (including normal operations and high-security operations).

[0081] Optionally, the sum of the resource requirements of started jobs, the resource requirements of jobs to be started in the execution queue, and the resource requirements of jobs to be started outside the execution queue whose start time belongs to the next time period can be determined as the first resource requirement of the next time period.

[0082] Optionally, if no new jobs are started in the first time period, the sum of the resource requirements of the jobs to be started in the execution queue and the resource requirements of the jobs to be started outside the execution queue whose start time belongs to the next time period is directly determined as the first resource requirement.

[0083] Optionally, if there are no jobs to be started in the execution queue, the sum of the resource requirements of the already started jobs and the resource requirements of the jobs to be started outside the execution queue whose start time belongs to the next time period can be directly determined as the first resource requirement. Similarly, if there are no jobs to be started outside the execution queue, or if there are no jobs to be started outside the execution queue whose start time belongs to the next time period, the sum of the resource requirements of the already started jobs and the resource requirements of the jobs to be started in the execution queue can be directly determined as the first resource requirement.

[0084] For example, assuming the current time is 3 o'clock, the next time period is from 3:30 to 4:30, the resource reserve is the amount of idle resources in the cluster at 3 o'clock, and the resource release amount is the amount of resources released by the jobs that finish running from 3 o'clock to 3:30. The started jobs are those that started from 3 o'clock to 3:30 and are still running at 3:30. The jobs to be started in the execution queue are those that were submitted to the execution queue before 3 o'clock and are still in the queue at 3:30 (including ordinary jobs and high-guarantee jobs). The jobs to be started outside the execution queue whose start time belongs to the next time period are those that were not submitted to the execution queue before 3 o'clock, but whose start time is between 3:30 and 4:30.

[0085] Step S24: Determine the throughput capacity of the first operation in the next time period based on the allocable resources and the first resource demand.

[0086] Optionally, the difference between the allocable resources and the first resource requirement can be determined as the first job throughput capacity in the next time period, representing the amount of resources available for the cluster to call after it has carried all jobs to be started in the next time period.

[0087] In this embodiment, by placing the supply and demand relationship of resources on the same time scale (the next time period) for precise quantitative comparison, the traditional scheduler makes decisions based only on the current state or a single dimension, thereby achieving optimal decision-making.

[0088] In one possible implementation, after step S20, the method further includes:

[0089] Step S25: Compare the throughput capacity of the first job with the resource threshold;

[0090] Resource threshold refers to a pre-configured system parameter that represents a critical point in the utilization of cluster resources. It can be dynamically adjusted based on factors such as job cycle. For example, at the end of the month or during holidays when resource demand is high and high-security jobs are in the peak of data reporting, the resource threshold can be appropriately lowered to provide more resources for job operation.

[0091] Understandably, setting reasonable resource thresholds allows for sufficient resource margins in the cluster, preventing resources from being pushed to their limits. This effectively prevents performance fluctuations and job failures caused by resource contention, ensuring the stability and high performance of the entire cluster. By comparing the throughput of the first job with the preset resource thresholds, it can be determined whether there are sufficient resources in the cluster to guarantee the operation of all jobs to be started in the next period, thus determining whether to prioritize the normal operation of high-security jobs.

[0092] Step S26: If the throughput capacity of the first job is greater than or equal to the resource threshold, submit the jobs to be started that belong to the next time period to the execution queue.

[0093] Step S27: If the throughput capacity of the first job is less than the resource threshold, perform the step of estimating the throughput capacity of the second job in the next time period based on the start time and resource occupation duration of the high-security job.

[0094] For example, when the throughput of the first job is greater than or equal to a preset resource threshold, it indicates that there are sufficient cluster resources in the next time period to ensure the operation of all jobs to be started that belong to the next time period. Therefore, all jobs to be started (including ordinary jobs and high-guarantee jobs) can be submitted to the execution queue for scheduling. The order of job submission can be determined according to FCFS (First Come First Service) policy, Round-Robin (RR) scheduling algorithm, etc. When the throughput of the second job is less than the preset resource threshold, it indicates that there are insufficient resources in the next time period to run all jobs to be started that belong to the next time period normally. To ensure the normal operation of high-guarantee jobs, it is necessary to further evaluate whether there are sufficient cluster resources in the next time period to run all high-guarantee jobs to determine whether resource pre-emption is necessary.

[0095] Step S30: Based on the start time and resource usage duration of the high-security operation, estimate the throughput capacity of the second operation in the next time period;

[0096] The second task throughput capacity represents the amount of resources remaining in the cluster after carrying all high-guarantee tasks in the next time period.

[0097] In one feasible implementation, step S30 includes:

[0098] Step S31: Determine the second resource requirement for the next time period based on the resource requirements of the started jobs, the resource requirements of high-security jobs in the execution queue, and the resource requirements of high-security jobs outside the execution queue whose start time belongs to the next time period.

[0099] In the execution queue, high-guarantee jobs represent those that have been submitted and are awaiting scheduling. High-guarantee jobs outside the execution queue whose start time belongs to the next time slot refer to those that have not yet been submitted and are not ready to run, but whose start time belongs to the next time slot. The resource requirements for each job can be pre-configured before job scheduling or dynamically estimated based on the historical execution data of each job.

[0100] Optionally, the sum of the resource quantity of started jobs, the resource quantity of high-security jobs in the execution queue, and the resource quantity of high-security jobs outside the execution queue whose start time belongs to the next time period can be determined as the second resource requirement for the next time period.

[0101] Step S32: Determine the throughput capacity of the second operation in the next time period based on the amount of allocable resources and the amount of second resource demand.

[0102] Optionally, the difference between the allocable resources and the second resource demand can be determined as the second job throughput capacity for the next time period, representing the amount of idle resources available for use after the cluster has carried all high-guarantee jobs in the next time period.

[0103] In one possible implementation, after step S30, the method further includes:

[0104] Step S33: If the throughput of the first job is less than the resource threshold and the throughput of the second job is greater than or equal to the resource threshold, add the high-security job whose start time belongs to the next time period to the preset execution queue.

[0105] Step S34: According to the preset scheduling strategy, add ordinary jobs whose start time belongs to the next time period to the execution queue.

[0106] A scheduling strategy is a set of rules or algorithms pre-configured in a scheduling system to determine which job is submitted first when multiple ordinary jobs compete for remaining resources. Common strategies include FCFS, Shortest Job First (SJF), priority queues, and Fair Sharing.

[0107] For example, when the throughput of the first job is less than a preset resource threshold, while the throughput of the second job is greater than or equal to the resource threshold, although the next time period is insufficient to guarantee the normal operation of all jobs to be started, there are enough cluster resources to run high-guarantee jobs whose start time belongs to the next time period. In this case, all high-guarantee jobs whose start time belongs to the next time period can be submitted to the execution queue to wait for scheduling. Then, based on the submission of high-guarantee jobs, ordinary jobs whose start time belongs to the next time period can be added to the execution queue according to the preset scheduling policy.

[0108] For example, in the process of adding a regular job to the execution queue, a candidate list can be formed by first filtering out jobs whose start time belongs to the next time period from all regular jobs; then, the candidate job list can be sorted according to a preset scheduling strategy (such as FCFS) to generate an ordered submission sequence; then, according to the order in the submission sequence, the resource request (required resource amount) of each job can be checked in turn to see if it is less than or equal to the idle resource amount of the next time period (the difference between the throughput capacity of the second job and the resource threshold); then, the regular jobs that pass the check are submitted to the execution queue, and the idle resource amount is reduced by the resource amount requested by the job, and the next job in the submission sequence is checked again until the idle resource amount cannot meet the needs of the next job, or all candidate jobs have been submitted.

[0109] In this implementation, by submitting high-security jobs first and then regular jobs, the normal operation of high-security jobs is effectively guaranteed, thereby improving the operational efficiency of high-security jobs. At the same time, the submission of regular jobs can effectively fill resource gaps and improve the utilization rate of cluster resources from a full-time perspective.

[0110] Step S40: If the throughput capacity of the first job and the throughput capacity of the second job are both less than the preset resource threshold, the target job is determined from the ordinary jobs running in the next time period, and the resources of the target job are preempted and reserved so that the job can be processed with high security based on the reserved resources.

[0111] The target job represents the job selected to perform a preemptive or occupancy operation.

[0112] For example, when the throughput capacity of the first job and the throughput capacity of the second job are both less than the preset resource threshold, the cluster in the next time period lacks sufficient resources to ensure the normal operation of all high-security jobs. In this case, in order to ensure the normal operation of high-security jobs whose start time belongs to the next time period, resources can be preempted for ordinary jobs running in the next time period in advance. When the start time of each high-security job arrives, resources can be preempted for ordinary jobs (target jobs) marked with resource reservations, so that high-security jobs can be processed according to the reserved resources, thereby ensuring the runtime efficiency of high-security jobs.

[0113] For example, after the preemption process is triggered, a target job can be selected from the regular jobs running in the next time slot according to a preset strategy. This could be done by sorting the regular jobs by their remaining runtime, selecting the regular job with the longest remaining runtime as the target job, and also identifying the resources occupied by each target job as resources to be released. The amount of resources to be released and the throughput capacity of the second job are then recalculated. It is then determined whether the recalculated throughput capacity of the second job is greater than or equal to a preset resource threshold. If it is less, the selected target job is removed from the regular jobs, and the regular job with the longest remaining runtime is selected again as the target job. This process is repeated until the recalculated throughput capacity of the second job is greater than or equal to the preset resource threshold. This process identifies the target job selected for resource preemption, and high-security jobs whose startup time belongs to the next time slot can be added to a preset execution queue to await scheduling.

[0114] For example, please refer to Figure 2 , Figure 2A flowchart illustrating job scheduling is provided. First, step B101 is executed to obtain jobs to be started, which can be obtained through lineage analysis of upstream jobs. Next, step B102 is executed to estimate the start time and resource usage duration of the jobs to be started. Based on this estimate, the throughput capacity of the first job in the next time period is further estimated (B103) to assess the resource availability of the cluster after it carries all jobs (including regular and high-guarantee jobs) whose start times belong to the next time period. Then, step B104 is executed to determine if the throughput capacity of the first job is greater than or equal to a preset resource threshold. If so, step B105 is executed to add all jobs to be started, i.e., according to a preset scheduling strategy, all jobs to be started are added to a preset execution queue. The queue is set up for execution. If the throughput of the first job is less than the resource threshold, step B106 is executed to further estimate the throughput of the second job in the next time period. This is to assess the resource availability of the cluster after it carries all high-guarantee jobs whose startup time belongs to the next time period, and to determine whether the throughput of the second job is greater than or equal to the resource threshold (B107). If so, the high-guarantee job is added to the execution queue (B108), and then the ordinary job is added to the execution queue (B109) to ensure that the high-guarantee job can run in a timely manner. If not, step B110 is executed to preempt resources for the ordinary job and add the high-guarantee job to the execution queue (B111). The specific preemption method can be referred to step S40 above, which will not be repeated here.

[0115] This embodiment provides a job scheduling method that, through timing prediction and resource preemption, accurately identifies and resolves resource bottlenecks before resource competition actually occurs, thereby preventing high-availability jobs from falling into resource competition. This priority processing mechanism can ensure the timely completion of high-availability jobs, avoid delays or failures caused by resource competition, and thus improve the operational efficiency of high-availability jobs.

[0116] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, the steps for estimating the start time and resource usage duration of the job to be started include:

[0117] Step S11: Determine the quantile of the historical start time of the job to be started as the start time;

[0118] Step S12: Determine the quantile of the historical runtime of the job to be started as the resource usage duration.

[0119] Historical start times and historical runtimes are both structured data stored in the cluster monitoring system or log database. Each record contains the actual start timestamp and actual runtime of a job for each past execution.

[0120] A quantile is a value that appears at a specific position when a set of data is arranged in ascending order. For example, the 75th percentile means that 75% of the values ​​in the set of data are less than or equal to this number.

[0121] Optionally, the corresponding quantiles can be set according to the time accuracy requirements of the jobs to be started. For example, for high-guarantee jobs, a more conservative quantile can be selected, such as the 25th percentile or even lower. This means that the high-guarantee job is expected to arrive earlier, avoiding the situation where the actual arrival time of the high-guarantee job is earlier than the estimated time, which would lead to being caught off guard and having to sacrifice more ordinary jobs. This would reduce the interruption of ordinary jobs while ensuring that the high-guarantee job can be completed on time in most cases. For ordinary jobs, a medium or high quantile can be selected, such as the 50th percentile (median) or higher. This ensures that the job is completed within a certain time while making reasonable use of system resources and avoiding excessive resource consumption that would affect the high-guarantee job. Furthermore, the quantile of the historical start time of the job to be started can be directly determined as its estimated start time, and the quantile of the historical runtime of the job can be determined as its estimated resource consumption duration.

[0122] For example, the quantile can be set to 100%, that is, the latest start time in the historical running record is determined as the estimated start time, but this may cause resource waiting. Alternatively, the quantile can be set to 0, that is, the earliest start time in the historical running record is determined as the estimated start time. Although this can cause resource waiting with the lowest probability, it may lead to insufficient resources after the job arrives, making it difficult for the job to run.

[0123] In this embodiment, timing prediction is performed using quantiles, which avoids the severe impact of extreme values ​​that can occur when using the mean. This improves the robustness of the prediction of startup time and resource usage duration, thereby enhancing the stability and reliability of scheduling decisions.

[0124] In one feasible implementation, the step of estimating the start time of the job to be started in step S10 includes:

[0125] Step E11: Obtain the upstream job of the job to be started;

[0126] Upstream operations refer to the preceding steps of a task to be started, which can provide the necessary data, status, or triggering conditions for the operation of the task to be started.

[0127] For example, when obtaining jobs to be started, the jobs to be started can be determined by parsing a preset job DAG (Directed Acyclic Graph) and already started jobs; then, when performing time series estimation, the upstream jobs of each job to be started can be traced back through the DAG. The DAG stores the dependencies between all jobs.

[0128] Step E12: Determine the start time of the job to be started based on the remaining runtime of the upstream job.

[0129] For example, firstly, the current status (whether completed) of all upstream jobs is queried; then, for completed jobs, their remaining runtime can be determined to be zero; for running upstream jobs, their remaining runtime can be estimated based on their historical runtime and current runtime. Furthermore, in the case of multiple jobs to be started, for any job to be started, its corresponding upstream job can be determined according to the DAG, and the remaining runtime of the upstream job can be added to the current time to determine the start time of the job to be started.

[0130] For example, please refer to Figure 3 , Figure 3This document provides a schematic diagram of phased job execution. For heterogeneous batch processing jobs using GPUs (Graphics Processing Units) for big data, the first phase is big data processing, primarily handling big data sub-jobs; the second phase is GPU data processing, primarily handling GPU sub-jobs. It can be determined that the big data sub-jobs are upstream jobs of the GPU sub-jobs. The big data sub-jobs primarily run on the big data computing cluster, and their runtime can be estimated using a corresponding job execution data collector. The GPU sub-jobs primarily run on the GPU computing cluster, and their running status can be monitored using its built-in job execution data collector to release cluster resources in a timely manner. The GPU computing cluster also has a built-in cluster resource collector, mainly used to collect cluster resource data to assess resource utilization. Job scheduling is primarily implemented through a scheduling system, which includes a resource estimator, a job scheduler, and a job scheduling information collector. The resource estimator acquires jobs to be started (i.e., determines the GPU sub-jobs to be scheduled based on the running status of upstream jobs) and estimates the start time, resource requirements, and resource occupation duration of the jobs based on job data. The job scheduler then decides whether to submit the GPU sub-jobs to a preset execution queue based on the estimated data. For example, if the throughput capacity of the second job (the cluster's throughput capacity after removing ordinary jobs from the GPU sub-jobs) is less than a preset resource threshold, resource preemption is performed in advance for high-guarantee jobs. The job scheduling information collector mainly records the scheduling information of GPU sub-jobs, including job data such as when the job was paused, when it was scheduled, and the job completion time, for use in the timing prediction of subsequent jobs.

[0131] It is important to note that Figure 3 The job data and cluster resource data shown are for illustrative purposes only and do not represent that all types of data are stored in the same database.

[0132] In this implementation, the start time of the job to be started is inferred by utilizing the dependencies between jobs, so that the prediction results have a strict logical basis, improving the accuracy of the start time estimation, and ensuring that the job to be started can immediately obtain resources and start at the earliest logically possible time, thereby improving the timeliness of all jobs in complex workflows, especially high-guarantee jobs.

[0133] In one feasible implementation, the job scheduling method further includes:

[0134] Step A10: If the current time is the start time of the next time period, interrupt the target job and re-determine the target job as a job to be started;

[0135] An interruption refers to stopping the execution of a job by sending a forced termination signal (such as SIGTERM) to the computing node where the target job resides, or by issuing a command to the resource manager to delete or terminate the target job. This allows the preempted job to perform some work to save its state (such as setting a checkpoint) so that it can resume execution from the breakpoint. After the process of the job is terminated, all computing resources and cluster resources it occupies will be reclaimed (released) and made available.

[0136] Step A20: Run the high-security jobs in the preset execution queue.

[0137] For example, if the estimated throughput capacity of both the first and second jobs in the next time period is less than a preset resource threshold, some ordinary jobs will be preempted for resources. When the next time period arrives, the selected target job is immediately interrupted, and the resources it occupies are released. The target job is then removed from the execution queue and re-identified as a job to be started, awaiting re-evaluation and submission. Furthermore, after the resources are released, it can be detected that high-guarantee jobs are ready (already in the execution queue) and have sufficient resources, and thus the high-guarantee jobs in the execution queue can be run through the compute nodes.

[0138] In this embodiment, by interrupting the previously performed logically occupied and resource-reserved ordinary jobs (target jobs) at the start of the next time period, the physical resources they occupy are released immediately, and these resources are used to start high-security jobs immediately. This achieves a seamless and timely switch of resources from low-priority jobs to high-priority jobs, thereby improving the runtime efficiency of high-security jobs.

[0139] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the job scheduling method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0140] This application also provides a job scheduling device, please refer to... Figure 4 The job scheduling device includes:

[0141] The timing prediction module 10 is used to obtain the jobs to be started and predict the start time and resource usage duration of the jobs to be started. The jobs to be started include high-security jobs and normal jobs.

[0142] The first throughput capacity estimation module 20 is used to estimate the throughput capacity of the first job in the next time period based on the start time and resource occupation duration of all jobs to be started.

[0143] The second throughput capacity estimation module 30 is used to estimate the throughput capacity of the second operation in the next time period based on the start time and resource occupation duration of the high-security operation.

[0144] The resource preemption module 40 is used to determine the target job from the ordinary jobs running in the next time period when the throughput capacity of the first job and the throughput capacity of the second job are both less than the preset resource threshold, and to preempt the resources of the target job so as to perform high-security job processing based on the preempted resources.

[0145] The job scheduling device provided in this application, employing the job scheduling method described in the above embodiments, can solve the technical problem of how to improve the operational efficiency of high-security jobs. Compared with the prior art, the beneficial effects of the job scheduling device provided in this application are the same as those of the job scheduling method provided in the above embodiments, and other technical features in the job scheduling device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0146] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the job scheduling method in the first embodiment described above.

[0147] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0148] like Figure 5As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0149] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0150] The electronic device provided in this application, employing the job scheduling method described in the above embodiments, can solve the technical problem of how to improve the runtime efficiency of high-security jobs. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the job scheduling method provided in the above embodiments, and other technical features of the electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0151] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0152] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0153] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the job scheduling method in the above embodiments.

[0154] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0155] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0156] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by an electronic device, the electronic device causes the electronic device to: acquire jobs to be started and estimate the start time and resource occupation duration of the jobs to be started, wherein the jobs to be started include high-security jobs and normal jobs; estimate the throughput capacity of a first job in the next time period based on the start time and resource occupation duration of all jobs to be started; estimate the throughput capacity of a second job in the next time period based on the start time and resource occupation duration of the high-security jobs; and, if both the throughput capacity of the first job and the throughput capacity of the second job are less than a preset resource threshold, determine a target job from the normal jobs running in the next time period, preempt the resources of the target job, and perform high-security job processing based on the preempted resources.

[0157] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0159] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0160] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described job scheduling method, thereby solving the technical problem of how to improve the runtime efficiency of high-security jobs. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the job scheduling method provided in the above embodiments, and will not be repeated here.

[0161] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the job scheduling method described above.

[0162] The computer program product provided in this application can solve the technical problem of how to improve the runtime efficiency of high-security jobs. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the job scheduling method provided in the above embodiments, and will not be repeated here.

[0163] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A job scheduling method, characterized in that, The job scheduling method includes: Obtain jobs to be started and estimate the start time and resource usage duration of the jobs to be started, wherein the jobs to be started include high-security jobs and normal jobs; Based on the start time and resource usage duration of all the jobs to be started, estimate the throughput capacity of the first job in the next time period; Based on the start time and resource usage duration of the high-security operation, the throughput capacity of the second operation in the next time period is estimated; If both the throughput capacity of the first job and the throughput capacity of the second job are less than the preset resource threshold, a target job is determined from the ordinary jobs running in the next time period, and the resources of the target job are preempted to ensure high-security job processing based on the preempted resources.

2. The job scheduling method as described in claim 1, characterized in that, The step of estimating the throughput capacity of the first job in the next time period based on the start time and resource usage duration of all the jobs to be started includes: The amount of resources available for allocation in the next time period is determined based on the current resource balance and the resource release amount at the start of the next time period. Determine the first time period between the current time and the start time of the next time period, and determine the jobs whose start time is within the first time period and whose resource usage duration is greater than the duration corresponding to the first time period as started jobs; Based on the resource requirements of the already started jobs, the resource requirements of jobs to be started in the preset execution queue, and the resource requirements of jobs to be started outside the execution queue whose start time belongs to the next time period, the first resource requirement of the next time period is determined. The first operation throughput capacity for the next time period is determined based on the allocable resources and the first resource demand.

3. The job scheduling method as described in claim 2, characterized in that, After the step of estimating the throughput capacity of the first job in the next time period based on the start time and resource usage duration of all the jobs to be started, the method further includes: Compare the throughput of the first job with the resource threshold; If the throughput capacity of the first job is greater than or equal to the resource threshold, submit the jobs to be started that belong to the next time period to the execution queue. If the throughput capacity of the first job is less than the resource threshold, the step of estimating the throughput capacity of the second job in the next time period based on the start time and resource occupation duration of the high-security job is executed.

4. The job scheduling method as described in claim 3, characterized in that, The step of estimating the throughput capacity of the second operation in the next time period based on the start time and resource occupation duration of the high-security operation includes: The second resource requirement for the next time period is determined based on the resource requirements of the started jobs, the resource requirements of the high-security jobs in the execution queue, and the resource requirements of the high-security jobs outside the execution queue whose start time belongs to the next time period. The second operation throughput capacity for the next time period is determined based on the allocable resources and the second resource demand.

5. The job scheduling method as described in claim 1, characterized in that, After the step of estimating the throughput capacity of the second operation in the next time period based on the start time and resource usage duration of the high-security operation, the method further includes: If the throughput of the first job is less than the resource threshold and the throughput of the second job is greater than or equal to the resource threshold, add high-security jobs whose start time belongs to the next time period to the preset execution queue. According to the preset scheduling strategy, ordinary jobs whose start time belongs to the next time period are added to the execution queue.

6. The job scheduling method as described in claim 1, characterized in that, The steps for estimating the start time and resource usage duration of the job to be started include: The quantile of the historical start time of the job to be started is determined as the start time; The resource usage duration will be determined based on the quantile of the historical runtime of the job to be started.

7. The job scheduling method as described in claim 1, characterized in that, The step of estimating the start time of the task to be started also includes: Obtain the upstream job of the job to be started; The start time of the job to be started is determined based on the remaining runtime of the upstream job.

8. The job scheduling method as described in claim 1, characterized in that, The method further includes: If the current time is the start time of the next time period, the target job is interrupted and the target job is redefined as a job to be started; Run high-security jobs from the pre-defined execution queue.

9. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the job scheduling method as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the job scheduling method as described in any one of claims 1 to 8.