Resource scheduling strategy determination method and device, storage medium, electronic equipment and program product
By determining the job type and resource requirements in a high-performance computing system, combining the resource usage of computing nodes, and formulating resource scheduling strategies, the problem of low scheduling efficiency in the existing technology is solved, and efficient resource utilization and job continuity are achieved.
Patent Information
- Application Number
- CN202412000527.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In the field of high-performance computing, it is difficult for the prior art to effectively schedule computing resources, resulting in low resource scheduling efficiency, especially during the elastic scaling and preemption of job resources, resulting in loss of computing progress.
By obtaining the pending jobs submitted by the client, determining the job type and resource requirements, and combining the resource scheduled data of the computing node, determining the resource scheduling strategy. This strategy includes utilizing scheduleable resources, shared resources and to-be-occupied resources to ensure that jobs can be executed at the minimum resource requirements and restore the preempted job status after resource release.
It improves the scheduling efficiency of computing resources, avoids idleness and waste of resources, ensures job continuity and resource utilization of high-performance computing clusters.
Smart Images

Figure CN120045315A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of high-performance computing clusters. Specifically, the present application relates to a method, apparatus, storage medium, electronic device, and program product for determining a resource scheduling strategy. Background Art
[0002] Currently, in the field of high-performance computing, especially when dealing with tasks such as large-scale scientific computing and deep learning, the resource scheduling of clusters faces severe challenges.
[0003] In related technologies, for the elastic scaling and preemption of job resources, it is usually handled through a simple job suspension or restart mechanism. After being preempted, a job is either completely suspended or directly terminated, resulting in the loss of existing computing progress. However, the above methods only manage resources in a coarse-grained manner and do not delve into the details of the process state, thus there is a technical problem of low scheduling efficiency of computing resources. Summary of the Invention
[0004] Embodiments of the present application provide a method, apparatus, storage medium, electronic device, and program product for determining a resource scheduling strategy, so as to at least solve the problem of low efficiency in recycling internal objects in related technologies.
[0005] According to an embodiment of the present application, a method for determining a resource scheduling strategy is provided. The method may include: obtaining a to-be-processed job submitted by a client to a scheduling system; determining the job type of the to-be-processed job and the resource requirement quantity of computing resources required by the to-be-processed job; determining the resource already scheduled data of multiple computing resources in a computing node, and based on the resource already scheduled data, job type, and resource requirement quantity, determining a resource scheduling strategy corresponding to the to-be-processed job, where the resource already scheduled data is used to represent the usage situation of multiple computing resources, and the resource scheduling strategy is used to represent the rule for scheduling the computing resources required by the to-be-processed job from the computing node.
[0006] In an exemplary embodiment, determining a resource scheduling strategy corresponding to the to-be-processed job based on the resource already scheduled data, job type, and resource requirement quantity includes: based on the resource already scheduled data, determining the quantity of schedulable resources in the computing node, where the quantity of schedulable resources is used to characterize the quantity of computing resources allowed to be scheduled among multiple computing resources; based on the job type, resource requirement quantity, and quantity of schedulable resources, determining the resource scheduling strategy.
[0007] In an exemplary embodiment, determining the resource scheduling strategy based on the job type, resource requirement quantity, and quantity of schedulable resources includes: comparing the resource demand quantity and the quantity of schedulable resources to obtain a comparison result; based on the comparison result and the job type, determining the scheduling strategy.
[0008] In an exemplary embodiment, a scheduling policy is determined based on a comparison result and a job type, including: in response to the job type being a non-real-time job type and the comparison result being that the number of schedulable resources is greater than or equal to the number of resource requirements, determining the resource scheduling policy as: using the computable resources allowed for scheduling to execute the pending job.
[0009] In an exemplary embodiment, a scheduling policy is determined based on a comparison result and a job type, including: in response to the job type being a non-real-time job type and the comparison result being that the number of schedulable resources is less than the number of resource requirements, determining whether there are shared resources in the computing nodes; in response to there being shared resources in the computing nodes, determining the sum of the shared resources and the number of schedulable resources, and in response to the sum of the shared resources and the number of schedulable resources being greater than or equal to the number of resource requirements, determining the resource scheduling policy as: using the computable resources allowed for scheduling and the shared resources to execute the pending job; in response to there being no shared resources in the computing nodes, or the sum of the shared resources and the number of schedulable resources being less than the number of resource requirements, determining the resources to be occupied; based on the number of resources to be occupied, determining the resource scheduling policy as: using the schedulable computable resources, shared resources, and the resources to be occupied to execute the pending job.
[0010] In an exemplary embodiment, determining the resources to be occupied includes: based on the resource scheduling data, determining at least one job being processed among multiple computable resources; from the at least one job, determining at least one target job, where the priority of the target job is lower than the priority of the pending job; determining the computable resources corresponding to the target job as the resources to be occupied.
[0011] In an exemplary embodiment, the method may further include: in response to the computable resources of a job being occupied by the pending job, creating a process snapshot for the target job, where the process snapshot is used to save the execution state of the computable resources before occupation.
[0012] In an exemplary embodiment, a scheduling policy is determined based on a comparison result and a job type, including: in response to the job type being a real-time job type and the comparison result being that the number of schedulable resources is greater than or equal to the number of resource requirements, determining the resource scheduling policy as: using the computable resources allowed for scheduling to execute the pending job.
[0013] In an exemplary embodiment, a scheduling policy is determined based on a comparison result and a job type, including: in response to the job type being a real-time job type and the comparison result being that the number of schedulable resources is less than the number of resource requirements, determining whether there are occupiable computable resources allowed for occupation among the scheduled computable resources; in response to there being occupiable computable resources, determining the scheduling policy as: using the schedulable computable resources and the occupiable computable resources to execute the pending job.
[0014] In an exemplary embodiment, the method may further include: sending an occupancy request to an agent process according to a resource scheduling policy, where the occupancy request includes identity information of the computable resources that can be occupied; obtaining a successful occupancy instruction returned by the agent process; and in response to the successful occupancy instruction, saving the execution status of the job process in the computable resources that can be occupied.
[0015] In an exemplary embodiment, the method may further include: executing a job to be processed according to a resource scheduling policy; releasing the computable resources occupied by the job to be processed in response to the completion of the execution of the job to be processed; and resuming the historical job process in the computable resources that can be occupied, where the historical job process is used to represent the execution status of the job executed by the computable resources that can be occupied before the computable resources that can be occupied are occupied.
[0016] According to another embodiment of the present application, there is also provided a device for determining a resource scheduling policy, which may include: an obtaining unit configured to obtain a job to be processed submitted by a client to a scheduling system; a first determining unit configured to determine the job type of the job to be processed and the quantity of resource requirements of the computable resources required by the job to be processed; and a second determining unit configured to determine the resource scheduled data of multiple computable resources in a computing node, and determine a resource scheduling policy corresponding to the job to be processed based on the resource scheduled data, the job type, and the quantity of resource requirements, where the resource scheduled data is used to represent the usage conditions of the multiple computable resources, and the resource scheduling policy is used to represent the rules for scheduling the computable resources required by the job to be processed from the computing node.
[0017] According to still another embodiment of the present application, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0018] According to still another embodiment of the present application, there is also provided an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0019] According to still another embodiment of the present application, there is also provided a computer program product including a computer program, where the computer program implements the steps in any one of the above method embodiments when executed by a processor.
[0020] Through this application, a to-be-processed job submitted by a client to a scheduling system is obtained; the job type of the to-be-processed job and the quantity of resource requirements for the computing resources required by the to-be-processed job are determined; the resource already-scheduled data of multiple computing resources in a computing node is determined, and based on the resource already-scheduled data, the job type, and the quantity of resource requirements, a resource scheduling policy corresponding to the to-be-processed job is determined, where the resource already-scheduled data is used to represent the usage conditions of the multiple computing resources, and the resource scheduling policy is used to represent the rules for scheduling the computing resources required by the to-be-processed job from the computing node. That is to say, in the embodiments of this application, after obtaining the to-be-processed job, the job type of the to-be-processed job and the quantity of resource requirements for executing the to-be-processed job can be determined, the resource already-scheduled data in the computing resources is determined, and based on the already-scheduled data, the job type, and the quantity of resource requirements, the resource scheduling policy corresponding to the to-be-processed job can be determined, thereby solving the technical problem of low scheduling efficiency of computing resources and achieving the technical effect of improving the scheduling efficiency of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a hardware structure block diagram of a server device for a method for determining a resource scheduling policy according to an embodiment of this application;
[0022] Figure 2 is a flowchart of a method for determining a resource scheduling policy according to an embodiment of this application;
[0023] Figure 3 is a schematic diagram of creating a process snapshot according to an embodiment of this application;
[0024] Figure 4 is a structure block diagram of a device for determining a resource scheduling policy according to an embodiment of this application;
[0025] Figure 5 is a computer system structure block diagram of an electronic device according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The embodiments of this application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0027] It should be noted that the terms "first", "second", etc. in the description, claims, and the above drawings of this application are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence.
[0028] First, some nouns or terms that appear during the description of the embodiments of this application are applicable to the following explanations:
[0029] A cluster manager and job scheduling system (Simple Linux Utility for Resource Management, abbreviated as Slurm) can be a highly scalable and fault-tolerant cluster manager and job scheduling system that can be used for large clusters of computing nodes;
[0030] Job: In a job scheduling system, a job can refer to a task or a set of tasks that require computing resources (such as CPU, memory, disk space, etc.) and time to execute. It can consist of one or more processes that execute on computing nodes and share resources. Jobs can be of various types, such as scientific simulations, data analysis, machine learning, etc. In the Slurm job scheduling system, after a user submits a job, Slurm can schedule the job to an appropriate computing node based on the job's requirements and the availability of system resources. Users can submit and manage jobs through Slurm command-line tools (such as sbatch, srun, etc.);
[0031] Real-time job: A real-time job can be a job that needs to be completed within a specific time; otherwise, it may lead to serious consequences, such as tasks like financial market simulations and real-time data analysis. Real-time jobs are usually scheduled prior to non-real-time jobs;
[0032] Non-real-time job: A non-real-time job can be a task that does not have a strict deadline and can be completed at any time. For example, tasks such as scientific simulations and data analysis. The scheduling of the above non-real-time jobs usually depends on resource availability and priority;
[0033] Preemption: In a high-performance computing (HPC) scheduling system, preemption can be a mechanism that allows the system to preempt running jobs when needed. It can be used to optimize resource utilization and improve the job completion speed. Preemption can occur between jobs, between nodes, or between computing resources. When a high-priority job is submitted to the system, the scheduling system may choose to preempt one or more running low-priority jobs to allocate resources to the newly submitted job. Among them, the preempted job will be suspended and rescheduled when resources are available;
[0034] Compute Unified Device Architecture (abbreviated as CUDA): It can be a general-purpose parallel computing architecture that enables GPUs to solve complex computing problems. It can include the CUDA instruction set architecture (ISA) and the parallel computing engine inside the GPU;
[0035] CUDA library: It can be a database containing a set of pre-compiled functions and classes, which can help developers more conveniently develop parallel computing applications on the CUDA platform;
[0036] A parallel computing (CUDA kernel) function can refer to a function used to perform parallel computing on a GPU in a CUDA program. A CUDA kernel function is a function designed to be executed in parallel on multiple threads, which are organized into a specific execution model, such as thread blocks, grids, and multi-dimensional indexing. A CUDA kernel function can be declared by using a special function call syntax (such as __global__), and can be launched on the GPU by calling the Application Programming Interface (API) of CUDA;
[0037] A software tool (checkpoint / restore in userspace, abbreviated as CRIU) can be a tool running on an operating system, which can implement the checkpoint / restore function in user space. Using this tool, a running program can be frozen and checkpointed to a series of files associated with the program, and then these files can be used to restore the program to the point when it was frozen on any host. In other words, it is a backup and restoration of the running program environment.
[0038] In this embodiment, a method for determining a resource scheduling policy is also provided. The system for implementing the embodiment and the preferred implementation manner has been described above and will not be repeated here. As used below, the terms "module" and "unit" are combinations of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0039] As an optional implementation manner, the method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking the execution on a server device as an example, Figure 1 is a hardware structure block diagram of a server device for a method of determining a resource scheduling policy according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in Figure 1The structure shown is only illustrative and does not limit the structure of the above server device. For example, the server device may further include more or fewer components than those shown in Figure 1 or have a different configuration from that shown in Figure 1 .
[0040] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the resource scheduling policy in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0041] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0042] In this embodiment, a method for determining a resource scheduling policy is provided. Figure 2 is a flowchart of the method for determining the resource scheduling policy according to the embodiments of the present application, as shown in Figure 2 . The method may include the following steps:
[0043] Step S202, obtaining a to-be-processed job submitted by a client to a scheduling system;
[0044] Step S204, determining the job type of the to-be-processed job and the quantity of resource requirements for the computing resources required by the to-be-processed job;
[0045] Step S206: Determine the resource scheduled data of multiple computing resources in the computing node, and based on the resource scheduled data, job type, and resource requirement quantity, determine the resource scheduling policy corresponding to the job to be processed, where the resource scheduled data is used to represent the usage of multiple computing resources, and the resource scheduling policy is used to represent the rules for scheduling the computing resources required for the job to be processed from the computing node.
[0046] In this embodiment, the above job to be processed can be a computing task (which can be simply referred to as a task) submitted by a user to the scheduling system and waiting for resource allocation and execution, and can be any task that requires computing resources (such as CPU, GPU, memory, etc.) for program or data processing. For example, the above job to be processed can be a deep learning model training job submitted by a user, and this job requires 8 CPU cores and 2 GPUs for training. The above job types can include real-time job types and non-real-time job types. Among them, real-time jobs have strict time response requirements and usually have higher priorities, and are used in application scenarios that require immediate feedback; non-real-time jobs do not have such time limitations, can tolerate delays or interruptions, and usually have lower priorities. The above resource requirement quantity can be used to determine the quantity of resources required when processing the job to be processed. The above computing resources can refer to the hardware devices in a high-performance computing cluster used to execute jobs, including but not limited to CPUs, GPUs, storage devices (such as hard disks, SSDs), network resources, etc. The above resource scheduling policy can be used to represent the rules for scheduling the computing resources required for the job to be processed from the computing node, and can refer to the rules and logic by which the scheduling system decides how to allocate and manage computing resources according to the current resource status, job type, and job requirements, and can be used to optimize resource usage efficiency and meet job execution requirements, and can include real-time job preemption policies or resource sharing policies. For example, the above resource scheduling policy can be used to determine the usage and scheduling status of computing resources in different locations in the computing node. It should be noted that only examples are given here, and there are no specific restrictions on the above job to be processed, the type of computing resources, and the content of the resource scheduling policy.
[0047] Optionally, in the scheduling system, the jobs to be processed can be divided into two types: real-time job types and non-real-time job types. Among them, the jobs to be processed of the above real-time job type exclusively occupy the CPU and do not share CPU resources with other jobs; the jobs to be processed of the non-real-time job type can share CPU resources among themselves. When a user submits a job to be processed to the scheduling system, the type of the job (whether it is a real-time job or a non-real-time job) and the quantity of CPU and GPU resources required by the job can be determined in advance. Through the above steps, the job type of the job to be processed and the resource requirement quantity of the computing resources required by the job to be processed can be obtained.
[0048] Optionally, the scheduling system can receive job requests submitted by users or automated tasks. The job request may include description information of the job, the types (such as CPU, GPU) and quantities of required resources. Based on this job request, the job to be processed can be obtained. By analyzing the job to be processed, the job type of the job to be processed and the quantity of resource requirements for the computing resources required by the job to be processed can be determined. At this time, it is necessary to confirm the resource scheduling data of the computing resources in the computing node to determine the available computing resources in the computing node. Further, the scheduling system is based on the current resource usage situation (i.e., the resource already scheduled data) on each computing node. Based on the already scheduled data, job type and quantity of resource requirements, the scheduling system can determine how to allocate resources for the job to be processed to obtain a resource scheduling policy. For example, if a real-time job requests resources, the scheduling system will give priority to preempting the resources of non-real-time jobs.
[0049] Optionally, the above job types can be divided into real-time jobs and non-real-time jobs, and can be classified according to the nature and requirements of the jobs. Among them, the above quantity of resource requirements may refer to the specific quantity and type of resources required for job completion. For example, the number of CPU cores, the number of GPUs and their models. It should be noted that this is only an example here, and the type of the quantity of resource requirements is not specifically limited.
[0050] For example, assume that in an HPC cluster, non-real-time job one is using 5 CPU cores and 1 GPU on node one, while real-time job two requires 10 CPU cores and 1 GPU to run. The scheduling system receives the submission of real-time job two. The scheduling system analyzes job two and determines that it is a real-time job that requires 10 CPU cores and 1 GPU. Then the scheduling system can check the resource already scheduled data and find out the resource usage situation of non-real-time job one on node one. The scheduling system can decide that since job two is a real-time job, it will pause or scale down job one to release the GPU and some CPU resources used by job one on node 1 to meet the requirements of job two. For example, job one can retain 2 CPU cores to continue running, while job two can use the remaining 8 CPU cores and this GPU.
[0051] In this embodiment, by allowing non-real-time jobs to share CPU resources, resource idle can be avoided and the overall efficiency of the cluster can be improved. The resource requirements of real-time jobs are preferentially met to ensure their completion within the specified time. The job elastic scaling solution allows dynamic adjustment of resources, enabling the scheduling system to more flexibly respond to fluctuations in resource requirements and improving the adaptability and response speed of the cluster. Through more effective resource management and scheduling, unnecessary resource waste can be reduced and operating costs can be lowered. Even in the case of resource shortage, non-real-time jobs can continue to run, avoiding the situation of completely canceling jobs, and improving the job completion rate and user satisfaction.
[0052] Optionally, through the above resource scheduling strategy, the cluster can ensure that real-time jobs obtain resources first, meet time-sensitive computing requirements, and improve the real-time performance of data processing. Through the above steps, the resources allocated to a job can be dynamically adjusted according to the job requirements, improving resource utilization efficiency and reducing waste. Through the above method, changes in computing resource requirements can be quickly responded to, sudden tasks can be effectively handled, and the overall computing power of the cluster can be improved to enhance the flexibility of the cluster.
[0053] Optionally, through an intelligent scheduling algorithm, the utilization rate of all computing resources in the cluster can be maximized, improving the overall performance and economic benefits of the cluster. Even in the case of resource constraints, it can ensure that the jobs submitted by users have the opportunity to be executed, reducing waiting time and job failure rate.
[0054] Through the above method, after obtaining the job to be processed, the job type of the job to be processed and the quantity of resource requirements for executing the job to be processed can be determined, and the resource scheduled data in the computing resources can be determined. Based on the scheduled data, job type, and quantity of resource requirements, the resource scheduling strategy corresponding to the job to be processed can be determined, thereby solving the technical problem of low scheduling efficiency of computing resources and achieving the technical effect of improving the scheduling efficiency of computing resources.
[0055] In an exemplary embodiment, step S206, determining the resource scheduling strategy corresponding to the job to be processed based on the resource scheduled data, job type, and quantity of resource requirements, includes: based on the resource scheduled data, determining the quantity of schedulable resources in the computing nodes, where the quantity of schedulable resources is used to represent the quantity of computing resources allowed to be scheduled among multiple computing resources; based on the job type, quantity of resource requirements, and quantity of schedulable resources, determining the resource scheduling strategy.
[0056] In this embodiment, the resource scheduled data is obtained, and based on the resource scheduled data, the quantity of schedulable resources in the computing nodes can be determined. The quantity of schedulable resources can be used to represent the quantity of computing resources allowed to be scheduled among multiple computing resources. Further, based on the job type, quantity of resource requirements, and quantity of schedulable resources, the resource scheduling strategy can be determined.
[0057] Optionally, the scheduling system determines the resource scheduling strategy of the job to be processed based on the resource scheduled data, job type, and quantity of resource requirements. This process can be divided into two sub-steps: based on the resource scheduled data, determining the quantity of schedulable resources in the computing nodes, and further, based on the job type, quantity of resource requirements, and quantity of schedulable resources, determining the resource scheduling strategy.
[0058] Optionally, the above-mentioned resource-scheduled data includes the usage of various computing resources (such as CPUs and GPUs) on the current computing node. By analyzing this data, the scheduling system can calculate which resources are idle, which resources can be further utilized, and which resources are occupied by existing jobs but may be partially released. The number of schedulable resources reflects the number of resources that the computing node can provide for the jobs to be processed based on the existing jobs. Therefore, based on this, the scheduling system can determine how to allocate the computing resources (which can be simply referred to as resources) in the computing node according to the type of the job (real-time or non-real-time), the number of resources required by the job, and the number of schedulable resources on the computing node. For real-time jobs, the scheduling system tends to monopolize resources and gives priority to the preemption mechanism; for non-real-time jobs, the scheduling system will more flexibly adopt resource sharing or partial resource allocation strategies.
[0059] For example, assume that there are 16 CPU cores on Node 1. Among them, 8 CPU cores are used by non-real-time Job 2, and Job 2 only uses 60% of the CPU core performance. At the same time, the GPU resources on Node 1 are not occupied by any job. The scheduling system analyzes the resource-scheduled data and determines that the number of resources that can be scheduled to real-time Job 1 on Node 1 is the remaining 8 CPU cores and 1 GPU. In addition, it can also consider temporarily sharing some CPU cores (for example, 2) from Job 2 to meet the requirements of Job 1. Further, it can be determined that real-time Job 1 requires 10 CPU cores and 1 GPU resource. After analyzing the number of schedulable resources on Node 1, the scheduling system finds that the resources directly meeting the requirements of Job 1 are insufficient. Therefore, the scheduling system formulates the following resource scheduling strategies: Real-time preemption strategy: The scheduling system decides to suspend or downsize non-real-time Job 2 and release some CPU resources occupied by Job 2 to meet the resource requirements of Job 1. For example, Job 1 can use the remaining 8 CPU cores, the additional 2 CPU cores shared from Job 2, and the idle GPU resources on Node 1. Resource dynamic allocation strategy: If the resource requirements of Job 1 cannot be fully met, the scheduling system can consider dynamically searching for or borrowing resources from other non-real-time jobs or nodes until the minimum resource requirements of Job 1 are met to ensure the timely execution of Job 1.
[0060] In summary, through the above steps, different resource scheduling strategies are provided for jobs of different job types. Thus, through precise resource scheduling strategies, the resources on the computing nodes can be utilized more effectively, avoiding resource idleness or over-allocation. Moreover, the priorities of real-time jobs are guaranteed to ensure that they can be completed in the shortest time to meet time-sensitive requirements. Non-real-time jobs can continue to execute through partial resource allocation or resource sharing mechanisms in case of resource tension, avoiding the situations of complete waiting or task cancellation and improving the job completion rate. Effective resource management and scheduling strategies reduce resource waste and help control the operation costs of the cluster. Even in an environment of highly tense resources, it can ensure that jobs are reasonably scheduled and executed, improving the system's support capacity for multiple users and multiple tasks and enhancing the availability and user satisfaction of the cluster.
[0061] Through the above steps, the scheduling system can achieve intelligent management of resources, dynamically adjust resource allocation according to job characteristics and resource status, and achieve the goals of optimizing resource usage, improving cluster efficiency, and meeting different job requirements.
[0062] In an exemplary embodiment, based on the job type, the number of resource requirements, and the number of schedulable resources, a resource scheduling strategy is determined, including: comparing the resource requirement quantity and the number of schedulable resources to obtain a comparison result; and determining a scheduling strategy based on the comparison result and the job type.
[0063] In this embodiment, the resource requirement quantity and the number of schedulable resources are compared to obtain a comparison result, which can be used to characterize the difference between the resource requirement quantity and the number of schedulable resources. Based on this comparison result, it can be determined whether the number of schedulable resources in the current computing node can meet the requirements of the job to be processed. Further, based on the comparison result and the job type, a scheduling strategy can be determined.
[0064] Optionally, the scheduling system first needs to evaluate whether the requirements of the job to be processed can be met by the currently schedulable resources. This involves comparing the resources required by the job (such as the number of CPU cores, the number of GPUs) with the unallocated resources available on the current node to obtain a comparison result. This comparison result can be used to guide the scheduling system to take corresponding actions. If the resource requirement quantity is less than or equal to the number of schedulable resources, the scheduling system can directly allocate resources to start the job. If the resource requirement quantity is greater than the number of schedulable resources, the scheduling system can decide whether to meet the resource requirements by preempting the resources of non-real-time jobs, sharing CPU resources, or waiting for resources to become available according to the job type (real-time or non-real-time).
[0065] In summary, through the intelligent comparison and decision-making in the above steps, the scheduling system can maximize the utilization of the computing resources of the cluster and avoid resource idleness and waste. The above scheduling strategy can be flexibly adjusted according to the changes in job types and resource requirements, improving the adaptability of the cluster in the face of different loads.
[0066] Optionally, through the above steps, even when resources are scarce, non-real-time jobs submitted by users can continue to execute in some form, reducing the waiting time and failure rate of jobs and enhancing users' satisfaction with the cluster resource allocation. The dynamic resource scheduling strategy helps reduce unnecessary resource allocation, avoiding over-investment in hardware to cope with occasional peak demands, thereby controlling operating costs.
[0067] The processing process of the jobs to be processed of the non-real-time job type will be further described below.
[0068] In an exemplary embodiment, based on the comparison result and the job type, a scheduling strategy is determined, including: in response to the job type being a non-real-time job type and the comparison result being that the number of schedulable resources is greater than or equal to the number of resource requirements, determining the resource scheduling strategy as: using the allowed schedulable computing resources to execute the job to be processed.
[0069] In an exemplary embodiment, based on the comparison result and the job type, a scheduling strategy is determined, including: in response to the job type being a non-real-time job type and the comparison result being that the number of schedulable resources is less than the number of resource requirements, determining whether there are shared resources in the computing nodes; in response to there being shared resources in the computing nodes, determining the sum of the shared resources and the number of schedulable resources, and in response to the sum of the shared resources and the number of schedulable resources being greater than or equal to the number of resource requirements, determining the resource scheduling strategy as: using the allowed schedulable computing resources and the shared resources to execute the job to be processed; in response to there being no shared resources in the computing nodes, or the sum of the shared resources and the number of schedulable resources being less than the number of resource requirements, determining the resources to be occupied; based on the number of resources to be occupied, determining the resource scheduling strategy as: using the schedulable computing resources, the shared resources, and the resources to be occupied to execute the job to be processed.
[0070] In an exemplary embodiment, determining the resources to be occupied includes: based on the resource scheduling data, determining at least one job being processed among multiple computing resources; from the at least one job, determining at least one target job, where the priority of the target job is lower than the priority of the job to be processed; determining the computing resources corresponding to the target job as the resources to be occupied.
[0071] In an exemplary embodiment, the method may further include: in response to the computing resources of a job being occupied by the job to be processed, creating a process snapshot for the target job, where the process snapshot is used to save the execution state of the computing resources before being occupied.
[0072] In this embodiment, it is determined whether the job type is a non-real-time job type. If the job type is a non-real-time job type and the comparison result is that the number of schedulable resources is greater than or equal to the number of resource requirements, then it can be determined that the current number of computing resources can meet the requirements of the job to be processed. Therefore, it can be determined that the scheduling policy is to use the computable resources allowed for scheduling to execute the job to be processed.
[0073] In this embodiment, if the comparison result is that the number of schedulable resources is less than the number of resource requirements, then it can be determined that the current number of computable resources allowed for scheduling cannot meet the requirements of the job to be processed. Further, it can be determined whether there are shared resources in the computing node. The shared resources can be shared resources for other non-real-time jobs. If there are shared resources in the computing node, then it can be determined whether the sum of the shared resources and the schedulable resources meets the number of required resources. If the sum of the two meets the number of required resources, then it can be determined that the scheduling policy is to use the shared resources and the schedulable resources to execute the job to be processed. If the sum of the shared resources and the schedulable resources does not meet the number of required resources, then the resources to be occupied in the computing node can be further determined. The resources to be occupied can be the resources allowed to be occupied in the computing node and can be the resources being used by non-real-time jobs with a lower priority than the job to be processed.
[0074] Optionally, after determining the resources to be occupied, the above-mentioned job to be processed can be processed among the schedulable resources, the resources to be occupied, and the shared resources. Among them, the above-mentioned resources to be occupied can be the computable resources allowed to be occupied in the computing node. The above-mentioned schedulable resources can be the resources in the computing node that are in an idle state and allowed to be called.
[0075] In this embodiment, the resources to be occupied can be determined through the following steps: Based on the resource scheduled data, at least one job being processed among multiple computing resources can be determined, and at least one target job can be determined from the at least one job. The target job can be a non-real-time job with a lower priority than the job to be processed, and the computing resources used by the target job can be determined as the resources to be occupied.
[0076] In this embodiment, before occupying the resources to be occupied, a process snapshot can be created for the target job. The process snapshot can be used to save the execution state of the computing resources before occupation, can be constructed using a CUDA agent, and can include the context information of the target job.
[0077] Optionally, in the scheduling system, the jobs to be processed can be divided into two types: real-time jobs and non-real-time jobs. Real-time jobs exclusively occupy the CPU and do not share CPU resources with other jobs; non-real-time jobs can share CPU resources with each other. Users can submit jobs to be processed to the scheduling system. When submitting a job to be processed, the user can specify whether the job to be processed is a real-time job or a non-real-time job, as well as the required amounts of CPU and GPU resources for the job. For non-real-time jobs, after submitting the job, the scheduling system can check whether there are idle CPU and GPU resources (i.e., computing resources) in the system. If there are idle resources, the job will be scheduled for execution immediately; if there are not enough resources, it will be considered whether it is possible to obtain running by sharing resources with other non-real-time jobs. If so, the job will be run by time-sharing and multiplexing the CPU resources with other non-real-time jobs.
[0078] Optionally, during job preemption, if only CRIU is used to save and restore the job process, there will be a problem that the context state of the GPU being used by the job cannot be saved. Therefore, in this embodiment, for jobs using the GPU for CUDA computing, the CUDA proxy can be used to complete the saving and restoring of the GPU context, that is, when calling the CUDA API in the job process, instead of directly calling the CUDA library, it is called through the CUDA proxy. Among them, the above CUDA proxy can include a dynamic link library (hereinafter referred to as the proxy link library) and a proxy process. Among them, the above dynamic link library can be used to control the job process's call to the CUDA API. For the above proxy process, the proxy link library can forward the intercepted call of the job process to the CUDA API to the proxy process, and the proxy process finally uses the CUDA library to issue requests.
[0079] For example, when the job process creates a CUDA context, the proxy link library intercepts the request and then forwards it to the proxy process. The proxy process can call the CUDA library to create a CUDA context. Subsequently, when the job process calls the CUDA API, the proxy link library will intercept and forward it to the proxy process, and the proxy process will forward the API call of the job process to the CUDA library through this context. Figure 3 It is a schematic diagram of creating a process snapshot according to an embodiment of the present application, as Figure 3As shown, when the job process 301 creates a CUDA context, the proxy link library 302 intercepts this request and then forwards the request to the proxy process 303. The proxy process 303 can call the CUDA library to create a CUDA context. When the job process 304 creates a CUDA context, the proxy link library 305 intercepts this request and then forwards the request to the proxy process 303. The proxy process 303 can call the CUDA library to create a CUDA context. And the created context can be transmitted to the image processor 306.
[0080] Optionally, for jobs that use the GPU for CUDA computing, when the job process starts, the dependent CUDA dynamic link library is not the provided CUDA library but the CUDA proxy link library; and when the job process starts, a proxy process also needs to be started simultaneously.
[0081] Optionally, to save resources, only one proxy process can be started on each computing node (physical machine), and the job processes running on that node all use this proxy process. The proxy process can create an independent CUDA context for each job process it proxies.
[0082] Optionally, when taking a snapshot of the job process to save the GPU state of the job process, the scheduling system can send an instruction to the proxy process; the proxy link library can also forward the CUDA API to the proxy process. All of these require communication with the proxy process. Therefore, after the proxy process starts, it can listen on a certain port of the host, and communicate with the proxy process by sending a request to this port.
[0083] Optionally, when the scheduling system starts a job process, it can set the port number listened on by the proxy process into the environment variable of the job process. When the proxy link library forwards the CUDA API to the proxy process, it can obtain the port number listened on by the proxy process through the environment variable and send a request to this port to complete the forwarding of the CUDA API request.
[0084] In this embodiment, when the job type is a non-real-time job type and the number of schedulable resources is greater than or equal to the number of resource requirements, the scheduling system can adopt a strategy of directly allocating resources.
[0085] Optionally, the scheduling system can also reserve computing resources for non-real-time jobs to ensure that the computing resources are available when the pending jobs start. At the same time, once the pending jobs are completed or paused, the resources are immediately released for other jobs to use, so as to improve the resource turnover rate.
[0086] Optionally, during the execution of a job to be processed, if other non-real-time jobs request resources and the scheduling system detects that there are sufficient resources on the computing node to meet the requirements of other non-real-time jobs, the scheduling system can dynamically adjust the resources without affecting the execution of the job to be processed, so as to allocate additional resources to Job Q and achieve flexible reallocation of resources.
[0087] Optionally, when determining the resources to be occupied, the scheduling system can also consider the following factors: the running state and remaining workload of the job to minimize the impact on the target job. The type of resources and their impact on job performance. For example, preferentially preempt the CPU rather than the GPU because saving and restoring the GPU state may be more complex and have a greater impact on the job. The dynamic adjustment of shared resources, that is, when resources are in short supply, allow running jobs to release some resources for other jobs to use without affecting critical performance. The job recovery plan after resource preemption to ensure that the target job can resume to the state before preemption after the resources are released, so as to reduce the losses and impacts caused by job interruption. It should be noted that the above methods for determining the resources to be occupied are only for illustrative purposes, and there is no specific limitation on the methods for determining the resources to be occupied here.
[0088] In summary, through the intelligent preemption mechanism, the system can ensure that the resource requirements of critical jobs (such as real-time jobs) are met, avoid important tasks from being delayed due to insufficient resources, and achieve efficient utilization of resources. Even in a situation of highly tense resources, the scheduling system can still provide necessary resources for the job to be processed through resource preemption, reducing the waiting time of jobs and the delay in job execution.
[0089] Optionally, the above resource preemption strategy ensures the effective allocation and utilization of resources, reduces the situation where jobs fail due to insufficient resources, and thus improves the success rate of jobs. The scheduling system can make intelligent resource preemption decisions based on the job priority, resource requirements, and current resource usage situation, avoiding waste of resources and unnecessary interruption of jobs. At the same time, by reducing job waiting time and increasing the success rate, the scheduling system enhances the availability and efficiency of the cluster and optimizes the experience of users submitting and executing jobs.
[0090] In this embodiment, through the above steps, not only a resource scheduling strategy with elastic scaling is provided, but also intelligent saving and restoration of job states are realized, greatly enhancing the resource scheduling ability of the high-performance computing cluster and the continuity of job execution, and providing a more efficient, flexible, and secure computing environment for users. And through the above steps, the scheduling system can not only handle the difference between resource requirements and schedulable resources, but also ensure the smooth execution of all jobs in the cluster through the intelligent resource preemption mechanism, while achieving efficient utilization of resources and optimizing the user experience.
[0091] The processing process of the jobs to be processed of the real-time job type will be further described below.
[0092] In an exemplary embodiment, based on the comparison result and the job type, a scheduling policy is determined, including: in response to the job type being a real-time job type and the comparison result being that the number of schedulable resources is greater than or equal to the number of resource requirements, determining the resource scheduling policy as: using the allowable schedulable computing resources to execute the job to be processed.
[0093] In an exemplary embodiment, based on the comparison result and the job type, a scheduling policy is determined, including: in response to the job type being a real-time job type and the comparison result being that the number of schedulable resources is less than the number of resource requirements, determining whether there are occupiable computing resources that are allowed to be occupied among the already scheduled computing resources; in response to the existence of occupiable computing resources, determining the scheduling policy as: using the schedulable computing resources and the occupiable computing resources to execute the job to be processed.
[0094] In an exemplary embodiment, the method may further include: sending an occupancy request to the agent process according to the resource scheduling policy, where the occupancy request includes the identity information of the occupiable computing resources; obtaining the successful occupancy instruction returned by the agent process; in response to the successful occupancy instruction, saving the execution status of the job process in the occupiable computing resources.
[0095] In this embodiment, if the job type is a real-time job type and the comparison result is that the number of schedulable resources is greater than or equal to the number of resource requirements, the resource scheduling policy can be determined to directly use the schedulable computing resources to process the job to be processed. If the comparison result is that the number of schedulable resources is less than the number of resource requirements, it can be determined whether there are occupiable computing resources that are allowed to be occupied in the computing nodes. In response to the existence of occupiable computing resources that are allowed to be occupied, the scheduling policy can be determined as: using the schedulable computing resources and the occupiable computing resources to execute the job to be processed. Further, the task to be executed can be processed on the corresponding computing resources according to the scheduling policy.
[0096] Optionally, after determining the resource scheduling policy, if the scheduling policy is to use the schedulable computing resources and the occupiable computing resources to execute the job to be processed, an occupancy request can be sent to the agent process. The occupancy request can include the identity information of the occupied computing resources, obtain the successful occupancy instruction returned by the agent process, and in response to the instruction, save the execution status of the job process in the occupiable computing resources. That is, in this embodiment, before occupying the computing resources, the execution status of the job process in the computing resources can be stored, so that after the occupiable computing resources finish processing the job to be processed, the stored job process can be used to restore the processing data before the occupiable computing resources.
[0097] Optionally, for real-time jobs, after submitting a job, the scheduling system checks whether there are idle CPU and GPU resources in the computing nodes. If there are idle resources, the job can be scheduled for execution immediately. If there are not enough resources, it can be considered whether sufficient resources can be obtained by pre-empting the resources being used by non-real-time jobs. If sufficient resources can be obtained by pre-empting one or more non-real-time jobs, the non-real-time jobs can be pre-empted to release the resources, and then the submitted real-time job can be started.
[0098] Optionally, the process of pre-empting the computing resources of non-real-time jobs may include the following: The scheduling system connects to the port of the proxy process and sends a request to the proxy process. Since the proxy process can proxy multiple job processes, the identity information of the job process is carried in the sent request, and this identity information can be used to identify the process that needs to take a snapshot. After receiving the request, the proxy process checks whether there is a currently executing CUDA kernel function. If not, the proxy process directly returns success to the scheduling system. If there is a currently executing CUDA kernel function, it waits for it to finish, and after it finishes, the proxy process directly returns success to the scheduling system. During this process, the proxy program may receive CUDA requests forwarded by the proxy library, but the proxy program does not process the CUDA requests forwarded by the proxy library and temporarily stores them in memory.
[0099] Optionally, after the scheduling system requests and receives success from the proxy process, it executes the command provided by CRIU to save the state of the job process. CRIU can save the CPU state of the process and the memory of the process to a file, and usually saves it to the network file system shared by each computing node for easy recovery. Further, the scheduling system can send another request to the proxy process, requesting the proxy process to save the CUDA requests forwarded by the proxy library that are temporarily stored (if any). The proxy process saves the CUDA requests forwarded by the proxy library that are temporarily stored to a file (hereinafter referred to as the GPU snapshot file). By the above steps, the preservation of the state of the pre-empted process is completed, and at this time, the resources occupied by the pre-empted job process can be released for the pre-empting process to use.
[0100] For the processing process of real-time jobs, the embodiments of the above claims provide two different strategies to determine the job scheduling strategy based on the comparison result of the resource requirements and the number of schedulable resources. Below, we will combine the content in the technical disclosure document to explain, exemplify, and expand these steps in detail and illustrate the beneficial effects.
[0101] Optionally, when the real-time job type is determined and the number of schedulable resources is greater than or equal to the number of resource requirements, the pending job can be directly executed using the computable resources allowed for scheduling. However, when the real-time job type is determined and the number of schedulable resources is less than the number of resource requirements, first, it can be checked whether there are occupiable computable resources that are allowed to be occupied among the already scheduled computable resources. If there are, the real-time job is executed using these occupiable computable resources and the schedulable resources. However, before occupying the above-mentioned occupiable computable resources, the scheduling system can send an occupation request to the agent process, and the request can include the identity information of the computable resources preempted from the non-real-time job. This information is used to indicate to the agent process which resources need to be snapshotted and saved. Further, after receiving the request, the agent process will pause the GPU context state of the non-real-time job, wait for the current CUDA kernel function to finish execution, and then, in cooperation with the CRIU tool, create a snapshot of the non-real-time job and save its state. After completion, the agent process returns a successful occupation instruction to the scheduling system. After receiving the successful occupation instruction from the agent process, the scheduling system will use CRIU to save the execution state of the non-real-time job, including the CPU state and memory data, to ensure that the non-real-time job can be restored to the state before being preempted after the resources are occupied.
[0102] Optionally, to further improve the efficiency of resource management and scheduling, a specific resource pool can also be reserved for real-time jobs. When a real-time job is submitted, resources are preferentially allocated from this resource pool to reduce the possibility of resource preemption. At the same time, the resource preemption strategy can be dynamically adjusted according to the load condition of the cluster. For example, resource preemption can be reduced when the load is low to avoid affecting non-real-time jobs. When the resources become available again, the non-real-time jobs that have had their resources preempted can be preferentially restored to reduce the impact brought by job interruption.
[0103] In summary, through the above steps, even in the case of insufficient resources, real-time jobs can obtain sufficient resources through resource preemption, ensuring the continuity of jobs and the elastic scaling ability of cluster resources. After the non-real-time jobs have their resources preempted, their states are saved, avoiding waste of resources. At the same time, the scheduling system can reallocate the released resources, improving the overall utilization rate of resources. The above strategy ensures the priority of real-time jobs and also takes into account the continuity of non-real-time jobs, achieving a balance between fairness and efficiency in cluster scheduling.
[0104] In an exemplary embodiment, the method may further include: executing a pending job according to a resource scheduling strategy; in response to the completion of the execution of the pending job, releasing the computable resources occupied by the pending job; and resuming the historical job processes in the occupiable computable resources, where the historical job processes are used to represent the execution states of the jobs executed by the occupiable computable resources before the occupiable computable resources were occupied.
[0105] In this embodiment, when the job to be processed is completed, the computing resources occupied by the job to be processed can be released, and the historical job process in the available computing resources can be restored from the released computing resources. The historical job process can be a pre-saved job process, which can be used to characterize the execution status of the job executed before the available computing resources are occupied. The execution status can be used to determine the running result, running progress, etc. of the above-mentioned job.
[0106] Optionally, the scheduling system can execute the CRIU command to restore the preempted job process from the previously saved file.
[0107] Optionally, when a job to be processed is completed, the resources it occupies are released, and the scheduling system can perform rescheduling. At this time, it can first check whether there are real-time jobs to be scheduled. If there are no real-time jobs to be scheduled or the running conditions for real-time jobs are not met, then non-real-time jobs are scheduled. If there are sufficient resources, the preempted job process saved previously can be restored and continue to execute.
[0108] Optionally, the logic for selecting the CPU when restoring the job process is as follows: If there are sufficient idle CPUs on the node (idle CPUs refer to CPUs not used by any job), then select the computing resources for executing the job to be processed from the idle CPUs. If there are not enough idle CPUs on the computing node (which can be simply referred to as the node), then sharing the CPU with other non-real-time jobs can be considered.
[0109] Optionally, assume that the number of idle CPUs on the node plus the number of shared CPUs is m, and the number of CPUs requested by the job is n. If m is less than n, then it can be considered to allocate m CPUs for the job to run. Doing so may result in poor performance of the job, but it can avoid waste of resources.
[0110] Optionally, the job process can be restored through the following steps: If there is no proxy process running on the computing node where the job process is to be restored, the scheduling system starts a proxy process on this node. The scheduling system sends a request to the proxy process, requesting the proxy process to restore the GPU context state, and specifying the GPU snapshot file in the request. The proxy process receives the request, calls the NVIDIA CUDA API to re-establish a new CUDA context. If there are previously saved but not yet executed CUDA Kernel functions in the GPU snapshot file, then the previously saved and unexecuted CUDA Kernel functions are executed in the newly created CUDA context. The proxy process returns success to the scheduling system.
[0111] Optionally, when the real-time job type is determined and the number of schedulable resources is greater than or equal to the number of resource requirements, the pending job can be directly executed using the computable resources allowed for scheduling. However, when the real-time job type is determined and the number of schedulable resources is less than the number of resource requirements, first, it is possible to check whether there are occupiable computable resources that are allowed to be occupied among the already scheduled computable resources. If so, these occupiable computable resources and the schedulable resources are used to execute the real-time job.
[0112] For example, before occupying the above-mentioned occupiable computable resources to execute the pending job, an occupation request can be sent to the proxy process first, and the scheduling system sends an occupation request to the proxy process. This request may include the identity information of 4 CPU cores and 1 GPU resource preempted from the non-real-time job. This information is used to indicate to the proxy process which resources need to be snapshot saved. After receiving the request, the proxy process can suspend the GPU context state of the non-real-time job, wait for the current CUDA kernel function to finish execution, and then, in cooperation with the CRIU tool, create a snapshot of the above non-real-time job and save its state. After completion, the proxy process can return a successful occupation instruction to the scheduling system. Further, after receiving the successful occupation instruction from the proxy process, the scheduling system uses CRIU to save the execution state of the above non-real-time job. This execution state may include the CPU state and memory data, and can be used to ensure that the non-real-time job can be restored to the state before being preempted after the resources are occupied.
[0113] In summary, through the above steps, the resource requirements of real-time jobs can be processed more intelligently and efficiently, while maintaining the flexibility and fairness of cluster resource scheduling, providing strong support for high-performance computing scenarios.
[0114] Optionally, even in the case of resource shortage, real-time jobs can obtain sufficient resources through resource preemption, ensuring the continuity of the jobs and the elastic scaling ability of cluster resources. After the non-real-time jobs have their resources preempted, their states are saved, avoiding waste of resources. At the same time, the scheduling system can reallocate the released resources, improving the overall utilization rate of resources.
[0115] Optionally, the determined resource scheduling strategy ensures the priority of real-time jobs and also takes into account the continuity of non-real-time jobs, achieving a balance between fairness and efficiency in cluster scheduling.
[0116] Optionally, the above method can be applied to the field of high-performance computing clusters. Through the above method, it can effectively support the shared use of GPU computing power resources by multiple users in high-performance computing scenarios.
[0117] In the embodiments of the present application, after obtaining a job to be processed, the job type of the job to be processed and the required quantity of resource requirements for executing the job to be processed can be determined, the scheduled data of resources in the computing resources can be determined, and based on the scheduled data, the job type, and the quantity of resource requirements, the resource scheduling policy corresponding to the job to be processed can be determined, thereby solving the technical problem of low scheduling efficiency of computing resources and achieving the technical effect of improving the scheduling efficiency of computing resources.
[0118] To facilitate the understanding of the embodiments of the present application, relevant scenarios are now explained, but they do not limit the present application.
[0119] Currently, high-performance computing plays a relatively extensive and important role in many industries such as life sciences, manufacturing simulation, chemical engineering, aerospace, materials, and meteorology. In these fields, a large amount of computation is generally involved, and it is required to complete the computing tasks within a certain period of time. It is impossible to meet the requirements with a single general-purpose server node. Generally, multiple high-performance servers are used and interconnected through a high-speed network to form a high-performance computing cluster to process a large amount of data in parallel at an extremely high speed.
[0120] For the high-performance computing cluster job scheduling system, such as Slurm, this system maintains a queue of user job scripts to be processed and manages the overall resource utilization of this job. It manages the available computing node resources in a shared or non-shared manner for users to execute jobs. This system will reasonably allocate resources to the job queue and monitor the job until it is completed.
[0121] In the related art, it is impossible to achieve elastic scaling of the nodes or CPU resources used by the job. In the existing cluster scheduling systems, although job preemption based on priority or other conditions is allowed, it is impossible to automatically scale down the preempted job or migrate it to other host nodes to continue execution. For the processing method of the preempted job, either the preempted job is suspended or the preempted job is cancelled.
[0122] However, if the preempted job is cancelled, the results of the job that have already been executed will be lost, and the job can only be resubmitted and executed from the beginning. If the preempted job is suspended, usually the suspended job can only release the CPU resources it uses, but cannot release the memory and GPU resources it uses. Although the CPU resources are released in this way, other jobs can use the released CPU resources, but it is possible that other jobs may not fully use the released CPU resources, and some CPU resources are not used, resulting in waste. Therefore, the above methods still have the technical problem of low scheduling efficiency of computing resources.
[0123] In view of the above problems, the present invention proposes a job scheduling method based on a CUDA agent to support elastic scaling of CPU / GPU jobs, which solves the problem that in the high-performance computing scenario, the resources used by nodes cannot be elastically scaled, resulting in the inability to perform more fine-grained scheduling of resources in the cluster.
[0124] Optionally, this method can be responsible for the management and use of resources such as CPUs and GPUs in the entire cluster by a job scheduling system. The scheduling system allocates CPU and GPU resources for user jobs that require GPU resources, provides a real-time job usage plan, and supports creating snapshots and restoring jobs that use GPUs, improving the utilization rate of resources in the cluster and ensuring the priority execution of real-time jobs.
[0125] In this embodiment, jobs are divided into two categories: real-time jobs and non-real-time jobs. Non-real-time jobs are allowed to share CPU resources with each other, while real-time jobs do not share CPU with other jobs. Considering that the priority of real-time jobs is higher than that of non-real-time jobs, real-time jobs are allowed to preempt the resources used by running non-real-time jobs. Further, the state of the preempted job (including the memory, registers, etc. of the CPU and GPU used by the job process) can be saved to a file through CRIU. When there are idle CPU resources in the system or by sharing CPU resources with other non-real-time jobs, the preempted job can be rescheduled and restored from the saved file to continue execution, thus solving the technical problem of low scheduling efficiency of computing resources and achieving the technical effect of improving the scheduling efficiency of computing resources.
[0126] Optionally, this embodiment realizes the migration of GPU job processes between nodes by using a CUDA agent to save and restore the state of GPU job processes, thereby achieving the purpose of elastic scaling of jobs. And it can be combined with scheduling strategies such as preemption in existing HPC scheduling systems, improving the accuracy and flexibility of scheduling and the utilization rate of resources in the cluster, thus solving the problem that in the high-performance computing scenario, the resources used by computing nodes cannot be elastically scaled, resulting in the inability to perform more fine-grained scheduling of resources in the cluster.
[0127] Optionally, this embodiment realizes the saving and restoring of job processes, especially GPU job processes, by using a CUDA agent, solving the problems of inability to migrate job processes and elastic scaling of jobs in traditional HPC scheduling systems. It has a significant effect in multiple scenarios of HPC scheduling, which can improve the accuracy of scheduling and the utilization efficiency of resources.
[0128] Optionally, one scenario where the above method can be used is as follows: During the execution of a job process, due to a node crashing abnormally, the job cannot be restored. By using the above method, the status of the job process can be saved regularly and restored to the previously saved state after the abnormal crash. Another scenario where it can be used is as follows: After a low-priority job is preempted by a high-priority job, the traditional scheduling system cannot migrate the job process, resulting in the low-priority job being unable to continue execution. With the solution of this patent, the preempted job can be migrated to other nodes and continue execution from the previously saved state.
[0129] Optionally, in this embodiment of the scheduling system, jobs are divided into two categories: real-time jobs and non-real-time jobs. Real-time jobs are allowed to preempt non-real-time jobs. Real-time jobs exclusively use the CPU, and non-real-time jobs can share the CPU among themselves.
[0130] Optionally, in this embodiment, through the CUDA proxy (including two parts: the proxy link library and the proxy process), the preservation and restoration of the GPU status are completed. By cooperating with CRIU, the checkpoint / restore function of the job process is completed. Of course, CRIU can also be replaced with other similar software with the checkpoint / restore function.
[0131] In the solution for job elastic scaling in the high-performance computing scenario provided in this embodiment. Based on this solution, real-time jobs can be completed in the shortest possible time. By setting exclusive CPU resources for real-time jobs and allowing real-time jobs to preempt non-real-time jobs, and ensuring that real-time jobs obtain resources prior to non-real-time jobs, real-time jobs can be completed as soon as possible.
[0132] Optionally, this embodiment allows non-real-time jobs to share CPU resources, improving resource utilization and avoiding resource waste. For example, a certain non-real-time job requires 10 CPUs, but there are only 5 idle CPUs in the system. At this time, if CPU resource sharing is not allowed, the non-real-time job cannot run due to insufficient resources, and the 5 idle CPUs cannot be utilized, resulting in resource idleness and waste; while in this embodiment, CPU resource sharing is allowed, so in addition to using the 5 idle CPUs, this job can also obtain 5 shared CPUs by sharing with other non-real-time jobs, thus meeting the requirement of 10 CPUs and enabling the job to run.
[0133] Optionally, this embodiment allows the number of CPUs allocated to non-real-time jobs to be less than the number of CPUs they request, thereby improving resource utilization and avoiding resource waste. For example, a non-real-time job requests 10 CPUs, and the scheduling system allocates 10 CPUs to it. Later, 5 of the 10 CPUs are preempted by real-time jobs, and at this time only the remaining 5 CPUs are available in the system. Then we can allow the non-real-time job to continue running using the remaining 5 non-preempted CPUs.
[0134] Optionally, when submitting a non-real-time job, a parameter can also be added: the minimum number of CPUs.
[0135] In this embodiment, the problem that CRIU cannot save the GPU state is solved by the above method. Moreover, the solution for saving and restoring the GPU state adopted is not related to a specific physical GPU, that is, this embodiment can be applied to multiple different models of GPUs.
[0136] Optionally, this embodiment is transparent to the job process. That is, it is not necessary to modify the code of the job program, and there is no intrusion into the job process.
[0137] In this embodiment, since if no proxy is used, after the job process is restored, the CUDA context held in the job process is no longer available and a new CUDA context needs to be recreated. Therefore, a CUDA proxy can be used to intercept and proxy the CUDA API calls of the job process. In this embodiment, by adopting the method of proxying CUDA, the proxy maintains the CUDA context. When the proxy process restores the GPU state, it can automatically recreate a new CUDA context. For the job process, this process is transparent and is automatically completed by the proxy process, and the job process cannot perceive the change in the context before and after the restoration. When the subsequent job process calls the CUDA API, it will automatically use the CUDA context newly created by the proxy.
[0138] In this embodiment, the CUDA proxy is divided into a proxy library and a proxy process. Thus, when the proxy process listens on a certain port, the system can send requests to the proxy process through this port, and multiple job processes on the same host can use the same proxy process. Moreover, the above method can start the proxy process before restoring the job process, and the proxy process creates a new CUDA context.
[0139] In this embodiment, by combining CRIU with the solution for saving and restoring the GPU state, the scheduling system completes the orchestration of the entire process. The scheduling system can first send a request to the agent to make the agent process stop executing subsequent CUDA kernel functions after finishing the current CUDA kernel function, then execute the CRIU command to save the job process state, and then request the agent process to save the unexecuted CUDA kernel functions to a file.
[0140] In this embodiment, jobs can be divided into real-time jobs and non-real-time jobs. Real-time jobs are allowed to preempt non-real-time jobs, so as to ensure that real-time jobs can be executed prior to non-real-time jobs. And this embodiment allows non-real-time jobs to share the CPU, thus improving the utilization rate of CPU resources.
[0141] In the embodiment of the present application, for real-time jobs, when submitting a job, a priority can be assigned to each job. Jobs with higher priorities will be scheduled first during scheduling. That is, for real-time jobs, scheduling is performed in descending order of priority. First, jobs with higher priorities are scheduled. After scheduling jobs with high priorities, if there are remaining resources (or resources used by non-real-time jobs can be preempted), then jobs with lower priorities are scheduled. If there are no remaining resources and resources used by non-real-time jobs cannot be preempted, then jobs with lower priorities wait to be scheduled and executed when resources are available. For real-time jobs with the same priority, they can be scheduled and executed in the order of job submission time, i.e., the first-submitted job is scheduled first.
[0142] Optionally, in order to prevent the over-scaling ratio from being too high when non-real-time jobs share CPU resources, a maximum over-scaling ratio can be set in the scheduling system. For example, if the maximum over-scaling ratio is set to 4, then one CPU can be used by at most 4 non-real-time jobs. When the maximum over-scaling ratio is reached, when submitting a non-real-time job, it will not be executed immediately but wait for available resources.
[0143] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0144] In this embodiment, a device for determining a resource scheduling policy is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0145] Figure 4 is a structural block diagram of a device for determining a resource scheduling policy according to an embodiment of the present application. As Figure 4 shown, the device may include: an acquisition unit 42, a first determination unit 44, and a second determination unit 46.
[0146] The acquisition unit 42 is used to acquire a to-be-processed job submitted by a client to a scheduling system.
[0147] The first determination unit 44 is used to determine the job type of the to-be-processed job and the quantity of resource requirements for the to-be-processed job's required computing resources.
[0148] The second determination unit 46 is used to determine the resource scheduled data of multiple computing resources in a computing node, and based on the resource scheduled data, job type, and quantity of resource requirements, determine a resource scheduling policy corresponding to the to-be-processed job, where the resource scheduled data is used to represent the usage conditions of multiple computing resources, and the resource scheduling policy is used to represent the rules for scheduling the computing resources required for the to-be-processed job from the computing node.
[0149] Through the above device, after acquiring the to-be-processed job, the job type of the to-be-processed job and the quantity of resource requirements for executing the to-be-processed job can be determined, the resource scheduled data in the computing resources can be determined, and based on the scheduled data, job type, and quantity of resource requirements, the resource scheduling policy corresponding to the to-be-processed job can be determined, thereby solving the technical problem of low scheduling efficiency of computing resources and achieving the technical effect of improving the scheduling efficiency of computing resources.
[0150] It should be noted that the above-mentioned respective modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: all the above modules are located in the same processor; or, the above-mentioned respective modules are separately located in different processors in any combination form.
[0151] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0152] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0153] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.
[0154] Optionally, Figure 5 is a block diagram of a computer system structure of an electronic device according to an embodiment of the present application. As Figure 5 shown, the computer system 500 includes a central processing unit 501 (CPU), which can perform various appropriate actions and processes according to a program stored in the read-only memory 502 (ROM) or a program loaded from the storage section 508 into the random access memory 503 (RAM). In the random access memory 503, various programs and data required for system operation are also stored. The central processing unit 501, the read-only memory 502, and the random access memory 503 are connected to each other via a bus 504. The input / output interface 505 (Input / Output interface, i.e., I / O interface) is also connected to the bus 504.
[0155] The following components are connected to the input / output interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a local area network card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disc, a magneto-optical disc, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read from it can be installed into the storage section 508 as needed.
[0156] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0157] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program. When the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0158] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0159] An embodiment of the present application also provides a computer program. The computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.
[0160] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0161] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0162] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for determining a resource scheduling strategy, characterized in that: include: Get pending jobs submitted by the client to the scheduling system; Determine the job type of the job to be processed and the resource requirement quantity of the computing resources required for the job to be processed; Determine resource scheduled data of multiple computing resources in a computing node, and determine a resource scheduling policy corresponding to the job to be processed based on the resource scheduled data, the job type and the resource demand quantity, wherein the resource scheduled data is used to represent the usage of the multiple computing resources, and the resource scheduling policy is used to represent the rules for scheduling the computing resources required for the job to be processed from the computing node.
2. The method according to claim 1, characterized in that The determining, based on the resource scheduled data, the job type and the resource demand quantity, a resource scheduling strategy corresponding to the job to be processed includes: Determine the number of schedulable resources in the computing node based on the resource scheduled data, wherein the number of schedulable resources is used to represent the number of computing resources that are allowed to be scheduled among the plurality of computing resources; The resource scheduling strategy is determined based on the job type, the resource requirement quantity and the schedulable resource quantity.
3. The method according to claim 2, characterized in that The determining the resource scheduling strategy based on the job type, the resource requirement quantity and the schedulable resource quantity includes: Comparing the resource demand and the amount of schedulable resources to obtain a comparison result; The scheduling strategy is determined based on the comparison result and the job type.
4. The method according to claim 3, characterized in that: The determining the scheduling strategy based on the comparison result and the job type includes: In response to the job type being a non-real-time job type and the comparison result being that the number of schedulable resources is greater than or equal to the number of resource requirements, the resource scheduling strategy is determined to be: executing the job to be processed using the computing resources that are allowed to be scheduled.
5. The method according to claim 4, characterized in that The determining the scheduling strategy based on the comparison result and the job type includes: In response to the job type being the non-real-time job type and the comparison result being that the number of schedulable resources is less than the number of resource requirements, determining whether there is a shared resource in the computing node; In response to the existence of the shared resource in the computing node, determining the sum of the shared resource and the schedulable resource quantity, and in response to the sum of the shared resource and the schedulable resource quantity being greater than or equal to the resource demand quantity, determining the resource scheduling strategy to be: executing the to-be-processed job using the computing resource and the shared resource that are allowed to be scheduled; In response to the shared resource not existing in the computing node, or the sum of the shared resource and the schedulable resource quantity is less than the resource demand quantity, determining the resource to be occupied; Based on the number of the resources to be occupied, the resource scheduling strategy is determined as: using the schedulable computing resources, the shared resources and the resources to be occupied to execute the job to be processed.
6. The method according to claim 5, characterized in that The determining of the resources to be occupied includes: Determining at least one job being processed in a plurality of the computing resources based on the resource scheduled data; Determine at least one target job from at least one of the jobs, wherein the priority of the target job is lower than the priority of the job to be processed; The computing resources corresponding to the target job are determined as the resources to be occupied.
7. The method according to claim 6, characterized in that The method further comprises: In response to the computing resources of the job being occupied by the to-be-processed job, a process snapshot is created for the target job, wherein the process snapshot is used to save the execution state of the computing resources before being occupied.
8. The method according to claim 3, characterized in that The determining the scheduling strategy based on the comparison result and the job type includes: In response to the job type being a real-time job type and the comparison result being that the number of schedulable resources is greater than or equal to the number of resource requirements, the resource scheduling strategy is determined to be: executing the job to be processed using the computing resources that are allowed to be scheduled.
9. The method according to claim 8, characterized in that The determining the scheduling strategy based on the comparison result and the job type includes: In response to the job type being a real-time job type and the comparison result being that the number of schedulable resources is less than the number of resource requirements, determining whether there are any occupiable computing resources that are allowed to be occupied among the scheduled computing resources; In response to the existence of the available computing resources, the scheduling strategy is determined as: using the schedulable computing resources and the available computing resources to execute the job to be processed.
10. The method according to claim 9, characterized in that The method further comprises: According to the resource scheduling strategy, an occupation request is sent to the proxy process, wherein the occupation request includes the identity information of the computing resource that can be occupied; Obtaining a successful occupation instruction returned by the proxy process; In response to the successful occupation instruction, the execution state of the job process in the occupyable computing resource is saved.
11. The method according to claim 1, characterized in that: The method further comprises: Execute the pending job according to the resource scheduling strategy; In response to the completion of the execution of the pending job, releasing the computing resources occupied by the pending job; Restoring a historical job process in the occupyable computing resource, wherein the historical job process is used to represent an execution state of a job executed by the occupyable computing resource before occupying the occupyable computing resource.
12. A device for determining a resource scheduling strategy, characterized in that: include, An acquisition unit, used to acquire pending jobs submitted by the client to the scheduling system; A first determining unit, configured to determine a job type of the job to be processed and a resource requirement quantity of computing resources required for the job to be processed; The second determination unit is used to determine the resource scheduled data of multiple computing resources in the computing node, and based on the resource scheduled data, the job type and the resource demand quantity, determine the resource scheduling policy corresponding to the job to be processed, wherein the resource scheduled data is used to represent the usage of the multiple computing resources, and the resource scheduling policy is used to represent the rules for scheduling the computing resources required for the job to be processed from the computing node.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 11 when executed by a processor.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 11 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
DCU-based resource scheduling method and device, and computer equipment
CN112612600A
Resource scheduling method and device
CN114579302A
Resource scheduling method, storage medium and electronic equipment
CN116456496A
Cluster resource scheduling method and device, equipment and medium
CN116643890A
Resource scheduling method and device, electronic equipment and storage medium
CN117992207A
Cited By
Data processing method and device, electronic equipment, storage medium and program product
CN121597461A