Job scheduling method and device, electronic equipment and storage medium
By determining the target computing cluster based on metadata and changing the job submission method in the job scheduling method of the computing center, and combining a two-level scheduling queue and dynamic strategy, the problem of the inability to globally manage job scheduling of heterogeneous computing clusters is solved, and efficient job processing is achieved.
Patent Information
- Application Number
- CN202511302064.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-12
AI Technical Summary
In the existing technology, the job scheduling of the heterogeneous computing clusters in the computing power center cannot be managed globally, resulting in low job processing efficiency.
A job scheduling method is provided, which obtains computing jobs to be processed from multiple scheduling queues, determines the target computing cluster from multiple computing clusters based on metadata, converts the jobs into computing jobs that match the job submission method of the target computing cluster, distributes them to the target computing cluster for execution, and releases resources when the jobs are completed. By combining two-level scheduling queues and dynamic scheduling strategies, joint scheduling of different cluster types can be achieved.
It improves job processing efficiency, enables joint scheduling of computing clusters of different cluster types, and improves the processing efficiency of computing jobs.
Smart Images

Figure CN120803749A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and in particular to a job scheduling method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the rapid development of big data and artificial intelligence technologies, computing power resources have become the core resources supporting complex computing tasks and intelligent applications. As a key computing hub, a computing power center integrates large-scale servers, GPU clusters, and high-speed interconnection components, and provides a unified platform for massive data processing and large model training, and has become the core of the infrastructure in the intelligent era. At present, the heterogeneous computing clusters of the computing power center mainly include two types of clusters: high-performance computing clusters (such as Slurm computing clusters) for high-performance computing and containerized clusters (such as Kubernetes computing clusters).
[0003] However, in the related art, the job scheduling of computing clusters of different cluster types is independent of each other, so that global scheduling management cannot be performed, resulting in low job processing efficiency. SUMMARY
[0004] To solve the problems of the prior art, the embodiments of the present application provide a job scheduling method, device, electronic equipment and storage medium. The technical solution is as follows: On the one hand, a job scheduling method is provided, the method comprising: obtaining a to-be-processed computing job from a first scheduling queue in a plurality of scheduling queues; the plurality of scheduling queues comprising the first scheduling queue and a second scheduling queue, the first scheduling queue being used to store computing jobs that can be scheduled, and the second scheduling queue being used to store computing jobs that are temporarily not schedulable; determining a target computing cluster for processing the to-be-processed computing job from a plurality of computing clusters based on metadata of the to-be-processed computing job; the plurality of computing clusters comprising heterogeneous computing power clusters of different cluster types; converting the to-be-processed computing job into a target computing job that is consistent with a job submission mode of the target computing cluster, and distributing the target computing job to the target computing cluster, so that the target computing cluster executes the target computing job, and releases heterogeneous computing power resources occupied by executing the target computing job when the execution of the target computing job ends; In response to a resource state update triggered based on the release of the heterogeneous computing power resources, performing scheduling state evaluation on a computing job in the second scheduling queue, and when the result of the scheduling state evaluation indicates that the computing job is schedulable, transferring the computing job to the first scheduling queue.
[0005] In some example embodiments, the metadata of the to-be-processed computing job comprises a cluster identifier; and determining, based on the metadata of the to-be-processed computing job, a target computing cluster for processing the to-be-processed computing job from a plurality of computing clusters comprises: obtaining resource pool configuration information, the resource pool configuration information comprising configuration files of the computing clusters in the plurality of computing clusters; determining, based on the resource pool configuration information, a target configuration file matching the cluster identifier; determining, based on the target configuration file, a target computing cluster for processing the to-be-processed computing job from the plurality of computing clusters.
[0006] In some example embodiments, the scheduling state evaluation on the computing jobs in the second scheduling queue comprises: obtaining updated current resource state information, the current resource state information indicating current resource states of the computing clusters in the plurality of computing clusters; taking a computing job at a head of the second scheduling queue as a to-be-evaluated computing job, determining, based on metadata of the to-be-evaluated computing job, a first computing cluster for processing the to-be-evaluated computing job and an expected resource state required for executing the to-be-evaluated computing job; in a case where the current resource state of the first computing cluster satisfies the expected resource state required for executing the to-be-evaluated computing job, determining that a result of the scheduling state evaluation is schedulable; in a case where the current resource state of the first computing cluster does not satisfy the expected resource state required for executing the to-be-evaluated computing job, determining that the result of the scheduling state evaluation is temporarily non-schedulable.
[0007] In some example embodiments, the method further comprises: obtaining a first computing job submitted by a target account based on a job configuration interface; determining, based on metadata of the first computing job, a second computing cluster for processing the first computing job from the plurality of computing clusters and an expected resource state required for executing the first computing job; determining whether the first computing job is currently schedulable according to a matching condition between a current resource state of the second computing cluster and the expected resource state required for executing the first computing job; in a case where the first computing job is currently schedulable, storing the first computing job to the first scheduling queue; and in a case where the first computing job is currently non-schedulable, storing the first computing job to the second scheduling queue.
[0008] In some example embodiments, the method further comprises: obtaining, based on a correspondence between the account and the computing cluster, a current resource state of the computing cluster corresponding to the target account from the current resource state information; presenting the current resource state of the computing cluster corresponding to the target account to the target account.
[0009] In some example embodiments, the converting the to-be-processed computing job into a target computing job consistent with a job submission mode of the target computing cluster and distributing the target computing job to the target computing cluster comprises: in a case where a cluster type of the target computing cluster is a containerized cluster, encapsulating the to-be-processed computing job as a workflow of the containerized cluster and writing the workflow into a local queue of the target computing cluster; in a case where the cluster type of the target computing cluster is a high-performance computing cluster, converting the to-be-processed computing job into a target command script for submitting a batch job and sending the target command script to the target computing cluster through an application programming interface of the high-performance computing cluster.
[0010] In another aspect, a job scheduling apparatus is provided, the apparatus comprising: a to-be-processed job obtaining module configured to obtain a to-be-processed computing job from a first scheduling queue in a plurality of scheduling queues; the plurality of scheduling queues comprising the first scheduling queue and a second scheduling queue, the first scheduling queue being configured to store computing jobs that can be scheduled, and the second scheduling queue being configured to store computing jobs that are temporarily unable to be scheduled; a target cluster determining module configured to determine, based on metadata of the to-be-processed computing job, a target computing cluster for processing the to-be-processed computing job from a plurality of computing clusters; the plurality of computing clusters comprising heterogeneous computing clusters of different cluster types; a job distributing module configured to convert the to-be-processed computing job into a target computing job consistent with a job submission mode of the target computing cluster and distribute the target computing job to the target computing cluster, so that the target computing cluster executes the target computing job and releases heterogeneous computing resources occupied by the execution of the target computing job when the execution of the target computing job ends; an activating module configured to, in response to a resource state update triggered based on the release of the heterogeneous computing resources, perform a scheduling state evaluation on a computing job in the second scheduling queue, and when the result of the scheduling state evaluation indicates that the computing job can be scheduled, transfer the computing job to the first scheduling queue.
[0011] In some example embodiments, the metadata of the to-be-processed computing job comprises a cluster identifier; and the target cluster determining module comprises: a resource pool configuration obtaining module, configured to obtain resource pool configuration information, the resource pool configuration information comprising configuration files of each computing cluster in the plurality of computing clusters; a target configuration file determining module, configured to determine, based on the resource pool configuration information, a target configuration file matching the cluster identifier; a target cluster determining sub-module, configured to determine, based on the target configuration file, a target computing cluster for processing the to-be-processed computing job from the plurality of computing clusters.
[0012] In some example embodiments, the activating module comprises: a current resource state obtaining module, configured to obtain updated current resource state information of each computing cluster in the plurality of computing clusters, the current resource state information indicating a current resource state of each computing cluster; a first determining module, configured to determine, based on metadata of a computing job at the head of the second scheduling queue, a first computing cluster for processing the computing job and an expected resource state required for executing the computing job; a first evaluating module, configured to determine that the result of the scheduling state evaluation is schedulable in a case where the current resource state of the first computing cluster satisfies the expected resource state required for executing the computing job; a second evaluating module, configured to determine that the result of the scheduling state evaluation is temporarily non-schedulable in a case where the current resource state of the first computing cluster does not satisfy the expected resource state required for executing the computing job.
[0013] In some example embodiments, the apparatus further comprises: a submitted job obtaining module, configured to obtain a first computing job submitted by a target account based on a job configuration interface; a second determining module, configured to determine, based on metadata of the first computing job, a second computing cluster for processing the first computing job from the plurality of computing clusters and an expected resource state required for executing the first computing job; a third determining module, configured to determine whether the first computing job is currently schedulable according to a matching condition between the current resource state of the second computing cluster and the expected resource state required for executing the first computing job; The job cache module is configured to store the first computing job to the first scheduling queue if the first computing job is currently schedulable, and store the first computing job to the second scheduling queue if the first computing job is currently not schedulable.
[0014] In some example embodiments, the apparatus further comprises: The account resource state acquisition module is configured to acquire, from the current resource state information, a current resource state of a computing cluster corresponding to the target account based on a correspondence between the account and the computing cluster. The resource state display module is configured to display the current resource state of the computing cluster corresponding to the target account to the target account.
[0015] In some example embodiments, the job distribution module is specifically configured to: in a case where the cluster type of the target computing cluster is a containerized cluster, encapsulate the to-be-processed computing job as a workflow of the containerized cluster, and write the workflow into a local queue of the target computing cluster; in a case where the cluster type of the target computing cluster is a high-performance computing cluster, convert the to-be-processed computing job into a target command script for submitting a batch processing job, and send the target command script to the target computing cluster through an application programming interface of the high-performance computing cluster.
[0016] In another aspect, an electronic device is provided, which includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the job scheduling method of any of the above aspects.
[0017] In another aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by a processor to implement the job scheduling method of any of the above aspects.
[0018] In another aspect, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the job scheduling method of any of the above aspects.
[0019] The embodiment of the present application obtains a to-be-processed computing job from a first scheduling queue in a plurality of scheduling queues, the plurality of scheduling queues including the first scheduling queue and a second scheduling queue, the first scheduling queue being used to store computing jobs that can be scheduled, and the second scheduling queue being used to store computing jobs that are temporarily unschedulable, determines a target computing cluster for processing the to-be-processed computing job from a plurality of computing clusters based on metadata of the to-be-processed computing job, the plurality of computing clusters including heterogeneous computing clusters of different cluster types, converts the to-be-processed computing job into a target computing job that is consistent with a job submission mode of the target computing cluster, and distributes the target computing job to the target computing cluster, so that the target computing cluster executes the target computing job, and releases heterogeneous computing resources occupied by the target computing cluster when the execution of the target computing job ends. At the same time, the computing job in the second scheduling queue is subjected to scheduling state evaluation in response to a resource state update triggered by the release of the heterogeneous computing resources, and when the result of the scheduling state evaluation indicates that the computing job is schedulable, the computing job is transferred to the first scheduling queue. Thus, based on the unified management of computing clusters of different cluster types, combined with two-level scheduling buffering and a dynamic scheduling strategy, the computing job is scheduled to computing clusters of different cluster types, the joint scheduling of computing clusters of different cluster types is realized, and the job processing efficiency is greatly improved. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0021] Figure 1 is a system architecture diagram of joint scheduling provided by an embodiment of the present application; Figure 2 is a flowchart of a job scheduling method provided by an embodiment of the present application; Figure 3 is a flowchart of another job scheduling method provided by an embodiment of the present application; Figure 4 is a structural block diagram of a job scheduling device provided by an embodiment of the present application; Figure 5 is a hardware structural block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0022] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0023] It should be noted that the terms "first", "second" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units need not be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0024] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as processing circuitry or memory) or a combination thereof. Similarly, one processor (or multiple processors or memory) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0025] It can be understood that in the specific embodiments of the present application, data related to user information and the like is involved, and when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions.
[0026] Please refer to Figure 1 which is a system architecture diagram of a joint scheduling system provided by an embodiment of the present application, including a joint scheduling node 110 and at least one computing power center 120. The joint scheduling node 110 and each computing power center 120 can communicate through network connection.
[0027] The computing clusters of the computing center 120 can include different cluster types, where the computing clusters can be heterogeneous computing clusters, and the heterogeneous computing power in the heterogeneous computing clusters can include, but is not limited to, a graphics processing unit (GPU), a neural network processing unit (NPU), a tensor processing unit (TPU), an FPGA (Field Programmable Gate Array) hardware, and the like.
[0028] Different cluster types can include a containerized cluster, such as a Kubernetes cluster as shown in Figure 1 , and a high-performance computing cluster for high-performance computing tasks, such as a Slurm cluster as shown in Figure 1 . The Kubernetes cluster is widely used in cloud-native scenarios, is good at automatic deployment, elastic scaling, and service orchestration of containerized applications, and has good scalability and ecological compatibility. The Slurm cluster supports large-scale batch job scheduling, resource reservation, and queue management, and is suitable for scenarios with high demand for computing resources, such as AI training and scientific computing.
[0029] The joint scheduling node 110 uniformly manages each computing cluster of at least one computing center 120 and can realize joint scheduling of computing clusters of different cluster types, so that each computing cluster of at least one computing center 120 constitutes a resource pool of the joint scheduling node 110. Specifically, the joint scheduling node 110 can synchronize configuration files of each computing cluster through a backend API (application programming interface) to generate resource pool configuration information. For example, for a Kubernetes cluster, the cluster is managed through a KubeConfig configuration file, and fields in the KubeConfig configuration file include: cluster name, cluster availability zone, platform type of the cluster, state of the cluster, server entry of the cluster, and user verification Token; for a Slurm cluster, the cluster is managed through a set SlurmConfig configuration file, and fields in the SlurmConfig configuration file include: cluster name, platform type of the cluster, state of the cluster, server entry of the cluster, username, and user verification Token.
[0030] The joint scheduling node 110 can synchronize resource states of each computing cluster in real time to obtain current resource state information, for example, each computing cluster can be monitored based on Prometheus, and resource state information of each computing cluster, including CPU utilization, memory occupancy, disk I / O occupancy, network traffic information, and GPU utilization, can be collected in real time.
[0031] In the embodiments of the present application, the joint scheduling node 110 is provided with a first scheduling queue and a second scheduling queue, wherein the first scheduling queue and the second scheduling queue are respectively used to store computing jobs in different scheduling states, the first scheduling queue is used to store computing jobs that can be scheduled, and the second scheduling queue is used to store computing jobs that cannot be temporarily scheduled. The joint scheduling node 110 obtains a to-be-processed computing job from the first scheduling queue, determines a target computing cluster for processing the to-be-processed computing job from a plurality of computing clusters based on the metadata of the to-be-processed computing job, converts the to-be-processed computing job into a target computing job consistent with the job submission mode of the target computing cluster, and distributes the target computing job to the target computing cluster, so that the target computing cluster executes the target computing job, and releases the heterogeneous computing power resources occupied by executing the target computing job when the execution of the target computing job ends, so that the joint scheduling node 110 responds to the resource state update triggered based on the release of the above-mentioned heterogeneous computing power resources, performs scheduling state evaluation on the computing job in the second scheduling queue, and when the result of the scheduling state evaluation indicates that the computing job can be scheduled, the computing job is transferred to the first scheduling queue, thereby realizing two-level scheduling buffering and dynamic scheduling strategy through the first scheduling queue and the second scheduling queue, realizing joint scheduling of computing clusters of different cluster types, and greatly improving the job processing efficiency.
[0032] In a specific application scenario, the joint scheduling node 110 shows a job configuration interface to the user in the form of a Web interface, the user can configure various job parameters of the required computing job through the job configuration interface, and submit a computing job generated based on the configured various job parameters through the job configuration interface, so that the joint scheduling node 110 obtains the submitted computing job, and combines the first scheduling queue and the second scheduling queue to perform job scheduling by using two-level scheduling buffering and dynamic scheduling strategy, which will be described in detail in the subsequent content of the embodiments of the present application.
[0033] It should be noted that the nodes / servers involved in the embodiments of the present application can be independent physical servers, or server clusters or distributed systems composed of multiple physical servers, or cloud servers providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network, content distribution network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal devices involved in the embodiments of the present application include but are not limited to mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle-mounted terminals, etc.
[0034] In an example embodiment, each computing cluster in the joint scheduling node 110 and the computing power center 120 can be a node device in a blockchain system, capable of sharing the obtained and generated information to other node devices in the blockchain system, realizing information sharing between multiple node devices. The multiple node devices in the blockchain system can be configured with the same blockchain, which is composed of multiple blocks, and the adjacent blocks have a correlation relationship, so that when the data in any block is tampered with, it can be detected through the next block, thereby avoiding the data in the blockchain from being tampered with, and ensuring the security and reliability of the data in the blockchain.
[0035] Referring to Figure 2 , which shows a flowchart of a job scheduling method provided by an embodiment of the application. The method can be applied to Figure 1 the joint scheduling node. It should be noted that the present specification provides method operation steps as described in the embodiments or flowcharts, but more or fewer operation steps can be included based on conventional or non-creative labor. The order of steps listed in the embodiments is only one of the many execution orders, and does not represent the only execution order. In actual system or product execution, the method order shown in the embodiments or the drawings can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically, as shown in Figure 2 , the job scheduling method of the embodiments of the application can include: S201, obtaining a to-be-processed computing job from a first scheduling queue in a plurality of scheduling queues.
[0036] The plurality of scheduling queues include the first scheduling queue and a second scheduling queue, the first scheduling queue is used to store schedulable computing jobs, and the second scheduling queue is used to store temporarily unschedulable computing jobs.
[0037] Specifically, for the obtained computing job, the joint scheduling node first determines whether the current resource state of the managed multiple computing clusters meets the demand of the computing job, and if so, determines that the scheduling state of the computing job is currently schedulable, and stores it to the first scheduling queue, otherwise, if not, determines that the scheduling state of the computing job is currently unschedulable, and stores it to the second scheduling queue.
[0038] The to-be-processed computing job can be a computing job at the head of the first scheduling queue.
[0039] Based on this, in some example embodiments, as shown in Figure 3 , another flowchart of a job scheduling method, the method can further include: S301, obtaining a first computing job submitted by a target account based on a job configuration interface.
[0040] The target account can be any computing resource tenant. The job configuration interface is used to configure job parameters of a required computing job, and then the required computing job can be generated through the configured job parameters. Specifically, the job parameters can include but are not limited to cluster identifier, expected resource state, job type, etc. The job type can include container instance, network mapping, and distributed job, etc.
[0041] In a specific implementation, the terminal device can display the job configuration interface to the target account. The terminal device can generate a corresponding computing job, referred to as a first computing job, based on the job parameters configured in the job configuration interface in response to a job submission instruction triggered based on the job configuration interface, and send the first computing job to the joint scheduling node. Thus, the joint scheduling node obtains the first computing job submitted by the target account based on the job configuration interface.
[0042] S303, determining a second computing cluster for processing the first computing job from the plurality of computing clusters based on the metadata of the first computing job, and an expected resource state required for executing the first computing job.
[0043] Specifically, the metadata of the first computing job is used to describe the first computing job. The metadata can include job basic information and resource requirement information, which can be configured by the target account based on the job configuration interface. The job basic information can include job identifier, specified cluster identifier, job type, etc. The resource requirement information can include the expected resource state, which can represent the computing resource required for executing the job. The joint scheduling node can obtain the specified cluster identifier and the expected resource state in the metadata of the first computing job by analyzing the metadata.
[0044] In the embodiments of the present application, the joint scheduling node can maintain resource pool configuration information, which includes configuration files of each computing cluster in the plurality of computing clusters of the computing center, such as KubeConfig configuration file and SlurmConfig configuration file. Each configuration file includes the cluster identifier of the corresponding cluster. It can be understood that each configuration file can also include other necessary information, such as cluster availability zone, platform type of the cluster, state of the cluster, server entrance of the cluster, and user verification Token in the KubeConfig configuration file. The platform type of the cluster, the state of the cluster, the server entrance of the cluster, the username, and the user verification Token in the SlurmConfig configuration file.
[0045] In a specific implementation, the joint scheduling node determines the second computing cluster used for processing the first computing job based on the cluster identifier specified in the metadata of the first computing job and a target configuration file matched with the cluster identifier in the resource pool configuration information.
[0046] S305, determining whether the first computing job is currently schedulable based on matching between the current resource state of the second computing cluster and the expected resource state required for executing the first computing job.
[0047] In the embodiments of the present application, the joint scheduling node can collect the current resource state of each computing cluster in the multiple computing clusters of the computing power center in real time, and then obtain the current resource state information of the resource pool. It can be understood that the current resource state information of the resource pool can be updated with real-time collected data. The current resource state of each computing cluster can represent the current usage of computing power resources in the corresponding computing cluster, for example, can include CPU utilization, memory occupancy, disk I / O occupancy, network traffic information, and GPU utilization, etc.
[0048] In a specific implementation, the joint scheduling node can obtain the current resource state of the second computing cluster from the current resource state information, and compare the current resource state of the second computing cluster with the expected resource state obtained by analyzing the metadata of the first computing job. If the current resource state of the second computing cluster meets the expected resource state required by the first computing job, it is determined that the scheduling state of the first computing job is currently schedulable. Otherwise, if the current resource state of the second computing cluster does not meet the expected resource state required by the first computing job, it is determined that the scheduling state of the first computing job is currently not schedulable.
[0049] S307, in the case that the first computing job is currently schedulable, storing the first computing job to the first scheduling queue; in the case that the first computing job is currently not schedulable, storing the first computing job to the second scheduling queue.
[0050] The above-mentioned embodiments determine the scheduling state of the first computing job submitted by the target account based on the job configuration interface according to the demand of the first computing job and the comprehensive analysis of the resource state of the computing power cluster, and then store it in the scheduling queue matched with its scheduling state, so as to perform dynamic scheduling based on the scheduling state subsequently, which is beneficial to improve the job processing efficiency of the computing job.
[0051] S203, determining a target computing cluster used for processing the to-be-processed computing job from multiple computing clusters based on the metadata of the to-be-processed computing job, wherein the multiple computing clusters include heterogeneous computing power clusters of different cluster types.
[0052] Specifically, the metadata of the to-be-processed computing job for describing the to-be-processed computing job can include a specified cluster identifier, and the step S203 can include, when implemented, the following steps: obtaining resource pool configuration information, the resource pool configuration information including configuration files of each computing cluster in the plurality of computing clusters; determining a target configuration file matched with the cluster identifier based on the resource pool configuration information; and determining a target computing cluster for processing the to-be-processed computing job from the plurality of computing clusters based on the target configuration file, so that the plurality of computing clusters are uniformly managed based on the resource pool configuration information, and the target computing cluster for processing the to-be-processed computing job is accurately determined based on the resource pool configuration information.
[0053] S205, converting the to-be-processed computing job into a target computing job consistent with a job submission mode of the target computing cluster, and distributing the target computing job to the target computing cluster, so that the target computing cluster executes the target computing job, and releases the heterogeneous computing resource occupied by the target computing job when the execution of the target computing job ends.
[0054] Specifically, since the job submission modes of computing clusters of different cluster types are different, after the target computing cluster for processing the to-be-processed computing job is determined, the to-be-processed computing job needs to be converted into a target computing job consistent with the job submission mode of the target computing cluster, so as to realize the joint scheduling of computing clusters of different cluster types.
[0055] In some exemplary embodiments, the cluster types can include containerized clusters (such as Kubernetes clusters) and high-performance computing clusters (such as Slurm clusters), and the step S205 can include, when implemented: in the case where the cluster type of the target computing cluster is a containerized cluster, encapsulating the to-be-processed computing job as a workflow of the containerized cluster, and writing the workflow into a local queue of the target computing cluster; in the case where the cluster type of the target computing cluster is a high-performance computing cluster, converting the to-be-processed computing job into a target command script for submitting a batch job, and sending the target command script to the target computing cluster through an application programming interface of the high-performance computing cluster.
[0056] In a specific implementation, if the cluster type of the target computing cluster is a Kubernetes cluster, the federated scheduling node can encapsulate the target computing job as a WorkLoad of the Kubernetes cluster through a Kueue submodule, and write the WorkLoad into a LocalQueue of the target computing cluster, so that the WorkLoad can perform WorkLoad admission, resource allocation, POD creation, and resource backfill operations under the unified quota constraints of the ClusterQueue of the target computing cluster, to execute the target computing job, and release the heterogeneous computing resources occupied during execution when the target computing job is completed.
[0057] If the cluster type of the target computing cluster is a Slurm cluster, the federated scheduling node can convert the target computing job into a Sbatch script, and distribute the Sbatch script to the target computing cluster through a REST backend API interface of the Slurm cluster, to execute the target computing job by the target computing cluster. When the Slurm cluster executes a computing job, the computing job usually enters a PENDING queue in the Slurm cluster and is scheduled for resources by a Slurm Scheduler based on a MultiFactor priority, and when the system resources of the Slurm cluster meet the requirements, the state of the computing job is converted to RUNNING, and when the system resources of the Slurm cluster cannot meet the requirements of the computing job, the computing job enters a BackFill window for waiting.
[0058] S207, in response to the resource state update triggered based on the release of the heterogeneous computing resources, performing scheduling state evaluation on the computing job in the second scheduling queue, and when the result of the scheduling state evaluation indicates that the computing job can be scheduled, transferring the computing job to the first scheduling queue.
[0059] Specifically, when the target computing cluster completes the execution of the target computing job and releases the heterogeneous computing resources occupied during the execution of the target computing job, the federated scheduling node can monitor the change of the resource state and update the current resource state information of the resource pool, and then perform scheduling state evaluation on the computing job in the second scheduling queue in response to the resource state update triggered based on the release of the heterogeneous computing resources, and the result of the scheduling state evaluation indicates whether the computing job can be scheduled, and when the result of the scheduling state evaluation indicates that the computing job can be scheduled, the federated scheduling node transfers the computing job in the second scheduling queue to the first scheduling queue, and specifically, the federated scheduling node can transfer the computing job to the tail of the first scheduling queue.
[0060] In a specific embodiment, the step S207 can include the following when performing the scheduling state evaluation on the computing jobs in the second scheduling queue: obtaining the updated current resource state information, the current resource state information indicating the current resource state of each computing cluster in the plurality of computing clusters; taking the computing job at the head of the second scheduling queue as a to-be-evaluated computing job, determining a first computing cluster for processing the to-be-evaluated computing job based on the metadata of the to-be-evaluated computing job, and determining the expected resource state required for executing the to-be-evaluated computing job; in a case where the current resource state of the first computing cluster meets the expected resource state required for executing the to-be-evaluated computing job, determining that the result of the scheduling state evaluation is schedulable; in a case where the current resource state of the first computing cluster does not meet the expected resource state required for executing the to-be-evaluated computing job, determining that the result of the scheduling state evaluation is temporarily unschedulable. Thus, when the heterogeneous computing power resource is released after the target computing job is executed, the scheduling state update of the second scheduling queue can be activated in time, and the admission condition of the computing job at the head of the second scheduling queue is re-evaluated, thereby realizing the dynamic heterogeneous computing power scheduling based on the two-level scheduling buffer, and improving the job processing efficiency.
[0061] In actual application, after executing the target computing job, the target computing cluster can return the corresponding computing job execution result to the joint scheduling node, and can display the computing job execution result to the account corresponding to the target computing job through the front-end web page.
[0062] In some exemplary embodiments, the joint scheduling node can store the correspondence between the accounts and the computing clusters in the resource pool, which can represent the resource group owned by the tenant of the corresponding account, and the resource group can include one or more computing clusters in the resource pool. Further, the job scheduling method of the embodiments of the present application can further include: obtaining the current resource state of the computing cluster corresponding to the target account from the current resource state information of the resource pool; and displaying the current resource state of the computing cluster corresponding to the target account to the target account. Specifically, the target account is any tenant, and the joint scheduling node can visually display the current resource state of the resource group owned by each tenant through the web interface of the front-end based on the current resource state information of the resource pool, thereby realizing the joint management of computing clusters of different cluster types.
[0063] From the above technical solutions of the embodiments of the present application, it can be seen that the embodiments of the present application can cooperatively manage the computing clusters of different cluster types (such as Kubernetes clusters and Slurm clusters) between the computing power centers, realize unified management and joint scheduling of heterogeneous computing power resources, and based on the first scheduling queue and the second scheduling queue corresponding to different scheduling states, adopt a two-level scheduling buffer and a dynamic scheduling strategy to schedule the computing jobs of users to the computing clusters of different cluster types, thereby greatly improving the job processing efficiency while realizing joint scheduling of the computing clusters of different cluster types.
[0064] Corresponding to the job scheduling method provided by the above several embodiments, the embodiments of the present application also provide a job scheduling device. Since the job scheduling device provided by the embodiments of the present application corresponds to the job scheduling method provided by the above several embodiments, the implementation manner of the foregoing job scheduling method is also applicable to the job scheduling device provided by the embodiments of the present application, which will not be described in detail in the embodiments.
[0065] Please refer to Figure 4 which is a structural schematic diagram of a job scheduling device provided by the embodiments of the present application. The device has the function of realizing the job scheduling method in the above method embodiments. The function can be realized by hardware, or the corresponding software can be executed by hardware. As Figure 4 shown, the job scheduling device 400 can include: a to-be-processed job obtaining module 410, configured to obtain a to-be-processed computing job from a first scheduling queue in a plurality of scheduling queues; the plurality of scheduling queues include the first scheduling queue and a second scheduling queue, the first scheduling queue is used to store a computing job that can be scheduled, and the second scheduling queue is used to store a computing job that is temporarily not schedulable; a target cluster determining module 420, configured to determine a target computing cluster for processing the to-be-processed computing job from a plurality of computing clusters based on metadata of the to-be-processed computing job; the plurality of computing clusters include heterogeneous computing power clusters of different cluster types; a job distribution module 430, configured to convert the to-be-processed computing job into a target computing job consistent with a job submission mode of the target computing cluster, and distribute the target computing job to the target computing cluster, so that the target computing cluster executes the target computing job, and releases the heterogeneous computing power resources occupied by executing the target computing job when the execution of the target computing job ends; an activation module 440, configured to, in response to a resource state update triggered based on the release of the heterogeneous computing power resources, perform scheduling state evaluation on a computing job in the second scheduling queue, and when the result of the scheduling state evaluation indicates that the computing job is schedulable, transfer the computing job to the first scheduling queue.
[0066] In some example embodiments, the metadata of the to-be-processed computing job comprises a cluster identifier; the target cluster determination module 420 comprises: a resource pool configuration obtaining module, configured to obtain resource pool configuration information, the resource pool configuration information comprising configuration files of each computing cluster in the plurality of computing clusters; a target configuration file determination module, configured to determine, based on the resource pool configuration information, a target configuration file matching the cluster identifier; a target cluster determination sub-module, configured to determine, based on the target configuration file, a target computing cluster for processing the to-be-processed computing job from the plurality of computing clusters.
[0067] In some example embodiments, the activation module 440 comprises: a current resource state obtaining module, configured to obtain updated current resource state information, the current resource state information indicating current resource states of each computing cluster in the plurality of computing clusters; a first determination module, configured to determine, as a to-be-evaluated computing job, a computing job at the head of the second scheduling queue, determine, based on metadata of the to-be-evaluated computing job, a first computing cluster for processing the to-be-evaluated computing job, and a desired resource state required for executing the to-be-evaluated computing job; a first evaluation module, configured to determine, in a case where the current resource state of the first computing cluster satisfies the desired resource state required for executing the to-be-evaluated computing job, that a result of the scheduling state evaluation is schedulable; a second evaluation module, configured to determine, in a case where the current resource state of the first computing cluster does not satisfy the desired resource state required for executing the to-be-evaluated computing job, that the result of the scheduling state evaluation is temporarily non-schedulable.
[0068] In some example embodiments, the apparatus 400 further comprises: a submitted job obtaining module, configured to obtain a first computing job submitted by a target account based on a job configuration interface; a second determination module, configured to determine, based on metadata of the first computing job, a second computing cluster for processing the first computing job from the plurality of computing clusters, and a desired resource state required for executing the first computing job; a third determination module, configured to determine, according to a matching condition between the current resource state of the second computing cluster and the desired resource state required for executing the first computing job, whether the first computing job is currently schedulable; The job cache module is configured to store the first computing job to the first scheduling queue if the first computing job is currently schedulable, and store the first computing job to the second scheduling queue if the first computing job is currently not schedulable.
[0069] In some example embodiments, the apparatus 400 further includes: The account resource state acquisition module is configured to acquire, from the current resource state information, a current resource state of a computing cluster corresponding to the target account based on a correspondence between the account and the computing cluster. The resource state display module is configured to display the current resource state of the computing cluster corresponding to the target account to the target account.
[0070] In some example embodiments, the job distribution module 430 is specifically configured to: in a case where the cluster type of the target computing cluster is a containerized cluster, encapsulate the to-be-processed computing job as a workflow of the containerized cluster, and write the workflow into a local queue of the target computing cluster; in a case where the cluster type of the target computing cluster is a high-performance computing cluster, convert the to-be-processed computing job into a target command script for submitting a batch processing job, and send the target command script to the target computing cluster through an application programming interface of the high-performance computing cluster. It should be noted that the apparatus provided in the above embodiments, in realizing its functions, only takes the above-mentioned division of each functional module as an example for illustration, and in actual application, the above-mentioned functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above-described functions. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0071] The electronic device provided in the embodiments of the present application includes a processor and a memory, the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program is loaded and executed by the processor to implement any one of the job scheduling methods provided in the above method embodiments.
[0072] The memory can be used to store software programs and modules, and the processor executes various function applications and data processing by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by functions, etc.; and the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory can also include a memory controller to provide access for the processor to the memory.
[0073] The method embodiments provided by the embodiments of the application can be executed in a computer terminal, a server or a similar computing device, that is, the above-mentioned electronic device can include a computer terminal, a server or a similar computing device. Taking the case of running on a server as an example, Figure 5 is a hardware structure block diagram of a server running a job scheduling method provided by the embodiments of the application, as Figure 5 shown, the server 500 can have a large difference due to different configurations or performances, and can include one or more central processing units (CPU) 510 (the central processing unit 510 can include but is not limited to a microprocessor MCU or a programmable logic device FPGA processing device), a memory 530 for storing data, one or more storage media 520 (such as one or more mass storage devices) for storing application programs 523 or data 522. Among them, the memory 530 and the storage medium 520 can be temporary storage or persistent storage. The programs stored in the storage medium 520 can include one or more modules, and each module can include a series of instruction operations in the server. Further, the central processing unit 510 can be configured to communicate with the storage medium 520 and execute a series of instruction operations in the storage medium 520 on the server 500. The server 500 can also include one or more power supplies 560, one or more wired or wireless network interfaces 550, one or more input and output interfaces 540, and / or one or more operating systems 521, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0074] The input / output interface 540 can be configured to receive or transmit data via a network. Examples of the network can include a wireless network provided by a communication provider of the server 500. In an example, the input / output interface 540 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In an example, the input / output interface 540 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.
[0075] Those skilled in the art can understand that, Figure 5 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 500 can further include more or fewer components than those shown, or have a different configuration of components than those shown. Figure 5 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 500 can further include more or fewer components than those shown, or have a different configuration of components than those shown. Figure 5 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 500 can further include more or fewer components than those shown, or have a different configuration of components than those shown.
[0076] The embodiments of the present application also provide a computer readable storage medium, which can be arranged in an electronic device to store at least one instruction or at least one program for implementing a job scheduling method. The at least one instruction or the at least one program is loaded and executed by the processor to implement any one of the job scheduling methods provided by the method embodiments.
[0077] The embodiments of the present application also provide a computer program product or computer program, which includes computer instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the electronic device execute any one of the job scheduling methods provided by the method embodiments.
[0078] Optionally, in the present embodiment, the storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0079] It should be noted that the above-mentioned order of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. And the above describes the specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.
[0080] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0081] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk.
[0082] The above only describes the preferred embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A job scheduling method, characterized in that: The method comprises: Obtaining a pending computing job from a first scheduling queue among a plurality of scheduling queues; the plurality of scheduling queues include the first scheduling queue and a second scheduling queue, the first scheduling queue being used to store computing jobs that can be scheduled, and the second scheduling queue being used to store computing jobs that cannot be scheduled temporarily; Determining, based on the metadata of the pending computing job, a target computing cluster for processing the pending computing job from a plurality of computing clusters; the plurality of computing clusters including heterogeneous computing clusters of different cluster types; Converting the pending computing job into a target computing job that complies with the job submission method of the target computing cluster, and distributing the target computing job to the target computing cluster so that the target computing cluster executes the target computing job, and releasing the heterogeneous computing resources occupied by executing the target computing job when the execution of the target computing job ends; In response to a resource status update triggered by the release of the heterogeneous computing resources, a scheduling status evaluation is performed on the computing job in the second scheduling queue, and when the result of the scheduling status evaluation indicates that it can be scheduled, the computing job is transferred to the first scheduling queue.
2. The method according to claim 1, characterized in that The metadata of the pending computing job includes a cluster identifier; and determining, based on the metadata of the pending computing job, a target computing cluster for processing the pending computing job from a plurality of computing clusters includes: Obtaining resource pool configuration information, the resource pool configuration information including a configuration file of each computing cluster in the multiple computing clusters; Determining a target configuration file that matches the cluster identifier based on the resource pool configuration information; Based on the target configuration file, a target computing cluster for processing the to-be-processed computing job is determined from the multiple computing clusters.
3. The method according to claim 1, characterized in that The performing scheduling status evaluation on the computing jobs in the second scheduling queue includes: Acquire updated current resource status information, where the current resource status information indicates a current resource status of each computing cluster in the plurality of computing clusters; The computing job at the head of the second scheduling queue is used as the computing job to be evaluated, and based on the metadata of the computing job to be evaluated, a first computing cluster for processing the computing job to be evaluated and an expected resource state required to execute the computing job to be evaluated are determined; When the current resource state of the first computing cluster meets the expected resource state required for executing the computing job to be evaluated, determining that the result of the scheduling state evaluation is that the cluster can be scheduled; When the current resource state of the first computing cluster does not meet the expected resource state required for executing the computing job to be evaluated, it is determined that the result of the scheduling state evaluation is that the job is temporarily unschedulable.
4. The method according to claim 3, characterized in that The method further comprises: Obtain the first computing job submitted by the target account based on the job configuration interface; determining, based on the metadata of the first computing job, a second computing cluster from the plurality of computing clusters for processing the first computing job and an expected resource state required to execute the first computing job; determining whether the first computing job can be currently scheduled based on a match between a current resource state of the second computing cluster and an expected resource state required for executing the first computing job; If the first computing job can be currently scheduled, the first computing job is stored in the first scheduling queue; if the first computing job cannot be currently scheduled, the first computing job is stored in the second scheduling queue.
5. The method according to claim 4, characterized in that The method further comprises: Based on the correspondence between the account and the computing cluster, obtaining the current resource status of the computing cluster corresponding to the target account from the current resource status information; The current resource status of the computing cluster corresponding to the target account is displayed to the target account.
6. The method according to claim 1, characterized in that The converting the pending computing job into a target computing job that is consistent with the job submission mode of the target computing cluster, and distributing the target computing job to the target computing cluster includes: In a case where the cluster type of the target computing cluster is a containerized cluster, encapsulating the to-be-processed computing job into a workflow of the containerized cluster, and writing the workflow into a local queue of the target computing cluster; In the case where the cluster type of the target computing cluster is a high-performance computing cluster, the computing job to be processed is converted into a target command script for submitting a batch job, and the target command script is sent to the target computing cluster through the application programming interface of the high-performance computing cluster.
7. A job scheduling device, characterized in that: The device comprises: a pending job acquisition module, configured to acquire pending computing jobs from a first scheduling queue among a plurality of scheduling queues; the plurality of scheduling queues including the first scheduling queue and a second scheduling queue, the first scheduling queue being configured to store computing jobs that can be scheduled, and the second scheduling queue being configured to store computing jobs that are temporarily unschedulable; a target cluster determination module, configured to determine, based on metadata of the pending computing job, a target computing cluster for processing the pending computing job from a plurality of computing clusters; the plurality of computing clusters including heterogeneous computing clusters of different cluster types; a job distribution module, configured to convert the pending computing job into a target computing job that conforms to the job submission method of the target computing cluster, and distribute the target computing job to the target computing cluster so that the target computing cluster executes the target computing job, and release the heterogeneous computing resources occupied by executing the target computing job when the execution of the target computing job is completed; An activation module is used to perform a scheduling status evaluation on the computing job in the second scheduling queue in response to a resource status update triggered by the release of the heterogeneous computing power resources, and transfer the computing job to the first scheduling queue when the result of the scheduling status evaluation indicates that it can be scheduled.
8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the job scheduling method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the job scheduling method according to any one of claims 1 to 6.
10. A computer program, characterized in that When the computer program is executed by a processor, the job scheduling method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Job processing method and device, computer equipment, storage medium and program product
CN116820758A
Distributed AI computing power resource scheduling method and system
CN117785400A
Resource scheduling method, job processing method, scheduler, system, and related device
WO2024221991A1
Cited By
Task arrangement and dynamic scheduling method oriented to multi-machine concurrency
CN121979645A