High-performance computing and intelligent computing fusion scheduling method
Patent Information
- Application Number
- CN202311706831.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-12-12
AI Technical Summary
[0004]本发明实施例提供了一种高性能计算与智能计算融合调度方法,以至少解决物理集群整体的吞吐率较低的技术问题
[0014] According to another aspect of the present invention, a processor is also provided. The processor is used to run a program, wherein the program is executed by the processor to perform the high-performance computing and intelligent computing fusion scheduling method of the present invention.
Smart Images

Figure CN117667356B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of resource scheduling, and more specifically, to a scheduling method that integrates high-performance computing and intelligent computing. Background Technology
[0002] Currently, in information scheduling, high-performance computing (HPLC) clusters and intelligent computing clusters are often deployed on different physical clusters. Even when HPLC and intelligent computing clusters are deployed on the same physical cluster, these two clusters are independent and unconnected, thus preventing resource exchange. For example, when the HPLC cluster is overloaded and queuing, while the intelligent computing cluster is idle, it is impossible to temporarily migrate computing resources from the intelligent computing cluster to the HPLC cluster to alleviate computing pressure. Conversely, when the intelligent computing cluster is overloaded and the HPLC cluster is idle, it is also impossible to temporarily migrate computing resources from the HPLC cluster to the intelligent computing cluster, resulting in a low overall throughput of the physical cluster.
[0003] There is currently no effective solution to the technical problem of low overall throughput of the aforementioned physical clusters. Summary of the Invention
[0004] This invention provides a high-performance computing and intelligent computing fusion scheduling method to at least solve the technical problem of low overall throughput of physical clusters.
[0005] According to one aspect of the present invention, a high-performance computing and intelligent computing fusion scheduling method is provided. The method may include: within a current scheduling cycle, determining a first control node of a first scheduling system and a second control node of a second scheduling system; determining, from the resource information of the first control node, first resource information occupied by at least one computing node in the first scheduling system, and from the resource information of the second control node, second resource information occupied by at least one computing node in the second scheduling system; based on the first and second resource information, determining target resource information of a target scheduling device between the first and second control nodes; and based on the scheduling strategy of the target scheduling device and the target resource information, running a job to be run and entering the next scheduling cycle of the current scheduling cycle, wherein the job to be run is determined from the task queues of the first and second scheduling systems.
[0006] Optionally, based on the first resource information and the second resource information, the target resource information of the target scheduling device between the first control node and the second control node is determined, including: summarizing the first resource information and the second resource information to obtain the target resource information.
[0007] Optionally, based on the scheduling policy of the target scheduling device and the target resource information, the pending jobs are run and the next scheduling cycle of the current scheduling cycle is entered, including: in response to the scheduling policy being the target scheduling policy, determining the pending jobs and the number of pending jobs in the task queues of the first scheduling system and the second scheduling system, wherein the target scheduling policy is used to run the pending jobs; and based on the number and target resource information, running the pending jobs and entering the next scheduling cycle.
[0008] Optionally, based on the quantity and target resource information, the job to be run is executed, including: determining a first proportion threshold for at least one computing node to be occupied in the first management system and a second proportion threshold for at least one computing node to be occupied in the second management system; and executing the job to be run based on the quantity, the first proportion threshold, the second proportion threshold and the target resource information.
[0009] Optionally, based on the quantity, a first proportional threshold, a second proportional threshold, and target resource information, running the pending jobs includes: in response to the quantity being the target quantity and the first occupancy ratio corresponding to the first resource information exceeding the first proportional threshold, comparing the first pending jobs in the task queue of the second scheduling system with the target resource information to obtain a first comparison result, and running the first pending jobs based on the first comparison result, wherein the first pending jobs represent the remaining pending jobs in the task queue of the second scheduling system when the first occupancy ratio exceeds the first proportional threshold; in response to the quantity being the target quantity and the second occupancy ratio corresponding to the second resource information exceeding the second proportional threshold, comparing the second pending jobs in the task queue of the first scheduling system with the target resource information to obtain a second comparison result, and running the second pending jobs based on the second comparison result, wherein the second pending jobs represent the remaining pending jobs in the task queue of the first scheduling system when the second occupancy ratio exceeds the second proportional threshold.
[0010] Optionally, based on the first comparison result, running the first job to be run includes: responding to the first comparison result that the running resource information for running the first job to be run meets the target resource information, running the first job to be run, and using the first remaining resource information in the target resource information to run the third job to be run in the task queue of the first scheduling system, wherein the first remaining resource information is used to represent the resource information in the target resource information other than the resource information occupied by running the first job to be run, and the third job to be run is used to represent the remaining jobs to be run in the task queue of the first scheduling system when the first occupancy ratio exceeds the first ratio threshold.
[0011] Optionally, based on the first comparison result, running the first job to be run includes: responding to the second comparison result that the running resource information for running the second job to be run meets the target resource information, running the second job to be run, and using the second remaining resource information in the target resource information to run the fourth job to be run in the task queue of the second scheduling system, wherein the second remaining resource information is used to represent the resource information in the target resource information other than the resource information occupied by running the second job to be run, and the fourth job to be run is used to represent the remaining jobs to be run in the task queue of the second scheduling system when the second occupancy ratio exceeds the second ratio threshold, and the jobs to be run include: the first job to be run, the second job to be run, the third job to be run, and the fourth job to be run.
[0012] According to one aspect of the present invention, a high-performance computing and intelligent computing fusion scheduling device is provided. The device may include: an acquisition unit, configured to determine, within a current scheduling cycle, a first control node of a first scheduling system and a second control node of a second scheduling system; a first determination unit, configured to determine, from the resource information of the first control node, first resource information occupied by at least one computing node in the first scheduling system, and from the resource information of the second control node, second resource information occupied by at least one computing node in the second scheduling system; a second determination unit, configured to determine, based on the first and second resource information, target resource information of a target scheduling device between the first and second control nodes; and a running unit, configured to run a job to be run based on the scheduling strategy of the target scheduling device and the target resource information, and enter the next scheduling cycle of the current scheduling cycle, wherein the job to be run is determined from the task queues of the first and second scheduling systems.
[0013] According to another aspect of the present invention, a computer-readable storage medium is also provided. The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to execute the high-performance computing and intelligent computing fusion scheduling method of the present invention.
[0014] According to another aspect of the present invention, a processor is also provided. The processor is used to run a program, wherein the program is executed by the processor to perform the high-performance computing and intelligent computing fusion scheduling method of the present invention.
[0015] In this embodiment of the invention, within the current scheduling cycle, a first control node of a first scheduling system and a second control node of a second scheduling system are determined. Then, from the resource information of the first control node, the first resource information occupied by at least one computing node in the first scheduling system can be determined, and from the resource information of the second control node, the second resource information occupied by at least one computing node in the second scheduling system can be determined. Based on the first and second resource information, the target resource information of the target scheduling device between the first and second control nodes can be determined. By comparing the scheduling strategy of the target scheduling device with the target scheduling strategy, and by comparing the pending jobs in the task queues of each scheduling system with the target resource information, the pending jobs can be run and enter the next scheduling cycle of the current scheduling cycle. This achieves the purpose of alleviating the pressure of computing demand, solves the technical problem of low overall throughput of the physical cluster, and realizes the technical effect of improving the overall throughput of the physical cluster. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0017] Figure 1 This is a flowchart of a high-performance computing and intelligent computing fusion scheduling method according to an embodiment of the present invention;
[0018] Figure 2 This is a schematic diagram illustrating the configuration of a scheduling system according to an embodiment of the present invention;
[0019] Figure 3 This is a schematic diagram of a single-node unified resource view and two scheduler resource views according to an embodiment of the present invention;
[0020] Figure 4 This is a flowchart illustrating the operation of a unified resource coordinator according to an embodiment of the present invention;
[0021] Figure 5 This is a schematic diagram of a high-performance computing and intelligent computing fusion scheduling device according to an embodiment of the present invention. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] Example 1
[0025] According to embodiments of the present invention, a high-performance computing and intelligent computing fusion scheduling method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0026] Figure 1 This is a flowchart of a high-performance computing and intelligent computing fusion scheduling method according to an embodiment of the present invention. The method may include the following steps:
[0027] Step S101: Within the current scheduling cycle, determine the first control node of the first scheduling system and the second control node of the second scheduling system.
[0028] In the technical solution provided by step S101 of the present invention, the first scheduling system can be used to represent a system for resource scheduling of a high-performance computing cluster, and the second scheduling system can be used to represent a system for resource scheduling of an intelligent computing cluster. For example, the first scheduling system can be a system implemented based on the open-source container orchestration and management platform (Kubernetes, or K8S for short), the first control node can be a K8S master node, the second scheduling system can be an open-source job scheduling system (Slurm), and the second control node can be a Slurm master node. This is only an example and is not specifically limited.
[0029] Optionally, within the current scheduling cycle, the first control node of the first scheduling system and the second control node of the second scheduling system are determined. For example, within the current scheduling cycle, the master control node of the first scheduling system and the master control node of the second scheduling system are determined.
[0030] Step S102: Determine the first resource information occupied by at least one computing node in the first scheduling system from the resource information of the first control node, and determine the second resource information occupied by at least one computing node in the second scheduling system from the resource information of the second control node.
[0031] In the technical solution provided by step S102 of the present invention, the first resource information may include at least the usage information of resources such as the central processing unit (CPU), memory, and computing card in the first scheduling system, and the second resource information may include at least the usage information of resources such as the CPU, memory, and computing card in the second scheduling system. For example, the computing card may be at least one of the following: graphics processing unit (GPU), network processing unit (NPU), etc. This is only an example and is not specifically limited.
[0032] Optionally, within the current scheduling cycle, after determining the first control node of the first scheduling system and the second control node of the second scheduling system, the first resource information occupied by at least one computing node in the first scheduling system is determined from the resource information of the first control node, and the second resource information occupied by at least one computing node in the second scheduling system is determined from the resource information of the second control node. For example, the usage information of CPU, memory, computing card, and other resources occupied by at least one computing node in the first scheduling system can be determined from the resource information of the master control node of the first scheduling system, and the usage information of CPU, memory, computing card, and other resources occupied by at least one computing node in the second scheduling system can be determined from the resource information of the master control node of the second scheduling system.
[0033] Optionally, the resource information of the master node of K8S can be used to determine the usage information of at least one compute node in K8S, such as CPU, memory, and compute card. Similarly, the resource information of the master node of Slurm can be used to determine the usage information of at least one compute node in Slurm, such as CPU, memory, and compute card. This is only an example and is not specifically limited.
[0034] Step S103: Based on the first resource information and the second resource information, determine the target resource information of the target scheduling device between the first control node and the second control node.
[0035] In the technical solution provided by step S103 of the present invention, the target scheduling device can be a unified resource coordinator, and the target resource information can be used to represent a unified resource view. This is only an example and is not specifically limited.
[0036] Optionally, after determining the first resource information occupied by at least one computing node in the first scheduling system from the resource information of the first control node, and determining the second resource information occupied by at least one computing node in the second scheduling system from the resource information of the second control node, the target resource information of the target scheduling device between the first control node and the second control node is determined based on the first and second resource information. For example, based on the usage information of CPU, memory, computing card, and other resources occupied by at least one computing node in the first scheduling system, and the usage information of CPU, memory, computing card, and other resources occupied by at least one computing node in the second scheduling system, the target resource information of the target scheduling device between the master control node of the first scheduling system and the master control node of the second scheduling system can be determined. For example, the unified resource view of the unified resource coordinator between the two master control nodes can be determined.
[0037] Optionally, based on the usage information of CPU, memory, computing card and other resources occupied by at least one computing node in K8S, and the usage information of CPU, memory, computing card and other resources occupied by at least one computing node in Slurm, the target resource information of the target scheduling device between the K8S master node and the Slurm master node can be determined. For example, the unified resource view of the unified resource coordinator between the two master nodes can be determined.
[0038] Step S104: Based on the scheduling strategy of the target scheduling device and the target resource information, run the job to be run and enter the next scheduling cycle of the current scheduling cycle.
[0039] In the technical solution provided by step S104 of the present invention, the job to be run can be determined from the task queue of the first scheduling system and the task queue of the second scheduling system. For example, the job to be run can include at least one of the following: the remaining jobs to be run in the task queue of the first scheduling system and the remaining jobs to be run in the task queue of the second scheduling system. This is only an example and is not specifically limited.
[0040] Optionally, after determining the target resource information of the target scheduling device between the first control node and the second control node based on the first resource information and the second resource information, the job to be run is executed based on the scheduling policy of the target scheduling device and the target resource information, and the next scheduling cycle of the current scheduling cycle is entered. For example, by comparing the scheduling policy of the target scheduling device with the target scheduling policy, and by comparing the job to be run in the task queue of each scheduling system with the target resource information, the job to be run can be executed and the next scheduling cycle of the current scheduling cycle can be entered. The target scheduling policy may include at least one of the following: the scheduler's native scheduling policy, the user-defined scheduling policy, etc.
[0041] In steps S101 to S104 of this application, within the current scheduling cycle, the first control node of the first scheduling system and the second control node of the second scheduling system are determined. Then, from the resource information of the first control node, the first resource information occupied by at least one computing node in the first scheduling system can be determined, and from the resource information of the second control node, the second resource information occupied by at least one computing node in the second scheduling system can be determined. Based on the first and second resource information, the target resource information of the target scheduling device between the first and second control nodes can be determined. By comparing the scheduling strategy of the target scheduling device with the target scheduling strategy, and by comparing the pending jobs in the task queues of each scheduling system with the target resource information, the pending jobs can be run and enter the next scheduling cycle of the current scheduling cycle. This achieves the purpose of alleviating the pressure of computing demand, solves the technical problem of low overall throughput of the physical cluster, and realizes the technical effect of improving the overall throughput of the physical cluster.
[0042] The method described in this embodiment will be further described below.
[0043] As an optional embodiment, step S103, based on the first resource information and the second resource information, determines the target resource information of the target scheduling device between the first control node and the second control node, including: summarizing the first resource information and the second resource information to obtain the target resource information.
[0044] In this embodiment, after determining the first resource information occupied by at least one computing node in the first scheduling system from the resource information of the first control node, and determining the second resource information occupied by at least one computing node in the second scheduling system from the resource information of the second control node, the target resource information can be obtained by summarizing the first and second resource information. For example, the target resource information can be obtained by summarizing the usage information of CPU, memory, computing card, and other resources occupied by at least one computing node in the first scheduling system and the usage information of CPU, memory, computing card, and other resources occupied by at least one computing node in the second scheduling system. For example, a unified resource view can be obtained.
[0045] Optionally, a unified resource view can be obtained by summarizing the usage information of CPU, memory, compute card and other resources occupied by at least one compute node in K8S and the usage information of CPU, memory, compute card and other resources occupied by at least one compute node in Slurm. This is only an example and is not specifically limited.
[0046] As an optional embodiment, step S104, based on the scheduling policy of the target scheduling device and the target resource information, runs the jobs to be run and enters the next scheduling cycle of the current scheduling cycle, including: in response to the scheduling policy being the target scheduling policy, determining the jobs to be run and the number of jobs to be run in the task queues of the first scheduling system and the second scheduling system; running the jobs to be run based on the number and the target resource information, and entering the next scheduling cycle.
[0047] In this embodiment, the target scheduling strategy described above can be used to run jobs to be run. For example, the target scheduling strategy may include at least one of the following: the scheduler's native scheduling strategy, the user-defined scheduling strategy, etc. This is only an example and is not specifically limited.
[0048] Optionally, after determining the target resource information of the target scheduling device between the first control node and the second control node based on the first resource information and the second resource information, the scheduling policy of the target scheduling device is compared with the target scheduling policy. The number of jobs to be run and the number of jobs to be run can be determined in the task queue of the first scheduling system and the task queue of the second scheduling system. If the scheduling policy is the target scheduling policy, the number of jobs to be run and the number of jobs to be run can be determined in the task queue of the first scheduling system and the task queue of the second scheduling system. Then, according to the determined number and the target resource information, the determined jobs to be run can be run and the next scheduling cycle can begin.
[0049] As an optional embodiment, running a job to be run based on quantity and target resource information includes: determining a first proportion threshold for at least one computing node to be occupied in a first management system and a second proportion threshold for at least one computing node to be occupied in a second management system; and running the job to be run based on the quantity, the first proportion threshold, the second proportion threshold, and the target resource information.
[0050] In this embodiment, the first proportional threshold can be 50%, and the second proportional threshold can also be 50%. This is only an example and is not a specific limitation.
[0051] Optionally, in response to the target scheduling strategy, after determining the jobs to be run and the number of jobs to be run in the task queues of the first scheduling system and the second scheduling system, by detecting the proportion threshold of at least one computing node in the first management system and the proportion threshold of at least one computing node in the second management system, a first proportion threshold and a second proportion threshold of at least one computing node in the second management system can be obtained. Then, based on the number, the first proportion threshold, the second proportion threshold and the target resource information, the jobs to be run can be executed.
[0052] As an optional embodiment, running a job to be run based on a quantity, a first proportional threshold, a second proportional threshold, and target resource information includes: in response to a quantity being a target quantity and a first occupancy ratio corresponding to the first resource information exceeding the first proportional threshold, comparing a first job to be run in the task queue of a second scheduling system with the target resource information to obtain a first comparison result, and running the first job to be run based on the first comparison result; in response to a quantity being a target quantity and a second occupancy ratio corresponding to the second resource information exceeding the second proportional threshold, comparing a second job to be run in the task queue of the first scheduling system with the target resource information to obtain a second comparison result, and running the second job to be run based on the second comparison result.
[0053] In this embodiment, the first job to be run can be used to represent the remaining jobs to be run in the task queue of the second scheduling system when the first occupancy ratio exceeds the first ratio threshold, and the second job to be run can be used to represent the remaining jobs to be run in the task queue of the first scheduling system when the second occupancy ratio exceeds the second ratio threshold.
[0054] Optionally, after determining the first proportion threshold for at least one computing node occupying in the first management system and the second proportion threshold for at least one computing node occupying in the second management system, by comparing the quantity with the target quantity and comparing the first occupancy ratio corresponding to the first resource information with the first proportion threshold, it can be determined whether to compare the first job to be run in the task queue of the second scheduling system with the target resource information to obtain a first comparison result, and run the first job to be run based on the first comparison result. If the quantity is the target quantity and the first occupancy ratio corresponding to the first resource information exceeds the first proportion threshold, then the first job to be run in the task queue of the second scheduling system is compared with the target resource information to obtain a first comparison result, and run the first job to be run based on the first comparison result. If the quantity is not the target quantity, or the first occupancy ratio corresponding to the first resource information does not exceed the first proportion threshold, then the first job to be run in the task queue of the second scheduling system is not compared with the target resource information.
[0055] Optionally, by comparing the quantity with the target quantity and comparing the second occupancy ratio corresponding to the second resource information with the second ratio threshold, it can be determined whether to compare the second job to be run in the task queue of the first scheduling system with the target resource information to obtain a second comparison result, and run the second job to be run based on the second comparison result. If the quantity is the target quantity and the second occupancy ratio corresponding to the second resource information exceeds the second ratio threshold, then the second job to be run in the task queue of the first scheduling system is compared with the target resource information to obtain a second comparison result, and run the second job to be run based on the second comparison result. If the quantity is not the target quantity, or the second occupancy ratio corresponding to the second resource information does not exceed the second ratio threshold, then the second job to be run in the task queue of the first scheduling system is not compared with the target resource information.
[0056] Optionally, when the first ratio threshold is 50%, by comparing the quantity with the target quantity, and by comparing the first occupancy ratio corresponding to the usage information of CPU, memory, compute card and other resources occupied by at least one compute node in K8S with 50%, it can be determined whether to compare the first job to be run in the Slurm task queue with the unified resource view to obtain a first comparison result, and run the first job to be run based on the first comparison result. If the quantity is the target quantity and the first occupancy ratio exceeds 50%, then compare the first job to be run in the Slurm task queue with the unified resource view to obtain a first comparison result, and run the remaining jobs to be run in the Slurm task queue based on the first comparison result. If the quantity is not the target quantity, or the first occupancy ratio does not exceed 50%, then do not compare the first job to be run in the Slurm task queue with the unified resource view.
[0057] Optionally, when the second ratio threshold is 50%, by comparing the quantity with the target quantity, and by comparing the second occupancy ratio corresponding to the usage information of CPU, memory, compute card, and other resources occupied by at least one compute node in Slurm with 50%, it can be determined whether to compare the second job to be run in the K8S task queue with the unified resource view to obtain a second comparison result, and run the second job to be run based on the second comparison result. If the quantity is the target quantity and the second occupancy ratio exceeds 50%, then the second job to be run in the K8S task queue is compared with the target resource information to obtain a second comparison result, and run the second job to be run based on the second comparison result. If the quantity is not the target quantity, or the second occupancy ratio does not exceed 50%, then the second job to be run in the K8S task queue is not compared with the target resource information.
[0058] As an optional embodiment, running a first job to be run based on a first comparison result includes: in response to the first comparison result indicating that the running resource information for running the first job to be run meets the target resource information, running the first job to be run, and using the first remaining resource information in the target resource information to run a third job to be run in the task queue of the first scheduling system.
[0059] In this embodiment, the first remaining resource information can be used to represent the resource information in the target resource information other than the resource information occupied by the first job to be run, and the third job to be run can be used to represent the remaining jobs to be run in the task queue of the first scheduling system when the first occupancy ratio exceeds the first ratio threshold.
[0060] Optionally, by analyzing the first comparison result, if the analysis shows that the running resource information for running the first job to be run meets the target resource information, then the first job to be run is run, and the third job to be run in the task queue of the first scheduling system is run using the first remaining resource information in the target resource information.
[0061] As an optional embodiment, running a first job to be run based on a first comparison result includes: in response to a second comparison result that the running resource information for running a second job to be run satisfies the target resource information, running the second job to be run, and using the second remaining resource information in the target resource information to run a fourth job to be run in the task queue of the second scheduling system.
[0062] In this embodiment, the second remaining resource information can be used to represent the resource information in the target resource information other than the resource information occupied by the second pending job. The fourth pending job can be used to represent the remaining pending jobs in the task queue of the second scheduling system when the second occupancy ratio exceeds the second ratio threshold. The pending jobs can include: the first pending job, the second pending job, the third pending job, and the fourth pending job.
[0063] Optionally, by analyzing the second comparison result, if the analysis shows that the running resource information for running the second job meets the target resource information, then the second job is run, and the fourth job in the task queue of the second scheduling system is run using the second remaining resource information in the target resource information. For example, in the entire fusion system, whether the resources occupied by the two schedulers, Slurm and K8S, are occupied in a certain proportion. For example, when the proportion is 50%, it means that Slurm and K8S each occupy half of the resources. If Slurm exceeds the occupied amount, the frequency of selecting jobs from Slurm is reduced until all jobs in the K8S scheduler are in running status, and then all remaining resources are allocated to jobs in Slurm.
[0064] In this embodiment, the switch information and operation information of at least one branch containing an execution device are acquired to determine the conduction state of the switch in the branch and the operation state of the branch. By analyzing the acquired switch information and operation information, it can be determined whether the switch in each branch containing the execution device is conducting and whether the branch is operating normally. Thus, from the branch set, at least one target branch can be determined based on the branches where the switch is conducting and the branch is operating normally. Then, based on the remaining branches in the branch set excluding the target branch, the output threshold of the remaining branches can be determined. Finally, by summing the output threshold and the component of the target branch, the target output quantity of the execution device can be obtained. This achieves the goal of standardizing the method for coordinated control of multiple controllers, solves the technical problem that process industry control methods are difficult to reuse quickly in the prior art, and realizes the technical effect of rapid reuse of process industry control methods.
[0065] Example 2
[0066] The technical solutions of the embodiments of the present invention will be illustrated below with reference to preferred embodiments.
[0067] In the process of information scheduling, high-performance computing clusters and intelligent computing clusters are often deployed on different physical clusters. Even if high-performance computing clusters and intelligent computing clusters are deployed on the same physical cluster, the two clusters are independent of each other and do not communicate with each other, so they cannot exchange resources. For example, when the high-performance computing cluster is overloaded and queued, and the intelligent computing cluster is idle and unused, it is impossible to temporarily migrate the computing resources of the intelligent computing cluster to the high-performance computing cluster to alleviate the computing demand pressure. Conversely, when the intelligent computing cluster is overloaded and the high-performance computing cluster is idle, it is also impossible to temporarily migrate the computing resources of the high-performance computing cluster to the intelligent computing cluster, thus leading to the technical problem of low overall throughput of the physical cluster.
[0068] However, this invention proposes a high-performance computing and intelligent computing fusion scheduling method. This method deploys an open-source job scheduling system (Slurm) and an open-source container orchestration and management platform (Kubernetes, or K8S) on each computing node, thus achieving a hybrid deployment of Slurm and K8S. This allows the high-performance computing cluster running Slurm to connect with the intelligent computing cluster running K8S, thereby alleviating the pressure of computing demand, solving the technical problem of low overall throughput of the physical cluster, and achieving the technical effect of improving the overall throughput of the physical cluster.
[0069] Figure 2 This is a schematic diagram illustrating the configuration of a scheduling system according to an embodiment of the present invention, as shown below. Figure 2 As shown, the unified resource coordinator 201 can transmit data with the Slurm master node 202 and the K8S master node 203. The Slurm master node 202 and the K8S master node 203 can transmit data with the compute nodes 2041 to 204n.
[0070] Optionally, a unified resource coordinator 201 is constructed to maintain a unified resource view, which can be above the K8S scheduler and the Slurm scheduler. The unified resource coordinator can be responsible for polling the jobs in the Slurm and K8S job queues and determining whether the jobs can be run. If the jobs can be run, the Slurm or K8S scheduler is notified to run the jobs.
[0071] Optionally, the compute nodes in the cluster can be scheduled by both the Slurm scheduler and the Kubernetes scheduler. Both the Slurm scheduler and the Kubernetes scheduler can obtain the current cluster resource information from the unified resource coordinator 201 and perform job resource scheduling and allocation.
[0072] Figure 3This is a schematic diagram of a single-node unified resource view and two scheduler resource views according to an embodiment of the present invention, as shown below. Figure 3 As shown, the unified resource view covers the resource usage of all compute nodes. For each compute node, the unified resource view can obtain the usage of resources such as CPU, memory, and compute cards (e.g., GPU, NPU) from the Slurm scheduler and the Kubernetes scheduler respectively. Then, it summarizes the resource view and sends the summarized and merged resource view to the Slurm scheduler and the Kubernetes scheduler.
[0073] Optionally, all jobs in the system can be managed and scheduled by a unified resource coordinator. After the system starts, the unified resource coordinator can build a unified resource view based on the resource usage in the Slurm scheduler and the Kubernetes scheduler.
[0074] Optionally, based on the scheduling policy, jobs to be run can be obtained from the task queues of the Slurm scheduler and the Kubernetes scheduler. The unified resource coordinator needs to determine whether the resources requested by the current task are satisfied based on the unified resource view. If satisfied, the task is allowed to execute and enter the next round of scheduling. If not satisfied, the task is not executed and directly enters the next round of scheduling. The scheduling policy may include at least one of the following: the scheduler's native scheduling policy, the user-defined scheduling policy, etc.
[0075] Optionally, the scheduling strategy can be specifically as follows: The job selection method for the Slurm scheduler and the K8S scheduler: When the converged scheduling system starts, it first determines which scheduler to select jobs from, and whether to select one or multiple tasks in each round of scheduling; Cluster resource allocation between the Slurm and K8S schedulers: In the entire converged system, whether the resources occupied by the Slurm and K8S schedulers are allocated according to a certain ratio. For example, when the ratio is 50%, it means that Slurm and K8S each occupy half of the resources. If Slurm exceeds its allocation, the frequency of selecting jobs from Slurm is reduced until all jobs in the K8S scheduler are running, and then the remaining resources are allocated to the jobs in Slurm.
[0076] Figure 4 This is a flowchart illustrating the operation of a unified resource coordinator according to an embodiment of the present invention, such as... Figure 4 As shown, this working method may include the following steps:
[0077] Step S401: Start the unified resource scheduling system.
[0078] After starting the unified resource scheduling system, proceed to step S402 to obtain the job running status of the Slurm scheduler and the K8S scheduler, and build or update the unified resource view.
[0079] After obtaining the job running status of the Slurm scheduler and the Kubernetes scheduler, and building or updating the unified resource view, proceed to step S403, and select the job to be scheduled according to the scheduling policy.
[0080] After selecting a job to be scheduled according to the scheduling strategy, proceed to step S404 to determine whether the job meets the resource requirements. If the job meets the resource requirements, proceed to steps S405 and S407 to run the job, end this round of scheduling, and proceed to the next round of scheduling. If the job does not meet the resource requirements, proceed to steps S406 and S407 to not run the job, end this round of scheduling, and proceed to the next round of scheduling.
[0081] In this embodiment, a unified resource scheduling system is started, and then the job running status of the Slurm scheduler and the Kubernetes scheduler is obtained. A unified resource view is built or updated, and then, according to the scheduling policy, a job to be scheduled is selected. It is determined whether the job meets the resource requirements. If the job meets the resource requirements, the job is run, the current round of scheduling ends, and the next round of scheduling begins. If the job does not meet the resource requirements, the job is not run, the current round of scheduling ends, and the next round of scheduling begins. This achieves the purpose of alleviating the pressure of computing demand, solves the technical problem of low overall throughput of the physical cluster, and realizes the technical effect of improving the overall throughput of the physical cluster.
[0082] Example 3
[0083] According to embodiments of the present invention, a high-performance computing and intelligent computing fusion scheduling device is also provided. It should be noted that this high-performance computing and intelligent computing fusion scheduling device can be used to execute a high-performance computing and intelligent computing fusion scheduling method as described in Embodiment 1.
[0084] Figure 5 This is a schematic diagram of a high-performance computing and intelligent computing fusion scheduling device according to an embodiment of the present invention. Figure 5 As shown, the high-performance computing and intelligent computing fusion scheduling device 500 may include: an acquisition unit 501, a first determination unit 502, a second determination unit 503, and a running unit 504.
[0085] The acquisition unit 501 is used to determine the first control node of the first scheduling system and the second control node of the second scheduling system within the current scheduling cycle.
[0086] The first determining unit 502 is used to determine, from the resource information of the first control node, the first resource information occupied by at least one computing node in the first scheduling system, and from the resource information of the second control node, the second resource information occupied by at least one computing node in the second scheduling system.
[0087] The second determining unit 503 is used to determine the target resource information of the target scheduling device between the first control node and the second control node based on the first resource information and the second resource information.
[0088] The running unit 504 is used to run the job to be run based on the scheduling strategy and target resource information of the target scheduling device, and enter the next scheduling cycle of the current scheduling cycle. The job to be run is determined from the task queue of the first scheduling system and the task queue of the second scheduling system.
[0089] Optionally, the second determining unit 503 may include: a summarizing module, used to summarize the first resource information and the second resource information to obtain the target resource information.
[0090] Optionally, the running unit 504 may include: a determining module, used to determine the jobs to be run and the number of jobs to be run in the task queues of the first scheduling system and the second scheduling system in response to the scheduling policy being the target scheduling policy, wherein the target scheduling policy is used to run the jobs to be run; and a running module, used to run the jobs to be run based on the number and target resource information, and enter the next scheduling cycle.
[0091] Optionally, the running module may include: a determining submodule, used to determine a first proportion threshold for at least one computing node to occupy in the first management system and a second proportion threshold for at least one computing node to occupy in the second management system; and a running submodule, used to run the job to be run based on the quantity, the first proportion threshold, the second proportion threshold and the target resource information.
[0092] Optionally, the running submodule can execute the following steps to run pending jobs based on quantity, a first proportional threshold, a second proportional threshold, and target resource information: In response to the quantity being the target quantity and the first occupancy ratio corresponding to the first resource information exceeding the first proportional threshold, the first pending job in the task queue of the second scheduling system is compared with the target resource information to obtain a first comparison result, and the first pending job is run based on the first comparison result, wherein the first pending job represents the remaining pending jobs in the task queue of the second scheduling system when the first occupancy ratio exceeds the first proportional threshold; In response to the quantity being the target quantity and the second occupancy ratio corresponding to the second resource information exceeding the second proportional threshold, the second pending job in the task queue of the first scheduling system is compared with the target resource information to obtain a second comparison result, and the second pending job is run based on the second comparison result, wherein the second pending job represents the remaining pending jobs in the task queue of the first scheduling system when the second occupancy ratio exceeds the second proportional threshold.
[0093] Optionally, the running submodule can execute the following steps to run a first job to be run based on the first comparison result: in response to the first comparison result indicating that the running resource information for running the first job to be run meets the target resource information, run the first job to be run, and use the first remaining resource information in the target resource information to run a third job to be run in the task queue of the first scheduling system, wherein the first remaining resource information is used to represent the resource information in the target resource information other than the resource information occupied by running the first job to be run, and the third job to be run is used to represent the remaining jobs to be run in the task queue of the first scheduling system when the first occupancy ratio exceeds the first ratio threshold.
[0094] Optionally, the running submodule can execute the following steps to run a first job to be run based on a first comparison result: in response to a second comparison result that the running resource information for running a second job to be run satisfies the target resource information, run the second job to be run, and use the second remaining resource information in the target resource information to run a fourth job to be run in the task queue of the second scheduling system, wherein the second remaining resource information is used to represent the resource information in the target resource information other than the resource information occupied by running the second job to be run, and the fourth job to be run is used to represent the remaining jobs to be run in the task queue of the second scheduling system when the second occupancy ratio exceeds the second ratio threshold, and the jobs to be run include: the first job to be run, the second job to be run, the third job to be run, and the fourth job to be run.
[0095] In this embodiment, the acquisition unit is used to determine the first control node of the first scheduling system and the second control node of the second scheduling system within the current scheduling cycle; the first determination unit is used to determine, from the resource information of the first control node, the first resource information occupied by at least one computing node in the first scheduling system, and from the resource information of the second control node, the second resource information occupied by at least one computing node in the second scheduling system; the second determination unit is used to determine, based on the first and second resource information, the target resource information of the target scheduling device between the first and second control nodes; and the running unit is used to run the job to be run based on the scheduling strategy of the target scheduling device and the target resource information, and enter the next scheduling cycle of the current scheduling cycle. The job to be run is determined from the task queues of the first and second scheduling systems, thus solving the technical problem of low overall throughput of the physical cluster and achieving the technical effect of improving the overall throughput of the physical cluster.
[0096] Example 4
[0097] According to an embodiment of the present invention, a computer-readable storage medium is also provided, the storage medium including a stored program, wherein the program executes the high-performance computing and intelligent computing fusion scheduling method of embodiment 1.
[0098] Example 5
[0099] According to an embodiment of the present invention, a processor is also provided for running a program, wherein the program is executed by the processor to perform the high-performance computing and intelligent computing fusion scheduling method in embodiment 1.
[0100] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0101] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0102] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0103] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0104] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0105] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0106] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A scheduling method integrating high-performance computing and intelligent computing, characterized in that, include: Within the current scheduling cycle, determine the first control node of the first scheduling system and the second control node of the second scheduling system; From the resource information of the first control node, determine the first resource information occupied by at least one computing node in the first scheduling system, and from the resource information of the second control node, determine the second resource information occupied by the at least one computing node in the second scheduling system. Based on the first resource information and the second resource information, the target resource information of the target scheduling device between the first control node and the second control node is determined. Based on the scheduling strategy of the target scheduling device and the target resource information, the job to be run is executed, and the next scheduling cycle of the current scheduling cycle is entered. The job to be run is determined from the task queue of the first scheduling system and the task queue of the second scheduling system. The process of running pending jobs and entering the next scheduling cycle of the current scheduling cycle, based on the scheduling policy of the target scheduling device and the target resource information, includes: in response to the scheduling policy being a target scheduling policy, determining the pending jobs and the number of pending jobs in the task queues of the first scheduling system and the second scheduling system, wherein the target scheduling policy is used to run the pending jobs; determining a first proportion threshold for the at least one computing node in the first management system and a second proportion threshold for the at least one computing node in the second management system; and running the pending jobs based on the number, the first proportion threshold, the second proportion threshold, and the target resource information.
2. The method according to claim 1, characterized in that, Based on the first resource information and the second resource information, the target resource information of the target scheduling device between the first control node and the second control node is determined, including: The target resource information is obtained by summarizing the first resource information and the second resource information.
3. The method according to claim 1, characterized in that, Based on the quantity, the first ratio threshold, the second ratio threshold, and the target resource information, the job to be run is executed, including: In response to the quantity being a target quantity and the first occupancy ratio corresponding to the first resource information exceeding the first ratio threshold, the first job to be run in the task queue of the second scheduling system is compared with the target resource information to obtain a first comparison result, and the first job to be run is run based on the first comparison result, wherein the first job to be run is used to represent the remaining jobs to be run in the task queue of the second scheduling system when the first occupancy ratio exceeds the first ratio threshold; In response to the target quantity being the quantity and the second occupancy ratio corresponding to the second resource information exceeding the second ratio threshold, the second job to be run in the task queue of the first scheduling system is compared with the target resource information to obtain a second comparison result, and the second job to be run is run based on the second comparison result, wherein the second job to be run represents the remaining job to be run in the task queue of the first scheduling system when the second occupancy ratio exceeds the second ratio threshold.
4. The method according to claim 3, characterized in that, Based on the first comparison result, the first job to be run is executed, including: In response to the first comparison result indicating that the runtime resource information for running the first job to be run satisfies the target resource information, the first job to be run is run, and the third job to be run in the task queue of the first scheduling system is run using the first remaining resource information in the target resource information. The first remaining resource information represents the resource information in the target resource information other than the resource information occupied by running the first job to be run, and the third job to be run represents the remaining jobs to be run in the task queue of the first scheduling system when the first occupancy ratio exceeds the first ratio threshold.
5. The method according to claim 3, characterized in that, Based on the second comparison result, the second job to be run is executed, including: In response to the second comparison result that the running resource information for running the second job to be run satisfies the target resource information, the second job to be run is run, and the fourth job to be run in the task queue of the second scheduling system is run using the second remaining resource information in the target resource information. The second remaining resource information is used to represent the resource information in the target resource information other than the resource information occupied by running the second job to be run. The fourth job to be run is used to represent the remaining jobs to be run in the task queue of the second scheduling system when the second occupancy ratio exceeds the second ratio threshold. The jobs to be run include: the first job to be run, the second job to be run, the third job to be run, and the fourth job to be run.
6. A high-performance computing and intelligent computing integrated scheduling device, characterized in that, include: The acquisition unit is used to determine the first control node of the first scheduling system and the second control node of the second scheduling system within the current scheduling cycle. The first determining unit is configured to determine, from the resource information of the first control node, first resource information occupied by at least one computing node in the first scheduling system, and from the resource information of the second control node, second resource information occupied by the at least one computing node in the second scheduling system. The second determining unit is configured to determine the target resource information of the target scheduling device between the first control node and the second control node based on the first resource information and the second resource information. The running unit is used to run the job to be run based on the scheduling policy of the target scheduling device and the target resource information, and enter the next scheduling cycle of the current scheduling cycle, wherein the job to be run is determined from the task queue of the first scheduling system and the task queue of the second scheduling system; The running unit is further configured to execute the following steps to run the pending jobs based on the scheduling policy of the target scheduling device and the target resource information, and enter the next scheduling cycle of the current scheduling cycle: in response to the scheduling policy being the target scheduling policy, determining the pending jobs and the number of pending jobs in the task queues of the first scheduling system and the second scheduling system, wherein the target scheduling policy is used to run the pending jobs; determining a first proportion threshold for the at least one computing node to occupy in the first management system and a second proportion threshold for the at least one computing node to occupy in the second management system; and running the pending jobs based on the number, the first proportion threshold, the second proportion threshold, and the target resource information.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to execute the high-performance computing and intelligent computing fusion scheduling method according to any one of claims 1 to 5.
8. A processor, characterized in that, The processor is used to run a program, wherein the program is executed by the processor to perform the high-performance computing and intelligent computing fusion scheduling method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Cross-cluster resource scheduling method and system and terminal equipment
CN116010111A
Quality of service tagging for computing jobs
US20170097851A1