Job scheduling method, server, and server cluster
By dynamically allocating computing and storage resources based on job requirements in HPC clusters, the problem of low resource utilization of existing schedulers is solved, and more efficient resource management and job execution is achieved.
Patent Information
- Application Number
- PCT/CN2024/099079
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-06-13
- Publication Date
- 2025-07-03
AI Technical Summary
The existing HPC cluster scheduler fails to flexibly schedule when allocating resources, resulting in low resource utilization, especially the failure to fully consider detailed resource requirements such as memory bandwidth, resulting in waste of resources.
The management node obtains the resource requirements of the target job, and based on the computing and storage resource requirements of the target job, determines the target computing nodes that meet the needs from multiple computing nodes, and distributes the jobs to these nodes for execution, taking into account storage resource requirements such as memory bandwidth for flexible allocation.
Improve resource utilization, reduce resource waste, and improve job execution efficiency and overall throughput of server clusters.
Smart Images

Figure CN2024099079_03072025_PF_FP_ABST
Abstract
Description
Job scheduling method, server and server cluster
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 25, 2023, with application number 202311798795.0 and application name “A job scheduling method, server and server cluster”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of server technology, and in particular to a job scheduling method, a server, and a server cluster. Background Art
[0003] The performance of parallel computing programs in the field of high-performance computing (HPC) is determined by the number of central processing unit (CPU) cores, floating-point performance per core, memory bandwidth, network bandwidth and latency, and storage bandwidth and latency. In non-interactive computing in HPC, jobs issued by upper-level applications are typically submitted to a scheduler. The scheduler first evaluates the job's required resources and execution priority, and then assigns the job to an appropriate node for execution.
[0004] The current scheduler for HPC clusters primarily considers the number of CPU cores and memory size, without giving much consideration to detailed resources such as memory bandwidth. This results in resource waste and low resource utilization when the HPC cluster runs multiple applications.
[0005] To address this issue, a scheduling solution has been proposed. In this solution, the management node in the HPC cluster categorizes applications into multiple types and allocates corresponding resources to each type. When an application issues a job, the management node allocates the corresponding resource type based on the application type. This solution takes into account specific resources such as memory bandwidth, but the various resources are pre-configured, making it difficult to flexibly schedule them based on actual application needs, resulting in low resource utilization.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a job scheduling method, a server, and a server cluster, which can flexibly allocate resources and improve resource utilization.
[0008] To achieve the above technical objectives, the embodiments of the present application adopt the following technical solutions:
[0009] In the first aspect, an embodiment of the present application provides a job scheduling method, which is applied to a management node in a server cluster, and the management node communicates with multiple computing nodes in the server cluster. The method includes: obtaining the resource requirements of the target job; the resource requirements include computing resource requirements and storage resource requirements; determining one or more target computing nodes based on the resource requirements of the target job, and the one or more target computing nodes meet the resource requirements of the target job; and sending the target job to one or more target computing nodes so that the target job is executed in one or more target computing nodes.
[0010] In the above method, when the target computing node executes the target job, the resources allocated to the target job are the resources required by the target job. In addition, storage resource requirements are also taken into account when allocating resources. Therefore, this method can flexibly allocate resources according to the actual job situation and improve resource utilization.
[0011] In one possible implementation, one or more target computing nodes are determined based on the resource requirements of the target job, and the one or more target computing nodes meet the resource requirements of the target job, including: obtaining available resource information of multiple computing nodes, the available resource information including available computing resource information and available storage resource information; based on the available resource information of each computing node, determining one or more target computing nodes; wherein the available resources of the one or more target computing nodes are greater than or equal to the resource requirements of the target job.
[0012] It is understandable that determining one or more target computing nodes based on the available resource information of each computing node can ensure that the available resources of the target computing nodes meet the requirements of the target job, so as to successfully execute the target job.
[0013] In another possible implementation, the management node stores a resource information table, which is used to record resource information of multiple computing nodes. The resource information includes available resource information. Obtaining the available resource information of multiple computing nodes includes: obtaining the available resource information of multiple computing nodes based on the resource information of the multiple computing nodes recorded in the resource information table.
[0014] It can be understood that the management node manages the resources of multiple computing nodes through the resource information table, can clearly grasp the resource allocation status of multiple computing nodes and the number of remaining available resources, and can allocate jobs to computing nodes with sufficient available resources based on the resource information table, which is conducive to improving the management efficiency of the management node on the resources on the computing nodes.
[0015] In another possible implementation, the method further includes: sending a first indication to multiple computing nodes respectively; the first indication is used to instruct the computing nodes to return their respective available resource information; based on the available resource information returned by the multiple computing nodes, correcting the available resource information of the multiple computing nodes recorded in the resource information table.
[0016] It is understandable that the management node can periodically or regularly send the first indication to multiple computing nodes to correct the available resource information of multiple computing nodes recorded in the resource information table to improve the accuracy of the available resource information of multiple computing nodes recorded in the resource information table.
[0017] In another possible implementation, after the target job is sent to one or more target computing nodes so that the target job is executed in the one or more target computing nodes, the method also includes: updating the available resource information of the one or more target computing nodes recorded in the resource information table based on the resources occupied by the one or more target computing nodes for processing the target job.
[0018] It can be understood that if the management node determines the target computing node, it means that the management node instructs to allocate the resources of the target computing node to the target job, and updates the available resource information of the target computing node in the resource information table. This is beneficial for subsequent other jobs to be allocated based on the latest resource information in a timely manner, avoiding job execution failure due to insufficient available resources.
[0019] In another possible implementation, the management node stores a resource information table, which is used to record resource information of multiple computing nodes. The method also includes: upon receiving a response message indicating that the target job has been executed, reclaiming the resources allocated to the target job on one or more target computing nodes, and updating the available resource information of the one or more target computing nodes recorded in the resource information table accordingly.
[0020] It is understandable that the management node timely updates the resource information of the target computing node recorded in the resource information table, which is beneficial for obtaining the latest resource information of multiple computing nodes when there are new jobs in the future, thereby improving the utilization rate of resources in the computing nodes.
[0021] In another possible implementation, the target job is sent to one or more target computing nodes, including: splitting the target job into multiple sub-jobs; sending task information to the one or more target computing nodes, where the task information includes the sub-jobs to be executed and the resources required to execute the sub-jobs to be executed.
[0022] It is understandable that if the available resources of each computing node do not meet the resources required by the target job, the target job can be divided into multiple sub-jobs, and the multiple sub-jobs can be distributed to multiple target computing nodes so that multiple target computing nodes can complete the target job together.
[0023] In another possible implementation, before sending the target job to one or more target computing nodes, the method also includes: determining one or more CPUs to execute the target job based on the target computing resources required by the target job; sending the target job to one or more target computing nodes, including: sending task information to the one or more target computing nodes, the task information including: resources required to execute the target job and one or more CPUs to execute the target job; the resources required by the target job include target computing resources and target storage resources, and the target computing resources include the target number of CPU cores.
[0024] It is understandable that after determining the target computing node, the CPU that executes the target job is determined based on the target computing resources required by the target job to ensure that the target job can be successfully executed.
[0025] In another possible implementation, the computing resource requirement includes a CPU core number requirement, and the storage resource requirement includes a memory bandwidth requirement.
[0026] In another possible implementation, the storage resource requirement includes a memory bandwidth requirement and an L3 cache requirement.
[0027] It can be understood that in the above two implementation methods, memory bandwidth requirements are taken into account in the storage resources. L3cache is also optional. In this method, storage resources such as memory bandwidth are subdivided as dynamically allocated resources, so that the storage resources of the target computing node can be more fully utilized, improving resource utilization, and further improving the performance of the computing node when executing the target job and the overall throughput of the cluster.
[0028] In the second aspect, an embodiment of the present application provides a job scheduling method, which is applied to a computing node in a server cluster, and the computing node communicates with a management node in the server cluster. The method includes: receiving task information sent by the management node; the task information includes a job to be executed, and target resources, the job to be executed is a target job, or a sub-job of the target job, the target resources are the resources required to execute the job to be executed, and the target resources include target computing resources and target storage resources; allocating target resources to the job to be executed, and executing the job to be executed.
[0029] It is understandable that the target computing node allocates the required target resources to the target job under the instruction of the management node, which can improve the flexibility of resource allocation and increase resource utilization.
[0030] In one possible implementation, the task information includes one or more CPUs that execute the job to be executed, and target resources are allocated to the job to be executed. Before executing the job to be executed, the steps include: distributing the job to be executed to one or more CPUs based on the task information; allocating target resources to the job to be executed, and executing the job to be executed, including: allocating target resources on one or more CPUs to the job to be executed, and executing the job to be executed on one or more CPUs.
[0031] It is understandable that, under the instruction of the management node, the target computing node distributes the job to be executed to the designated CPU to successfully execute the job to be executed and improve the execution efficiency of the job.
[0032] In another possible implementation, the target storage resource includes a target memory bandwidth. After allocating the target resource to the job to be executed, the method further includes: physically isolating the target memory bandwidth.
[0033] It is understandable that the target computing node physically isolates the target memory bandwidth, so that when the target computing node executes multiple jobs simultaneously, each job only uses its allocated memory bandwidth, effectively avoiding resource competition among multiple jobs.
[0034] In a third aspect, an embodiment of the present application provides a server, wherein the server is applied to various modules of the job scheduling method of the first aspect or any possible implementation of the first aspect, or the server is applied to various modules of the job scheduling method of the second aspect or any possible implementation of the second aspect.
[0035] In a fourth aspect, embodiments of the present application provide a server comprising a memory and a processor. The memory and the processor are coupled; the memory is configured to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the server executes the job scheduling method according to the first aspect and any possible implementation thereof; alternatively, when the processor executes the computer instructions, the server executes the job scheduling method according to the second aspect and any possible implementation thereof.
[0036] In a fifth aspect, an embodiment of the present application provides a server cluster, comprising a plurality of servers, wherein the plurality of servers are divided into a management node and a plurality of computing nodes; the management node comprises: a job receiving module, a resource management center module and a storage resource management module; the job receiving module is used to receive a job request and send the job request to the resource management center module; the job request is used to request allocation of resources for a target job and to execute the target job; the resource management center module is used to call the storage resource management module based on the resource requirements of the target job to determine the storage resources of the plurality of computing nodes; the resource requirements include computing resource requirements and storage resource requirements; the resource management center module is also used to determine one or more target nodes for executing the target job based on the resources of the plurality of computing nodes Computing node; a storage resource management module for managing the storage resources of multiple computing nodes; the computing node includes: a job execution module, a storage resource allocation module and a memory bandwidth interface; the job execution module is used to receive jobs distributed by the resource management center module, and under the instruction of the resource management center module, call the storage resource allocation module to allocate the indicated storage resources to the job; the storage resources include memory bandwidth; the storage resource allocation module is used to allocate the indicated storage resources to the job through the memory bandwidth interface under the instruction of the job execution module; the storage resource allocation module is also used to release the storage resources allocated to the job after the job execution is completed; the memory bandwidth interface is used to allocate memory bandwidth and physically isolate the allocated memory bandwidth.
[0037] In a sixth aspect, embodiments of the present application provide a computer-readable storage medium comprising computer instructions. When the computer instructions are executed on a server, the server executes the job scheduling method according to the first aspect and any possible implementation thereof; or, when the computer instructions are executed on the server, the server executes the job scheduling method according to the second aspect and any possible implementation thereof.
[0038] In a seventh aspect, embodiments of the present application provide a computer program product comprising computer instructions. When the computer instructions are executed on a server, the server executes the job scheduling method according to the first aspect and any possible implementation thereof; or, when the computer instructions are executed on the server, the server executes the job scheduling method according to the second aspect and any possible implementation thereof.
[0039] For the specific description of the third to seventh aspects and their various implementations in the embodiments of the present application, reference can be made to the detailed description in the first or second aspect and their various implementations; and for the beneficial effects of the third to seventh aspects and their various implementations, reference can be made to the analysis of the beneficial effects in the first or second aspect and their various implementations, which will not be repeated here.
[0040] These and other aspects of the embodiments of the present application will be more clearly understood in the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] FIG1 is a diagram of a server cluster architecture involved in a job scheduling method provided in an embodiment of the present application;
[0042] FIG2 is another server cluster architecture diagram involved in a job scheduling method provided in an embodiment of the present application;
[0043] FIG3 is a flow chart of a job scheduling method provided in an embodiment of the present application;
[0044] FIG4 is a flow chart of another job scheduling method provided in an embodiment of the present application;
[0045] FIG5 is a schematic diagram of a resource configuration interface provided in an embodiment of the present application;
[0046] FIG6 is a diagram of a job operation performance provided by an embodiment of the present application;
[0047] FIG7 is another operation performance diagram provided by an embodiment of the present application;
[0048] FIG8 is a flowchart of another job scheduling method provided in an embodiment of the present application;
[0049] FIG9 is a schematic diagram of the structure of a management node provided in an embodiment of the present application;
[0050] FIG10 is a schematic diagram of the structure of a computing node provided in an embodiment of the present application;
[0051] FIG11 is a schematic diagram of the structure of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] For ease of understanding, the following briefly introduces the relevant terms involved in the embodiments of this application:
[0053] (1) Slurm (simple Linux utility for resource management) is an open source, fault-tolerant, highly scalable cluster management and job scheduling system for large and small Linux clusters.
[0054] (2) PBS (portable batch system) is a software system used to schedule batch jobs on computer clusters or supercomputers.
[0055] (3) Job: A job is a collection of tasks that a user requests a computer system to perform during a problem solving process or a transaction. In some operating systems, a job is an execution unit assigned to the operating system by a computer operator (or job scheduler). A job consists of a program, corresponding data, and a job description.
[0056] (4) Resource director technology (RDT), which helps system administrators allocate and manage resources on the processor.
[0057] In the following, the terms "first," "second," and "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Thus, a feature designated as "first," "second," or "third," etc., may explicitly or implicitly include one or more of the features.
[0058] In the description of the embodiments of this application, unless otherwise specified, " / " means "or." For example, A / B can mean A or B. "And / or" in this document is simply a description of an association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more.
[0059] Related technology In order to solve the problem of resource waste in the current HPC cluster, a scheduling solution is proposed. In this solution, the management node in the HPC cluster divides the applications into multiple types and pre-configures the corresponding resources for each type of application. When an application issues a job, the management node allocates corresponding resources to the job based on the type of application that issues the job. For example, each type of application has different requirements for resources. Computation-intensive applications have higher requirements for the number of CPU cores, memory-intensive applications have higher requirements for memory bandwidth, and communication-intensive applications and IO-intensive applications have higher requirements for network performance. This solution divides resources into several categories based on the needs of different applications and establishes a corresponding relationship between resources and applications. Although this solution takes into account detailed resources such as memory bandwidth, various resources have been pre-configured and cannot be flexibly scheduled based on the needs of actual applications, resulting in low resource utilization.
[0060] For example, if the job type is memory-intensive, the requirement for the number of CPU cores is 2 CPU cores, and the requirement for single-core memory bandwidth is 8GB / s, the management unit has pre-set the number of CPU cores to 4 CPU cores and the single-core memory bandwidth to 10GB / s for memory-intensive jobs. At this time, the management unit cannot allocate resources to the job based on the requirements of the job. If 4 CPU cores and 10GB / s single-core memory bandwidth are allocated to the job, there will be a waste of resources, which reduces resource utilization.
[0061] Based on this, an embodiment of the present application proposes a job scheduling method, which is applied to a management node in a server cluster. After the management node obtains the resource requirements of the target job, it determines one or more target computing nodes that meet the target job requirements from multiple computing nodes based on the target resources required by the target job, and distributes the target job to one or more target computing nodes, so that one or more target computing nodes use the target resources to execute the target job. It can be understood that in this method, when the target computing node executes the target job, the resources allocated to the target job are the resources required by the target job. In addition, the storage resource requirements are also considered when allocating resources. Therefore, this method can flexibly allocate resources according to the actual job situation and improve resource utilization.
[0062] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0063] Please refer to Figure 1 , which shows a server cluster architecture diagram involved in the job scheduling method provided in an embodiment of the present application. As shown in Figure 1 , the server cluster 100 includes: a management node 110 and a computing node 120 .
[0064] The management node 110 and the computing nodes 120 are connected via a network to form a server cluster 100, such as an HPC cluster.
[0065] The management node 110 is a computing node in the server cluster 100 that manages resources in the multiple computing nodes 120. For example, it manages the resource usage of the multiple computing nodes 120 and allocates or reclaims resources in the multiple computing nodes 120. In addition, the management node 110 is also used to receive jobs issued by applications running on the upper-layer client 130 and distribute the jobs to one or more of the managed computing nodes 120 based on the resources in the managed computing nodes 120.
[0066] Exemplarily, resources include computing resources and storage resources. Computing resources include the number of CPU cores on a CPU chip, and may also include computing resources on chips such as graphics processing units (GPUs), embedded neural-network processing units (NPUs), or field programmable gate arrays (FPGAs); storage resources include memory capacity, and may further include resources such as memory bandwidth and L3 cache.
[0067] The computing node 120 is a computing device in the server cluster 100 that is used to allocate corresponding resources to a job distributed by the management node 110 and execute the job under the scheduling of the management node 110 .
[0068] For example, the management node 110 or the computing node 120 may be composed of one or more servers with multi-core processors, which may be rack servers, blade servers, tower servers, or other different types of servers.
[0069] The number of management nodes 110 or computing nodes 120 in the server cluster 100 can be one or more. The embodiment of the present application does not limit the number of servers included in each management node 110 or each computing node 120, nor the number of management nodes 110 and computing nodes 120.
[0070] Optionally, the server cluster 100 proposed in the embodiment of the present application is connected to a client 130 via a network.
[0071] Client 130 is a program that provides local services to users. Generally, client 130 and server work in conjunction with each other. Client 130 can be installed on a common client computer, such as a mobile phone, tablet computer, desktop computer, laptop computer, notebook computer, and netbook computer. The server can be installed on management node 110. Client 130 contains multiple applications (APPs). These applications generate jobs during operation. Client 130 sends these jobs to management node 110, where the server is installed. Management node 110 then processes the jobs sent by the applications on client 130.
[0072] In one embodiment, the management node 110 in the server cluster 100 receives a job request (job request) from the upper-layer client 130 and selects a compute node 120 from the multiple compute nodes 120 it manages whose available resources meet the job's execution requirements to process the job. The management node 110 distributes the job to the selected one or more compute nodes 120. After receiving the job, the compute node 120 allocates the target resources indicated by the management node 110 to the job and executes the job.
[0073] In one embodiment, the server cluster system in which the server cluster 100 resides includes a scheduler. The scheduler is a distributed process running in the server cluster system, which is used to schedule jobs received by the management node 110 and allocate corresponding target resources to the jobs. Exemplarily, the scheduler can be a scheduler such as Slurm or PBS.
[0074] Specifically, as shown in Figure 2, the scheduler has the following modules:
[0075] The job receiving module (slurm cli) 201 is used to receive job requests sent by the upper-layer client and forward the job requests to the resource management center module 202, which then schedules the jobs.
[0076] The resource management center module (slurmctrld) 202 is used to call the storage resource management module based on the resource requirements of the target job to determine the storage resources of multiple computing nodes. In addition, the resource management center module 202 is also used to distribute the job to one or more computing nodes 120 and instruct one or more computing nodes 120 on the specific resources to be allocated when executing the job, and optionally, the CPU to execute the job.
[0077] The storage resource management module 203 is used to manage the storage resources (such as memory bandwidth and / or L3 cache) of multiple computing nodes 120 under the instruction of the resource management center module 202, for example: querying or recording the used storage resources and available storage resources (i.e., unused storage resources) in each computing node 120.
[0078] The job execution module (slurmd) 204 receives jobs issued by the job reception module 201 and, in accordance with the instructions of the job reception module 201, calls the storage resource allocation module 205 to allocate the indicated memory bandwidth to the jobs. Furthermore, the job execution module 204 is responsible for starting and stopping jobs and returning job execution results and resource usage information to the job reception module 201.
[0079] The storage resource allocation module 205 is configured to allocate the required memory bandwidth to the job through the memory bandwidth interface 206 under the instruction of the job execution module 204 and release the allocated memory bandwidth after the job execution is completed.
[0080] The memory bandwidth interface (Intel RDT) 206 is a driver interface for allocating memory bandwidth, which can allocate memory bandwidth and physically isolate the allocated memory bandwidth.
[0081] The job receiving module 201 , resource management center module 202 and storage resource management module run in the management node 110 , and the job execution module 204 , storage resource allocation module 205 and memory bandwidth interface 206 run in the computing node 120 .
[0082] The following describes the job scheduling method provided in the embodiment of the present application:
[0083] As shown in Figure 3, the job scheduling method proposed in an embodiment of the present application is applied to a management node in a server cluster, and mainly includes the following steps: Step 1: The management node obtains the resource requirements of the target job; Step 2: The management node determines one or more target computing nodes based on the resource requirements of the target job, and the one or more target computing nodes meet the resource requirements of the target job; Step 3: The management node sends the target job to one or more target computing nodes so that the target job is executed in one or more target computing nodes.
[0084] The following describes in detail the job scheduling method proposed in an embodiment of the present application. Please refer to Figure 4, which is a flowchart of a job scheduling method provided in an embodiment of the present application. This method is applied to management nodes and computing nodes in a server cluster, such as management node 110 and computing node 120 in Figure 1. As shown in Figure 4, this method may include S101-S109.
[0085] S101: The management node obtains the resource requirements of the target job.
[0086] A job can be a computing task issued by an application when a user runs the application in an upper-layer client.
[0087] In one embodiment, the upper-level client has a resource configuration interface (e.g., a portal). Before running an application on the client, the user configures the resources required for the application in the resource configuration interface. After the configuration is complete, the client runs the application and sends the target job issued during the application's execution, along with the target job's resource requirements, to the management node. The resources required to execute the target job are referred to as target resources. The management node receives the target job and its resource requirements.
[0088] Resource requirements include computing and storage resource requirements. Computing resource requirements include, but are not limited to, the number of CPU cores and GPU cores. Storage resource requirements include memory capacity requirements and may further include memory bandwidth and L3 cache requirements. The following examples use the number of CPU cores, memory capacity, and memory bandwidth as examples.
[0089] For example, as shown in Figure 5, a resource configuration interface is used to configure resources for application 1, such as memory bandwidth, number of CPU cores, and memory capacity. If the target resources required by application 1 to run a target job include: target memory bandwidth of 80 GB / s, target number of CPU cores of 16, and target memory capacity of 1 GB, the corresponding values can be configured in the resource configuration interface.
[0090] In the above example, the target memory bandwidth of 80 GB / s represents the total bandwidth of the target CPU cores, and the single-core memory bandwidth is 5 GB / s. In the embodiment of the present application, the target memory bandwidth information carried by the job request can also be the single-core memory bandwidth. The required total bandwidth is the single-core memory bandwidth multiplied by the target number of CPU cores. For the convenience of subsequent calculations, the following descriptions of memory bandwidth and target memory bandwidth are all based on the total bandwidth as an example.
[0091] S102: The management node determines one or more target computing nodes based on the resource requirements of the target job, and the one or more target computing nodes meet the resource requirements of the target job.
[0092] When the management node determines a target computing node, the available computing resources and available storage resources of the target computing node meet the resource requirements of the target job. When the management node determines multiple target computing nodes, the sum of the available computing resources and the sum of the available storage resources of the multiple target computing nodes respectively meet the resource requirements of the target job.
[0093] The available computing resources mentioned above are unallocated computing resources, and the available storage resources are unallocated storage resources.
[0094] A possible implementation scheme for S102 is proposed, including: S102a-S102b.
[0095] S102a: The management node obtains available resource information of multiple computing nodes.
[0096] Available resource information includes available computing resource information and available storage resource information.
[0097] Exemplarily, the available computing resource information includes the number of available CPU cores, and the available storage resource information includes available memory capacity and available memory bandwidth.
[0098] In one embodiment, the management node stores a resource information table, which is used to record resource information of multiple computing nodes. The resource information includes available resource information. The management node obtains available resource information of multiple computing nodes based on the resource information of multiple computing nodes recorded in the resource information table.
[0099] Optionally, the resource information table also includes allocated resource information, which includes allocated computing resource information and allocated storage resource information. For example, the allocated computing resource information includes the allocated number of CPU cores, and the allocated storage resource information includes the allocated memory capacity and allocated memory bandwidth.
[0100] For example, if the total memory bandwidth of each computing node is 200 GB / s, the total number of CPU cores is 32, and the total memory capacity is 16 GB, the management node records the available resource information and allocated resource information of each computing node through a resource information table. As shown in Table 1, Table 1 shows a resource information table that records the available resource information and allocated resource information of multiple computing nodes. The resource information table includes: "Computing node identification", "Memory bandwidth", "CPU core number" and "Memory capacity", where "Memory bandwidth" includes "Available memory bandwidth" and "Allocated memory bandwidth", "CPU core number" includes "Available CPU core number" and "Allocated CPU core number", and "Memory capacity" includes "Available memory capacity" and "Allocated memory capacity".
[0101] Table 1
[0102] The identifier of the computing node may be a computing node number, a host number, an IP address, a MAC address, or other content that can represent the identity of the computing node. Table 1 above uses the computing node number as an example for explanation.
[0103] S102b: The management node determines one or more target computing nodes based on the available resource information of each computing node.
[0104] The available resources of one or more target computing nodes are greater than or equal to the resource requirements of the target job.
[0105] If there is a computing node whose available resources are greater than or equal to the resource requirements of the target job, then the computing node is the target computing node.
[0106] If the available resources of each of the multiple computing nodes are less than the resource requirements of the target job, and the sum of the available resources of the multiple computing nodes is greater than or equal to the resource requirements of the target job, then the multiple computing nodes are all target computing nodes.
[0107] For example, if the target resources include: target memory bandwidth of 80GB / S, target number of CPU cores of 16, and target memory capacity of 1G, based on the resource information recorded in the resource information table shown in Table 1, the management node queries the available memory bandwidth, available number of CPU cores and available memory capacity of each computing node, and determines that the available memory bandwidth of "computing node 1" is 80GB / s, the available number of CPU cores is 16, and the available memory capacity is 8G, which can meet all requirements greater than or equal to the target resources, while the available memory bandwidth of "computing node 2" is 60GB / s, which is less than the target memory bandwidth, and the available number of CPU cores is 8, which is less than the target number of CPU cores, and does not meet the requirements of the target memory bandwidth and target number of CPU cores in the target resources; the available number of CPU cores of "computing node 3" is 12, which is less than the target number of CPU cores, and does not meet the requirements of the target number of CPU cores in the target resources. Therefore, the management node can determine that the target computing node is "computing node 1" from the multiple computing nodes recorded in the resource information table.
[0108] For example, if the target resources include: target memory bandwidth of 120 GB / S, target number of CPU cores of 24, and target memory capacity of 2G, the management node queries the available memory bandwidth, available number of CPU cores, and available memory capacity of each computing node, and determines that the available resources of each computing node are less than the target resources, and the sum of the available memory bandwidth of "computing node 1" and "computing node 2" is 140 GB / S, the sum of the available number of CPU cores is 24, and the sum of the available memory capacity is 18G, or the sum of the available memory bandwidth of "computing node 1" and "computing node 3" is 180 GB / S, the sum of the available number of CPU cores is 2 8, and the total available memory capacity is 16 GB. Alternatively, the total available memory bandwidth of Compute Node 2 and Compute Node 3 is 160 GB / S, the total number of available CPU cores is 20, and the total available memory capacity is 16 GB. Therefore, the total available resources of Compute Node 1 and Compute Node 2, or Compute Node 1 and Compute Node 3, or Compute Node 2 and Compute Node 3, are greater than or equal to the target resources. Therefore, Compute Node 1 and Compute Node 2, or Compute Node 1 and Compute Node 3, or Compute Node 2 and Compute Node 3 are determined as the target compute nodes.
[0109] It should be noted that if the available resources of each of the multiple computing nodes in Table 1 meet the target resource requirements of the target job, the management node can select one of the multiple computing nodes that meet the requirements as the target computing node based on the principle of load balancing. Alternatively, the management node can randomly select one as the target computing node. This embodiment of the application does not limit the specific strategy of how the management node determines the target computing node from the multiple computing nodes that meet the target resources.
[0110] Optionally, after determining one or more target computing nodes, the management node updates the available resource information of the one or more target computing nodes recorded in the resource information table based on the resources occupied by the one or more target computing nodes for processing the target job.
[0111] In one embodiment, the management node updates the target memory bandwidth portion from the available memory bandwidth to the allocated memory bandwidth, updates the target CPU core count portion from the available CPU core count to the allocated CPU core count, and updates the target memory capacity from the available memory capacity to the allocated memory capacity in the resource information table.
[0112] For example, if the target memory bandwidth is 80 GB / s, the target number of CPU cores is 16, and the target memory capacity is 1 GB, the available memory bandwidth of "Compute Node 1" in Table 1 is 80 GB / s, the available number of CPU cores is 16, and the available memory capacity is 8 GB. After the management node determines that the target compute node is "Compute Node 1", it updates the available memory bandwidth of "Compute Node 1" from 80 GB / s to 0 GB / s, updates the allocated memory bandwidth from 120 GB / s to 200 GB / s, updates the available number of CPU cores from 16 to 0, updates the allocated number of CPU cores from 16 to 32, updates the available memory capacity from 8 GB to 7 GB, and updates the allocated memory capacity from 8 GB to 9 GB. As shown in Table 2, Table 2 shows the updated resource information table.
[0113] Table 2
[0114] It can be understood that if the management node determines the target computing node, it means that the management node instructs to allocate the resources of the target computing node to the target job, and updates the available resource information of the target computing node in the resource information table. This is beneficial for subsequent other jobs to be allocated based on the latest resource information in a timely manner, avoiding job execution failure due to insufficient available resources.
[0115] In the above optional implementation method, the management node may update the resource information table after determining the target computing node, or may update the resource information table after issuing the target job. The embodiment of the present application does not limit the execution order.
[0116] The above-mentioned management node manages the resources of multiple computing nodes through a resource information table, can clearly grasp the resource allocation status of multiple computing nodes and the number of remaining available resources, and can allocate jobs to computing nodes with sufficient available resources based on the resource information table, which is conducive to improving the management efficiency of the management node on the resources on the computing nodes.
[0117] Optionally, the management node sends a first indication to the multiple computing nodes respectively; the first indication is used to instruct the computing nodes to return their respective available resource information; based on the available resource information returned by the multiple computing nodes, the available resource information of the multiple computing nodes recorded in the resource information table is corrected.
[0118] The management node may periodically or regularly send a first instruction to the plurality of computing nodes to correct the available resource information of the plurality of computing nodes recorded in the resource information table, thereby improving the accuracy of the available resource information of the plurality of computing nodes recorded in the resource information table.
[0119] S103: The management node sends the target job to one or more target computing nodes, so that the target job is executed in the one or more target computing nodes.
[0120] In one embodiment, if the management node determines a target computing node, the management node sends the target job to the target computing node.
[0121] In another embodiment, if the management node determines multiple target computing nodes, the following embodiment is proposed for S103, including S103a-S103b.
[0122] S103a: The management node splits the target job into multiple sub-jobs.
[0123] When splitting a target job, the management node splits the target job into multiple sub-jobs based on the available resources of each target computing node, wherein the available resources of the target computing node executing each sub-job are greater than or equal to the resources required by the sub-job to be executed.
[0124] For example, if the management node determines that two target computing nodes jointly execute the target job, the management node can split the target job into two sub-jobs, wherein the management node determines that sub-job 1 is executed by target computing node 1 and sub-job 2 is executed by target computing node 2, the available resources of target computing node 1 are greater than or equal to the resources required for sub-job 1, and the available resources of target computing node 2 are greater than or equal to the resources required for sub-job 2.
[0125] S103b: The management node sends task information to one or more target computing nodes. The task information includes the sub-jobs to be executed and the resources required to execute the sub-jobs to be executed.
[0126] For example, if the target job requires resources including a target memory bandwidth of 120 GB / S, a target number of CPU cores of 24, and a target memory capacity of 2 GB, the management node splits the target job into two sub-jobs, namely sub-job 1 and sub-job 2. The resources required for sub-job 1 include: a memory bandwidth of 40 GB / S, a target number of CPU cores of 12, and a memory capacity of 1 GB; the resources required for sub-job 2 include: a memory bandwidth of 80 GB / S, a target number of CPU cores of 12, and a memory capacity of 1 GB. The management node sends task information to target computing node 1. The task information includes "sub-job 1" and the resources required to execute "sub-job 1": a memory bandwidth of 40 GB / S, a target number of CPU cores of 12, and a memory capacity of 1 GB. The management node also sends task information to target computing node 2. The task information includes "sub-job 2" and the resources required to execute "sub-job 2": a memory bandwidth of 80 GB / S, a target number of CPU cores of 12, and a memory capacity of 1 GB.
[0127] Optionally, after determining the target computing node, the management node further determines one or more CPUs on the target computing node that execute the target job based on the target computing resources required by the target job.
[0128] At this time, the task information sent by the management node to one or more target computing nodes includes: the resources required to execute the target job and one or more CPUs to execute the target job; the resources required for the target job include target computing resources and target storage resources, and the target computing resources include the target number of CPU cores.
[0129] For example, if the target computing node includes CPU1 and CPU2, the number of available CPU cores of CPU1 is 32, the number of available CPU cores of CPU2 is 20, and the resources required by the target job include: target memory bandwidth of 120GB / S, target number of CPU cores is 24, and target memory capacity is 2G, after the management node determines to send the target job to the target computing node, it determines that the CPU executing the target job on the target computing node is CPU1 based on the target number of CPU cores 24 required by the target job. Then the task information sent by the management node to the target computing node includes: target job, resources required by the target job include: target memory bandwidth of 120GB / S, target number of CPU cores is 24, target memory capacity is 2G, and the CPU executing the target job is CPU1.
[0130] S104: The target computing node receives the task information sent by the management node.
[0131] Task information includes the job to be executed and target resources. The job to be executed is the target job, or a sub-job of the target job. The target resources are the resources required to execute the job to be executed, and the target resources include target computing resources and target storage resources.
[0132] Optionally, the task information also includes one or more CPUs that execute the job to be executed.
[0133] It can be understood that if the management node determines a target computing node to execute the target job, the job to be executed received by the target computing node is the target job. If the management node determines multiple target computing nodes to jointly execute the target job, the target computing node is any one of the multiple target computing nodes, and the job to be executed received by the target computing node is a sub-job of the target job.
[0134] S105: The target computing node distributes the job to be executed to one or more CPUs.
[0135] Exemplarily, the task information includes: the job to be executed is CPU1, and the target computing node distributes the job to be executed to CPU1 based on the task information.
[0136] In some other implementation scenarios, if the management node does not specify the CPU to execute the job to be executed, that is, the task information does not include one or more CPUs to execute the job to be executed, the target computing node determines one or more CPUs to execute the job to be executed based on the target computing resources required by the target job.
[0137] S106: The target computing node allocates target resources to the job to be executed, and executes the job to be executed.
[0138] Optionally, the target computing node allocates target resources on one or more CPUs to the job to be executed, and executes the job to be executed on the one or more CPUs.
[0139] For example, the target resources include: a target memory bandwidth of 120 GB / s, a target number of CPU cores of 24, a target memory capacity of 2 GB, the CPU that executes the target job is CPU1, the target computing node distributes the target job to CPU1, the 120 GB / s memory bandwidth, 24 CPU cores, and 2 GB memory capacity on CPU1 are allocated to the target job, and the target job is executed.
[0140] Optionally, after allocating target resources to the target job, the target computing node physically isolates the target memory bandwidth.
[0141] It is understandable that the target computing node physically isolates the target memory bandwidth, so that when the target computing node executes multiple jobs simultaneously, each job only uses its allocated memory bandwidth, effectively avoiding resource competition among multiple jobs.
[0142] S107: After the target computing node completes executing the target job, it releases the target resources.
[0143] After the target computing node releases the target resources, the physical isolation of the target memory bandwidth is lifted.
[0144] The target computing node releases the target resources in a timely manner, and can subsequently continue to allocate these resources to other jobs, which is beneficial to improving resource utilization.
[0145] S108: The target computing node sends a response message indicating that the execution of the target job is completed to the management node, so that the management node reclaims the target resources.
[0146] The above S107 can be executed before S108, or after S108, or can be executed simultaneously. The embodiment of the present application does not limit the execution order of S107 and S108.
[0147] S109: Upon receiving a response message indicating that the target job has been completed, the management node reclaims resources allocated to the target job on one or more target computing nodes, and updates the available resource information of the one or more target computing nodes recorded in the resource information table accordingly.
[0148] In one embodiment, in one embodiment, the management node updates the target memory bandwidth portion from the allocated memory bandwidth to the available memory bandwidth, updates the target CPU core number portion from the allocated CPU core number to the available CPU core number, and updates the target memory capacity from the allocated memory capacity to the available memory capacity in the resource information table.
[0149] For example, if the target memory bandwidth is 80 GB / s, the target number of CPU cores is 16, and the target memory capacity is 1 GB, the available memory bandwidth of "Compute Node 1" in Table 2 is 0 GB / s, the available number of CPU cores is 0, and the available memory capacity is 7 GB. After the management node receives the response message indicating that the target job has been executed, it reclaims the resources allocated to the target job in "Compute Node 1" and restores the available memory bandwidth from 0 GB / s to 80 GB / s, updates the allocated memory bandwidth from 200 GB / s to 120 GB / s, restores the available number of CPU cores from 0 to 16, updates the allocated number of CPU cores from 32 to 16, restores the available memory capacity from 7 GB to 8 GB, and updates the allocated memory capacity from 9 GB to 8 GB. That is, if the resources of other compute nodes do not change, the resource information table is restored from Table 2 to Table 1.
[0150] In this method, the management node promptly updates the resource information of the target computing node recorded in the resource information table, which is beneficial for obtaining the latest resource information of multiple computing nodes when there are new jobs in the future, thereby improving the utilization rate of resources in the computing nodes.
[0151] In the job scheduling method proposed in the embodiment of the present application, after the management node obtains the resource requirements of the target job, it determines one or more target computing nodes that meet the requirements of the target job from multiple computing nodes based on the target resources required by the target job, and distributes the target job to one or more target computing nodes, so that one or more target computing nodes use the target resources to execute the target job. It can be understood that in this method, when the target computing node executes the target job, the resources allocated to the target job are the resources required by the target job. In addition, the storage resource requirements are also taken into account when allocating resources. Therefore, this method can flexibly allocate resources according to the actual job situation and improve resource utilization.
[0152] In one example, Job 1 is a memory-intensive job requiring 8 GB / s of single-core memory bandwidth and 2 CPU cores. Job 2 is a compute-intensive job requiring 2 GB / s of single-core memory bandwidth and 6 CPU cores. If the target computing node has 8 CPU cores and a single-core memory bandwidth of 10 GB / s, running Job 1 and Job 2 without considering the memory bandwidth requirement is shown in FIG6 , which shows the performance graphs of Job 1 and Job 2 before using the method of the embodiment of the present application. In the embodiment of the present application, running Job 1 and Job 2 with the memory bandwidth requirement in mind is shown in FIG7 , which shows the performance graphs of Job 1 and Job 2 after using the method of the embodiment of the present application.
[0153] As can be seen from Figure 6, if the memory bandwidth allocation is not considered, the running time of Job 1 is 1000 seconds, and the running time of Job 2 is 1000 seconds. As can be seen from Figure 7, if the memory bandwidth allocation is considered, the running time of Job 1 is 650 seconds, and the running time of Job 2 is 1000 seconds. Therefore, considering the memory bandwidth allocation, the running time of Job 1 is shortened, and the running efficiency of the job is improved. Therefore, the job calling method proposed in the embodiment of the present application can improve resource utilization. In the same time, the computing node can process more jobs, thereby improving the overall throughput and performance of the server cluster.
[0154] In conjunction with FIG2 , an embodiment of the present application proposes a flowchart of another job scheduling method, as shown in FIG8 , the method includes: S201 - S209 .
[0155] S201: The job receiving module of the management node obtains the resource requirements of the target job and sends them to the resource management center module.
[0156] S202: The resource management center module in the management node determines one or more target computing nodes based on the resource requirements of the target job, and the one or more target computing nodes meet the resource requirements of the target job.
[0157] Specifically, the resource management center module in the management node calls the storage resource management module, and the storage resource management module determines the computing node whose available memory bandwidth is greater than the target memory bandwidth from multiple computing nodes, and sends the identifier of the determined computing node to the resource management center module, which further determines the target computing node.
[0158] In this method, the management node manages the memory bandwidth of multiple compute nodes through the storage resource management module. This information is fed back to the resource management center module, enabling the resource management center to distribute target jobs to appropriate compute nodes, thereby improving memory bandwidth utilization. This method allows the addition of a separate memory bandwidth management module without changing the original storage resource management module's functionality, making it simple to operate and easy to implement.
[0159] S203: The resource management center module in the management node sends the target job to one or more target computing nodes, so that the target job is executed in the one or more target computing nodes.
[0160] In one embodiment, the resource management center module in the management node sends the task information.
[0161] S204: The job execution module in the target computing node receives the task information sent by the management node.
[0162] S205: The job execution module in the target computing node distributes the job to be executed to one or more CPUs.
[0163] S206: The job execution module in the target computing node allocates target resources to the job to be executed, and executes the job to be executed.
[0164] The target resources include: computing resources and storage resources required by the target job, wherein the computing resources are allocated by the job execution module and the storage resources are allocated by the storage resource allocation module.
[0165] Specifically, the storage resource allocation module allocates storage resources including S206a-S206c:
[0166] S206a: The job execution module in the target computing node sends the storage resources required for executing the target job to the storage resource allocation module, so that the storage resource allocation module allocates storage resources for the target job.
[0167] The storage resources required by the target job include target memory bandwidth.
[0168] S206b: The storage resource allocation module in the target computing node allocates target memory bandwidth to the target job through the memory bandwidth interface, and physically isolates the target memory bandwidth.
[0169] The storage resource allocation module is used to allocate memory bandwidth, and the memory bandwidth interface is the driver for implementing memory bandwidth allocation. The computing node implements the actual allocation of memory bandwidth through these two modules.
[0170] S206c: After allocating the storage resources required by the target job to the target job, the storage resource allocation module in the target computing node sends a message indicating that the storage resource allocation is completed to the job execution module.
[0171] S207: After the job execution module in the target computing node finishes executing the target job, it sends a response message indicating that the target job has been completed to the storage resource allocation module and the management node respectively.
[0172] S208: After receiving the response message indicating that the target job has been executed, the storage resource allocation module in the target computing node releases the target memory bandwidth through the memory bandwidth interface.
[0173] S209: After receiving the response message indicating that the target job has been executed, the resource management center module in the management node reclaims the target resources allocated to the target job on the target computing node.
[0174] Specifically, the resource management center module in the management node calls the storage resource management module to reclaim the target memory bandwidth allocated to the target job on the target computing node.
[0175] The above S208 can be executed before S209, or after S209, or can be executed simultaneously. The embodiment of the present application does not limit the execution order of S208 and S209.
[0176] For detailed description of the above S201-S209, please refer to S101-S109.
[0177] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. It should be easy to realize that the technical goals in this field are combined with the units and algorithm steps of each example described in the embodiments disclosed herein, and the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technical goals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0178] The embodiment of the present application further provides a management node 300, such as the management node 110 in Figure 1. Figure 9 is a schematic diagram of the structure of a management node 300 provided in the embodiment of the present application.
[0179] Among them, the management node 300 includes: an acquisition unit 301, used to obtain the resource requirements of the target job; the resource requirements include computing resource requirements and storage resource requirements; a processing unit 302, used to determine one or more target computing nodes based on the resource requirements of the target job, and one or more target computing nodes meet the resource requirements of the target job; a sending unit 303, used to send the target job to one or more target computing nodes so that the target job is executed in one or more target computing nodes.
[0180] In some embodiments, the processing unit 302 is specifically used to obtain available resource information of multiple computing nodes, where the available resource information includes available computing resource information and available storage resource information; based on the available resource information of each computing node, determine one or more target computing nodes; wherein the available resources of the one or more target computing nodes are greater than or equal to the resource requirements of the target job.
[0181] In some implementations, the management node stores a resource information table, which is used to record resource information of multiple computing nodes. The resource information includes available resource information. The acquisition unit 301 is specifically used to obtain available resource information of multiple computing nodes based on the resource information of multiple computing nodes recorded in the resource information table.
[0182] In some embodiments, the sending unit 303 is further used to send a first indication to multiple computing nodes respectively; the first indication is used to instruct the computing nodes to return their respective available resource information; the processing unit 302 is further used to correct the available resource information of multiple computing nodes recorded in the resource information table based on the available resource information returned by the multiple computing nodes.
[0183] In some embodiments, after the target job is sent to one or more target computing nodes so that the target job is executed in the one or more target computing nodes, the processing unit 302 is also used to update the available resource information of the one or more target computing nodes recorded in the resource information table based on the resources occupied by the one or more target computing nodes for processing the target job.
[0184] In some embodiments, the management node stores a resource information table, which is used to record resource information of multiple computing nodes. The processing unit 302 is also used to, upon receiving a response message indicating that the target job has been executed, reclaim the resources allocated to the target job on one or more target computing nodes, and update the available resource information of the one or more target computing nodes recorded in the resource information table accordingly.
[0185] In some implementations, the sending unit 303 is specifically configured to split the target job into multiple sub-jobs; and send task information to one or more target computing nodes, where the task information includes the sub-jobs to be executed and the resources required to execute the sub-jobs to be executed.
[0186] In some embodiments, before sending the target job to one or more target computing nodes, the processing unit 302 is also used to determine one or more CPUs to execute the target job based on the target computing resources required by the target job; the sending unit 303 is specifically used to send task information to one or more target computing nodes, and the task information includes: the resources required to execute the target job and one or more CPUs to execute the target job; the resources required for the target job include target computing resources and target storage resources, and the target computing resources include the target number of CPU cores.
[0187] In some implementations, the computing resource requirement includes a CPU core count requirement, and the storage resource requirement includes a memory bandwidth requirement.
[0188] In some implementations, the storage resource requirement includes a memory bandwidth requirement and an L3 cache requirement.
[0189] The embodiment of the present application further provides a computing node 400, such as the computing node 120 in Figure 1. Figure 10 is a schematic diagram of the structure of a computing node 400 provided in the embodiment of the present application.
[0190] Among them, the computing node 400 includes: a receiving unit 401, used to receive task information sent by the management node; the task information includes the job to be executed, and target resources, the job to be executed is the target job, or a sub-job of the target job, the target resources are the resources required to execute the job to be executed, and the target resources include target computing resources and target storage resources; a processing unit 402, used to allocate target resources to the job to be executed and execute the job to be executed.
[0191] In some embodiments, the task information includes one or more CPUs that execute the job to be executed, and target resources are allocated to the job to be executed. Before executing the job to be executed, the processing unit 402 is also used to distribute the job to be executed to one or more CPUs based on the task information; allocate target resources on one or more CPUs to the job to be executed, and execute the job to be executed on one or more CPUs.
[0192] In some implementations, the target storage resources include target memory bandwidth. After allocating the target resources to the job to be executed, the processing unit 402 is further configured to physically isolate the target memory bandwidth.
[0193] Of course, the management node 300 or computing node 400 provided in the embodiment of the present application includes but is not limited to the above modules.
[0194] FIG11 is a schematic diagram of the structure of a server 500 provided in an embodiment of the present application. As shown in FIG11 , the server 500 includes a processor 501 , a memory 502 , and a network interface 503 .
[0195] The processor 501 includes one or more CPUs, which may be single-core CPUs (single-CPU) or multi-core CPUs (multi-CPU).
[0196] The memory 502 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or optical storage.
[0197] In some implementations, the processor 501 implements the job scheduling method provided in the embodiments of the present application by reading instructions stored in the memory 502, or the processor 501 implements the job scheduling method provided in the embodiments of the present application by internally stored instructions. In the case where the processor 501 implements the method in the above embodiment by reading instructions stored in the memory 502, the memory 502 stores instructions that implement the job scheduling method provided in the embodiments of the present application.
[0198] Network interface 503, a device including a transmitter and a receiver, is used to communicate with other devices or communication networks. It can be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface. Alternatively, network interface 503 can be a wireless interface. It should be understood that network interface 503 includes multiple physical ports and is used for communication, etc.
[0199] In some implementations, the server 500 further includes a bus 504 , and the processor 501 , memory 502 , and network interface 503 are typically interconnected via the bus 504 , or are interconnected in other ways.
[0200] In actual implementation, the acquisition unit 301, the processing unit 302, and the sending unit 303, or the receiving unit 401 and the processing unit 402, can be implemented by a processor calling computer program codes in a memory. The specific execution process can be referred to the description of the method part above and will not be repeated here.
[0201] Another embodiment of the present application provides a server comprising a memory and a processor. The memory and the processor are coupled; the memory is configured to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the server executes the steps of the job scheduling method described in the above method embodiment.
[0202] Another embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on a server, the server executes each step executed by the server in the job scheduling method process shown in the above method embodiment.
[0203] Another embodiment of the present application provides a chip system for use in a server. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via circuits. The interface circuits are configured to receive signals from the server's memory and send signals to the processors, the signals including computer instructions stored in the memory. When the server's processor executes the computer instructions, the server executes the steps performed by the server in the job scheduling method flow shown in the above method embodiment.
[0204] In another embodiment of the present application, a computer program product is provided. The computer program product includes computer instructions. When the computer instructions are executed on a server, the server executes each step executed by the server in the job scheduling method process shown in the above method embodiment.
[0205] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer-executable instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a server, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).
[0206] The above is only a specific embodiment of the present application. Those skilled in the art may conceive of changes or substitutions based on the specific embodiment provided in this application, and all such changes or substitutions shall fall within the scope of protection of this application.
Claims
1. A job scheduling method, characterized in that, Applied to a management node in a server cluster, the management node communicates with multiple computing nodes in the server cluster, and the method includes: Obtaining the resource requirements of a target job; the resource requirements include computing resource requirements and storage resource requirements; Determining one or more target computing nodes based on the resource requirements of the target job, where the one or more target computing nodes meet the resource requirements of the target job; Issuing the target job to the one or more target computing nodes so that the target job is executed in the one or more target computing nodes.
2. The method according to claim 1, wherein The determining one or more target computing nodes based on the resource requirements of the target job, where the one or more target computing nodes meet the resource requirements of the target job, includes: Obtaining the available resource information of the multiple computing nodes, where the available resource information includes available computing resource information and available storage resource information; Determining one or more target computing nodes based on the available resource information of each computing node; where the available resources of the one or more target computing nodes are greater than or equal to the resource requirements of the target job.
3. The method according to claim 2, characterized in that, The management node stores a resource information table, which is used to record the resource information of the multiple computing nodes, and the resource information includes the available resource information. The obtaining the available resource information of the multiple computing nodes includes: Obtaining the available resource information of the multiple computing nodes based on the resource information of the multiple computing nodes recorded in the resource information table.
4. The method according to claim 3, wherein The method further includes: Sending a first instruction to each of the multiple computing nodes respectively; the first instruction is used to instruct the computing node to return its respective available resource information; Based on the available resource information returned by the multiple computing nodes, correcting the available resource information of the multiple computing nodes recorded in the resource information table.
5. The method according to claim 3 or 4, characterized in that, After the issuing the target job to the one or more target computing nodes so that the target job is executed in the one or more target computing nodes, the method further includes: Updating the available resource information of the one or more target computing nodes recorded in the resource information table based on the resources occupied by the one or more target computing nodes for processing the target job.
6. The method according to any one of claims 1 to 5, characterized in that, The management node stores a resource information table, which is used to record the resource information of the multiple computing nodes, and the method further includes: When receiving a response message indicating that the target job has been executed, reclaiming the resources allocated to the target job on the one or more target computing nodes and correspondingly updating the available resource information of the one or more target computing nodes recorded in the resource information table.
7. The method according to any one of claims 1 to 6, characterized in that, The issuing the target job to the one or more target computing nodes includes: Splitting the target job into multiple sub-jobs; Sending task information to the one or more target computing nodes, where the task information includes the sub-jobs to be executed and the resources required to execute the sub-jobs to be executed.
8. The method according to any one of claims 1 to 6, characterized in that Before the issuing the target job to the one or more target computing nodes, the method further includes: Determine one or more central processing units (CPUs) for executing the target job based on the target computing resources required by the target job; The step of sending the target job to the one or more target computing nodes includes: Sending task information to the one or more target computing nodes, where the task information includes: the resources required to execute the target job and one or more CPUs for executing the target job; the resources required by the target job include the target computing resources and target storage resources, and the target computing resources include the target number of CPU cores.
9. The method according to any one of claims 1 to 8, characterized in that, The computing resource requirement includes the requirement for the number of CPU cores, and the storage resource requirement includes the requirement for memory bandwidth.
10. The method according to any one of claims 1 to 8, characterized in that The storage resource requirement includes the requirement for memory bandwidth and the requirement for level 3 cache (L3 cache).
11. A job scheduling method, characterized in that, The method is applied to a computing node in a server cluster, and the computing node communicates with a management node in the server cluster. The method includes: Receiving task information sent by the management node; the task information includes the job to be executed and target resources, the job to be executed is the target job or a sub-job of the target job, and the target resources are the resources required to execute the job to be executed, and the target resources include target computing resources and target storage resources; Allocating the target resources for the job to be executed and executing the job to be executed.
12. The method according to claim 11, wherein The task information includes one or more central processing units (CPUs) for executing the job to be executed. Before allocating the target resources for the job to be executed and executing the job to be executed, the method further includes: Distributing the job to be executed to the one or more CPUs based on the task information; The step of allocating the target resources for the job to be executed and executing the job to be executed includes: Allocating the target resources on the one or more CPUs for the job to be executed and executing the job to be executed on the one or more CPUs.
13. The method according to claim 11 or 12, characterized in that, The target storage resources include the target memory bandwidth. After allocating the target resources for the job to be executed, the method further includes: Physically isolating the target memory bandwidth.
14. A server cluster, characterized in that, It includes multiple servers, and the multiple servers are divided into a management node and multiple computing nodes; The management node includes: a job receiving module, a resource management center module, and a storage resource management module; The job receiving module is configured to receive a job request and send the job request to the resource management center module; the job request is used to request resource allocation for a target job and execute the target job; The resource management center module is configured to call the storage resource management module based on the resource requirements of the target job to determine the storage resources of the multiple computing nodes; the resource requirements include computing resource requirements and storage resource requirements; the resource management center module is further configured to determine one or more target computing nodes for executing the target job based on the resources of the multiple computing nodes; The storage resource management module is configured to manage the storage resources of the multiple computing nodes; The computing node includes: a job execution module, a storage resource allocation module, and a memory bandwidth interface; The job execution module is configured to receive jobs distributed by the resource management center module, and under the instruction of the resource management center module, call the storage resource allocation module to allocate the indicated storage resources for the jobs; the storage resources include memory bandwidth; The storage resource allocation module is configured to, under the instruction of the job execution module, allocate the indicated storage resources for the jobs through the memory bandwidth interface; The storage resource allocation module is further configured to release the storage resources allocated for the jobs after the jobs are completed; The memory bandwidth interface is configured to allocate the memory bandwidth and physically isolate the allocated memory bandwidth.
15. A server, characterized in that, It includes a memory and a processor; the memory is coupled to the processor; the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the server executes the computer instructions, the server is caused to execute the method according to any one of claims 1-13.
Citation Information
Patent Citations
Resource distribution method and device
CN105893142A
Task processing method and device, computing equipment and computer storage medium
CN111866043A
Cluster resource scheduling method and device, computer equipment and storage medium
CN113535332A
Job scheduling method, server and server cluster
CN117950825A
Task scheduling method and apparatus, electronic device, storage medium, and program product
WO2022198853A1