Task scheduling method and apparatus
By setting up computing resources and scheduling tasks for each tenant in the computing cluster, the problem of poor tenant isolation in the multi-tenant shared scheduler is solved, logical isolation between tenants and efficient task scheduling is realized, and the performance and fairness of the computing cluster are improved.
Patent Information
- Application Number
- PCT/CN2024/116401
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2024-09-02
- Publication Date
- 2025-07-24
AI Technical Summary
In the scenario where multi-tenant shared scheduler is shared, the isolation between tenants in the prior art is poor, resulting in mutual influence of job scheduling processes and affecting the task scheduling performance of the computing cluster.
Set up computing resources for each tenant in the computing cluster and start scheduling tasks for it. By executing tenants' scheduling tasks, scheduling job tasks to the set computing resources, logical isolation between tenants, and task scheduling is performed according to the tenant's scheduling strategy and priority.
It realizes logical isolation between tenants in the multi-tenant shared scheduler scenario, improves the performance and fairness of multi-tenant task scheduling in the computing cluster, and ensures that high-priority tenants have more scheduling opportunities.
Smart Images

Figure CN2024116401_24072025_PF_FP_ABST
Abstract
Description
Task scheduling method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 19, 2024, with application number 202410082980.8 and application name “A Task Scheduling Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a task scheduling method and device. Background Art
[0003] With the in-depth development of cloud computing, cloud services have become the development direction. Multi-tenancy is the core architecture of cloud services. Multi-tenant architecture refers to an architecture in which multiple tenants share a software instance.
[0004] In a high-performance computing (HPC) cluster, one software instance is the scheduler. As the core management software for an HPC cluster, the scheduler primarily helps tenants manage resources within the HPC cluster and schedule jobs for execution. For example, the scheduler receives tenant jobs, assigns them to the cluster's compute nodes using a specific scheduling policy so that the compute nodes can process the user's jobs, and monitors and manages the job lifecycle.
[0005] Currently, in a multi-tenant architecture, a single scheduler schedules jobs for multiple tenants. In this multi-tenant shared scheduler solution, the isolation between multiple tenants is poor, and the job scheduling processes of different tenants will affect each other.
[0006] Summary of the Invention
[0007] The present application provides a task scheduling method and device, which can independently schedule tasks for tenants in a scenario where multiple tenants share a scheduler, thereby achieving logical isolation between tenants.
[0008] This application adopts the following technical solutions:
[0009] In the first aspect, the present application provides a task scheduling method, which is applied to a management node in a computing cluster (specifically a scheduler deployed on the management node), the method comprising: setting a first computing resource for a first tenant in the computing cluster; when the job task of the first tenant needs to be executed, starting a first scheduling task for the first tenant, and the first scheduling task is used to schedule computing resources for the job task of the first tenant; then obtaining the job task of the first tenant, and sending the job task of the first tenant to the task queue of the first tenant; and executing the first scheduling task to schedule the first computing resource for the job task in the task queue of the first tenant to execute the job task of the first tenant.
[0010] In this application, since computing resources can be set for tenants of the computing cluster and scheduling tasks can be started for tenants, the tenants' job tasks can be scheduled to the computing resources set for them by executing the tenants' scheduling tasks to execute the tenants' job tasks. In the scenario where multiple tenants share a scheduler, tasks can be scheduled independently for tenants, thereby achieving logical isolation between tenants and improving the performance of task scheduling for multiple tenants in the computing cluster.
[0011] In a possible implementation, the task scheduling method provided in the present application further includes: setting a scheduling policy for the first tenant.
[0012] In a possible implementation, the scheduling of the first computing resource for the job task in the task queue of the first tenant includes: scheduling the first computing resource for the job task in the task queue according to a scheduling policy.
[0013] This application supports configuring scheduling strategies for tenants. Different scheduling strategies are customized for different tenants according to their different job scenarios, so as to schedule tasks for tenants based on their scheduling strategies.
[0014] In a possible implementation, the task scheduling method provided in the present application further includes: setting a tenant priority for the first tenant, and the tenant priority can be used to determine the scheduling priority of the first scheduling task.
[0015] In a possible implementation, the executing the first scheduling task includes: executing the first scheduling task according to a scheduling priority.
[0016] This application supports configuring tenant priorities for tenants so that tasks can be scheduled based on tenant priorities, ensuring fairness in task scheduling between tenants and ensuring that high-priority tenants have more scheduling opportunities.
[0017] In one possible implementation, executing the first scheduling task according to the scheduling priority includes: determining whether the first scheduling task in the priority queue is the highest scheduling priority, and the priority queue also includes the second scheduling task of the second tenant; if the first scheduling task is the highest scheduling priority, executing the first scheduling task; if the first scheduling task is not the highest scheduling priority, executing the second scheduling task.
[0018] In one possible implementation, the task scheduling method provided herein further includes dynamically updating the scheduling priority of the first scheduled task based on the elapsed execution time of the first scheduled task. Dynamically updating the scheduling priority of the scheduled task can ensure fairness in task scheduling among tenants and guarantee that high-priority tenants receive more scheduling opportunities.
[0019] In one possible implementation, the computing cluster also includes a second tenant, and the second computing resources of the second tenant share the physical computing resources of the computing cluster with the first computing resources of the first tenant. The task scheduling method provided in this application also includes: dividing the first scheduling task and the second scheduling task of the second tenant into multiple subtasks respectively. In this way, the first scheduling task is executed according to the scheduling priority, including: if the first scheduling task has the highest scheduling priority, executing the first subtask of the first scheduling task, scheduling the first computing resource for the job task in the task queue of the first tenant to execute the job task of the first tenant; then updating the scheduling priority of the first scheduling task; if the first scheduling task after the scheduling priority is updated is not the highest scheduling priority, executing the second subtask of the second scheduling task, scheduling the second computing resource for the job task in the task queue of the second tenant to execute the job task of the second tenant. By dividing the scheduling task, it is possible to switch the scheduling between the scheduling tasks of different tenants, so that the scheduling tasks of different tenants are smoothly scheduled and the task scheduling opportunities of each tenant are guaranteed.
[0020] In the second aspect, the present application provides a task scheduling device, which includes various modules for implementing the method described in the first aspect and one of its possible implementation methods, such as a configuration module, a creation module, an acquisition module, a scheduling module, an update module, and a splitting module.
[0021] The task scheduling device has the function of implementing the behavior of the method example of any one of the above-mentioned first aspect and its possible implementation methods. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions.
[0022] In a third aspect, the present application provides a computing device comprising a memory and at least one processor connected to the memory, wherein the memory is used to store computer program code, and the computer program code comprises computer instructions. When the computer instructions are executed by at least one processor, the computing device executes the method of the first aspect and any one of its possible implementations.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium storing computer instructions. When the computer instructions are run on a computer, the method of the first aspect and any one of its possible implementations is executed.
[0024] In a fifth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are run on a computer, the method of the first aspect and any one of its possible implementations is executed.
[0025] In a sixth aspect, the present application provides a chip system, comprising: a processor for calling and running a computer program from a memory, so that a storage device equipped with the chip system executes the method of the first aspect and any one of its possible implementations.
[0026] It should be understood that the beneficial effects achieved by the technical solutions of the second to sixth aspects of this application and the corresponding possible implementation methods can be referred to the technical effects of the first aspect and its corresponding possible implementation methods mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] FIG1 is a schematic diagram of an HPC cluster architecture according to an embodiment of the present application;
[0028] FIG2 is a second schematic diagram of an HPC cluster architecture provided in an embodiment of the present application;
[0029] FIG3 is a third schematic diagram of an HPC cluster architecture provided in an embodiment of the present application;
[0030] FIG4 is a hardware diagram of a management node provided in an embodiment of the present application;
[0031] FIG5 is a schematic diagram of a software architecture of a scheduler provided in an embodiment of the present application;
[0032] FIG6 is a schematic diagram of a process for creating a logical cluster in a task scheduling method provided in an embodiment of the present application;
[0033] FIG7 is a flowchart of a task scheduling method according to an embodiment of the present application;
[0034] FIG8 is a schematic diagram of a framework of a task scheduling method according to an embodiment of the present application;
[0035] FIG9 is a second schematic diagram of a framework of a task scheduling method provided in an embodiment of the present application;
[0036] FIG10 is a schematic diagram of a process of splitting a scheduling process in a task scheduling method provided in an embodiment of the present application;
[0037] FIG11 is a second flow chart of a task scheduling method provided in an embodiment of the present application;
[0038] FIG12 is a schematic diagram of state switching of a scheduling task in a task scheduling method provided in an embodiment of the present application;
[0039] FIG13 is a structural diagram of a task scheduling device according to an embodiment of the present application;
[0040] FIG14 is a second structural diagram of a task scheduling device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0042] The terms "first" and "second" and the like in the description and claims of the embodiments of the present application are used to distinguish different objects rather than to describe a specific order of objects.
[0043] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0044] In the description of the embodiments of the present application, unless otherwise specified, “plurality” means two or more.
[0045] First, some technical terms involved in a task scheduling method and device provided in an embodiment of the present application are explained.
[0046] 1. Multi-tenancy
[0047] Multi-tenancy is a software architecture that uses multi-tenancy technology (also known as multi-tenancy) to enable a single software instance (or application instance) to provide services to multiple tenants. This means that multiple tenants use the same application. Software-as-a-service (SaaS) is an example of a multi-tenant architecture.
[0048] It should be noted that in a multi-tenant architecture, it is necessary to ensure that multiple tenants are unaware of each other, that is, to achieve data isolation between multiple tenants under the same set of programs.
[0049] In the embodiments of the present application, the lessee of a software instance may be referred to as a tenant or a customer. In the following embodiments, tenants and customers are the same concept.
[0050] It can be understood that each tenant has one or more users, and the users are the actual users of the tenant's software instance.
[0051] With the continued development of cloud computing, multi-tenancy has become a core architecture within cloud services. Currently, computing clusters (such as HPC clusters) are also rapidly evolving towards cloud services. HPC clusters also require multi-tenancy, meaning that multiple tenants share a single software instance within the HPC cluster. For example, the software instance could be the HPC cluster's scheduler.
[0052] For HPC clusters, the scheduler is a core management software that helps tenants manage the computing resources of computing nodes in the HPC cluster and schedule user-submitted jobs (or job tasks) to be executed on the corresponding computing resources.
[0053] 2. Multi-tenant task scheduling
[0054] Multi-tenant task scheduling, also known as tenant scheduling, refers to the switching strategies and processes between multiple tenants when their scheduling processes share a scheduler. In the embodiments of the present application, the job scheduling process within a tenant is abstracted as a task (or scheduling task), and the scheduling of tasks by the scheduler for multiple tenants is multi-tenant task scheduling.
[0055] Special attention should be paid to the difference between the above-mentioned task scheduling (i.e., tenant scheduling) and job scheduling (i.e., job task scheduling). Job scheduling refers to the matching algorithm and process between multiple jobs submitted by tenants and the computing resources owned by the tenants.
[0056] Taking an HPC cluster as an example, referring to Figure 1, the HPC cluster includes one or more management nodes 101 and multiple computing nodes (or worker nodes) 102. A scheduler is deployed on management node 101, which allocates computing resources to jobs submitted by users under tenants.
[0057] In some implementations, the management node 101 may include a login node and a master control node. The login node is used for users to log in to the scheduler. For example, a scheduler client is deployed on the login node, and user-level and management-level users can log in to the scheduler through the client in the login node. The master control node is used to implement the management function of the scheduler. Other functional modules of the scheduler are deployed on the master control node, such as a user management module, a job management module, a resource management module, and a scheduling policy management module.
[0058] Currently, in an HPC cluster, one HPC cluster is usually configured for one tenant, and each tenant has an independent scheduler. Most current HPC schedulers (schedulers in HPC clusters), such as LSF, Slurm, openPBS, as well as the AI cluster scheduling software Kubernetes and the big data scheduler YARN, do not support multi-tenancy. That is, these schedulers only support one scheduler for one tenant and do not support task scheduling in a multi-tenant architecture.
[0059] Referring to Figure 2, one method for implementing multi-tenant task scheduling is to divide the original HPC cluster into multiple independent subclusters, one for each tenant. Each subcluster is configured with an independent management node, and a scheduler is deployed on the management node to schedule jobs submitted by that tenant. In this method, each tenant is assigned an independent physical cluster (i.e., the aforementioned subcluster). The physical resources of the subclusters corresponding to different tenants are isolated, thus providing good isolation between tenants. However, a separate management node must be configured for each tenant's subcluster. As the number of tenants increases, the number of management nodes also increases, making the cost of implementing multi-tenant task scheduling high.
[0060] Referring to Figure 3, another multi-tenant task scheduling approach involves multiple tenants still sharing a single HPC cluster. The computing resources of the HPC cluster's compute nodes are pooled, logically divided, and managed. In other words, multiple tenants share the physical computing resources of the HPC cluster. In this approach, the computing resources used to process jobs corresponding to the multiple tenants are isolated, and a scheduler deployed on the HPC cluster's management node performs unified task scheduling for the multiple tenants. This approach only isolates the computing resources among the multiple tenants, resulting in a low degree of isolation.
[0061] In response to the above problems, an embodiment of the present application provides a task scheduling method, which is applied to a management node in a computing cluster. In this method, the management node sets a first computing resource for a first tenant in the computing cluster; when it is necessary to execute the job task of the first tenant, a first scheduling task is started for the first tenant that can schedule computing resources for the job task of the first tenant; then the job task of the first tenant is obtained, and the job task of the first tenant is sent to the task queue of the first tenant; and the first scheduling task is executed to schedule the first computing resource for the job task in the task queue of the first tenant to execute the job task of the first tenant. In this method, since a scheduling task is started for each tenant and computing resources are set, independent task scheduling can be performed on the tenants. Therefore, in a scenario where multiple tenants share a scheduler, logical isolation between tenants can be achieved.
[0062] The task scheduling method provided in the embodiments of the present application is executed by a management node. Optionally, the management node can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the management node can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0063] For example, with continued reference to Figure 4 , the management node may include: one or more processors 401, memory 402, and a communication interface 403. The processor 401, memory 402, and communication interface 403 may be connected via a bus 404, or in other ways. Alternatively, the various components of the management node may be implemented in hardware, including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.
[0064] The processor 401 is the control center of the management node, and the processor 401 may be a CPU or other general-purpose processors, such as a microprocessor or any conventional processor. Optionally, the general-purpose processor 401 may include one or more processing cores.
[0065] The controller in processor 401 is the nerve center and command center of the management node. Based on instruction opcodes and timing signals, the controller generates operational control signals to control instruction fetching and execution. Optionally, processor 401 may also include a memory for storing instructions and data.
[0066] The memory 402 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. In the embodiment of the present application, the memory 402 can store information such as computer instructions.
[0067] In one possible implementation, memory 402 may exist independently of processor 401. Memory 402 may be connected to the processor via bus 404 and used to store data, instructions, or program code. When the processor calls and executes the instructions or program code stored in the memory, the relevant steps of the method provided in the embodiment of the present application can be implemented.
[0068] In another possible implementation, the memory 402 may also be integrated with the processor.
[0069] The communication interface 403 may be a transceiver module for communicating with other devices or communication networks, such as Ethernet, RAN, wireless local area networks (WLAN), etc. The communication interface 403 may receive instructions, messages, or data. The transceiver module may be a device such as a transceiver or a transceiver. Alternatively, the communication interface 403 may be a transceiver circuit located within the processor, for implementing signal input and signal output of the processor. The communication interface 403 may be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface, or the communication interface 403 may be a wireless interface.
[0070] Bus 404 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be classified as an address bus, a data bus, a control bus, etc. Buses can also be classified as serial buses and parallel buses. For ease of illustration, FIG4 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0071] Optionally, the management node in the embodiment of the present application may further include an input / output interface 405, which is used to connect to an input device and receive information input by a user through the input device (e.g., a tenant's login information, a job submitted by a tenant). Input devices include, but are not limited to, keyboards, touch screens, microphones, and the like. The input / output interface 405 is also used to connect to an output device and output the processing results of the processor 401, including, but not limited to, a display, a printer, and the like.
[0072] It should be noted that the management node in FIG4 is merely an example of a management node, and the management node may have more or fewer components than those shown in FIG4 , may combine two or more components, or may have different component configurations.
[0073] It should be understood that in the embodiments of the present application, the task scheduling method is performed by a scheduler deployed on a management node in a computing cluster. Referring to Figure 5 , the software architecture of a scheduler deployed on a management node provided in an embodiment of the present application is shown. The scheduler includes: a command line interface (CLI), a graphical user interface (Portal), a REST API, tenant services, cluster services, scheduling services, and unified resource management. The scheduler shown in Figure 5 supports multi-tenancy.
[0074] Among them, the command line interface (CLI) is a tool used by users to log in to the scheduler and is deployed on the login node in the management node.
[0075] Graphical interface (Portal): It is an interface for users to submit jobs and is deployed on the login node in the management node.
[0076] REST API: An application programming interface (API) that follows the Representational state transfer (REST) architectural specification. Applications or devices can connect and communicate with each other based on REST API.
[0077] Tenant Services: These services are used for tenant management, cluster management, and dynamic scaling management. Tenant management includes, but is not limited to, managing tenant permissions; cluster management includes, but is not limited to, maintaining the relationship between tenants and logical clusters; and dynamic scaling management includes, but is not limited to, adding or removing nodes in a compute cluster.
[0078] The cluster service is used to generate logical clusters, encapsulate data and status information during task scheduling, and maintain relationships between tenants and logical clusters. The cluster service can create multiple logical clusters, with each tenant corresponding to one or more logical clusters. Each logical cluster includes functional modules for job management, user management, resource management, and scheduling policy management.
[0079] It should be noted that in the embodiments of the present application, the logical clusters of tenants are isolated from each other, which means that each tenant has its own logical cluster, and the job management, user management, resource management, and scheduling policy management in the logical cluster manage each tenant separately. In other words, each tenant is managed independently. For example, under a tenant, the job numbers submitted by users are continuous, and the job numbers of different tenants are unrelated; the users under each tenant are also managed independently; each tenant has its own independent scheduling policy (the policy or method for processing job tasks), which supports customized scheduling policies for tenants; further, the computing resources used to process each tenant's jobs are also different.
[0080] In the embodiments of the present application, since the logical clusters of tenants are isolated from each other, tenants are unaware of each other. For each tenant, it is equivalent to corresponding to a separate computing cluster. Multiple logical clusters share the resources of a scheduler deployed on a management node. The scheduler schedules the computing resources (set for tenants) for the job tasks of multiple tenants, thus realizing the scheduler's native support for multi-tenancy.
[0081] Scheduling service: It is used to provide some other services during the task scheduling process. For example, in large-scale multi-tenant scenarios, the scheduling service can adjust tenant priorities, split the scheduling process, build the scheduling context (that is, the execution status of the scheduling task), call context switching, and manage the parallel or serial scheduling process.
[0082] Unified Resource Management: This allows users to manage resources, such as creating or deleting resources corresponding to logical clusters. Unified Resource Management shields differences between resource providers, provides a standardized interface for tenant services, and can be integrated with any cloud or offline cluster resource management software.
[0083] Based on the above, the scheduler shown in Figure 5 can support multiple tenants. Multiple tenants share a software instance (such as a scheduler), which greatly reduces the occupancy of management nodes. The more tenants there are, the lower the cost. And since there is no need to deploy a software instance for each tenant, the tenant's operation and maintenance work on the software instance can be reduced.
[0084] With the scheduler software architecture shown in Figure 5, a logical cluster can be created for each tenant in the computing cluster. This logical cluster can also be deleted, modified, or queried during operation. The following, using Figure 6 as an example, describes the process of creating a logical cluster for a tenant, taking one tenant (hereinafter referred to as the first tenant) as an example.
[0085] S601: The tenant service module receives a cluster creation request triggered by a tenant administrator.
[0086] When a new tenant (e.g., the first tenant) is added to the computing cluster, the tenant administrator (operation and maintenance personnel) of the computing cluster sends a cluster creation request to the tenant service. This creation request is used to request the creation of a logical cluster for the first tenant. Optionally, the cluster creation request may include information about the first tenant (e.g., the first tenant's identifier) and the first tenant's resource requirement information. The resource requirement information indicates the first tenant's demand for computing resources for processing jobs. For example, the resource requirement information indicates the number of computing nodes and the configuration requirements for the computing nodes (e.g., the number of cores, memory size, etc.).
[0087] S602: The tenant service creates a correspondence between tenants and logical clusters.
[0088] In one implementation, the tenant service may generate a logical cluster identifier for the first tenant and save the correspondence between the first tenant identifier and the logical cluster identifier, thereby establishing a correspondence between the first tenant and the logical cluster.
[0089] Optionally, after the tenant service receives the cluster creation request, the tenant service also verifies whether the first tenant has the permission to create a cluster. For example, the tenant management function in the tenant service verifies the permission of the first tenant to determine whether the first tenant has the permission to create a logical cluster. If the first tenant has the permission to create a logical cluster, the tenant service establishes a corresponding relationship between the first tenant and the logical cluster.
[0090] S603: The tenant service sends a resource creation request to the unified resource management.
[0091] The resource creation request is used to request the first tenant to set up a first computing resource, which is used to process the first tenant's job. The resource creation request includes the first tenant's resource requirement information.
[0092] S604: After receiving the resource creation request, the unified resource management sets a first computing resource for the first tenant.
[0093] The unified resource management sets computing resources for the first tenant, specifically allocating computing resources (such as processing cores and memory) of computing nodes in the computing cluster to the first tenant.
[0094] S605: The unified resource management sends resource information to the cluster service.
[0095] The resource information is information about the first computing resource set for the first tenant. For example, if the allocated first computing resource is a computing node, the resource information may be the IP address or name of the computing node.
[0096] S606: The tenant service sends a cluster creation request to the cluster service.
[0097] S607: After receiving the cluster creation request sent by the tenant service, the cluster service creates a logical cluster.
[0098] The cluster service creates a logical cluster, including creating functional modules for job management, user management, scheduling policy management, and resource management. The first tenant's logical cluster is logically isolated from other tenants' logical clusters. The logical cluster is used to manage the first tenant's jobs, users, scheduling policies, and computing resources used to execute the first tenant's jobs.
[0099] S608: After receiving the resource information sent by the unified resource management, the cluster service establishes a mapping relationship between the first tenant and the first computing resource.
[0100] After the cluster service creates the logical cluster of the first tenant, it stores the mapping relationship between the first tenant and the first computing resource, which is equivalent to storing the mapping relationship between the logical cluster and the first computing resource.
[0101] Through the above steps S601-S608, the logical cluster for the first tenant is created. It will be appreciated that in this embodiment of the present application, during or after the logical cluster is created, the logical cluster can be configured, i.e., parameters can be configured for the first tenant. Tasks can then be scheduled for multiple tenants based on the created logical cluster.
[0102] In this embodiment, each tenant in the computing cluster has corresponding tenant configuration items (or configuration parameters). Some of these tenant configuration items are shared by all tenants and configured by operations and maintenance personnel; some are unique to each tenant and configured independently by each tenant. See Table 1 for an example of tenant configuration items.
[0103] Table 1
[0104] As can be seen from Table 1, a scheduling policy and tenant priority are configured for the first tenant. This application supports configuring independent scheduling policies for tenants. For example, different scheduling policies can be customized for different tenants based on their different operating scenarios, thereby scheduling tasks for tenants based on their scheduling policies. In addition, this application supports configuring tenant priorities for tenants, so as to facilitate task scheduling based on tenant priorities, ensure fairness in task scheduling between tenants, and ensure that high-priority tenants have more scheduling opportunities.
[0105] Optionally, you can delete, modify, and query the created logical cluster later.
[0106] When the first tenant exits the computing cluster, a cluster deletion request can be sent to the tenant service, and the cluster deletion request can include the identifier of the first tenant; thereby, the tenant service deletes the corresponding relationship between the first tenant and the logical cluster, and sends cluster deletion requests to the unified resource management and cluster service respectively, the unified resource management deletes the first computing resource corresponding to the logical cluster, the cluster service deletes the mapping relationship between the first tenant and the first computing resource, and deletes the logical cluster and the configuration items of the first tenant.
[0107] When you need to query relevant information about the first tenant or logical cluster, you can send a query request to the tenant service, such as querying the configuration items of the first tenant or querying the information of the logical cluster corresponding to the first tenant, so that the tenant service or cluster service can return the configuration items or cluster information.
[0108] When it is necessary to modify the relevant information of the first tenant or the logical cluster, a modification request can be sent to the tenant service, such as requesting to modify the configuration items of the first tenant or modify the information of the logical cluster (such as modifying the configuration of the first computing resource), so that the relevant modules can be modified according to the modification requirements.
[0109] Combined with the architectural introduction of the scheduler provided in the embodiment of the present application in Figure 4, the computing cluster can support multiple tenants, multiple tenants share one scheduler, and the scheduling process of each tenant is defined as a scheduling task, and computing resources and scheduling tasks for processing job tasks are set for each tenant. After the scheduler obtains the job tasks of multiple tenants, it executes the tenant's scheduling tasks and schedules the tenant's jobs to the corresponding computing resources as the tenant's job tasks.
[0110] The task scheduling method provided in the embodiment of the present application can be applied to the scenario of a computing cluster with a multi-tenant architecture based on cloud services, and can also be applied to the scenario of a computing cluster with a multi-tenant architecture based on physical resources, without specific limitation.
[0111] In conjunction with the description of the computing cluster, management node, and scheduler in the above embodiments, the process of providing a task scheduling method in the embodiment of the present application is described below. This task scheduling method is applied to the management node in the computing cluster, that is, it is executed by the management node in the computing cluster (specifically, it is executed by the scheduler deployed on the management node). The following description is based on the scheduler as the execution subject. As shown in Figure 7, the method includes S701-S704.
[0112] S701. Set a first computing resource for a first tenant.
[0113] In an embodiment of the present application, a computing cluster may include multiple tenants, and a scheduler may configure (i.e., allocate) corresponding computing resources for each tenant of the computing cluster. For example, a computing cluster may include a first tenant and a second tenant, a first computing resource may be configured for the first tenant, and a second computing resource may be configured for the second tenant, and the first computing resource and the second computing resource may share the physical computing resources of the computing cluster.
[0114] Taking the first tenant as an example, the specific process of the scheduler setting the first computing resource for the first tenant can refer to the description of S603 to S605 in the process of creating the logical cluster of the first tenant described in the above embodiment, which will not be repeated here.
[0115] S702: When the job task of the first tenant needs to be executed, start a first scheduling task for the first tenant, where the first scheduling task is used to schedule computing resources for the job task of the first tenant.
[0116] In the embodiment of the present application, the first scheduling task refers to the scheduling process of the job task of the first tenant, and each tenant corresponds to a scheduling task. Starting the first scheduling task for the first tenant can be understood as: creating the first scheduling task for the first tenant, or generating the first scheduling task for the first tenant.
[0117] S703: Obtain the job task of the first tenant, and send the job task to the task queue of the first tenant.
[0118] The tenant's task queue includes job tasks submitted by multiple users under the tenant. Optionally, the job tasks in the first tenant's task queue can be sorted according to the time the job tasks were received, or can be sorted according to other methods, which are not limited in this embodiment of the application.
[0119] S704: Execute a first scheduling task to schedule a first computing resource for the job task in the task queue to execute the job task.
[0120] The scheduler executing the scheduling task specifically includes executing the scheduling task of the tenant based on the scheduling resources allocated for the scheduling task (resources in the scheduler used to execute the scheduling task).
[0121] It is understandable that the scheduling resource can be a process, thread, or coroutine, and the specific selection is based on actual needs and is not limited in this application. The relationship between processes, threads, and coroutines is that a process can include multiple threads, and a thread can include multiple coroutines. For ease of description, in the embodiments of this application, the scheduling resource of the scheduler is explained by taking the scheduling thread as an example.
[0122] In an embodiment of the present application, the scheduler executing the scheduling task of the tenant is actually the scheduler calling the thread to execute the scheduling task. The scheduler can create multiple scheduling threads to execute scheduling tasks. When the resources of the scheduler are sufficient (i.e., the resources used by the scheduler in the management node), the scheduler can allocate independent scheduling resources for each of the multiple scheduling tasks. For example, one scheduling thread is used to execute one scheduling task (one tenant corresponds to an independent scheduling thread), and the scheduler executes the scheduling tasks of multiple tenants in parallel. When the scheduling resources of the scheduler are insufficient, multiple scheduling tasks share the scheduling resources of the scheduler. For example, one scheduling thread is used to execute multiple scheduling tasks (multiple tenants share one scheduling thread), and the scheduler executes the scheduling tasks of multiple tenants sequentially.
[0123] In some implementations, the scheduler's resources can be evenly distributed to multiple scheduling tasks, or different proportions of resources can be allocated to multiple scheduling tasks based on other factors (such as the tenant weights and / or service quality configured for the tenants as mentioned above). This is not limited to the embodiments of the present application.
[0124] To sum up, in the task scheduling method provided in the embodiment of the present application, since computing resources can be set for each tenant of the computing cluster and scheduling tasks can be started for the tenant, the tenant's job tasks can be scheduled to the computing resources set for it by executing the tenant's scheduling tasks to execute the tenant's job tasks. In the scenario where multiple tenants share a scheduler, tasks can be independently scheduled for tenants, logical isolation between tenants can be achieved, and the performance of task scheduling for multiple tenants in the computing cluster can be improved.
[0125] According to the description in the above embodiment, the scheduler can set a scheduling policy for each tenant. Based on this, in one implementation, in S704 above, scheduling the first computing resource for the job task in the task queue of the first tenant includes: scheduling the first computing resource for the job task in the task queue according to the scheduling policy.
[0126] For example, the scheduling policy may be to schedule computing resources for a job task according to the job task's reception time or job task priority, etc. For example, a first computing resource may be scheduled according to the job task priority. If the first job task has a higher priority than a second job task, the first computing resource will be scheduled preferentially for the first job task.
[0127] In an embodiment of the present application, a scheduler is deployed on the management node. For ease of description, in the following embodiments, the resources of the management node are the resources of the scheduler, and the resources of the management node are used to schedule tasks for tenants.
[0128] In computing clusters, the management node used to deploy the scheduler typically has limited resources. In different scenarios, when scheduling tasks for multiple tenants, it's important to consider resource allocation for the management node. Management node resources include processing resources (such as processors / cores) and storage resources (such as memory).
[0129] The task scheduling method provided in the embodiments of the present application can be applied to scenarios with different tenant scales. The following is an illustration of three different tenant scales. The three scenarios with different tenant scales are Scenario 1 (small-scale multi-tenant scenario), Scenario 2 (medium-scale multi-tenant scenario), and Scenario 3 (large-scale multi-tenant scenario).
[0130] Scenario 1: Small-scale multi-tenant scenario
[0131] For small-scale multi-tenant scenarios, the number of tenants is relatively small and the resources of the management node are relatively sufficient. In this case, when the scheduler executes the scheduling task, it can create a scheduling thread for each scheduling task, that is, allocate independent resources in the management node as scheduling resources for each tenant, and allocate scheduling resources that can meet the needs for each scheduling task to execute the tenant's scheduling task.
[0132] Referring to Figure 8 , tenants 1 through n each correspond to a scheduling thread. Multiple scheduling tasks are scheduled by having multiple scheduling threads run in parallel. This means the scheduler can call multiple scheduling threads to execute multiple tenants' scheduling tasks in parallel, parallelizing the scheduling process between tenants and improving scheduling efficiency. For a tenant, a scheduling thread executes the tenant's scheduling tasks, dispatching the jobs in the tenant's task queue to the computing resources allocated to the tenant to execute the jobs in the tenant's task queue. This serializes the scheduling process within the tenant, using time-sharing scheduling (serial scheduling).
[0133] In one implementation, the range of processing and storage resources that a scheduling thread occupies on a management node can be configured based on operating system kernel mechanisms. For example, the resources of a management node can be divided based on the Cgroup mechanism in the Linux kernel to create different scheduling threads. It will be appreciated that Cgroups can control the allocation of processing and storage resources to processes or threads.
[0134] For example, referring to some configuration items of Cgroup shown in Table 2 below, the allocation of processing resources and storage resources can be achieved by using key configurations such as cpuset.cpus and memory.limit_in_bytes in Cgroup.
[0135] Table 2
[0136] Scenario 2: Medium-scale multi-tenant scenario
[0137] For medium-sized multi-tenant scenarios, the resources of the management node may be insufficient. In this case, scheduling resources can be allocated to each scheduling task based on the tenant priority. For example, a scheduling thread is created for each scheduling task based on the tenant priority to execute the tenant's scheduling task.
[0138] Referring to Figure 9 , in an embodiment of the present application, a tenant priority is configured for each tenant. The tenant priority may include a tenant weight and / or quality of service. The relationship between tenant priority and tenant weight is that the greater the tenant weight, the higher the tenant priority. For example, for tenants with urgent tasks, the tenant weight is larger and the tenant priority is higher. The relationship between tenant priority and tenant quality of service is that the higher the tenant quality of service, the higher the tenant priority.
[0139] Optionally, there is a functional relationship between the tenant priority, the tenant weight, and the service quality, and the tenant priority can be calculated according to a preset functional relationship.
[0140] In one implementation, when allocating management node resources to tenants, a resource allocation ratio may be determined according to tenant priorities of multiple tenants, and then corresponding ratios of resources may be allocated to different scheduling tasks according to the resource allocation ratio.
[0141] For example, if the tenant priority includes the tenant weight, the computing cluster includes 3 tenants, and the tenant weights of the three tenants are 100, 20, and 10 respectively, then the resource allocation ratio is determined to be 10:2:1. Then, the processing resources and storage resources of the management node are allocated according to 10:2:1, and three scheduling threads are created to execute the scheduling tasks of the 3 tenants.
[0142] For example, referring to some of the configuration items of Cgroup shown in Table 3 below, the allocation of processing resources and storage resources can be achieved by using key configurations such as cpuset.cpus, cpu.cfs_quota_us & cpu.cfs_period_us, and cpu.shares in Cgroup.
[0143] Table 3
[0144] In an embodiment of the present application, allocating management node resources to multiple scheduling tasks according to tenant priority can ensure that high-priority customers are allocated more scheduling resources for executing scheduling tasks, thereby allowing high-priority tenants to obtain more scheduling opportunities.
[0145] Scenario 3: Large-scale multi-tenant scenario
[0146] For large-scale multi-tenant scenarios, the management node has insufficient resources (such as resource shortage). In this case, multiple scheduling tasks share the resources of the management node. For example, among the multiple scheduling threads created, one scheduling thread is used to execute multiple scheduling tasks.
[0147] It can be understood that there is a corresponding relationship (or mapping relationship) between the tenant (or the tenant's scheduling task) and the scheduling thread. In the embodiment of the present application, the tenant name can be used as the key value to calculate the hash value of all scheduling tasks. Different hash values correspond to different queues (each scheduling thread corresponds to a queue, that is, different hash values correspond to different scheduling threads), and then the tenant's scheduling task is mapped to the corresponding queue. The queue is a queue of scheduling tasks for multiple tenants. Multiple scheduling tasks are arranged in the queue according to the scheduling priority of the scheduling task, so the queue can be called a priority queue. The above-mentioned tenant priority can be used to determine the scheduling priority of the tenant's scheduling task, which will be described in detail in the following embodiments.
[0148] In one implementation, the amount of management node resources occupied by multiple scheduling threads can be the same or different. For example, the management node resources are allocated based on information such as the number of tenants corresponding to a scheduling thread and / or the tenant priorities of at least two tenants corresponding to the scheduling thread.
[0149] In an embodiment of the present application, in a large-scale multi-tenant scenario, a priority queue includes a first scheduled task for a first tenant and a second scheduled task for a second tenant. During the process of scheduling tasks for multiple tenants, the scheduler calls a scheduling thread to execute the scheduled tasks for the multiple tenants, including executing the scheduled tasks according to the scheduling priorities of the multiple scheduled tasks.
[0150] Taking the first tenant as an example, executing the first scheduling task of the first tenant specifically includes: executing the first scheduling task according to the scheduling priority of the first scheduling task. Specifically, determine whether the first scheduling task in the priority queue has the highest scheduling priority; if the first scheduling task has the highest scheduling priority, execute the first scheduling task; if the first scheduling task does not have the highest scheduling priority, execute the second scheduling task. In other words, the scheduling tasks in the priority queue are executed in descending order of scheduling priority. When the scheduling priority of the first scheduling task is higher than the scheduling priority of the second scheduling task, the first scheduling task is executed first; when the scheduling priority of the first scheduling task is lower than the scheduling priority of the second scheduling task, the second scheduling task is executed first.
[0151] In one implementation, for each scheduled task in a priority queue, the scheduled task can be segmented (i.e., the tenant's job scheduling process can be segmented), and each scheduled task can be divided into multiple slices. Each slice can be called a subtask. In this way, a scheduled task includes multiple subtasks. For example, for the first and second scheduled tasks in the priority queue, the first scheduled task and the second scheduled task of the second tenant can be divided into multiple subtasks respectively.
[0152] Optionally, the above-mentioned process of segmenting the tenant's scheduling tasks specifically includes: using the interruptible execution points in each tenant's scheduling process (each tenant's scheduling process is each tenant's scheduling task) as safe points, and then dividing the tenant's job scheduling process into multiple slices based on the safe points. The size of the slice is the execution time of the slice, and the size of a slice can range from 0.1ms (milliseconds) to 5ms. For example, each successful scheduling of a job takes about 1ms, so the scheduling process of a job can be treated as a subtask.
[0153] Optionally, the size of the slice is adjusted according to the tenant priority, such as the tenant's quality of service (QoS). The higher the tenant's quality of service, the smaller the slice, ie, the shorter the execution time of the slice.
[0154] In some embodiments, a tenant's scheduling process is segmented according to the characteristics of the scheduling process. For example, the scheduling process includes multiple different stages, and it is determined whether each stage can be segmented; for another example, the scheduling process includes multiple jobs, and the point where each job scheduling ends is used as the segmentation point.
[0155] For example, referring to Figure 10, assume that a tenant's scheduling process includes five phases, some of which may include multiple sub-phases, which may include serial, parallel, and loop processes. The dashed line in Figure 10 represents a tenant's scheduling process. By segmenting the scheduling process, a tenant's scheduling tasks are divided into multiple sub-tasks.
[0156] In some embodiments, when splitting the scheduling process at the source code level, split markers are added at the execution points where splitting is required. When the scheduling thread executes the scheduled task, the execution code from the current safepoint to the next safepoint is automatically generated and cached. For example, the @safepoint annotation is added to the source code level, and scheduling subtasks are generated through bytecode enhancement technology.
[0157] For example, a code example for adding a segmentation mark is as follows:
[0158] In one implementation, when a scheduling task is divided into multiple subtasks, the above-mentioned execution of the first scheduling task according to the scheduling priority includes: if the first scheduling task has the highest scheduling priority, executing the first subtask of the first scheduling task, scheduling the first computing resource for the job task in the task queue of the first tenant, so as to execute the job task of the first tenant; then updating the scheduling priority of the first scheduling task; if the first scheduling task after the scheduling priority update is not the highest scheduling priority, executing the second subtask of the second scheduling task, scheduling the second computing resource for the job task in the task queue of the second tenant, so as to execute the job task of the second tenant. By dividing the scheduling task, the scheduling can be switched between the scheduling tasks of different tenants, so that the scheduling tasks of different tenants are smoothly scheduled and the task scheduling opportunities of each tenant are guaranteed.
[0159] For each scheduling thread created by the scheduler, the process of the scheduler calling the scheduling thread to perform task scheduling is similar. The following embodiment takes one scheduling thread as an example to describe in detail the task scheduling process for tenants.
[0160] In one implementation, the scheduler invokes a scheduling thread to periodically execute multiple scheduled tasks in a priority queue. During a scheduling cycle, the scheduler iterates through multiple scheduled tasks in a priority queue. For each scheduled task, during the scheduling cycle, the scheduled task includes several status parameters as shown in Table 4 below.
[0161] Table 4
[0162] It should be understood that the scheduling priority of a scheduling task is related to the execution time (n_runtime) of the above-mentioned scheduling task (n_runtime is related to the tenant priority). Within a scheduling cycle, the scheduling priority of a scheduling task decreases as the execution time of the scheduling task increases. The shorter the execution time of the scheduling task, the higher the scheduling priority of the scheduling task.
[0163] Optionally, in a priority queue, scheduled tasks are arranged in order of n_runtime from small to large. In some cases, if n_runtime of multiple scheduled tasks is equal, the scheduled tasks of the tenants can be sorted by tenant name or randomly sorted, which is not limited in the embodiments of the present application.
[0164] Within a scheduling cycle, each scheduled task is assigned an executable duration based on the tenant's quality of service. The executable duration of a scheduled task is T*QoS, where T is the duration of the scheduling cycle and QoS is the tenant's quality of service. The sum of the executed duration n_runtime and the remaining executable duration r_runtime of the scheduled task is T*QoS.
[0165] In one implementation, the embodiment of the present application can also set a task state for each scheduled task. For example, the state of the scheduled task can include ready state, running state, sleeping state and blocked state, wherein the ready state indicates that the scheduled task is in a waiting state, the running state indicates that the scheduled task is in an executing state, the sleeping state indicates that the scheduled task is in a dormant state, indicating that the tenant currently has no job, and the blocked state indicates that the scheduled task is in a locked state, indicating that the executable time of the scheduled task in the current scheduling cycle has been used up, that is, the scheduled task cannot be executed.
[0166] With reference to Table 4, it can be understood that the various state parameters of the scheduling task change dynamically within a scheduling cycle.
[0167] In one implementation, during initialization (ie, when starting a scheduling task), various state parameters of the scheduling task may be set to: curr_num=0, n_runtime=0, r_runtime=T*QoS, state=ready.
[0168] As shown in FIG11 , when multiple scheduling tasks share one scheduling thread, the process of the scheduler scheduling one scheduling thread to execute scheduling tasks of multiple tenants in one scheduling cycle includes S1101 - S1104 .
[0169] S1101. Add scheduling tasks of multiple tenants to a priority queue in order of scheduling priority.
[0170] It should be understood that during task scheduling, tasks in the priority queue are tasks waiting to be scheduled. That is, tasks in the priority queue are in the ready state. When a task's state changes to sleeping or blocked, it is removed from the priority queue. This gives other tenants scheduling opportunities, for example, increasing the scheduling opportunities for tenants with a small number of tasks.
[0171] 12 , taking the first scheduling task of the first tenant as an example, the switching process between various states of the first scheduling task is described.
[0172] In a scheduling cycle, when the first scheduling task is moved into the priority queue, the state of the first scheduling task is ready. When the scheduling priority of the first scheduling task is the highest, the scheduling thread executes the first scheduling task, and the state of the first scheduling task switches from ready to running.
[0173] When a subtask of the first scheduled task is completed, the first scheduled task still has remaining executable time, and the scheduling priority of the first scheduled task is not the highest, the state of the first scheduled task is switched from the running state to the ready state.
[0174] When a subtask of the first scheduled task is completed and there is no remaining executable time for the first scheduled task, the state of the first scheduled task is switched from the running state to the blocked state.
[0175] When a subtask of the first scheduled task is completed, if the first tenant has no jobs to execute, the state of the first scheduled task is switched from running to sleeping. The tenant can be considered an idle tenant.
[0176] When the first tenant submits a new job, the first tenant changes from having no job to being able to execute to having a job to be able to execute, and the state of the first scheduled task is switched from the sleeping state to the ready state.
[0177] When the timer (used for timing scheduling period) is reset and triggered again, if the first scheduled task was switched to the blocked state in the previous cycle, the state of the first scheduled task is switched from the blocked state to the ready state when the timer is triggered again.
[0178] In the embodiment of the present application, starting a timer and adding multiple scheduled tasks to a priority queue in the order of scheduling priority includes the following situations:
[0179] 1. When the scheduler calls the scheduling thread for the first time, it sets the status of all scheduled tasks that the scheduling thread needs to execute to the ready state and arranges them in descending order of scheduling priority. In addition, since the n_runtime of all scheduled tasks is initialized to 0 during initialization, they can be arranged by tenant name or randomly.
[0180] 2. Add the scheduled task in the blocked state to the priority queue and switch the state of the scheduled task to the ready state.
[0181] 3. For tenants without jobs (the status of their scheduled tasks is sleeping), when the tenant submits a new job, the tenant's scheduled tasks are added to the priority queue and the status of the scheduled tasks is switched to ready.
[0182] S1102: Obtain the scheduling task with the highest scheduling priority in the priority queue.
[0183] After obtaining the scheduling task with the highest scheduling priority in the priority queue, the status of the scheduling task is switched from ready to running.
[0184] In one implementation, if there are no scheduled tasks in the priority queue, scheduled tasks can be obtained from other priority queues and executed, thereby improving the resource utilization of the management node. The status of the scheduled tasks obtained from the priority queues of other scheduling threads is updated to the running state.
[0185] S1103: Execute the subtask of the scheduling task with the highest scheduling priority.
[0186] For a scheduled task, the multiple subtasks contained in the scheduled task are arranged in the order of the subtask numbers, for example, the subtasks are executed in the order from small to large according to the subtask numbers. In the embodiment of the present application, the scheduler records the execution time of the subtask (denoted as delta) during the execution of a subtask. After a subtask is completed, the number curr_num of the current subtask of the scheduled task, the execution time n_runtime of the scheduled task, and the remaining executable time r_runtime of the scheduled task are updated.
[0187] Among them, curr_num=curr_num+1;
[0188] n_runtime=n_runtime+delta*1024 / weight;
[0189] r_runtime=max(0, r_runtime-delta).
[0190] Weight indicates the tenant weight of the tenant.
[0191] S1104: Dynamically update the scheduling priority of the scheduling task.
[0192] In one implementation, taking the first scheduled task as an example, the scheduling priority of the first scheduled task can be dynamically updated based on the execution time (n_runtime) of the first scheduled task. Specifically, the execution time (n_runtime) of the first scheduled task is updated, and then the scheduling priority of the first scheduled task is updated based on the execution time of each scheduled task in the priority queue.
[0193] The above-mentioned dynamic update of the scheduling priority of a scheduled task actually updates the priority queue. If the order of the execution time of each scheduled task in the priority queue changes, the scheduling priority of the scheduled task will also change. For example, the first scheduled task has the highest scheduling priority in the priority queue. After executing a subtask of the first scheduled task, the execution time of the first scheduled task is no longer the scheduled task with the smallest execution time in the priority queue. In this case, the scheduling priority of the first scheduled task is updated, and the scheduling priority of the first scheduled task is no longer the highest scheduling priority.
[0194] During a scheduling cycle, after executing a scheduled task's subtasks, the status of the scheduled task also needs to be updated. It should be noted that after the status of the scheduled task is updated, the scheduled tasks in the blocked and sleeping states are removed from the priority queue, and the remaining scheduled tasks in the priority queue are reordered according to the updated scheduling priority.
[0195] Taking the first scheduled task as an example, if the remaining executable time r_runtime of the first scheduled task is equal to 0, the state of the first scheduled task is updated to the blocked state, the executed time n_runtime = T*QoS, and the first scheduled task is removed from the priority queue.
[0196] If the r_runtime of the first scheduled task is not equal to 0 and the first scheduled task has no job, the state of the first scheduled task is updated to the sleeping state, and the execution time of the first scheduled task is updated to n_runtime = n_runtime-min(n_runtime1,...), where min(n_runtime1,...) represents the minimum execution time of all scheduled tasks that the scheduling thread needs to execute. It should be noted that, subsequently, when the first scheduled task has another job submitted, the state of the first scheduled task is updated from the sleeping state to the ready state, and the n_runtime of the first scheduled task is updated to n_runtime+min(n_runtime1,...).
[0197] It is understandable that when the first scheduling task has no job, the first scheduling task will no longer be executed. In the current scheduling cycle, the execution time of the first scheduling task will remain unchanged. As other scheduling tasks are executed, the execution time of other scheduling tasks will increase. When the first scheduling task has a job submitted, the scheduling priority of the first scheduling task is the highest compared with other scheduling tasks. The scheduling thread will give priority to executing the first scheduling task and will not execute other scheduling tasks for a long time. In order to ensure the fairness of task scheduling for multiple tenants, the execution time of the first scheduling task is updated to appropriately increase the execution time of the first scheduling task to avoid other scheduling tasks not being scheduled for a long time. To a certain extent, other scheduling tasks can obtain more scheduling opportunities.
[0198] If the r_runtime of the first scheduled task is not equal to 0 and the first scheduled task has a job, the status of the scheduled task is determined according to the scheduling priority of the first scheduled task. According to the above embodiment, after executing a subtask, the n_runtime of the first scheduled task is n_runtime + delta * 1024 / weight, and the status of the first scheduled task is updated according to n_runtime. If the n_runtime of the first scheduled task is the smallest among the n_runtimes of multiple scheduled tasks, the status of the first scheduled task remains in the running state. If the n_runtime of the first scheduled task is not the smallest among the n_runtimes of multiple scheduled tasks, the status of the first scheduled task is updated to the ready state.
[0199] It should be noted that when a second tenant is newly added to the computing cluster, a second scheduling task is started for the second tenant, and the second scheduling task is mapped to a scheduling thread of the scheduler. In this case, the n_runtime of the second scheduling task is set to: n_runtime = min(n_runtime1,...) + 0.5ms, the state of the priority queue is set to the ready state, and the second scheduling task is added to the priority queue.
[0200] In the embodiment of the present application, within one scheduling cycle, S1102 - S1104 are repeatedly executed until the timer expires.
[0201] It is understandable that the above method is executed by a task processing device (the task processing device can be a scheduler deployed on the management node), and the task processing device includes a hardware structure and / or software module corresponding to the execution of each function in order to realize the above functions. It should be easy for those skilled in the art to appreciate that, in combination with the method steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0202] The embodiment of the present application can divide the functional modules of the above-mentioned task processing device according to the above-mentioned method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0203] 13 shows a possible structural diagram of the task processing device involved in the above embodiment, in which each functional module is divided into corresponding functional modules. The task processing device includes a configuration module 1301 , a creation module 1302 , an acquisition module 1303 and a scheduling module 1304 .
[0204] Among them, the configuration module 1301 is used to execute S701 in the above method embodiment; the creation module 1302 is used to execute S702 in the above method embodiment; the acquisition module 1303 is used to execute S703 and S1102 in the above method embodiment; and the scheduling module 1304 is used to execute S704 and S1103 in the above method embodiment.
[0205] Optionally, the task processing device provided in the embodiment of the present application further includes an updating module 1305 and a splitting module 1306. The updating module 1305 is used to execute S1104 in the above method embodiment; the splitting module 1306 is used to split the scheduled task into multiple subtasks.
[0206] The various modules of the above-mentioned task processing device can also be used to perform other actions in the above-mentioned method embodiment. All relevant contents of each step involved in the above-mentioned method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0207] In the case of adopting an integrated unit, Figure 14 shows another possible structural diagram of the task processing device involved in the above embodiment. As shown in Figure 14, the task processing device provided in the embodiment of the present application may include: a processing module 1401 and a communication module 1402. The processing module 1401 can be used to control and manage the actions of the task processing device. For example, the processing module 1401 can be used to support the configuration module 1301, creation module 1302, acquisition module 1303, scheduling module 1304, update module 1305 and segmentation module 1306 in Figure 14 above to perform corresponding steps, and / or other processes for the technology described in this article. The communication module 1402 can be used to support communication between the task processing device and other network entities. As shown in Figure 14, the task processing device may also include a storage module 1403 for storing computer instructions and data.
[0208] Among them, the processing module 1401 can be a processor or a controller (for example, the processing module 1401 can be the processor 401 in Figure 4). The above-mentioned processing module 1401 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication module 1402 can be a communication interface (for example, the communication module 1402 can be the communication interface 403 in Figure 4). The storage module 1403 can be a memory (for example, the storage module 1403 can be the memory 402 in Figure 4). When the processing module 1401 is a processor, the communication module 1402 is a communication interface, and the storage module 1403 is a memory, the processor, transceiver, and memory can be connected via a bus.
[0209] For more details on how the modules included in the task processing device implement the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here. The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments.
[0210] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions in accordance with the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a magnetic disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state drive (SSD)).
[0211] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0212] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0213] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0214] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0215] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0216] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A task scheduling method, characterized in that, Applied to a management node in a computing cluster, the method includes: Set the first computing resource for the first tenant; When it is necessary to execute the job task of the first tenant, start the first scheduling task for the first tenant, where the first scheduling task is used to schedule computing resources for the job task of the first tenant; Obtain the job task of the first tenant and send the job task to the task queue of the first tenant; Execute the first scheduling task to schedule the first computing resource for the job task in the task queue to execute the job task.
2. The method according to claim 1, wherein The method further includes: Set a scheduling policy for the first tenant; Scheduling the first computing resource for the job task in the task queue includes: Scheduling the first computing resource for the job task in the task queue according to the scheduling policy.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Set a tenant priority for the first tenant; the tenant priority is used to determine the scheduling priority of the first scheduling task; Executing the first scheduling task includes: Executing the first scheduling task according to the scheduling priority.
4. The method according to claim 3, wherein The executing the first scheduling task according to the scheduling priority includes: Determine whether the first scheduling task in the priority queue is the highest scheduling priority, and the priority queue also includes a second scheduling task of a second tenant; If the first scheduling task is the highest scheduling priority, execute the first scheduling task; if the first scheduling task is not the highest scheduling priority, execute the second scheduling task.
5. The method according to claim 3, wherein The method further includes: Dynamically update the scheduling priority of the first scheduling task according to the executed duration of the first scheduling task.
6. The method according to claim 3, wherein The computing cluster further includes a second tenant, and the second computing resource of the second tenant shares the physical computing resources of the computing cluster with the first computing resource of the first tenant. The method further includes: Divide the first scheduling task and the second scheduling task of the second tenant into multiple subtasks respectively; The executing the first scheduling task according to the scheduling priority includes: If the first scheduling task is the highest scheduling priority, execute the first subtask of the first scheduling task, schedule the first computing resource for the job task in the task queue of the first tenant to execute the job task of the first tenant; Update the scheduling priority of the first scheduling task; If the first scheduling task after the scheduling priority is updated is not the highest scheduling priority, execute the second subtask of the second scheduling task, schedule the second computing resource for the job task in the task queue of the second tenant to execute the job task of the second tenant.
7. A task scheduling device, characterized in that, Applied to a management node in a computing cluster, the task scheduling device includes: a configuration module, a creation module, an acquisition module, and a scheduling module; The configuration module is used to set the first computing resource for the first tenant; The creation module is used to start the first scheduling task for the first tenant when it is necessary to execute the job task of the first tenant, where the first scheduling task is used to schedule computing resources for the job task of the first tenant; The obtaining module is configured to obtain the job tasks of the first tenant and send the job tasks to the task queue of the first tenant; The scheduling module is configured to execute the first scheduling task, and schedule the first computing resources for the job tasks in the task queue to execute the job tasks.
8. The task scheduling device according to claim 7, wherein The configuration module is further configured to set a scheduling policy for the first tenant; The scheduling module is specifically configured to schedule the first computing resources for the job tasks in the task queue according to the scheduling policy.
9. The task scheduling device according to claim 7 or 8, wherein The configuration module is further configured to set a tenant priority for the first tenant; the tenant priority is used to determine the scheduling priority of the first scheduling task; The scheduling module is specifically configured to execute the first scheduling task according to the scheduling priority.
10. The task scheduling device according to claim 9, wherein The scheduling module is specifically configured to determine whether the first scheduling task in the priority queue is the highest scheduling priority, and the priority queue further includes a second scheduling task of a second tenant; if the first scheduling task is the highest scheduling priority, execute the first scheduling task; if the first scheduling task is not the highest scheduling priority, execute the second scheduling task.
11. The task scheduling device according to claim 9, characterized in that, It further includes an updating module; The updating module is configured to dynamically update the scheduling priority of the first scheduling task according to the executed duration of the first scheduling task.
12. The task scheduling device according to claim 9, wherein The computing cluster further includes a second tenant, and the second computing resources of the second tenant and the first computing resources of the first tenant share the physical computing resources of the computing cluster; the task scheduling device further includes a splitting module and an updating module; The splitting module is configured to divide the first scheduling task and the second scheduling task of the second tenant into multiple subtasks respectively; The scheduling module is specifically configured to, if the first scheduling task is the highest scheduling priority, execute the first subtask of the first scheduling task, and schedule the first computing resources for the job tasks in the task queue of the first tenant to execute the job tasks of the first tenant; The updating module is configured to update the scheduling priority of the first scheduling task; The scheduling module is further configured to, if the first scheduling task after the scheduling priority is updated is not the highest scheduling priority, execute the second subtask of the second scheduling task, and schedule the second computing resources for the job tasks in the task queue of the second tenant to execute the job tasks of the second tenant.
13. A computer-readable storage medium, characterized in that, Stores computer instructions, which, when running on a computer, execute the method according to any one of claims 1 to 6.
14. A computer program product, characterized in that, Contains instructions that, when run on a computing device, cause the computing device to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Task scheduling method and device
CN120353545A
GPU cluster scheduling method and device
CN116431329A
Cluster resource scheduling method and device, equipment and medium
CN116643890A
Multi-tenant resource scheduling method and device and storage medium
CN117112199A
Throttling queue for a request scheduling and processing system
US20180375784A1