Task scheduling method and device
By setting up computing resources and scheduling tasks for each tenant in the computing cluster, the problem of poor tenant isolation in the multi-tenant shared scheduler is solved, and the logical isolation and task scheduling performance are improved, ensuring fairness among tenants and the scheduling opportunities of high-priority tenants are guaranteed.
Patent Information
- Application Number
- CN202410082980.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2025-07-22
AI Technical Summary
In the scenario where multi-tenant shared scheduler is shared, the isolation between tenants in the prior art is poor, resulting in mutual influence of job scheduling processes and affecting the task scheduling performance of the computing cluster.
Set up computing resources for each tenant in the computing cluster and start scheduling tasks for it. By executing tenants' scheduling tasks, scheduling job tasks to the set computing resources, logical isolation between tenants, and support the configuration of customized scheduling policies and priorities for tenants to ensure the scheduling opportunities of high-priority tenants.
It realizes logical isolation between tenants in the multi-tenant shared scheduler scenario, improves the task scheduling performance of the computing cluster, and ensures the fairness of task scheduling between tenants and the scheduling opportunities of high-priority tenants.
Smart Images

Figure CN120353545A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a task scheduling method and apparatus. Background Art
[0002] With the in-depth development of cloud computing, cloud services have become the development direction. Multi-tenancy is the core architecture in cloud services. The multi-tenant architecture refers to an architecture where multiple tenants share a software instance.
[0003] In a high performance computing (HPC) cluster, one software instance is a scheduler. As the core management software of the HPC cluster, the scheduler mainly helps tenants manage the resources within the HPC cluster and schedules jobs to the corresponding resources for execution. For example, the scheduler receives jobs from tenants, distributes the jobs to the computing nodes of the cluster through certain scheduling policies so that the computing nodes process the users' jobs, and monitors and manages the life cycle of the jobs.
[0004] Currently, in the multi-tenant architecture, a single scheduler schedules jobs for multiple tenants. In this multi-tenant shared scheduler solution, the isolation between multiple tenants is relatively poor, and the job scheduling processes of different tenants will affect each other. Summary of the Invention
[0005] This application provides a task scheduling method and apparatus, which can independently perform task scheduling for tenants in the scenario of multi-tenant shared scheduler, and achieve logical isolation between tenants.
[0006] This application adopts the following technical solutions:
[0007] In a first aspect, this application provides a task scheduling method, which is applied to a management node in a computing cluster (specifically, a scheduler deployed on the management node). The method includes: setting first computing resources for a first tenant in the computing cluster; when it is necessary to execute the job task of the first tenant, starting a first scheduling task for the first tenant, where the first scheduling task is used to schedule computing resources for the job task of the first tenant; then obtaining the job task of the first tenant and sending the job task of the first tenant to the task queue of the first tenant; and executing the first scheduling task to schedule the first computing resources for the job tasks in the task queue of the first tenant to execute the job task of the first tenant.
[0008] In this application, since computing resources can be set for tenants of a computing cluster and scheduling tasks can be started for the tenants, the job tasks of the tenants are scheduled to the computing resources set for them by executing the scheduling tasks of the tenants, so as to execute the job tasks of the tenants. It is possible to independently perform task scheduling for tenants in a scenario where multiple tenants share a scheduler, achieve logical isolation between tenants, and further improve the performance of task scheduling for multiple tenants in the computing cluster.
[0009] In a possible implementation manner, the task scheduling method provided by this application further includes: setting a scheduling policy for the first tenant.
[0010] In a possible implementation manner, scheduling the first computing resource for the job tasks in the task queue of the first tenant includes: scheduling the first computing resource for the job tasks in the task queue according to the scheduling policy.
[0011] This application supports configuring scheduling policies for tenants. According to the different job scenarios of the tenants, different scheduling policies are customized for different tenants, so as to perform task scheduling for the tenants based on the scheduling policies of the tenants.
[0012] In a possible implementation manner, the task scheduling method provided by this application further includes: setting a tenant priority for the first tenant, and this tenant priority can be used to determine the scheduling priority of the first scheduling task.
[0013] In a possible implementation manner, executing the first scheduling task includes: executing the first scheduling task according to the scheduling priority.
[0014] This application supports configuring tenant priorities for tenants, so as to perform task scheduling according to the tenant priorities, ensure the fairness of task scheduling between tenants, and ensure that high-priority tenants can obtain more scheduling opportunities.
[0015] In a possible implementation manner, executing the first scheduling task according to the scheduling priority includes: determining whether the first scheduling task in the priority queue is the highest scheduling priority, and the priority queue further includes the second scheduling task of the second tenant; if the first scheduling task is the highest scheduling priority, then execute the first scheduling task; if the first scheduling task is not the highest scheduling priority, then execute the second scheduling task.
[0016] In a possible implementation manner, the task scheduling method provided by this application further includes: dynamically updating the scheduling priority of the first scheduling task according to the executed duration of the first scheduling task. Dynamically updating the scheduling priority of the scheduling task can ensure the fairness of task scheduling between tenants and ensure that high-priority tenants can obtain more scheduling opportunities.
[0017] In a possible implementation, the computing cluster further includes a second tenant. The second computing resources of the second tenant share the physical computing resources of the computing cluster with the first computing resources of the first tenant. The task scheduling method provided in this application further includes: dividing the first scheduling task and the second scheduling task of the second tenant into multiple subtasks respectively. In this way, executing the first scheduling task according to the scheduling priority includes: if the first scheduling task has the highest scheduling priority, then execute the first subtask of the first scheduling task, schedule the first computing resources for the job tasks in the task queue of the first tenant to execute the job tasks of the first tenant; then update the scheduling priority of the first scheduling task; if the first scheduling task after the scheduling priority is updated does not have the highest scheduling priority, then execute the second subtask of the second scheduling task, schedule the second computing resources for the job tasks in the task queue of the second tenant to execute the job tasks of the second tenant. Splitting the scheduling tasks can switch and schedule between the scheduling tasks of different tenants, enabling the smooth scheduling of the scheduling tasks of different tenants and ensuring the task scheduling opportunities of each tenant.
[0018] In a second aspect, this application provides a task scheduling device. The task scheduling device includes various modules for implementing the method described in the first aspect and any one of its possible implementations, such as a configuration module, a creation module, an acquisition module, a scheduling module, an update module, a splitting module, etc.
[0019] The task scheduling device has the function of implementing the behaviors in the method examples described in the first aspect and any one of its possible implementations. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0020] In a third aspect, this application provides a computing device, including a memory and at least one processor connected to the memory. The memory is used to store computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the at least one processor, the computing device is caused to execute the method described in the first aspect and any one of its possible implementations.
[0021] In a fourth aspect, this application provides a computer-readable storage medium storing computer instructions. When the computer instructions are run on a computer, the method described in the first aspect and any one of its possible implementations is executed.
[0022] In a fifth aspect, this application provides a computer program product. The computer program product includes computer instructions. When the computer instructions are run on a computer, the method described in the first aspect and any one of its possible implementations is executed.
[0023] In a sixth aspect, the present application provides a chip system, including: a processor, configured to call and run a computer program from a memory, such that a storage device installed with the chip system executes the method according to any one of the first aspect and its possible implementations.
[0024] It should be understood that for the beneficial effects achieved by the technical solutions of the second to sixth aspects of the present application and their corresponding possible implementation manners, reference may be made to the technical effects of the first aspect and its corresponding possible implementation manners described above, and details are not repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 One of the schematic diagrams of the architecture of an HPC cluster provided by an embodiment of the present application;
[0026] Figure 2 Another schematic diagram of the architecture of an HPC cluster provided by an embodiment of the present application;
[0027] Figure 3 Another schematic diagram of the architecture of an HPC cluster provided by an embodiment of the present application;
[0028] Figure 4 A schematic diagram of the hardware of a management node provided by an embodiment of the present application;
[0029] Figure 5 A schematic diagram of the software architecture of a scheduler provided by an embodiment of the present application;
[0030] Figure 6 A schematic diagram of the creation process of a logical cluster in a task scheduling method provided by an embodiment of the present application;
[0031] Figure 7 One of the schematic diagrams of the process of a task scheduling method provided by an embodiment of the present application;
[0032] Figure 8 One of the schematic diagrams of the framework of a task scheduling method provided by an embodiment of the present application;
[0033] Figure 9 Another schematic diagram of the framework of a task scheduling method provided by an embodiment of the present application;
[0034] Figure 10 A schematic diagram of the process of splitting a scheduling process in a task scheduling method provided by an embodiment of the present application;
[0035] Figure 11 Another schematic diagram of the process of a task scheduling method provided by an embodiment of the present application;
[0036] Figure 12 A schematic diagram of the state transition of a scheduling task in a task scheduling method provided by an embodiment of the present application;
[0037] Figure 13 One of the schematic structural diagrams of a task scheduling device provided by an embodiment of the present application;
[0038] Figure 14 Another schematic structural diagram of a task scheduling device provided by an embodiment of the present application. Detailed implementation manners
[0039] The term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0040] The terms "first", "second", etc. in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order of the objects.
[0041] In the embodiments of the present application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0042] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more.
[0043] First, some technical terms involved in a task scheduling method and device provided by an embodiment of the present application are explained.
[0044] 1. Multi-tenant
[0045] Multi-tenant is a software architecture. The multi-tenant architecture uses multi-tenant technology (or called multiple rental technology) to enable a single software instance (or called application instance) to provide services for multiple tenants, that is, multiple tenants use the same set of programs. For example, software-as-a-service (SaaS) is a multi-tenant architecture.
[0046] It should be noted that in the multi-tenant architecture, it is necessary to ensure that multiple tenants are unaware of each other, that is, data isolation between multiple tenants is achieved under the same set of programs.
[0047] In the embodiments of the present application, the renter of the software instance can be called a tenant or a customer. In the following embodiments, tenant and customer are the same concept.
[0048] It is understandable that each tenant has one or more users, and the user is the actual user of the software instance of the tenant.
[0049] With the in-depth development of cloud computing, multi-tenancy has become the core architecture in cloud services. Currently, computing clusters (such as HPC clusters) are also rapidly developing towards the direction of cloud services. In an HPC cluster, there is also a need for multi-tenancy, that is, multiple tenants share a software instance in the HPC cluster. For example, the software instance can be the scheduler of the HPC cluster.
[0050] For an HPC cluster, the scheduler is a core management software that can help tenants manage the computing resources of the computing nodes in the HPC cluster and schedule the jobs (or job tasks) submitted by users to the corresponding computing resources for execution.
[0051] 2. Task Scheduling for Multi-Tenancy
[0052] The task scheduling for multi-tenancy can also be called tenant scheduling, which refers to the switching strategies and processes among multiple tenants when the scheduling processes of multiple tenants share a scheduler. In the embodiments of the present application, the job scheduling process within a tenant is abstracted as a task (or scheduling task), and the scheduler's scheduling of the tasks of multiple tenants is the task scheduling for multi-tenancy.
[0053] It should be particularly noted the difference between the above-mentioned task scheduling (i.e., tenant scheduling) and job scheduling (i.e., job task scheduling). Job scheduling refers to the matching algorithm and process between multiple jobs submitted by a tenant and the computing resources owned by the tenant.
[0054] Taking the computing cluster as an HPC cluster as an example, refer to Figure 1 , the HPC cluster includes one or more management nodes 101 and multiple computing nodes (or multiple working nodes) 102. Among them, the scheduler is deployed on the management node 101, and the scheduler allocates computing resources to the jobs submitted by users under the tenant.
[0055] In some implementation manners, the management node 101 may include a login node and a master control node. The login node is used for users to log in to the scheduler. For example, the client of the scheduler is deployed on the login node, and users at the user level and management level can log in to the scheduler through the client in the login node. The master control node is used to implement the management functions of the scheduler, and other functional modules of the scheduler are deployed on the master control node. For example, a user management module, a job management module, a resource management module, and a scheduling policy management module, etc.
[0056] Currently, in an HPC cluster, usually one HPC cluster is configured for one tenant, and each tenant has an independent scheduler. Currently, most HPC schedulers (schedulers in the HPC cluster), such as LSF, Slurm, openPBS, and the scheduling software Kubernetes for AI clusters and the big data scheduler YARN, do not support multi-tenancy, that is, these schedulers only support one scheduler corresponding to one tenant and do not support task scheduling in a multi-tenant architecture.
[0057] Reference Figure 2 , one method to implement task scheduling for multi-tenancy is to divide the original HPC cluster into multiple independent sub-clusters, with each tenant corresponding to an independent sub-cluster. Each sub-cluster is configured with an independent management node, and a scheduler is deployed on the management node to schedule the jobs submitted by the tenant. In this method, an independent physical cluster (i.e., the above-mentioned sub-cluster) is allocated for each tenant, and the physical resources between the sub-clusters corresponding to different tenants are isolated. Therefore, the isolation between tenants is relatively good. However, a separate management node needs to be configured for each tenant's sub-cluster, and as the number of tenants increases, the number of management nodes also increases, resulting in a relatively high cost for implementing multi-tenant task scheduling.
[0058] Reference Figure 3 , another method for multi-tenant task scheduling is that multiple tenants still share one HPC cluster. The computing resources of the computing nodes in the HPC cluster are pooled, logically partitioned, and managed, that is, multiple tenants share the physical computing resources of the HPC cluster. In this method, the computing resources used by multiple tenants to process jobs are isolated, and the scheduler deployed on the management node of the HPC cluster performs unified task scheduling for multiple tenants. In this method, only the computing resources are isolated for multiple tenants, and the isolation degree is relatively low.
[0059] To address the above problems, the embodiments of the present application provide a task scheduling method. This method is applied to the management node in the computing cluster. In this method, the management node sets the first computing resources for the first tenant in the computing cluster; when it is necessary to execute the job task of the first tenant, a first scheduling task that can schedule the computing resources for the job task of the first tenant is started for the first tenant; then the job task of the first tenant is obtained and sent to the task queue of the first tenant; and the first scheduling task is executed to schedule the first computing resources for the job tasks in the task queue of the first tenant to execute the job task of the first tenant. In this method, since a scheduling task is started and computing resources are set for each tenant, independent task scheduling can be performed for tenants. Therefore, in the scenario of multi-tenants sharing a scheduler, logical isolation between tenants can be achieved.
[0060] The task scheduling method provided by the embodiments of this application is executed by a management node. Optionally, the management node can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the management node can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0061] Exemplarily, continuing to refer to Figure 4 , the management node may include: one or more processors 401, a memory 402, and a communication interface 403. Among them, the processors 401, the memory 402, and the communication interface 403 may be connected through a bus 404, or connected to each other in other ways. Optionally, various components included in the management node may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.
[0062] Among them, the processor 401 is the control center of the management node. The processor 401 may be a CPU, or other general-purpose processors, such as a microprocessor or any conventional processor, etc. Optionally, the general-purpose processor 401 may include one or more processing cores.
[0063] The controller in the processor 401 is the nerve center and command center of the management node. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions. Optionally, a memory may also be provided in the processor 401 for storing instructions and data.
[0064] The memory 402 includes, but is not limited to, a random access memory (RAM), a read only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or an optical memory, a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer. In the embodiments of this application, the memory 402 can store information such as computer instructions.
[0065] In a possible implementation, the memory 402 can exist independently of the processor 401. The memory 402 can be connected to the processor through the bus 404 for storing data, instructions, or program code. When the processor calls and executes the instructions or program code stored in the memory, the relevant steps in the method provided by the embodiments of this application can be implemented.
[0066] In another possible implementation, the memory 402 can also be integrated with the processor.
[0067] The communication interface 403 can be a transceiver module for communicating with other devices or communication networks, such as Ethernet, RAN, wireless local area networks (WLAN), etc. The communication interface 403 can receive instructions, messages, data, etc. The transceiver module can be a device such as a transceiver or a transceiver unit. Optionally, the communication interface 403 can also be a transceiver circuit located within the processor to implement signal input and signal output of the processor. The communication interface 403 can be a wired interface (port), such as a fiber distributed data interface (FDDI), a gigabit ethernet (GE) interface, or the communication interface 403 can also be a wireless interface.
[0068] The bus 404 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus can also be divided into a serial bus and a parallel bus. For ease of representation, Figure 4 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0069] Optionally, the management node in the embodiments of the present application can further include an input / output interface 405. The input / output interface 405 is used to connect to an input device and receive information input by the user through the input device (such as the login information of the tenant, the job submitted by the tenant). The input device includes but is not limited to a keyboard, a touch screen, a microphone, etc. The input / output interface 405 is also used to connect to an output device and output the processing result of the processor 401. The output device includes but is not limited to a display, a printer, etc.
[0070] It should be noted that Figure 4 the management node in the figure is only an example of the management node, and the management node can have more or fewer components than those Figure 4 shown in the figure, can combine two or more components, or can have different component configurations.
[0071] It should be understood that in the embodiments of the present application, the task scheduling method is executed by a scheduler deployed on a management node in a computing cluster. Refer to Figure 5 , which shows a software architecture of a scheduler deployed on a management node provided in the embodiments of the present application. The scheduler includes: a command line interface (CLI), a graphical interface (Portal), a REST API, a tenant service, a cluster service, a scheduling service, and a unified resource management. Figure 5 The scheduler shown supports multiple tenants.
[0072] Among them, the command line interface (CLI): is a tool for users to log in to the scheduler and is deployed on the login node in the management node.
[0073] The graphical interface (Portal): is an interface for users to submit jobs and is deployed on the login node in the management node.
[0074] The REST API: is an application programming interface (API) that follows the Representational state transfer (REST) architectural specification, and applications or devices can be connected and communicate with each other based on the REST API.
[0075] The tenant service: is used for functions such as tenant management, cluster management, and dynamic scaling management. Among them, tenant management may include, but is not limited to, the management of tenant permissions, cluster management includes, but is not limited to, maintaining the correspondence between tenants and logical clusters, and dynamic scaling management includes, but is not limited to, adding or deleting nodes in the computing cluster.
[0076] The cluster service is used to generate logical clusters, encapsulate data and status information during the task scheduling process, and maintain the relationship between tenants and logical clusters. The cluster service can create multiple logical clusters. Each tenant corresponds to one or more logical clusters, and each logical cluster includes functional modules for job management, user management, resource management, and scheduling policy management.
[0077] It should be noted that in the embodiments of the present application, the logical clusters of tenants are isolated from each other. That is to say, each tenant has its own logical cluster, and job management, user management, resource management, and scheduling policy management in the logical cluster manage each tenant respectively. That is to say, each tenant is managed independently. For example, under one tenant, the numbers of jobs submitted by users are consecutive, and the job numbers of different tenants are not related; the users under each tenant are also managed independently; each tenant corresponds to an independent scheduling policy (strategy or method for processing job tasks) for that tenant, supporting the configuration of customized scheduling policies for the tenant; furthermore, the computing resources used to process the jobs of each tenant are also different.
[0078] In the embodiments of the present application, since the logical clusters of tenants are isolated from each other, tenants do not perceive each other. For one tenant, it is equivalent that the tenant corresponds to a separate computing cluster, and multiple logical clusters share the resources of the scheduler deployed on a management node. The scheduler schedules the computing resources (set for the tenant) of the job tasks of multiple tenants, that is, the function of the scheduler natively supporting multi-tenancy is realized.
[0079] Scheduling service: used to provide some other services during task scheduling. For example, in a large-scale multi-tenant scenario, the scheduling service can adjust the priorities of tenants, split the scheduling process, construct a scheduling context (i.e., the execution status of the scheduling task), retrieve context switching, and manage the parallel or serial nature of the scheduling process, etc.
[0080] Unified resource management: used for users to manage resources, such as creating or deleting resources corresponding to logical clusters. Unified resource management can shield the differences of different resource providers, provide a standardized interface for tenant services, and can be docked with any cloud or offline cluster resource management software.
[0081] Based on the above, Figure 5 The shown scheduler can support multi-tenancy. Multiple tenants share one software instance (such as the scheduler), greatly reducing the occupation of management nodes. The more the number of tenants, the lower the cost; and since there is no need to deploy one software instance for each tenant, therefore, the operation and maintenance work of tenants for software instances can be reduced.
[0082] Combined with Figure 5 the software architecture of the shown scheduler, logical clusters can be created for each tenant of the computing cluster, and during operation, the created logical clusters can also be deleted, modified, or queried. The following combines Figure 6 , taking one tenant (hereinafter referred to as the first tenant) as an example, to introduce the process of creating a logical cluster for the tenant.
[0083] S601. The tenant service module receives a cluster creation request triggered by the tenant administrator.
[0084] When a new tenant (e.g., called the first tenant) is added to the computing cluster, the tenant administrator (operation and maintenance personnel) of the computing cluster sends a cluster creation request to the tenant service. This creation request is used to request the creation of a logical cluster for the first tenant. Optionally, the cluster creation request may include information about the first tenant (such as the identifier of the first tenant) and the resource requirement information of the first tenant. The resource requirement information is used to indicate the demand of the first tenant for computing resources used to process jobs. For example, the resource requirement information indicates the number of computing nodes and the configuration requirements for the computing nodes (such as the number of cores, memory size, etc.).
[0085] S602. The tenant service creates the correspondence between the tenant and the logical cluster.
[0086] In one implementation, the tenant service may generate an identifier for a logical cluster for the first tenant and save the correspondence between the identifier of the first tenant and the identifier of the logical cluster, thereby establishing the correspondence between the first tenant and the logical cluster.
[0087] Optionally, after receiving the cluster creation request, the tenant service also verifies whether the first tenant has the permission to create a cluster. For example, the tenant management function in the tenant service verifies the permissions of the first tenant to determine whether the first tenant has the permission to create a logical cluster. If the first tenant has the permission to create a logical cluster, the tenant service establishes the correspondence between the first tenant and the logical cluster.
[0088] S603. The tenant service sends a resource creation request to the unified resource management.
[0089] This resource creation request is used to request to set the first computing resources for the first tenant, and these computing resources are used to process the jobs of the first tenant. The above resource creation request includes the resource requirement information of the first tenant.
[0090] S604. After receiving the resource creation request, the unified resource management sets the first computing resources for the first tenant.
[0091] Specifically, the unified resource management sets the computing resources for the first tenant by allocating the computing resources of the computing nodes in the computing cluster (such as allocating processing cores and memory) to the first tenant.
[0092] S605. The unified resource management sends resource information to the cluster service.
[0093] The resource information is the information of the first computing resources set for the first tenant. For example, if the allocated first computing resources are computing nodes, the resource information may be the IP address of the computing node or the name of the computing node, etc.
[0094] S606. The tenant service sends a cluster creation request to the cluster service.
[0095] S607. After the cluster service receives the cluster creation request sent by the tenant service, it creates a logical cluster.
[0096] The cluster service creating a logical cluster includes creating function modules for job management, user management, scheduling policy management, and resource management. The logical cluster of the first tenant is logically isolated from the logical clusters of other tenants, and this logical cluster is used to manage the jobs of the first tenant, manage the users of the first tenant, manage the scheduling policy of the first tenant, and manage the computing resources for executing the jobs of the first tenant.
[0097] S608. After the cluster service receives the resource information sent by the unified resource management, it establishes a mapping relationship between the first tenant and the first computing resource.
[0098] After the cluster service creates the logical cluster of the first tenant, it stores the mapping relationship between the first tenant and the first computing resource, which is equivalent to storing the mapping relationship between the logical cluster and the first computing resource.
[0099] Through the above S601 - S608, the creation of the logical cluster of the first tenant is completed. It can be understood that in the embodiments of the present application, during the process of creating the logical cluster or after the logical cluster is created, the logical cluster can be configured, that is, parameter configuration is performed on the first tenant. After that, task scheduling can be performed on multiple tenants based on the created logical cluster.
[0100] In the embodiments of the present application, each tenant of the computing cluster corresponds to tenant configuration items (or called configuration parameters). Some of the tenant configuration items are configuration items shared by all tenants, and the configuration items shared by tenants are configured by the operation and maintenance personnel; some configuration items are unique to the tenant, and the unique configuration items of the tenant are independently configured by each tenant. Referring to Table 1, it is an example of the tenant's configuration items.
[0101] Table 1
[0102]
[0103] Combined with Table 1, it can be seen that a scheduling policy and tenant priority are configured for the first tenant. The present application supports configuring an independent scheduling policy for the tenant. For example, according to the different job scenarios of the tenant, different scheduling policies are customized for different tenants, so as to perform task scheduling on the tenant based on the tenant's scheduling policy. And, the present application supports configuring tenant priority for the tenant, so as to perform task scheduling according to the tenant priority, ensuring fairness in task scheduling among tenants and ensuring that high-priority tenants can obtain more scheduling opportunities.
[0104] Optionally, for the created logical cluster, subsequent deletion, modification, and query operations can also be performed on the logical cluster.
[0105] When the first tenant exits the computing cluster, a cluster deletion request can be sent to the tenant service. The cluster deletion request can include the identifier of the first tenant. As a result, the tenant service deletes the correspondence between the first tenant and the logical cluster, and sends a cluster deletion request to the unified resource management and the cluster service respectively. The unified resource management deletes the first computing resource corresponding to the logical cluster, and the cluster service deletes the mapping relationship between the first tenant and the first computing resource, and deletes the logical cluster and the configuration items of the first tenant.
[0106] When it is necessary to query the relevant information of the first tenant or the logical cluster, a query request can be sent to the tenant service. For example, query the configuration items of the first tenant, or query the information of the logical cluster corresponding to the first tenant. As a result, the tenant service or the cluster service can return the configuration items or the cluster information.
[0107] When it is necessary to modify the relevant information of the first tenant or the logical cluster, a modification request can be sent to the tenant service. For example, request to modify the configuration items of the first tenant or modify the information of the logical cluster (such as modifying the configuration of the first computing resource). As a result, the relevant modules can make modifications according to the modification requirements.
[0108] Combined with Figure 4 Regarding the architecture introduction of the scheduler provided in the embodiments of the present application, the computing cluster supports multiple tenants. Multiple tenants share one scheduler, and the scheduling process of each tenant is defined as a scheduling task. Computing resources and scheduling tasks for processing job tasks are set for each tenant. After the scheduler obtains the job tasks of multiple tenants, it executes the scheduling tasks of the tenants, and schedules the jobs of the tenants to the corresponding computing resources for the job tasks of the tenants.
[0109] The task scheduling method provided in the embodiments of the present application can be applied to the scenario of a computing cluster with a multi-tenant architecture based on cloud services, or can also be applied to the scenario of a computing cluster with a multi-tenant architecture based on physical resources, without specific limitation.
[0110] Combined with the relevant descriptions of the computing cluster, the management node, and the scheduler in the above embodiments, the process of the task scheduling method provided in the embodiments of the present application is described below. This task scheduling method is applied to the management node in the computing cluster, that is, it is executed by the management node in the computing cluster (specifically, executed by the scheduler deployed on the management node). The following describes it with the scheduler as the execution subject. As Figure 7 shown, this method includes S701 - S704.
[0111] S701. Set the first computing resource for the first tenant.
[0112] In the embodiments of the present application, a computing cluster may include multiple tenants, and the scheduler may set (i.e., allocate) corresponding computing resources for each tenant of the computing cluster. For example, the computing cluster includes a first tenant and a second tenant. The first computing resources are set for the first tenant, and the second computing resources are set for the second tenant, and the first computing resources and the second computing resources share the physical computing resources of the computing cluster.
[0113] Taking the first tenant as an example, the specific process of the scheduler setting the first computing resources for the first tenant may refer to the descriptions of S603 to S605 in the process of creating the logical cluster of the first tenant described in the above embodiments, which will not be elaborated here.
[0114] S702. When it is necessary to execute the job task of the first tenant, start a first scheduling task for the first tenant, and the first scheduling task is used to schedule computing resources for the job task of the first tenant.
[0115] In the embodiments of the present application, the first scheduling task refers to the scheduling process of the job task of the first tenant, and each tenant corresponds to a scheduling task. Starting the first scheduling task for the first tenant can be understood as: creating a first scheduling task for the first tenant, or generating a first scheduling task for the first tenant.
[0116] S703. Obtain the job task of the first tenant and send the job task to the task queue of the first tenant.
[0117] The task queue of the tenant includes job tasks submitted by multiple users under the tenant. Optionally, the job tasks in the task queue of the first tenant may be sorted according to the reception time of the job tasks, or sorted by other methods, which are not limited in the embodiments of the present application.
[0118] S704. Execute the first scheduling task to schedule the first computing resources for the job tasks in the task queue to execute the job tasks.
[0119] The scheduler executing the scheduling task specifically includes executing the scheduling task of the tenant based on the scheduling resources (resources used by the scheduler to execute the scheduling task) allocated for the scheduling task.
[0120] It can be understood that the scheduling resources may be processes, threads or coroutines, which are specifically selected according to actual needs and are not limited in the present application. The relationship between processes, threads and coroutines is that a process may include multiple threads, and a thread may include multiple coroutines. For the sake of convenience of description, in the embodiments of the present application, the scheduling resources of the scheduler are described by taking scheduling threads as an example.
[0121] In the embodiments of the present application, when the scheduler executes the scheduling tasks of tenants, in fact, the scheduler calls threads to execute the scheduling tasks. The scheduler can create multiple scheduling threads to execute the scheduling tasks. When the resources of the scheduler are sufficient (i.e., the resources used by the scheduler in the management node), the scheduler can allocate independent scheduling resources to each of the multiple scheduling tasks. For example, one scheduling thread is used to execute one scheduling task (one tenant corresponds to one independent scheduling thread), and the scheduler executes the scheduling tasks of multiple tenants in parallel. When the scheduling resources of the scheduler are insufficient, multiple scheduling tasks share the scheduling resources of the scheduler. For example, one scheduling thread is used to execute multiple scheduling tasks (multiple tenants share one scheduling thread), and the scheduler executes the scheduling tasks of multiple tenants sequentially.
[0122] In some implementation manners, the resources of the scheduler can be evenly distributed to multiple scheduling tasks, or different proportions of resources can be allocated to multiple scheduling tasks according to other factors (such as the tenant weights and / or quality of service configured for the tenants as described above). The embodiments of the present application do not make any limitations.
[0123] In summary, in the task scheduling method provided by the embodiments of the present application, since computing resources can be set for each tenant of the computing cluster, and scheduling tasks are started for the tenants, the job tasks of the tenants are scheduled to the computing resources set for them by executing the scheduling tasks of the tenants, so as to execute the job tasks of the tenants. It is possible to perform independent task scheduling for tenants in a scenario where multiple tenants share a scheduler, realize logical isolation between tenants, and further improve the performance of task scheduling for multiple tenants in the computing cluster.
[0124] According to the description in the above embodiments, the scheduler can set a scheduling policy for each tenant. Based on this, in one implementation manner, in the above S704, scheduling the first computing resource for the job tasks in the task queue of the first tenant includes: scheduling the first computing resource for the job tasks in the task queue according to the scheduling policy.
[0125] Exemplarily, the scheduling policy can be to schedule computing resources for job tasks according to the reception time of the job tasks or the priority of the job tasks, etc. For example, scheduling the first computing resource according to the priority of the job tasks. Assume that the priority of the first job task is higher than that of the second job task, then the first computing resource is preferentially scheduled for the first job task.
[0126] In the embodiments of the present application, a scheduler is deployed on the management node. For the convenience of description, in the following embodiments, the resources of the management node are the resources of the scheduler, and the resources of the management node are used for task scheduling of tenants.
[0127] In a computing cluster, generally, the resources of the management node used to deploy the scheduler are limited. In different scenarios, during the process of task scheduling for multiple tenants, the resource allocation of the management node needs to be considered. The resources of the management node include processing resources (such as processors / processing cores) and storage resources (such as memory).
[0128] The task scheduling method provided by the embodiments of this application can be applied to scenarios with different tenant scales. The following will be described with three different tenant scales. The three scenarios with different tenant scales are Scenario 1 (small-scale multi-tenant scenario), Scenario 2 (medium-scale multi-tenant scenario), and Scenario 3 (large-scale multi-tenant scenario).
[0129] Scenario 1, Small-scale multi-tenant scenario
[0130] For the small-scale multi-tenant scenario, the number of tenants is relatively small and the resources of the management node are relatively sufficient. In this case, when the scheduler executes the scheduling task, a scheduling thread can be created for each scheduling task, that is, independent resources in the management node are allocated to each tenant as scheduling resources, and scheduling resources that can meet the requirements are allocated to each scheduling task to execute the scheduling task of the tenant.
[0131] Reference Figure 8 , Tenant 1 to Tenant n respectively correspond to a scheduling thread. The scheduling method for multiple scheduling tasks is that multiple scheduling threads run in parallel, that is, the scheduler can call multiple scheduling threads to execute the scheduling tasks of multiple tenants in parallel, realizing the parallelization of the scheduling process between tenants and improving the scheduling efficiency. For a tenant, the scheduling thread executes the scheduling task of the tenant, and schedules the job tasks in the task queue of the tenant to the computing resources allocated to the tenant to execute the job tasks in the task queue of the tenant, that is, the scheduling process within the tenant is sequentialized, that is, time-sharing scheduling (serial scheduling) is adopted.
[0132] In one implementation, the range of the processing resources and storage resources of the management node occupied by the scheduling thread can be configured based on the operating system kernel mechanism. For example, the resources of the management node can be divided based on the Cgroup mechanism of the linux kernel to create different scheduling threads. It can be understood that Cgroup can control the allocation of processing resources and storage resources of processes or threads.
[0133] Exemplarily, referring to the partial configuration items of Cgroup shown in Table 2 below, the allocation of processing resources and storage resources can be realized by using key configurations such as cpuset.cpus and memory.limit_in_bytes in Cgroup.
[0134] Table 2
[0135] Parameter Name Parameter Description cpuset.cpus CPU Usage Range memory.limit_in_bytes Maximum Memory Usage Limit
[0136] Scenario 2: Medium-scale multi-tenant scenario
[0137] In a medium-scale multi-tenant scenario, the resources of the management node may be insufficient. In this case, scheduling resources can be allocated for each scheduling task according to the tenant priority. For example, a scheduling thread can be created for each scheduling task according to the tenant priority to execute the scheduling tasks of the tenant.
[0138] Reference Figure 9 , in the embodiments of the present application, a tenant priority is configured for each tenant, and the tenant priority may include tenant weight and / or quality of service. The relationship between the tenant priority and the tenant weight is that the greater the tenant weight of the tenant, the higher the tenant priority. For example, for a tenant with urgent tasks, the tenant weight is larger and the tenant priority is high. The relationship between the tenant priority and the quality of service of the tenant is that the higher the quality of service of the tenant, the higher the tenant priority.
[0139] Optionally, there is a functional relationship between the tenant priority, the tenant weight, and the quality of service, and the tenant priority can be calculated according to the preset functional relationship.
[0140] In one implementation, when allocating the resources of the management node to the tenants, the resource allocation ratio can be determined according to the tenant priorities of multiple tenants, and then the corresponding proportion of resources can be allocated to different scheduling tasks according to the resource allocation ratio.
[0141] Exemplarily, if the tenant priority includes the tenant weight, the computing cluster includes 3 tenants, and the tenant weights of the three tenants are 100, 20, and 10 respectively, then the determined resource allocation ratio is 10:2:1. Then, the processing resources and storage resources of the management node are allocated according to 10:2:1, and three scheduling threads are created to execute the scheduling tasks of the 3 tenants.
[0142] Exemplarily, referring to some configuration items of Cgroup shown in Table 3 below, the allocation of processing resources and storage resources can be realized by using key configurations such as cpuset.cpus, cpu.cfs_quota_us & cpu.cfs_period_us, and cpu.shares in Cgroup.
[0143] Table 3
[0144]
[0145] In the embodiments of the present application, allocating the resources of the management node to multiple scheduling tasks according to the tenant priority can ensure that high-priority customers are allocated more scheduling resources for executing scheduling tasks, so that high-priority tenants can obtain more scheduling opportunities.
[0146] Scenario 3: Large-scale multi-tenant scenario
[0147] In a large-scale multi-tenant scenario, the resources of the management node are insufficient (such as resource shortage). In this case, multiple scheduling tasks share the resources of the management node. For example, among multiple created scheduling threads, one scheduling thread is used to execute multiple scheduling tasks.
[0148] It can be understood that there is a corresponding relationship (or mapping relationship) between the tenant (or the tenant's scheduling task) and the scheduling thread. In the embodiments of the present application, the tenant name can be used as the key value to calculate the hash values of all scheduling tasks. Different hash values correspond to different queues (each scheduling thread corresponds to a queue, that is, different hash values correspond to different scheduling threads), and then the scheduling tasks of the tenant are mapped to the corresponding queues. This queue is the queue of the scheduling tasks of multiple tenants, and multiple scheduling tasks are arranged in the queue according to the scheduling priorities of the scheduling tasks. Therefore, this queue can be called a priority queue. The above tenant priority can be used to determine the scheduling priority of the scheduling tasks of the tenant, which will be described in detail in the following embodiments.
[0149] In one implementation, the resource amounts of the management node occupied by multiple scheduling threads can be the same or different. For example, the resources of the management node are allocated based on information such as the number of tenants corresponding to one scheduling thread and / or the tenant priorities of at least two tenants corresponding to the scheduling thread.
[0150] In the embodiments of the present application, in a large-scale multi-tenant scenario, the priority queue includes the first scheduling task of the first tenant and the second scheduling task of the second tenant. During the process of the scheduler performing task scheduling for multiple tenants, the scheduler calls one scheduling thread to execute the scheduling tasks of multiple tenants, including: executing the scheduling tasks according to the scheduling priorities of the multiple scheduling tasks.
[0151] Taking the first tenant as an example, executing the first scheduling task of the first tenant specifically includes: executing the first scheduling task according to the scheduling priority of the first scheduling task. Specifically, it is determined whether the first scheduling task in the priority queue is the highest scheduling priority; if the first scheduling task is the highest scheduling priority, the first scheduling task is executed; if the first scheduling task is not the highest scheduling priority, the second scheduling task is executed. That is, the scheduling tasks in the priority queue are executed in the order from high to low scheduling priority. When the scheduling priority of the first scheduling task is higher than that of the first scheduling task, the first scheduling task is executed first; when the scheduling priority of the first scheduling task is lower than that of the second scheduling task, the second scheduling task is executed first.
[0152] In one implementation, for each scheduling task in the priority queue, the scheduling task can be split (i.e., splitting the job scheduling process of the tenant), and each scheduling task is divided into multiple slices. One slice can be called a subtask. In this way, one scheduling task includes multiple subtasks. For example, for the first scheduling task and the second scheduling task in the priority queue, the first scheduling task and the second scheduling task of the second tenant are respectively divided into multiple subtasks.
[0153] Optionally, the process of splitting the scheduling tasks of the tenant specifically includes: taking the execution points that can be interrupted in the scheduling process of each tenant (the scheduling process of each tenant is the scheduling task of each tenant) as safe points, and then using the safe points as splitting points to divide the job scheduling process of the tenant into multiple slices. The size of the slice is the execution duration of the slice, and the size range of one slice can be 0.1ms (milliseconds) - 5ms. For example, it takes about 1ms to successfully schedule one job, so the scheduling process of one job can be used as one subtask.
[0154] Optionally, the size of the slice is adjusted according to the tenant priority, such as the quality of service (QoS) of the tenant. The higher the quality of service of the tenant, the smaller the slice, that is, the shorter the execution duration of the slice.
[0155] In some embodiments, the scheduling process of one tenant is split according to the characteristics of the scheduling process. For example, the scheduling process includes multiple different stages, and it is determined whether each stage can be split; or, for example, the scheduling process includes multiple jobs, and the point where each job scheduling ends is used as the splitting point.
[0156] Exemplarily, referring to Figure 10 , assume that the scheduling process of one tenant includes 5 stages, where some stages may further include multiple sub - stages, and the multiple sub - stages may include processes such as serial, parallel, and loop. Figure 10 The dotted line in
[0157] represents the scheduling process of one tenant. By splitting the scheduling process, the scheduling task of one tenant is divided into multiple subtasks.
[0158] Exemplarily, a code example for adding a splitting identifier is as follows:
[0159] public void schedule(){
[0160] while (running) {
[0161] @safepoint preShcedule();
[0162] @safepoint
[0163] for (job job : jobs) {
[0164] schedulejob(job);
[0165] }
[0166] @safepoint postSchedule()
[0167] }
[0168] }
[0169] In one implementation, when the scheduling task is split into multiple subtasks, the above-mentioned execution of the first scheduling task according to the scheduling priority includes: if the first scheduling task has the highest scheduling priority, execute the first subtask of the first scheduling task, schedule the first computing resource for the job tasks in the task queue of the first tenant to execute the job tasks of the first tenant; then update the scheduling priority of the first scheduling task; if the first scheduling task after the scheduling priority is updated does not have the highest scheduling priority, execute the second subtask of the second scheduling task, schedule the second computing resource for the job tasks in the task queue of the second tenant to execute the job tasks of the second tenant. Splitting the scheduling task can switch the scheduling between the scheduling tasks of different tenants, enabling the smooth scheduling of the scheduling tasks of different tenants and ensuring the task scheduling opportunities of each tenant.
[0170] For each scheduling thread created by the scheduler, the process of the scheduler calling the scheduling thread to execute task scheduling is similar. The following embodiments will describe in detail the process of task scheduling for tenants taking one scheduling thread as an example.
[0171] In one implementation, the scheduler calls the scheduling thread to periodically execute multiple scheduling tasks in the priority queue. In one scheduling cycle, the scheduler traverses multiple scheduling tasks in a priority queue. For each scheduling task, within the scheduling cycle, the scheduling task includes several state parameters as shown in Table 4 below.
[0172] Table 4
[0173] Status Parameter of Scheduled Task Description curr_num Current Sub-task Number (or Index) of Scheduled Task n_runtime Executed Duration of Scheduled Task (or Normalized Runtime) r_runtime Remaining Executable Duration of Scheduled Task state Task Status of Scheduled Task
[0174] It should be understood that the scheduling priority of a scheduling task is related to the executed duration (n_runtime) of the above scheduling task (n_runtime is related to the tenant priority). Within a scheduling period, the scheduling priority of a scheduling task decreases as the executed duration of the scheduling task increases. The smaller the executed duration of the scheduling task, the higher the scheduling priority of the scheduling task.
[0175] Optionally, in a priority queue, the scheduling tasks are arranged in ascending order of n_runtime. In some cases, if the n_runtime of multiple scheduling tasks is equal, the scheduling tasks of the tenant can be sorted according to the tenant name or randomly sorted. The embodiments of the present application do not make any limitations.
[0176] Within a scheduling period, the executable duration can be allocated to each scheduling task according to the quality of service of the tenant. The executable duration of a scheduling task is T * QoS, where T is the duration of the scheduling period and QoS is the quality of service of the tenant. The sum of the executed duration n_runtime and the remaining executable duration r_runtime of the above scheduling task is T * QoS.
[0177] In one implementation, the embodiments of the present application can also set a task state for each scheduling task. For example, the state of the scheduling task can include the ready state, the running state, the sleeping state, and the blocked state. Among them, the ready state indicates that the scheduling task is in a pending execution state, the running state indicates that the scheduling task is in an executing state, the sleeping state indicates that the scheduling task is in a dormant state, indicating that there is no job for the tenant currently, and the blocked state indicates that the scheduling task is in a locked state, indicating that the executable duration of the scheduling task within the current scheduling period has been used up, that is, the scheduling task cannot be executed.
[0178] Combined with Table 4, it can be understood that the various state parameters of the scheduling task change dynamically within a scheduling period.
[0179] In one implementation, during initialization (i.e., when starting the scheduling task), the various state parameters of the scheduling task can be set as: curr_num = 0, n_runtime = 0, r_runtime = T * QoS, state = ready.
[0180] As Figure 11 shown, when multiple scheduling tasks share a scheduling thread, within a scheduling period, the process of the scheduler scheduling a scheduling thread to execute the scheduling tasks of multiple tenants includes S1101 - S1104.
[0181] S1101. Add the scheduling tasks of multiple tenants to the priority queue in the order of scheduling priorities.
[0182] It should be understood that during the task scheduling process, the scheduling tasks in the priority queue are tasks to be scheduled, that is, the status of the scheduling tasks in the priority queue is the ready status. When the status of a scheduling task is updated to the sleeping status and the blocked status, the scheduling task is removed from the priority queue. Thereby giving other tenants a chance to schedule. For example, it can increase the chance of task scheduling for tenants with a small number of job tasks.
[0183] Reference Figure 12 , taking the first scheduling task of the first tenant as an example, describe the switching process between various statuses of the first scheduling task.
[0184] Within a scheduling period, when the first scheduling task is moved into the priority queue, the status of the first scheduling task is the ready status. When the scheduling priority of the first scheduling task is the highest, the scheduling thread executes the first scheduling task, and the status of the first scheduling task switches from the ready status to the running status.
[0185] After a subtask of the first scheduling task is executed, if the first scheduling task still has remaining executable duration and the scheduling priority of the first scheduling task is not the highest, then the status of the first scheduling task is switched from the running status to the ready status.
[0186] After a subtask of the first scheduling task is executed, if the first scheduling task has no remaining executable duration, then the status of the first scheduling task is switched from the running status to the blocked status.
[0187] When a subtask of the first scheduling task is executed, if the first tenant has no job to execute, the status of the first scheduling task is switched from running to sleeping. This tenant can be regarded as an idle tenant.
[0188] When a new job is submitted by the first tenant, the first tenant changes from having no job to execute to having a job to execute, and the status of the first scheduling task is switched from the sleeping status to the ready status.
[0189] When the timer (used to time the scheduling period) is reset and triggered again, if the first scheduling task was switched to the blocked status in the previous period, then when the timer is triggered again, the status of the first scheduling task is switched from the blocked status to the ready status.
[0190] In the embodiments of the present application, starting a timer and adding multiple scheduling tasks to a priority queue in the order of scheduling priorities includes the following situations:
[0191] 1. If the scheduler initially calls the scheduling thread, set the status of all scheduling tasks that the scheduling thread needs to execute to the ready state, and arrange them in descending order of scheduling priorities. Additionally, during initialization, since the n_runtime of all scheduling tasks is initialized to 0, they can be arranged according to the tenant name or randomly.
[0192] 2. Add the scheduling tasks in the blocked state to the priority queue, and switch the status of the scheduling tasks to the ready state.
[0193] 3. For a tenant without jobs (the status of whose scheduling tasks is sleeping), when a new job is submitted by the tenant, add the tenant's scheduling tasks to the priority queue, and switch the status of the scheduling tasks to the ready state.
[0194] S1102. Obtain the scheduling task with the highest scheduling priority in the priority queue.
[0195] After obtaining the scheduling task with the highest scheduling priority in the priority queue, switch the status of the scheduling task from ready to running.
[0196] In one implementation, if there is no scheduling task in the priority queue, scheduling tasks can be obtained from other priority queues and executed, so as to improve the resource utilization rate of the management node. Update the status of the scheduling tasks obtained from the priority queues of other scheduling threads to the running state.
[0197] S1103. Execute the subtasks of the scheduling task with the highest scheduling priority.
[0198] For a scheduling task, the multiple subtasks included in the scheduling task are arranged in the order of the subtask numbers, for example, execute the subtasks in ascending order of the subtask numbers. In the embodiments of the present application, during the process of the scheduler executing a subtask, record the execution time of the subtask (denoted as delta). After a subtask is executed, update the current subtask number curr_num, the already executed duration n_runtime, and the remaining executable duration r_runtime of the scheduling task.
[0199] Among them, curr_num = curr_num + 1;
[0200] n_runtime = n_runtime + delta * 1024 / weight;
[0201] r_runtime = max(0, r_runtime - delta).
[0202] weight represents the tenant weight of the tenant.
[0203] S1104. Dynamically update the scheduling priority of the scheduling task.
[0204] In one implementation, taking the first scheduling task as an example, the scheduling priority of the first scheduling task can be dynamically updated according to the executed duration (n_runtime) of the first scheduling task. Specifically, update the executed duration (n_runtime) of the first scheduling task, and then update the scheduling priority of the first scheduling task according to the executed durations of each scheduling task in the priority queue.
[0205] The above dynamic update of the scheduling priority of the scheduling task actually updates the priority queue. If the sorting of the executed durations of each scheduling task in the priority queue changes, the scheduling priority of the scheduling task also changes. For example, if the first scheduling task is the scheduling task with the highest scheduling priority in the priority queue, after executing a subtask of the first scheduling task, the executed duration of the first scheduling task is no longer the scheduling task with the smallest executed duration in the priority queue, then update the scheduling priority of the first scheduling task, and the scheduling priority of the first scheduling task is no longer the highest scheduling priority.
[0206] Within a scheduling period, after executing a subtask of a scheduling task, it is also necessary to update the status of the scheduling task. It should be noted that after the status of the scheduling task is updated, the scheduling tasks in the blocked state and the sleeping state are removed from the priority queue, and the remaining scheduling tasks in the priority queue are re-sorted according to the updated scheduling priority.
[0207] Taking the first scheduling task as an example, if the remaining executable duration r_runtime of the first scheduling task is equal to 0, then update the status of the first scheduling task to the blocked state, the executed duration n_runtime = T * QoS, and remove the first scheduling task from the priority queue.
[0208] If the r_runtime of the first scheduling task is not equal to 0 and the first scheduling task has no job, update the status of the first scheduling task to the sleeping state, and update the executed duration n_runtime of the first scheduling task as n_runtime = n_runtime - min(n_runtime1,...), where min(n_runtime1,...) represents the minimum of the executed durations of all scheduling tasks that the scheduling thread needs to execute. It should be noted that subsequently, when a job is submitted to the first scheduling task again, update the status of the first scheduling task from the sleeping state to the ready state, and update n_runtime of the first scheduling task as n_runtime = n_runtime + min(n_runtime1,...).
[0209] It can be understood that when the first scheduling task has no job, the first scheduling task is no longer executed. During the current scheduling cycle, the executed duration of the first scheduling task will remain unchanged. As other scheduling tasks are executed, the executed durations of other scheduling tasks will increase. When a job is submitted to the first scheduling task, compared with other scheduling tasks, the scheduling priority of the first scheduling task is the highest, and the scheduling thread will give priority to executing the first scheduling task, and other scheduling tasks will not be executed for a long time. To ensure the fairness of task scheduling for multiple tenants, update the executed duration of the first scheduling task to appropriately increase the executed duration of the first scheduling task and avoid other scheduling tasks not being scheduled for a long time, which can, to a certain extent, enable other schedulings to obtain more scheduling opportunities.
[0210] If the r_runtime of the first scheduling task is not equal to 0 and the first scheduling task has a job, determine the status of the scheduling task according to the scheduling priority of the first scheduling task. According to the above embodiments, after executing a sub-task, n_runtime of the first scheduling task is n_runtime = n_runtime + delta * 1024 / weight. Update the status of the first scheduling task according to n_runtime. If n_runtime of the first scheduling task is the smallest among the n_runtimes of multiple scheduling tasks, the status of the first scheduling task remains the running state. If n_runtime of the first scheduling task is not the smallest among the n_runtimes of multiple scheduling tasks, update the status of the first scheduling task to the ready state.
[0211] It should be noted that when a second tenant is newly added to the computing cluster, a second scheduling task is started for the second tenant, and the second scheduling task is mapped to a scheduling thread of the scheduler. In this case, the n_runtime of the second scheduling task is set to: n_runtime = min(n_runtime1, …) + 0.5ms, the status of the priority queue is set to the ready state, and the second scheduling task is added to the priority queue.
[0212] In the embodiment of the present application, within a scheduling period, S1102 - S1104 are repeatedly executed until the timer expires.
[0213] It can be understood that the above method is executed by a task processing device (the task processing device can be a scheduler deployed on a management node). In order to implement the above functions, the task processing device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the method steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0214] The embodiments of the present application can divide the function modules of the above task processing device according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software function module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0215] In the case of dividing each function module corresponding to each function, Figure 13 A possible structural schematic diagram of the task processing device involved in the above embodiments is shown. The task processing device includes a configuration module 1301, a creation module 1302, an acquisition module 1303, and a scheduling module 1304.
[0216] Among them, the configuration module 1301 is used to execute S701 in the above method embodiment; the creation module 1302 is used to execute S702 in the above method embodiment; the acquisition module 1303 is used to execute S703 and S1102 in the above method embodiment; the scheduling module 1304 is used to execute S704 and S1103 in the above method embodiment.
[0217] Optionally, the task processing device provided in the embodiments of the present application further includes an update module 1305 and a splitting module 1306. Among them, the update module 1305 is used to execute S1104 in the above method embodiment; the splitting module 1306 is used to execute splitting the scheduling task into multiple subtasks.
[0218] Each module of the above task processing device can also be used to execute other actions in the above method embodiment. All relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here.
[0219] In the case of adopting an integrated unit, Figure 14 Fig. shows another possible structural schematic diagram of the task processing device involved in the above embodiment. As Figure 14 shown, the task processing device provided in the embodiments of the present application may include: a processing module 1401 and a communication module 1402. The processing module 1401 can be used to control and manage the actions of the task processing device. For example, the processing module 1401 can be used to support the Figure 14 configuration module 1301, creation module 1302, acquisition module 1303, scheduling module 1304, update module 1305, and splitting module 1306 in the above to execute corresponding steps, and / or for other processes of the technology described herein. The communication module 1402 can be used to support the communication of the task processing device with other network entities. As Figure 14 shown, the task processing device may further include a storage module 1403 for storing computer instructions and data.
[0220] Among them, the processing module 1401 can be a processor or a controller (for example, the processing module 1401 can be the Figure 4 processor 401 in), and the above processing module 1401 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of DSP and a microprocessor, and so on. The communication module 1402 can be a communication interface (for example, the communication module 1402 can be the Figure 4 communication interface 403 in). The storage module 1403 can be a memory (for example, the storage module 1403 can be the Figure 4 memory 402 in). When the processing module 1401 is a processor, the communication module 1402 is a communication interface, and the storage module 1403 is a memory, the processor, transceiver, and memory can be connected through a bus.
[0221] For more details on how the modules included in the above task processing device implement the above functions, please refer to the descriptions in the foregoing method embodiments, and will not be repeated here. Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
[0222] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state drive (SSD)), etc.
[0223] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. For the specific working processes of the systems, devices, and units described above, reference can be made to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0224] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0225] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0226] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0227] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disc and other various media that can store program codes.
[0228] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A task scheduling method, characterized in that, Applied to a management node in a computing cluster, the method includes: Setting first computing resources for a first tenant; When it is necessary to execute the job tasks of the first tenant, starting a first scheduling task for the first tenant, where the first scheduling task is used to schedule computing resources for the job tasks of the first tenant; Obtaining the job tasks of the first tenant and sending the job tasks to the task queue of the first tenant; Executing the first scheduling task to schedule the first computing resources for the job tasks in the task queue to execute the job tasks.
2. The method according to claim 1, characterized in that, The method further includes: Setting a scheduling policy for the first tenant; The scheduling of the first computing resources for the job tasks in the task queue includes: Scheduling the first computing resources for the job tasks in the task queue according to the scheduling policy.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Setting a tenant priority for the first tenant; the tenant priority is used to determine the scheduling priority of the first scheduling task; The executing of the first scheduling task includes: Executing the first scheduling task according to the scheduling priority.
4. The method according to claim 3, wherein The executing of the first scheduling task according to the scheduling priority includes: Determining whether the first scheduling task in the priority queue is the highest scheduling priority, and the priority queue further includes a second scheduling task of a second tenant; If the first scheduling task is the highest scheduling priority, then execute the first scheduling task; if the first scheduling task is not the highest scheduling priority, then execute the second scheduling task.
5. The method according to claim 3, characterized in that, The method further includes: Dynamically updating the scheduling priority of the first scheduling task according to the executed duration of the first scheduling task.
6. The method according to claim 3, characterized in that, The computing cluster further includes a second tenant, and the second computing resources of the second tenant and the first computing resources of the first tenant share the physical computing resources of the computing cluster. The method further includes: Dividing the first scheduling task and the second scheduling task of the second tenant into multiple subtasks respectively; The executing of the first scheduling task according to the scheduling priority includes: If the first scheduling task is the highest scheduling priority, then execute the first subtask of the first scheduling task, schedule the first computing resources for the job tasks in the task queue of the first tenant to execute the job tasks of the first tenant; Updating the scheduling priority of the first scheduling task; If the first scheduling task after the scheduling priority is updated is not the highest scheduling priority, then execute the second subtask of the second scheduling task, schedule the second computing resources for the job tasks in the task queue of the second tenant to execute the job tasks of the second tenant.
7. A task scheduling device, characterized in that, Applied to a management node in a computing cluster, the task scheduling device includes: a configuration module, a creation module, an acquisition module, and a scheduling module; The configuration module is used to set first computing resources for a first tenant; The creation module is used to start a first scheduling task for the first tenant when it is necessary to execute the job tasks of the first tenant, where the first scheduling task is used to schedule computing resources for the job tasks of the first tenant; The obtaining module is configured to obtain the job tasks of the first tenant and send the job tasks to the task queue of the first tenant; The scheduling module is configured to execute the first scheduling task, schedule the first computing resource for the job tasks in the task queue, and execute the job tasks.
8. The task scheduling device according to claim 7, wherein: The configuration module is further configured to set a scheduling policy for the first tenant; The scheduling module is specifically configured to schedule the first computing resource for the job tasks in the task queue according to the scheduling policy.
9. The task scheduling device according to claim 7 or 8, wherein: The configuration module is further configured to set a tenant priority for the first tenant; the tenant priority is used to determine the scheduling priority of the first scheduling task; The scheduling module is specifically configured to execute the first scheduling task according to the scheduling priority.
10. The task scheduling device according to claim 9, wherein: The scheduling module is specifically configured to determine whether the first scheduling task in the priority queue is the highest scheduling priority, and the priority queue further includes a second scheduling task of a second tenant; if the first scheduling task is the highest scheduling priority, execute the first scheduling task; if the first scheduling task is not the highest scheduling priority, execute the second scheduling task.
11. The task scheduling device according to claim 9, wherein It further includes an updating module; The updating module is configured to dynamically update the scheduling priority of the first scheduling task according to the executed duration of the first scheduling task.
12. The task scheduling device according to claim 9, characterized in that, The computing cluster further includes a second tenant, and the second computing resource of the second tenant and the first computing resource of the first tenant share the physical computing resources of the computing cluster; the task scheduling device further includes a splitting module and an updating module; The splitting module is configured to divide the first scheduling task and the second scheduling task of the second tenant into multiple subtasks respectively; The scheduling module is specifically configured to, if the first scheduling task is the highest scheduling priority, execute the first subtask of the first scheduling task, schedule the first computing resource for the job tasks in the task queue of the first tenant, and execute the job tasks of the first tenant; The updating module is configured to update the scheduling priority of the first scheduling task; The scheduling module is further configured to, if the first scheduling task after the scheduling priority is updated is not the highest scheduling priority, execute the second subtask of the second scheduling task, schedule the second computing resource for the job tasks in the task queue of the second tenant, and execute the job tasks of the second tenant.
13. A computer-readable storage medium, characterized in that, Stores computer instructions, which when run on a computer, execute the method according to any one of claims 1 to 6.
14. A computer program product, characterized in that, Contains instructions that, when run on a computing device, cause the computing device to execute the method according to any one of claims 1 to 6.
Citation Information
Cited By
Task scheduling method and apparatus
WO2025152456A1