Multi-tenant-based scheduling platform, resource scheduling method and related equipment
Through a multi-tenant scheduling platform, the business platform and heterogeneous cluster resources of the autonomous driving model are unifiedly managed, and the problems of low resource utilization and high cost are solved, efficient sharing and unified scheduling of resources are realized, and the cost of the production cycle is reduced.
Patent Information
- Application Number
- CN202410091914.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
During the production cycle of the autonomous driving model, each business platform independently manages network resources, resulting in low resource utilization and high cost, and the inability to effectively share and unified scheduling of resources.
It provides a multi-tenant-based scheduling platform, including a task receiving module, a scheduling module and a state synchronization module, and uniformly manages tasks and resource scheduling of multiple business platforms, forms a resource sharing pool, and realizes a unified scheduling strategy across clusters through the k8s system.
It improves the utilization rate of network resources, reduces the cost of production cycles, and reduces the needs of maintenance exclusive task scheduling and resource management systems for various business platforms.
Smart Images

Figure CN120378488A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving, and in particular, to a scheduling platform based on multi-tenants, a resource scheduling method, and related devices. Background Art
[0002] During the entire production cycle of an autonomous driving model, there are multiple links including data processing, model training, and simulation evaluation. Among them, each link is executed by an independent business platform, and each business platform has its own task scheduling and resource management system. Specifically, each business platform independently manages its own network resources and allocates network resources to each business task of the business platform based on its own resource scheduling strategy. In this way, the business platform can complete business tasks based on the allocated network resources.
[0003] It can be seen from the above description that the resources between business platforms are completely isolated, and each platform's business needs to maintain a dedicated task scheduling and resource management system. In this way, resources cannot be shared between business platforms, and it is ineffective to construct a global shared resource pool, resulting in low utilization rate of network resources. At the same time, the system tasks and strategies of each task scheduling and resource management system corresponding to different business platforms are similar, such as resource management, task management, priority scheduling, queuing, and preemption. If each business platform maintains a dedicated task scheduling and resource management system, the cost of the entire production cycle will become very high.
[0004] Therefore, how to flexibly and reasonably manage the network resources of each business platform while reducing the cost of the entire production cycle of the autonomous driving model has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a scheduling platform based on multi-tenants, a resource scheduling method, and related devices, which can reasonably and efficiently manage network resources, improve the utilization rate of network resources, and reduce the cost of the production cycle while alleviating network resource anxiety. The specific technical solutions are as follows.
[0006] In a first aspect, an embodiment of this application provides a scheduling platform for multi-tenants, and the scheduling platform includes:
[0007] A task receiving module, a scheduling module, and a status synchronization module.
[0008] Among them, the first end of the task receiving module is connected to at least one business platform, the second end of the task receiving module is connected to the first end of the scheduling module, and the second end of the scheduling module is connected to at least one heterogeneous cluster.
[0009] The task receiving module is used to receive business tasks submitted by the business platform and deliver the successfully submitted business tasks to the corresponding task queue.
[0010] The scheduling module is used to schedule the service tasks in the task queue and allocate network resources for the service tasks in the task queue.
[0011] The status synchronization module is used to monitor the resource scheduling status of the service tasks and feedback the resource scheduling status to the service platform.
[0012] In an optional implementation manner, the task receiving module is specifically used for:
[0013] Obtain the first task level of the service task submitted by the service platform.
[0014] Determine the first network resources in at least one heterogeneous cluster corresponding to the first task level.
[0015] If all the first network resources are occupied, it is determined that the submission of the service task by the service platform fails.
[0016] If not all the first network resources are occupied, it is determined that the submission of the service task by the service platform is successful.
[0017] In an optional implementation manner, the task receiving module is further used for:
[0018] If all the first network resources are occupied, downgrade the task level of the service task to the second task level according to the downgrading rule. Among them, not all the second network resources in at least one heterogeneous cluster corresponding to the second task level are occupied.
[0019] Determine that the submission of the service task after the task level is downgraded is successful.
[0020] In an optional implementation manner, the task receiving module is further used for:
[0021] Receive the service task parameters submitted by the service platform and perform data format conversion on the service task parameters.
[0022] Render the service task according to the converted service task parameters.
[0023] In an optional implementation manner, the scheduling module is specifically used for:
[0024] Determine the task priority of the service task with a successful submission.
[0025] Deliver the service task with a successful submission to the task queue according to the task priority. The task queue is used to input the service tasks in the task queue to the scheduling module in sequence.
[0026] In an optional implementation manner, the scheduling module is specifically used for:
[0027] Determine the task priority of the successfully submitted service tasks according to the task level of the service tasks, the initial submission time of the service task parameters corresponding to the service tasks, and / or the number of service tasks already submitted by the service platform.
[0028] In an optional implementation manner, the scheduling module is specifically configured to:
[0029] Before the task receiving module conveys the successfully submitted service tasks to the task queue, determine the available network resources corresponding to the task queue.
[0030] If the target network resources required by the successfully submitted service tasks are less than or equal to the available network resources, input the successfully submitted service tasks into the task queue.
[0031] If the target network resources required by the successfully submitted service tasks are greater than the available network resources, instruct the scheduling module to schedule the idle network resources corresponding to other task queues for the task queue.
[0032] In an optional implementation manner, the scheduling module is further configured to:
[0033] If the scheduling module successfully schedules idle network resources for the task queue, input the successfully submitted service tasks into the task queue.
[0034] If the scheduling module fails to schedule idle network resources for the task queue, preempt the positions of other service tasks in the task queue according to the task priority of the successfully submitted service tasks.
[0035] In an optional implementation manner, the status synchronization module is specifically configured to:
[0036] Monitor the resource scheduling status of the service tasks in real time and feedback the resource scheduling status to the service platform. And obtain the task status of all service tasks submitted by the service platform according to a preset period and feedback the task status of all service tasks to the service platform.
[0037] In an optional implementation manner, the status synchronization module is further configured to:
[0038] Update the first network resources according to the resource scheduling status of the service tasks or the task status of all service tasks.
[0039] In a second aspect, an embodiment of the present invention provides a resource scheduling method based on multi-tenancy, which is applied to any one of the scheduling platforms described in the first aspect. The method includes:
[0040] Receive the service task parameters of at least one service task submitted by the service platform.
[0041] Convert the data format of the business task parameters and render the business task according to the converted business task parameters.
[0042] After the business task is successfully submitted, allocate network resources of at least one heterogeneous cluster to the business task in the task queue.
[0043] In an optional implementation, the method further includes:
[0044] Obtain the first task level of the rendered business task.
[0045] Determine the first network resources in at least one heterogeneous cluster corresponding to the first task level.
[0046] If all the first network resources are occupied, determine that the business platform fails to submit the business task.
[0047] If not all the first network resources are occupied, determine that the business platform successfully submits the business task.
[0048] In an optional implementation, the method further includes:
[0049] If all the first network resources are occupied, downgrade the task level of the business task to the second task level according to the downgrading rule. Among them, not all the second network resources in at least one heterogeneous cluster corresponding to the second task level are occupied.
[0050] Determine that the business task after the task level is downgraded is successfully submitted.
[0051] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for resource scheduling based on multi-tenancy described in the second aspect above is implemented.
[0052] In a fourth aspect, an embodiment of the present application provides an electronic device, and the electronic device includes:
[0053] One or more processors;
[0054] The processor is coupled to the storage device, and the storage device is used to store one or more programs,
[0055] When the one or more programs are executed by the one or more processors, the electronic device implements the method for resource scheduling based on multi-tenancy described in the second aspect above.
[0056] In a fifth aspect, the present application provides a computer program product, and the computer program product includes a computer program, and when the computer program is executed by a processor, the method for resource scheduling based on multi-tenancy described in the second aspect is implemented.
[0057] As can be seen from the above, the scheduling platform based on multi-tenancy provided by the embodiments of the present application provides access interfaces for multiple service platforms at one end, allowing multiple service platforms to access the scheduling platform simultaneously. And at the other end, it is connected to multiple heterogeneous clusters to form a resource sharing pool. In this way, the tasks of each service platform are uniformly scheduled by the scheduling platform, and the scheduling platform allocates the network resources of the heterogeneous clusters to the tasks of each service platform based on a unified scheduling policy. Therefore, all the network resources in the heterogeneous clusters can be flexibly allocated to each service platform as shared network resources, rather than being independent network resources exclusive to one service platform. In this way, while improving the throughput rate of each task, the utilization rate of network resources will also be greatly improved. At the same time, each service platform does not need to maintain a dedicated task scheduling and resource management system, but can uniformly utilize the scheduling platform. In this way, the costs of each service platform will also be reduced, thereby reducing the costs of the entire production cycle. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0059] Figure 1 It is a system architecture diagram of a scheduling system provided by an embodiment of the present application;
[0060] Figure 2 It is a schematic structural diagram of a scheduling platform provided by an embodiment of the present application;
[0061] Figure 3 It is a schematic flowchart of a task scheduling method provided by an embodiment of the present application;
[0062] Figure 4 It is a schematic flowchart of another task scheduling method provided by an embodiment of the present application;
[0063] Figure 5 It is a schematic flowchart of another task scheduling method provided by an embodiment of the present application;
[0064] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] In view of this, the present application provides a scheduling platform based on multi-tenancy, a resource scheduling method and related devices, which can manage network resources reasonably and efficiently, improve the utilization rate of network resources, and reduce the costs of the production cycle while alleviating network resource anxiety.
[0066] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0067] It should be noted that the terms "including" and "having" in the embodiments of the present application and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0068] The embodiments of the present application disclose a scheduling platform based on multi-tenancy, a resource scheduling method and related devices. The scheduling platform can provide general resource management and task scheduling capabilities for multiple business platforms. One end of it is connected to each business platform, and the other end is connected to multiple large-scale heterogeneous clusters. It can provide a unified business platform access solution upward and implement a cross-cluster resource scheduling strategy downward. In this way, on the one hand, the cost of the upstream business platform can be reduced, and on the other hand, a downstream resource sharing pool can be established to improve the utilization rate of network resources. The embodiments of the present application will be described in detail below.
[0069] Before introducing the scheduling platform provided by the embodiments of the present application, the entire scheduling system will be introduced first. Figure 1 It is a system architecture diagram of a scheduling system provided for the embodiments of the present application. As Figure 1 shown, the scheduling system includes: a business platform 101, a scheduling platform 102, a heterogeneous cluster 103, and a monitoring and alarm platform 104.
[0070] Among them, the upstream of the scheduling platform 102 is connected to each business platform 101, and the downstream is connected to the heterogeneous cluster 103.
[0071] Specifically, the scheduling platform 102 provides a unified platform access interface, and multiple business platforms 101 can access the scheduling platform 102 simultaneously. It can be understood that the business platform 101 can submit business tasks through the platform access interface. After the business tasks are successfully submitted, the scheduling platform 102 uniformly schedules the business tasks of each business platform 101 and allocates network resources for each business task. Exemplarily, in the production cycle of an autonomous driving model. The business platform 101 can be a model training platform, a model testing platform, an agile orchestration system platform, etc., which is not specifically limited.
[0072] The scheduling platform 102 is used to manage the service tasks submitted by each service platform and allocate network resources of the heterogeneous cluster 103 for each service task. In the embodiment of the present application, the scheduling platform 102 implements the scheduling process based on the k8s system. Among them, k8s is an open-source container orchestration technology used for automatic deployment, scaling, and management of containerized applications. It deploys and manages microservices architecture applications by forming an abstraction layer over the cluster. Its main functions include controlling and managing the use of resources by applications, automatically load-balancing requests between multiple instances of applications, monitoring resource usage and resource limits, automatically preventing applications from consuming excessive resources, and migrating application instances from one host to another when the host resources are exhausted or the host crashes, etc. In the embodiment of the present application, the k8s system can allocate network resources for the service tasks submitted by the service platform based on a unified scheduling policy, that is, deploy the service tasks on the network devices in the heterogeneous cluster 103.
[0073] A cluster refers to a computer cluster, which is a computer system. Specifically, it refers to a collection of computers that connect a group of loosely integrated computer software or hardware and cooperate closely to complete computing work. The heterogeneous cluster 103 refers to multiple large-scale k8s clusters with different physical machine hardware or a single cluster with different internal physical machine hardware. It is used to provide network resources for the service platform to calculate and execute each service task. The differences include, but are not limited to, differences in CPU or GPU models, network cards, or video memory of each computer, etc.
[0074] The monitoring and alarm platform 104 is used to monitor various performance indicators during the service task scheduling process. Specifically, it obtains the service indicators of the service tasks from the scheduling platform 102, then monitors the system indicators provided by the heterogeneous cluster 103, and then makes subsequent adjustment work according to the monitoring results. For example, if a service task does not meet the service indicators of the service task during the execution process, an alarm needs to be issued to adjust the scheduling policy subsequently.
[0075] In an embodiment of the present application, one end of the scheduling platform 102 provides access interfaces for multiple service platforms, allowing multiple service platforms to access simultaneously. The other end is connected to multiple heterogeneous clusters to form a resource sharing pool. In this way, the tasks of each service platform are uniformly scheduled by the scheduling platform, and the scheduling platform allocates the network resources of the heterogeneous clusters to the tasks of each service platform based on a unified scheduling policy. Therefore, all network resources in the heterogeneous clusters can be flexibly allocated to each service platform as shared network resources, rather than being independent network resources exclusive to one service platform. In this way, while improving the throughput rate of each task, the utilization rate of network resources will also be greatly improved. At the same time, each service platform does not need to maintain a dedicated task scheduling and resource management system, but can uniformly utilize the scheduling platform. In this way, the costs of each service platform will also be reduced, thereby reducing the costs of the entire production cycle.
[0076] Based on the above description, the structure of the scheduling platform content will be described in detail below. Figure 2 A schematic structural diagram of a scheduling platform provided by an embodiment of the present application, as Figure 2 shown, the scheduling platform includes:
[0077] A task receiving module 201, a scheduling module 202, and a status synchronization module 203.
[0078] Among them, one end of the task receiving module 201 is connected to the service platform, and the other end is connected to its corresponding task queue. One end of the task queue is connected to one end of the scheduling module 202. The other end of the scheduling module 202 is connected to the heterogeneous cluster.
[0079] The task receiving module is used to receive service tasks submitted by the service platform. Specifically, the service platform inputs service task parameters, and the task receiving module converts the data format of the service task parameters, uniformly renders them into the yaml format of k8s, and then performs the submission work of the service task. After the submission is successful, the service task will be sent to the task queue corresponding to the service platform.
[0080] Next, the task queue will be sequentially input into the scheduling module 202 in order. The scheduling module 202 determines the order (priority) of each service task according to the priority of each service task, and then allocates network resources to the service tasks in the task queue according to the priority of each service task.
[0081] Finally, the status synchronization module 203 is used to monitor the entire scheduling process, determine the resource scheduling status of the service task, that is, whether it is queuing, whether it has entered the task queue, whether resources have been allocated, etc. Then, the resource scheduling status of the service task is fed back to the service platform.
[0082] For the above scheduling platform, the working processes of each module in the scheduling platform are introduced in detail below:
[0083] (1) Task receiving module 201:
[0084] Figure 3 It is a schematic flowchart of a task scheduling method provided by an embodiment of the present application; as Figure 3 shown, the method includes:
[0085] 301. The task receiving module receives the service task parameters submitted by the service platform.
[0086] First, the service platform submits the service task parameters to the scheduling platform. The task receiving module receives the service task parameters of each platform through a unified submission entry and enters the task submission link.
[0087] 302. The task receiving module performs data format conversion on the service task parameters and renders the service task according to the converted service task parameters.
[0088] It can be understood that the scheduling platform is based on the k8s system for task deployment. Therefore, the service task parameters should follow the data input format of k8s to be successfully received. And the corresponding yaml data format of k8s is a very complex data format. If it is required that the service platform strictly follows this format to provide service task parameters, it will bring great difficulties to the task submission process of the service platform. In the embodiment of the present application, the task receiving module of the scheduling platform uniformly performs data format conversion, so that the k8s details are shielded from the service platform. Users of each service platform can input service task parameters according to the data format they are familiar with, greatly reducing the difficulty of service task submission.
[0089] 303. Obtain the first task level of the service task submitted by the service platform.
[0090] It can be understood that the user of the service platform can specify the priority of the service task, that is, the task level of the service task. Generally, the higher the task level of a certain service task, the more important the service task is, and the service task should be scheduled first and network resources should be allocated to it. And the task level of the service task also specifies the allocation level of network resources. Exemplarily, if the user specifies a certain service task as a high-level task, then the scheduling module needs to allocate the network resources corresponding to this high level to it, such as higher-performance GPU resources.
[0091] 304. Determine whether the first network resources in the heterogeneous cluster corresponding to the first task level are completely occupied. If not, execute step 305; if so, execute step 306.
[0092] After obtaining the task level of the service task, it is necessary to check whether there is any remaining network resource in the heterogeneous cluster corresponding to this task level in the database, that is, whether it is fully occupied. If it is not occupied, it means that there are still remaining network resources in the heterogeneous cluster that can be allocated to this service task. If it is fully occupied, it means that the heterogeneous cluster temporarily does not have the ability to execute this service task.
[0093] Exemplarily, the network resource can be GPU time. Among them, high-level service tasks require high-level GPUs to execute. If the task level of a service task submitted by a certain service platform is high level, then it is necessary to obtain the GPU time corresponding to the high-level GPUs in the heterogeneous cluster. Suppose there are 10 high-level GPUs in the cluster, then the total GPU time is 240 hours. At this time, it is necessary to judge the remaining amount of GPU time in the heterogeneous cluster. If there is a remaining amount, it means that the resources of the high-level GPUs are not fully occupied. If there is no remaining amount, it means that the resources of the high-level GPUs are fully occupied.
[0094] 305. Determine that the service task is successfully submitted to the business platform.
[0095] When the first network resource is not fully occupied, it means that the heterogeneous cluster still has the ability to execute this service task, so it is determined that the service task is successfully submitted.
[0096] 306. Degrade the task level of the service task according to the downgrading rule.
[0097] If the first network resource is fully occupied, it means that the network resource corresponding to this level cannot be allocated to this service task. At this time, it is necessary to automatically downgrade the service task until the network resource corresponding to the downgraded level is not fully occupied. Specifically, it is necessary to downgrade the task level of the service task from the first task level to the second task level. If there are remaining network resources corresponding to the second task level in the heterogeneous cluster, then allocate the network resources corresponding to the second task level to this service task.
[0098] 307. Determine that the service task is stored in the database and send a storage notice to the business platform.
[0099] After the service task is successfully submitted, it is determined that the service task is stored in the database. Wait for the subsequent scheduling module to perform the task scheduling and resource allocation for the next link. At the same time, it is necessary to feedback to the business platform that the storage is successful to inform the business platform that this service task has been successfully received and avoid repeated submission.
[0100] In the above embodiment, the task receiving module uniformly receives the service tasks and adjusts the task levels of the service tasks according to the resource inspection, laying a foundation for the subsequent task scheduling and resource allocation.
[0101] (2) Scheduling Module 202:
[0102] After the task receiving module successfully receives a service task and the task is successfully stored in the database, the scheduling module needs to perform the processes of task scheduling and resource allocation. Figure 4 It is a schematic flowchart of a task scheduling method provided by an embodiment of the present application. As Figure 4 shown, the entire process includes the following steps:
[0103] 401. Determine the task priority of the successfully submitted service task.
[0104] It can be understood that users of the service platform can specify the priority for service tasks, that is, the task level. And the task priority is the actual priority calculated by the scheduling system according to the scheduling policy. It can be understood that the task level specified by the user is the most important factor affecting the task priority. Also, the initial submission time (queuing time) of the service task parameters corresponding to the service task and the number of service tasks already submitted by the service platform (the water level already used by the service platform) will also affect the task priority.
[0105] Among them, the higher the task level specified by the user, the higher the task priority of the service task. When the queuing time of the service task is too long, the task priority of the service task can be appropriately increased to improve the speed of allocating network resources for the service task. Also, when the water level already used by a service platform is higher, it means that the number of service tasks already submitted by the service platform is larger. To ensure the normal operation of other service platforms, the priority of the service tasks submitted by the service platform can be appropriately reduced. It can be understood that the priority policy can be specified according to the actual situation and is not specifically limited.
[0106] 402. Obtain the distributed lock of the task queue.
[0107] The distributed lock is the lock of the distributed system. Using the distributed lock can solve the problem of controlling access to shared resources. In the embodiment of the present application, the distributed lock is used to ensure that only one service task enters the task queue each time, avoiding conflicts between multiple service tasks.
[0108] Among them, service tasks enter the task queue in sequence according to the priority level, and then the task queue outputs service tasks in the order of priority.
[0109] 403. Determine the available network resources corresponding to the task queue.
[0110] After determining the task priority of a business task, it is necessary to first determine whether it is the turn of this business task to enter the task queue. If there is a business task with a higher priority than this business task, then it is necessary to wait. And if it is the turn of this business task to enter the task queue, it is also necessary to determine whether there are still available allocation resources remaining in the task queue. That is, whether there is still space in the task queue to accommodate this business task.
[0111] 404. Determine whether the available network resources are greater than the target network resources required by the business task. If so, execute step 405. If not, execute step 406.
[0112] If the available network resources in the task queue are greater than the target network resources required by the business task, it means that the task queue can accommodate this business task. And if the available network resources in the task queue are less than the target network resources required by the business task, it means that the task queue cannot accommodate this business task, and this business task still needs to queue up and wait.
[0113] 405. Input the successfully submitted business task into the task queue.
[0114] If the task queue can accommodate this business task, then input this business task into the task queue, and then the task queue will be input into the k8s system in sequence, and the k8s system will allocate network resources for it.
[0115] 406. Determine whether it is possible to schedule the idle network resources corresponding to other task queues for the task queue. If so, execute step 407. If not, execute step 408.
[0116] It can be understood that this step reflects the elastic scheduling strategy supported by the scheduling platform provided in the embodiments of the present application. That is, the queue resources corresponding to multiple business platforms can be shared. If the network allocation resources corresponding to a task queue are insufficient, the remaining resources of other queues can be borrowed.
[0117] 407. Schedule the idle network resources corresponding to other task queues for the task queue, and input the successfully submitted business task into the task queue.
[0118] If the elastic scheduling is successful, then input this business task into the task queue.
[0119] 408. According to the task priority of the successfully submitted business task, preempt the positions of other business tasks in the task queue.
[0120] If the elastic scheduling fails, then the preemption process needs to be started. That is, preempt the positions of other business tasks that already exist in the preemption task queue. Specific preemption strategies may include: high-priority business tasks can preempt the positions of low-priority business tasks in the queue, while medium-priority tasks do not support preemption. Or for two business tasks with the same priority, the business task with a longer submission time can preempt the position of the business task with a shorter submission time, etc., which are not specifically limited.
[0121] It can be understood that if the preemption is successful, then the business task successfully enters the task queue, and the preempted business task will be eliminated. And if the preemption is not successful, then it is necessary to recalculate the task priority of this business task and enter the queuing and waiting link again. It can be understood that the business task that fails in preemption can appropriately increase its task priority to reduce the task queuing and waiting time.
[0122] 409. Release the distributed lock.
[0123] It can be understood that when the task scheduling process is ended and the business task is selected, it is necessary to input the business task into the k8s system, and the k8s system creates a k8s task corresponding to the business task and allocates network resources in the heterogeneous cluster for this business task.
[0124] The above embodiments have introduced the task scheduling process in detail. The scheduling module schedules each business task based on a unified scheduling strategy. Unified regulations are made for task priority scheduling, elastic scheduling, preemption queuing, etc. In this way, the business tasks submitted by each business platform can be scheduled based on a unified scheduling rule, and resources can be flexibly allocated to each business task.
[0125] (Three) Status synchronization module 203:
[0126] The status synchronization module provides two mechanisms. One is the list mechanism to ensure the consistency and accuracy of the final task status. Specifically, the status synchronization module needs to periodically and regularly obtain the resource scheduling status of all business tasks submitted by the business platform to update the task status of all business tasks.
[0127] The watch mechanism is to monitor the resource scheduling status of business tasks in real time to ensure the real-time update of task status. When the task status changes, the status of each database is updated in real time. Exemplarily, the remaining amount of each network resource in the heterogeneous cluster can be updated according to the resource scheduling status of each business task.
[0128] The following gives an overall description of the entire scheduling process. Figure 5 It is a flowchart of a task scheduling method provided by an embodiment of the present application. As Figure 5As shown, the business platform submits the business task parameters through a unified task submission entrance. Then the scheduling platform first checks the parameters and converts the data, renders the business task into the data format corresponding to k8s and stores the business task. Then the scheduling platform determines the priority of the business task and delivers the business task to the task queue according to the priority of the business task. Next, the scheduling module determines whether the business task is successfully enqueued. If the enqueuing is passed, it is input into the k8s system, and after constructing the k8s task for it, the k8s task is submitted to the k8s system. After the k8s system allocates network resources for it, the downstream cluster executes the business task. If the enqueuing fails, the preemption process is started. If the preemption is successful, the process after the enqueuing continues. If the preemption fails, the task priority of the business task needs to be recalculated and the queue continues.
[0129] Among them, the status synchronization module monitors the task status of each business task in real time and feeds back the task status to the upstream business platform.
[0130] In an embodiment of the present application, one end of the scheduling platform provides access interfaces for multiple business platforms, allowing multiple business platforms to access at the same time. The other end is connected to multiple heterogeneous clusters to form a resource sharing pool. In this way, the tasks of each business platform are uniformly scheduled by the scheduling platform, and the scheduling platform allocates network resources of the heterogeneous clusters to the tasks of each business platform based on a unified scheduling strategy. Therefore, all network resources in the heterogeneous cluster can be flexibly allocated to each business platform as shared network resources, and no longer as independent network resources exclusively belonging to a business platform. In this way, while improving the throughput of each task, the utilization rate of network resources will also be greatly improved. At the same time, each business platform does not need to maintain a dedicated task scheduling and resource management system, but can uniformly use the scheduling platform. In this way, the cost of each business platform will also be reduced, thereby reducing the cost of the entire production cycle.
[0131] Based on the above description, an electronic device provided in an embodiment of the present application is shown in FIG. Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1000 can be specifically a virtual reality VR device, a mobile phone, a tablet, a laptop computer, a smart wearable device, a monitoring data processing device, or a radar data processing device, etc., which is not limited here. Among them, the electronic device 1000 can be deployed with Figures 1 to 2 The scheduling platform described in the corresponding embodiment is used to implement Figures 3 to 5 Specifically, the electronic device 1000 includes: a receiver 1001, a transmitter 1002, a processor 1003 and a memory 1004 (wherein the number of the processor 1003 in the execution device 1000 may be one or more, Figure 6Take a processor as an example. Among them, the processor 1003 may include an application processor 10031 and a communication processor 10032. In some embodiments of the present application, the receiver 1001, the transmitter 1002, the processor 1003, and the memory 1004 may be connected through a bus or other means.
[0132] The memory 1004 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1003. A part of the memory 1004 may also include a non-volatile random access memory (NVRAM). The memory 1004 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions may include various operation instructions for implementing various operations.
[0133] The processor 1003 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system. Among them, the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clarity, all kinds of buses are referred to as the bus system in the figure.
[0134] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1003. The processor 1003 can be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 1003 or instructions in software form. The above-mentioned processor 1003 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1003 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1004, and the processor 1003 reads the information in the memory 1004 and combines its hardware to complete the steps of the above method.
[0135] The receiver 1001 can be used to receive input digital or character information and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1002 can be used to output digital or character information through the first interface; the transmitter 1002 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1002 can also include a display device such as a display screen.
[0136] In the embodiments of the present application, the application processor 10031 in the processor 1003 is used to execute Figures 3 to 5 the task scheduling method in the corresponding embodiment. It should be noted that the specific manner in which the application processor 10031 executes each step is based on the same concept as the corresponding method embodiments in the present application, and the technical effects brought by it are the same as those of the corresponding method embodiments in the present application. For the specific content, reference can be made to the description in the method embodiments shown above in the present application, and details will not be elaborated here. Figures 3 to 5 corresponding, and the technical effects brought by it are the same as those of the corresponding method embodiments in the present application. For the specific content, reference can be made to the description in the method embodiments shown above in the present application, and details will not be elaborated here. Figures 3 to 5 corresponding, and the technical effects brought by it are the same as those of the corresponding method embodiments in the present application. For the specific content, reference can be made to the description in the method embodiments shown above in the present application, and details will not be elaborated here.
[0137] An embodiment of the present application provides a computer-readable storage medium, which includes computer instructions that are used to implement the technical solution of any one of the task scheduling methods in the embodiments of the present application when executed by a processor.
[0138] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0139] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0140] Among them, the computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media, such as modulated data signals and carrier waves.
[0141] The present application provides a computer program product, which includes a computer program that implements the technical solution of any one of the task scheduling methods in the embodiments of the present application when executed by a processor.
[0142] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present application.
[0143] Those of ordinary skill in the art can understand that the modules in the device in the embodiment can be distributed in the device in the embodiment according to the description of the embodiment, or can be correspondingly changed and located in one or more devices different from this embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A scheduling platform for multi-tenants, characterized in that The scheduling platform includes: a task receiving module, a scheduling module, and a status synchronization module; wherein, the first end of the task receiving module is connected to at least one service platform, the second end of the task receiving module is connected to the first end of the scheduling module, and the second end of the scheduling module is connected to at least one heterogeneous cluster; The task receiving module is configured to receive service tasks submitted by the service platform and deliver the successfully submitted service tasks to the task queue corresponding to the service platform; The scheduling module is configured to schedule the service tasks in the task queue and allocate network resources to the service tasks in the task queue; The status synchronization module is configured to monitor the resource scheduling status of the service tasks and feedback the resource scheduling status to the service platform.
2. The scheduling platform according to claim 1, wherein The task receiving module, specifically, is configured to: obtain a first task level of the service task submitted by the service platform; determine first network resources in the at least one heterogeneous cluster corresponding to the first task level; if the first network resources are all occupied, determine that the service platform fails to submit the service task; if the first network resources are not all occupied, determine that the service platform successfully submits the service task.
3. The scheduling platform according to claim 2, wherein The task receiving module is further configured to: if the first network resources are all occupied, downgrade the task level of the service task to a second task level according to a downgrading rule; wherein, second network resources in the at least one heterogeneous cluster corresponding to the second task level are not all occupied; determine that the service task with the downgraded task level is successfully submitted.
4. The scheduling platform according to any one of claims 2 to 3, characterized in that, The task receiving module is further configured to: receive service task parameters submitted by the service platform and perform data format conversion on the service task parameters; render the service task according to the converted service task parameters.
5. The scheduling platform according to claim 4, characterized in that The scheduling module, specifically, is configured to: determine the task priority of the successfully submitted service task; deliver the successfully submitted service task to the task queue according to the task priority; The task queue is configured to sequentially input the service tasks in the task queue to the scheduling module.
6. The scheduling platform according to claim 5, characterized in that The scheduling module, specifically, is configured to: determine the task priority of the successfully submitted service task according to the task level of the service task, the initial submission time of the service task parameters corresponding to the service task, and / or the number of service tasks already submitted by the service platform.
7. The scheduling platform according to claim 6, wherein The scheduling module, specifically, is configured to: determine the available network resources corresponding to the task queue before the task receiving module delivers the successfully submitted service task to the task queue; if the target network resources required by the successfully submitted service task are less than or equal to the available network resources, input the successfully submitted service task into the task queue; if the target network resources required by the successfully submitted service task are greater than the available network resources, instruct the scheduling module to schedule idle network resources corresponding to other task queues for the task queue.
8. The scheduling platform according to claim 7, characterized in that, The scheduling module is further configured to: If the scheduling module successfully schedules the idle network resources for the task queue, the submitted business task that is successfully submitted is input into the task queue; If the scheduling module fails to schedule the idle network resources for the task queue, it preempts the positions of other business tasks in the task queue according to the task priority of the submitted business task that is successfully submitted.
9. The scheduling platform according to claim 8, characterized in that, The status synchronization module is specifically configured to: Monitor the resource scheduling status of the business task in real time and feedback the resource scheduling status to the business platform; and Obtain the task status of all business tasks submitted by the business platform according to a preset period and feedback the task status of all business tasks to the business platform.
10. The scheduling platform according to claim 9, characterized in that, The status synchronization module is further configured to: Update the first network resources according to the resource scheduling status of the business task or the task status of all business tasks.
11. A resource scheduling method based on multi-tenancy, applied to the scheduling platform according to any one of claims 1 to 10, characterized in that, The method includes: Receiving business task parameters of at least one business task submitted by a business platform; Converting the data format of the business task parameters and rendering the business task according to the converted business task parameters; When the business task is successfully submitted, allocate network resources of at least one heterogeneous cluster to the business tasks in the task queue.
12. The scheduling method according to claim 11, wherein The method further includes: Obtaining a first task level of the rendered business task; Determining first network resources in the at least one heterogeneous cluster corresponding to the first task level; If the first network resources are all occupied, it is determined that the business platform fails to submit the business task; If the first network resources are not all occupied, it is determined that the business platform successfully submits the business task.
13. The scheduling method according to claim 12, wherein The method further includes: If the first network resources are all occupied, downgrade the task level of the business task to a second task level according to a downgrading rule; wherein, second network resources in the at least one heterogeneous cluster corresponding to the second task level are not all occupied; Determine that the business task after the task level is downgraded is successfully submitted.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the multi-tenant based resource scheduling method according to any one of claims 11 to 13.
15. An electronic device, characterized in that, The electronic device includes: One or more processors; The processor is coupled to a storage device, and the storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the electronic device implements the multi-tenant based resource scheduling method according to any one of claims 11 to 13.